Bookmarks
2026-08-312
2026-07-162
2026-05-283
2026-05-063
2026-04-29
2026-04-16
2026-03-172
2026-02-06
2026-01-22
2026-01-09
2025-12-132
2025-12-02
2025-10-13
2025-09-01
2025-08-262
2025-08-19
2025-08-11
2025-07-242
2025-07-2223
- Street Fighting Transformers
- How might LLMs store facts | Deep Learning Chapter 7
- Re-thinking Transformers: Searching for Efficient Linear Layers over a Continuous Space of...
- LSTM: The Comeback Story?
- How DeepSeek Rewrote the Transformer [MLA]
- What is the Transformers’ Context Window in Deep Learning? (and how to make it LONG)
- DeepMind’s AlphaEvolve AI: History In The Making!
- What is a Transformer? (Transformer Walkthrough Part 1/2)
- Fireside Chat With Ilya Sutskever and Jensen Huang AI Today and Vision of the Future March 2023
- Matt Squire - Diving into Transformer Model Internals | PyData London 25
- I Visualised Attention in Transformers
- Energy-Based Transformers are Scalable Learners and Thinkers (Paper Review)
- The Attention Mechanism in Large Language Models
- Large Language Models in Five Formulas
- Stanford CS25: V2 I Introduction to Transformers w/ Andrej Karpathy
- LoRA explained (and a bit about precision and quantization)
- How to Build an LLM from Scratch | An Overview
- Transformer Neural Network: Visually Explained
- Sitan Chen - Provably learning a multi-head attention layer - IPAM at UCLA
- Tutorial | LLMs in 5 Formulas (360°)
- Let's build GPT: from scratch, in code, spelled out.
- What's next for AI agentic workflows ft. Andrew Ng of AI Fund
- If we don’t get AGI by GPT-7 (~$1T), will we just never get it? – Sholto Douglas & Trenton Bricken
2025-07-212
2025-07-15
2025-07-09
2025-07-02
2025-06-28
2025-06-27
2025-06-26
2025-05-29
2025-05-163
2025-05-032
2025-04-24
2025-04-224
2025-04-192
2025-04-17
2025-04-15
2025-04-10
2025-04-07
2025-04-05
2025-03-29
2025-03-24
2025-03-23
2025-03-17
2025-03-092
2025-03-03
2025-02-19
2025-02-153
2025-01-27
2025-01-25
2025-01-22
2025-01-18
2025-01-17
2025-01-07
2024-12-26
2024-12-24
2024-12-173
2024-12-16
2024-12-05
2024-11-24
2024-10-31
2024-09-30
2024-07-292
2024-06-27
2024-06-18
2024-06-11
2024-06-06
2024-05-27
2024-05-26
2024-05-254
2024-05-19
2024-05-01
2024-04-25
2024-04-10
2024-03-28
2024-03-06
2024-02-28
2024-02-18
2024-02-08
2024-02-07
2024-01-28
2024-01-20
2024-01-15
2024-01-142
2024-01-12
2024-01-105
- This project is about how to systematically persuade LLMs to jailbreak them.
- mlx-examples/lora at main · ml-explore/mlx-examples · GitHub
- Mixtral of Experts
- Paper page - Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models
- From LLM to Conversational Agent: A Memory Enhanced Architecture with Fine-Tuning of Large Language Models
2024-01-09
2024-01-08
2024-01-073
2024-01-05
2024-01-046
2024-01-034
Subcategories
- applications (9)
- architectures (76)
- efficiency (38)
- evaluation (6)
- inference (4)
- model_reports (5)
- theory (7)
- training (10)
- usage (6)