Bookmarks
2026-08-31
2026-05-06
2026-04-29
2026-04-16
2026-01-22
2026-01-09
2025-12-02
2025-08-262
2025-07-242
2025-07-2222
- Street Fighting Transformers
- How might LLMs store facts | Deep Learning Chapter 7
- Re-thinking Transformers: Searching for Efficient Linear Layers over a Continuous Space of...
- LSTM: The Comeback Story?
- How DeepSeek Rewrote the Transformer [MLA]
- What is the Transformers’ Context Window in Deep Learning? (and how to make it LONG)
- What is a Transformer? (Transformer Walkthrough Part 1/2)
- Fireside Chat With Ilya Sutskever and Jensen Huang AI Today and Vision of the Future March 2023
- Matt Squire - Diving into Transformer Model Internals | PyData London 25
- I Visualised Attention in Transformers
- Energy-Based Transformers are Scalable Learners and Thinkers (Paper Review)
- The Attention Mechanism in Large Language Models
- Large Language Models in Five Formulas
- Stanford CS25: V2 I Introduction to Transformers w/ Andrej Karpathy
- LoRA explained (and a bit about precision and quantization)
- How to Build an LLM from Scratch | An Overview
- Transformer Neural Network: Visually Explained
- Sitan Chen - Provably learning a multi-head attention layer - IPAM at UCLA
- Tutorial | LLMs in 5 Formulas (360°)
- Let's build GPT: from scratch, in code, spelled out.
- What's next for AI agentic workflows ft. Andrew Ng of AI Fund
- If we don’t get AGI by GPT-7 (~$1T), will we just never get it? – Sholto Douglas & Trenton Bricken
2025-07-212
2025-07-15
2025-07-02
2025-06-28
2025-05-29
2025-05-162
2025-04-22
2025-04-192
2025-04-17
2025-04-15
2025-04-05
2025-03-24
2025-03-23
2025-03-092
2025-03-03
2025-01-27
2024-12-17
2024-12-05
2024-10-31
2024-09-30
2024-06-18
2024-06-11
2024-06-06
2024-05-27
2024-04-10
2024-01-142
2024-01-10
2024-01-073
2024-01-05
2024-01-046
2024-01-03
Subcategories
- context (3)
- transformers (73)