LLM — Large Language Models
Status: Basic Tags: LLM AI MachineLearning NLP Created: 2026-08-09 Related: Deep Learning, Transformer Architecture, Fine-Tuning
Overview
A Large Language Model (LLM) is an AI model trained on massive text corpora to understand and generate human language. LLMs use neural network architectures, primarily Transformers, to process sequences of text and predict subsequent tokens.
Key Characteristics
- Scale: Billions to trillions of parameters
- Training Data: Trillions of tokens from internet text, books, code
- Capabilities: Text generation, translation, summarization, coding, reasoning
- Architecture: Transformer-based with self-attention mechanisms
Architecture
Transformer Model
The Transformer architecture (Vaswani et al., 2017) is the foundation of modern LLMs.
Core Components
- Self-Attention Mechanism — Models relationships between all tokens in a sequence
- Multi-Head Attention — Parallel attention layers capturing different aspects
- Position-Wise Feed-Forward — Neural network applied to each position
- Layer Normalization & Residual Connections — Stabilize training
- Position Encoding — Injects sequence order information
Attention Formula
Attention(Q, K, V) = softmax(QK^T / √d_k)V
Where:
- Q = Query matrix
- K = Key matrix
- V = Value matrix
- d_k = dimension of key vectors
Variants
- Decoder-only: GPT series, Llama series — autoregressive generation
- Encoder-decoder: T5, BERT — understanding and generation
- Encoder-only: BERT, RoBERTa — text classification, NLU tasks
Parameter Efficiency
| Model | Parameters | Architecture | Notes |
|---|---|---|---|
| GPT-2 | 1.5B | Decoder-only | First GPT released |
| GPT-3 | 175B | Decoder-only | Breakthrough scale |
| LLaMA | 7-70B | Decoder-only | Open weight |
| Mistral | 7-8x7B | Decoder-only | MoE architecture |
| Mixtral | 46.7B | MoE | Sparse Mixture of Experts |
| Claude | 1T+ | Decoder-only | Anthropic’s proprietary |
| Gemini | 1.8T | Decoder-only | Google’s largest |
| Qwen | 110B | Decoder-only | Alibaba’s open model |
Training Process
Phases of LLM Training
1. Pre-training (Foundation Model)
- Goal: Learn general language understanding
- Data: Internet text, books, Wikipedia, code repositories
- Objective: Next token prediction (causal language modeling)
- Compute: Thousands of GPU/TPU hours
- Duration: Weeks to months
2. Supervised Fine-Tuning (SFT)
- Goal: Align model with human instructions
- Data: Human-written prompt-response pairs
- Method: Continue pre-training on instruction data
- Result: Model follows instructions reliably
3. Human Feedback
- RLHF (Reinforcement Learning from Human Feedback):
- Train reward model on human preferences
- Optimize LLM with PPO (Proximal Policy Optimization)
- DPO (Direct Preference Optimization):
- Direct optimization on preference pairs
- More stable than RLHF
- ORPO (Odds Ratio Preference Optimization):
- Combines SFT and preference optimization
Training Loss
L = -Σ log P(w_t | w_<t; θ)
Key Training Techniques
- Learning Rate Scheduling: Cosine decay with warmup
- Batch Size: Global batch size of 2-8M tokens
- Optimizer: AdamW with β₁=0.9, β₂=0.95
- Normalization: RMSNorm (more efficient than LayerNorm)
- Activation: SwiGLU (better performance)
Inference
Inference Methods
Decoding Strategies
| Strategy | Description | Quality | Speed |
|---|---|---|---|
| Greedy | Always pick highest probability token | Low | Fastest |
| Beam Search | Keep top-K sequences | Medium | Medium |
| Sampling | Pick from probability distribution | High | Fast |
| Top-k | Sample from top-k most likely tokens | High | Fast |
| Top-p (nucleus) | Sample from cumulative probability p | High | Fast |
| Temperature | Control randomness (higher = more random) | Variable | Fast |
Optimization Techniques
- KV Cache: Cache key-value pairs for efficient generation
- Quantization: FP16, INT8, INT4, GPTQ, AWQ
- Speculative Decoding: Draft model + verify model
- FlashAttention: Optimized attention implementation
- PagedAttention: Efficient memory management (vLLM)
Hardware Requirements
| Model Size | FP16 VRAM | INT4 VRAM | Recommended GPU |
|---|---|---|---|
| 7B | ~14GB | ~5GB | RTX 3090/4090 |
| 13B | ~26GB | ~10GB | RTX 4090 |
| 30B | ~60GB | ~20GB | A100 80GB |
| 70B | ~140GB | ~40GB | A100 80GB × 2 |
| 405B | ~810GB | ~200GB | A100 80GB × 8+ |
Applications
Common Use Cases
- Chatbots: Conversational AI assistants
- Code Generation: GitHub Copilot, Cursor
- Document Analysis: Summarization, extraction
- Translation: Machine translation between languages
- Content Creation: Marketing, articles, emails
- Research: Literature review, hypothesis generation
- Education: Tutoring, personalized learning
Prompt Engineering Techniques
- Zero-shot: Direct prompt without examples
- One-shot: Single example in prompt
- Few-shot: Multiple examples in prompt
- Chain-of-Thought: “Let’s think step by step”
- ReAct: Reasoning + acting pattern
- Tree of Thoughts: Explore multiple reasoning paths
- Self-Consistency: Sample multiple paths, vote on answer
Limitations & Challenges
Known Issues
- Hallucination: Confidently generates incorrect information
- Context Window Limits: Finite token limit for input
- Bias: Reflects biases in training data
- Copyright: Training data raises legal concerns
- Compute Cost: Massive resources required for training
- Safety: Potential for misuse (deepfakes, misinformation)
- Reasoning: Struggles with complex logical deduction
- Factual Accuracy: Outdated knowledge beyond training data
Mitigation Strategies
- Retrieval-Augmented Generation (RAG): Augment with external knowledge
- Constitutional AI: Self-correction via principles
- Red-teaming: Adversarial testing before deployment
- Content Filtering: Safety classifiers
- Fact-checking: Cross-reference with reliable sources
Key Papers & Resources
Foundational Papers
- “Attention Is All You Need” — Vaswani et al. (2017)
- “BERT: Pre-training of Deep Bidirectional Transformers” — Devlin et al. (2019)
- “GPT-2: Language Models are Unsupervised Multitask Learners” — Radford et al. (2019)
- “GPT-3: Language Models are Few-Shot Learners” — Brown et al. (2020)
- “LLaMA: Open and Efficient Foundation Language Models” — Touvron et al. (2023)
- “LoRA: Low-Rank Adaptation of Large Language Models” — Hu et al. (2021)
- “FlashAttention: Fast and Memory-Efficient Exact Attention” — Dao et al. (2022)
Learning Resources
- The Illustrated Transformer — Jay Alammar
- Language Model Zoo — Hugging Face
- Lilian Weng’s LLM blog — Lilian Weng
- Stanford CS324 — Stanford NLP