Language as tokens and probabilities
A large language model divides text into tokens and represents them as numbers. It learns to estimate likely sequences from large collections of language and related data. During generation, it repeatedly predicts a distribution over the next token.
This simple objective can support broad capabilities because language contains patterns of facts, reasoning, style, code, and human interaction. It does not create an internal guarantee that generated claims are true.
Transformers and attention
Most modern LLMs use the transformer architecture. Self-attention lets token representations incorporate information from other positions in the context, while stacked layers build increasingly complex representations.
Scale in parameters, data, and computation contributed to broad capability, but architecture, data quality, training choices, tools, and evaluation remain important.
From model to useful system
Pretraining produces a broad model. Instruction tuning, preference optimization, retrieval, tool use, safety controls, and product design shape how people experience it. The deployed system is therefore more than model weights alone.
An LLM can summarize, translate, draft, answer questions, and write code. Reliability varies across tasks and versions, and fluent language can conceal unsupported statements called hallucinations.
Verification and limits
Important outputs should be checked against controlled sources, deterministic calculations, tests, or qualified human review. A model's own confidence statement is not independent verification.
LLMs are generative systems specialized in language; not every generative model is an LLM. A category describes how a system is built or used; it does not by itself prove accuracy, safety, intelligence, or consciousness. Real performance must be tested on the actual task and conditions of use.
Primary research and institutions
Sources are linked directly so readers can examine the underlying evidence. Numerical statements identify the reporting organization and year. Projections are presented as estimates, not established future facts.