The Evolution of Large Language Models: From ELIZA to AI Agents
AIIn 1966, a computer program called ELIZA could imitate a therapist by recognizing a few keywords and turning a user’s statements into scripted questions. It had no language model in the modern sense, no learned representation of language, and no real understanding of what it was saying.
Six decades later, AI systems can process text, images, audio and video, generate software, operate development environments, retrieve information from external systems, call APIs, and perform multi-step tasks with limited human supervision.
The distance between these two systems is enormous, but the path connecting them is more coherent than it first appears.
Modern large language models still perform the deceptively simple operation at the center of language modeling: given the tokens that came before, estimate what should come next. The difference is that today’s models perform this operation using billions of learned parameters, enormous training corpora, sophisticated post-training procedures, and increasingly large amounts of inference-time computation.
The evolution of LLMs can therefore be understood not simply as a race toward larger models, but as a sequence of solutions to different technical bottlenecks.
Early systems struggled to represent language. Recurrent networks struggled with long-range dependencies and sequential computation. Transformers removed the recurrence bottleneck and made large-scale training practical. Scaling turned general-purpose pre-training into a viable paradigm. Instruction tuning and preference optimization made models useful as assistants rather than merely text generators. Multimodal training expanded what could be represented. Reasoning-focused training made inference-time computation part of the problem-solving process. Finally, tools, memory, retrieval and execution environments transformed language models into components of larger agentic systems.
This article follows that progression from the 1960s to the current generation of AI systems.
What Is an LLM?
A large language model is a neural network trained to model sequences of tokens. In the most common formulation, the model receives a sequence of tokens and predicts a probability distribution for the next token.
A token is not necessarily a complete word. Depending on the tokenizer, a word may be represented as one token or several. Punctuation, spaces, numbers and parts of words can also become individual tokens.
For example, a word such as “unbelievable” might be represented by several subword tokens rather than a single unit.
During training, the model repeatedly compares its prediction with the actual next token in the training data and adjusts its parameters through gradient-based optimization. Over an enormous number of examples, the network learns statistical relationships between tokens and increasingly complex structures in the data.
The interesting part is what emerges from this process.
A sufficiently capable language model does not merely memorize which words tend to follow one another. Its internal representations can encode syntax, semantic relationships, factual associations, programming patterns, mathematical structures and many other regularities present in the training data.
This does not mean that a language model “understands” text in the same way a human does. It means that predicting language at sufficient scale requires learning representations that are useful for many different language tasks.
The term “large” is also relative. Parameter counts have grown from millions to billions and, in some architectures, hundreds of billions or more. But parameter count alone is not a reliable measure of capability. Training data, architecture, optimization, post-training, inference-time computation and the quality of the training process all matter.
This distinction becomes increasingly important later in the history of LLMs.
Era 1: Before the Transformer
From symbolic rules to statistical language models
The history of language-processing machines began long before modern LLMs.
In the early decades of computing, natural language was primarily approached through manually designed rules. Researchers attempted to encode grammar, dictionaries and linguistic structures explicitly.
This approach could work reasonably well in constrained environments, but language is full of exceptions, ambiguity and context-dependent meaning. Writing enough rules to cover real-world language quickly became impractical.
ELIZA and the illusion of understanding
In 1966, Joseph Weizenbaum at MIT introduced ELIZA, one of the most famous early conversational programs.
Its best-known script, DOCTOR, simulated a psychotherapist. It searched a user’s input for recognizable patterns and generated responses according to predefined rules.
A statement such as:
“My mother does not understand me.”
could produce a response along the lines of:
“Tell me more about your family.”
There was no statistical learning and no internal representation comparable to those used by modern neural networks. ELIZA was essentially a sophisticated pattern-matching system.
Yet people sometimes reacted to it as though they were interacting with an intelligent entity.
This phenomenon became known as the ELIZA effect: people tend to attribute more understanding and intentionality to a system when its language behavior appears convincing.
That observation remains relevant in the age of generative AI. Fluent language is not, by itself, evidence that a system has a human-like understanding of the world.
Statistical language modeling
During the 1990s and 2000s, researchers increasingly replaced manually written linguistic rules with statistical models.
N-gram language models estimated the probability of a word based on a limited number of preceding words. A bigram considers one previous word, a trigram two, and so on.
For example, after seeing:
“See you”
a model might assign a high probability to:
“tomorrow”
if that combination occurred frequently in its training corpus.
This is conceptually similar to the autocomplete systems found in modern keyboards, although modern systems are vastly more sophisticated.
N-gram models were useful in speech recognition, spelling correction, machine translation and other applications. However, they had an obvious limitation: their view of context was short.
A sentence can depend on information dozens or hundreds of words earlier. A fixed-size statistical window cannot represent such relationships efficiently.
Word embeddings
A major conceptual change came from distributed representations of words.
In 2013, researchers at Google introduced Word2Vec, demonstrating that words could be represented as dense numerical vectors learned from their surrounding context.
Instead of treating a word as an isolated symbol, the system represented it using a position in a high-dimensional vector space.
Words appearing in similar contexts tended to acquire similar representations.
This was an important transition because neural networks could now manipulate continuous representations of linguistic concepts rather than operating exclusively on discrete word identities.
The famous vector arithmetic examples were particularly illustrative. Certain semantic relationships appeared to correspond approximately to directions in the learned vector space.
Word embeddings did not solve language understanding, but they established an important principle that remains fundamental to modern LLMs: language can be represented numerically in a way that allows a neural network to learn relationships between concepts.
RNNs and LSTMs
The next major step was the use of recurrent neural networks.
An RNN processes a sequence one element at a time while maintaining a hidden state intended to represent what it has seen so far.
Conceptually, the process looks like:
token 1 -> state
token 2 -> state
token 3 -> state
token 4 -> state
...
The hidden state acts as a form of running memory.
This architecture was useful, but it created a serious problem. During training, gradients could become extremely small as they were propagated backward through long sequences. This became known as the vanishing gradient problem.
Long short-term memory networks, or LSTMs, introduced by Sepp Hochreiter and Jürgen Schmidhuber in 1997, addressed this problem using gated memory mechanisms.
Instead of treating every piece of previous information equally, an LSTM could learn which information should be retained, updated or discarded.
LSTMs became extremely important in speech recognition, machine translation and sequence modeling.
Attention mechanisms later improved this architecture further by allowing a model to selectively inspect relevant parts of an input sequence instead of relying entirely on a single recurrent state.
But RNN-based systems had another fundamental limitation.
The problem with sequential computation
An RNN has to process a sequence in order.
It cannot calculate the representation for token 100 before it has processed token 99.
That dependency makes training difficult to parallelize.
Modern GPUs are extremely good at performing large numbers of matrix operations simultaneously. An architecture that forces computation to proceed strictly from one token to the next cannot exploit that hardware as efficiently as a highly parallel architecture.
This limitation became increasingly important as researchers began considering much larger datasets and models.
The industry needed an architecture that could process relationships across a sequence while taking much greater advantage of parallel hardware.
That architecture arrived in 2017.
Era 2: The Transformer Revolution
Attention Is All You Need
In 2017, Ashish Vaswani and colleagues at Google published “Attention Is All You Need.”
The paper introduced the Transformer, an architecture based on attention rather than recurrence.
This was one of the most consequential changes in the history of machine learning.
The Transformer was not simply another improvement to RNNs. It removed recurrence from the central sequence-processing mechanism altogether.
Instead, the model could process many tokens in parallel and use self-attention to determine how strongly different tokens should interact.
What self-attention actually does
Consider the sentence:
“The trophy did not fit in the suitcase because it was too large.”
To interpret the final “it”, a language model needs to establish which earlier noun is relevant.
Self-attention gives the model a mechanism for calculating relationships between tokens directly.
Each token produces representations used to determine:
- what information it should attend to,
- how strongly it should attend to other tokens,
- and what information should be incorporated into its updated representation.
Multiple attention heads can learn different relationships simultaneously.
One head might capture syntactic dependencies. Another might focus on positional relationships. Another may learn patterns that are useful for semantic associations.
The exact behavior of individual attention heads is more complicated than these simplified descriptions suggest, but the important architectural property is straightforward: the model can establish direct interactions between tokens rather than passing information through a long recurrent chain.
Why Transformers scaled
The Transformer solved two major problems at once.
First, much of the computation during training could be parallelized.
Second, the architecture provided more direct paths between distant tokens.
This combination made it practical to train much larger models on much larger datasets.
The original Transformer paper demonstrated strong machine-translation results while also showing substantially improved training efficiency compared with contemporary recurrent architectures.
The implications went far beyond translation.
Once researchers discovered that Transformer architectures could continue benefiting from more data and compute, language modeling began changing from a collection of specialized NLP tasks into a general-purpose pre-training problem.
The central question was no longer:
“How do we build a separate model for this language task?”
It increasingly became:
“How large can we make a general model, and what capabilities emerge when we train it on enough data?”
That question defined the next era.
Era 3: The Scaling Era
GPT-1 and BERT
The Transformer architecture contains two major conceptual components: an encoder and a decoder.
Different research groups explored different ways of using them.
OpenAI’s GPT-1 used a decoder-only Transformer trained autoregressively, meaning it learned to predict the next token from previous tokens.
Google’s BERT took a different approach. It used the encoder architecture and trained the model to reconstruct masked tokens using information from both sides of the missing word.
These designs reflected two different goals.
BERT was extremely effective for understanding-oriented tasks such as classification and information extraction.
GPT’s autoregressive formulation was naturally suited to generating text.
Both approaches were important, but the GPT lineage would eventually become the dominant foundation for general-purpose conversational models.
Pre-training changes the economics of NLP
Before large-scale pre-training, building a language application often meant collecting a task-specific dataset and training a model for that particular task.
A sentiment classifier required one dataset.
A translation system required another.
A question-answering system required another.
Pre-training changed this workflow.
A large model could first learn general language patterns from enormous amounts of unlabeled text. It could then be adapted to many downstream tasks.
The model was no longer trained from scratch for every application.
This was one of the most important economic changes in machine learning.
GPT-2
GPT-2, released in 2019, demonstrated what happened when the same basic approach was scaled significantly.
At 1.5 billion parameters, it was substantially larger than GPT-1 and generated remarkably coherent passages for its time.
OpenAI initially staged the release of the largest version because of concerns about potential misuse.
The debate surrounding the release was important because it anticipated a problem that became much larger later: publishing increasingly capable models creates both research benefits and security risks.
Model release policy subsequently became a major issue involving open weights, licensing, safety testing and responsible deployment.
GPT-3 and few-shot learning
GPT-3, introduced in 2020, pushed scale dramatically further with 175 billion parameters.
More important than the number itself was what researchers observed.
GPT-3 could perform many tasks from instructions and a small number of examples supplied directly in the prompt, without updating its parameters.
This is known as few-shot learning.
For example, a prompt could demonstrate several input-output pairs and then provide a new input. The model could infer the intended task from the examples.
This changed the relationship between developers and machine-learning systems.
Previously, solving a new task often meant building or fine-tuning another model.
With sufficiently capable foundation models, part of that work moved into the prompt.
Prompt engineering emerged from this shift.
The model itself became reusable infrastructure, while the instructions supplied at inference time determined how it was applied.
OpenAI’s GPT-3 research demonstrated that scaling could substantially improve few-shot performance across many tasks without task-specific gradient updates.
Scaling laws
Research into scaling laws helped explain why this strategy worked.
Model performance often improved predictably as researchers increased the amount of training compute, model capacity and training data.
This did not mean that simply adding parameters guaranteed useful intelligence. The proportions between parameters, data and compute mattered.
That became particularly clear with DeepMind’s Chinchilla work in 2022.
Researchers showed that many large language models had been trained with too little data relative to their parameter count. A smaller model trained on substantially more tokens could outperform a larger but undertrained model.
This changed how frontier labs thought about compute allocation.
The goal was no longer simply to build the model with the largest parameter count.
The real optimization problem became:
How should a fixed amount of compute be divided between model size, training data and optimization?
That distinction remains fundamental today.
Era 4: Alignment and Instruction Following
Raw language modeling was not enough.
A model trained to predict internet text does not automatically become a good assistant.
Internet documents contain questions, answers, arguments, jokes, misinformation, instructions, advertisements and countless other forms of text.
If a model simply learns to continue that distribution, it may generate plausible language without understanding what the user actually wants.
From completion to instruction following
The next major step was post-training.
Instead of asking only:
“Can the model predict the next token?”
researchers increasingly asked:
“Can the model produce the kind of response a person actually wants?”
Instruction tuning provided one answer.
Human annotators supplied examples of desirable behavior. The model was then fine-tuned to reproduce that behavior.
Reinforcement learning from human feedback, or RLHF, extended the idea.
A simplified RLHF pipeline looks like this:
- Collect demonstrations of desired responses.
- Fine-tune a language model on those demonstrations.
- Generate multiple candidate answers to the same prompts.
- Have humans rank the candidates.
- Train a reward model to approximate those preferences.
- Optimize the language model against that reward signal.
The details vary between systems, and modern post-training pipelines are considerably more sophisticated than this simplified description.
But the conceptual transition is important.
Pre-training teaches a model about the statistical structure of its data.
Post-training teaches it how to behave in a particular product or interaction setting.
OpenAI’s InstructGPT work demonstrated that a much smaller model trained with human feedback could be preferred by evaluators over a much larger base GPT-3 model.
This showed that capability and usefulness were not the same thing.
ChatGPT
ChatGPT was released publicly in November 2022 as a conversational system based on the GPT-3.5 family and trained using RLHF techniques.
Its importance was not that it suddenly created large language models.
GPT-3 already existed.
The breakthrough was product-level accessibility.
Instead of interacting with a model through carefully designed API prompts, users could simply write natural-language requests and continue a conversation.
This removed much of the technical barrier between the underlying model and the end user.
The result changed the public perception of language models almost overnight.
The industry also learned an important lesson:
A model does not have to be fundamentally new at every layer to create a major product breakthrough. Sometimes the critical innovation is making an existing capability reliable, steerable and accessible.
Era 5: Multimodal Models and the Foundation-Model Economy
The next stage expanded the definition of what a language model could process.
Text was no longer the only important modality.
Models increasingly began accepting images, audio and eventually video alongside text.
GPT-4 and multimodal input
GPT-4 represented an important step in this direction.
OpenAI described GPT-4 as a multimodal model capable of accepting text and image inputs while producing text output. The system demonstrated substantial improvements over GPT-3.5 on a range of professional and academic evaluations.
The architectural details were not fully disclosed, which also marked another important change in the industry.
As models became commercially valuable, leading AI companies became increasingly reluctant to publish complete information about training data, model architecture, compute budgets and parameter counts.
The scientific literature had traditionally rewarded detailed disclosure.
The commercial AI industry increasingly operated under a different set of incentives.
Open weights
At the same time, another movement was developing.
Meta released the original LLaMA family in 2023, followed by Llama 2 and a rapidly expanding ecosystem of open-weight models.
Other organizations released models with permissive licenses, while researchers developed techniques that made fine-tuning large models substantially cheaper.
This created two parallel development models.
The first was the closed frontier model:
Large proprietary model
|
v
API access
|
v
Applications and agents
The second was the open-weight model:
Open model weights
|
+----> Self-hosting
|
+----> Fine-tuning
|
+----> Quantization
|
+----> Local inference
|
+----> Custom deployment
Neither approach eliminated the other.
Closed models offered access to expensive frontier-scale infrastructure without requiring the customer to operate it.
Open models provided greater control over deployment, customization, privacy and hardware choices.
Parameter-efficient fine-tuning
Another important development was the ability to adapt large models without updating every parameter.
LoRA and later QLoRA allowed developers to fine-tune large models using a much smaller number of trainable parameters and significantly less memory.
This helped transform fine-tuning from something reserved for organizations with large clusters into a technique available to much smaller engineering teams.
Quantization pushed this trend further.
A model originally stored using high-precision weights could be represented with lower-bit numerical formats, reducing memory requirements and making local inference practical for increasingly capable models.
Longer context windows
Another major direction was context length.
Early language models could only process relatively short sequences. Modern systems expanded this limit dramatically.
Long context changed the kinds of applications that became possible.
Instead of giving a model a few paragraphs, developers could provide:
- entire technical manuals,
- large source-code repositories,
- legal documents,
- research papers,
- conversation histories,
- databases of internal documentation.
However, a larger context window did not automatically mean perfect use of that information.
Models can still miss relevant information, confuse similar passages, or give disproportionate attention to some parts of a long context.
This led to a new discipline around context engineering, retrieval and information selection.
The model might have a large context window, but the application still needs to decide what information belongs inside it.
Era 6: The Reasoning Revolution
For years, language models primarily spent computation during training.
At inference time, the basic pattern was comparatively simple:
prompt -> model -> next token -> next token -> next token
The reasoning-model era changed this balance.
Inference-time computation
Some problems cannot be solved reliably by producing the first plausible answer.
Mathematics, software engineering, planning and complex analysis often benefit from intermediate computation.
A reasoning-oriented system can spend additional inference-time compute exploring a problem before producing its final response.
This leads to a useful distinction:
Training compute improves the model before deployment.
Inference compute gives the model more resources while solving a particular problem.
The second category became increasingly important in 2024 and 2025.
Chain-of-thought as a stepping stone
Research in 2022 demonstrated that prompting models to work through problems in multiple steps could improve performance on certain reasoning tasks.
This produced the popular “chain-of-thought” prompting technique.
However, it is important not to confuse chain-of-thought prompting with the entire reasoning-model paradigm.
A modern reasoning system may use reinforcement learning, specialized training data, search-like procedures, verification mechanisms, hidden intermediate computation and other techniques that are not equivalent to simply asking a model to “think step by step.”
For production systems, the goal is usually not to expose every internal reasoning trace. What matters is whether additional computation improves the reliability of the resulting answer or action.
OpenAI o1 and the new scaling dimension
OpenAI’s o1, introduced in 2024, became an important example of this approach.
Instead of treating inference as a fixed-cost generation process, the model could spend additional computation on difficult problems.
This created a new dimension of scaling.
Earlier generations were primarily associated with:
more parameters
+
more training data
+
more training compute
Reasoning models added:
more inference-time computation
The practical consequence is significant.
A small, fast model can be appropriate for extracting a date from an invoice.
A reasoning model may be preferable for debugging a difficult algorithm, analyzing a complicated proof or developing a multi-step implementation plan.
This introduced a new engineering trade-off between latency, cost and problem difficulty.
The most capable model is not necessarily the most appropriate model for every request.
Era 7: Open Models Change the Economics
The next major shift was not primarily architectural.
It was economic.
By 2025, the assumption that frontier-level capabilities required access to a small number of proprietary US model providers became increasingly difficult to maintain.
Open-weight models became much more competitive, while inference optimization, mixture-of-experts architectures and model distillation reduced the cost of deploying sophisticated systems.
DeepSeek-R1
DeepSeek-R1, released in January 2025, became one of the most visible examples.
DeepSeek reported that R1 used large-scale reinforcement learning during post-training and achieved performance comparable to leading reasoning systems on several mathematical, coding and reasoning evaluations. The model and related materials were released under the MIT license.
The significance was larger than any individual benchmark.
R1 demonstrated that high-end reasoning behavior could emerge from a development process that was substantially different from the proprietary frontier-lab model pipeline.
It also highlighted the importance of post-training.
The model did not need to derive all of its capabilities exclusively from an enormous supervised dataset. Reinforcement learning could be used to encourage behaviors associated with successful problem solving.
Mixture-of-experts architectures
Another important technique became increasingly prominent: mixture of experts, or MoE.
Instead of activating every parameter for every token, an MoE model contains multiple expert networks and a routing mechanism that selects a subset of experts for each input.
Conceptually:
Input token
|
Router
/ | \
/ | \
Expert A Expert C Expert F
\ | /
\ | /
Output
An MoE model can therefore have a very large total parameter count while activating only a fraction of those parameters for each token.
This can improve the relationship between model capacity and inference cost, although real-world efficiency depends heavily on architecture, routing, hardware and serving implementation.
Distillation
Another important trend was distillation.
A large model can generate training data for a smaller model, allowing some of its behavior to be transferred into a much more compact system.
This creates a cascade:
Large reasoning model
|
v
Synthetic examples
|
v
Smaller distilled model
|
v
Local or cheaper deployment
This matters because not every application requires a frontier-scale model.
For many workloads, a smaller model with the right post-training and domain data can provide better economics.
The new bottleneck
As model access became easier, the engineering problem shifted.
Getting a capable model was no longer necessarily the hardest part.
The difficult questions became:
- Which model is appropriate for the workload?
- How should it be evaluated?
- How much context should it receive?
- Which data should be retrieved?
- How should inference be served?
- How should tool access be controlled?
- How should failures be detected?
- How can the system be monitored in production?
This transition leads directly to the agentic era.
Era 8: From Language Models to AI Agents
The most important conceptual change in the current generation is that the language model is increasingly becoming a component rather than the entire application.
A chatbot generates an answer.
An agent is expected to accomplish a goal.
That difference sounds small, but architecturally it is enormous.
What is an AI agent?
A practical AI agent can be thought of as a system built around a language model that can:
- interpret a goal,
- decide what action to take,
- call external tools,
- inspect the result,
- update its working state,
- continue with another action,
- detect failure,
- and eventually return a result.
The model is the reasoning component, but the surrounding software provides the capabilities required to interact with the outside world.
A simplified architecture looks like this:
User goal
|
v
+---------------+
| Language model|
+---------------+
|
tool decision
|
+-----------+-----------+
| | |
v v v
Search Code API
| | |
+-----------+-----------+
|
Results
|
v
Language model
|
next action
|
...
This is fundamentally different from ordinary text generation.
A model that tells you how to fix a failing test is a chatbot.
A system that opens the repository, runs the tests, examines the failure, modifies the code, runs the tests again and reports the result is an agentic system.
The agent loop
At its simplest, an agent can be reduced to a loop:
def run_agent(goal, llm, tools, max_steps=10):
history = [{"role": "user", "content": goal}]
for _ in range(max_steps):
decision = llm(history, tools=tools)
if decision["type"] == "final_answer":
return decision["content"]
tool = tools[decision["tool_name"]]
result = tool(**decision["arguments"])
history.append({
"role": "assistant",
"content": str(decision)
})
history.append({
"role": "tool",
"content": str(result)
})
return "Stopped: step limit reached"
Real systems are substantially more complicated.
They need authentication, authorization, retries, timeouts, state management, tool schemas, error handling, observability, sandboxing, human approval and protection against malicious inputs.
But the underlying control structure is often surprisingly close to this.
The model proposes an action.
The environment executes it.
The result returns to the model.
The process repeats.
ReAct and the roots of agentic systems
The basic idea predates the current generation.
The ReAct research paradigm demonstrated a way of combining reasoning with actions, allowing a model to alternate between internal reasoning and external tool use.
What changed later was the quality of the models and the surrounding infrastructure.
Better reasoning models became more capable of planning multi-step tasks.
Tool calling became a standard capability.
Longer context windows made it easier to maintain task state.
Retrieval systems gave models access to external knowledge.
Coding environments gave them access to files, terminals and test suites.
The result was a transition from isolated model inference toward interactive computational systems.
Tools change what an LLM can do
A language model has no inherent ability to:
- read your private database,
- browse a website,
- execute Python,
- send an email,
- query an ERP system,
- access a Git repository,
- control a robot,
- or modify a production system.
It can only generate tokens.
Tools provide the missing connection.
For example:
LLM
|
+-- search()
+-- database_query()
+-- execute_python()
+-- read_file()
+-- write_file()
+-- browser()
+-- send_email()
This distinction is important because the capability of an agent is no longer determined exclusively by the model.
The surrounding tool environment can dramatically increase what the system can accomplish.
At the same time, it creates new security problems.
Giving a model access to a database is fundamentally different from giving it access to a read-only documentation search.
Giving it permission to create a file is different from giving it permission to deploy software.
Giving it permission to send an email is different from giving it permission to transfer money.
Agent design therefore becomes partly an authorization problem.
Memory and state
Another important distinction is between model knowledge and application memory.
The model’s parameters contain what it learned during training.
An agent may additionally maintain:
- conversation history,
- task state,
- retrieved documents,
- previous tool results,
- user preferences,
- intermediate files,
- execution logs,
- long-term memory.
These are not the same thing.
A model can be stateless while the application around it is stateful.
This distinction is increasingly important as agents perform tasks lasting minutes, hours or longer.
Context engineering
The growth of agentic systems also produced a new discipline: context engineering.
Prompt engineering asks:
“What instructions should I give the model?”
Context engineering asks a broader question:
“What information should the model receive at this particular step?”
An agent may have access to thousands of documents, hundreds of previous actions and dozens of tools.
Putting all of that into every model call would be inefficient and can actually reduce performance.
The system therefore needs to select relevant context dynamically.
This can involve:
- retrieval,
- summarization,
- state compression,
- memory selection,
- tool-result filtering,
- conversation pruning,
- structured task state.
In long-running agents, context management can become as important as the prompt itself.
Model Context Protocol
The Model Context Protocol, introduced by Anthropic in late 2024, was designed to standardize connections between AI systems and external data sources and tools.
The broader significance of protocols such as MCP is that agent ecosystems need interoperability.
Without a common interface, every AI application has to build custom integrations for every database, development environment, business application and information source.
Standardized tool interfaces can reduce this integration cost.
The analogy is similar to what happened with networking.
The value of the Internet does not come only from individual computers. It comes from standardized protocols that allow many different systems to communicate.
Agentic AI is moving toward a similar model of interoperable tools and services.
The Agentic Architecture Is Bigger Than the Model
It is tempting to describe an AI agent as:
LLM + tools.
In production, that is incomplete.
A serious agent typically requires several additional layers.
Model
The model interprets language, reasons about the current state and proposes actions.
Tools
Tools allow the system to interact with external environments.
Orchestration
The orchestration layer determines how model calls, tools, retries and state transitions are executed.
Memory and state
These mechanisms preserve information across steps and, in some cases, across sessions.
Retrieval
Retrieval systems provide relevant external information that is not necessarily encoded in model parameters.
Guardrails
Guardrails restrict dangerous or unauthorized actions.
Permissions
Authorization determines what the agent is actually allowed to do.
Observability
Logs and traces allow developers to reconstruct what happened during a long agent run.
Evaluation
Evaluation determines whether the agent actually accomplishes its goal rather than merely producing plausible text.
This leads to a more useful mental model:
AI application
|
+--------------+--------------+
| |
Agent UI/API
|
+------+------+------+------+------+
| | | | | |
Model Tools Memory Retrieval State Guardrails
|
v
External systems and data
The language model remains important, but it is only one layer.
Why AI Agents Are Much Harder Than Chatbots
A chatbot can fail with a bad answer.
An agent can fail by taking a bad action.
That distinction changes the engineering requirements.
Suppose an assistant incorrectly summarizes a document.
That is undesirable.
Now suppose an agent incorrectly interprets an instruction and deletes the wrong files, modifies a production database or sends an unauthorized email.
The failure has moved from the informational domain into the operational domain.
This is why agent security increasingly revolves around:
- least-privilege permissions,
- sandboxed execution,
- tool-level authorization,
- approval gates,
- audit logs,
- isolation between trusted and untrusted data,
- prompt-injection defenses,
- rate limits,
- rollback mechanisms,
- human review for high-impact actions.
The model itself cannot be treated as a security boundary.
A malicious document can contain instructions intended to manipulate an agent. A webpage can contain hidden or misleading instructions. Retrieved content can be untrusted.
The system therefore has to distinguish between:
instructions from trusted control channels
and
data supplied by external sources
This is one of the fundamental security problems of agentic AI.
The Evolution of What “Scale” Means
Looking at the entire history reveals that the definition of scaling has changed several times.
First: scale the dataset
Statistical models improved when they were exposed to more examples.
Then: scale the neural network
Transformers made much larger neural models practical.
Then: scale training compute
Large-scale pre-training became an industrial process.
Then: scale context
Models became capable of processing increasingly large amounts of information in a single interaction.
Then: scale inference
Reasoning systems began spending more computation on difficult problems.
Finally: scale the environment
Agents can now increase their effective capability by using more tools, more external data, more execution steps and more sophisticated workflows.
This last transition is particularly important.
A model with a fixed set of parameters can perform very different tasks depending on the environment around it.
A model with access only to text generation is fundamentally different from the same model connected to:
- a browser,
- a compiler,
- a database,
- a code repository,
- a search engine,
- a simulation environment,
- or a robotics controller.
The future of AI may therefore depend as much on scaling the environment as on scaling the model itself.
What Comes After Language Models?
The next stage is already extending beyond purely digital environments.
Vision-language-action models
Robotics requires more than recognizing an image or generating a sentence.
A robot has to connect perception with action.
Vision-language-action systems attempt to map combinations of visual observations and natural-language instructions into physical actions.
The basic loop becomes:
Camera
|
v
Visual representation
|
+---- Language instruction
|
v
Reasoning / policy model
|
v
Action
|
v
Physical environment
|
+---- new observation
|
v
next action
This is structurally similar to an agent using software tools, except that the environment is physical.
World models
Another research direction involves models that learn representations of environments and predict how those environments change.
A sufficiently capable world model could help an agent evaluate possible actions before executing them.
For robotics, autonomous systems and simulations, this could become extremely important.
Instead of:
observe -> act -> hope
the system could increasingly move toward:
observe -> simulate possibilities -> select action -> act -> observe
AI at the edge
Another trend is moving inference away from centralized data centers.
Smaller language and multimodal models can increasingly run on:
- smartphones,
- laptops,
- industrial computers,
- vehicles,
- robots,
- embedded systems.
This creates a different set of engineering constraints.
Cloud systems can use enormous compute resources but require network connectivity and expose data to a remote service.
Edge systems offer lower latency, greater privacy and offline operation, but must work within strict memory, power and compute limits.
This is one reason model compression, quantization, distillation and efficient architectures remain important even as frontier models become larger.
Eight Eras at a Glance
| Era | Approx. period | Representative technologies | Main breakthrough | Main limitation removed |
|---|---|---|---|---|
| 1 | 1960s-2016 | ELIZA, N-grams, Word2Vec, RNNs, LSTMs | Statistical and neural language representations | Hand-written rules and weak sequence memory |
| 2 | 2017 | Transformer | Self-attention and parallel training | Sequential computation |
| 3 | 2018-2020 | GPT-1, BERT, GPT-2, GPT-3 | Large-scale pre-training and few-shot learning | Task-specific models |
| 4 | 2021-2022 | InstructGPT, RLHF, ChatGPT | Instruction following and preference optimization | Gap between language modeling and useful assistance |
| 5 | 2023 | GPT-4, Llama 2, multimodal models, QLoRA | Multimodal input, open weights, long context | Text-only interaction and expensive customization |
| 6 | 2024-2025 | Reasoning models | Inference-time computation and reasoning-oriented training | Fixed-cost generation |
| 7 | 2025 | DeepSeek-R1, MoE models, distillation | Competitive open models and lower deployment costs | Concentration of capability in proprietary systems |
| 8 | 2025-2026 | Tool use, coding agents, MCP, computer-use systems | Multi-step autonomous task execution | Isolation from external tools and environments |
These boundaries are not absolute.
Research from one era often overlaps another, and many technologies continued evolving after the point at which they became historically important.
The purpose of the eight-era model is therefore not to establish eight clean technical generations, but to show the major changes in the dominant engineering problem.
Four Patterns That Explain the Entire History
Looking across six decades, several recurring patterns become visible.
1. Every major architecture removed a bottleneck
RNNs improved sequence modeling but suffered from long-range dependencies and limited parallelism.
Transformers removed much of that sequential constraint.
Mixture-of-experts architectures addressed the relationship between model capacity and active computation.
Efficient attention mechanisms, quantization and smaller architectures continue the same pattern.
The history of AI is therefore not simply a story of bigger models.
It is a story of removing constraints that previously prevented larger or more useful systems.
2. Scale moved through different dimensions
At first, scale meant more training examples.
Then it meant more parameters.
Then more training compute.
Then longer context.
Then more inference-time computation.
Now it increasingly means more interaction with external environments.
This is a major conceptual shift.
An AI system’s capability is no longer determined exclusively by the number of parameters inside the neural network.
3. Post-training became as important as pre-training
Pre-training creates general capabilities.
Post-training shapes behavior.
Instruction tuning, preference optimization, reinforcement learning and reasoning-oriented training determine how the underlying model is used.
The distinction between the pretrained model and the deployed model is therefore increasingly important.
The model users interact with is usually not simply the raw neural network produced at the end of pre-training.
4. The unit of progress became larger than the model
In the early history of NLP, the model itself was the application.
Today, the model is increasingly one component of a system.
A production AI application may include:
Foundation model
+
Post-training
+
Prompt / context management
+
Retrieval
+
Tools
+
Memory
+
Orchestration
+
Security
+
Evaluation
+
Observability
This is perhaps the most important change in the entire timeline.
What the History of LLMs Actually Tells Us
The history of large language models is often presented as a sequence of increasingly large neural networks.
That description is incomplete.
The more useful interpretation is a sequence of engineering problems and solutions.
ELIZA demonstrated that superficial language behavior could create the impression of understanding.
Statistical models showed that language could be modeled probabilistically.
Word embeddings introduced distributed representations.
RNNs and LSTMs demonstrated the power of neural sequence modeling but exposed the limitations of sequential computation.
The Transformer made large-scale parallel training practical.
GPT and BERT demonstrated the power of pre-training.
GPT-3 showed that sufficiently large models could perform many tasks directly from instructions and examples.
RLHF and instruction tuning transformed those models into more useful assistants.
Multimodal systems expanded the input space beyond text.
Open-weight models changed the economics of deployment and experimentation.
Reasoning models introduced inference-time computation as another way to scale performance.
Agents connected models to tools, memory and external environments.
The result is a progression that can be summarized as:
Rules
↓
Statistical models
↓
Neural sequence models
↓
Transformers
↓
Foundation models
↓
Instruction-following models
↓
Multimodal models
↓
Reasoning models
↓
Tool-using agents
↓
Systems that interact with the physical world
The underlying language-model objective did not suddenly disappear during this evolution.
Instead, an increasingly sophisticated stack was built around it.
That is why the original idea behind language modeling remains surprisingly relevant. Predicting the next token sounds trivial. But when that operation is performed by a sufficiently capable neural network, trained on sufficiently broad data, combined with post-training, retrieval, tools, external computation and feedback from the environment, it can become the core component of a system capable of solving tasks that would have seemed impossible when ELIZA was introduced.
The important lesson is therefore not that every new generation of AI is simply “smarter” than the previous one.
Each generation changes the conditions under which the model operates.
The Transformer changed how language models were computed.
Scaling changed how they were trained.
RLHF changed how they behaved.
Multimodality changed what they could perceive.
Reasoning changed how much computation they could spend on a problem.
Open models changed who could deploy and modify them.
Agents changed what they could do outside the model itself.
And that final transition may prove to be the most consequential one.
A language model that only generates text is a powerful information-processing system.
A language model connected to tools, memory, software environments and physical actuators becomes something different: an interface between computation and the world.
That is the direction in which the next chapter of AI is being written.