A living field guide to the AI frontier.
YLEM tracks where machine intelligence actually stands: field by field, revised as it moves, with every claim we can test put through the lab.
Where each field stands now
Each guide is a living page: a one-line state of the field, how mature it is, which way it's moving, and what to watch.
Foundation Models
Capability gains now come as much from post-training and inference-time compute as from raw pre-training scale.
Reasoning
Spending more compute at inference time, through long internal reasoning trained with reinforcement learning, is the most important capability lever since scaling pre-training.
Agents
Agents are reliable inside well-instrumented environments such as codebases and browsers with clear goals. Open-ended autonomy over long horizons is still fragile.
AI Coding
Coding is the first domain where agents do substantial delegated work. The bottleneck has moved from writing code to reviewing and verifying it.
Multimodal
Understanding images and documents is routine, and real-time voice works. Video generation is impressive but hard to control precisely.
Memory & Retrieval
Retrieval-augmented generation is the standard enterprise pattern. The hard problem has shifted from finding text to deciding what an agent should remember.
Robotics & Embodied AI
Vision-language-action models show real generalisation in research. Deployment is held back by data, hardware cost and reliability.
Compute & Infrastructure
Inference, not training, increasingly dominates compute demand, driven by reasoning models and agents that use many tokens per task.
Alignment & Safety
Alignment is now part of the product. Techniques like RLHF and constitutional training shape every assistant, while interpretability and evaluation of dangerous capabilities are active frontiers.
Tested, not reported
How much does it take to give an agent your own data over MCP?
We built a 31-line MCP server that searches local Markdown notes and benchmarked it. The protocol overhead is negligible; the real work is tool design.
Latest notes
How we got here
Claude Code preview
An agent that works in the terminal across a whole codebase. Coding assistants move from autocomplete to delegation.
Jan 2025DeepSeek-R1
An open-weight reasoning model trained largely with RL, competitive with closed models at a fraction of the reported cost.
Nov 2024Model Context Protocol
An open protocol for connecting AI apps to tools and data. Integrations become reusable.
Oct 2024Computer use
Claude operates a desktop through screenshots, mouse and keyboard: agents leave the API sandbox.
Sept 2024o1 and test-time compute
A model trained with RL to reason at length before answering. More thinking at inference buys accuracy.
May 2024GPT-4o
One model natively handles text, audio and images, enabling real-time voice conversation.
Mar 2024Claude 3
A three-tier family (Haiku, Sonnet, Opus) with long context and vision.
Feb 2024Sora preview
Minute-long coherent video from text. Video generation joins the frontier.
Dec 2023Mixtral 8x7B
An open sparse mixture-of-experts model. Only a fraction of parameters are active per token.
Dec 2023Gemini 1.0
Google's first model family trained multimodal from the start.
One email when the frontier moves.
A weekly digest of field-guide revisions, new lab verdicts and the log entries that mattered. No hype, no sponsors.