Thank you for being part of over 70,000+ ML professionals and enthusiasts who receive weekly articles & tutorials on Machine Learning & MLOps 🤖 You can join the newsletter /newsletter/ ⭐
If you like the content please support the newsletter by sharing with your friends via ✉️ Email, 🐦 Twitter, 💼 Linkedin and 📕 Facebook!
This week in ML Engineering:
- Observability in the Age of AI Agents
- Gergely on the State of Tech
- Hugging Face on Multi-Harness RL
- Pinterest on LLM User Journeys
- State of AI Report 2026
- Open Source ML Frameworks
- Awesome AI Guidelines to check out this week
- + more 🚀
Observability in the Age of AI Agents
Last month I had the pleasure to deliver the opening keynote at Signals Conference 2026 on "The New Failure Modes: Observability in the age of AI Agents" 🚀 This talk goes through how agents break across observability, evals, memory, identity, security and more.
I really enjoyed this one, especially as it was done the week OpenAI Astra came out so the slides are a full on 3D rendering engine instead of slides!
On observability, an agent can be 100% available, 100% within latency and 100% wrong. And none of our current dashboards would notice. This is why we need semantic tracing with a span per reasoning step, and the same evals running offline and online.
On memory, agents can store hallucinations or a prompt injection as facts, which then resurface months later many hops downstream; every memory needs provenance and scoped reads.
On identity, a request that hops from a user to an agent to its sub-agents can lose who it acts for or widen its authority, so we need workload identity, token exchange and a strong fail-closed default.
On security, incidents like the Replit agent deleting a production database or the GitHub MCP exfiltration show none of this is hypothetical anymore.
It is clear that the next decade of production ML will be about understanding what our agents are actually doing; do check it out and let me know your thoughts!
Gergely on the State of Tech
Agents are now opening more PRs on GitHub than humans, and Gergely Orosz brought the numbers in his LDX3 New York keynote on the state of the tech industry:
He lists 16 things that have changed, from most engineers no longer hand-writing code and running several agents in parallel, to IDEs fading, companies building their own agent harnesses and a lot less junior hiring.
The numbers are pretty wild; he reports that agent-authored PRs grew 9x in 8 months, and migrations that used to take years now take weeks, with Bun moving from Zig to Rust in 11 days and Uber converting 600k JUnit tests in 4 months.
The list of what is broken was the most relevant one for me; the volume of code has outgrown our old assumptions, one engineer he quotes describes human code review as "theater", and quality and reliability are going down.
Most of the data comes from vendors like GitHub, Linear and Factory, and his sample leans towards AI-heavy companies, but it matches what we are seeing in practice.
Hugging Face on Multi-Harness RL
Does the harness matter as much as the model? Hugging Face and Liquid AI published a super detailed guide on multi-harness RL, where they train a model inside Claude Code, Codex, OpenCode and Mini-SWE-Agent without modifying any of them, and their conclusion is that training across several harnesses improves the model in all of them.
The starting point shows why this matters: the same base model scores 62% under Mini-SWE-Agent and 33% under Claude Code. Training LFM2.5-2.6B across all four harnesses lifts its average from 42% to 54%, and the gains show up where single-harness training falls short; it reaches 49% in Claude Code and 54% in Codex, against 42% and 43% when trained on OpenCode alone. They also report it beats SFT on a larger teacher's rollouts (47.5%) while making 31% fewer tool calls.
The setup is an OpenEnv capture proxy that sits between the harness and vLLM and records the token ids and logprobs, with Harbor for the tasks and sandboxes and TRL's async GRPO for training. The reward is correctness plus a small bonus for using fewer tool calls.
They are upfront that this is early work; the models are 2B-scale with a single seed, and the overall average gap with OpenCode-only RL (54.2% vs 52.3%) is within the noise, so the benefit is mainly in Claude Code and Codex. The checkpoints, datasets and code are all open.
It is interesting to see this next to the harness results in the State of AI report below, as it seems the harness is becoming as important as the weights.
Pinterest on LLM User Journeys
Pinterest shared how they replaced their multi-stage user journey pipeline with a single LLM generation step:
The old system extracted keywords, clustered them, named and ranked the clusters and then classified each journey stage with heuristics. This split one goal into several journeys, produced generic names like "Art" and only worked in English.
Now a fine-tuned Qwen3 4B reads a 360-day activity log of up to 150 events and returns a ranked JSON list of up to around 12 journeys in the user's language. They distilled it from a larger teacher model, and found that quality plateaued after a few thousand examples; diverse and hard examples mattered more than volume.
For production ML practitioners the serving setup is worth a look: around 100 L40S GPUs with NVIDIA Dynamo in front of vLLM, serving about 775 requests per second at a p95 of about 3 seconds. The online A/B test shows +1.1% email click-through and +1.3% push opens against the clustering system.
The gains are modest, but it is great to see a whole pipeline of heuristics replaced by one small fine-tuned model that also works in every language.
State of AI Report 2026
The State of AI Report 2026 is out! Nathan Benaich and the Air Street team cover research, industry, politics and safety, and quite a lot of interesting trends:
On research, changing only the tools, context and feedback around GLM-5.1 lifted it from 52.5% to 65.5% on a set of SWE-bench Verified tasks, with no change to the weights. On industry, they estimate that the inference cost to hit a fixed benchmark score falls about 13x a year.
The safety section covers the OpenAI and Hugging Face incident in detail, and the nine predictions for 2027 range from card networks assigning liability for agent purchases to "AGI in 2027". Their track record is a 53% hit rate since 2018, so it is definitely worth keeping an eye on next year's scorecard!
Upcoming MLOps Events
The MLOps ecosystem continues to grow at break-neck speeds, making it ever harder for us as practitioners to stay up to date with relevant developments. A fantsatic way to keep on-top of relevant resources is through the great community and events that the MLOps and Production ML ecosystem offers. This is the reason why we have started curating a list of upcoming events in the space, which are outlined below.
You can also find our upcoming events and past talk recordings on the talks and events page.
Events we are speaking at this year:
- Code.Talks 2026 - November @ Hamburg
Other relevant events:
- MLOps World 2026 - Nov @ Austin
- WAICF 2027 - Feb @ Cannes
- KubeCon + CloudNativeCon Europe 2027 - March @ Barcelona
- QCon London 2027 - April @ London
- MLSys 2027 - May @ Santa Clara
- The AI Summit London 2027 - June @ London
- AI Engineer World's Fair 2027 - June @ San Francisco
- RAISE Summit 2027 - July @ Paris
- WeAreDevelopers World Congress 2027 - July @ Berlin
- AGNTCon + MCPCon Europe 2027 - Sept @ London
- HumanX EMEA 2027 - Sept @ Amsterdam
In case you missed our talks, check our recordings below:
- Observability in the Age of AI Agents - Signals Berlin 2026
- The State of AI in 2025 - WeAreDevelopers 2025
- Prod Generative AI in 2024 - KubeCon AI Day 2025
- The State of AI in 2024 - WeAreDevelopers 2024
- Responsible AI Workshop Keynote - NeurIPS 2021
- Practical Guide to ML Explainability - PyCon London
- ML Monitoring: Outliers, Drift, XAI - PyCon Keynote
- Metadata for E2E MLOps - Kubecon NA 2022
- ML Performance Evaluation at Scale - KubeCon Eur 2021
- Industry Strength LLMs - PyData Global 2022
- ML Security Workshop Keynote - NeurIPS 2022
Open Source MLOps Tools
Check out the fast-growing ecosystem of production ML tools & frameworks at the github repository which has reached over 20,000 ⭐ github stars. We are currently looking for more libraries to add - if you know of any that are not listed, please let us know or feel free to add a PR. Here's a few featured open source libraries that we maintain:
- SARC - Provides wrappers for popular agentic frameworks to enable guardrails and constraints that are enforced through the flow.
- KAOS - K8s Agent Orchestration Service for managing the KAOS in large-scale distributed agentic systems.
- Kompute - Blazing fast, lightweight and mobile phone-enabled GPU compute framework optimized for advanced data processing usecases.
- Production ML Tools - A curated list of tools to deploy, monitor and optimize machine learning systems at scale.
- AI Policy List - A mature list that maps the ecosystem of artificial intelligence guidelines, principles, codes of ethics, standards, regulation and beyond.
- Agentic Systems Tools - A new list that aims to map the emerging ecosystem of agentic systems with tools and frameworks for scaling this domain
Please do support some of our open source projects by sharing, contributing or adding a star ⭐
About us
The Institute for Ethical AI & Machine Learning is a European research centre that carries out world-class research into responsible machine learning.