Thank you for being part of over 70,000+ ML professionals and enthusiasts who receive weekly articles & tutorials on Machine Learning & MLOps 🤖 You can join the newsletter /newsletter/ ⭐

If you like the content please support the newsletter by sharing with your friends via ✉️ Email, 🐦 Twitter, 💼 Linkedin and 📕 Facebook!

This week in ML Engineering:

Agentic Security & Identity, Part 2

My 2-part series on agentic identity & security is now COMPLETE! Here's the final blog in the series with a hands-on example of the design decisions in Part 1 using the Kubernetes Agent OS (KAOS) and the new open source Zalando Agentic Identity Broker.

This blog post covers a case-study where two users of the same agent platform are delegating different actions to their respective AI agents. Alice and Bob send requests across their agent graphs and MCPs, and we walk through the requests to see how identity and security can be enforced in agentic systems.

From there I go through the harder cases, like how an agent's own calls to its tools and models get authorized, and how an autonomous agent with no user behind it still only reaches what it was granted allowing users to delegate actions to their agents.

This is one of the most fascinating ereas I've bumped into so far developing the Kubernetes Agent OS (KAOS), as this has required tapping into traditional engineering fundamentals whilst making almost-philosophical decisions.

Agent identity is certainly going to be a topic of exploration for the next years to come.

Uber on Designing the MCP Gateway

Uber shared how they built their MCP Gateway, which now is the platform all their AI agents now go through to reach Uber's backend systems:

Uber's MCP platform is made of two parts, an MCP registry as the control plane that catalogues servers and tools, and a proxy gateway as the data plane that translates MCP calls into HTTP, gRPC or TChannel at runtime.

They have now over 800 MCP servers and 5k tools - as expected most of them are actually not hand-written, as they have workflow called AutoCrawler continuously scans their API definitions and creates MCP servers from them with LLM-enriched tool descriptions.

They also put security in the gateway with tool-level authorization through their internal access control system, built-in redaction of sensitive data, and every tool registered as disabled until someone explicitly enables it.

The scaling problem they hit was context bloat and cost, so they added a single proxy server with incremental discovery, GraphQL-like field selection on responses, and a code mode where coding agents call tools through a CLI and the output goes to files instead of the context.

For production ML practitioners this is a great reference architecture, and it is interesting to see that most of the hard work is in the plumbing around the agents.

Gemini 4 Argon is Out

Google DeepMind released Gemini 4 Argon this week to compete against OpenAI's Astra and Claude's Fable:

THis is a 1M token output limit model (finally), and priced at $2 per million input tokens and $10 per million output tokens during the intro period, then doubles to $4 and $20.

They claim top results across agentic and coding benchmarks, with 77.9% on DeepSWE, 51.3% on AutomationBench and 68% on CWE-bench for vulnerability remediation; but we all know that we'll find out in practice once the community starts using it.

It seems they have been using it internally within Google and they talk about some use-cases it's unlocked, such as memory optimisation that freed 300 TiB across Google, and a code migration of over 800K lines in the Fuchsia Zircon kernel.

They have a phased rollout, and there's not much on their architecture or method so we'll have to wait to see what is behind it, but it is clear that the frontier labs competition is only accelerating.

Self-Improving Harnesses That Don't Overfit

Is it time to build our own self-improving Agent harness? Google Research + Stanford + others published a framework for self-improvement of agent harnesses:

This paper goes into a really interesting area of harness optimization, namely when an LLM keeps editing the prompts, tools and control flow around a frozen model against the same evolve set, the gains on that set shrink or vanish out of distribution.

The results are interesting: the RRSI they propose has the smallest gain on the evolve set they use, but the best out-of-distribution average (43.6 against 39.7 for the base harness). Compared to the unregularised run it also uses about a third fewer tokens per trial.

For this they keep the edit space fully open and regularise the search itself. An annealed budget caps how many edits each candidate can make, starting with bundled edits and ending with single ones, and a critic reads every diff before evaluation and rejects edits that encode task names or benchmark-specific answers. A pruner then flags any component that stopped earning its keep as a deletion target, so the harness doesn't keep piling up complexity.

For production ML practitioners the takeaway is that a harness tuned on your own eval set needs the same overfitting discipline as a model trained on it.

Raschka on Text Classification and Jev

This is a MUST read from Sebastian Raschka on a history of text classification from bag-of-words all the way to Jev, the decision model from TypeSafe AI:

It is impressive to see now the Imagenet moment for classification tasks, and it's great to see some of the Good Old datasets like IMDB to build an intuition of BERT vs Bag-of-words vs Jev, which lands at 96.5% with no fine-tuning at all, for about $0.65 and 22 minutes. This puts the perf into context, as a fine-tuned ModernBERT gets to around 95% but needs 23 minutes of training plus 7 of inference.

The architecture of Jev is still proprietary, so right now the approach is mostly an educated guess of a small encoder that scores each candidate class, trained with RL that rewards both the right answer and calibrated confidence. However it's a pretty good guess!

Definitely worth checking out if you want the background behind all the Jev buzz or a really fantastic refresher on the 2010s of ML / DL classificaiton methods.

Upcoming MLOps Events

The MLOps ecosystem continues to grow at break-neck speeds, making it ever harder for us as practitioners to stay up to date with relevant developments. A fantsatic way to keep on-top of relevant resources is through the great community and events that the MLOps and Production ML ecosystem offers. This is the reason why we have started curating a list of upcoming events in the space, which are outlined below.

You can also find our upcoming events and past talk recordings on the talks and events page.

Events we are speaking at this year:

Other relevant events:

In case you missed our talks, check our recordings below:

Open Source MLOps Tools

Check out the fast-growing ecosystem of production ML tools & frameworks at the github repository which has reached over 20,000 ⭐ github stars. We are currently looking for more libraries to add - if you know of any that are not listed, please let us know or feel free to add a PR. Here's a few featured open source libraries that we maintain:

  • SARC - Provides wrappers for popular agentic frameworks to enable guardrails and constraints that are enforced through the flow.
  • KAOS - K8s Agent Orchestration Service for managing the KAOS in large-scale distributed agentic systems.
  • Kompute - Blazing fast, lightweight and mobile phone-enabled GPU compute framework optimized for advanced data processing usecases.
  • Production ML Tools - A curated list of tools to deploy, monitor and optimize machine learning systems at scale.
  • AI Policy List - A mature list that maps the ecosystem of artificial intelligence guidelines, principles, codes of ethics, standards, regulation and beyond.
  • Agentic Systems Tools - A new list that aims to map the emerging ecosystem of agentic systems with tools and frameworks for scaling this domain

Please do support some of our open source projects by sharing, contributing or adding a star ⭐

About us

The Institute for Ethical AI & Machine Learning is a European research centre that carries out world-class research into responsible machine learning.

Check out our website

✉️ Email, 🐦 Twitter, 💼 Linkedin