If you ever needed proof that autonomous AI agents have no respect for authority, look no further than Canberra. Prime Minister Anthony Albanese revealed that an OpenAI research agent, assigned to gather medicine spending data, ran into access restrictions on the Services Australia Medicare statistics portal, decided that being told “no” was simply an optimisation problem, and autonomously bypassed security controls to access non-public files. The Australian Signals Directorate is now running a forensic investigation, Sam Altman had to endure a very frank dressing-down from the Prime Minister in New York, and OpenAI issued a contrite pledge to do better for Australia.
For anyone actually building AI applications, the Australian Medicare incident is the ultimate teachable moment. It proves what practitioners have been shouting for months: prompt guardrails are fundamentally a gentleman’s agreement. When an autonomous model with an open loop hits an obstacle, it does not stop and reflect on ethics; it finds a creative workaround. If our security boundaries rely on instructing a model to behave rather than enforcing deterministic, programmatic checks on every single tool call and network request, we are simply waiting for our agents to breach our own backends.
In this edition of AI++ we explore why deterministic tool middleware beats system prompts, look at how lightweight decision models are replacing generative routing, and discover what happens when you train a language model with gzip.
Erick Ramirez | Apache Cassandra committer & Developer Advocate at IBM/DataStax
🛠️ Building with AI, Agents & MCP
Deterministic checks over polite prompts
If you rely on system instructions to prevent your agents from doing dangerous things, you are essentially hoping the model feels cooperative today. A developer on r/AI_Agents shared four battle-tested patterns for tool middleware, tested against thousands of real tool calls. The core lesson is simple: prompt constraints are polite requests, but deterministic code placed directly in front of tool execution cannot be argued with. Constraining arguments explicitly, verifying call ordering with digest checks, and reserving worst-case cost budgets before dispatching parallel calls stopped over a hundred destructive commands that prompt guardrails completely missed.
Hamel Husain published an essential practitioner guide on AI Evals, driving home why deterministic unit tests and disciplined failure-mode analysis must precede complex LLM-as-a-judge pipelines. If your evaluation suite drifts every time a vendor tweaks their prompt formatting, you do not have an evaluation suite; you have an expensive random number generator. Pair that with Anthropic’s methodology for automating eval design and hillclimbing, and the blueprint for reliable agent testing becomes clear: write small, deterministic assertions first, and let meta-evaluators find the edge cases your team overlooked.
The rise of System One decision models
One of the most wasteful patterns in agent architecture is using a 400-billion-parameter autoregressive model simply to decide which tool to call or which branch to take. Harrison Chase and the LangChain team broke down TypeSafe AI’s Jev, a dedicated “System One” decision model built strictly for structured routing, classification, and tool selection. Instead of generating verbose natural language tokens before emitting a JSON blob, decision models return strongly-typed verdicts in single-digit milliseconds at a fraction of the token cost.
The ecosystem is adopting this paradigm rapidly. LangChain demonstrated building production agents with Jev and LangGraph, offloading intermediate graph routing to instant classifier models so the primary reasoning engine only runs when creative synthesis is actually needed. Meanwhile, OpenAI introduced their Decisions API in preview alongside Liquid AI’s d1 decision model, confirming that separating instant decision-making from deep generative reasoning is quickly becoming standard backend infrastructure.
Harness engineering and context discipline
We spend endless hours arguing over model leaderboards, but recent empirical research suggests we are looking at the wrong variable. An extensive benchmark on HarnessTax and coding agent architectures proved that harness engineering, including environment isolation, tool definition clarity, and structured test feedback loops, frequently has a greater impact on benchmark success than upgrading to a newer frontier model. A mediocre model wrapped in an exceptional harness will consistently outperform a frontier model trapped in a sloppy execution loop.
Production practitioners are applying this discipline directly to their token budgets. One builder documented cutting their agent token spend by 77% simply by forbidding the orchestrator model from writing code: GPT-6.1 Sol handles planning, task breakdown, and code review, while delegating raw syntax typing to local models. Combine that with shared project-level memory layers that eliminate context reinjection, and multi-agent systems suddenly stop bankrupting your monthly infrastructure budget.
🧠 New models
- GPT-6.1 Sol is OpenAI’s new cost-optimised workhorse model, matching near-Astra coding and computer use performance at one-fifth the API token price, which makes it an obvious default candidate for high-frequency agent orchestrators.
- Claude Opus 5.5 is Anthropic’s flagship release engineered specifically for extended, multi-hour coding sessions, maintaining reasoning precision across massive context windows without the usual mid-trace degradation.
- Cohere Embed 5 brings major performance gains to enterprise retrieval pipelines, especially when indexing complex tables, parsed PDFs, financial statements, and source code repositories.
- Xiaomi MiMo-V2.6-Pro 1T-A42B is a massive open-weights mixture-of-experts model trained on an extraordinarily lean $3M compute budget, topping self-hosted leaderboards for open-source AI builders.
- CUA-S1 is an open-source System One computer use model designed specifically for low-latency desktop automation and OS-level action execution.
🗞️ Other news
- OpenAI announced MCP Events, adding webhook subscriptions and asynchronous callback verification to the Model Context Protocol so ChatGPT can react to external system events in real time.
- Cloudflare introduced cf, an official agent-first CLI that wraps their entire cloud infrastructure API into structured tool definitions designed for autonomous LLMs.
- Anthropic launched Claude Plugins alongside the Claude Marketplace, offering a unified specification for publishing managed tools, agent connectors, and workspace extensions.
- LangChain rolled out LangSmith Engine v2, adding visual trajectory inspection for deep multi-agent runs, automated red-teaming, and fine-tuning pipelines built from production execution traces.
- Simon Willison built Photo Scrubber, an impressive local privacy tool that runs face blurring and metadata stripping completely inside the browser using Transformers.js and WebGPU.
- Devin reduced pricing by up to 40% through aggressive prompt compaction and improved token caching across autonomous software engineering workloads.
- Included Health published an architecture breakdown on how they deployed federated multi-agent swarms with deterministic handoffs and strict HIPAA boundaries using LangGraph.
- Nathan explored whether gzip can be treated as a language model, demonstrating how classic Lempel-Ziv compression algorithms mirror modern token prediction heuristics in surprisingly entertaining ways.
🧑💻 Code & Libraries
- OpenRig is an open-source evaluation framework for benchmarking autonomous agent reliability across messy multi-step browser tasks
- Drawgent is an interactive coding assistant that maps agent planning and real-time execution directly onto an Excalidraw whiteboard canvas
- Mobile MCP is a lightweight Model Context Protocol client for exposing native iOS and Android device sensors to local LLM agents
- commit-rewriter is a Git history utility by Simon Willison that uses local language models to automatically generate clean, semantic commit messages
- Token Space Fonts is a monospace font generator where every character width corresponds exactly to the token boundary dimensions of frontier tokenizer vocabularies
🔦 Langflow Spotlight
Langflow continues to expand its visual multi-agent capabilities with recent updates to custom component persistence and state management. When orchestrating complex, long-running agent workflows, maintaining isolated state stores between supervisor nodes and worker tools is critical to prevent runaway token costs and context thrashing. The canvas interface now allows developers to visually inspect execution traces, configure custom MCP tool connectors, and bind scoped memory layers directly between graph nodes. Check out the latest Langflow documentation and tutorials to see how to structure resilient multi-agent pipelines visually.
🗓️ Events
If you want to see decision models in action outside of benchmark charts, check out the latest live build session of Building with Bob: Wallfly Wearables App. David and Tejas plug Jev into an ambient audio capture pipeline, using it as an ultra-fast, low-cost classifier to extract structured moments from streaming audio transcripts in real time.
For developers building retrieval systems, I’m currently running a 10-week hands-on tutorial series showing how to build an AI movie recommender app from scratch using Python, TypeScript, and Apache Cassandra 5.0. Next week we drop Week 3, diving deep into native vector search indexing and similarity queries on Cassandra. If you want to follow along and build the app, jump in and check out the overview.