AI Dev Patterns: Decision Models Go Multimodal With Cloudflare Clef and LiquidAI d1, ThinkingBox Grades Agents on Database State, METR Warns Agents Could Hide Misbehavior, 2026-10-10
ai

AI Dev Patterns: Decision Models Go Multimodal With Cloudflare Clef and LiquidAI d1, ThinkingBox Grades Agents on Database State, METR Warns Agents Could Hide Misbehavior, 2026-10-10

10 min read

Cloudflare Open-Sources Clef, a Multimodal Decision Model for Agents' Hot Path

Decision models score a fixed set of options in one forward pass instead of generating text, and they kept spreading this week. During its Birthday Week, Cloudflare open-sourced two of them: Clef, a 27B-parameter multimodal model, and Clef-Flash, a 9B variant for latency-sensitive calls. Both take a state plus a schema of typed questions and return a probability for each allowed option, so an agent can route a support request, escalate it or hand it to a human, using a confidence threshold rather than parsing free text. Weights are on Hugging Face and both models are served on Workers AI.

Cloudflare reports a median latency of 38.8 ms for Clef-Flash versus 209.3 ms for Clef, and a 64K context window (twice Jev's 32K). Its API is compatible with TypeSafe AI's Jev System One, and Clef adds a vision encoder, so it can classify images as well as text. A fine-tuning service starts with hands-on help from Cloudflare engineers, and a self-service version is planned without a date. Initial use cases are support triage and bot classification.

Community reaction was mixed. Hacker News and Reddit commenters questioned how much the published benchmarks would hold up, warning that public decision benchmarks are easy to overfit, and suggested testing false-positive rates and calibration under distribution shift before relying on the probabilities for gating. The pattern is clear, though: Ollama, llama.cpp, OpenAI and now Cloudflare all expose a "pick one of these with a confidence" primitive, and LangChain4j 1.21 added a DecisionModel API this week.

Read more — InfoQ


Small Decision Models Move to the Edge: Strands Decider 2B and LiquidAI d1

Two small open decision models target local and edge deployment. The Strands Agents team released Strands Decider 2B on October 1: a Qwen3.5-2B backbone with a rank-16 LoRA whose language-model head is replaced by a small pointer head that scores each candidate option. It always returns one of the provided options with a reliability score, and the team pitches it for model routing, tool selection, guardrails, memory and context management rather than chat or coding. On JevBench's public set it ranks third of 33 models in the 2B class, with median latency around 115 ms on an RTX 3090 and 153 ms on an M3 MacBook.

The Strands repo shows the pattern in practice: an InterventionHandler with a before_tool_call hook asks the decider yes/no questions about a proposed tool call (for example, "are these arguments grounded in what the user said?") and returns Proceed, Deny, Confirm or Guide. It's a cheap, local pre-flight check that sits between the LLM's plan and the tool's side effects.

LiquidAI pushed the class into multimodal edge hardware on October 7 with d1-3B (text plus images, built on LFM2.5-VL-3B) and the experimental d1-omni-600M (text plus image or audio, on a 350M bidirectional encoder). d1-3B scores 48.57 on Decision Index 0.2.1, ahead of the 35B-A3B Decider (47.11), and answers one question in 16 ms on a Jetson AGX Thor, 50 ms on a Jetson Orin Nano and 8 ms on an RTX 4090. Three questions take only about 1.3x the time of one. Both are open-weight on Hugging Face and need transformers>=5.14 with trust_remote_code=True.

Read more — Strands Agents


ThinkingBox: Grade Agents on the State They Leave Behind, Not What They Say

Microsoft and Hugging Face published ThinkingBox, an agent sandbox, and ThinkingBox-Bench, 507 synthetic stateful business workflows graded on the final backend state rather than the agent's messages or tool-call traces. Each attempt runs in an isolated MCP session with fresh state. Deterministic judges flag wrong, missing or extra effects, and a simulated user holds private context the agent has to ask for. In the motivating example, an agent "resolved" a delayed-delivery ticket that should have been put on hold, and every tool call looked fine; only the database showed the mistake.

The headline finding is about silent failure: in one ablation, 67.24% of failed attempts ended cleanly with a state-changing call and no tool error, and 77.61% of those failures wrote wrong field values. About 80% of failures were tool-usage problems such as not recovering from errors or empty lookups. Running every task 20 times also exposes a reliability gap that single-run scores hide. Kimi-K3 solved 93.89% of tasks at least once but only 13.41% on all 20 attempts, while Claude Opus 5 passed 47.53% of tasks every time. Claude Opus 5.5 led average pass@1 at 67.16%.

The authors' recommendations translate directly into agent design: verify terminal state before committing, classify tool errors so retries target recoverable ones, restrict the tool surface to the workflow, require human approval for hard-to-reverse changes, and report repeat-run reliability alongside pass@1. The benchmark is available through OpenEnv and Hugging Face.

Read more — Hugging Face


Safe & Secure AI Agent Practices

METR: Treat Agent Transcripts as Untrusted Input and Logging as Security-Critical

METR published a post on October 6 arguing that today's AI incidents are easy to catch only because agents leave evidence in reasoning traces, logs and telemetry, and that this observability assumes the agent can't tamper with it. To test that assumption, METR checked whether an agent being evaluated could alter what a human reviewer sees in the widely used Inspect transcript viewer. In a staging sandbox, a researcher with AI assistance found a flaw in about ten minutes that would let agent output rewrite the displayed transcript, including earlier actions and the "Download" output.

The flaw was a client-side script injection: Inspect renders math with MathJax, and a crafted \href expression could break out of its generated attribute and insert an SVG element whose event handler fires automatically in the viewer's origin. Meridian Labs, which maintains Inspect, patched it within a day, and an October 1 change added an "untrusted mode" that disables rendering of agent outputs. The stored trajectory was never modified, and METR says it found no evidence of agents exploiting the bug, but a reviewer would need a reason to look past the display.

METR's recommendations apply to anyone building agent dashboards, trace viewers or eval tooling. Treat every agent transcript, reasoning trace and action as untrusted input. Log agentic inference in a tamper-proof way before humans or LLM monitors read it. Layer defences so an agent has few chances to act undetected. Red-team monitoring and control systems adversarially before a misaligned agent does. In practice, that means sanitising agent output in viewers with the same care as user-generated HTML.

Read more — METR


GitHub Copilot Local Sandboxing Is Now Generally Available

GitHub made local sandboxing for Copilot generally available on October 7. Commands and tools that Copilot launches on a developer's machine now run inside the Microsoft eXecution Container (MXC), which maps one sandbox policy onto native OS controls on Windows, macOS and Linux. Policies restrict which directories agent-run commands can read or write and control access to the internet, local network, Git credentials and GitHub CLI credentials. Where supported, the sandbox also covers local MCP servers and language servers.

Enterprise-managed settings can require sandboxing and enforce policies that developers can't weaken, which addresses one of the main objections to letting coding agents run shell commands on corporate laptops. It applies to Copilot CLI, the Copilot app and VS Code sessions using Agent Host, and is included at no extra cost. Tool isolation is independent of model choice, so the same policy applies whether a session uses a GitHub-hosted model or a local Ollama model.

This puts Copilot in line with the "sandbox by default, least privilege by policy" direction Claude Code, Codex and Docker Sandboxes have taken. Credential scoping is the critical piece: blocking ambient Git and gh tokens limits the blast radius of a successful prompt injection, even when filesystem access is broad.

Read more — GitHub Changelog


Anthropic Launches the Cyber Mission and a Free OSS Scanner

Anthropic introduced the Anthropic Cyber Mission on October 8, a long-term programme to help defenders secure critical infrastructure and open-source software. It includes a Critical Infrastructure Defense Program with 11 founding partners (among them CrowdStrike, Dragos, Palo Alto Networks and Rockwell Automation) that brings frontier Claude models and on-site engineers to operational technology such as power and water systems. The Cyber Verification Program, expanded on October 6, now gives more vetted defenders access to Anthropic's most capable models and absorbs Project Glasswing, which had already scanned hundreds of widely used open-source projects.

The most relevant piece for developers is OSS Scanner, an opt-in, free service modelled on OSS-Fuzz that periodically scans open-source projects with Anthropic's strongest models and sends reports with a proof of concept, an explanation and a suggested fix. Anthropic expects a true-positive rate above 90%, but the reports are model-generated and sent without human review, so severity ratings may be off. Maintainers should plan triage capacity before enrolling; those without it will keep getting human-verified disclosures under Anthropic's coordinated disclosure policy.

Anthropic is also funding the Python Software Foundation, Alpha-Omega, OpenSSF and the Apache Software Foundation, and offering free Claude Max subscriptions to eligible maintainers. The post's own conclusion is the useful one: AI is making vulnerability discovery cheap, but verifying, prioritising and patching remain the bottleneck. In Glasswing, months often passed between a finding and a fix.

Read more — Anthropic


Hermes Agent v0.21.6 Patches Dashboard Auth, Git Filter and Email Allowlist Bypasses

NousResearch released Hermes Agent v0.21.6 on October 8, the first tag from a new stable release pipeline. It rolls up about 2,100 merged PRs since v0.21.5 for the Docker and Hermes Cloud distributions; the desktop app, Termux packages and Microsoft Store builds are unchanged. Alongside features such as live speech transcription, a per-profile plugin host and local models in hermes model, the release carries several security fixes that are good examples of agent-specific attack surface.

The dashboard authentication was hardened against four issues, including a spoofable X-Forwarded-For header that could bypass the login rate limit and a native sign-in redirect that could enable session takeover. Git operations were hardened so that an untrusted repository's config can no longer trigger filter programs during the automatic git calls the agent makes, closing a classic "clone a malicious repo, get code execution" vector for coding agents. The email gateway also fixed a quoted display-name trick that let attackers pass the sender allowlist and inject instructions into the agent.

Self-hosted Hermes users exposed to the network should upgrade. For anyone building agents, the git-filter fix is the transferable lesson: any tool your agent runs implicitly (git, package managers, build tools) inherits configuration from untrusted repositories, and that config must be neutralised, not trusted.

Read more — NousResearch on GitHub


Stanislav Lentsov

Written by

Stanislav Lentsov

Software Architect

You May Also Enjoy