Claude Haiku 5.5: The Cheapest Claude Yet, With Effort Control
Anthropic released Claude Haiku 5.5 (claude-haiku-5-5) on October 7, positioning it as its cheapest, fastest and most capable small model. Pricing is $0.10 input / $0.50 output per million tokens for prompts up to 100K tokens, and $0.50 / $2.50 above that. Cache reads cost $0.01 per MTok. Anthropic puts the average saving at about 75% versus Haiku 4.5 ($1/$5). It is also the first Haiku with an adjustable effort setting, and it's available on the Anthropic API, AWS, Google Cloud and Azure from day one. Claude Code's changelog lists a 1M-token context window.
The benchmark jump over Haiku 4.5 is large: OSWorld 2.1 (offline subset) goes from 15.7% to 72.4%, Terminal-Bench 4.0 from 0.0% to 39.2%, Humanity's Last Exam (no tools) from 10.2% to 45.9%, and GDPval-AA v2.1 from 735 to 1620. Sonnet 5.5 still leads on hard agentic coding (70.6% on Terminal-Bench 4.0, 83.9% on OSWorld 2.1), and Anthropic explicitly recommends Haiku 5.5 for high-volume, narrowly scoped work: summarisation, compaction, classification, database queries, browser use and subagent tasks.
Alongside the launch, Sonnet 5.5 cache-read prices were halved to $0.10 per MTok, which Anthropic estimates cuts the cost of most agentic Sonnet work by about 20%. The Python and TypeScript SDKs also gained beta support for computer use and browser use. GitHub Copilot added Haiku 5.5 the same day. For agent builders, the practical move is to push subagents, routing and context-compaction calls down to Haiku 5.5 and reserve Sonnet or Opus for the planning and coding loop.
Read more — Anthropic
Claude Code 2.1.289–2.1.296: Fail-Closed Hooks, Per-Agent Effort, and Marketplace Installs
Eight Claude Code releases shipped between October 3 and 9. Version 2.1.293 made Haiku 5.5 the default Haiku model, and 2.1.296 picked up the cheaper Sonnet 5.5 cache-read pricing. The hooks system got the most meaningful changes: 2.1.294 fixed a bug where prompt and agent hooks written as natural-language instructions (for example "Block commands that…") could allow what they were meant to block, and 2.1.295 added onFailure: "block" for command and HTTP hooks, so a hook that crashes or times out now fails closed instead of silently allowing the action.
Agent and plugin tooling also moved forward. 2.1.292 added an effort parameter on the Agent tool, so a parent agent can run cheap subagents at low effort, plus claude plugin install --marketplace <source>. Local stdio MCP servers and claude.ai connectors now negotiate MCP protocol 2026-07-28 by default. 2.1.295 raised the MCP tool-description cutoff from 2,048 to 16,384 characters and capped subagent skill preloading at 32. 2.1.296 added autoCompactWindow for subagents and a CLAUDE_CODE_WORKFLOW_SUBAGENT_MODEL variable.
Several changes tighten what a repository can do to you: project settings can no longer enable Claude in Chrome, CLAUDE_CODE_DISABLE_ATTACHMENTS can't be set from a repo's .claude/settings*.json, and Bash now asks before running pyright and more forms of ps. Security fixes in 2.1.292 closed a symlink-based read escape and a case where PreToolUse approvals bypassed UNC path prompts. The WebSearch budget also changed from a hard 200-call session cap to refilling at 100 calls per hour.
Read more — Anthropic
Codex CLI 0.161–0.162: Managed Worktrees, MCP Login, and GPT-6.1 Sol Ultrafast
OpenAI shipped Codex CLI 0.161.0 (October 7) and 0.162.0 (October 8), followed by bug-fix release 0.162.1. Version 0.161 makes GPT-6.1 Sol the default model in both the bundled and Amazon Bedrock catalogs and enables multi-agent V2 and Ultra reasoning on Bedrock. It adds /mcp login <name> to authenticate to MCP servers from the terminal, microphone and speaker selection for voice, and an opt-in "Daybreak" mode via --enable cli_daybreak. A sandbox fix lets an approved filesystem escalation grant broader write access while keeping denied reads and network restrictions intact.
Version 0.162 adds tools for creating and listing managed Git worktrees from trusted local projects, which makes parallel agent tasks safer without manual git worktree bookkeeping. It also adds task pinning in the Command Center (p), /copy for transcript blocks, clickable URLs in approval prompts, and live web access and remote compaction for custom Responses-compatible providers. apply_patch now preserves CRLF line endings, a long-standing Windows annoyance.
On October 8 OpenAI also brought GPT-6.1 Sol Ultrafast mode to Codex and ChatGPT Work for Pro and eligible Enterprise/Edu plans. It is off by default for Enterprise until a workspace owner enables it, and it supports inference residency in the US and Europe (EEA plus Switzerland).
Read more — OpenAI
GitHub Rewrites the Copilot Runtime in Rust and Adds Local Models to Copilot CLI
GitHub has migrated the runtime behind Copilot CLI, the Copilot app and the Copilot SDK from TypeScript/Node.js to Rust, replacing more than 800,000 lines of production code in about 14.5 weeks across 128 pull requests and 135 releases. AI agents wrote most of the Rust. Compilation, tests and human review caught regressions in state and lifetime handling, library semantics and lost optimisations, and 87.1% of 4,478 cargo check runs passed cleanly. Instead of a big-bang rewrite, GitHub replaced components one at a time behind a temporary N-API bridge (2,019 exports, 3,356 TypeScript call sites) so existing end-to-end tests kept working, then removed the bridge.
The payoff is mostly in embedding. The Rust runtime exposes a C ABI and can run in-process in host applications, and a startup-plus-single-turn scenario dropped from 5.25 seconds to 292 milliseconds. The old Node.js design added about 100MB of working set per client. Commentators noted that compilation alone doesn't prove behavioural equivalence for cancellation, retries and backpressure, but the case is a useful data point for incremental, agent-driven language migrations.
Separately, Copilot CLI 1.0.94+ can now discover models from a running Ollama instance via /model. Discovered models appear alongside cloud models and must be explicitly confirmed. They must support tool calling and streaming, and choosing a local model does not enable offline mode or disable telemetry; that still requires COPILOT_OFFLINE=true.
Read more — InfoQ
JetBrains Air Arrives Inside the IDE as an Agent-Agnostic Workspace
JetBrains opened the EAP for "Air in IDEs", bringing the agentic workspace it introduced on September 22 into IntelliJ-platform IDEs, either as a plugin or natively in 2026.3 EAP builds. Air is deliberately not an AI provider and ships with no agents installed. Instead it hosts the agents and subscriptions you already have: Codex, Gemini, GitHub Copilot, Claude Agent and Junie are supported out of the box, and others (JetBrains names Cursor) connect via the Agent Client Protocol (ACP).
The workspace shows parallel sessions across projects with activity, changed files, outgoing commits and per-session cost. Double-tapping Ctrl opens a prompt with the current editor context attached. Sessions can run on temporary Git worktrees, with results cherry-picked back, and supported agents get access to IDE tools for debugging, profiling, database exploration and semantic code search. That last point is the main reason to run an external agent inside the IDE rather than in a terminal.
Air itself doesn't require a JetBrains AI subscription if you bring your own agent subscriptions, and Junie Lite is free with a JetBrains Account. Cloud runs for long tasks need a JetBrains AI subscription and are currently limited to organisations with AI seats. Note that installing Air from the AI Assistant notification disables AI Assistant, which you can re-enable under Settings | Plugins.
Read more — JetBrains