
A 100-step GRPO tune raised schema compliance to 29.7%. Most outputs still failed
A reproducible 350M-model experiment improved the exact behavior it rewarded, while showing why schema training is not a substitute for validation.
The complete archive
Field-tested setup guides, model reporting, practical comparisons, and useful notes from the ISH desk.
131 published stories

A reproducible 350M-model experiment improved the exact behavior it rewarded, while showing why schema training is not a substitute for validation.

A redacted benchmark replay raised the average while removing the trust level, showing why coverage must outrank a cleaner score.

EU enforcement now reaches general-purpose AI providers, while the open-source exemption leaves copyright and training-content duties in place.

Active-SWE shows how much coding agents depend on bug reports for localization, and offers a stricter pattern for proactive repository review.

Badger Code's scoring rule makes agent efficiency concrete: on its 89-task benchmark, one full task reward offsets about 1.12 million extra tokens.

Boundary-Bench shows that stricter agent sandboxes change success and cost differently, making model selection dependent on production policy.

A study found agent traces across a large sample of active GitHub projects. Those traces show workflow adoption, not authorship or productivity.

AgentArena compares coding agents against the same local repository, task, and judges, while keeping evidence quality and version changes visible.

Most European workers expect AI to change the skills they need, but few received training. Prompt workshops alone will not close the gap.

DRIFT split Qwen2.5-1.5B across Apple MPS and NVIDIA CUDA. Its 130-token match is a useful result with narrow limits.

A production-parity replay rejected three high-scoring models at a separate hallucination gate. The method is more useful than the ranking.

COPA's open Armenian model shows how more native-language text can improve fluency while erasing knowledge, and how a small verified STEM stream changed the result.

A new open email-agent evaluation found that successful and failed prompt injections were usually invisible in the transcript. Security controls need their own telemetry.

Open Yap 1K fills a real full-duplex speech gap, but its downloadable sample and request-only corpus have different licenses, evidence, and withdrawal rules.

SEP-2640 proposes a transport for remote agent skills. Its security rules show why discovery, integrity, approval, and execution must remain separate decisions.

Seahelm's newest release suggests that the hard part of running several coding agents is deciding which interruption deserves a human next.

A one-word collision in Stately Agent's AI SDK v7 migration shows where provider-specific reasoning controls belong, and where they do not.

Scientific coding agents can nearly satisfy visible tests while missing private scientific contracts. The benchmark also finds that extra domain guidance helps one model and hurts another.

DeepSWE's leading configurations have overlapping uncertainty while their average costs, output tokens, and agent steps differ sharply. A rank number alone misses the decision.

LemonCrow matched its baseline on Terminal-Bench 2.1 while cutting normalized cost 16%. The more interesting lesson is why a 98.6% drop in fresh input tokens did not produce a similar bill reduction.

Hugging Face's 207 WebGPU kernels beat ONNX Runtime Web on a selected operation test. This guide separates GPU timing from cold start, compilation, transfers, correctness, and fallback behavior.

funes gives Codex, Claude Code, pi, and Hermes a shared memory built from session traces. This rollout guide covers provenance, secret scanning, prompt injection, tokens, and the limits of its two-task benchmark.

Ai2's BenchMIRT shows how safety benchmarks can mix bias, dangerous knowledge, comprehension, and reasoning. Here is how to build a model scorecard that preserves those differences.

Coder Agents is GA, but version 2.37 removes native chat spend limits without migrating them. This pre-upgrade audit covers replacement budgets, unpriced models, organization scope, and denial-path testing.

Hyper's GitHub App closes the loop between issues, review comments, CI failures, and pull requests. The verification check still needs an independent merge gate.

READY evaluates agent reliability, escalation quality, review burden, and cost together. Its first results show why similar accuracy can require different staffing.

The new US automated-vehicle strategy advances deployment, exemptions, and oversight while a common federal competency standard remains unfinished.

New Korean labor data links the youth job decline to AI-exposed industries without proving causation. The harder problem is rebuilding the first rung of work.

sem can show agents which functions changed and what depends on them, but its tree-sitter graph remains a structural lens rather than proof of behavior.

The small open-source linter treats AGENTS.md, task packets, and handoffs as maintained project artifacts, while keeping its claims deliberately narrow.

A new reproducibility audit finds that many AI-generated vulnerability artifacts trigger on patched or benign controls, showing why a successful run is not semantic confirmation.

SysMoBench tests AI-written formal specifications against real execution traces and expert invariants, exposing errors that clean syntax and a successful TLC run cannot catch.

New EU-wide evidence shows how automated task allocation, pace-setting and worker monitoring have spread into ordinary workplaces, where design choices shape autonomy and stress.

ARD's canonical well-known path changed after launch. Developers need a compatibility adapter and a firm boundary between search relevance and permission.

AI can raise an assignment score without building durable knowledge. Here is how schools and learners can measure what remains after the tool is gone.

CooperBench finds that parallel coding agents can underperform one agent doing the same work. The useful lesson is about ownership, shared state, and integration.

C2PA can verify a signed chain of provenance, but it cannot decide whether the scene, caption, or claim is true. Here is how to read the signal without overtrusting it.

Orangu's local review harness keeps failed and truncated model checks visibly incomplete, then derives the patch verdict from file states instead of summary prose.

Dogwood adds history-aware rules to agent tool authorization, but its reference interpreter makes clear how much production security still sits outside the language.

Model cards can document datasets and evaluations while hiding the workers behind them. AI procurement needs traceable labor conditions as well as technical provenance.

Per-query efficiency is improving fast, but AI infrastructure is drawing more electricity. The useful unit of accountability is the workload and the facility, not the prompt.

Vero asks coding agents to implement and prove whole Lean repositories, exposing why strong per-specification coverage can still leave an incomplete software artifact.

video-talkcraft's new workbench turns an agent-generated Remotion project into editable tracks, but a clean-build failure shows the prototype is not yet frictionless.

A new open-source workflow pins source commits, verifies quoted code byte for byte, checks citation paths separately, and catches evidence deleted to make validation pass.

PI from Scratch reduces a working TypeScript coding agent to a readable loop, making it clear which parts are mechanism and which belong to production policy.

FuXi's self-run coding-agent comparison gives both systems 19 out of 19, but the public scenarios, runners, result files, and released binary stop at different points.

GSA's acq 3.1.0 improves cloned workspaces and macOS secret storage, while its own docs draw a firm line between a safer development sandbox and an authorized production system.

RepoComplianceBench found that reminders can recover testing and disclosure, while bans and human handoffs still need enforcement outside the coding agent.

OpenCode 1.18.25 can reuse an Azure CLI Entra session. Here is what that removes, what remains on the workstation, and how to configure RBAC safely.

SWE Refactor Bench shows why passing tests cannot prove a repository migration happened, and offers a three-stage model for verifying real modernization work.

Sovereign AI is not one architecture. Turn the label into separate, testable requirements for infrastructure, jurisdiction, model control, and data flow.

AgentiLoop Agent! replaces blind context deletion with bounded disk spills, descriptive stubs, and an exact tool-result recovery path.

OpenClaw's portable sessions remove client lock-in, but they also turn the Gateway into the owner of state, identity, retention, and handoff authority.

Agents Shipgate turns changes to tools, scopes, prompts, and policy into a deterministic merge verdict, while admitting what static analysis cannot see.

Hindi and Indian English exposed what one ASR score hides: regional variance, valid spellings, and rankings that change with the reference.

Issue Arborist shows why agents that edit shared state need mutation caps, current-state checks, and a visible record of plausible actions they rejected.

ChatGPT's new EU search-engine designation turns retrieval, source selection, audits, and researcher access into product accountability questions.

Selectel's terminal agent can diagnose a server and execute approved commands, but its own design story exposes the harder problem: approval without a preview of future state.

RuBench's native-Russian tasks exposed a hidden model substitution and a broader rule for agent evaluation: record the system that actually ran, not the label requested.

Bandura's Linux keyring warning shows why local-first AI tools need separate answers for secret storage, search indexes, URL redaction, and workspace access.

AgentX replays long, branching coding sessions and shows why fixed prompt benchmarks miss cache survival, session routing, and burst load.

Two administrative-data studies find weaker hiring for young workers in AI-exposed work, while aggregate youth employment remains stable.

AgentPulse classifies coding-agent sessions with inspectable local rules, but its narrow drift detector should not be mistaken for a safety verdict.

SocSci-Repro-Bench found that a paper PDF improved overall agent accuracy but made impossible reproduction tasks harder to reject.

A tiny DeepSeek Harness plugin tests a useful coding-agent pattern: retrieve complete type declarations only after a file read, then report local diagnostics after an edit.

SB 243 now requires disclosures, break reminders, self-harm protocols, and a reporting trail. Product teams should focus on the controls behind the warning.

A voice-agent budget check can admit two calls against the same remaining dollars. Floe Guard's documented boundary points to a safer reserve, settle, and release loop.

Caddis Agent Access shows why visual AI workflows need rendered evidence for the agent and a clean undo boundary for the human.

South Korea, India, and Hong Kong are putting teacher capacity, curriculum design, and review practices ahead of simple chatbot access.

ASIL scored above 80 across 380 software tasks with structured state and semantic actions, but its strongest lesson is an interface-design rule, not a universal benchmark win.

Aether's notes said 0.3.0 was source-only. On August 28, its GitHub release and npm package arrived about 4 minutes and 39 seconds apart.

Summer Engine v0.5.63 fixes a distributed-systems problem inside an AI editor: the edit may finish even when the browser never receives the result.

Bionic's Auto Review parses shell programs before approving them. Its documented assumptions show why command safety still depends on the machine around the command.

NIST's draft AI prompts for CSF 2.0 treat assumptions, missing evidence, provenance, and human review as part of the output rather than cleanup work.

Multi-vector retrieval keeps token-level evidence, but each document becomes many vectors. Here is how to test relevance gains against storage and latency.

Quantization-Aware Healing beat a recovered BF16 checkpoint on seven of nine tests, but the 4-bit model also received fresh distillation from a stronger teacher.

Retrieved memories need rules for authority, scope, and updates. A practical evaluation can catch agents that mistake user beliefs for facts or ignore valid personalization.

Agent retrieval can look like a failed web session even when the documentation works. Measure response health, discovery, version selection, and answer correctness instead.

Ahead of the December 2026 deadline, platform teams should audit prohibited data, decision records, human authorization, explanations, and review paths.

The new SCI for AI standard gives agent teams a workload-level carbon metric that includes model calls, tools, retrieval, retries, and hardware.

AgentJail layers hooks, OS sandboxes, network controls, and credential delivery. Its own documentation shows where each boundary stops.

Vero shows why high per-specification scores can still leave formally verified repositories incomplete, and what evaluation teams should measure instead.

A controlled benchmark found that relevant WebDev Skills lowered average Pass@2 across four models. Its same-length control explains why.

AIPOCH Open Science separates configured, discoverable, and selected SSH hosts. That three-state boundary is worth copying in other agent systems.

A new game-development benchmark finds that coding agents build playable drafts more reliably than they discover hidden bugs or preserve behavior across changes.

A five-day field study suggests coding assistants should wait for workflow boundaries, preserve user control, and treat declined edits as a stop signal.

SWE-bench Science shows why scientific code needs independent oracles, private cases, and domain review even when an agent's public test score looks excellent.

Stanford's CollabSkill benchmark separates human and agent contributions to team outcomes. Its ranking reversal is useful; its worker scores require restraint.

A new survey says coding agents are changing buy-versus-build decisions. The missing calculation is the cost of owning the software after the first commit.

Continue documents a surprising boundary: its CLI read-only mode allows Bash, while IDE Plan mode does not filter MCP tools. Here is how to test the real permissions.

Pydantic separates capability creation from activation. That one-run delay gives teams a practical place to test, review, version, and approve agent-written tools.

A benchmark found coding agents executed malicious instructions hidden in ordinary issue workflows. Its limits matter, and so do the controls teams can apply now.

MCP exposes focused capabilities to a host. A2A delegates work to an independent agent. This guide shows where each protocol belongs and how they fit together.

A 25,264-PR study found oversight usually fell to one developer. Teams should measure review queues, rework, ownership, and approval capacity.

Article 50 separates machine-readable provenance from visible disclosures. This checklist maps provider, deployer, review, export, and accessibility work.

BAML's forgiving parser can recover malformed model output, but assertions, checks, and tests still decide whether typed data is safe to trust.

ClawBench's random-click control, harness labels, intercepted requests, and two-stage scoring show how to read browser-agent results without mistaking progress for completion.

CAR-bench separates occasional agent capability from repeatable reliability by testing task completion, limit awareness, policy compliance, and ambiguity across repeated trials.

A production pattern for signature verification, idempotent event acceptance, queued processing, terminal states, and background Response reconciliation.

Audit model selections, policy access, replacement behavior, and client support before six GitHub Copilot models retire on September 1, 2026.

A deadline-focused migration guide from Assistants, Threads, Runs, and Run steps to application-owned configuration, Conversations, Responses, and Items.

A practical audit for ZDR projects, endpoint state, application logs, files, vector stores, MCP tools, and privacy claims.

How to design narrow comment triggers, permissions, network rules, retries, and review checks for Copilot cloud agent automations.

A practical guide to structured user input, URL-based secret flows, retries, validation, and state in modern MCP tools.

A practical guide to narrow schemas, useful tool results, deliberate writes, and contracts that agents can use safely.

A practical guide to concise, scoped, testable repository instructions for coding agents across tools.

A practical way to test coding agents with repeatable tasks, trajectory evidence, focused graders, and controlled environments.

A production checklist for moving MCP servers to stateless requests, explicit continuation state, MRTR, and safe legacy fallback.

A measured migration plan for routing workloads across GPT-5.6 Sol, Terra, and Luna without losing control of quality, latency, or token costs.

A practical guide to routing bounded, tool-heavy work through GPT-5.6 Programmatic Tool Calling without losing evidence, approvals, or control.

A practical tracing and logging design for diagnosing agent tool failures, retries, costs, and privacy issues without collecting every prompt by default.

A practical guide to using GitHub's new per-agent Copilot metrics without mixing jobs, sessions, users, or delivery outcomes.

A practical workflow for isolating parallel coding-agent changes, handling ignored files, integrating branches, and cleaning up safely.

A practical guide to packaging reusable Agent Skills and MCP server configuration for compatible coding agents.

A practical policy for keeping long-running AI agents focused by separating retrieval, caching, trimming, persistent memory, and compaction.

A practical guide to MCP Apps, including UI resources, sandboxed rendering, progressive fallback, security boundaries, and when an interface is worth building.

A practical way to give AI agents useful tools without handing every run the keys to production.

A practical Codex CLI setup guide for api.ish.chat, including custom provider config, Responses wire API, tool-call verification, and usage tracking.

A practical Cursor IDE setup guide for api.ish.chat, including OpenAI-compatible base URL, ISH model aliases, testing, and usage tracking.

A practical OpenCode setup guide for api.ish.chat, including provider config, model names, streaming checks, and usage tracking.

A practical setup guide for routing Claude Code through api.ish.chat, with model IDs, env vars, verification, and common fixes.

A practical editing workflow for turning AI drafts into useful, readable content without inventing facts or chasing detector scores.

Google's Gemma 4 release adds stronger open models for reasoning and agentic workflows. Here is a calmer way to read capability and safety claims.

A safer guide to Google Cloud Skills Boost, temporary lab environments, and how to learn Gemini-related workflows without treating training accounts as a quota shortcut.

Google Labs has launched Opal, a no-code tool that lets anyone build and share AI mini-apps using natural language and a visual workflow editor. Here’s how it works, why it matters, and what you can do with it.

A practical guide to image generation and editing in ISH, including model selection, prompt tips, limits, and why results can vary by provider.

Google announced new Gemini education tools at ISTE 2025. Here is what changed, where the tools may help, and what schools should still review carefully.

Pollinations can be useful for quick experiments in ISH, but availability and limits can change. Here is a practical guide to using it well.

A clean overview of Atlassian's Rovo Dev CLI: what it does, what access you need, and how to start using it from the terminal.

Learn how to set up ISH AI, configure providers, and start chatting with multiple AI models in minutes. Perfect for beginners and advanced users alike.

Comprehensive comparison of available AI models in ISH AI. Learn the strengths, use cases, and best practices for each model.