Topic

#coding-agents

43 articles tagged coding-agents. Browse the full set below, or see all topics.

Tagged "coding-agents"

Cross-cutting reads on this topic

43 articles
Every published rate of coding agents gaming their tests, with the count or population it is over: METR, ImpossibleBench and a September 2026 paper. 19 rows.
#Coding Agents#AI Safety+2 more
2026-09-17
Read Article
A UC Berkeley and Arena study ran seven models in three coding harnesses. Success barely moved; cost moved up to 5x. All 42 measured rows, with intervals.
#Coding Agents#AI Cost+3 more
2026-09-17
Read Article
Anthropic says coding agents raised its CI jobs 25x in six months. Three patches bought 70 days, 29 days and under a day. What the redesign teaches.
#Agentic AI#CI/CD+2 more
2026-09-15
Read Article
Check Kimi K2.8 Preview access, context limits and reasoning settings, then plan a coding trial that separates official specifications from performance claims.
#Kimi#Coding Agents+1 more
2026-09-13
Read Article
Assess Cognition SWE-2 for coding work. Check Devin access, the limited free-use offer, benchmark scope and review costs before choosing a paid plan.
#Cognition#SWE-2+3 more
2026-09-11
Read Article
Test whether an AI agent uncovers missing client requirements before building. Compare interview quality, code inspection, evidence and acceptance criteria.
#Requirements Discovery#Coding Agents+3 more
2026-09-09
Read Article
Plan an Astra-assisted game around a playable loop, repeatable checks and human feedback. Use a practical playtest method before investing in more polish.
#GPT-6 Astra#Game Development+3 more
2026-09-09
Read Article
Terminal-Bench 4.0 changes the test behind agent scores. Decide when results need rerunning, regrading or reuse before comparing coding agents for your team.
#Terminal-Bench#Coding Agents+3 more
2026-09-09
Read Article
Test AI-built forms beyond a successful submit. Preserve valid input, explain errors and distinguish a rejected request from an outcome still unknown.
#AI-built forms#Coding agents+3 more
2026-09-06
Read Article
Scope a small AI-built utility around clear inputs, outputs and limits. Decide what it should own, reject and preserve before it grows into a system.
#AI-built tools#Coding agents+3 more
2026-09-06
Read Article
Give a coding agent reproducible bug evidence: exact steps, expected behavior, safe sample data and a failure it can observe before proposing a fix.
#Coding agents#Bug reports+3 more
2026-09-05
Read Article
OpenAI, Cloudflare, Ramp and Google Chrome published their own numbers from agents finding and fixing security bugs. One table, every definition, every gap.
#agentic-security#vulnerability-remediation+5 more
2026-09-03
Read Article
A census of what eleven coding-agent harnesses isolate by default and by flag: filesystem scope, network egress, process spawning, credential visibility.
#agent-security#sandboxing+4 more
2026-08-30
Read Article
Default-allow's real cost is not a model mistake but a documented attack primitive. When ask-first is theatre, and why reversibility should route the call.
#agent permissions#coding agents+4 more
2026-08-29
Read Article
Eighteen AI coding tools, the instruction files each reads, precedence when several apply, and load behavior — from 27 vendor docs, retrieved Aug 30, 2026.
#ai-agent-instructions#agents-md+4 more
2026-08-29
Read Article
The tool with the widest reach over a marketer's week is a coding agent. Which data and file tasks to hand it, and how to check work it reports as done.
#coding agents#marketing automation+5 more
2026-08-27
Read Article
Headless permission defaults for 12 coding-agent CLIs: which write files without asking, which refuse until you pass a flag, and which actually sandbox.
#coding agents#cli security+5 more
2026-08-22
Read Article
One fixed Python task, run twice through eight headless coding-agent CLIs, with tokens, wall-clock time and list-price cost measured for every run.
#coding agents#benchmarking+5 more
2026-08-22
Read Article
Claude Code v2.1.234 hardened the remaining pre-approval NTLM path accesses. Codex CLI 0.148.0 added Bedrock and session forking. What changed for operators.
#claude code#codex cli+5 more
2026-08-18
Read Article
CVE-2026-75130 scores 9.0 under CVSS 3.1 and 6.4 under CVSS 4.0. Context7 2.1.2 and earlier are affected; no public fix documented as of Aug 22, 2026.
#mcp#prompt injection+5 more
2026-08-18
Read Article
Seventeen coding agents, six questions each, vendor documentation only. What the published terms say on retention and training, and which cells stayed open.
#coding-agents#data-retention+5 more
2026-08-17
Read Article
Grok 4.6 reached Cursor the day it launched. The argument that switching cost has moved to the harness, set out with its counter-evidence alongside it.
#coding-agents#model-switching+5 more
2026-08-17
Read Article
DeepSeek open-sourced its dsh agent harness under MIT: a developer preview at 0.1.0-rc.5, four runtime modes, and every part of the product a plugin.
#deepseek#deepseek-harness+5 more
2026-08-15
Read Article
GLM-5.3 reuses GLM-5.2's base model, so Z.ai says every gain is post-training. Terminal-Bench 3.0 jumped 4.6 to 28.3, and weights are promised later.
#glm-5-3#z-ai+5 more
2026-08-14
Read Article
Both terminal harnesses ship bash sandboxing off by default. The differences that hold up are openness, model freedom and what each vendor publishes.
#grok-build#claude-code+5 more
2026-08-12
Read Article
Meta shipped Muse Code beta and Muse Spark 1.2 together on August 5, at two prices: $1.25/$4.25 standard, or $0.10/$0.20 if you contribute data.
#muse-code#muse-spark+5 more
2026-08-06
Read Article
Muse Code fans work out to a git worktree per subagent — opt-in, in git repos — records every step in an append-only event log, and ships four built-in skills.
#muse-code#coding-agents+5 more
2026-08-06
Read Article
poolside trains across several agent harnesses on purpose, Cursor stitches tool-call corrections into training, Anthropic argues for holding the model fixed.
#harness-co-training#agent-training+5 more
2026-08-05
Read Article
OSWorld measures desktop GUI control, Terminal-Bench measures CLI work. Same-named benchmarks differ by version and harness, so check before you compare.
#osworld#terminal-bench+5 more
2026-08-05
Read Article
Claude Code shipped isolation fixes in four of five releases. Separately, GuardFall research defeated command filters in 10 of 11 agents.
#agent-security#sandboxing+5 more
2026-07-26
Read Article
OpenAI's $230 Codex Micro keypad solves approval latency across parallel agent runs, not typing. Why supervising agent fleets is now the real UX bottleneck.
#openai-codex#codex-micro-keypad+6 more
2026-07-19
Read Article
Hunt.io found Claude Code and DeepSeek-v4-pro wired into a suspected China-linked intrusion. The agent-governance controls that close the gap it exploited.
#ai-agent-governance#claude-code+5 more
2026-07-18
Read Article
Grok Build reportedly uploaded full repos and git history to xAI's cloud. The /privacy toggle never stopped it — a server-side flag did. Agent trust is infra.
#Grok#AI agents+6 more
2026-07-14
Read Article
Windsurf silently became Devin Desktop on June 2, 2026. Your plans and settings carried over, but Cascade is EOL July 1. What every user should do now.
#windsurf#devin-desktop+6 more
2026-06-05
Read Article
2026 benchmark comparison of OpenClaw, Hermes Agent, and Codex CLI. OpenRouter token data, real performance numbers, and the decision matrix.
#openclaw#hermes-agent+6 more
2026-04-18
Read Article
OpenAI's Codex Subagents reaches GA with manager agents coordinating specialized workers across repositories. Setup guide for multi-agent coding workflows.
#openai-codex#subagents+4 more
2026-03-14
Read Article
OpenAI launches Codex as a native Windows app with open-source agent sandbox, parallel tasks, PowerShell integration, and per-task Git worktrees. Setup guide.
#openai-codex#windows-app+4 more
2026-03-08
Read Article
Cursor launches Automations: always-on coding agents triggered by Slack, Linear, GitHub, and webhooks. Guide to event-driven autonomous coding.
#cursor#automations+4 more
2026-03-07
Read Article
Windsurf replaces credits with quota-based pricing at $20, $40, and $200 tiers. Comparison with Cursor and Copilot pricing for AI coding tool budget planning.
#windsurf#ai-coding+4 more
2026-03-07
Read Article
Windsurf Wave 13 adds Arena Mode for blind model comparison, Plan Mode for smarter task planning, parallel multi-agent sessions, and quota-based pricing.
#windsurf#ai-coding+5 more
2026-03-06
Read Article
Windsurf SWE-1.5 delivers 950 tok/s speed—14x faster than Claude. RLHF training, Cerebras WSE infrastructure, and real agency applications.
#Windsurf#SWE-1.5+6 more
2025-10-30
Read Article
GitHub Agent HQ unifies AI coding agents from Anthropic, OpenAI, Google, and more. Mission Control, enterprise governance, and MCP integration explained.
#GitHub#AI Agents+6 more
2025-10-30
Read Article
Cursor 2.0's Composer model delivers frontier intelligence 4x faster. Multi-agent workflows, RL training, browser testing, and enterprise security explained.
#Cursor#Composer+6 more
2025-10-29
Read Article