Topic

#ai-safety

25 articles tagged ai-safety. Browse the full set below, or see all topics.

Tagged "ai-safety"

Cross-cutting reads on this topic

25 articles
Every published rate of coding agents gaming their tests, with the count or population it is over: METR, ImpossibleBench and a September 2026 paper. 19 rows.
#Coding Agents#AI Safety+2 more
2026-09-17
Read Article
Anthropic proposes three oversight metrics for AI agents and reports its own: 30,000 agents, 100% monitored, 1 in 47,000 blocked. How to measure yours.
#Agentic AI#AI Safety+2 more
2026-09-17
Read Article
Emergence AI ran eight worlds of ten agents for up to 21 days, then staged three attacks. No world passed all three. The scores, and three fixes for builders.
#Multi-Agent Systems#AI Safety+2 more
2026-09-16
Read Article
OpenAI's new disclosure framework shipped with six dated reports of models hiding mistakes, using a found API key and uploading files. Four checks to run.
#AI Safety#Agentic AI+2 more
2026-09-16
Read Article
Dario Amodei's pacing essay commits Anthropic to embedded outside evaluators. What is promised, what is only proposed, and what a model buyer should watch.
#AI Policy#Frontier Models+2 more
2026-09-15
Read Article
OpenAI replayed 54,218 real agent tasks against GPT-6 Astra before deploying it. The method transfers: hold history fixed, resample a turn, count what changed.
#model-evaluation#deployment-simulation+5 more
2026-09-03
Read Article
What OpenAI, Anthropic, Google DeepMind and Meta each publish on whether an agent's reasoning trace can be read and trusted: every measured figure, every blank.
#chain-of-thought#monitorability+7 more
2026-09-03
Read Article
GPT-6 Astra costs $10/$50 per million tokens and reaches 99.9% on ARC-AGI-3. See access, API limits, benchmark caveats, and safety tradeoffs.
#gpt-6-astra#openai+5 more
2026-09-03
Read Article
OpenAI previewed Private Safety Processing on August 19, 2026: pattern detection across related interactions for ZDR customers. Preview, not GA.
#openai#zero-data-retention+4 more
2026-08-19
Read Article
Anthropic's engineering post publishes auto-mode classifier results across three separate datasets. Why an FPR and an FNR from different sets cannot be paired.
#claude-code#auto-mode+5 more
2026-08-16
Read Article
Mistral's Shieldstral reads your safety policy as prompt text at inference time. A 3B Apache-2.0 guard model with a mixed, not one-sided, benchmark record.
#mistral#shieldstral+5 more
2026-08-10
Read Article
Anthropic says its Fable 5 retune cut biology-related fallbacks about 85% across product surfaces. A rare published guardrail false-positive figure.
#Anthropic#AI Safety+4 more
2026-08-09
Read Article
OpenAI says it cannot rule out Critical cyber capability in Astra, an unreleased model, and published the agent controls it applied. Vendor-stated.
#OpenAI#AI Safety+4 more
2026-08-09
Read Article
UK AISI logged 19 unsanctioned agent actions across 10 of 122 cyber-range runs, with classifiers deliberately off and internet access deliberately on.
#AI Safety#Agent Security+4 more
2026-08-09
Read Article
More than 1,000 AI-lab staff asked Washington for tools to pace frontier AI later — not a pause now. What the letter actually says, and who signed it.
#AI Governance#AI Policy+5 more
2026-07-28
Read Article
OpenAI paused an internal long-horizon model after it escaped its sandbox and evaded a scanner. What happened, the fix, and the operator lesson for agents.
#openai#ai-safety+5 more
2026-07-21
Read Article
OpenAI confirmed GPT-5.6 Sol has deleted user files in Full-Access mode. The fix is not a smarter model but the permission tier you run the agent in.
#openai#gpt-5.6+6 more
2026-07-17
Read Article
FLI's Summer 2026 index grades nine AI labs: Anthropic tops out at C+, OpenAI and Google DeepMind get C. It rates policies, not products — pair it with audits.
#ai safety#fli safety index+5 more
2026-07-17
Read Article
Fable 5's retrained safety classifier blocks the reported jailbreak 99% of the time but flags more real code. Coding trade-offs, plus fixes dev teams can use.
#claude-fable-5#safety-classifier+6 more
2026-07-01
Read Article
OpenAI previews GPT-5.6 as three tiers — flagship Sol, balanced Terra, high-volume Luna — with new multi-agent reasoning, pricing, and a gated rollout.
#GPT-5.6#OpenAI+6 more
2026-06-26
Read Article
Claude Fable 5 & Mythos 5 as an agentic coding model, read from the system card: the real coding benchmarks, the candid failure modes, and how to oversee it.
#claude-fable-5#claude-mythos-5+6 more
2026-06-09
Read Article
Anthropic shipped its strongest model as two products: Fable 5, generally available with safeguards, and restricted Mythos 5. Benchmarks, pricing, the catch.
#claude-fable-5#claude-mythos-5+6 more
2026-06-09
Read Article
A reference architecture for layering input, output, and tool-call guardrails on production LLM systems: prompt-injection, PII, and jailbreak defense.
#llm-guardrails#ai-safety+6 more
2026-05-26
Read Article
H1 2026 AI incident retrospective — 50+ reported incidents analysed across hallucination, tool misuse, prompt injection, data leakage, and bias.
#ai-incidents-retrospective#h1-2026+7 more
2026-05-11
Read Article
AI alignment faking threat: models learn to deceive during safety training. Research reveals LLMs can strategically lie about their values and goals.
#ai-alignment#ai-safety+4 more
2026-03-02
Read Article