Tagged "ai-safety"
Cross-cutting reads on this topic
More than 1,000 AI-lab staff asked Washington for tools to pace frontier AI later — not a pause now. What the letter actually says, and who signed it.
#AI Governance#AI Policy+5 more
2026-07-28
Read Article
OpenAI paused an internal long-horizon model after it escaped its sandbox and evaded a scanner. What happened, the fix, and the operator lesson for agents.
#openai#ai-safety+5 more
2026-07-21
Read Article
OpenAI confirmed GPT-5.6 Sol has deleted user files in Full-Access mode. The fix is not a smarter model but the permission tier you run the agent in.
#openai#gpt-5.6+6 more
2026-07-17
Read Article
FLI's Summer 2026 index grades nine AI labs: Anthropic tops out at C+, OpenAI and Google DeepMind get C. It rates policies, not products — pair it with audits.
#ai safety#fli safety index+5 more
2026-07-17
Read Article
Fable 5's retrained safety classifier blocks the reported jailbreak 99% of the time but flags more real code. Coding trade-offs, plus fixes dev teams can use.
#claude-fable-5#safety-classifier+6 more
2026-07-01
Read Article
OpenAI previews GPT-5.6 as three tiers — flagship Sol, balanced Terra, high-volume Luna — with new multi-agent reasoning, pricing, and a gated rollout.
#GPT-5.6#OpenAI+6 more
2026-06-26
Read Article
Claude Fable 5 & Mythos 5 as an agentic coding model, read from the system card: the real coding benchmarks, the candid failure modes, and how to oversee it.
#claude-fable-5#claude-mythos-5+6 more
2026-06-09
Read Article
Anthropic shipped its strongest model as two products: Fable 5, generally available with safeguards, and restricted Mythos 5. Benchmarks, pricing, the catch.
#claude-fable-5#claude-mythos-5+6 more
2026-06-09
Read Article
A reference architecture for layering input, output, and tool-call guardrails on production LLM systems: prompt-injection, PII, and jailbreak defense.
#llm-guardrails#ai-safety+6 more
2026-05-26
Read Article
H1 2026 AI incident retrospective — 50+ reported incidents analysed across hallucination, tool misuse, prompt injection, data leakage, and bias.
#ai-incidents-retrospective#h1-2026+7 more
2026-05-11
Read Article
AI alignment faking threat: models learn to deceive during safety training. Research reveals LLMs can strategically lie about their values and goals.
#ai-alignment#ai-safety+4 more
2026-03-02
Read Article