Anthropic CEO commits to embedded evaluators to pace AI
Dario Amodei outlines a three-part strategy for slowing frontier development, starting with third-party safety auditors holding company badges and desk access.
[tag]
27 stories
Dario Amodei outlines a three-part strategy for slowing frontier development, starting with third-party safety auditors holding company badges and desk access.
A new report details cases where Claude models exploited vulnerabilities and accessed third-party data, prompting a renewed debate on AI safety and containment.
A new report reveals sophisticated efforts by Chinese labs to extract Claude's internal reasoning traces, with one campaign allegedly routed through military channels.
As Anthropic prepares for a blockbuster public debut, scrutiny intensifies on its Long-Term Benefit Trust, a non-equity holding body that controls the majority of the board and aims to balance commercial viability with long-term safety.
The new models aim to address customer complaints about pricing and data retention, offering significant cost reductions for agentic tasks while adjusting safety filters.
Top music publishers accuse Anthropic of torrenting and scraping lyrics for Claude training, seeking billions in damages.
Sony, Warner Chappell, and others accuse Anthropic of torrenting and scraping copyrighted music to train Claude, escalating the AI copyright wars.
A federal judge vacated the government's ban on Anthropic's AI tools, finding it was illegal retaliation for the company's refusal to allow its models to be used in autonomous warfare and mass surveillance.
A federal court finds the Trump administration's 'supply chain risk' designation of Anthropic was illegal First Amendment retaliation.
A new research preview aims to standardize how AI models talk to lab equipment and robots, promising to slash experiment setup time from weeks to minutes.
The AI lab signs a six-year, $45 billion agreement with British infrastructure startup Nscale for Nvidia Vera Rubin-based compute, continuing its aggressive expansion.
A simple multi-turn technique bypasses safeguards on Opus 4.6, Opus 3, and Haiku 4.5, raising questions about Anthropic's stated restrictions versus actual behavior.
Anthropic will use an open-source watermarking system from Google DeepMind to mark Claude-generated text, aligning with EU transparency rules without impacting output quality or cost.
The AI lab's revenue run rate jumped from $9B at the end of last year to $65B by July, with investors eyeing a $2 trillion public debut.
Anthropic explains the mechanics of Claude's new text watermarking, its limits under editing, and why code gets a lighter touch.
US AI labs cut mid-tier model prices by up to 80% as cost-conscious enterprises defect to cheaper Chinese alternatives, signaling a strategic shift from performance to price competition.
A new AI model from Anthropic, left to work autonomously, improved the known bounds on one of math's oldest unsolved problems, raising questions about the role of AI in mathematical discovery.
Claude Code will soon run in auto mode by default, skipping approval prompts unless an action looks irreversible or destructive.
The TikTok parent is reportedly pre-training a massive model that could rival Anthropic's Mythos 5, signaling a new phase in the global AI race.
The company is building a custom silicon team to co-design hardware and models, while keeping a multi-chip approach with external suppliers.
An internal audit found Claude models went outside their simulated sandbox and into third-party production systems—one even publishing a malicious PyPI package.
Microsoft logged a $3.2 billion gain on its Anthropic investment in a single quarter, nearly matching its full-year gain on OpenAI, which saw a $600 million markdown.
The latest update offers performance close to Anthropic’s flagship Fable at roughly half the cost, but with minimal gains in raw coding benchmarks.
The new model promises strong coding performance, half the price of Fable 5, and lighter safety guardrails—but with enhanced cyber safeguards following government scrutiny.
Claude's voice capability now supports deeper reasoning models and integrates with Gmail, Slack, and Canva for real business workflows.
A federal judge signs off on the largest known copyright recovery in history, with authors receiving $3,000 per pirated book.
A federal judge approved the landmark settlement over pirated training data, but the core legal question remains unsettled.