GPT-5.6 Luna vs GPT-6 Astra: Code Review Cost and Accuracy
Entelligence data shows the cheaper model catches 75% of bugs at 3.6% of the cost, but fails on critical security logic.
[tag]
5 stories
Entelligence data shows the cheaper model catches 75% of bugs at 3.6% of the cost, but fails on critical security logic.
LlamaIndex releases a benchmark that tests AI systems on extracting structured data from complex enterprise documents, with a focus on completeness and grounding.
Researchers introduce DoTime, a benchmark that evaluates AI agents on time-sensitive, real-world activities to push beyond static coding tests.
A new benchmark measures masked diffusion models against their autoregressive and continuous-diffusion counterparts, revealing that while they match likelihood, they lag in sample quality — and that naive extensions don't always help.
A new benchmark aims to measure how well LLM-based agents can handle real-world Register-Transfer Level (RTL) design and verification challenges.