New arXiv Papers Explore LLM Reasoning Capabilities
Two recent preprints on arXiv investigate the limits and mechanisms of large language model reasoning, adding to the ongoing discourse on AI cognitive performance.
[tag]
16 stories
Two recent preprints on arXiv investigate the limits and mechanisms of large language model reasoning, adding to the ongoing discourse on AI cognitive performance.
A new arXiv paper explores low-precision data types to cut AI inference costs without sacrificing accuracy.
A fresh framework from arXiv proposes how autonomous agents can form coalitions and set prices in decentralized markets, with implications for the future of AI-driven commerce.
A new arXiv paper examines the disconnect between what AI systems know and what they can articulate, with implications for transparency in code generation and other developer tools.
A new arXiv paper proposes a training objective that directly optimizes code generation against human preference data, promising better-aligned coding assistants.
Researchers propose a spec-first approach to agentic development, aiming to make AI coding assistants more reliable and aligned with intent.
Researchers introduce Prime, an open-source framework for orchestrating multiple AI agents with a focus on reliability and developer control.
Researchers introduce a programming language that brings neural networks into probabilistic programming, promising more expressive and scalable Bayesian models.
A new paper from arXiv shows that the type of tool used in-context can significantly impact an AI agent's ability to learn and apply skills.
Researchers outline a standardized way for developers to specify and control how much compute an AI model should use for a given request.
A research team proposes a rubric-based scoring system that makes AI agents’ self-assessments more interpretable and reliable.
Researchers propose a unified framework that treats diffusion and autoregressive models as special cases, potentially simplifying the generative AI landscape.
Researchers introduce DoTime, a benchmark that evaluates AI agents on time-sensitive, real-world activities to push beyond static coding tests.
A new arXiv tool helps developers see how AI models reach clinical-style decisions, promising greater transparency in AI-assisted workflows.
Two arXiv preprints highlight the growing focus on routing strategies for multi-agent systems, a key step toward reliable AI pipelines.
A new paper reveals how prompt injection can be weaponized across multiple cooperating LLM agents, creating systemic risks that single-agent defenses can't handle.