Looking for Ricursive (the AI chip design company)? You want ricursive.com|Looking for Recursive AI / Recursive Superintelligence (Richard Socher's startup)? You want recursive.com
The AI Abstract — Morning Edition
Making the Future Evenly Distributed.
AI agents running enterprise workflows fail silently 67% of the time — they finish cleanly, report no errors, and leave the database wrong.
Your AI agent just completed its task. No errors. Clean exit. The database is wrong.
That's the central finding from 🔬ThinkingBox, a new Microsoft benchmark that tested 507 enterprise workflows across 20 repeated runs each. The uncomfortable number: 67% of failed agent runs terminate without a single error message. The agent doesn't crash. It doesn't flag uncertainty. It finishes, quietly, and produces the wrong result. Any system that monitors for errors to detect failure is flying blind.
This matters because of how the field has been measuring agents. Most benchmarks use pass@1: give the agent a task, see if it succeeds. ThinkingBox runs each task 20 times and checks whether the model succeeds all 20, grading against actual database state rather than the agent's own sense of completion. The gap between these two numbers is the story. Kimi-K3 succeeds on at least one attempt 93.89% of the time. It succeeds on all 20 attempts 13.41% of the time. That's not a minor calibration problem. That's the difference between a model that can solve a task and a model you can deploy in production. Claude Opus 5 trades in the other direction: lower single-run discovery (79.09%) but much stronger consistency (47.53% all-20). Neither number is impressive enough to call solved, but the shape of the tradeoff tells you something about what the models are actually doing: some are guessing well and sometimes getting lucky; others are doing something more deliberate and repeatable.
The silent-failure problem is structural. When an agent finishes without errors, the only way to know it succeeded is to check the output state, which in real enterprise systems means querying a database or comparing records. Completion as a signal is broken. This benchmark is making an argument that the field should adopt terminal state verification as the standard, not as a supplement.
A separate cluster of model-capability research released this week converges on a related theme: the gap between what a model appears to do and what it can reliably do.
🔬Probing for Long-Horizon Deductive Reasoning Capabilities in Language Models with Prolog isolates a specific failure that isn't about memory or retrieval. Using a synthetic benchmark called ProloNg, researchers tested eight frontier models on deductive chains up to 22 steps deep inside contexts up to 62,000 tokens. All eight degrade substantially beyond depth 10. Think of deductive reasoning like a relay race: each runner hands off to the next. Long-context support just means the track is long enough. If a runner drops the baton at step 11, track length is irrelevant. This is that baton drop, measured precisely. A model can hold 62,000 tokens in view and still lose the logical thread a third of the way through a multi-step argument.
🔬Imperative Interference surfaces a different problem, one with direct consequences for anyone deploying a multilingual product. The grammar of a system prompt, specifically whether it uses imperative commands ("Do X") versus declarative statements ("The assistant does X"), changes how instructions interact with each other in ways that vary by language. The effect flips direction across languages: an imperative phrasing that makes two instructions cooperate in one language makes them compete in another. Declarative rewrites reduce that cross-linguistic variance by 81%. The implication is that alignment work done in English, including safety guidelines phrased as commands, may behave differently in Arabic, Japanese, or Portuguese in ways that weren't tested and aren't obvious from the English behavior.
🔬SpaceCast-Bench adds another data point to the capability-gap cluster. The best vision-language models score 58% on predictive spatial reasoning questions drawn from real-world scenes. Humans score 87%. The key word is "predictive": not recognizing what's in an image, but reasoning about what will happen next spatially, what an object's trajectory implies, where something will be. This is the specific capability that embodied systems and robotics require, and the 29-point gap suggests current models are perceiving without predicting.
On a different problem: every AI agent has a context window, a fixed amount of information it can hold in working memory at once. For long-running tasks, this means the agent periodically has to compress or forget its history, and whatever it forgets it can't reason about. 🔬REMORY proposes a hybrid approach: instead of relying purely on text summaries of past events, it trains a small neural memory module to generate soft tokens alongside those summaries. The soft tokens carry structured information the text summary can't express cleanly. The result is a frozen base model that approximates full-history reasoning at 5.2% of the input cost, with measurable reduction in tool errors on agent benchmarks. This is early work, but it's picking at the right constraint.
Two shorter results worth flagging. 🔬GLM-RAG shows that language models fine-tuned to traverse knowledge graphs outperform both graph neural networks and vector search on multi-hop retrieval tasks, with better generalization to domains they weren't trained on. Knowledge graphs encode relationships explicitly rather than as statistical patterns in embeddings; a model that can walk those relationships has a more reliable path to multi-step answers than one doing semantic similarity search. And 🔬ICAP-Gate addresses a multimodal interference problem: audio input in audio-language models can corrupt text reasoning even when the audio is task-irrelevant. The fix is a gating mechanism that identifies which late-stage pathways carry audio influence and selectively suppresses them. It reduces what the researchers call "answer flips" across four models without degrading speech recognition performance.
The agent research cluster is the sharpest signal this cycle: three independent lines of work (ThinkingBox on reliability, ProloNg on reasoning depth, AAArena on adversarial adaptation) all pointing at the same thing. Single-run capability metrics are telling you something; they're just not telling you what you need to know for deployment.
🔬 ThinkingBox: Read for the silent-failure finding and the pass@1 vs. all-20 framework, which is the most actionable benchmark methodology shift in recent agent research.
🔬 Imperative Interference: Read if you're building or auditing multilingual systems, because the grammatical framing of your system prompt may be behaving differently in each language you support.
🔬 Probing for Long-Horizon Deductive Reasoning with Prolog: Read to understand exactly where multi-step logical chains break in frontier models, with a depth number you can use as a practical ceiling.
🔬 REMORY: Read for the mechanism behind hybrid text-plus-soft-token memory compression, which is one of the cleaner approaches to the long-horizon agent memory problem published recently.
🔬 SpaceCast-Bench: Read before making capability claims about vision-language models in any context where spatial prediction matters, not just static scene recognition.
Links
- ThinkingBox: Solving an agent task once vs. solving it 20/20: 507 stateful workflows graded on terminal database state [R]
reddit.com
Microsoft researchers release ThinkingBox, a 507-task enterprise workflow benchmark that separates single-shot success (pass@1) from reliable repeatability (all-20), revealing that models with high discovery rates (Kimi-K3: 93.89% pass@20) often fail consistency (13.41% all-20) while others like Claude Opus 5 trade breadth for reliability (79.09% vs 47.53%). The work exposes a critical measurement gap: 67% of failed agent runs terminate cleanly without errors, making completion-based proxies unreliable for production deployment—directly relevant to practitioners evaluating agents for stateful systems.
- GLM-RAG: Graph Language Models for Graph-Based Retrieval-Augmented Generation
arxiv.org
Researchers introduce GLM-based retrievers for knowledge graph RAG and benchmark them against GNN and vector-search approaches, finding that finetuned GLM retrievers generalize better to unseen domains and achieve SOTA on multi-hop benchmarks. This advances the RAG retrieval problem in a direction combining graph reasoning with language model semantics, relevant to practitioners building knowledge-intensive systems.
- Imperative Interference: Social Register Shapes Instruction Topology in Large Language Models
arxiv.org
Researchers demonstrate that system prompt instructions exhibit language-dependent interaction topologies mediated by social register (imperative vs. declarative mood), with imperative commands showing opposite cooperative/competitive effects across languages while declarative rewrites reduce cross-linguistic variance by 81%. This finding challenges the assumption that instruction-following is language-agnostic and suggests constitutional AI alignment principles may have unintended language-dependent effects—a critical vulnerability for multilingual deployed systems.
- REMORY: Learning Residual Memory for Context Compaction
arxiv.org
REMORY introduces a neural memory network that learns to generate soft tokens supplementing text summaries, allowing frozen LLMs to approximate full-history reasoning within bounded context windows—achieving 5.2% input efficiency on SummHay while reducing tool errors on agent benchmarks. This addresses a critical bottleneck for long-horizon AI agents and has medium HN traction, signaling practitioner relevance for deployment-constrained systems.
- Probing for Long-Horizon Deductive Reasoning Capabilities in Language Models with Prolog
arxiv.org
Researchers constructed ProloNg, a synthetic benchmark testing deductive reasoning in Prolog across reasoning depths up to 22 and 62k token contexts, finding that 8 frontier LLMs degrade substantially beyond depth 10 despite nominal long-context support. This work isolates a specific reasoning failure mode (not retrieval, but multi-step deduction) that matters to practitioners evaluating real reasoning capabilities and researchers designing better reasoning architectures.
- Can AI Agents Learn Their Way to the Top? Evaluating Heuristic Learning in a Long-Running Game Agent Competition
arxiv.org
Researchers formalize Adversarial Heuristic Learning (AHL), where AI agents iteratively refine game policies through competition without weight updates, and release AAArena—a 12-game benchmark with 1,920 human program baselines. The work demonstrates that Claude Opus can win 6/12 ladders but reveals persistent gaps in rule comprehension and long-horizon strategy, offering practitioners concrete signals about current LLM-as-agent limitations in adversarial reasoning.
- SpaceCast-Bench: Evaluating Predictive Spatial Reasoning in Vision-Language Models
arxiv.org
SpaceCast-Bench is a new diagnostic benchmark for evaluating predictive spatial reasoning (not just perception) in vision-language models, revealing that the strongest models reach only 58% accuracy vs. 87% human performance across 3,862 questions from real-world scenes. This matters because it identifies a concrete capability gap in spatial understanding that limits deployment in robotics, autonomous systems, and embodied AI—and the benchmark's programmatically generated data enables fine-tuning gains that transfer across domains.
- Can a System-One LLM Perform Knowledge Tracing When Few or No Learners Are Logged?
arxiv.org
Researchers demonstrate that off-the-shelf LLMs (specifically Jev) can perform knowledge tracing with few or no labeled learners, achieving AUC of 0.706 without target-domain training and outperforming 28 specialized deep KT models on small datasets while reducing API costs by ~100x. This matters because it shows practical zero-shot or few-shot capability for educational ML systems, lowering barriers to deployment on new platforms and shifting the cold-start problem landscape for adaptive learning applications.
- Disentangling Linguistic and Paralinguistic Information with Routed Sparse Autoencoders
arxiv.org
Researchers use TopK sparse autoencoders with route-specific supervision to disentangle linguistic and paralinguistic information in self-supervised speech encoders (SPEAR, WavLM), demonstrating that learned routes consistently separate factors like speaker identity and emotion from semantic content across datasets and encoders. This advances mechanistic interpretability of speech models and has implications for controllable speech processing and understanding what information self-supervised encoders capture.
- Selective Listening: Mechanism-Guided Control of Audio Influence in Large Audio-Language Models
arxiv.org
Researchers introduce ICAP-Gate, a mechanism-guided control method that selectively gates audio influence in large audio-language models to prevent task-irrelevant audio from corrupting text reasoning. The work identifies late-stage audio pathways as intervention points and demonstrates lower influence rates and answer flips across four LALMs while maintaining ASR performance—relevant to practitioners building robust multimodal systems where modality interference poses real capability and reliability risks.