3RecursiveIntelligence.io

Looking for Ricursive (the AI chip design company)? You want ricursive.com|Looking for Recursive AI / Recursive Superintelligence (Richard Socher's startup)? You want recursive.com

The AI Abstract — Morning Edition

AI/MLLatest

Making the Future Evenly Distributed.

A new agent architecture matches the performance of today's best long-horizon AI research systems while using 84% fewer tokens — not by being smarter, but by throwing away memory entirely.

The most expensive thing an AI agent does during a long task isn't reasoning. It's remembering. Every step an agent takes gets appended to a growing context window, and by the time a research session is hours old, most of what the model is "reading" at each step is historical noise it can't act on anyway. The context bloats, the cost compounds, and eventually the whole thing collapses under its own weight. That's not a bug in any particular system. It's the architecture.

🔬Stateless Language Agents proposes a structural fix: stop storing agent state in the conversation history at all. Instead, maintain research state in an external harness, a piece of scaffolding that lives outside the model, and at each new step reconstruct only the context that step actually requires. The model never "sees" the whole session. It sees a fresh, minimal snapshot built from the structured state. Think of it like the difference between a surgeon who reads the entire patient chart aloud before every incision versus one who reviews only the current procedure notes. The surgery is the same. The reading time is not.

The results across software engineering, kernel optimization, and algorithm design tasks are worth taking seriously: SLAs match or beat recent agent frameworks on task performance and hit baseline-equivalent outcomes with 84% fewer tokens. That's not a compression trick. It's a consequence of never accumulating irrelevant history in the first place. The architecture is stateless at the model level precisely because state is managed deliberately at the system level. The signal tracker shows this "long-horizon agent" problem clustering across multiple concurrent papers, which suggests the field has converged on it as a real constraint worth attacking from several directions simultaneously.

The alignment picture looks worse the closer researchers measure it. 🔬Toward Alignment Scaling Laws is an attempt to put numbers on a question that has mostly been argued qualitatively: as models get larger, does the safety problem get easier or harder? The paper formalizes "alignment burden" as a scaling law with an exponent. If that exponent is less than 1, alignment gets proportionally easier as you scale. If it exceeds 1, you're accumulating what the authors call alignment debt: the safety problem grows faster than the model does. The worst-case exponent across all risk categories determines which regime you're in, not the average.

The preregistered measurements on Pythia and Qwen models give early, sobering numbers. Truthfulness improves slightly with scale (exponent roughly -0.05, meaning it gets marginally easier). Sycophancy, the tendency to tell users what they want to hear rather than what's true, scales with an exponent of 0.89, close enough to 1 to be concerning. Backdoor vulnerability is undetermined, which is its own finding. What the paper establishes is methodology: a way to measure these exponents empirically before deploying at frontier scale, rather than discovering the regime you're in after the fact. That's the contribution. The numbers themselves are preliminary, and the authors are explicit about that, but having a formal framework for this question is categorically different from not having one.

Meta just open-sourced a piece of production infrastructure that most organizations don't know they need until they hit the wall it solves. 🔬Rebalancer is a C++ library for assignment problems: given a set of objects and a set of bins, find the optimal allocation. That sounds abstract. It is, at the problem-formulation level. At the application level it's every resource scheduling decision a large system makes: which server handles which jobs, which cache stores which data, which compute node gets which model shard. Meta has been running this in production for nine years. It now handles 40 million of these problems per day. The release includes a dual-solver architecture combining exact optimization with local search heuristics, Python bindings, and a debugging interface. It's available via pip install rebalancer. The class of problem it solves is technically NP-hard, meaning no algorithm is guaranteed to find the perfect answer in reasonable time at scale. Rebalancer doesn't claim to solve that. It claims to find very good answers very fast, which is what nine years of production use confirms.

Three architectural papers round out the payload, each solving a different scaling bottleneck.

🔬RAM-Net addresses a known failure mode in efficient sequence models. Standard linear attention, the approach used by architectures like Mamba to process sequences in linear rather than quadratic time, suffers from what you might call channel bleed: different pieces of information stored in the recurrent state interfere with each other, degrading recall over long sequences. RAM-Net adds address-based access to that state, so retrievals are targeted rather than diffuse. The analogy is the difference between reading an entire book to find a fact versus having an index. It outperforms Mamba2 on both fine-grained recall and general language modeling quality, which is an unusual combination since those objectives often trade off.

🔬STAVE attacks in-context learning costs from the other direction. When you show a model examples of what you want at query time, you pay a token tax on every single request: the examples get re-encoded from scratch every time. STAVE compresses the information from those examples into two task-specific vectors injected into the model's embedding layer once, then reused. Two vectors. Across six multimodal and five language models, performance matches or beats the demo-heavy baseline. For production systems running millions of queries per day, the cost reduction is not marginal.

🔬Pleias-RAG makes a narrower but practically important point: you can bake source attribution into a small model architecturally rather than prompting for it. The 350M and 1B parameter models in this family are purpose-built for retrieval-augmented generation with citation grounding built into the training objective. They outperform comparable small models on multi-hop question answering benchmarks. The relevance for anyone deploying AI in resource-constrained environments where hallucination has real consequences is that "small model, trustworthy output" is no longer a contradiction in terms.

Two safety-adjacent papers worth flagging briefly. 🔬DMAST identifies a specific attack surface in multimodal web agents: adversaries who coordinate corrupted visual and textual inputs simultaneously to hijack agent behavior. The proposed defense uses a three-stage training process combining imitation learning, supervised fine-tuning, and self-play. The cross-modal attack vector is appearing in multiple concurrent papers, suggesting practitioners building browser automation agents should treat it as a near-term deployment concern rather than a theoretical one. Separately, 🔬DIR-R challenges the standard assumption in model unlearning: that the parameters most statistically associated with a piece of knowledge are the right ones to edit when you want to remove it. The paper shows that dynamic parameter selection outperforms localization-based approaches on standard unlearning benchmarks. If you're building systems that need to reliably forget specific information, the "find the relevant weights and edit them" approach is not as clean as it looks.


🔬 Stateless Language Agents: Read for the architectural mechanism behind the token reduction — the external harness design is the part worth understanding.

🔬 Toward Alignment Scaling Laws: Read for the formal framework and the sycophancy exponent specifically — that number will matter more as frontier models get larger.

🔬 RAM-Net: Read for the address-based state access design if you're following efficient architecture development — the comparison with Mamba2 is the empirical anchor.

🔬 Pleias-RAG: Read if you're evaluating small models for production RAG — the citation grounding benchmark methodology is what to scrutinize.

📰 Meta AI Open-Sources Rebalancer: Read for the dual-solver architecture description before deciding whether the library fits your allocation problem class.

Links

  1. Stateless Language Agents: Scaling Long-Horizon Automated Research

    arxiv.org

    Researchers introduce Stateless Language Agents (SLAs), a framework that decouples agent state from conversation history by maintaining research state in a harness and reconstructing fresh contexts for each agent invocation. On software engineering, kernel optimization, and algorithm design tasks, SLAs outperform recent frameworks and achieve baseline performance with 84% fewer tokens, addressing fundamental scaling failures in long-horizon automated research systems.

  2. Meta AI Open-Sources Rebalancer: A C++ Assignment Solver That Runs About 40 Million Placement Problems a Day

    marktechpost.com

    Meta open-sourced Rebalancer, a C++ library for solving large-scale assignment problems (object-to-bin allocation) that has run production workloads for 9+ years and now handles 40 million problems daily across Meta's infrastructure. The release includes a dual-solver architecture (MIP + local search), Python bindings, PyPI distribution, and a debugging UI; it directly addresses NP-hard resource allocation at hyperscale—a capability gap practitioners and infrastructure teams can now deploy via `pip install rebalancer`.

  3. Toward Alignment Scaling Laws: A Framework and First Preregistered Measurements

    arxiv.org

    Researchers propose a mathematical framework modeling alignment burden as scaling laws B_r(N)=a_rN^alpha_r across risk categories, proving that the worst-case exponent determines long-run regime dynamics and that alignment debt accumulates when exponents exceed 1. Preregistered experiments on Pythia and Qwen models reveal heterogeneous scaling (truthfulness -0.05, sycophancy 0.89, backdoors undetermined), establishing empirical methodology for predicting whether alignment becomes easier or harder at frontier scale—a foundational signal for AI safety research.

  4. RAM-Net: Linear-Time Sequence Modeling with Sparsely Addressable State

    arxiv.org

    RAM-Net proposes sparse, address-based access to recurrent state for linear-time sequence modeling, eliminating inter-token interference that degrades long-range retrieval in standard linear attention. This advances the core efficiency frontier of sequence models with demonstrated improvements over Mamba2 on both fine-grained recall and perplexity—directly relevant to practitioners building next-generation efficient architectures.

  5. Dual-Modality Multi-Stage Adversarial Safety Training: Robustifying Multimodal Web Agents Against Cross-Modal Attacks

    arxiv.org

    Researchers identify a critical cross-modal attack surface in multimodal web agents and propose DMAST, a three-stage adversarial training framework combining imitation learning, supervised fine-tuning, and self-play RL to defend against coordinated visual-textual attacks. This matters to practitioners building web automation agents and the broader field studying robustness of multi-input AI systems against adversarial corruption at deployment time.

  6. Minimal Witness Reinforcement Learning

    arxiv.org

    Researchers introduce Minimal-Witness Reinforcement Learning (MWRL), a framework that recovers entire families of minimal sufficient explanations rather than single solutions, using credit assignment derived from set-theoretic coverage principles. This addresses a fundamental gap in RL—the ability to discover diverse, parsimonious solutions to verification problems—with applications to interpretability, causal discovery, and LLM-scale policy optimization.

  7. Evidence-Bound Reasoning: Neuro-Semantic Verification of Biomedical AI in Glioblastoma Radiogenomics

    arxiv.org

    Researchers developed a neuro-semantic verification framework that converts radiomic features into machine-checkable evidence records, decoupling verifiability from predictive performance and achieving 100% accuracy on corruption benchmarks while maintaining deterministic claim verification independent of model drift. This matters because it operationalizes accountability in high-stakes biomedical AI by treating explainability as an engineered, auditable property rather than a black-box byproduct—a methodological advance with implications for regulatory compliance and clinical deployment of AI systems.

  8. Even Small Reasoners Should Quote Their Sources: Introducing the Pleias-RAG Model Family

    arxiv.org

    Researchers introduce Pleias-RAG-350m and Pleias-RAG-1B, small reasoning models (350M-1B params) purpose-built for RAG workflows with native citation grounding and multi-lingual support, outperforming comparable SLMs on HotPotQA and 2wiki benchmarks. Significant for practitioners because it demonstrates that factuality and source attribution can be architecturally baked into compact, deployable models—unlocking RAG use cases in resource-constrained environments where hallucination risk demands provenance.

  9. Storage Is Not Strategy: State-Conditioned Support Control for LLM Unlearning

    arxiv.org

    Researchers challenge the assumption that parameters most associated with target knowledge are best for unlearning, proposing Intervention Score and Dynamic Intervention Re-ranking (DIR-R) to dynamically select editable parameter groups. The work demonstrates substantial improvements over existing localization-based unlearning methods on TOFU and LACUNA benchmarks, with direct relevance to practitioners building reliable model editing and alignment systems.

  10. Two Vectors Replace In-Context Demos: Structured Task Adaptation via Embeddings

    arxiv.org

    STAVE replaces demo-heavy in-context learning with two task-specific vectors injected into embeddings, reducing parameters while maintaining or exceeding performance across six multimodal and five language models. This is relevant to practitioners scaling LMMs in production where reencoding demos at every query is a significant bottleneck, and offers a cleaner alternative to prior token-insertion methods with theoretical justification.