3RecursiveIntelligence.io

Looking for Ricursive (the AI chip design company)? You want ricursive.com|Looking for Recursive AI / Recursive Superintelligence (Richard Socher's startup)? You want recursive.com

The AI Abstract — Morning Edition

AI/MLLatest

Making the Future Evenly Distributed.

An open-weight computer-use model just matched frontier-class performance at one-seventh the cost per task, while a separate fix cut a speech transcription model's hallucination rate on silence from 62% to 2%.

A Paris-based startup just made it significantly cheaper to deploy an AI that can actually use a computer. H Company's 📰Holo4 is a family of open-weight models, 27B and 35B parameters, that can click through a desktop interface, write and run code, call APIs, and navigate Android apps, all within a single unified architecture. The 35B variant ships under an Apache 2.0 license, meaning you can run it on your own hardware without paying per call. On OSWorld 2.0, the standard benchmark for computer-use agents, the 35B model scores 61.7% at $1.22 per task. Claude Opus 5.5 costs $8.48 for the same task. That's not a marginal difference in pricing; it's a 7x gap, and the open model is competitive on the benchmark. The cost-capability frontier for agents that can autonomously operate software just moved in a meaningful direction for anyone not building at hyperscale. The "models" cluster has seen 215 stories since May, but the open-weight price-performance story here is what separates Holo4 from the general release noise.

Whisper, OpenAI's widely deployed speech transcription model, has a known failure mode that's worse than most users realize. When you feed it silence or near-silence, it generates text anyway. Not occasionally: in controlled testing, it hallucinated on 61.9% of silent clips. That means a transcription pipeline running on any audio with gaps is quietly fabricating content at a rate you'd never accept if you saw the raw number. A new 🔬paper from arXiv cuts that rate to 2.4%. The method works by doing the opposite of what you might expect. Instead of teaching the model to recognize silence and stay quiet, the researchers collected 40,891 phrases that Whisper commonly hallucinates across 100 languages, then synthesized those phrases as training examples, explicitly teaching the model to discriminate between real speech that sounds like those phrases and the nothing that was actually there. Think of it like showing a proofreader every common forgery so they stop mistaking a blank page for a document. The fix also recovers 82.7% of genuine speech that naive silence-suppression methods would have silenced. The benchmark, the phrase lexicon, and the synthetic corpus are all released on Hugging Face, so this is a deployable fix today, not a research direction.

Google Research released 🔬RRSI, a framework that lets an LLM agent rewrite its own prompts, tools, and decision logic without the self-improvement process overfitting to the training tasks it practiced on. The problem it solves is subtle but real: when an agent is allowed to optimize its own configuration, it tends to get better at the specific tasks it sees during tuning while getting worse on anything new, the same problem a student has when they memorize past exams instead of learning the material. RRSI borrows from classical regularization, the same family of techniques used to prevent neural networks from memorizing training data. It applies edit budgets that cap how many changes the agent can make, pruning rules that strip unnecessary complexity, and cost penalties that push toward simpler solutions. The result is +3.5 to +4.9 point gains on held-out benchmarks with a 30% reduction in inference cost. The code ships Apache 2.0. This is the "agents" cluster's second day of coverage; it's worth watching whether self-improving agent frameworks start appearing as a sustained research category rather than isolated releases.

Two new attention mechanisms, 🔬CoWindow and MassAlloc, deliver 2.2x to 8.6x speedups on the attention computation step for long-context models, cutting training compute by 23 to 28.5% at 14 billion parameters. Attention is the part of a transformer that figures out which words in a long document are relevant to each other. For short text it's fine, but it scales quadratically with length: double the context, quadruple the work. CoWindow handles this by splitting context into overlapping sparse windows so every position is covered by at least one window without any single window doing all the work. MassAlloc goes further by watching the attention weights as they form, then routing more compute toward positions that are generating high-confidence signal and less toward positions that aren't. The author is transparent that neither method is lossless relative to full dense attention. That caveat matters for use cases where precision on rare long-range dependencies is critical, but for most long-context training and inference workloads the tradeoff is favorable.

Standard fine-tuning with human preference data has a quiet defect. The method called DPO, which trains a model to prefer one response over another by comparing pairs, gives outsized weight to words like "the," "a," and "is." Those tokens appear at roughly the same rate in both the preferred and dispreferred response, so they contribute equal signal in both directions and dilute the gradient that's supposed to capture what actually made one response better. A new 🔬paper proposes masking those high-frequency tokens out of the loss calculation entirely. The fix requires no additional parameters and shows consistent gains across AlpacaEval, MT-Bench, and Arena-Hard. If your organization fine-tunes on human preference data, this is the kind of zero-cost improvement worth testing before your next training run.

Two more research results worth flagging briefly. The 🔬TaH2 architecture applies extra computation passes only to the tokens in a sequence that benefit from them, rather than uniformly iterating over everything. On AIME math benchmarks it improves test-time scaling efficiency by 53%, which matters for any deployment where you want to spend more compute on harder problems without paying that cost on easy ones. Separately, the 🔬Ongiini-Eval-OW benchmark is being built to evaluate machine translation for Oshindonga and Oshikwanyama, two related languages spoken by roughly one million people in Namibia and Angola, for which current commercial and open-source systems have essentially no coverage. It's a concept paper announcing a 600-item benchmark, not a finished tool, but the "language" signal has appeared in 634 stories since February, and low-resource African language evaluation infrastructure is clearly an active area.


🔬 How to Reduce Whisper Hallucination: Read for the concrete deployment fix and the released lexicon, especially if you run transcription pipelines on real-world audio with silence gaps.

📰 Holo4: Open-Weight Computer-Use Models: Read for the benchmark comparisons and cost-per-task numbers, which are the real story beneath the model release.

🔬 RRSI: AI Agents That Improve Their Own Harness Without Overfitting: Read for the regularization mechanism, which explains why most prior self-improving agent attempts failed on transfer.

🔬 Frequency-Hard DPO: Read for the mechanism section, which makes the gradient dilution problem intuitive and shows why the fix works without changing model architecture.

🔬 CoWindow and MassAlloc Attention: Read for the author's honest limitations section, which is a useful template for evaluating attention efficiency claims more generally.

Links

  1. CoWindow and MassAlloc Attention: collective causal coverage and distribution-adaptive compute [R]

    reddit.com

    Two novel attention mechanisms—CoWindow (complementary sparse windows with collective causal coverage) and MassAlloc (adaptive compute allocation via softmax statistics)—achieve 2.2x-8.6x attention operator speedups on long-context training/inference while reducing training FLOPs by 23-28.5% at 14B scale. Both represent genuine advances in attention efficiency for practitioners building long-context models, with the author transparently acknowledging that neither achieves lossless equivalence to dense attention.

  2. How to Reduce Whisper Hallucination

    arxiv.org

    Researchers address Whisper's well-known hallucination problem—the model generates text on 61.9% of silent audio clips—by collecting 40,891 hallucination phrases across 100 languages, synthesizing them as training positives to teach discrimination rather than suppression. The approach reduces hallucination on silence from 61.9% to 2.4% while recovering 82.7% of genuine speech, with benchmark, lexicon, and synthetic corpus released on Hugging Face for reproducible deployment improvements.

  3. Google Research Open-Sources RRSI: AI Agents That Improve Their Own Harness Without Overfitting

    marktechpost.com

    Google Research released RRSI, a regularized framework for LLM agents to autonomously improve their own prompts, tools, and control flow without overfitting. The method applies classical regularization principles (L0 edit budgets, L1 pruning, L2 cost rules) to the search process itself, achieving measurable out-of-distribution gains (+3.5 to +4.9 points) across JobBench, GDPval, and APEX-Agents while reducing inference cost by 30%, representing a practical advance in agentic self-improvement that practitioners can deploy today.

  4. H Company Releases Holo4: Open-Weight Computer-Use Models That Click, Code and Call Tools Across Desktop, Web, Android and APIs

    marktechpost.com

    H Company released Holo4, a family of open-weight computer-use vision-language models (27B and 35B-MoE) that unify GUI interaction, code execution, and tool calling across desktop, web, Android, and APIs on a single architecture. The 35B variant ships Apache 2.0 weights for self-hosting, with 61.7% OSWorld 2.0 performance at $1.22/task—dramatically cheaper than frontier models (Claude Opus 5.5 at $8.48)—making agentic automation accessible to practitioners and shifting the cost-capability frontier for open deployment.

  5. Masking Frequent Tokens Sharpens Direct Preference Optimization

    arxiv.org

    Researchers identify that high-frequency tokens dominate DPO's sequence-level reward signal symmetrically across preferred and dispreferred responses, diluting preference gradients. They propose Frequency-Hard DPO, which masks these tokens to sharpen alignment, showing consistent improvements on AlpacaEval, MT-Bench, and Arena-Hard without additional parameters—a technique immediately applicable to any organization fine-tuning with DPO.

  6. Improving Test-Time Scaling with Adaptive Looped Transformers

    arxiv.org

    Researchers introduce TaH2, an adaptive looped transformer that selectively applies extra computation iterations only to tokens that benefit from them, improving test-time scaling efficiency by 53% on AIME benchmarks. This addresses a gap in parameter-efficient architectures by demonstrating that adaptive depth allocation yields steeper accuracy-compute slopes than both fixed looping and non-looped baselines, with implications for deployment under variable compute budgets.

  7. Rethinking Token Reweighting for SFT: Suppress, Reverse, and Extrapolate Learned Features

    arxiv.org

    Researchers propose SCALE, an entropy-guided method that reweights tokens during supervised fine-tuning by freezing pretrained weights and learned deltas, then learning gated controls that can suppress, reverse, or extrapolate features. Results show consistent gains over baselines on mathematical reasoning and code generation across Qwen2.5 and Qwen3 models, advancing understanding of how to correct harmful SFT behaviors without catastrophic forgetting.

  8. Functional Gradient Descent with Adaptive Representations [R]

    reddit.com

    Researchers propose adaptive representations framework for functional gradient descent that provably converges to global minimizers while remaining implementable, achieving order-of-magnitude improvements over neural networks in tested settings. Addresses fundamental approximation problem in functional GD—a core algorithmic contribution relevant to practitioners exploring alternatives to standard deep learning optimization.

  9. The Ongiini-Eval-OW Benchmark: A Concept Paper for the Planned Benchmarking of Machine Translation and Large Language Models on Oshindonga and Oshikwanyama

    arxiv.org

    Researchers are designing Ongiini-Eval-OW, a 600-item machine translation benchmark for Oshindonga and Oshikwanyama (Oshiwambo cluster languages spoken by ~1M people), addressing a complete coverage gap in commercial and open-source MT systems. This matters to the field as it signals growing attention to non-Latin-script, low-resource language evaluation infrastructure and democratizes benchmarking for underserved linguistic communities.

  10. Beyond Solo and Consistency: Vindicating Multi-Agent Debate via Conditional Progressive Pruning

    arxiv.org

    Researchers propose Conditional Progressive Pruning (CPP), a pruning framework that enables multi-agent debate (MAD) to outperform single-agent and consistency-based baselines for the first time under strict computational budgets. This resolves a key empirical challenge in the MAD field and advances test-time scaling techniques that improve LLM reasoning on complex tasks.