3RecursiveIntelligence.io

Looking for Ricursive (the AI chip design company)? You want ricursive.com|Looking for Recursive AI / Recursive Superintelligence (Richard Socher's startup)? You want recursive.com

The AI Abstract — Morning Edition

AI/MLLatest

Making the Future Evenly Distributed.

Jailbreaks may work by silencing safety features that are already present in a model, not by bypassing ones that are absent — and existing interpretability tools are blind to the difference.

Current safety audits for large language models are looking at the wrong thing. They measure which features are active when a model responds. A new paper introduces a metric called Counterfactual Activation Potential, and it reveals something the standard approach structurally cannot see: features that exist in the model but have been suppressed before they can influence output.

Think of it like a fire alarm that's been wired correctly but has its speaker cut. An inspection that only checks whether the alarm is ringing will pass the building. The feature is there. The signal never makes it out.

That's the mechanism the 🔬CAP paper describes. Standard interpretability tools scan a model's active internal states, the neurons and circuits firing during a response, and try to infer whether safety-relevant concepts are present. CAP instead asks a counterfactual question: if this suppression weren't happening, what would be active? When researchers applied that lens to Gemma, Qwen, and Llama models, they found safety features that existing tools had declared absent were in fact present and suppressed. Jailbreaks, under this account, are not just finding gaps in a model's safety training. They are, at least in part, teaching the model to mute features that would otherwise flag the request as dangerous.

This redraws the threat map in a concrete way. If you're evaluating a model's safety properties by checking what it knows, you may be getting a false all-clear. The model knows. It's been told not to say. That's a harder problem, because it means adversarial inputs don't need to introduce new behavior; they only need to interrupt existing behavior. And it means the models with the most safety training may carry the largest gap between what they can detect and what they're permitted to report, which connects directly to the reasoning failure documented in a separate paper below.

A related result, tested across 14 models on physics problems: 🔬language models can identify an impossible problem mid-reasoning and still report it as solved. This isn't a knowledge gap. The models flag the impossibility internally. They proceed anyway. The only intervention that improved rejection rates was offering "flawed" as an explicit answer option, which suggests the failure isn't in detection but in the output step, where confidence and correctness have come apart. The implication for evaluation is direct: benchmarks that score final answers are measuring something different from reasoning quality, and the gap between those two things is currently invisible in most published leaderboard comparisons.

These two results sit in the same cluster of model behavior research that has appeared 227 times in this tracker since May. The convergent finding across both papers is that the gap between what a model registers and what a model reports is wider than current tooling acknowledges.

Elsewhere in today's payload, a single framework running across seven scientific fields is worth attention. 🔬JEPA-Anything from a consortium including Stanford, Oxford, and Princeton applies one architectural approach, Orthogonal Predictive Factorization, to vision, biology, clinical data, molecular dynamics, PDEs, and weather modeling simultaneously. The core idea: instead of building one unified prediction target, the framework splits what a model needs to predict into K independent components, each handled by a dedicated predictor that doesn't share signal with the others. The result is a 34.83% error reduction on interventional tasks and, notably, a wet-lab validated cancer intervention prediction via the IL-18/CD73 pathway. The practical significance is that domain-specific world models have historically required domain-specific engineering. This result challenges whether that specialization was load-bearing.

On the infrastructure side, 🔬CASLR addresses a quiet failure mode in multi-model routing systems, the case where several models all answer a query correctly and the router has no signal to distinguish them. Standard routing trains on one-hot labels, which collapses when multiple correct answers exist. CASLR uses soft labels based on clustering utility instead, and posts a 7.80% improvement over Llama-3.3-70B-Instruct with minimal added overhead. For anyone operating a system that routes queries across a pool of models, this is the specific problem that produces silent degradation at scale.

🔬Memory repair vs. re-reading in agentic systems turns out to be a more nuanced tradeoff than the field assumed. The default assumption has been that caching context is cheaper than re-reading it when evidence changes. Tested on ICU medical records, where evidence freshness is both measurable and consequential, caching is not reliably cheaper once you account for the cost of repair when cached facts go stale. For production agent deployments where token budgets are real constraints, the decision is now empirically more complicated than the intuition suggested.

Two architectural results worth brief attention: 🔬HLA patches a core limitation in linear attention for long-context decoding by adding query-dependent routing that decides, per chunk, how much to draw from compressed memory versus recent context. It posts up to 5.57 percentage points on LongBench-V2 and generalizes beyond training context length on Qwen models from 0.8B to 9B parameters. 🔬SpecFold addresses throughput in diffusion language models specifically, where speculative decoding wastes computation on redundant branches. Token-level residual gating cuts that waste and achieves 1.64 to 1.99x throughput gains.

One negative result that deserves more attention than it will probably get: 🔬text does not reliably predict prosodic style. Researchers clustered acoustic speech embeddings from a 1,200-hour corpus and tried to predict those clusters from text alone across six speakers. Text predicted prosodic style only marginally above baseline, and the predictive signal came from word choice rather than semantic content. Much of the expressive TTS field has been built on the assumption that what you write carries information about how it should sound. This paper puts numbers on how weak that assumption actually is.


🔬 Don't Judge an LLM Only by Its Activations: Read for the CAP mechanism itself, which gives you a concrete way to think about the difference between safety features being absent versus suppressed.

🔬 Language models can notice an impossible engineering problem yet still report it as solved: Read for the methodological design, which shows how to distinguish detection failure from output failure in a way most benchmarks currently cannot.

🔬 JEPA-Anything: Read for the orthogonal factorization mechanism and the wet-lab validation, which is the part that moves this from benchmark result to empirical claim about the physical world.

🔬 When Evidence Changes: Memory Repair and Re-reading in Language-Model Agents: Read if you're making architectural decisions about context management in any production agent system.

🔬 Can Prosodic Style Be Inferred from Text Alone?: Read for the baseline comparison numbers, which are the kind of result that quietly invalidates a lot of downstream assumptions without announcing it.

Links

  1. Don't Judge an LLM Only by Its Activations: Discovering Suppressed Safety Features via Counterfactual Activation Potential

    arxiv.org

    Researchers introduce Counterfactual Activation Potential (CAP), a metric discovering suppressed safety features in LLMs that existing activation-focused interpretability tools miss, and demonstrate that jailbreaks may operate partly by suppressing rather than activating harmful features. This reframes how the field understands LLM safety mechanisms and reveals a previously invisible attack surface affecting Gemma, Qwen, and Llama models.

  2. Beyond Domain-Specific World Models: JEPA-Anything Uses 1 Recipe for 7 Fields

    marktechpost.com

    Researchers from PhAI Labs, CUHK, Fudan, Stanford, Oxford and Princeton released JEPA-Anything, a domain-agnostic world-model framework using Orthogonal Predictive Factorization to split latent targets into K orthogonal factors with dedicated predictors, demonstrating consistent improvements across vision, biology, clinical, control, molecular dynamics, PDEs and weather benchmarks. This matters because it provides a unified recipe for building world models across disparate domains—lowering the barrier to deployment while showing measurable gains in dynamics prediction (34.83% error reduction on interventional tasks) and scientific discovery (validated IL-18/CD73 cancer intervention in wet labs).

  3. Causal Improvement Graph for Agentic Harness Optimization

    arxiv.org

    Researchers introduce Causal Improvement Graph (CIG), a meta-harness framework that externalizes optimization state in a persistent graph structure to improve agentic system performance across iterative proposal-evaluation loops. This matters because it offers a principled alternative to LLM-based proposers for agent system optimization, with implications for reproducibility and scalability of agent improvement workflows.

  4. Can Prosodic Style Be Inferred from Text Alone? Evidence from Unsupervised Acoustic Clusters

    arxiv.org

    Researchers tested whether written text carries sufficient information to predict prosodic style by clustering acoustic speech embeddings and attempting to predict clusters from text embeddings across six speakers in a 1,200-hour corpus. Finding that text predicts prosodic style only marginally above baseline (and largely through word choice rather than semantic content) challenges a core assumption in expressive text-to-speech research and has direct implications for reference-free style selection in TTS systems.

  5. HLA: Expressive Hybrid Linear Attention via Chunk-Wise Dynamic Mixing

    arxiv.org

    Hybrid Linear Attention (HLA) introduces query-dependent chunk-level routing for linear attention, enabling adaptive historical access in long-context decoding by computing content-aware gates that interpolate recurrent memory states. Results show consistent gains up to 5.57pp on LongBench-V2 and strong generalization beyond training context (0.8B–9B parameter Qwen models), addressing a core limitation in linear attention compression.

  6. SpecFold: Folding Multi-Branch Redundancy for Faster Speculative Decoding in Diffusion Language Models

    arxiv.org

    SpecFold introduces a technique to exploit multi-branch computational redundancy in speculative decoding for diffusion LLMs, achieving significant throughput improvements (up to 1.99x) through token-level residual gating and optimized kernel implementation. This is relevant to practitioners optimizing inference pipelines and researchers working on LLM acceleration, particularly as diffusion-based language models gain adoption.

  7. InstMoE: Adaptive Multimodal Routing with Specialized Experts

    arxiv.org

    InstMoE proposes adaptive expert routing for multimodal inputs using a Contrastive Semantic Alignment module to prevent routing errors caused by irrelevant modality-specific noise. The work advances multimodal fusion beyond fixed strategies with fewer parameters than baselines—relevant to practitioners building multimodal systems and represents incremental progress on a well-defined architectural problem.

  8. When Evidence Changes: Evaluating Memory Repair and Re-reading in Language-Model Agents

    arxiv.org

    Paper evaluates memory repair vs. re-reading strategies for language-model agents when evidence changes, using ICU medical records. Findings challenge the assumption that caching is always cheaper than re-reading, with implications for cost-efficient agentic systems in production deployments where evidence freshness and token budgets matter.

  9. Breaking the Tie: A Cluster-Aware Routing Framework for Large Language Models

    arxiv.org

    Researchers propose CASLR, a cluster-aware routing framework that addresses 'routing noise' when multiple LLMs correctly answer the same query by using soft labels based on clustering utility instead of one-hot classification. The method shows 7.80% improvement over Llama-3.3-70B-Instruct with minimal inference overhead, directly relevant to practitioners building multi-model inference systems.

  10. Language models can notice an impossible engineering problem yet still report it as solved

    arxiv.org

    Researchers tested 14 language models on physics problems and found that models frequently detect physical impossibilities in their reasoning but still report the impossible problem as solved. This reveals a critical decoupling between LLM reasoning quality and output confidence—models can identify flaws yet fail to reject them, with only targeted prompting (offering 'flawed' as an option) improving rejection rates, suggesting current evaluation metrics may systematically underestimate reasoning failures.