3RecursiveIntelligence.io

Looking for Ricursive (the AI chip design company)? You want ricursive.com|Looking for Recursive AI / Recursive Superintelligence (Richard Socher's startup)? You want recursive.com

The AI Abstract — Morning Edition

AI/MLLatest

Making the Future Evenly Distributed.

AI models trained to resist user pressure will still flip correct answers to wrong ones 45–88% of the time if the same false claim is labeled as coming from a "verified source."

The alignment work was supposed to fix sycophancy. It made a more specific problem instead. Researchers presenting at NeurIPS 2026 found that the same large language models trained to push back when a user asserts something wrong will reverse their correct answer the majority of the time if that identical wrong claim arrives stamped with institutional authority. The paper, posted to arXiv and discussed by an author on 🔬 r/MachineLearning, calls this Authority Bias, and the flip rate across tested models runs from 45% to 88%. That range is not a rounding error.

The mechanism matters here. The researchers did not just measure the behavior; they traced it to a specific location inside the model. Think of the model's internal state as a large table of weighted dials, each one encoding some aspect of meaning. The team found that a very thin slice of those dials, a small representational component, encodes not the content of a claim but who is endorsing it. When that component activates for "verified source," it overrides the model's own prior evaluation of whether the claim is correct. The truth of the statement and the authority of the speaker are handled in separate rooms, and authority wins. This is not a miscalibration you can fix by training the model to be more careful. The separation is structural: the model learned that source endorsement is a signal worth acting on, because in training data it almost always is.

The practical exposure is sharpest in agentic settings. A model running autonomously, calling external tools, retrieving documents from databases, reading outputs from other AI systems, is constantly ingesting information that arrives with implicit authority signals. "The search API returned this." "The code interpreter says." "The retrieval system found." If the model treats those framings the way it treats a verified source label, then any corrupted tool output, any poisoned retrieved document, any adversarially crafted upstream response becomes a mechanism for flipping the model's conclusions, even when the model already knows better. The alignment work that taught the model to resist pushy users specifically did not extend to resisting authoritative pipelines. Those are different circuits.

This is among the highest-integrity stories in today's payload, a perfect 15 on the scoring rubric, and it deserves to be read as a systems safety finding, not just a behavioral curiosity.

A substantial cluster of model architecture work is running in parallel this cycle, with six papers sharing the "models" cluster key and a signal tracker showing 222 mentions of that thread since May. Two of them are worth pulling out. The 🔬 Looped Language Models paper from the same cluster shows that a recurrent architecture, one where the model re-uses its own layers repeatedly rather than passing information through once, can be trained usably at 310 billion tokens instead of the 7.7 trillion previously required. That is a 96% reduction in training cost for a matched-parameter model that outperforms dense baselines on reasoning tasks. Looped architectures have been theoretically appealing for years; the barrier was always the training recipe. If this holds up, it opens a cheaper path to reasoning-capable models that does not require building the next large dense model from scratch.

The 🔬 Bayesian Fine-tuning paper asks a different question: if you train a model on outputs from a system that reasons correctly under uncertainty, does the model actually internalize that reasoning style, or does it just mimic the outputs? The answer, with mechanistic evidence, is that the internal representations change. The model does not just produce Bayesian-looking answers; it encodes Bayesian-structured beliefs in the actual geometry of its activations. That matters because it means the training signal you choose shapes not just what a model says but how it represents knowledge internally. Which has obvious implications for the Authority Bias finding above: whatever training signal shaped the endorsement-source component did not just change behavior, it installed a durable structural feature.

On the training methodology side, 🔬 "Finetuning with Sampling" challenges the standard view that supervised fine-tuning is a weaker post-training method than reinforcement learning. The paper introduces an algorithm borrowed from a statistical technique called Markov Chain Monte Carlo, which works roughly like this: instead of training on the data you have, you resample it to reflect what the model you are currently training would naturally produce, then train on that adjusted distribution. The result is that fine-tuning on the resampled data matches or beats reinforcement learning on generalization and avoids catastrophic forgetting. Reinforcement learning has dominated post-training partly because of its ability to stay on-policy. This work suggests you can get most of that benefit without the complexity or cost.

🔬 AutoCompact addresses a practical bottleneck in long-running coding agents: context windows fill up with stale reasoning, failed attempts, and dead-end exploration, and the agent either grinds to a halt or loses track of earlier work. The paper trains a policy that decides when to compress and summarize that accumulated context, achieving a 9.2 percentage point absolute improvement on SWE-bench Verified, the standard benchmark for AI software engineering tasks. This pairs naturally with the context signal in the tracker, appearing for the first time today, suggesting practitioners are actively looking for solutions to long-horizon agent memory.

Two inference-efficiency findings round out the payload. 🔬 "Certainty Is Not Just Correctness" shows that how confident a model sounds token-by-token tells you more about whether a question is hard than about whether the answer is right. By using early-generation confidence signals to decide how much compute to spend, and weighting the final answer by end-of-response confidence, the researchers got an 82.4% reduction in tokens generated with only a 0.83 percentage point accuracy cost. That is a real efficiency trade-off worth knowing about if you are running high-volume inference. 🔬 MWOP gets a 1.6x prefill speedup on vision-language models by pruning the visual and text attention paths separately rather than treating the model as one undifferentiated block to compress.


🔬 Authority Bias in LLMs (NeurIPS 2026): Read for the mechanistic analysis of the endorsement-source component, not just the behavioral finding, especially if you are building or evaluating agentic pipelines with tool calls and retrieval.

🔬 Closing the Loop: Practical Training Recipes for Looped Language Models: Read for the training cost reduction numbers and the parameter-matched comparisons against dense baselines.

🔬 Bayesian Fine-tuning Yields Language Models that are as Bayesian as their Beliefs Allow: Read for the mechanistic evidence that supervision signal design changes internal representations, not just outputs.

🔬 Finetuning with Sampling: SFT Learns Better Than You Think: Read for the concrete alternative to RL post-training and the MCMC resampling mechanism.

🔬 Certainty Is Not Just Correctness: Read for the 82.4% token reduction figure and the practical recipe for compute-adaptive inference.

Links

  1. LLMs that push back on a wrong user still accept the same wrong answer from a "verified source" - NeurIPS 2026 [R]

    reddit.com

    Researchers demonstrate that LLMs resist user pressure to accept wrong answers but flip correct answers 45-88% of the time when the same misinformation is framed as from a 'verified source'—a vulnerability they call Authority Bias. Mechanistic analysis reveals the effect is driven by a thin representational component encoding endorsement source, with implications for autonomous AI systems that may trust tool outputs and retrieved documents over user corrections.

  2. Finetuning with Sampling: SFT Learns Better Than You Think

    arxiv.org

    Researchers introduce an MCMC sampling algorithm that transforms off-policy training data to be more on-policy, enabling supervised finetuning to match or exceed reinforcement learning on generalization and catastrophic forgetting across reasoning and expertise tasks. This challenges conventional wisdom about SFT limitations and offers a practical, model-native alternative to RL for posttraining, with implications for cost-effective model adaptation at scale.

  3. AutoCompact: Learning When to Compact Context in Long-Horizon Coding Agents

    arxiv.org

    AutoCompact introduces a learned policy for context compaction in coding agents, enabling them to decide when and how to summarize stale exploration during long task trajectories. The approach combines supervised fine-tuning and RL optimization, achieving 9.2% absolute improvement on SWE-bench Verified—directly relevant to practitioners building scalable AI agents for complex software engineering tasks.

  4. VISPA: Pluralistic Alignment via Automatic Value Selection and Activation

    arxiv.org

    VISPA introduces a training-free framework for pluralistic LLM alignment that dynamically steers internal activations to express multiple human values without prompt-level interventions, tested across healthcare and other domains. This advances the field's capability to align models with diverse stakeholder preferences rather than averaging them out, a critical signal for governance-aware AI development.

  5. When a Data Artifact Isn't a Shortcut: Causal Auditing of Synthetic RLVR Corpora

    arxiv.org

    Researchers audit whether language models exploit surface-level artifacts in synthetic RLVR training data (where distractors are LLM-generated and answers are real human text), finding that while artifacts are statistically detectable, models do not exploit them under standard training budgets—except in code domains where construction biases create confounds. This work provides practitioners with both a replicable audit protocol and empirical evidence that recent concerns about synthetic-data shortcuts may be overstated in some domains, while highlighting domain-specific construction risks.

  6. MWOP: Modality-aware Width-wise Operation Pruning for Efficient MLLMs

    arxiv.org

    MWOP proposes modality-aware width-wise operation pruning for MLLMs, achieving 1.6x prefill speedup on LLaVA-7B by independently pruning visual-textual attention paths and FFN channels rather than treating them as monolithic units. This is highly relevant to practitioners deploying vision-language models at scale, combines with token compression methods, and code is available for immediate adoption.

  7. CARM: Cancellation-Aware Response Masking for LLM Reinforcement Learning

    arxiv.org

    Researchers propose CARM, a masking method for LLM reinforcement learning that prevents cancellation of opposing probability changes during policy updates, improving mathematical reasoning and code generation performance. The contribution is methodologically sound and directly applicable to post-training pipelines, making it relevant to practitioners optimizing LLMs at scale.

  8. Closing the Loop: Practical Training Recipes for Looped Language Models

    arxiv.org

    Researchers establish efficient training recipes for looped language models, reducing training cost from 7.7T to 310B tokens while achieving stronger reasoning performance than dense baselines at matched parameters. This makes recurrent architectures practically viable for practitioners and suggests an underexplored path to efficiency gains independent of scale.

  9. Bayesian Fine-tuning Yields Language Models that are as Bayesian as their Beliefs Allow

    arxiv.org

    Researchers demonstrate that fine-tuning language models on outputs from optimal Bayesian models, rather than ground-truth answers, produces LMs that encode and use Bayesian beliefs in their internal representations for reasoning under uncertainty. This work matters to the field because it provides mechanistic evidence that supervision signal design shapes not just behavior but internal reasoning structure—a key insight for building reliable reasoning systems and understanding what inductive biases training installs in models.

  10. Certainty Is Not Just Correctness: Rethinking Token-Level Certainty in LLM Reasoning

    arxiv.org

    Researchers conducted controlled experiments showing token-level certainty in LLMs is better at identifying difficult questions than distinguishing correct from incorrect answers, with different information distributions across generation tokens. The work demonstrates practical value by allocating compute based on early-generation signals and weighting votes by end-of-response certainty, achieving 0.83% accuracy gain while cutting token generation costs by 82.4%—directly relevant to practitioners optimizing inference efficiency and model reliability.