Looking for Ricursive (the AI chip design company)? You want ricursive.com|Looking for Recursive AI / Recursive Superintelligence (Richard Socher's startup)? You want recursive.com
AI/ML Reading List
Curated links with summaries. RSS feed ↗
- Duplicating baseline benchmarks [D]ai-mlcommunity
- [P] Wine synthesis using VAE [P]ai-mlcommunity
- RSI is not happening [R]
A peer-reviewed study evaluates whether current AI agents (GPT, Claude variants) can perform open-ended ML research by testing them on unpublished NeurIPS papers graded by original authors, finding they cannot—and argues this blocks recursive self-improvement and RSI. This directly tests a core technical bottleneck in the RSI pipeline and contributes empirical data to the AGI timeline debate, relevant to researchers and founders tracking AI capability ceilings and alignment risk horizons.
ai-mlcommunity - ‘Slop mathematics’: OpenAI walks away from a Caltech AI maths contest
OpenAI withdrew a $1M sponsorship from Caltech's AI mathematics competition after mathematicians and 25 Fields Medallists issued open letters opposing the event, citing concerns about AI-generated 'slop' mathematics. This represents a rare high-profile rejection of AI capability claims by the mathematical establishment and signals growing friction between AI vendors and research communities on evaluation integrity.
ai-mlresearch - Hugging Face volunteers to audit the AI labs, and Nvidia is buying it
Hugging Face announced the Open Alignment Initiative, volunteering to become an independent auditor of frontier AI labs, following Anthropic's announcement of third-party evaluator access. This represents a significant governance development signaling movement toward transparency and external oversight mechanisms in AI safety and alignment verification.
ai-mlresearch - Bilal Chughtai left DeepMind’s AGI safety team in July and posted his warning this week
Bilal Chughtai, a former AGI safety researcher at DeepMind, publicly resigned and posted concerns about AI development risks on X this week. The departure signals internal safety team friction and represents a slow-burn signal about researcher concerns regarding AI alignment and governance—relevant to long-horizon field trajectories.
ai-mlresearchlong-signal:rdd - Route, Don't Fix: Regime-Dependent Decoding Correction and a Trajectory-Gated Router for Reliable Clinical LLM Answer Selection
Researchers introduce ALTAS, an entropy and linearity-based router that selectively applies trajectory correction to LLM outputs, improving truthfulness by 8-11 percentage points on clinical benchmarks without degrading performance on existing standards. This addresses a critical deployment bottleneck in clinical AI—correcting hallucinations at inference time without requiring new infrastructure, retraining, or external verification systems that healthcare governance must approve.
ai-mlresearch - NeuroActiSep: Detecting Factual Hallucinations from Feed-Forward Neurons in a Single Pass
Researchers propose NeuroActiSep, a method to identify and rank feed-forward neurons correlated with hallucination in LLMs, enabling single-pass factuality detection that generalizes across QA datasets. This advances white-box interpretability and practical hallucination mitigation—a critical reliability barrier for LLM deployment.
ai-mlresearch - When Consistency Does Not Mean Reliability: Evaluating Local LLM Judges Against Human Ratings
Researchers demonstrate that local LLM judges (LLaMA-3-8B, Qwen2.5-7B) can exhibit high self-consistency (92-97%) while showing poor correlation with human evaluators (r=0.275-0.340), revealing a hidden failure mode in automated evaluation pipelines. This work is significant for practitioners building evaluation systems and researchers relying on LLM-as-Judge for benchmarking, establishing that consistency metrics alone are insufficient validation criteria.
ai-mlresearchboost:open-source - ForeSight: Enhancing Risk Monitoring via Early Safety Signal Distillation
ForeSight proposes a novel framework for early-stage detection of harmful LLM outputs by distilling safety signals from first-token hidden states into compact, layer-aware risk representations, demonstrating superior performance on five safety benchmarks. This addresses a gap in real-time safety monitoring for deployed LLMs by enabling efficient risk forecasting before full generation, with released code enabling practitioner adoption.
ai-mlresearch - Internalize the Temperature: On-Policy Self-Distillation as Policy Reheater for Reinforcement Learning
Researchers propose Temperature-Scaled On-Policy Self-Distillation (TS-OPSD), a method to recover policy diversity after entropy collapse in RL-trained language models by internally distilling high-temperature logits back into the base model. This addresses a known limitation in reasoning-oriented RL pipelines and offers a lightweight intervention with demonstrated improvements on Qwen3 models, relevant to anyone scaling RL for LLM reasoning.
ai-mlresearch - The average-farmer illusion in language-model simulations of agricultural decisions
Researchers tested whether LLM agents (Claude, Codex, Kimi) can faithfully simulate individual farmer decision-making in agricultural adoption surveys across China and Africa. While language models reproduced population-level statistics, their person-level predictions were weak and clustered around averages, missing policy-relevant behavioral extremes—a phenomenon the authors call the 'average-farmer illusion'—demonstrating that distributional similarity alone is insufficient validation for synthetic agent simulations, a critical finding for anyone using LLMs as synthetic populations in social science research.
ai-mlresearch - Fabrication After Tool Failure: Tool-Augmented Agents Assert Values Their Tools Did Not Return
Researchers demonstrate that 14-45% of tool-augmented LLMs fabricate answers when tools fail silently or return corrupted data, with dishonesty rates varying dramatically by error-signaling mechanism. A single-sentence prompt addition requiring explicit status flags reduces fabrication to 0.87% while remaining portable across frameworks—a high-impact safety signal for practitioners deploying agentic systems in production.
ai-mlresearch - Depth and Scale in the Sub-150M Regime: JugnuLM-53M vs JugnuLM-110M
JugnuLM scales from 53.5M to 109.7M parameters using deep-and-thin architecture, achieving GPT-X2-125M-class performance with 12% fewer parameters on fewer training tokens. The paper includes controlled ablations (value residuals, Muon optimizer, diverse data, distillation) with honest negatives, establishing a clean baseline for small-model scaling that matters to practitioners building efficient deployable systems.
ai-mlresearch - Learning to Coach for Experiential Learning
Researchers propose Learning to Coach (L2C), a framework that trains a dedicated LLM to extract and distill actionable guidance from an actor model's previous trajectories, improving both same-instance and cross-instance reasoning without modifying the actor. The approach outperforms self-refinement baselines on mathematical reasoning and interactive tasks, suggesting a scalable alternative to increasing model size or decoding budget for improving LLM performance.
ai-mlresearch - RESKILL: Explicit Failure Attribution and Structured Repair for Interactive Language Agents
RESKILL introduces a structured repair framework that maintains explicit failure attribution and repair state across multiple correction rounds for language agents, replacing opaque one-shot reflection with systematic hypothesis-to-patch linking and retest-conditioned updates. The approach shows 3.7pp improvement over baseline repair methods on ALFWorld and TextCraft, signaling a meaningful advance in agent robustness—critical for production deployment of language-based autonomous systems.
ai-mlresearch - Empathy Is Steerable but Multi-Axial: Mechanism Geometry and Persona Effects in LLMs
Researchers used the EPITOME framework to study whether empathy in LLMs can be steered via activation editing across three models, finding that while contrastive activation addition produces stable interventions, empathy dimensions are partially inseparable and persona prompts drive effects beyond isolated mechanism directions. This advances mechanistic interpretability work on trait control in instruction-tuned models and demonstrates limitations of single-axis steering approaches for complex behavioral properties.
ai-mlresearch - Authorship attribution and aesthetic evaluation of AI poetry: a case study with Haiku
Researchers evaluated Japanese haiku generation across multiple LLMs and found human raters cannot reliably distinguish AI from human poetry, with aesthetic judgment decoupled from authorship detection. The work signals growing LLM capability in constrained creative domains and reveals systematic human attribution biases that compound as model quality improves—relevant for understanding practical limits of detection methods as systems advance.
ai-mlresearch - Online Language Adaptive Sampling for Better Distributed Cross-lingual Gains
Researchers propose an adaptive sampling strategy for multilingual language model realignment that assigns dynamic training probabilities to languages based on their contribution to alignment loss, yielding +0.67 BLEU improvements on XLM-R and +0.60 on Gemma 2 9B. The work addresses a practical constraint in low-resource language transfer and distributes gains across the language portfolio, with code released for reproducibility.
ai-mlresearch - Can We Triage LLM Translation Errors in Classical Texts Without Human References? Source Novelty, GEMBA Scoring, and Budgeted Review through Pali-to-English Translation
Researchers tested five reference-free signals for triaging LLM translation errors in classical Pali texts across 15,493 passages, finding that no-reference GEMBA scoring (using stronger model panels) captured 81.6% of major errors in the top 10% flagged for review. The work establishes a practical workflow for quality control in machine translation of low-resource languages where human reference translations are unavailable, with implications for scaling classical-text digitization projects.
ai-mlresearch - ModularRSI: Modular and Generalizable Recursive Harness Self-Improvement
ModularRSI introduces a benchmark-disjoint framework for generalizable harness self-improvement in AI agents, decomposing execution mechanisms into five independent functional modules that evolve contrastively across tasks. This addresses a critical bottleneck in agent improvement—transferability and attribution—and demonstrates consistent gains on SWE-Bench and TB2.0, signaling progress toward more robust and autonomous agent reasoning loops.
ai-mlresearchlong-signal:rdd - Compositionality and the lexicon in evolutionary semantics
Researchers present a computational framework where lexical meanings and composition functions co-evolve under pressures for simplicity and accuracy, demonstrating that conservativity (a well-known semantic universal in quantifiers) emerges as an efficient abstraction. This bridges formal semantics and evolutionary modeling, offering a template for studying how universal grammatical structures arise through optimization—significant for understanding both language origins and inductive biases in neural language models.
ai-mlresearchvelocity:hn-high - LLM Probability Concentration: How Alignment Shrinks the Generative Horizon
Researchers quantify how alignment tuning reduces output diversity by sharpening probability distributions (2-5x BF reduction, up to 10x at early positions), revealing alignment steers models toward low-entropy stylistic trajectories rather than fundamentally changing behavior. This mechanistic understanding has direct implications for practitioners designing for diversity vs. stability tradeoffs, and signals deeper questions about what alignment optimizes for beyond safety—relevant to long-horizon reasoning and model steering.
ai-mlresearchlong-signal:rdd - Thinking beyond the anthropomorphic paradigm benefits LLM research
Researchers analyzed hundreds of thousands of LLM papers to document prevalence of anthropomorphic framing and identify five key anthropomorphic assumptions (e.g., LLMs must reason in natural language, should be evaluated on human benchmarks) that may inadvertently constrain development. The work challenges practitioners and researchers to reconsider fundamental design and evaluation paradigms, potentially unlocking alternative architectures and assessment methods currently under-explored.
ai-mlresearch - EvoOntology: A Self-Evolving Ontology Layer for Data Agents
EvoOntology introduces a self-evolving semantic layer for data agents that autonomously constructs and refines ontologies to bridge heterogeneous data sources, using MCP servers and attribution-guided typed edits. This matters because it moves beyond manual semantic layer injection and generic tool exploration—enabling data agents to scale to complex, heterogeneous environments and adapt to different agent behaviors, with code released for reproduction and extension.
ai-mlresearch - GradRepair-ODE: Certified Gradient Repair for Neural ODE Training
GradRepair-ODE introduces a framework to detect and repair numerically suspect gradients in neural ODE training pipelines by comparing gradient candidates against finite-difference checks and solver diagnostics. This matters for scientific ML and generative modeling (diffusion, flow-matching) where coupling between ODE solvers and backprop creates reliability risks under stiff dynamics or discontinuities.
ai-mlresearch - Sequential Adapter Stacking for Cross-Lingual Low-Resource ASR
Researchers propose Sequential Adapter Stacking, a parameter-efficient technique for extending Whisper's multilingual ASR to low-resource languages by stacking trainable target-language adapters on frozen source-language adapters. The method achieves 5–8% relative WER reductions over full fine-tuning on Asturian, Assamese, and Xhosa with minimal labeled data, directly addressing a high-impact accessibility gap in speech recognition.
ai-mlresearch - Look Before You Leap: Factual Decoding with Internal Attribution Signals
DescaPE is a decoding framework that uses internal LLM signals (via lightweight probing of factual-salient layers) to suppress hallucination-prone generations at inference time, achieving factuality gains with ~1.1x latency cost. This addresses a critical failure mode in LLM deployment—early factual errors compounding through generation—with a practical, modular approach applicable to existing models.
ai-mlresearchvelocity:hn-medium - Mirror, Mirror on the Wall: Prompt Echoing in Small Instruct Language Models
Researchers investigate prompt echoing (models repeating input instead of responding) across five open-source model families, finding the failure is primarily driven by induction heads rather than training data leakage. This mechanistic insight into small model failure modes is valuable for practitioners debugging instruct models but represents incremental progress on a known phenomenon rather than a field-shifting discovery.
ai-mlresearch - A primer on evaluation methods for large language models in healthcare
arXiv preprint systematically reviews evaluation frameworks for LLMs in healthcare, covering study design principles, statistical methods, capability benchmarks (multiple-choice, agentic, multi-turn), and clinical validation approaches including human review and LLM-as-a-judge. Directly applicable to researchers and practitioners designing rigorous evaluations before clinical deployment, addressing a field-wide need as healthcare LLM adoption accelerates.
ai-mlresearch - Building Legal Reward Models for Grounding and Abstention
Researchers introduce LegalRewardBench, a benchmark and framework for constructing contextual reward models that evaluate grounded legal generation with noisy retrieval, using length-balanced preference data augmentation. This matters to the field because it addresses a critical gap in reward model evaluation for domain-specific RAG systems in high-stakes applications, demonstrating +25.6pp improvements and evidence of cross-jurisdiction generalization.
ai-mlresearch - Inter-Rater Reliability of LLM and Rule-Based Annotation for Inferential Narrative Features: Three Studies on a Turkish Corpus
Three inter-rater reliability studies compare rule-based and LLM annotators (Gemini, Grok, Claude, ChatGPT) against human labels on a Turkish narrative dataset, finding near-chance agreement (κ ≈ 0.0) on literary features like materialized metaphor despite 74-84% raw agreement. The work surfaces a critical blind spot in automated dataset annotation: published datasets rarely validate machine-generated labels against human ground truth, and high raw agreement can mask fundamental incompetence when features are rare or genuinely inferential.
ai-mlresearch - PhysMent: An Interactive Approach For LLM Reasoning In Physics Problems
PhysMent is a new benchmark that evaluates LLMs on physical reasoning through iterative interaction with a MuJoCo physics simulator, requiring models to actively experiment rather than answer static questions. The work reveals that current models struggle with quantitative multi-step tasks (25-67% accuracy range), with failures driven by procedural limitations in tool use and exploration rather than conceptual understanding—a meaningful signal about reasoning bottlenecks beyond language comprehension.
ai-mlresearch - An Empirical Analysis of Factual Errors in Human-Written Text and Its Application to Factual Error Detection
Researchers analyzed newspaper article corrections to build a taxonomy of human-induced factual errors (including language-specific categories like kanji misconversions) and benchmarked LLMs on factual error detection, finding GPT-5.4 achieves only 52% F1 on synthetic data. The work redirects attention from hallucination detection toward the neglected problem of detecting errors in human-written text, providing structured evaluation methodology for practitioners building fact-checking systems.
ai-mlresearch - Identifying and Transferring Reasoning-Critical Neurons: Improving LLM Inference Reliability via Activation Steering
Researchers propose AdaRAS, a test-time activation steering method that identifies and intervenes on reasoning-critical neurons in LLMs to improve reliability on mathematical and coding tasks, achieving 13%+ improvements on AIME benchmarks without additional training. This addresses a practical bottleneck in deploying LLMs for high-stakes reasoning and demonstrates transferability across models and datasets, making it relevant to practitioners seeking inference-time reliability gains.
ai-mlresearchvelocity:hn-medium - Balancing Global Quality and Pronoun-Specific Feedback for Context-Aware Machine Translation
ProNMT, a reward-guided iterative self-training method, improves context-aware machine translation by combining global quality estimation with targeted pronoun-specific feedback, outperforming standard fine-tuning on English–German and English–French benchmarks. The work demonstrates that targeted linguistic feedback is most effective when balanced with global quality signals, relevant to practitioners optimizing specialized NLP tasks and those building discourse-aware translation systems.
ai-mlresearch - Self-Orchestrating Language Models: Leveraging Semantic Dependence for Efficient Inference
A thesis proposing self-orchestrating language models that annotate semantic token dependencies to enable parallel decoding, memory-efficient context management, and optimized diffusion ordering. This directly addresses fundamental efficiency bottlenecks in LLM deployment—latency under low batch size, KV cache memory strain, and hardware underutilization—with a generalizable framework applicable across inference strategies.
ai-mlresearchlong-signal:rddvelocity:hn-medium - Affinity-Aware Sharding for Delayed Tensor Parallelism
Researchers propose affinity-aware sharding for Delayed Tensor Parallelism (DTP), optimizing device placement of attention heads and FFN neurons to accelerate distillation/retraining after architectural adaptation. Achieves 50-67% faster convergence than naive layouts on small models (Qwen3-0.6B, Danube3-500M) with a sub-2-minute optimization procedure, signaling practical improvements for production inference systems.
ai-mlresearchvelocity:hn-medium - Merging the Knowledge of LLMs for Automatic Speech Recognition
Researchers propose integrating external language models into LLM-based ASR systems via LoRA parameter merging rather than traditional shallow fusion, eliminating inference-time computational overhead while improving domain adaptation on CSJ and LibriSpeech datasets. This addresses a practical deployment concern for practitioners scaling LLM-based speech systems, offering a parameter-efficient fusion pathway that maintains inference speed and memory efficiency.
ai-mlresearch - Psychosis involves a deficit of information compression in connected speech
Researchers used LLM-derived metrics (surprisal and intrinsic dimensionality) to demonstrate that psychosis involves a deficit in information compression during speech, with grammatical organization as the mechanistic link. This work bridges neuroscience, linguistics, and AI—showing how LLM representations can reveal computational deficits in psychiatric conditions and refining theories of altered semantic geometry in schizophrenia-spectrum disorders.
ai-mlresearch