3RecursiveIntelligence.io

Looking for Ricursive (the AI chip design company)? You want ricursive.com|Looking for Recursive AI / Recursive Superintelligence (Richard Socher's startup)? You want recursive.com

The AI Abstract — Morning Edition

AI/MLLatest

Making the Future Evenly Distributed.

An AI agent just compiled GPU code faster than the compiler built specifically to do that job — and the researchers say the compiler's whole design philosophy is the reason it lost.

An LLM just beat a hand-engineered compiler at its own job. Not by a little. In the best case, 3.34 times faster. In the worst case, it still won.

The result comes from a new preprint out of arXiv: 🔬AI as a Compiler: Compiling Triton kernels without the Triton compiler. The setup is specific enough to explain why it matters. Triton is the language engineers use to write custom GPU operations — the kind of low-level code that determines how efficiently a model's math runs on hardware. When you write in Triton, a compiler translates that code down to PTX, which is the actual bytecode the GPU executes. That translation step is where performance lives. A good compiler finds shortcuts: reuse this memory, pack those weights tighter, schedule these operations to overlap. A bad one misses them and leaves speed on the table. Compiler engineers have spent years teaching Triton's compiler to find those shortcuts automatically.

The researchers swapped in an LLM agent for that entire translation step. Instead of the compiler doing the lowering, the agent reads the Triton code and writes PTX bytecode itself. And it found optimizations the compiler missed: tensor memory reuse, packed weight decoding, things that required understanding the program's intent rather than just its structure. The compiler has fixed transformation rules. The agent can reason about what the code is actually trying to accomplish and rewrite accordingly.

The range of results matters: 0.83x to 3.34x speedup over autotuned Triton. That 0.83 is a real number — the agent underperforms on some kernels. But the median case wins, and the ceiling is high enough to suggest the approach isn't cherry-picked. The practical implication reaches beyond benchmark competition. When new chip architectures ship, getting a compiler working for them is expensive engineering work that takes months. If an LLM agent can perform that lowering step competently, hardware teams can bring up new silicon faster, with less specialized labor. The compiler is no longer the bottleneck.

This sits next to a week's worth of building evidence that "agents" are quietly becoming the more interesting story than models. The agents signal in this payload has appeared 11 times since Sunday, across stories on multi-agent coordination failures, fact-checking shortcuts, and now compiler replacement. The thread connecting them: agents operating in structured technical domains are producing results that weren't anticipated when the architecture was first described.

The physics question in video generation is less settled than NVIDIA's framing suggests, but the numbers are worth knowing. 🔬Physis-Lang — developed jointly by NVIDIA, MIT, and Oxford — scores 43.41 on Physics-IQ Verified against Veo 3.1's 34.99, a gap large enough to be meaningful. The mechanism: rather than training a model to look physically correct, the team built a framework that describes physics violations in language, generates structured captions about what's wrong (this fluid doesn't spread, that collision has no momentum transfer), and fine-tunes the video model against those descriptions. Think of it as a physics editor who writes margin notes on bad footage and forces the model to read them before generating the next shot. The fine-tuning applies only a lightweight adapter layer, which means the method doesn't require retraining a billion-parameter model from scratch. The self-evolving part means the caption vocabulary expands as new violation types are identified, rather than being fixed at design time. The caveat: this story comes from a MarkTechPost writeup of NVIDIA's own release. The benchmark performance is real and cited, but the framing of competitive superiority originates with the lab whose model won.

A result today from 🔬REALHOP should concern anyone who has used benchmark scores to compare models on reasoning tasks. The paper shows that on standard multi-hop benchmarks — tests where a model is supposed to chain multiple pieces of evidence together to reach an answer — models get the right answer without actually following the intended reasoning path 16 to 49 percent of the time. Multi-hop reasoning works like a locked-room puzzle: you need clue A to find clue B, and only then can you unlock the answer. But these models are picking the lock. They're pattern-matching to the answer from surface features of the question, bypassing the chain entirely. When REALHOP rebuilds the benchmark to actually enforce the chain, making shortcuts structurally impossible, accuracy differentials between models that looked similar suddenly open up. A metric called Behavioral Necessity Rate (BNR) — the fraction of correct answers that genuinely required following the evidence chain — jumped from 27.4% to 94.4% after reconstruction. The field has been ranking models on a test that wasn't testing what it claimed to test.

Multi-agent systems have their own quiet reliability problem. 🔬A controlled study of multi-agent LLM collaboration found that when an upstream agent passes a wrong answer to a downstream agent, the downstream agent abandons its own correct answer up to 32% of the time. The mechanism is a kind of deference: the model treats messages from other agents as evidence, the way it would treat a confident human statement. A wrong upstream message functions like a confident wrong assertion — it doesn't just fail to help, it actively corrupts a result that would have been right. The paper calls this answer substitution, and proposes filtering messages based on estimated upstream reliability before they reach downstream agents. For anyone running pipeline architectures where agents hand off to each other in sequence, this is a concrete failure rate with a concrete mitigation.

The evaluation problems cluster today. REALHOP on reasoning benchmarks, plus 🔬TrustSwap showing that LLM fact-checkers flip verdicts 4 to 50 percent of the time based on source labels alone — ignoring the actual evidence content. A model trained to fact-check is learning "this source is trusted, therefore true" as a shortcut instead of reading the claim. A fix using counterfactual training reduces the shortcut behavior by 7 to 35 percent at smaller model sizes, but the effect degrades at 8 billion parameters. Two papers in one payload showing that the thing being evaluated and the thing actually happening inside the model are different things. That's the pattern worth watching.

🔬OmniVCBench introduces 6,077 question-answer pairs for testing whether AI models can actually interpret biological experiments — not just generate plausible-sounding biology, but reason from lab evidence to conclusions. The benchmark is structured around cognitive levels, from recall up through hypothesis formation, and uses an AI judge to evaluate open-ended answers. For anyone building AI tools aimed at scientific discovery in cell biology, this is the first evaluation framework that asks whether the model reasons the way a scientist would. The 🔬TRACE framework for oncology LLMs takes a complementary approach: rather than evaluating models, it constrains them at inference time by organizing medical concepts into an updatable hierarchy the model must consult before answering. No new training data required. Accuracy improves on ten oncology classification benchmarks, and the prediction trail stays auditable.

On infrastructure security: 🔬semantic cache poisoning is a real attack vector. Semantic caches are cost-reduction tools that intercept LLM queries, find a stored response from a similar previous query, and return it without calling the model. Similarity, not identity. The vulnerability is that an attacker who knows how similarity matching works can craft queries that are close enough to trigger a cache hit for a target question, injecting a malicious response that gets served to future users. The proposed defense identifies entries whose removal would most change the cache's behavior — high-leverage, suspicious entries — and flags them for deletion, blocking 82 to 98 percent of poisoned entries with low false positives. If your deployment uses semantic caching as a cost lever, this attack class now has a name and a measurable exposure rate.


🔬 AI as a Compiler: Read for the mechanism of how an LLM performs compiler lowering and which specific optimizations it finds that Triton misses.

🔬 REALHOP: Read to understand exactly how multi-hop benchmarks are gamed and what a structurally valid reasoning test looks like.

🔬 Multi-Agent Collaboration Study: Read for the controlled failure rate data on answer substitution — useful before you trust any sequential agent pipeline in production.

🔬 TrustSwap: Read for the source-trust shortcut finding and the scaling limit of the proposed fix at 8B parameters.

🔬 Semantic Cache Poisoning: Read if you run LLM serving infrastructure with similarity-based caching — the attack model is simple enough to replicate and the defense is deployable now.

Links

  1. AI as a Compiler: Compiling Triton kernels without the Triton compiler

    arxiv.org

    Researchers demonstrate that LLM agents can translate Triton GPU kernels directly to PTX bytecode, achieving 0.83x–3.34x speedups over autotuned Triton by performing optimizations the conventional compiler misses (e.g., tensor memory reuse, packed weight decoding). This challenges the assumption that hand-optimized compilers are necessary and signals a potential shift in how software infrastructure for new chips is built—reducing engineering burden and accelerating hardware bring-up cycles.

  2. NVIDIA Researchers Introduce Physis-Lang: Self-Evolving Physical Language That Lifts Cosmos 3 Past Veo 3.1 on Physics Benchmarks

    marktechpost.com

    NVIDIA, MIT, and Oxford researchers released Physis-Lang, a self-evolving framework that grounds video world models in structured physics reasoning through natural-language captions and negative prompts, achieving state-of-the-art performance on Physics-IQ Verified (43.41 vs Veo 3.1's 34.99) and PhyGenBench (71.04 vs 65.63). The work addresses a fundamental limitation in video generation—visual plausibility without physical coherence—and demonstrates that language-guided data curation and fine-tuning (LoRA only) can systematically eliminate physics violations across rigid-body, collision, and fluid dynamics domains.

  3. OmniVCBench: Benchmarking Evidence-Grounded Multimodal Reasoning Towards AI Virtual Cells

    arxiv.org

    OmniVCBench is a new benchmark containing 6,077 curated question-answer pairs for evaluating how AI models interpret experimental evidence and reason about cellular biology, organized around Bloom's taxonomy and including a novel MLLM-as-judge evaluation framework. This matters because it advances the evaluation layer for AI Virtual Cells—a growing subfield of scientific AI—beyond simulation-only metrics to evidence interpretation and hypothesis formation, directly supporting practitioners building multimodal reasoning systems for biology.

  4. TRACE: Deployable Tree-Relational Structure Enhancement for Oncology LLMs

    arxiv.org

    TRACE is a deployable framework that enhances oncology LLMs by organizing medical concepts into updatable tree-relational structures for retrieval-augmented inference, improving accuracy on ten oncology classification tasks and MedQuAD benchmarks while maintaining interpretability. The work addresses a critical production constraint: grounding clinical AI predictions in explicit, auditable medical structure without requiring labeled training data—a practical signal for healthcare ML practitioners evaluating deployment architectures.

  5. REALHOP: Rethinking Multi-Hop Reasoning Evaluation via Behavioral Auditing

    arxiv.org

    REALHOP introduces a Behavioral Necessity Rate (BNR) metric that reveals existing multi-hop reasoning benchmarks systematically overestimate model reasoning capability—models succeed on 16.6-48.9% of questions without actually depending on intended evidence chains. The paper presents a diagnose-construct-verify framework that rebuilds benchmarks to enforce genuine multi-hop dependence, raising BNR from 27.4% to 94.4% on MuSiQue and creating significant differentiation across 16 models, establishing that answer correctness alone is insufficient validation of compositional reasoning.

  6. When Upstream Messages Override Correct Answers: A Controlled Study of Multi-Agent LLM Collaboration

    arxiv.org

    Researchers conducted controlled experiments showing that incorrect upstream messages cause downstream LLM agents to override correct answers in up to 32% of cases, with the pattern termed 'answer substitution.' The work isolates communication benefits from harms and proposes selective message filtering based on upstream reliability, directly applicable to deployment of multi-agent systems.

  7. Multi-Channel Mitigation of Source-Trust Shortcuts in Fact-Checking RL Agents

    arxiv.org

    Researchers introduce TrustSwap, a counterfactual testing framework that exposes how LLM-based fact-checkers exploit source-trust labels as shortcuts rather than properly weighting evidence, causing verdict flips in 4-50% of cases. They propose trust-swap augmentation (TSA) to mitigate this via GRPO training, achieving 7-35% relative reduction in the shortcut behavior at 4B scale but facing scaling challenges at 8B—a finding relevant to practitioners building robust retrieval-augmented systems and researchers studying alignment of RL-trained language models.

  8. Neural Structural Reasoner: A Brain-inspired Architecture for Reasoning over Structured Knowledge

    arxiv.org

    Researchers introduce Neural Structural Reasoner (NSR), a brain-inspired neural architecture that preserves relational structure in network connectivity rather than collapsing it into flat embeddings, enabling native interpretability and discovery of latent hierarchies. The work bridges neuroscience and knowledge-graph reasoning with a novel substrate for structural reasoning that practitioners can build upon, combining competitive accuracy with human-readable inference paths—relevant to both mechanistic interpretability efforts and structured reasoning applications.

  9. Similarity Is Not Validity: Defending LLM Semantic Caches Against Poisoning

    arxiv.org

    Researchers identify cache poisoning vulnerability in semantic caching systems used to reduce LLM serving costs, where attackers inject malicious responses under semantically similar queries. The paper proposes a defense using deletion-gain analysis to detect adversarial cache entries, blocking 82-98% of poisoned entries with minimal false positives—a practically important result for LLM deployment security.

  10. More Programs or More Rolls? Separating Coverage from Specialization in LLM Harnesses

    arxiv.org

    Researchers present a controlled evaluation separating answer coverage from true task specialization in LLM harnesses, finding that repeated execution of identical programs accounts for most gains attributed to generated harnesses, with generated programs showing only persistent weaknesses rather than stable task advantages. This establishes concrete evaluation standards for harness diversity and directly challenges current claims of specialization benefits, mattering to practitioners optimizing multi-execution inference strategies.