CompoWorld: Compositional Environment Scaling for General Agents

arXiv:2609.33665v2 Announce Type: replace Abstract: Automatically generated environments provide a scalable source of interaction data for training general agents. However, existing approaches mainly generate tasks within a single environment, while real-world workflows require agents to connect information and actions across multiple services. We introduce Compositional…

Source: cs.AI updates on arXiv.org

Automatically aggregated summary — full article and all rights belong to the original publisher.

Instruction Retrieval at Inference Time for Small Language Models

arXiv:2510.13935v3 Announce Type: replace-cross Abstract: The facts a language model stores are tied to its parameter count, so small models that fit on edge devices fail on expert problems, which need specialized knowledge and follow multi-step procedures. Fine-tuning for a specific domain or task writes…

Source: cs.AI updates on arXiv.org

Automatically aggregated summary — full article and all rights belong to the original publisher.

PhantomEnvironments: Training LLM Agents in Fictional Worlds

arXiv:2609.40221v1 Announce Type: cross Abstract: Training LLM agents with reinforcement learning (RL) is bottlenecked by environments, which must provide verifiable rewards, support long-horizon interaction, and scale cheaply. Existing approaches rely on costly human-curated data or on LLM-generated environments that risk hallucinations and benchmark contamination. We…

Source: cs.AI updates on arXiv.org

Automatically aggregated summary — full article and all rights belong to the original publisher.

The Download: OpenAI’s chief research officer explains its hacking response

This is today’s edition of The Download, our weekday newsletter that provides a daily dose of what’s going on in the world of technology. “We’re not going to shoot ourselves in the foot” over hack fallout, says OpenAI’s chief research officer Two months after OpenAI’s…

Source: MIT Technology Review

Automatically aggregated summary — full article and all rights belong to the original publisher.

“We’re not going to shoot ourselves in the foot” over hack fallout, says OpenAI’s chief research officer

Two months after the bombshell news that a swarm of its agents had broken their containment and hacked into the computers of the AI company Hugging Face, OpenAI is still putting out fires. A steady drip of disclosures about other hacks in the weeks since…

Source: MIT Technology Review

Automatically aggregated summary — full article and all rights belong to the original publisher.

Independent Verification Paths Are Not Independent: A Case Study of Common-Mode Failure in a Satellite Catalogue Pipeline

arXiv:2609.37603v1 Announce Type: cross Abstract: A common safeguard for a data pipeline is redundant computation: derive each published number by two routes built on different technology and refuse to exit when they disagree. We report one such gate failing, in a cross-catalogue integrity study of…

Source: cs.AI updates on arXiv.org

Automatically aggregated summary — full article and all rights belong to the original publisher.