FM-Bench: A Benchmark for Long-Horizon Management with Competing Agents

arXiv:2608.18423v2 Announce Type: replace Abstract: Language model agents now execute bounded tasks reliably. Whether they can sustain effective decision-making over long horizons, where actions have cumulative consequences and the environment responds to their choices, remains largely unmeasured. FM-Bench (Football Management Benchmark) measures this. An LLM…

Source: cs.AI updates on arXiv.org

Automatically aggregated summary — full article and all rights belong to the original publisher.

This IEEE Senior Member Develops AI Tools for E-Commerce Sites

Balaji Ingole rarely saw televisions while growing up in Udgir, India. No one in the small Maharashtra village had computers or phones. Only one household owned a television, and neighbors often gathered there to watch shows together.Ingole never even saw a computer growing up. It…

Source: IEEE Spectrum

Automatically aggregated summary — full article and all rights belong to the original publisher.

Stop Hunting, Start Solving: Accelerating Root Cause Analysis with Agentic AI

About this WebinarTurn Yield Excursions into Faster, More Confident Root Cause AnalysisWhen a yield issue emerges, the answer rarely lives in a single system. Critical clues are spread across metrology data, tool traces, chemical analysis, and facilities systems, while growing data volumes make traditional dashboards…

Source: IEEE Spectrum

Automatically aggregated summary — full article and all rights belong to the original publisher.

The Download: threats from space mirrors and credit for AI drugs

This is today’s edition of The Download, our weekday newsletter that provides a daily dose of what’s going on in the world of technology. This company’s plans to deploy space mirrors could jeopardize the night sky for many A company that plans to beam sunlight…

Source: MIT Technology Review

Automatically aggregated summary — full article and all rights belong to the original publisher.

This company’s plans to deploy space mirrors could jeopardize the night sky for many

A company that plans to beam sunlight from space to Earth on demand might unintentionally brighten the night sky for many more people than intended, according to a new study. Later this year, the US company Reflect Orbital plans to launch a test satellite called…

Source: MIT Technology Review

Automatically aggregated summary — full article and all rights belong to the original publisher.

When AI designs a drug, who gets the credit?

When the biotech company Insilico Medicine used its computer models to propose a promising drug for pulmonary fibrosis, it enthusiastically claimed in a press release that the molecule had been “discovered by” its generative AI platform. Insilico leads a pack of companies using AI to…

Source: MIT Technology Review

Automatically aggregated summary — full article and all rights belong to the original publisher.

A knowledge-guided agentic framework for mitigating patient-context ambiguity in health queries

arXiv:2608.19875v1 Announce Type: cross Abstract: Patients often submit short, underspecified queries to healthcare chatbots that lack the patient-specific information needed to determine an appropriate response. Although these queries may be linguistically clear, they can support multiple plausible answers depending on undisclosed factors such as symptoms,…

Source: cs.AI updates on arXiv.org

Automatically aggregated summary — full article and all rights belong to the original publisher.

SoftVTBench: A Safety-Aware Visuo-Tactile Benchmark for Physically Constrained Robotic Manipulation of Deformable Objects (Early Version)

arXiv:2607.04234v2 Announce Type: replace-cross Abstract: Deformable object manipulation poses challenges beyond task completion: successful execution must also maintain safe physical interaction, holding the object stably without slip or drop while avoiding excessive deformation. However, existing manipulation benchmarks are predominantly success-oriented and rarely evaluate whether a…

Source: cs.AI updates on arXiv.org

Automatically aggregated summary — full article and all rights belong to the original publisher.

Distilling Aggregate Mobility Statistics into a Language Model Policy for Post-Event Crowd Simulation

arXiv:2608.19778v1 Announce Type: cross Abstract: Pedestrian simulators need a behaviour rule for every agent, but privacy usually limits the data for setting one to aggregate statistics, namely zone-level device counts and origin-to-destination (OD) flows, with no individual trajectories. Such aggregates under-determine individual behaviour, because many…

Source: cs.AI updates on arXiv.org

Automatically aggregated summary — full article and all rights belong to the original publisher.