Program Semantic Inequivalence Game with Large Language Models

arXiv:2505.03818v3 Announce Type: replace-cross Abstract: Large Language Models (LLMs) can achieve strong performance on everyday coding tasks, but they can fail on complex tasks that require non-trivial reasoning about program semantics. Finding training examples to teach LLMs to solve these tasks can be challenging. In…

Source: cs.AI updates on arXiv.org

Automatically aggregated summary — full article and all rights belong to the original publisher.

Scientists just created female clones of male mice

Scientists have deliberately turned male mouse embryos into females for the first time. A team based in Japan used a CRISPR-based approach to remove the Y chromosome from male cells and create female clones of male mice.  “No one has done this before,” says Monika…

Source: MIT Technology Review

Automatically aggregated summary — full article and all rights belong to the original publisher.

Scaling AI agents with trustworthy data

Business and technology leaders need no convincing that the time of agentic AI is here. Organizations are rapidly adopting agents, and few executives doubt the technology’s potential to transform work. But many organizations find that realizing the desired return on investment (ROI) from AI hinges…

Source: MIT Technology Review

Automatically aggregated summary — full article and all rights belong to the original publisher.

The Download: our 35 young innovators and the “censorship-industrial complex”

This is today’s edition of The Download, our weekday newsletter that provides a daily dose of what’s going on in the world of technology. How we picked 35 of the world’s top young scientists and engineers On September 8, MIT Technology Review will reveal its…

Source: MIT Technology Review

Automatically aggregated summary — full article and all rights belong to the original publisher.

Pakistani Judges Give Their Verdict on JudgeGPT

Judges around the world have made headlines for illicitly using generative AI in their work. But in Pakistan, a large-scale trial of a specially designed AI tool for judges found the technology–together with appropriate training–boosted the number of cases resolved by 6.3% with no obvious…

Source: IEEE Spectrum

Automatically aggregated summary — full article and all rights belong to the original publisher.

How we picked 35 of the world’s top young scientists and engineers

Next month, on September 8, MIT Technology Review will reveal its 2026 list of Innovators Under 35, recognizing 35 young people from around the world who are doing groundbreaking scientific work and building clever technical fixes for sticky problems.  By finding the top young innovators…

Source: MIT Technology Review

Automatically aggregated summary — full article and all rights belong to the original publisher.

Memory-Augmented Reinforcement Learning Agent for CAD Generation

arXiv:2605.19748v2 Announce Type: replace Abstract: Automatic generation of computer-aided design (CAD) models is a core technology for enabling intelligence in advanced manufacturing. Existing generation methods based on large language models (LLMs) often fall short when handling complex CAD models characterized by long operation sequences, diverse…

Source: cs.AI updates on arXiv.org

Automatically aggregated summary — full article and all rights belong to the original publisher.

Test-Time Self-Evolving GUI Visual Grounding via Reflection-Guided On-Policy Self-Distillation

arXiv:2608.11191v1 Announce Type: cross Abstract: GUI Visual Grounding is a fundamental capability for GUI agents. Existing models typically freeze their parameters after deployment, limiting their ability to adapt to unseen interfaces. Although recent methods attempt to adapt models via test-time reinforcement learning, they cannot reflect…

Source: cs.AI updates on arXiv.org

Automatically aggregated summary — full article and all rights belong to the original publisher.