Meet GPT-Red: an LLM super-hacker OpenAI built to make its models safer

OpenAI has built an LLM super-hacker called GPT-Red that it uses as a sparring partner to help its other models boost their defenses against cyberattacks. Last week the company released the latest version of its flagship LLM, GPT-5.6. OpenAI says that training it against GPT-Red…

Source: MIT Technology Review

Automatically aggregated summary — full article and all rights belong to the original publisher.

The First Chatbot’s Multiple Personalities

ELIZA is remembered as the world’s first AI star, a kindly therapist in chatbot form that gently probed users’ worries. Even its creator, Joseph Weizenbaum, was surprised by the warm reception given to his experiment in human-machine interaction. For some, it heralded an age of…

Source: IEEE Spectrum

Automatically aggregated summary — full article and all rights belong to the original publisher.

This AI Folds DNA Into Mini Masterpieces

Shaped like dogs, stars, and the Mona Lisa, you could mistake these DNA structures for fun-shaped macaroni if they weren’t only nanometers wide. South Korean scientists made the constructions using a technique called DNA origami, which can bend genetic material into any form. Designing DNA…

Source: IEEE Spectrum

Automatically aggregated summary — full article and all rights belong to the original publisher.

The Download: a useful quantum machine and a record-breaking subsea tunnel

This is today’s edition of The Download, our weekday newsletter that provides a daily dose of what’s going on in the world of technology. PsiQuantum has a plan to make a massive quantum computer out of light The machine that could change the world will…

Source: MIT Technology Review

Automatically aggregated summary — full article and all rights belong to the original publisher.

Best AI Models 2026: Ratings by Task & Strength

There Is No Best AI Model—Only the Best Model for the Right Job

Every few months, a new benchmark appears claiming that one model has become the world’s “best.”

Today it’s GPT.

Tomorrow it’s Claude.

Next month it may be Gemini, DeepSeek, or Qwen.

The reality is much more interesting.

There is no universal winner.

The latest independent benchmark reports consistently show that each frontier model dominates different dimensions of intelligence, cost, reasoning, coding, multimodal understanding, or context handling.


AI Is Becoming Specialized

Think about elite athletes.

You wouldn’t ask:

“Who’s the best athlete?”

Instead you ask:

• Who’s the fastest sprinter?
• Who’s the strongest weightlifter?
• Who’s the best marathon runner?

Large Language Models have entered the same stage of evolution.

Each excels at different work.


The Emerging Strengths

Claude

Best for:

• agentic software engineering
• autonomous coding
• complex development workflows


GPT-5.6

Best for:

• enterprise reasoning
• mathematics
• business analysis
• balanced performance across many tasks


Gemini

Best for:

• multimodal reasoning
• long-context analysis
• video, images, documents
• enterprise knowledge retrieval

Its extremely large context window allows it to process information that would overwhelm many other models.


DeepSeek

Best for:

• high reasoning performance
• cost efficiency
• open deployment strategies

Organizations can achieve near-frontier performance while dramatically reducing inference costs.


Qwen

Best for:

• multilingual applications
• open-weight deployment
• enterprise customization
• strong reasoning with infrastructure flexibility


This Isn’t New

Every major technology revolution followed exactly the same pattern.

Cloud didn’t replace everything.

It created:

• Infrastructure as a Service
• Platform as a Service
• Software as a Service

Each solved different problems.

AI is following the same trajectory.

Instead of one model replacing every other model…

We’re moving toward AI ecosystems.


The Leadership Shift

The executive question is no longer:

Which AI should we buy?

The better questions are:

  • Which workflow requires deep reasoning?
  • Which process needs multimodal understanding?
  • Which tasks require the lowest possible cost?
  • Which activities should remain human-led?

The competitive advantage is shifting from model selection to AI orchestration.


Executive Coaching in the AI Era

Leadership is becoming less about knowing the answers.

It’s becoming about designing intelligent systems.

Future executives must learn to:

  • orchestrate multiple AI models
  • redesign workflows
  • Govern AI responsibly
  • Combine human judgment with machine intelligence

This is remarkably similar to earlier transformation waves.

The organizations that succeeded in cloud, Agile, and digital transformation weren’t those that adopted technology first.

They were the ones who redesigned how work flowed through the organization.

AI is no different.


The future won’t belong to the organization using one “perfect” AI.

It will belong to the organization that knows which intelligence to apply, at the right moment, for the right problem.

Because in the age of AI, the competitive advantage is no longer the model.

It’s the leader who knows how to orchestrate them.

The post Best AI Models 2026: Ratings by Task & Strength first appeared on Agile Agilist | SAFe® Gold Partner for Agile Transformation, Innovation & Leadership Training.

Does Topic Sentiment Cause Perceived Ideology? Comparing Human and LLM Annotations in Political News Articles

arXiv:2606.06715v2 Announce Type: replace-cross Abstract: We ask whether topic sentiment has a causal effect on perceived political ideology, and whether the answer depends on who assigns the ideology label. Using articles from AllSides, paired with shared sentiment annotations from Llama-3.3-70b-versatile, we compare ideology labels from…

Source: cs.AI updates on arXiv.org

Automatically aggregated summary — full article and all rights belong to the original publisher.

Declarative by Design, Assistable Only by Convention: Benchmarking Multi-Agent Frameworks for AI-Assistability

arXiv:2602.11198v2 Announce Type: replace-cross Abstract: Multi-agent frameworks (MAFs) promise to simplify LLM-driven software development, yet no principled metric captures how well AI coding assistants can generate correct, framework-specific code. We introduce textit{AI-assistability} ($mathcal{AI}$), a composite metric that quantifies a framework's amenability to AI-assisted development by…

Source: cs.AI updates on arXiv.org

Automatically aggregated summary — full article and all rights belong to the original publisher.

FFAvatar: Feed-Forward 4D Head Avatar Reconstruction from Sparse Portrait Images

arXiv:2606.30347v2 Announce Type: replace-cross Abstract: We present FFAvatar, a Transformer-based 3D Gaussian framework for fast construction of high-quality and animatable 4D head avatars from one or more reference portrait images. Unlike existing feed-forward approaches that require a fixed number of input views, FFAvatar supports incremental…

Source: cs.AI updates on arXiv.org

Automatically aggregated summary — full article and all rights belong to the original publisher.