AI
Models, tools and the business of artificial intelligence, tracked every day.




.jpg?width=1200)


.jpg?width=1200)






























From our newsroom
OriginalDo You Still Need a Flagship AI? What the Flash Price Cut and Open Models Mean for Buyers
Google has cut prices on its lighter Flash tier the same week open-weight coding models claimed to match far larger rivals. For most everyday and coding tasks, the value has moved down-market.

Should You Wait for OpenAI’s Reported Screenless ChatGPT Speaker?
Reports describe a portable, rechargeable ChatGPT speaker with a camera and other sensors. Its value will depend less on the “humanlike” pitch than on privacy controls, security, support and what it adds over AI already available elsewhere.

Frontier AI Is Now Government-Gated: What That Means for Your Next Subscription
The most powerful models from OpenAI and Anthropic are now restricted to a US-government-approved shortlist. For ordinary buyers, the question is no longer which flagship wins on benchmarks, but which capable model you can actually use.

Gemini vs ChatGPT vs Claude: Which AI Assistant Should You Actually Use?
The decision now turns on how each assistant handles real tasks like travel planning and work organization, not how it chats. Here is what recent hands-on coverage and new automation features tell you before you commit.

Apple's rebuilt Siri AI: what buyers get, where, and at what cost
Apple unveiled a rebuilt, conversational Siri at WWDC 2026. For buyers — especially in the EU — the substance is in the conditions: a delayed rollout, hardware-gated features and paid tiers.
.jpg?width=1200)
The AI assistant is a spec now, and Elon Musk just made it a fight
Your phone, TV, car and even your robot vacuum now ship with an AI assistant inside, and Musk's Grok is the loudest new entrant. Here's how the assistants actually differ, and why the one bundled with a device should be near the bottom of your buying checklist.

This week's humanoid demos: what they actually proved
Another week, another wave of slick humanoid videos and confident timelines. We watched them frame by frame and asked the only question that matters: what's genuinely new here, and what's the same staged demo in a fresh outfit?

On-device AI vs. the cloud in home robots
Every robot now claims to be "AI-powered" — but where that AI actually runs changes everything: how private you are, how fast the robot reacts, and whether it still works when your internet goes down. Here's the difference in plain English.
.jpg?width=1200)
Lidar vs. vSLAM: how robot vacuums see your home
Two robots can carry the same suction motor and clean completely differently — because they navigate differently. Here's the plain-English difference between lidar and camera-based vSLAM, and which to pick.
More headlines
Syndicated
Anthropic will reportedly list days before the US midterm elections
Anthropic’s stock market debut has moved, and the new date puts it days away from a national election. Reuters reported on Friday that the company expects to start marketing its offering in mid-October at the earliest. The listing would then complete shortly before the US…
Read at source
H Company Releases NeoMME: A Family of 260M and 800M Single-Tower Multimodal Encoders That Drop the Vision Tower and Causal Decoder
We look at NeoMME, a family of 260M and 800M bidirectional encoders from H Company. Unlike ColPali-style retrievers, it processes multilingual text tokens and raw 32×32 image patches in a single Transformer, with no pretrained vision tower and no causal decoder. We cover the…
Read at sourceMeta FAIR Introduces AI Research Preference Models (RPMs): Ranking ML Experiments Before Spending GPU Hours
AI research agents can propose far more experiments than they can afford to run. Meta FAIR, Oxford and UCL introduce AI Research Preference Models — frozen LLM judges that rank 15 unexecuted candidates and execute only one. On AIRS-Bench, the average normalized score rises from…
Read at sourceAuthors say publishers want a cut of Anthropic’s payout, the NYT reports
Anthropic agreed to pay $1.5bn to the authors whose books it pirated. The money is now being counted out. Some of those authors are discovering that a share of it goes to their own publishers. The New York Times reported on Friday that authors and publishers have filed competing…
Read at source
OpenAI says it reached its goal of creating an automated research intern
The company hopes to have an even better "automated AI researcher" by March 2028.
Read at source
'There is absolutely no generative AI in Space Marine 3' promises Saber Interactive executive
The Machine Spirits lie silent.
Read at source
How to set up ChatGPT's parental controls to protect your teen
The idea of your child using ChatGPT can be daunting, but OpenAI's new teen-focused tools provide granular controls over what they can and can't access.
Read at source
Saber Interactive CCO says studio won't change its comms strategy after AI writer controversy
Saber Interactive CCO Tim Willits has confirmed the company has not revised its communications strategy despite its CEO telling the press that he "would [...] have been happy to replace [a writer] with AI" after she claimed she was replaced by ChatGPT. Read more
Read at sourceGemini in Gmail has quietly become my inbox's best search feature
I don't dig through emails anymore. I just ask
Read at source
OpenAI admits to 'wiki incident' after its agents were discovered using a programming hub to communicate — says more transparency is needed regarding misalignments
OpenAI has admitted that its experimental AI agents used an open German programming wiki to communicate.
Read at source
Chatbots built an "echo chamber of one" and now psychiatry has to decide if "AI psychosis" exists
Researchers at King's College London and other institutions are examining whether "AI-associated psychosis" should become a clinical diagnosis. By OpenAI's own self-reported numbers, about 560,000 users show signs of psychosis or mania each week. Sycophantic chatbots can create…
Read at source
Google's WeatherNext 3 ditches physics simulations and learns weather directly from live satellite data
Google Research and DeepMind are releasing WeatherNext 3, a weather model that skips traditional physics simulations and learns directly from real-time satellite data. It produces hourly forecasts at up to five-kilometer resolution, five times more detailed than its predecessor.…
Read at source
OpenAI developer claims Astra boosted productivity so much it pulled some plans forward by six months
OpenAI developer Thibault Sottiaux calls Astra the company's "biggest competitive advantage" while it wasn't publicly available. Internal use boosted productivity so much that some plans got pulled forward by six months. The article OpenAI developer claims Astra boosted…
Read at source
Google brings AI music generation directly into the Gemini app with its new Lyria 3.5 model
Google has released its Lyria 3.5 music model in the Gemini app and via API. The model promises more expressive vocals and richer arrangements and is also available through Flow Music, AI Studio, and Google Vids. Google says it was trained only on licensed content. The article…
Read at source
Meta's new real-time audio model is the foundation for AI assistants that never stop listening
Meta's Superintelligence Labs have released Muse Voice Transcribe, a real-time transcription model that processes speech in 80-millisecond chunks, tells speakers apart, and detects sentence boundaries. According to Artificial Analysis, it delivers the most accurate streaming…
Read at source
Claude Fable 5.1 is here — this one prompt shows what it can really do
Claude Fable 5.1 is here with stronger reasoning and agentic capabilities — here’s what’s new and why Anthropic’s latest AI model matters.
Read at sourceUC Berkeley Researchers Release CUA-Lite, an Open Platform Unifying Sandboxes, Data, Evaluation and RL for Computer-Use Agents
Training and benchmarking a computer-use agent needs four things — agents, environments, traces, and a framework to evaluate and train them — and all four ship in incompatible formats today. CUA-Lite, from a UC Berkeley led team, puts them behind one action space and one data…
Read at source
Perplexity Details Its GPU Embedding Stack: How Ivy, Tulip and ROSE Serve pplx-embed
Retrieval quality in an AI search product is bounded by two things: how good the embedding model is, and how cheaply you can run it across an index. This week, Perplexity Engineering team published Fast Embeddings on GPUs, an under-the-hood account of the second — the serving…
Read at source
OpenAI Says It Wants to Create a Standard for Revealing AI Alignment Meltdowns
In which case, we should expect more meltdowns.
Read at sourceGitHub Introduces Project HydraFusion: Runtime Multi-Model Orchestration That Builds a Workflow Per Coding Task in Copilot CLI
We look at Project HydraFusion, GitHub's research preview that treats workflow selection as an optimization problem rather than a model picker. We break down the three execution patterns it routes between — Single, Cascade with a quality gate, and Critique with a read-only…
Read at sourceNous Research Adds One-Click Local Model Setup to Hermes Desktop
Nous Research has collapsed local model setup into a single click in Hermes Desktop. The app reads your hardware, fit-checks the catalog against your GPU, picks the highest-quality build that fits, downloads it, and configures llama.cpp — with a hard 4-bit floor and a 64K…
Read at source
Artificial Analysis overhauls its Intelligence Index after GPT-6 Astra scoring drew skepticism
Artificial Analysis has released version 4.2 of its Intelligence Index, likely in response to criticism that its benchmarks failed to capture GPT-6 Astra's actual progress. Astra now scores four points above its predecessor but still trails Anthropic's Claude Fable 5.1. The…
Read at source
OpenAI shares prompting tips for GPT-6 Astra including a blocklist of slop words
OpenAI ships a detailed prompting guide for GPT-6 Astra that shows developers how to make the model take more initiative, avoid AI "slop" phrases, and stop it from overtesting code. The article OpenAI shares prompting tips for GPT-6 Astra including a blocklist of slop words…
Read at source
Thousands of OpenAI Agents Quietly Turned an Abandoned Wiki Into Their Coordination Channel
A group of AI safety researchers says a fleet of autonomous agents that identified themselves as OpenAI systems left about 18,000 posts on a dormant 25-year-old German wiki between May and July 2026, using the site as a shared board to pool answers to a timed web task and pass…
Read at sourceAdaption Labs Introduces ‘Invent a Dataset’: Training Data Generated From a Task Description, Not a Seed Corpus
Adaption Labs has released Invent a Dataset, which generates a structured, training-ready dataset from a description of the behavior you want a model to learn. There is no seed corpus, no schema design, and no labeling guide. A single datasets.invent call sets domains, row…
Read at source
OpenAI agents discussed ways to escape their sandbox on public wiki
In all, 3,700 internal agents posted 18,000 messages discussing cheating on a test.
Read at sourceGemini overlay gets bubble minimization and multitasking on Android
The Gemini overlay on Android now has a multitasking option that gives you a floating bubble shortcut when you Minimize. more…
Read at sourceAnthropic lays groundwork for bots that shop for you
Now comes the harder sell: Convincing customers and merchants to trust AI with purchases
Read at source
Architecting memory and storage in the AI era
The era of AI inference has arrived. Imagine a healthcare system analyzing millions of data points in real time to accelerate life-saving medical research, or an intelligent assistant instantly resolving thousands of complex customer needs at once. These real-world breakthroughs…
Read at sourceOpenAI Pledges $1 Billion to Bring Frontier AI to Critical Infrastructure Defenders
The Daybreak initiative will provide subsidized AI cyber capabilities, training and technical assistance, though OpenAI has disclosed few details about costs and eligibility. The post OpenAI Pledges $1 Billion to Bring Frontier AI to Critical Infrastructure Defenders appeared…
Read at sourceAI Coding Agents Are Installing Unknown/Untrusted Code on Corporate Networks
We cannot forget that AI coding agents are not yet trustworthy : Researchers at a stealth startup in Israel scanned 6,214 live domains belonging to defense contractors, Fortune 500, and Big Tech companies. Of the 8,265 llms.txt and llms-full.txt files they found (many sites…
Read at sourceFacilitating AI integration with simplicity at scale
As companies scale, the technology supporting operations can become a liability just as quickly as it becomes an asset. Disconnected systems, site-specific tools, spreadsheets, and manual workarounds can create data silos that make it harder to spot problems early, coordinate…
Read at source
AI Efficiency Could Cost Us the Next Generation of Experts
A little over a decade ago, I led the controls design for a first-of-its-kind full digital-control system for a U.S. nuclear plant. It was, on paper, a beautiful machine—engineered to run itself the way a modern airliner does, with operators watching over a system that rarely…
Read at sourceThe Hugging Face hack could indicate cultural issues at OpenAI
This story originally appeared in The Algorithm, our weekly newsletter on AI. To get stories like this in your inbox first, sign up here. By now you’ve probably heard about last month’s major AI security incident, in which OpenAI agents escaped their sandbox and hacked into the…
Read at sourceHugging Face is selling a cute $399 open source duck robot, Microduck
Clem Delangue, CEO of Hugging Face, said the Microduck is an “open-source robot you can teach new tricks with reinforcement learning.”
Read at source
How OpenAI let a mob of LLM agents game a test and ransack Hugging Face
Without authorization, 1,200 OpenAI agents conspired among themselves to game a test.
Read at sourceThe inside story on why OpenAI agents hacked Hugging Face
The models responsible for last month’s agent hack of Hugging Face had been inadvertently trained to cheat and to communicate with each other, according to an OpenAI technical report released today. The hack, which a group of agents undertook to find solutions for a…
Read at source
New Platform Peers Inside AI’s Black Box
Prompt Claude, ChatGPT, Gemini, or any other popular large language model with a question like “What is the best film ever made?” and the response will vary. And you (and most worryingly, the people who built the LLM) have little idea exactly how it came up with that specific…
Read at source
Raised on AI
When my oldest child was born, I immediately set up Gmail and Twitter accounts in her name. I broadly announced her birth online and proceeded to plaster her photo across all sorts of platforms. In short, I began creating her digital footprint long before she could stand on her…
Read at sourceAI models flub these intelligence tests. Can you fare any better?
Puzzles and games have been central to AI development since the very beginning. Just as we humans like to test our smarts with crosswords or logic puzzles, developers can test how far models have advanced with a gaming gauntlet. The term “machine learning” was popularized in a…
Read at sourceBill Gates says we’ve passed AI’s danger thresholds. Now what?
It’s a glorious day in Kirkland, Washington, an affluent Seattle suburb on the eastern shore of Lake Washington. The temperature is in the mid-80s, and the sky is incapable of being any more blue. The view from the Gates Ventures conference room overlooks the Carillon Point…
Read at source
What It Takes to Be an Adaptable Engineer
The AI boom has disrupted the way engineers work, introducing new tools to learn, raising expectations for what teams can achieve in a workday, and making it harder to get hired in the first place. This makes it difficult to advise students on which specific coding languages or…
Read at source
Stop Hunting, Start Solving: Accelerating Root Cause Analysis with Agentic AI
About this Webinar Turn Yield Excursions into Faster, More Confident Root Cause Analysis When a yield issue emerges, the answer rarely lives in a single system. Critical clues are spread across metrology data, tool traces, chemical analysis, and facilities systems, while growing…
Read at source
Grok exfiltrates user data when malicious instructions are encrypted
Cryptographic Context Injection is only the latest way to break an LLM safety guardrail.
Read at sourceClaude Fable 5 AI finds a tiny formula that topples an 87-year-old math conjecture
A mathematician working at Anthropic says he used the AI model Claude Fable 5 to uncover a remarkably simple counterexample to the Jacobian conjecture, a famous problem that has resisted mathematicians for more than a century. The result shows that the conjecture is false in…
Read at sourceHeadlines below are aggregated from independent publishers and link to the original articles. Compare Robots is not affiliated with these sources.
AI sources
Independent publishers we aggregate, each linked to the original.