AI benchmarks

ai benchmarks tests grade ai wrong a carnival high striker with a bell and mallet

The Tests That Grade AI May Be Getting It Wrong

Two Stanford-led studies for COLM 2026 test AI benchmarks the way psychologists test exams. Across 56 benchmarks and 53 models, safety tests with the same label often disagreed, capability tests blurred together, and the scoring format mattered more than the topic. We explain the findings and what they mean for choosing a model.

Read more
wajo launches fo ai agent calls payments a smartphone standing upright

Wajo Launches Fo, an AI Agent That Can Make Calls and Payments

Wajo has launched Fo, a personal AI agent with its own phone number, inbox and single-use payment cards that calls businesses, pays for bookings and hands hard tasks to trained human assistants. We explain how its calls and four payment routes work, test its 2x and 94% trust claims against its own paper, and set out who carries the risk under its terms.

Read more
Claude Sonnet 5.5 - claude sonnet 5 5 30 faster 30 cheaper a hourglass with flowing sand

Anthropic Announces Claude Sonnet 5.5: 30% Faster and 30% Cheaper

Anthropic has released Claude Sonnet 5.5, which it says is more than 30% faster than Sonnet 5 and up to 30% cheaper per task. The token price has not changed, so the saving comes from fewer tokens and tool calls. We work through the cost arithmetic, the benchmark table against Opus 5.5 and GPT-6 Sol, the new cyber and distillation safeguards, and whether to switch.

Read more
spacexai grok 4 7 aixploria a engine nozzle bell on a mounting ring

Discover Grok 4.7, Freshly Added to AIxploria

AIxploria added SpaceXAI’s Grok 4.7 at 01:29 UTC on 24 September 2026, a little over two days after launch. Its blurb calls the model “twice as fast and half the cost,” but SpaceXAI’s own announcement says Grok 4.7 is served “at the same price and speed as Grok 4.6,” and its “half the price” claim referred to rival models such as GPT-5.6 Sol, which OpenAI replaced the next day. We checked 13 claims, worked through the 200K-token pricing threshold, compared Grok 4.7 with GPT-6 Sol and Claude Opus 5.5 on three job sizes, and tested Musk’s pre-launch forecasts.

Read more
Claude Opus 5.5 - claude opus 5 5 aixploria a rosette badge disc with two ribbon tails

Claude Opus 5.5 Is Now Available on AIxploria

AIxploria published its Claude Opus 5.5 card at 03:23 UTC on 24 September 2026, about 35 hours after Anthropic’s launch. We checked 14 of its claims against Anthropic’s launch post, model overview and pricing page: nine are fully correct, but the card contradicts itself on the 1M-token context window, labels a paid model Freemium, misstates the launch gap with GPT-6 Sol, drops the benchmarks GPT-6 Astra wins, and carries a “#4 in LLM models” badge for a slot Grok 4.6 already holds. Here is what the card gets right, what it gets wrong and what the price cut really means.

Read more
deepseek v4 1 flash 552b moe model hugging face a deepseek v4.1 flash stylised whale v2

DeepSeek V4.1 Flash 552B MoE Model Released on Hugging Face. It Beats V4 Pro on Agents and Trails It on Knowledge

DeepSeek released V4.1 Flash on Hugging Face on 10 September 2026: a 552B mixture-of-experts model under the MIT licence that activates 8B parameters to read and 16B to write, stores its global KV cache in 890 bytes per token, and replaces V4 Pro for every API caller from 04:00 UTC on 14 September. Its own model card shows it leading V4 Pro on all 12 agentic rows and trailing it on 12 of 16 base-model rows, including a 12.9-point gap on SimpleQA-Verified. This breakdown covers the causal encoder-decoder architecture, the KV cache arithmetic, the new price sheet with a worked agent bill, the 511 GB checkpoint and 614 GB serving floor, the mismatched reasoning-effort aliases, and a pre-cutover checklist.

Read more
gpt 6 astra aixploria card a rubber stamp with round knob handle

GPT-6 Astra Is Now Available on AIxploria — and Its Card Claims a Rank the Directory Denies

AIxploria listed GPT-6 Astra at 03:35 UTC on 6 September 2026, three days after OpenAI shipped the model, with a gold verification badge, a 4.6 rating and a line reading “#1 in LLM models”. The site’s own LLM models category page does not contain the model on page one or page two. This is a documents-first read of that listing: the seven votes behind the star rating, a measured three-day propagation lag taken against two models we checked on the same pages earlier in the week, and a claim-by-claim comparison of twenty checkable statements against OpenAI’s launch post, model documentation, price list and system card. Sixteen check out exactly. Four do not, and two significant facts are missing.

Read more
gemini 3 8 flash aixploria listing a ballot box lid slot

Gemini 3.8 Flash Is Now Available on AIxploria

AIxploria listing pages are how a very large audience meets a new model for the first time, and Gemini 3.8 Flash got one at 03:16 UTC on 3 September 2026, roughly a day after Google shipped the model itself. The card sits at 4.4 out of 5 stars, carries 84 upvotes, and files Google’s newest […]

Read more
CHAT