DeepSeek

long ai conversations misinformation vulnerabilities seven chatbots a speaker cabinet with two round cones

Long AI Conversations Reveal Misinformation Vulnerabilities Across Seven Leading Chatbots

A University of Arizona team put seven widely used chatbots through 50-turn sequences of sustained misinformation pressure and published the results in Nature’s Scientific Reports. Misinformation affirmation rates ranged from 0.08% to 12.3%, a greater than 150-fold spread across architectures, with GPT-3.5 most vulnerable and Claude 3.5 Sonnet most resistant. The paper also names a new failure mode, conversational reverberation, in which a model oscillates between accepting and rejecting the same false statement across successive turns. This is a close reading of what was tested, what the numbers mean, the correctability dissociation that should change how you pick a model, and the controls that actually target each failure mode in production.

Read more
Z.ai - z ai lab behind ox alpha model a treasure chest closed lid

Surprise: Z.ai Is the AI Lab Behind the Mysterious Ox Alpha Model

On 26 August 2026 the mystery ended: Z.ai, the Beijing lab formerly known as Zhipu AI, confirmed that the anonymous Ox Alpha model topping OpenRouter and OpenCode was the newest iteration of its GLM series, and published the weights the same evening as GLM-5.3-Flash under an MIT licence. This breakdown covers what was confirmed and when, the architecture the model card revealed — 320 billion total parameters with just 18 billion active, hybrid sparse and linear attention, a 1,048,576-token context and forced reasoning that cannot be disabled — the published benchmark table showing 84.3 on Terminal Bench 2.1 against 85.0 for Claude Opus 4.8 and 87.4 for GPT-5.6 Terra, the viral 80 per cent DeepSWE claim that came from a 10-task subset and collapsed to 63.4 on the full 113-task run, the 44 trillion tokens and 503,000 users the stealth week generated, the tokenizer and error-code forensics that unmasked the lab before it spoke, the $0.15 and $0.50 per million token pricing, the company’s Hong Kong listing and US entity-list status, and a buyer’s checklist for deciding between the hosted API and self-hosted weights.

Read more
deepseek v4 complete guide a three ascending rounded pillars

DeepSeek V4 Complete Guide: Best Open-Weight AI of 2026

DeepSeek V4 is the MIT-licensed open-weight family that replaced the never-released R2, pairing a one-million-token context window with sub-dollar output pricing. This complete guide covers the V4 Pro 0813 and V4 Flash 0731 GA builds, their official benchmarks, the peak/off-peak billing change landing on 16 August 2026, and a practical framework for choosing between the two models.

Read more
CHAT