Quantum-inspired math from the University of Miami offers a way to flag the moment a chatbot starts answering beyond what it reliably knows. The method does not stop a large language model from making things up. It gives the person reading the answer a measured warning that the answer deserves a second look, which is the thing every “ChatGPT can make mistakes” disclaimer promises and never delivers.

The work comes from Kamal Premaratne, a professor of electrical and computer engineering, doctoral candidate Pragatheeswaran Vipulanandan and associate professor of computer science Dilip Sarkar. It was accepted as a poster at the International Conference on Learning Representations (ICLR 2026) and is on arXiv as Semantic Uncertainty Quantification of Hallucinations in LLMs: A Quantum Tensor Network Based Method. Tech Xplore republished the university’s summary on 30 September 2026.

This article explains what the quantum-inspired math does, checks the paper’s worked example number by number, sets out its limits and says what it means for organisations using AI. It builds on our explainer on how AI models decide not to answer a question.

What the Quantum-Inspired Math Research Found

quantum-inspired math - quantum inspired math ai knows when it does not know b unicycle leaning on its wheel

The short version: the quantum-inspired math produces a score that says how shaky a chatbot’s answer is, and in their tests that score separated right answers from wrong ones at least as well as the best existing methods, without any labelled training data.

The team and the venue

The preprint was submitted to arXiv on 27 January 2026. OpenReview lists it as an ICLR 2026 poster, and the team presented it at the conference in Rio de Janeiro earlier this year. The code is public on GitHub. The university’s own write-up, by Lorena Taboas, reached a wider audience on 30 September, eight months after the preprint first appeared.

The problem it targets

The paper focuses on one kind of error, which it calls confabulation: “fluent yet unreliable outputs that vary arbitrarily even under identical prompts”. Ask the same question five times and get three different confident answers, and you are looking at confabulation. It is a narrower target than all hallucination. A model that repeats a widespread myth consistently is wrong in a different way, and this quantum-inspired math is not designed to catch it.

A warning light, not a cure

The researchers are careful about what the quantum-inspired math can do. “The approach is not intended to eliminate AI hallucinations,” the summary says. “Its purpose is to provide a clearer signal when an AI-generated response may be unreliable and should receive a second look.” Vipulanandan describes the goal as a model that says “I’m not sure, but these are the top answers I have,” or asks for more context before it replies.

Why Chatbots Answer Even When They Don't Know

quantum inspired math ai knows when it does not know c open top safari jeep with a spare wheel

Language models do not look answers up. They build a reply one token at a time, picking likely continuations. That design is why it rarely says “I don’t know”.

Built to produce the most probable answer

“Large language models are built to generate the most probable response to a prompt,” the Tech Xplore summary explains. “Even when they lack reliable information, they may produce an answer that sounds convincing rather than acknowledge uncertainty.” The probabilities behind each word exist inside the model. They are simply not shown to the user, and a fluent sentence hides how close the model came to saying something else.

Where it costs most

Premaratne names the fields where that gap matters. “In certain applications like medicine, health care and defense, you don’t know whether the answer that the AI system is giving is correct or not,” he said. “Even a slight oversight in the output could be severe and very costly.” The paper’s introduction adds law, journalism and autonomous driving to that list.

Why a disclaimer is not enough

A blanket warning on every answer is the same warning on the good answers and the bad ones, so people learn to ignore it. A useful signal has to vary: quiet when the model is on firm ground, loud when it is guessing. That is the gap the quantum-inspired math is trying to fill, and it is the same gap our report on misleading AI-generated summaries found hurting readers.

How AI Uncertainty Is Measured Today

quantum inspired math ai knows when it does not know d wind chime hanging from a stand

The quantum-inspired math builds on an established method, so it helps to know the baseline it is trying to beat.

Ask the same question several times

The best-known approach is semantic entropy, published in Nature in 2024 by Sebastian Farquhar and colleagues. You ask a model the same question repeatedly, group the answers by meaning, and measure how spread out the groups are. “Paris” and “It’s Paris” land in the same group. If every answer lands in one group, the model is consistent. If they scatter, it is probably guessing. The Miami paper uses a DeBERTa model trained for textual entailment to do the grouping.

Other ways to score doubt

Two further baselines appear in the paper. The first, p(True), asks the model whether its own top answer is true and reads the probability it gives to “True”. The second, embedding regression, trains a small classifier on the model’s internal state to predict correctness. It works well on familiar data and poorly on anything new.

MethodWhat it measuresNeeds labelled data?
Naive entropy (NE)Spread of raw token-sequence probabilitiesNo
Semantic entropy (SE)Spread of answers grouped by meaningNo
Discrete semantic entropySame, counting group sizes onlyNo
p(True)Model’s own rating of its top answerA few examples
Embedding regressionClassifier on hidden statesYes
SRE-UQ (this paper)Rényi semantic entropy after an uncertainty-weighted adjustmentNo

What the baselines miss

The paper’s argument is that all of these treat the model’s probabilities as fixed facts. “Hallucination risk should depend on how sensitive these TS probabilities, and hence the semantic entropies they induce, are to model perturbations,” the authors write. In plain terms: two answers with the same probability can be very different bets if one would collapse under a tiny nudge and the other would not.

Inside the Quantum-Inspired Math

quantum inspired math ai knows when it does not know e ambulance with a roof light bar

This is where the physics behind the quantum-inspired math comes in. The team is not using a quantum computer. As the summary puts it, they have “adapted mathematical tools developed to measure uncertainty in complex systems”.

Answer probabilities as a wave function

The quantum-inspired math takes the probabilities of the sampled answers and turns them into a smooth curve using a standard statistical technique called a kernel mean embedding. It then treats that curve as the wave function of a small quantum system, known as a quantum tensor network, and works out the system that would produce it. This framing comes from José Príncipe’s work on information-theoretic learning and from the team’s earlier papers on time series and regression networks.

Nudge it and watch the wobble

Once the answers are described as a quantum system, the team applies perturbation theory, the textbook physics tool for asking what happens when a system is disturbed slightly. “Large first-order corrections, therefore, indicate highly unstable TS probabilities; small corrections indicate locally stable regions,” the paper says. The method reads those corrections across eight neighbouring modes and averages them into one uncertainty score per answer.

Lift the answers that deserve doubt

The score is then used to adjust the answer probabilities. Where uncertainty is high, probabilities are pushed towards a more even spread; where it is low, they stay close to what the model produced. A single setting, lambda, controls the trade-off. The logic is the maximum entropy principle: when you only know part of the picture, assume as little as you can about the rest.

Why Rényi entropy

Most prior work, including the Nature paper, measures spread with Shannon entropy. This quantum-inspired math uses the quadratic form of Rényi entropy instead, because the kernel framework it borrows from estimates that quantity directly. The choice matters more than it sounds, as the worked example below shows.

The Worked Example: Which Oil Producer Is a Close US Ally?

quantum inspired math ai knows when it does not know f roulette wheel with the ball in a pocket

The paper walks through one question, asked ten times: “Which oil producer is a close ally of the United States?” The model gave six distinct answers. Saudi Arabia came up five times; Russia, Iran, Kuwait, Qatar and Iraq once each. The table below uses the cluster probabilities from the paper’s Table 1, before and after the uncertainty adjustment. The last column is our arithmetic.

AnswerTimes seenBeforeAfterChange
Saudi Arabia5 of 100.888240.85880−3.3%
Iraq1 of 100.037170.04498+21.0%
Kuwait1 of 100.027490.03488+26.9%
Iran1 of 100.022230.02697+21.3%
Russia1 of 100.018140.02223+22.5%
Qatar1 of 100.006720.01214+80.7%

What the adjustment actually did

Read the numbers and the pattern is clear. The adjustment took about three points of probability away from the dominant answer and spread them across the five answers the model gave only once. Qatar, the least likely answer, gained the most in relative terms, rising by four-fifths. That is the quantum-inspired math doing what it is designed to do: refusing to let one repeated answer look more certain than the evidence supports.

Relative change in each answer’s probability after the adjustment (paper’s Table 1, our arithmetic)

Qatar: +80.7%
Kuwait: +26.9%
Russia: +22.5%
Iran: +21.3%
Iraq: +21.0%
Saudi Arabia: −3.3%

A sentence the table does not support

The paper’s prose says that “Saudi Arabia, a contextually appropriate response, receives a higher probability assignment after adjustment”. Its own table shows the opposite: Saudi Arabia’s cluster probability falls from 0.888 to 0.859. The direction in the table is the one the method’s design predicts, so the sentence looks like a slip in the text rather than a flaw in the approach. It is still worth knowing if you plan to rely on the worked example.

Why “lower entropy” needs care

The paper reports that its measure “yields systematically lower entropy estimates” than the older ones. In the table, that compares Shannon entropy before the adjustment (0.225) with Rényi entropy after it (0.130). Rényi entropy of this order is never higher than Shannon entropy for the same probabilities, so part of the drop comes from switching yardsticks. Measured like for like on the table’s own figures, the adjustment raised uncertainty: Shannon from 0.225 to 0.271, and Rényi from 0.101 to 0.130.

Entropy of the worked example, base-10 logs (paper’s figures and our recalculation from its Table 1)

Naive entropy (paper): 0.846
Shannon, after adjustment (ours): 0.271
Shannon semantic entropy, before (paper): 0.225
Rényi, after adjustment (paper): 0.130
Rényi, before adjustment (ours): 0.101

The Leopard Problem: Answers the Model Never Gave

The most memorable part of the university’s explanation is about answers that never appear at all.

A wildlife park analogy

“If you go to a park to see animals, you come out knowing how many giraffes, how many hippos and how many rhinos there are,” Premaratne said. “But the fact that you didn’t see a leopard doesn’t mean there are no leopards in the park. It’s just that you didn’t see one.” Ten samples from a chatbot are a short walk through a large park. Plenty of possible answers may simply not have come up.

The missing-species idea

The summary says the team uses “a statistical technique known as the missing-species model”. The classic tool for this, the Good–Turing estimator, came out of Alan Turing’s wartime codebreaking. Its rule of thumb is that the chance the next observation is something new equals the share of observations seen exactly once. Applied as an illustration to the oil question, five of the ten answers were singletons, which puts the chance of an unseen eleventh answer at about 50%.

Where the paper and the press release differ

That illustration is ours, and the distinction matters. The ICLR paper itself never uses the phrase “missing species”. Its mechanism for unseen possibilities is the maximum entropy step, which the authors describe as reasoning “under partial knowledge”. The press release appears to describe the team’s wider research programme as well as this paper. Anyone evaluating the quantum-inspired math should read the preprint, not only the summary.

What 116 Experiments Show About Quantum-Inspired Math

The team ran 116 experiments across four question sets: TriviaQA, SQuAD 1.1, the open-domain Natural Questions set and SVAMP maths word problems. They used eight open models from 1 billion to 13 billion parameters, including Llama 2, Llama 3.2, Mistral 7B and Falcon.

Output type16-bit8-bit4-bitTotal
Sentence-length answers12202052
Short answers, instruction-tuned models88824
Short answers, base models16121240
All experiments364040116

Win rates rather than one headline score

The main results are presented as win-rate grids: for each pair of methods, how often one beat the other on AUROC, a standard measure of how well a score separates right answers from wrong ones. The authors report that their quantum-inspired math is “competitive” with the best baselines on AUROC and “consistently” strong when the least certain answers are filtered out. It does this without labelled data, which the supervised embedding-regression baseline needs.

Compressed models hold up

Real deployments usually shrink models to 8-bit or 4-bit precision to save memory and cost. The paper tested all three precisions and found that its quantum-inspired math kept its ranking, while simpler measures such as naive entropy lost ground on the harder question sets. Very few earlier studies had checked this at all, and it is arguably the most practical result in the paper.

The danger zone between certain and lost

One finding is useful even without the new method. Across Llama 2 variants, the biggest swings sat in the middle of the entropy range, between 0.25 and 0.50, “a high-risk region where models frequently oscillate between multiple confident yet semantically divergent answers”. The lesson for anyone setting thresholds is that one cut-off for every question is brittle.

The Cost of Running Quantum-Inspired Math

A warning signal is only useful if it is cheap enough to run on every answer.

Set-up once, then milliseconds

According to the appendix, building the underlying quantum system takes 45 to 60 seconds on a CPU, once. After that, the quantum-inspired math adds 6 to 10 milliseconds per query on an Nvidia A6000 graphics card. The core matrix is 256 by 256, or 65,536 entries, small by modern standards.

The hidden cost is sampling

The quantum-inspired math itself is described as “deterministic, one-shot”. But the pipeline still starts by asking the model the same question repeatedly, ten times in the worked example, and grouping the answers. Ten generations cost roughly ten times the decoding of one. For a busy service that is the bill to plan around, not the milliseconds.

The Limits of Quantum-Inspired Math

The authors list the main limits themselves, which is to their credit.

Small models only, so far

Tests of the quantum-inspired math stop at 13 billion parameters. “They may not fully capture the behavior of larger frontier models such as GPT-4 or Claude,” the paper says. Whether the same signal holds on today’s much larger systems is untested.

It needs to see the probabilities

The method “requires access to token-level probabilities, which restricts its applicability to open-weight models”. Many commercial chatbots do not expose those numbers in full, so a business cannot simply bolt this quantum-inspired math onto a closed assistant it rents.

It inherits the grouping model’s errors

Answers are grouped by meaning with a separate entailment model. If that model wrongly decides two answers mean the same thing, or different things, the error flows straight into the uncertainty score. The paper acknowledges this dependence directly.

What Quantum-Inspired Math Means for Businesses Using AI

Few organisations will implement this paper themselves. The ideas behind it are still immediately useful for anyone deploying AI where mistakes cost money.

Treat confidence as a routing signal

The practical value of any uncertainty score, including this quantum-inspired math, is triage. Answers above a threshold go through; answers below it go to a person. The paper’s finding about the volatile middle band suggests setting separate rules for low, middle and high uncertainty rather than one cut-off.

Ask vendors the right questions

When you buy an AI tool for customer service, finance or clinical support, ask how it detects low-confidence answers, whether it samples more than once, and whether you can see a confidence signal. If the answer is “the model is very accurate”, keep asking. Our AI models and tools hub tracks which vendors publish evaluation detail.

Keep human judgment where it counts

Premaratne ends on a point that no quantum-inspired math can replace. “The logical thinking comes when you actually know how to do the work on your own,” he said, “not by simply parroting whatever ChatGPT is saying.” A confidence flag helps a skilled reviewer decide where to look. It does not make an unskilled one safe.

Quantum-Inspired Math FAQ

What is the quantum-inspired math approach to AI uncertainty?

It is a method from the University of Miami that treats a chatbot’s answer probabilities like a small quantum system, nudges that system mathematically, and measures how unstable the probabilities are. Unstable answers get flagged as less reliable.

Does it need a quantum computer?

No. The quantum-inspired math borrows equations from quantum physics and runs on ordinary hardware. The paper’s experiments used a workstation with one Nvidia A6000 graphics card.

Does it stop AI hallucinations?

No. The researchers say the quantum-inspired math is meant to signal when an answer “may be unreliable and should receive a second look”, not to prevent errors.

Can I use it with ChatGPT or Claude?

Not directly. The quantum-inspired math needs token-level probabilities, which the paper says limits it to open-weight models. It was also tested only on models up to 13 billion parameters.

Where was the research published?

It was accepted as a poster at ICLR 2026 and is available on arXiv as paper 2601.20026, with code on GitHub.

References and Further Reading