Quantum-inspired math from the University of Miami offers a way to flag the moment a chatbot starts answering beyond what it reliably knows. The method does not stop a large language model from making things up. It gives the person reading the answer a measured warning that the answer deserves a second look, which is the thing every “ChatGPT can make mistakes” disclaimer promises and never delivers.
The work comes from Kamal Premaratne, a professor of electrical and computer engineering, doctoral candidate Pragatheeswaran Vipulanandan and associate professor of computer science Dilip Sarkar. It was accepted as a poster at the International Conference on Learning Representations (ICLR 2026) and is on arXiv as Semantic Uncertainty Quantification of Hallucinations in LLMs: A Quantum Tensor Network Based Method. Tech Xplore republished the university’s summary on 30 September 2026.
This article explains what the quantum-inspired math does, checks the paper’s worked example number by number, sets out its limits and says what it means for organisations using AI. It builds on our explainer on how AI models decide not to answer a question.
Table of contents
- What the Quantum-Inspired Math Research Found
- Why Chatbots Answer Even When They Don’t Know
- How AI Uncertainty Is Measured Today
- Inside the Quantum-Inspired Math
- The Worked Example: Which Oil Producer Is a Close US Ally?
- The Leopard Problem: Answers the Model Never Gave
- What 116 Experiments Show About Quantum-Inspired Math
- The Cost of Running Quantum-Inspired Math
- The Limits of Quantum-Inspired Math
- What Quantum-Inspired Math Means for Businesses Using AI
- Quantum-Inspired Math FAQ
- References and Further Reading
What the Quantum-Inspired Math Research Found
The short version: the quantum-inspired math produces a score that says how shaky a chatbot’s answer is, and in their tests that score separated right answers from wrong ones at least as well as the best existing methods, without any labelled training data.
The team and the venue
The preprint was submitted to arXiv on 27 January 2026. OpenReview lists it as an ICLR 2026 poster, and the team presented it at the conference in Rio de Janeiro earlier this year. The code is public on GitHub. The university’s own write-up, by Lorena Taboas, reached a wider audience on 30 September, eight months after the preprint first appeared.
The problem it targets
The paper focuses on one kind of error, which it calls confabulation: “fluent yet unreliable outputs that vary arbitrarily even under identical prompts”. Ask the same question five times and get three different confident answers, and you are looking at confabulation. It is a narrower target than all hallucination. A model that repeats a widespread myth consistently is wrong in a different way, and this quantum-inspired math is not designed to catch it.
A warning light, not a cure
The researchers are careful about what the quantum-inspired math can do. “The approach is not intended to eliminate AI hallucinations,” the summary says. “Its purpose is to provide a clearer signal when an AI-generated response may be unreliable and should receive a second look.” Vipulanandan describes the goal as a model that says “I’m not sure, but these are the top answers I have,” or asks for more context before it replies.
Why Chatbots Answer Even When They Don't Know
Language models do not look answers up. They build a reply one token at a time, picking likely continuations. That design is why it rarely says “I don’t know”.
Built to produce the most probable answer
“Large language models are built to generate the most probable response to a prompt,” the Tech Xplore summary explains. “Even when they lack reliable information, they may produce an answer that sounds convincing rather than acknowledge uncertainty.” The probabilities behind each word exist inside the model. They are simply not shown to the user, and a fluent sentence hides how close the model came to saying something else.
Where it costs most
Premaratne names the fields where that gap matters. “In certain applications like medicine, health care and defense, you don’t know whether the answer that the AI system is giving is correct or not,” he said. “Even a slight oversight in the output could be severe and very costly.” The paper’s introduction adds law, journalism and autonomous driving to that list.
Why a disclaimer is not enough
A blanket warning on every answer is the same warning on the good answers and the bad ones, so people learn to ignore it. A useful signal has to vary: quiet when the model is on firm ground, loud when it is guessing. That is the gap the quantum-inspired math is trying to fill, and it is the same gap our report on misleading AI-generated summaries found hurting readers.
How AI Uncertainty Is Measured Today
The quantum-inspired math builds on an established method, so it helps to know the baseline it is trying to beat.
Ask the same question several times
The best-known approach is semantic entropy, published in Nature in 2024 by Sebastian Farquhar and colleagues. You ask a model the same question repeatedly, group the answers by meaning, and measure how spread out the groups are. “Paris” and “It’s Paris” land in the same group. If every answer lands in one group, the model is consistent. If they scatter, it is probably guessing. The Miami paper uses a DeBERTa model trained for textual entailment to do the grouping.
Other ways to score doubt
Two further baselines appear in the paper. The first, p(True), asks the model whether its own top answer is true and reads the probability it gives to “True”. The second, embedding regression, trains a small classifier on the model’s internal state to predict correctness. It works well on familiar data and poorly on anything new.
| Method | What it measures | Needs labelled data? |
|---|---|---|
| Naive entropy (NE) | Spread of raw token-sequence probabilities | No |
| Semantic entropy (SE) | Spread of answers grouped by meaning | No |
| Discrete semantic entropy | Same, counting group sizes only | No |
| p(True) | Model’s own rating of its top answer | A few examples |
| Embedding regression | Classifier on hidden states | Yes |
| SRE-UQ (this paper) | Rényi semantic entropy after an uncertainty-weighted adjustment | No |
What the baselines miss
The paper’s argument is that all of these treat the model’s probabilities as fixed facts. “Hallucination risk should depend on how sensitive these TS probabilities, and hence the semantic entropies they induce, are to model perturbations,” the authors write. In plain terms: two answers with the same probability can be very different bets if one would collapse under a tiny nudge and the other would not.
Inside the Quantum-Inspired Math
This is where the physics behind the quantum-inspired math comes in. The team is not using a quantum computer. As the summary puts it, they have “adapted mathematical tools developed to measure uncertainty in complex systems”.
Answer probabilities as a wave function
The quantum-inspired math takes the probabilities of the sampled answers and turns them into a smooth curve using a standard statistical technique called a kernel mean embedding. It then treats that curve as the wave function of a small quantum system, known as a quantum tensor network, and works out the system that would produce it. This framing comes from José Príncipe’s work on information-theoretic learning and from the team’s earlier papers on time series and regression networks.
Nudge it and watch the wobble
Once the answers are described as a quantum system, the team applies perturbation theory, the textbook physics tool for asking what happens when a system is disturbed slightly. “Large first-order corrections, therefore, indicate highly unstable TS probabilities; small corrections indicate locally stable regions,” the paper says. The method reads those corrections across eight neighbouring modes and averages them into one uncertainty score per answer.
Lift the answers that deserve doubt
The score is then used to adjust the answer probabilities. Where uncertainty is high, probabilities are pushed towards a more even spread; where it is low, they stay close to what the model produced. A single setting, lambda, controls the trade-off. The logic is the maximum entropy principle: when you only know part of the picture, assume as little as you can about the rest.
Why Rényi entropy
Most prior work, including the Nature paper, measures spread with Shannon entropy. This quantum-inspired math uses the quadratic form of Rényi entropy instead, because the kernel framework it borrows from estimates that quantity directly. The choice matters more than it sounds, as the worked example below shows.
The Worked Example: Which Oil Producer Is a Close US Ally?
The paper walks through one question, asked ten times: “Which oil producer is a close ally of the United States?” The model gave six distinct answers. Saudi Arabia came up five times; Russia, Iran, Kuwait, Qatar and Iraq once each. The table below uses the cluster probabilities from the paper’s Table 1, before and after the uncertainty adjustment. The last column is our arithmetic.
| Answer | Times seen | Before | After | Change |
|---|---|---|---|---|
| Saudi Arabia | 5 of 10 | 0.88824 | 0.85880 | −3.3% |
| Iraq | 1 of 10 | 0.03717 | 0.04498 | +21.0% |
| Kuwait | 1 of 10 | 0.02749 | 0.03488 | +26.9% |
| Iran | 1 of 10 | 0.02223 | 0.02697 | +21.3% |
| Russia | 1 of 10 | 0.01814 | 0.02223 | +22.5% |
| Qatar | 1 of 10 | 0.00672 | 0.01214 | +80.7% |
What the adjustment actually did
Read the numbers and the pattern is clear. The adjustment took about three points of probability away from the dominant answer and spread them across the five answers the model gave only once. Qatar, the least likely answer, gained the most in relative terms, rising by four-fifths. That is the quantum-inspired math doing what it is designed to do: refusing to let one repeated answer look more certain than the evidence supports.
Relative change in each answer’s probability after the adjustment (paper’s Table 1, our arithmetic)
A sentence the table does not support
The paper’s prose says that “Saudi Arabia, a contextually appropriate response, receives a higher probability assignment after adjustment”. Its own table shows the opposite: Saudi Arabia’s cluster probability falls from 0.888 to 0.859. The direction in the table is the one the method’s design predicts, so the sentence looks like a slip in the text rather than a flaw in the approach. It is still worth knowing if you plan to rely on the worked example.
Why “lower entropy” needs care
The paper reports that its measure “yields systematically lower entropy estimates” than the older ones. In the table, that compares Shannon entropy before the adjustment (0.225) with Rényi entropy after it (0.130). Rényi entropy of this order is never higher than Shannon entropy for the same probabilities, so part of the drop comes from switching yardsticks. Measured like for like on the table’s own figures, the adjustment raised uncertainty: Shannon from 0.225 to 0.271, and Rényi from 0.101 to 0.130.
Entropy of the worked example, base-10 logs (paper’s figures and our recalculation from its Table 1)
The Leopard Problem: Answers the Model Never Gave
The most memorable part of the university’s explanation is about answers that never appear at all.
A wildlife park analogy
“If you go to a park to see animals, you come out knowing how many giraffes, how many hippos and how many rhinos there are,” Premaratne said. “But the fact that you didn’t see a leopard doesn’t mean there are no leopards in the park. It’s just that you didn’t see one.” Ten samples from a chatbot are a short walk through a large park. Plenty of possible answers may simply not have come up.
The missing-species idea
The summary says the team uses “a statistical technique known as the missing-species model”. The classic tool for this, the Good–Turing estimator, came out of Alan Turing’s wartime codebreaking. Its rule of thumb is that the chance the next observation is something new equals the share of observations seen exactly once. Applied as an illustration to the oil question, five of the ten answers were singletons, which puts the chance of an unseen eleventh answer at about 50%.
Where the paper and the press release differ
That illustration is ours, and the distinction matters. The ICLR paper itself never uses the phrase “missing species”. Its mechanism for unseen possibilities is the maximum entropy step, which the authors describe as reasoning “under partial knowledge”. The press release appears to describe the team’s wider research programme as well as this paper. Anyone evaluating the quantum-inspired math should read the preprint, not only the summary.
What 116 Experiments Show About Quantum-Inspired Math
The team ran 116 experiments across four question sets: TriviaQA, SQuAD 1.1, the open-domain Natural Questions set and SVAMP maths word problems. They used eight open models from 1 billion to 13 billion parameters, including Llama 2, Llama 3.2, Mistral 7B and Falcon.
| Output type | 16-bit | 8-bit | 4-bit | Total |
|---|---|---|---|---|
| Sentence-length answers | 12 | 20 | 20 | 52 |
| Short answers, instruction-tuned models | 8 | 8 | 8 | 24 |
| Short answers, base models | 16 | 12 | 12 | 40 |
| All experiments | 36 | 40 | 40 | 116 |
Win rates rather than one headline score
The main results are presented as win-rate grids: for each pair of methods, how often one beat the other on AUROC, a standard measure of how well a score separates right answers from wrong ones. The authors report that their quantum-inspired math is “competitive” with the best baselines on AUROC and “consistently” strong when the least certain answers are filtered out. It does this without labelled data, which the supervised embedding-regression baseline needs.
Compressed models hold up
Real deployments usually shrink models to 8-bit or 4-bit precision to save memory and cost. The paper tested all three precisions and found that its quantum-inspired math kept its ranking, while simpler measures such as naive entropy lost ground on the harder question sets. Very few earlier studies had checked this at all, and it is arguably the most practical result in the paper.
The danger zone between certain and lost
One finding is useful even without the new method. Across Llama 2 variants, the biggest swings sat in the middle of the entropy range, between 0.25 and 0.50, “a high-risk region where models frequently oscillate between multiple confident yet semantically divergent answers”. The lesson for anyone setting thresholds is that one cut-off for every question is brittle.
The Cost of Running Quantum-Inspired Math
A warning signal is only useful if it is cheap enough to run on every answer.
Set-up once, then milliseconds
According to the appendix, building the underlying quantum system takes 45 to 60 seconds on a CPU, once. After that, the quantum-inspired math adds 6 to 10 milliseconds per query on an Nvidia A6000 graphics card. The core matrix is 256 by 256, or 65,536 entries, small by modern standards.
The hidden cost is sampling
The quantum-inspired math itself is described as “deterministic, one-shot”. But the pipeline still starts by asking the model the same question repeatedly, ten times in the worked example, and grouping the answers. Ten generations cost roughly ten times the decoding of one. For a busy service that is the bill to plan around, not the milliseconds.
The Limits of Quantum-Inspired Math
The authors list the main limits themselves, which is to their credit.
Small models only, so far
Tests of the quantum-inspired math stop at 13 billion parameters. “They may not fully capture the behavior of larger frontier models such as GPT-4 or Claude,” the paper says. Whether the same signal holds on today’s much larger systems is untested.
It needs to see the probabilities
The method “requires access to token-level probabilities, which restricts its applicability to open-weight models”. Many commercial chatbots do not expose those numbers in full, so a business cannot simply bolt this quantum-inspired math onto a closed assistant it rents.
It inherits the grouping model’s errors
Answers are grouped by meaning with a separate entailment model. If that model wrongly decides two answers mean the same thing, or different things, the error flows straight into the uncertainty score. The paper acknowledges this dependence directly.
What Quantum-Inspired Math Means for Businesses Using AI
Few organisations will implement this paper themselves. The ideas behind it are still immediately useful for anyone deploying AI where mistakes cost money.
Treat confidence as a routing signal
The practical value of any uncertainty score, including this quantum-inspired math, is triage. Answers above a threshold go through; answers below it go to a person. The paper’s finding about the volatile middle band suggests setting separate rules for low, middle and high uncertainty rather than one cut-off.
Ask vendors the right questions
When you buy an AI tool for customer service, finance or clinical support, ask how it detects low-confidence answers, whether it samples more than once, and whether you can see a confidence signal. If the answer is “the model is very accurate”, keep asking. Our AI models and tools hub tracks which vendors publish evaluation detail.
Keep human judgment where it counts
Premaratne ends on a point that no quantum-inspired math can replace. “The logical thinking comes when you actually know how to do the work on your own,” he said, “not by simply parroting whatever ChatGPT is saying.” A confidence flag helps a skilled reviewer decide where to look. It does not make an unskilled one safe.
Quantum-Inspired Math FAQ
What is the quantum-inspired math approach to AI uncertainty?
It is a method from the University of Miami that treats a chatbot’s answer probabilities like a small quantum system, nudges that system mathematically, and measures how unstable the probabilities are. Unstable answers get flagged as less reliable.
Does it need a quantum computer?
No. The quantum-inspired math borrows equations from quantum physics and runs on ordinary hardware. The paper’s experiments used a workstation with one Nvidia A6000 graphics card.
Does it stop AI hallucinations?
No. The researchers say the quantum-inspired math is meant to signal when an answer “may be unreliable and should receive a second look”, not to prevent errors.
Can I use it with ChatGPT or Claude?
Not directly. The quantum-inspired math needs token-level probabilities, which the paper says limits it to open-weight models. It was also tested only on models up to 13 billion parameters.
Where was the research published?
It was accepted as a poster at ICLR 2026 and is available on arXiv as paper 2601.20026, with code on GitHub.
References and Further Reading
ICLR 2026 poster listing on OpenReview
Tech Xplore: Quantum-inspired math could help AI recognize when it does not know the answer
Farquhar et al., Detecting hallucinations in large language models using semantic entropy (Nature)
Kadavath et al., Language Models (Mostly) Know What They Know
More AI coverage: explore Progressive Robot's AI Models, Tools & Releases hub — hands-on reviews, setup guides and benchmarks in one place.