DigBat is a free web platform from Tohoku University that has quietly become the largest curated collection of solid-state electrolyte data anywhere, and on 27 July 2026 the team behind it published the paper explaining what it now holds. The short version: 3,816 experimental electrolytes, 27,645 ionic conductivity entries, 852 computational materials, and a set of tools sitting on top that let you interrogate all of it without writing a line of code.

That last part is what makes DigBat interesting beyond materials science. Plenty of research groups publish a dataset. Far fewer publish a dataset that a machine can use — where measurements from hundreds of separate papers share one identifier, one unit convention and one schema, and where a large language model assistant sits on the front to answer questions in plain English. The gap between “we released the data” and “the data is usable by software” is where most scientific datasets die, and DigBat is a rare, well-documented example of a group crossing it.

The subject matter matters too. Solid-state batteries are the technology most of the automotive industry is betting on for the next decade, and the bottleneck is not enthusiasm but knowledge: nobody has a reliable map of which solid electrolytes conduct ions well, at what temperature, and why. If you follow the tooling side of this field through our AI models, tools and releases hub, DigBat is the kind of release that looks small in a news cycle and turns out to matter for years.

What follows is a full account of what the platform contains, how it grew from roughly 600 records to nearly 4,000 in three years, the ion migration models that were added in this release, how the curation was actually done, and the honest limits of what a database like this can and cannot tell you.

What DigBat Actually Is

digbat ai platform solid state battery research b conical lab flask

Start with the plain description, because the name gives away less than it should.

A dynamic database that outgrew its own name

DigBat began life as the DDSE — the Dynamic Database of Solid-State Electrolytes — described in a 2023 paper in Nano Materials Science. That first release covered inorganic materials only, with just over 600 entries. The “dynamic” part was the point: rather than freezing a supplementary spreadsheet alongside a paper, the group committed to updating the resource continuously as new measurements were published. Three years on, DigBat is what that commitment produced.

Three modules under one roof

The curated experimental data in DigBat is organised into three modules on the web interface: Inorganic Solid Electrolytes (ISEs), Solid Polymer Electrolytes (SPEs) and Gel Polymer Electrolytes (GPEs). The 2023 database only covered the first of those. Adding polymers and gels is what turned a specialist inorganic dataset into something closer to a general map of the field, because the three families compete directly on the same design problem and were previously documented in almost completely separate literatures.

Who built DigBat

The platform comes from the Advanced Institute for Materials Research (WPI-AIMR) at Tohoku University in Sendai, and specifically from the group of Hao Li, Distinguished Professor and Principal Investigator there. Li founded the Digital Catalysis and Battery Lab in 2022; DigBat is its battery arm. The July 2026 paper lists ten authors — Qian Wang, Hanghui Liu, Seong Hoon Jang, Di Zhang, Ryuhei Sato, Kazuaki Kisu, Yusuke Hashimoto, Eric Jianfeng Cheng, Shin-ichi Orimo and Li himself — which is a reasonable indication of how much labour sits behind a resource that looks, from the outside, like a website.

What the team says DigBat is for

Li frames it as navigation rather than storage. “A database can be much more than a place to store information,” he said of the release. “We see DigBat as a map for navigating the growing world of solid-state electrolyte research.” That framing is worth taking seriously, because it explains design decisions that would otherwise look like overkill — the shared identifiers, the migration models, the assistant layer. A warehouse needs shelves. A map needs everything on it to be in the same coordinate system.

What Is Inside DigBat Today

digbat ai platform solid state battery research c stack of round discs

Here are the numbers, and then what they actually mean.

What DigBat holdsFigure (June 2026)Note
Experimental solid-state electrolytes3,816Curated from the published literature
Ionic conductivity data entries27,645Multiple measurements per material
Computational materials852Linked to experimental records by identifier
Experimental modules3Inorganic, solid polymer, gel polymer
Temperature range covered132.4–1261.6 KReported in the original DDSE release
Mobile cations represented7Li, Na, K, Ag, Ca, Mg, Zn
Cost to accessFree, open webdigbat.org, no licence fee stated

Roughly seven measurements per material

Divide the 27,645 conductivity entries by the 3,816 experimental materials and DigBat is carrying about 7.2 measurements for every material it lists. That ratio is the whole argument for building this thing. Ionic conductivity is not a single number; it varies with temperature, with synthesis route, with how the pellet was pressed and with which laboratory did the measuring. A database that stored one headline figure per compound would be actively misleading. DigBat stores the spread, which is what lets you see disagreement rather than average it away.

How fast DigBat has grown

The growth curve is steep and recent. Just over 600 materials in the 2023 DDSE paper, roughly 3,000 experimental materials by February 2026, and 3,816 by June 2026. That is a jump of about 27% in four months.

Experimental solid-state electrolytes in DigBat, by snapshot (bars relative to 3,816)
June 2026 3,816
February 2026 ~3,000
2023, DDSE release ~600

The count that went down

One figure moved the other way. The computational materials count was reported at 863 in February 2026 and 852 in June 2026 — eleven fewer, over a period when everything else grew. Neither publication explains the difference, so treat any specific reason as speculation. It is worth noting anyway, because a count that can fall is a count somebody is checking. Most research datasets only ever accumulate.

Beyond lithium

The seven mobile cations in DigBat include divalent species — calcium, magnesium and zinc — that get a fraction of the attention lithium does. That coverage is not decoration. Divalent conductors are chronically under-explored precisely because their literature is thin and scattered, which makes them exactly the case where a curated database changes what questions are askable. The review literature has singled out this coverage as one of the platform’s more consequential contributions.

Why Solid-State Battery Research Needed DigBat

digbat ai platform solid state battery research d arched bridge two spans

To understand why a database counts as news, you have to look at the state it replaced.

The data was scattered across thousands of papers

Ionic conductivity measurements live in the results sections of individual papers, in units the authors chose, at temperatures the authors chose, with sample preparation described in prose. There has never been a central register. A researcher wanting to know whether a given sulfide family outperforms a given halide family at room temperature has historically had to read their way to an answer, which is slow, partial and unrepeatable. DigBat replaces that reading task with a query.

Measurements that could not be compared

Worse than scattered is incomparable. Two laboratories reporting the same compound at different temperatures are not disagreeing, but a naive reading makes them look like they are. Activation energy — the barrier an ion has to clear to hop between sites — is what makes measurements at different temperatures commensurable, and DigBat carries it alongside conductivity. That single editorial decision is what turns a pile of numbers into something you can fit a model to.

The commercialisation clock is running

This is not a leisurely field. Toyota and Samsung SDI have both publicly targeted 2027 for solid-state cells, and QuantumScape has moved from B-sample lithium-metal cells into equipment installation on its Eagle production line, with its PowerCo partnership behind it. Whether those dates hold is a separate argument; what is not in doubt is that money is being committed now against materials choices that are still open questions. A resource like DigBat is worth more at this moment than it would have been in 2019 or will be in 2032.

PlayerPublicly stated positionWhy DigBat is relevant
ToyotaTargeting a solid-state EV around 2027Electrolyte choice still drives cell design
Samsung SDIMass production target of 2027Sulfide chemistry, heavily represented in the data
QuantumScapeQSE-5 cells, Eagle line equipment from Feb 2026Separator and interface behaviour remain open
Academic groupsThousands of papers a year, no shared registerThe gap DigBat was built to close

What data scarcity costs a model

Machine learning on materials fails for a specific and boring reason: there is not enough clean labelled data. A model trained on a few hundred inconsistent records will happily produce confident predictions that mean nothing. The value of DigBat is not that it is clever; it is that it is large, consistent and honest about provenance, which is the only foundation on which a predictive model is worth training at all.

The Ion Migration Models Behind DigBat

digbat ai platform solid state battery research e upright ladder five rungs

The July 2026 release is not just more rows. The subtitle of the paper — “extending the dynamic database of solid-state electrolytes to a diversified electrolyte database with ion migration models” — names the actual novelty.

Why activation energy is the number that matters

Conductivity tells you how a material performed in one measurement. Activation energy tells you something closer to why. It is the energy barrier a mobile ion must overcome to move from one site in the crystal to the next, and it governs how the material will behave at temperatures nobody has yet tested. A database that carries both, across 3,816 materials and a temperature range from 132.4 K to 1261.6 K, is a database you can extrapolate from rather than merely look things up in.

Where the standard simulation methods fall short

Computational chemistry has two workhorses for ion migration: climbing-image nudged elastic band (CI-NEB), which finds the lowest-energy path between two known positions, and ab initio molecular dynamics (AIMD), which simulates the atoms moving. Both assume, in effect, that you already know roughly what the migration looks like. When the real mechanism is unusual, both can give you a confident wrong answer.

MethodWhat it does wellWhere it struggles
CI-NEBCheap barriers along an assumed pathYou must guess the path first
AIMDRealistic thermal motion, no path assumedShort timescales; rare events go unseen
Ab initio metadynamicsForces rare events to happen; finds unexpected routesExpensive; needs careful setup
ML interatomic potentialsNear-DFT accuracy at a fraction of the costOnly as good as the training set

The two-step mechanism in hydride conductors

The clearest demonstration of that failure mode came from the same Tohoku circle. Coupling the database with metadynamics simulations uncovered a two-step ion migration mechanism in hydride conductors, in which molecular groups mediate an unconventional hopping pathway. The review literature is explicit that ab initio metadynamics gave a more reliable description of cation transport there than CI-NEB or AIMD alone. Embedding migration models in DigBat is a direct response to that result: the mechanism is part of the record, not a separate paper you have to find.

One identifier, two kinds of evidence

The structural change that makes all of this work is unglamorous. DigBat assigns a common material identifier that links experimental and computational records for the same compound, so a conductivity measurement and a simulated migration barrier can be examined side by side instead of as two unrelated facts. Anyone who has tried to join two scientific datasets on a chemical formula string will recognise how much work that one decision represents.

Universal potentials change the economics

Alongside the metadynamics work sits a newer class of tool: universal machine learning interatomic potentials such as CHGNet, M3GNet and MACE, which approximate density functional theory at a small fraction of the compute. They make it plausible to simulate thousands of candidate materials rather than dozens. They are also entirely dependent on the quality of the data they were fitted to, which loops straight back to why a curated resource like DigBat matters.

How DigBat Makes Battery Data AI-Ready

digbat ai platform solid state battery research f three interlocking rings

“AI-ready” is a phrase that usually means nothing. Here it means four specific things.

Curation is the expensive part, and it was done by hand and machine together

Turning 27,645 measurements scattered across the literature into comparable rows is not a scripting job. Units differ, temperatures differ, sample preparation is described in prose, and a substantial share of published records do not even state clearly whether a measured conductivity is ionic or electronic. Automated pipelines built on natural language processing have been reported to cut the share of unclassified ionic-conductivity records in this literature from 93% to 24.3% — a reduction of 68.7 percentage points, or roughly three-quarters of the original ambiguity.

Share of ionic-conductivity records left unclassified, before and after automated extraction
Before extraction 93.0%
After extraction 24.3%
Ambiguity removed 68.7 points

The assistant on the front

DigBat ships with a large language model assistant, so a researcher can ask a question in ordinary English instead of composing a query. That is a smaller feature than it sounds and a bigger one than it looks. Smaller, because the assistant does not do the science. Bigger, because the practical barrier to using most research databases is not the data but the interface, and an assistant that can translate “which sodium conductors beat 10⁻³ S/cm at room temperature” into a filter removes an afternoon of learning.

Machine-readable, not merely human-readable

The distinction that does most of the work here is between a dataset a person can read and a dataset software can consume without a human in the loop. Shared identifiers, consistent units, explicit temperatures and structured provenance are what put DigBat in the second category. Every one of them is dull. Together they are the difference between a resource that gets cited and a resource that gets used.

PropertyRaw literatureDigBat
UnitsAuthor’s choiceNormalised
TemperatureOften implicitExplicit, 132.4–1261.6 K
Ionic vs electronicFrequently unstatedClassified
Experiment ↔ simulationSeparate papersJoined by one identifier
Query interfaceReadingFilters plus an LLM assistant
Update cadenceFrozen on publicationContinuous

Interpretability was a design goal, not an afterthought

Li is blunt about the point of all this. “What matters is not only whether a model can make a good prediction, but whether we can understand what stands behind that prediction,” he said. Storing migration mechanisms next to measurements is how that principle shows up in the schema. A model that predicts high conductivity is a lottery ticket; a model that predicts it and points at a migration pathway is a hypothesis somebody can test in a laboratory next week.

Where DigBat Sits In The Digital Materials Ecosystem

It is not the only materials database, and the comparison is instructive.

General-purpose repositories versus specialist ones

The Materials Project, OQMD, AFLOW and NOMAD are the giants of computational materials data, and they are broad by design — hundreds of thousands of density functional theory calculations across the whole periodic table. Breadth is their strength and their limitation. None of them was built to answer “which of these actually conducts sodium ions at 300 K, and did anyone measure it or only calculate it?” DigBat is narrow and experimental-first, which is precisely the axis those repositories leave open.

ResourcePrimary contentScope
Materials ProjectComputed structures, energetics, electronic structureBroad inorganic
OQMDAround 300,000 DFT calculationsBroad inorganic
AFLOWHigh-throughput first-principles calculationsBroad inorganic
NOMADLarge open repository of computed materials dataBroad, multi-code
DigCat>0.4M experimental, >0.3M computational catalysis recordsCatalysis
DigBat3,816 experimental electrolytes, 27,645 conductivity entriesSolid-state electrolytes

DigCat, DigBat and DigHyd are one method applied three times

The Tohoku group runs three platforms on the same pattern: DigCat for catalysis, DigBat for batteries, DigHyd for hydrogen storage, under the umbrella of a Digital Materials Lab. DigCat is by far the largest, with more than 400,000 experimental and 300,000 computational records, and it functions as proof that the approach scales when the underlying literature is bigger. The three-way split is a useful signal for anyone building something similar: the method is domain-agnostic, the curation is not.

Where autonomous laboratories fit

The eventual destination for all of this is a closed loop — a system that reads the literature, proposes a candidate, synthesises it and feeds the result back. Elements already exist. A-Lab, an autonomous synthesis platform, ran 355 experiments in 17 days and succeeded on 71% of its targets, an average of roughly 21 experiments a day with no human at the bench. Screening campaigns have gone further still, sifting over 10 million candidate compositions in one halide study. What has been missing is the trustworthy, curated substrate those loops read from, and that is the role a platform like DigBat plays.

The agent framing

Researchers in the same orbit have started describing this explicitly in terms of AI agents — systems with perception, reasoning, action and learning, rather than a static predictive model. Eric Jianfeng Cheng, an Associate Professor at AIMR and a co-author on the DigBat paper, put the appeal plainly: “AI agents allow us to move from isolated predictions to coordinated, multi-step research strategies that evolve as new information becomes available.” A database with stable identifiers and machine-readable provenance is the precondition for any of that working.

Concrete results already on the board

Agent-assisted screening has produced named candidates, not just methodology papers. One campaign explored 4,375 hypothetical sodium argyrodite structures and identified Na₆SiS₄Cl₂ with a theoretical room-temperature conductivity of 2.9 × 10⁻² S cm⁻¹. Another screened over 10⁷ compositions and experimentally validated Na₂LiYCl₆. Neither result would have been reachable by reading papers one at a time.

What DigBat Means Outside The Laboratory

You are probably not synthesising electrolytes. The pattern still transfers, and it transfers well.

The asset is the curation, not the algorithm

Every organisation with a decade of operational history is sitting on the equivalent of scattered literature: measurements in inconsistent units, in systems that were never designed to talk to each other, described in free text by people who have since left. The lesson of DigBat is that the expensive, valuable, defensible work is normalising that into something machine-readable. Models are commodities and get cheaper every quarter. A clean, identifier-joined domain dataset does not.

Shared identifiers are the whole game

If there is one thing to copy from DigBat, it is the common identifier that joins experimental and computational records. In a commercial setting that is a customer key that survives across billing, support and product telemetry, or a part number that means the same thing in engineering and in the warehouse. It is unglamorous plumbing that nobody gets promoted for, and it determines whether analytics and machine learning are possible at all.

Publish the disagreements

DigBat stores about 7.2 measurements per material rather than one averaged figure, which preserves the fact that laboratories disagree. Most internal reporting does the opposite, collapsing variation into a single number before anyone can see it. Keeping the spread is more work and more honest, and it is what allows a model to learn uncertainty instead of pretending it does not exist.

DigBat design choiceCommercial equivalentWhat it unlocks
One material identifier across sourcesMaster data key across systemsJoins that do not need a human
Explicit measurement conditionsRecorded context on every metricComparability across periods
Multiple records per entityKeep raw events, not just totalsUncertainty and drift become visible
Continuous updatingPipelines, not annual extractsDecisions on current data
Assistant over the query layerNatural-language access for non-analystsWider internal adoption

Open by default has compounding returns

DigBat is free to use on the open web, which is why it is cited, corrected and extended by people who will never visit Sendai. The equivalent internal move is making the curated dataset available across the business rather than fencing it inside the team that built it. The cost is governance work. The return is that everybody stops maintaining their own private spreadsheet of the same numbers.

What DigBat Does Not Do

A fair account has to include the limits, and there are real ones.

A database is not a discovery

DigBat does not invent materials. It narrows a search space and makes hypotheses cheaper to form. Every candidate it helps surface still has to be synthesised, characterised and cycled, and most will fail on something the database never held — mechanical stability against a lithium metal anode, interfacial resistance after a hundred cycles, or the plain question of whether the compound can be manufactured at scale.

Literature bias goes in and comes out

Any resource curated from published work inherits publication bias. Materials that performed poorly are under-reported. Negative results are largely absent. Fashionable chemistries are over-sampled relative to their promise. A model trained on DigBat will reproduce those biases faithfully, and no amount of curation quality fixes a distortion that happened before the data was written down.

A moving target is hard to cite

The “dynamic” property is a genuine benefit with a genuine cost: a query run today may not reproduce next year. Anyone using DigBat in published analysis needs to record which snapshot they used. That is a solved problem in principle — versioned releases, dated exports — but it is a discipline the user has to supply, and the difference between the February and June 2026 figures shows how quickly the ground moves.

Coverage is uneven by construction

Three modules do not mean three equally deep modules. The inorganic module descends from the original DDSE and has had years of curation; the polymer and gel modules are newer. Similarly, lithium and sodium are far better represented than calcium, magnesium or zinc, because that is how the field has spent its time. Absence of data in DigBat is evidence about the literature, not about the chemistry.

Your Next Steps

If you work on batteries, the practical move is small: open digbat.org, run a query you already know the answer to, and see whether the platform reproduces it. That is the cheapest possible test of whether a curated resource is trustworthy, and it tells you more than any paper about it will.

If you do not work on batteries, take the structural lesson instead. Look at one dataset your organisation depends on and ask the DigBat questions of it. Is there a single identifier that joins it to everything else? Are the conditions under which each figure was recorded still attached to that figure? Can software read it without a person interpreting a column heading first? If the answer to any of those is no, that is the work — and it is the same work whether the subject is ionic conductivity or invoice lines.

Either way, the release worth watching next is not a bigger row count. It is whether the assistant layer turns into something that proposes experiments rather than answering questions about ones already done.

Frequently Asked Questions

What is DigBat?

DigBat is the Digital Battery Platform, a free web resource from the Advanced Institute for Materials Research at Tohoku University that curates solid-state electrolyte data. As of June 2026 it holds 3,816 experimental electrolytes, 27,645 ionic conductivity entries and 852 computational materials, with tools for simulation, machine learning and an LLM-based assistant on top.

Is DigBat free to use?

Yes. DigBat is published on the open web at digbat.org with no stated licence fee, in keeping with the group’s other platforms, DigCat for catalysis and DigHyd for hydrogen storage.

What is the difference between DigBat and DDSE?

DDSE — the Dynamic Database of Solid-State Electrolytes, published in 2023 with just over 600 entries — is the direct ancestor. DigBat is the expanded platform: more than six times the materials, three electrolyte families instead of one, linked computational records, and the ion migration models added in the July 2026 release.

Which materials does DigBat cover?

Inorganic solid electrolytes, solid polymer electrolytes and gel polymer electrolytes, spanning seven mobile cations — lithium, sodium, potassium, silver, calcium, magnesium and zinc — with measurements reported across a temperature range of 132.4 to 1261.6 K.

Can DigBat predict a new battery material?

Not on its own. It supplies the curated, machine-readable substrate that predictive models and screening workflows need, and it stores migration mechanisms so a prediction can be interrogated rather than merely trusted. Synthesis and testing still decide the outcome.

Where was the DigBat research published?

In Nano Materials Science on 27 July 2026, under the title “Digital battery platform: extending the dynamic database of solid-state electrolytes to a diversified electrolyte database with ion migration models”, DOI 10.1016/j.nanoms.2026.07.001.

References