AI content moderation has a new kind of tool. On Tuesday 6 October 2026, the trust and safety company Musubi released PolicyLM-1.7B, a small open-weights model that reads a platform’s own content policy and scores every message against it in a fraction of a second. TechCrunch’s AI editor Russell Brandom framed it as the next use for decision models, the class of AI that took off in September when TypeSafe AI released Jev.
The pitch is simple. Today most platforms that run AI content moderation choose between fast, cheap classifiers that cannot read their rules and a large language model, which reads rules well but is too slow and costly for live chat. Musubi says PolicyLM sits between the two: it reads the policy like an LLM and answers at classifier speed.
This article explains what Musubi shipped, how decision models work, what the published benchmarks really show, and where a 1.7-billion-parameter model falls short. It closes with a routing design and a pilot plan for teams that run AI content moderation on a live service, including the UK and EU rules that shape those decisions.
Table of contents
- What Musubi Announced for AI Content Moderation
- What a Decision Model Is, and Why It Suits AI Content Moderation
- Three Ways to Run AI Content Moderation Today
- How PolicyLM Reads a Policy
- The Benchmarks Behind This AI Content Moderation Model
- Throughput: What Classifier-Speed AI Content Moderation Means
- The Limits of AI Content Moderation With a 1.7B Model
- Routing: Small Models First, Big Models for Appeals
- AI Content Moderation Under UK and EU Rules
- How to Pilot AI Content Moderation With PolicyLM
- AI Content Moderation: Frequently Asked Questions
- References
What Musubi Announced for AI Content Moderation
Musubi published the AI content moderation model with open weights under the Apache 2.0 licence on Hugging Face, alongside a launch post by co-founder and chief AI officer Filip Jankovic. A hosted version runs on Baseten. The model card shows it was uploaded on 2 October, four days before the announcement.
PolicyLM-1.7B in one paragraph
You write your content policy as a short list of categories with plain-English rules, then send a message. The model returns a score from 0 to 1 for every category in a single pass, without generating any text. You set a cutoff per category, and anything above it is flagged. Because the policy travels with every message, a policy edit takes effect on the next request, with no retraining and no new labelled data.
Who Musubi is
Musubi, also known as Musubi Labs, was founded in 2023 by Tom Quisel and Filip Jankovic. Its website describes a trust and safety suite covering AI content moderation, fraud and fake-account detection, AI guardrails and a review console. It lists Bluesky, Grindr, Bumble, Feeld, Muzz, Hornet, Rakuten Viber and Stocktwits among its customers and claims to protect more than 850 million users. Those are the company’s own figures.
Why Musubi frames it as AI content moderation for product teams
Jankovic told TechCrunch that this style of AI content moderation gives platform managers a way to label content proactively. “Product teams just want a better understanding of what’s happening on their platform, especially as the amount of content is exponentially increasing,” he said. Labelling everything “in a very scalable, customizable way is extremely useful.” That framing matters: Musubi is selling insight into all traffic, not only removals.
What a Decision Model Is, and Why It Suits AI Content Moderation
A decision model is built from the body of a transformer language model, but it does not write sentences. It outputs probabilities over a fixed set of choices that the user defines in advance. TechCrunch explains that PolicyLM’s choice is binary for each category: the content is either in it or not.
Decisions instead of text
Limiting the output to known answers removes the slowest part of an LLM, which is generating tokens one at a time. It also makes the result easy to use in software, because a score can be compared with a threshold. For AI content moderation that is exactly the shape of the job: one message in, a yes or no per rule out, thousands of times a second.
The Jev wave that made decision models news
TypeSafe AI, founded by former OpenAI researcher Diogo Almeida, released Jev in mid-September and called its output “calibrated decisions”. We covered the launch in our report on how developers are adopting Jev. OpenAI followed at DevDay on 29 September with a preview Decisions API, which we examined in OpenAI’s Jev clone and agent monitoring. Amazon then open-sourced Strands Decider 2B, built on Qwen3.5-2B.
Older roots than the headlines suggest
Jankovic says his interest predates Jev. He traces it to GLiNER, a 2024 research model for named entity recognition that used many of the same techniques: a compact bidirectional encoder that scores labels you supply at run time. Musubi is still happy to borrow the attention. Its launch post says: “If Jev caught your eye, PolicyLM-1.7B is the same kind of model, trained specifically for content moderation, that you can run yourself.”
Three Ways to Run AI Content Moderation Today
Musubi’s own comparison puts PolicyLM between the two tools most teams already use. The table below restates it in plain terms. It is the vendor’s framing, so treat the “best for” row as a claim to test rather than a finding.
Fixed classifiers
Fixed classifiers are fast and cheap, often tens of milliseconds, but they only know the categories they were trained on. Changing a rule means relabelling data and retraining, or waiting for a vendor to update its taxonomy. They are strong on stable, well-defined harms such as spam or known slurs.
Large language models reading the policy
An LLM can read a full written policy and explain its verdict. Musubi notes it usually takes hundreds of milliseconds or more, returns text you must parse, and is hard to threshold. LLMs are the better tool for appeals, bans and novel judgement calls, where a written reason matters more than speed.
Decision models in the middle
PolicyLM reads your categories with every message, returns a 0 to 1 score per category, and takes under 100 milliseconds, according to Musubi. It gives no written reason. That trade, flexibility with speed but no explanation, is the core of how decision models could change AI content moderation.
| Question | Fixed classifier | Large language model | PolicyLM-1.7B |
|---|---|---|---|
| How it decides | Categories fixed at training | Writes a verdict token by token | Scores your categories in one pass |
| Speed | Tens of ms | Hundreds of ms or more | Under 100 ms |
| Reads your policy? | No, retrain to change | Yes, from the prompt | Yes, with every message |
| Thresholds | Per category | Hard on a text answer | Per category, two presets |
| Written reason | No | Yes | No |
| Where it runs | Vendor API or in-house | Usually a hosted API | Open weights, in-house or Musubi |
| Best for (vendor view) | Stable, well-defined harms | Appeals, bans, novel cases | Every live message, DMs, usernames |
How PolicyLM Reads a Policy
The model card is unusually detailed about how a policy should be written, and the details matter for anyone planning AI content moderation with it. The policy and the message share a 2,048-token window.
Categories, rules and exceptions
Each category has a name, a violation rule, a “not a violation” rule and an optional exception override. A policy can hold up to 16 categories and 1,662 policy tokens. Musubi advises one category per violation type, with the decisive clause first, and keeping unrelated categories in separate calls, because categories read together affect each other’s scores.
Anchored meanings, not blank labels
Musubi draws a careful line. PolicyLM lets you pick your own labels, but it starts from meanings learned in training and layers your definitions on top. Instructions “can shift scores, but aren’t meant to redefine abuse as support.” Models that let labels be inverted are more flexible, Musubi argues, but easier for user content or an accidental edit to subvert.
Two cutoff presets
The default “precision” cutoff is 0.335 for your own policy and suits live chat, where violations are rare. The “balanced” cutoff is 0.275 and catches more when violations are common or a miss costs more than a false flag. In the built-in taxonomy mode the cutoffs are 0.69 and 0.45. You can also set a number per category.
A built-in taxonomy as a fallback
For generic screening, an “aegis” mode applies 23 harm categories from NVIDIA’s Aegis 2.0 dataset, from violence and hate to fraud, malware and unauthorised advice. Teams can start there, then move to their own categories once they know where the generic list misses their community.
A text cleaner before scoring
The helper code undoes look-alike characters, spaced-out letters, zero-width characters and base64 before scoring. Musubi says this caught 10 to 12 points more disguised violations on its test set at the balanced cutoff, with no rise in false flags. Leetspeak “mostly gets through”, which is a useful admission for any AI content moderation team facing determined evaders.
The Benchmarks Behind This AI Content Moderation Model
Musubi published a head-to-head on its own custom-policy benchmark, built from synthetic business policies. Every model got the same policy text and was timed on the same messages on an NVIDIA H100. The card says every evaluation set except OR-Bench was also used during development, so these are not blind tests.
The custom-policy results
The table reproduces the card’s figures. “Policy edits followed” measures how often a single-clause change to the policy flips the decision the way it should. Latency is the median time until every category has an answer, with six categories.
| Model | Parameters | Accuracy | Edits followed | Median latency |
|---|---|---|---|---|
| gpt-oss-safeguard-20B | 21.5B (3.6B active) | 0.909 | 0.716 | 349 ms |
| PolicyLM-1.7B | 1.7B | 0.842 | 0.528 | 22 ms |
| CoPE-B-A4B | 25.2B (3.8B active) | 0.829 | 0.460 | 53 ms |
| Granite Guardian 4.1-8B | 8.4B | 0.722 | 0.188 | 186 ms |
| Nemotron-3.5-Content-Safety | 4.3B | 0.668 | 0.068 | 50 ms |
| Shieldstral-1.0-3B | 3.8B | 0.643 | 0.030 | 44 ms |
| SingGuard-4B | 4.4B | 0.598 | 0.012 | 120 ms |
| GLiGuard-300M | 0.2B | 0.526 | 0.042 | 11 ms |
PolicyLM does not win on accuracy
Read the table closely and the headline claim, “beats every other model we ran under 20B parameters”, is accurate but narrow. OpenAI’s gpt-oss-safeguard-20B scores higher on accuracy (0.909 against 0.842) and follows policy edits far better (0.716 against 0.528). It is just much slower: 349 milliseconds against 22, roughly 16 times longer. For AI content moderation on live chat that gap decides the choice; for an appeals queue it may not.
Speed is where it pulls ahead
The chart below plots the card’s median latency per message on the H100. Each bar is the model’s time divided by the slowest result, 349 milliseconds, so 22 divided by 349 gives PolicyLM a bar of 6.3%.
Policy edits followed about half the time
The 0.528 figure deserves attention. The card says plainly that “not every edit changes the decision: about 1 in 2 single-clause edits did on our benchmark, at the balanced cutoff.” For a policy team, that means a rewording in the document may not change outcomes in production. Every edit needs testing on example messages before it goes live.
Generic safety benchmarks
On five public benchmarks, PolicyLM’s AUROC ranges from 0.835 on RabakBench to 0.968 on XSTest. It does not top any column: SingGuard-4B leads three and CoPE-B-A4B leads RabakBench. That is a respectable result for a model a fraction of the size, but it confirms that the selling point of this approach to AI content moderation is the custom policy, not generic safety screening.
Three numbers that do not agree
TechCrunch wrote that PolicyLM applies a policy “in under 50 milliseconds”. Musubi’s own post says “under 100 ms”, with a median of 35 milliseconds on a 24 GB NVIDIA L4, and the card gives 22 milliseconds on an H100. All three can be true on different hardware, but plan on the L4 figure. The blog also says a fine-tuned version runs on a platform handling over a million messages a day, while the card says the released weights are “not yet tested on live traffic”.
Throughput: What Classifier-Speed AI Content Moderation Means
Speed is only useful if it translates into volume. The card publishes batched throughput for three single-GPU set-ups, using short chat messages and the built-in taxonomy.
Messages per second and per day
The table below multiplies each published messages-per-second figure by 86,400 seconds in a day. These are best-case batched numbers, not live-traffic numbers.
| One GPU, bfloat16 | Median per message | Messages per second (batched) | Messages per day (x 86,400) |
|---|---|---|---|
| NVIDIA L4 (24 GB) | 35 ms | 39.3 | 3,395,520 |
| NVIDIA L40S | 34 ms | 127.2 | 10,990,080 |
| NVIDIA H100 PCIe | 22 ms | 134.1 | 11,586,240 |
The same numbers as a chart
Bars are scaled to the H100’s 11,586,240 messages a day, so the L4’s 3,395,520 is 29.3% of the track and the L40S’s 10,990,080 is 94.9%.
Live traffic is slower than the batch figures
With messages arriving at random times and a five-category policy, the card says one L40S sustained 132.6 messages per second at a 95th-percentile latency of 150 milliseconds or less, and one L4 sustained 34.4. At 34.4 a second, a single L4 covers about 2.97 million messages a day. A service with a million messages a day averages only 11.6 a second, so peaks, not averages, should drive sizing for AI content moderation.
Score everything instead of sampling
The practical change is that a mid-sized platform can afford to score every message rather than a sample. That turns AI content moderation from a reactive queue into a measurement system: you can see how often each rule is breached, in which features, and how that shifts after a policy change.
The Limits of AI Content Moderation With a 1.7B Model
Musubi is candid about what PolicyLM cannot do, and the list belongs in any procurement note. Several limits are design choices rather than bugs, but they still shape where the model can sit in an AI content moderation workflow.
Text only, one message at a time
The model reads text only and has no conversation history. It cannot see an image, a voice note or the three messages that turned a joke into a threat. Grooming, coordinated harassment and slow-burn scams often only appear across a thread, so a per-message scorer needs a second layer that looks at accounts and patterns.
No written reasons
PolicyLM returns a score, not a rationale. That is fine for routing and dashboards. It is a gap for user-facing decisions, where regulators increasingly expect a reason. We return to this in the regulation section, because it decides where automated AI content moderation can act alone.
Over-flagging and language gaps
The card warns that benign content that sounds harmful, identity mentions under a hate-speech policy, and long messages can raise false flags. It was evaluated in 19 languages, with English strongest and Tamil weakest, and every custom policy tested was written in English. Code-mixed slang and leetspeak remain weak spots for this kind of AI content moderation.
Out of scope by design
Musubi lists four uses the model is not for: child-safety enforcement, which should go to dedicated tooling; being the only self-harm safeguard; acting as a security boundary against adversarial users; and moderating an AI assistant’s own responses. Teams should write those exclusions into their internal policy so nobody quietly stretches the AI content moderation tool.
Routing: Small Models First, Big Models for Appeals
Musubi does not argue that one model should do all of AI content moderation. Its launch post says big models are “great for reasoning through appeals and nuanced policies”, while smaller ones are “great for analyzing and labeling content at scale”. The expected design is a cascade.
The ROOST connection
PolicyLM joins the ROOST Model Community, the open-source safety initiative’s catalogue of open models. To mark the launch, ROOST and Musubi published a guide, “Choosing and Routing Open Safety Models”, covering how to pick a model on accuracy, cost, speed and steerability, and when to cascade from a small model to a larger one. The two are running a workshop on 27 October called “Does a Model Follow Your Rules?”.
A score-band design for AI content moderation
The table shows one way to turn PolicyLM’s two published cutoffs into actions. The cutoffs come from the card; the actions are our suggestion, and every team should calibrate them on its own labelled sample.
| Score on a category | Meaning | Suggested action |
|---|---|---|
| Below 0.275 | Under both presets | Publish, keep the score for analytics |
| 0.275 to 0.335 | Flagged only at “balanced” | Route to an LLM or a human reviewer |
| 0.335 and above | Flagged at the default “precision” | Hold or limit reach, then review |
| Any score, child-safety signal | Out of scope for the model | Send to dedicated tooling at once |
| Any removal or ban | User-facing decision | Generate a written reason before acting |
Where agents fit
TechCrunch notes that an early use for decision models is checking AI agents‘ actions, and that “it’s only natural to apply the same technology to human misbehavior.” The same cascade works both ways: a cheap scorer on every message or agent step, with the expensive model reserved for the cases near the line. Our coverage of Jev’s rivals and LLM alternatives tracks how fast that pattern is spreading.
AI Content Moderation Under UK and EU Rules
A model choice is also a compliance choice. Two regimes matter most for UK and European services, and both reward the cascade design above more than a single model acting alone.
The UK Online Safety Act
Ofcom’s illegal content duties took full effect on 17 March 2025, and the children’s safety duties followed on 25 July 2025. Services in scope must assess risks, use proportionate measures to stop users encountering illegal content, and remove it swiftly. Faster AI content moderation helps with “swiftly”, but the risk assessment still has to show why the chosen tools fit your service.
The EU Digital Services Act
Under the DSA, a platform that removes or restricts content must give the user a statement of reasons, including whether automated means were used. A model that returns only a score cannot write that statement. In practice, PolicyLM can triage and hold content, but a reason-writing step, human or LLM, has to sit in front of the user-facing decision.
Data protection still applies
AI content moderation that scores every message means processing every message. Running open weights on your own infrastructure keeps that data in-house, which can simplify a data protection impact assessment compared with sending traffic to a third-party API. It does not remove the need for one, especially where special category data or children’s messages are involved.
How to Pilot AI Content Moderation With PolicyLM
A careful AI content moderation pilot takes a few weeks and protects users while you learn. If you need help structuring one, our AI strategy team works through exactly these steps with clients.
Write the policy as categories
Turn your community guidelines into short categories with a violation rule, a not-a-violation rule and exceptions. Keep each category to one harm and put unrelated categories in separate calls. Use the card’s check_policy helper to confirm you are inside the 16-category and 1,662-token limits.
Build a labelled sample
Collect a few thousand real messages, labelled by your own moderators, with plenty of borderline cases and the languages your users actually write in. Without this you cannot set cutoffs, and the card is explicit that scores shift with wording, language, device and numeric format.
Calibrate the cutoffs
Start from the precision preset, then measure precision and recall per category against your sample. Move each category’s cutoff on its own. Record the numeric format, because the card reports that bfloat16 and float32 runs can disagree on a few decisions.
Run in shadow mode first
Score live traffic without acting on it for two to four weeks, and compare against your current process. This is where AI content moderation earns trust: you see real false-flag rates on real messages before any user is affected.
Re-test after every policy edit
Because only about half of single-clause edits changed decisions in Musubi’s own benchmark, treat each policy change like a code change. Keep a regression set of messages for every category and check that the edit moves the scores you expected before it ships.
AI Content Moderation: Frequently Asked Questions
What is PolicyLM-1.7B?
It is an open-weights decision model for AI content moderation from Musubi, released on 6 October 2026 under Apache 2.0. It reads a content policy and a message together and returns a 0 to 1 score for each category, without generating text.
How is a decision model different from an LLM?
An LLM writes text token by token. A decision model scores a fixed set of choices you define, which makes it faster, cheaper and easy to threshold, but it cannot explain its answer.
Is PolicyLM the most accurate AI content moderation model?
No. On Musubi’s own benchmark, gpt-oss-safeguard-20B scored higher on accuracy and on following policy edits. PolicyLM’s advantage is speed: 22 milliseconds against 349 on an H100.
Can it moderate images or whole conversations?
Not yet. It is text only and scores one message at a time, with no conversation history, so AI content moderation for images and threads needs other tools.
Is it free to use?
The weights are free under Apache 2.0 and run on a single 24 GB GPU or a laptop. You pay for your own hardware, or for the hosted version on Baseten or a managed service from Musubi.
References
How AI decision models could change content moderation (TechCrunch)
Introducing PolicyLM-1.7B (Musubi)
musubilabs/policylm-1.7b model card (Hugging Face)
Choosing and Routing Open Safety Models (ROOST)
A new kind of AI model from a ChatGPT inventor is thrilling developers (TechCrunch)
Amazon releases its own Jev clone as decision models flood the web (TechCrunch)
Aegis 2.0: a diverse AI safety dataset and risks taxonomy (arXiv)
More AI coverage: explore Progressive Robot's AI Models, Tools & Releases hub — hands-on reviews, setup guides and benchmarks in one place.