AI content moderation has a new kind of tool. On Tuesday 6 October 2026, the trust and safety company Musubi released PolicyLM-1.7B, a small open-weights model that reads a platform’s own content policy and scores every message against it in a fraction of a second. TechCrunch’s AI editor Russell Brandom framed it as the next use for decision models, the class of AI that took off in September when TypeSafe AI released Jev.

The pitch is simple. Today most platforms that run AI content moderation choose between fast, cheap classifiers that cannot read their rules and a large language model, which reads rules well but is too slow and costly for live chat. Musubi says PolicyLM sits between the two: it reads the policy like an LLM and answers at classifier speed.

This article explains what Musubi shipped, how decision models work, what the published benchmarks really show, and where a 1.7-billion-parameter model falls short. It closes with a routing design and a pilot plan for teams that run AI content moderation on a live service, including the UK and EU rules that shape those decisions.

What Musubi Announced for AI Content Moderation

ai content moderation policylm decision model b fanning mill grain winnower with a crank wheel

Musubi published the AI content moderation model with open weights under the Apache 2.0 licence on Hugging Face, alongside a launch post by co-founder and chief AI officer Filip Jankovic. A hosted version runs on Baseten. The model card shows it was uploaded on 2 October, four days before the announcement.

PolicyLM-1.7B in one paragraph

You write your content policy as a short list of categories with plain-English rules, then send a message. The model returns a score from 0 to 1 for every category in a single pass, without generating any text. You set a cutoff per category, and anything above it is flagged. Because the policy travels with every message, a policy edit takes effect on the next request, with no retraining and no new labelled data.

Who Musubi is

Musubi, also known as Musubi Labs, was founded in 2023 by Tom Quisel and Filip Jankovic. Its website describes a trust and safety suite covering AI content moderation, fraud and fake-account detection, AI guardrails and a review console. It lists Bluesky, Grindr, Bumble, Feeld, Muzz, Hornet, Rakuten Viber and Stocktwits among its customers and claims to protect more than 850 million users. Those are the company’s own figures.

Why Musubi frames it as AI content moderation for product teams

Jankovic told TechCrunch that this style of AI content moderation gives platform managers a way to label content proactively. “Product teams just want a better understanding of what’s happening on their platform, especially as the amount of content is exponentially increasing,” he said. Labelling everything “in a very scalable, customizable way is extremely useful.” That framing matters: Musubi is selling insight into all traffic, not only removals.

What a Decision Model Is, and Why It Suits AI Content Moderation

ai content moderation policylm decision model c railway signal box with a glazed lever room

A decision model is built from the body of a transformer language model, but it does not write sentences. It outputs probabilities over a fixed set of choices that the user defines in advance. TechCrunch explains that PolicyLM’s choice is binary for each category: the content is either in it or not.

Decisions instead of text

Limiting the output to known answers removes the slowest part of an LLM, which is generating tokens one at a time. It also makes the result easy to use in software, because a score can be compared with a threshold. For AI content moderation that is exactly the shape of the job: one message in, a yes or no per rule out, thousands of times a second.

The Jev wave that made decision models news

TypeSafe AI, founded by former OpenAI researcher Diogo Almeida, released Jev in mid-September and called its output “calibrated decisions”. We covered the launch in our report on how developers are adopting Jev. OpenAI followed at DevDay on 29 September with a preview Decisions API, which we examined in OpenAI’s Jev clone and agent monitoring. Amazon then open-sourced Strands Decider 2B, built on Qwen3.5-2B.

Older roots than the headlines suggest

Jankovic says his interest predates Jev. He traces it to GLiNER, a 2024 research model for named entity recognition that used many of the same techniques: a compact bidirectional encoder that scores labels you supply at run time. Musubi is still happy to borrow the attention. Its launch post says: “If Jev caught your eye, PolicyLM-1.7B is the same kind of model, trained specifically for content moderation, that you can run yourself.”

Three Ways to Run AI Content Moderation Today

ai content moderation policylm decision model d wastepaper basket overflowing with crumpled paper

Musubi’s own comparison puts PolicyLM between the two tools most teams already use. The table below restates it in plain terms. It is the vendor’s framing, so treat the “best for” row as a claim to test rather than a finding.

Fixed classifiers

Fixed classifiers are fast and cheap, often tens of milliseconds, but they only know the categories they were trained on. Changing a rule means relabelling data and retraining, or waiting for a vendor to update its taxonomy. They are strong on stable, well-defined harms such as spam or known slurs.

Large language models reading the policy

An LLM can read a full written policy and explain its verdict. Musubi notes it usually takes hundreds of milliseconds or more, returns text you must parse, and is hard to threshold. LLMs are the better tool for appeals, bans and novel judgement calls, where a written reason matters more than speed.

Decision models in the middle

PolicyLM reads your categories with every message, returns a 0 to 1 score per category, and takes under 100 milliseconds, according to Musubi. It gives no written reason. That trade, flexibility with speed but no explanation, is the core of how decision models could change AI content moderation.

QuestionFixed classifierLarge language modelPolicyLM-1.7B
How it decidesCategories fixed at trainingWrites a verdict token by tokenScores your categories in one pass
SpeedTens of msHundreds of ms or moreUnder 100 ms
Reads your policy?No, retrain to changeYes, from the promptYes, with every message
ThresholdsPer categoryHard on a text answerPer category, two presets
Written reasonNoYesNo
Where it runsVendor API or in-houseUsually a hosted APIOpen weights, in-house or Musubi
Best for (vendor view)Stable, well-defined harmsAppeals, bans, novel casesEvery live message, DMs, usernames

How PolicyLM Reads a Policy

ai content moderation policylm decision model e trommel drum screen on a roller frame

The model card is unusually detailed about how a policy should be written, and the details matter for anyone planning AI content moderation with it. The policy and the message share a 2,048-token window.

Categories, rules and exceptions

Each category has a name, a violation rule, a “not a violation” rule and an optional exception override. A policy can hold up to 16 categories and 1,662 policy tokens. Musubi advises one category per violation type, with the decisive clause first, and keeping unrelated categories in separate calls, because categories read together affect each other’s scores.

Anchored meanings, not blank labels

Musubi draws a careful line. PolicyLM lets you pick your own labels, but it starts from meanings learned in training and layers your definitions on top. Instructions “can shift scores, but aren’t meant to redefine abuse as support.” Models that let labels be inverted are more flexible, Musubi argues, but easier for user content or an accidental edit to subvert.

Two cutoff presets

The default “precision” cutoff is 0.335 for your own policy and suits live chat, where violations are rare. The “balanced” cutoff is 0.275 and catches more when violations are common or a miss costs more than a false flag. In the built-in taxonomy mode the cutoffs are 0.69 and 0.45. You can also set a number per category.

A built-in taxonomy as a fallback

For generic screening, an “aegis” mode applies 23 harm categories from NVIDIA’s Aegis 2.0 dataset, from violence and hate to fraud, malware and unauthorised advice. Teams can start there, then move to their own categories once they know where the generic list misses their community.

A text cleaner before scoring

The helper code undoes look-alike characters, spaced-out letters, zero-width characters and base64 before scoring. Musubi says this caught 10 to 12 points more disguised violations on its test set at the balanced cutoff, with no rise in false flags. Leetspeak “mostly gets through”, which is a useful admission for any AI content moderation team facing determined evaders.

The Benchmarks Behind This AI Content Moderation Model

ai content moderation policylm decision model f drive through car wash gantry with spinning brushes

Musubi published a head-to-head on its own custom-policy benchmark, built from synthetic business policies. Every model got the same policy text and was timed on the same messages on an NVIDIA H100. The card says every evaluation set except OR-Bench was also used during development, so these are not blind tests.

The custom-policy results

The table reproduces the card’s figures. “Policy edits followed” measures how often a single-clause change to the policy flips the decision the way it should. Latency is the median time until every category has an answer, with six categories.

ModelParametersAccuracyEdits followedMedian latency
gpt-oss-safeguard-20B21.5B (3.6B active)0.9090.716349 ms
PolicyLM-1.7B1.7B0.8420.52822 ms
CoPE-B-A4B25.2B (3.8B active)0.8290.46053 ms
Granite Guardian 4.1-8B8.4B0.7220.188186 ms
Nemotron-3.5-Content-Safety4.3B0.6680.06850 ms
Shieldstral-1.0-3B3.8B0.6430.03044 ms
SingGuard-4B4.4B0.5980.012120 ms
GLiGuard-300M0.2B0.5260.04211 ms

PolicyLM does not win on accuracy

Read the table closely and the headline claim, “beats every other model we ran under 20B parameters”, is accurate but narrow. OpenAI’s gpt-oss-safeguard-20B scores higher on accuracy (0.909 against 0.842) and follows policy edits far better (0.716 against 0.528). It is just much slower: 349 milliseconds against 22, roughly 16 times longer. For AI content moderation on live chat that gap decides the choice; for an appeals queue it may not.

Speed is where it pulls ahead

The chart below plots the card’s median latency per message on the H100. Each bar is the model’s time divided by the slowest result, 349 milliseconds, so 22 divided by 349 gives PolicyLM a bar of 6.3%.

Median latency per message, NVIDIA H100, six categories (Musubi model card)
gpt-oss-safeguard-20B 349 ms
Granite Guardian 4.1-8B 186 ms
SingGuard-4B 120 ms
CoPE-B-A4B 53 ms
Nemotron-3.5-Content-Safety 50 ms
Shieldstral-1.0-3B 44 ms
PolicyLM-1.7B 22 ms
GLiGuard-300M 11 ms

Policy edits followed about half the time

The 0.528 figure deserves attention. The card says plainly that “not every edit changes the decision: about 1 in 2 single-clause edits did on our benchmark, at the balanced cutoff.” For a policy team, that means a rewording in the document may not change outcomes in production. Every edit needs testing on example messages before it goes live.

Generic safety benchmarks

On five public benchmarks, PolicyLM’s AUROC ranges from 0.835 on RabakBench to 0.968 on XSTest. It does not top any column: SingGuard-4B leads three and CoPE-B-A4B leads RabakBench. That is a respectable result for a model a fraction of the size, but it confirms that the selling point of this approach to AI content moderation is the custom policy, not generic safety screening.

Three numbers that do not agree

TechCrunch wrote that PolicyLM applies a policy “in under 50 milliseconds”. Musubi’s own post says “under 100 ms”, with a median of 35 milliseconds on a 24 GB NVIDIA L4, and the card gives 22 milliseconds on an H100. All three can be true on different hardware, but plan on the L4 figure. The blog also says a fine-tuned version runs on a platform handling over a million messages a day, while the card says the released weights are “not yet tested on live traffic”.

Throughput: What Classifier-Speed AI Content Moderation Means

Speed is only useful if it translates into volume. The card publishes batched throughput for three single-GPU set-ups, using short chat messages and the built-in taxonomy.

Messages per second and per day

The table below multiplies each published messages-per-second figure by 86,400 seconds in a day. These are best-case batched numbers, not live-traffic numbers.

One GPU, bfloat16Median per messageMessages per second (batched)Messages per day (x 86,400)
NVIDIA L4 (24 GB)35 ms39.33,395,520
NVIDIA L40S34 ms127.210,990,080
NVIDIA H100 PCIe22 ms134.111,586,240

The same numbers as a chart

Bars are scaled to the H100’s 11,586,240 messages a day, so the L4’s 3,395,520 is 29.3% of the track and the L40S’s 10,990,080 is 94.9%.

Batched messages per day on one GPU (card figures x 86,400)
NVIDIA H100 PCIe 11.59 million
NVIDIA L40S 10.99 million
NVIDIA L4 3.40 million

Live traffic is slower than the batch figures

With messages arriving at random times and a five-category policy, the card says one L40S sustained 132.6 messages per second at a 95th-percentile latency of 150 milliseconds or less, and one L4 sustained 34.4. At 34.4 a second, a single L4 covers about 2.97 million messages a day. A service with a million messages a day averages only 11.6 a second, so peaks, not averages, should drive sizing for AI content moderation.

Score everything instead of sampling

The practical change is that a mid-sized platform can afford to score every message rather than a sample. That turns AI content moderation from a reactive queue into a measurement system: you can see how often each rule is breached, in which features, and how that shifts after a policy change.

The Limits of AI Content Moderation With a 1.7B Model

Musubi is candid about what PolicyLM cannot do, and the list belongs in any procurement note. Several limits are design choices rather than bugs, but they still shape where the model can sit in an AI content moderation workflow.

Text only, one message at a time

The model reads text only and has no conversation history. It cannot see an image, a voice note or the three messages that turned a joke into a threat. Grooming, coordinated harassment and slow-burn scams often only appear across a thread, so a per-message scorer needs a second layer that looks at accounts and patterns.

No written reasons

PolicyLM returns a score, not a rationale. That is fine for routing and dashboards. It is a gap for user-facing decisions, where regulators increasingly expect a reason. We return to this in the regulation section, because it decides where automated AI content moderation can act alone.

Over-flagging and language gaps

The card warns that benign content that sounds harmful, identity mentions under a hate-speech policy, and long messages can raise false flags. It was evaluated in 19 languages, with English strongest and Tamil weakest, and every custom policy tested was written in English. Code-mixed slang and leetspeak remain weak spots for this kind of AI content moderation.

Out of scope by design

Musubi lists four uses the model is not for: child-safety enforcement, which should go to dedicated tooling; being the only self-harm safeguard; acting as a security boundary against adversarial users; and moderating an AI assistant’s own responses. Teams should write those exclusions into their internal policy so nobody quietly stretches the AI content moderation tool.

Routing: Small Models First, Big Models for Appeals

Musubi does not argue that one model should do all of AI content moderation. Its launch post says big models are “great for reasoning through appeals and nuanced policies”, while smaller ones are “great for analyzing and labeling content at scale”. The expected design is a cascade.

The ROOST connection

PolicyLM joins the ROOST Model Community, the open-source safety initiative’s catalogue of open models. To mark the launch, ROOST and Musubi published a guide, “Choosing and Routing Open Safety Models”, covering how to pick a model on accuracy, cost, speed and steerability, and when to cascade from a small model to a larger one. The two are running a workshop on 27 October called “Does a Model Follow Your Rules?”.

A score-band design for AI content moderation

The table shows one way to turn PolicyLM’s two published cutoffs into actions. The cutoffs come from the card; the actions are our suggestion, and every team should calibrate them on its own labelled sample.

Score on a categoryMeaningSuggested action
Below 0.275Under both presetsPublish, keep the score for analytics
0.275 to 0.335Flagged only at “balanced”Route to an LLM or a human reviewer
0.335 and aboveFlagged at the default “precision”Hold or limit reach, then review
Any score, child-safety signalOut of scope for the modelSend to dedicated tooling at once
Any removal or banUser-facing decisionGenerate a written reason before acting

Where agents fit

TechCrunch notes that an early use for decision models is checking AI agents‘ actions, and that “it’s only natural to apply the same technology to human misbehavior.” The same cascade works both ways: a cheap scorer on every message or agent step, with the expensive model reserved for the cases near the line. Our coverage of Jev’s rivals and LLM alternatives tracks how fast that pattern is spreading.

AI Content Moderation Under UK and EU Rules

A model choice is also a compliance choice. Two regimes matter most for UK and European services, and both reward the cascade design above more than a single model acting alone.

The UK Online Safety Act

Ofcom’s illegal content duties took full effect on 17 March 2025, and the children’s safety duties followed on 25 July 2025. Services in scope must assess risks, use proportionate measures to stop users encountering illegal content, and remove it swiftly. Faster AI content moderation helps with “swiftly”, but the risk assessment still has to show why the chosen tools fit your service.

The EU Digital Services Act

Under the DSA, a platform that removes or restricts content must give the user a statement of reasons, including whether automated means were used. A model that returns only a score cannot write that statement. In practice, PolicyLM can triage and hold content, but a reason-writing step, human or LLM, has to sit in front of the user-facing decision.

Data protection still applies

AI content moderation that scores every message means processing every message. Running open weights on your own infrastructure keeps that data in-house, which can simplify a data protection impact assessment compared with sending traffic to a third-party API. It does not remove the need for one, especially where special category data or children’s messages are involved.

How to Pilot AI Content Moderation With PolicyLM

A careful AI content moderation pilot takes a few weeks and protects users while you learn. If you need help structuring one, our AI strategy team works through exactly these steps with clients.

Write the policy as categories

Turn your community guidelines into short categories with a violation rule, a not-a-violation rule and exceptions. Keep each category to one harm and put unrelated categories in separate calls. Use the card’s check_policy helper to confirm you are inside the 16-category and 1,662-token limits.

Build a labelled sample

Collect a few thousand real messages, labelled by your own moderators, with plenty of borderline cases and the languages your users actually write in. Without this you cannot set cutoffs, and the card is explicit that scores shift with wording, language, device and numeric format.

Calibrate the cutoffs

Start from the precision preset, then measure precision and recall per category against your sample. Move each category’s cutoff on its own. Record the numeric format, because the card reports that bfloat16 and float32 runs can disagree on a few decisions.

Run in shadow mode first

Score live traffic without acting on it for two to four weeks, and compare against your current process. This is where AI content moderation earns trust: you see real false-flag rates on real messages before any user is affected.

Re-test after every policy edit

Because only about half of single-clause edits changed decisions in Musubi’s own benchmark, treat each policy change like a code change. Keep a regression set of messages for every category and check that the edit moves the scores you expected before it ships.

AI Content Moderation: Frequently Asked Questions

What is PolicyLM-1.7B?

It is an open-weights decision model for AI content moderation from Musubi, released on 6 October 2026 under Apache 2.0. It reads a content policy and a message together and returns a 0 to 1 score for each category, without generating text.

How is a decision model different from an LLM?

An LLM writes text token by token. A decision model scores a fixed set of choices you define, which makes it faster, cheaper and easy to threshold, but it cannot explain its answer.

Is PolicyLM the most accurate AI content moderation model?

No. On Musubi’s own benchmark, gpt-oss-safeguard-20B scored higher on accuracy and on following policy edits. PolicyLM’s advantage is speed: 22 milliseconds against 349 on an H100.

Can it moderate images or whole conversations?

Not yet. It is text only and scores one message at a time, with no conversation history, so AI content moderation for images and threads needs other tools.

Is it free to use?

The weights are free under Apache 2.0 and run on a single 24 GB GPU or a laptop. You pay for your own hardware, or for the hosted version on Baseten or a managed service from Musubi.

References