For platforms, marketplaces and communities

Content moderation services

Most moderation problems are not classifier problems. They are queue design, policy drift and staffing decisions made once and never revisited, and the supplier you pick locks in all three.

Moderation gets bought late, usually after an incident, and the shape of the purchase is set by whoever is loudest in the room that week. That is how platforms end up with a single vendor contract covering four different jobs that needed four different answers.

Worth separating them first, because the supplier question has no sensible answer until the queue question does.

4distinct jobs usually bought as one
$10–$35hourly, depending almost entirely on language
0illegal-material queues that belong in a distributed pool

Four jobs, not one

These differ in latency tolerance, in who is qualified to do them, and in what happens when they are wrong. A contract that prices them identically has already made a mistake.

Pre-publication review

Nothing goes live until somebody clears it. Highest cost, highest latency, and the only option where the harm genuinely does not occur. Suits small volumes and high stakes — marketplace listings, verified profiles, anything where publication is difficult to undo.

Reactive queues

Items reach a human because a user reported them or a classifier flagged them. The bulk of the work at most platforms, and the queue where throughput targets do the most damage, because the composition changes as the automated layer improves.

Appeals

A decision is contested. This needs a different person from the one who decided originally, more time per item, and someone who can read the policy rather than apply it by pattern. Under-resourcing appeals is the most common way a moderation system loses the trust of its own users.

Sampling for measurement

Reviewing a random subset against a known answer to find out how accurate the rest of it is. Not a queue anybody clears, and the first thing cut under pressure, which is why so many platforms cannot say what their error rate is.

Only the second and fourth distribute well. Appeals need consistency and accountability that a rotating pool does not provide. Pre-publication review is usually low enough in volume that an internal team handles it better.

A platform that cannot say what its moderation error rate is has not been measuring it, and the reason is almost always that sampling was the first thing cut.

The three ways to buy it

Each is right for something. The failure is picking on price and discovering the constraint eighteen months later.

StructureSuitsReal constraint
Internal teamPolicy work, appeals, escalationSlow to scale, hard to cover time zones
BPO vendorSteady high volume, defined processContracted capacity, limited language spread
Distributed poolMarket-specific judgment, surge coverageNeeds the task to be genuinely decomposable

An internal team is where the policy should live regardless of what else you buy. Somebody has to own what the rules mean, decide the cases nobody anticipated, and answer for the decisions publicly. That does not outsource, and platforms that have tried it discover during their first serious incident that nobody in the building can explain their own enforcement.

A managed vendor is the right answer for a large, steady, always-on English queue. The economics are good, the process can be trained, and the operational maturity of the better providers is genuinely hard to reproduce.

Distribution answers the case neither of the others handles well: judgment that depends on being from somewhere. Whether a phrase reads as a threat or a joke in Brazilian Portuguese. Whether an image carries a meaning in one market that it does not carry in another. Whether a claim about a local regulation is plausible. Country targeting explains how to recruit that audience without turning a digital task into an in-person job.

What actually degrades moderation quality

Consistently, across platform sizes, and almost none of it is about the people doing the work.

Throughput targets set from the wrong sample. Somebody times a batch of easy items, derives a rate, and applies it to a live queue. The target then becomes the policy, because a reviewer who cannot examine carefully will approve, and approval is what the numbers will show.

Interface asymmetry. If clearing an item takes one click and escalating takes a form with a free-text field, the system has expressed a preference the written policy never stated. This is measurable and almost never measured.

Policy drift. Guidelines get revised, the revisions reach some reviewers and not others, and six weeks later two people are applying different rules to the same content with equal confidence. Version the policy, date it, and test against it.

And feedback absence underneath all of it. Reviewers who never learn their own error rate become confidently inconsistent, and the inconsistency compounds because nothing corrects it. Some proportion of items has to carry a known answer, and the reviewer has to see how they did.

The queues that should not be distributed

Being explicit about this matters more than any capability claim.

Illegal material, and particularly child sexual abuse material, belongs with a specialist provider operating under a defined legal protocol with trained staff, controlled access, and mandatory reporting handled correctly. A distributed workforce cannot meet the evidence-handling requirements, cannot provide the support the work requires, and should never be routed material of that kind. If a supplier offers to take that queue as part of a general moderation contract, that is the point to stop the conversation.

The same reasoning applies more mildly to sustained exposure to graphic violence and self-harm content. That work needs an employment relationship with real support attached, not a task rate.

What it costs, and what the number hides

General review in English runs roughly $10 to $30 an hour. Work requiring a particular language or market sits at $15 to $35, with scarcer languages at the top. Anything needing a verified professional background runs considerably higher, because the eligible pool is small.

Per-item pricing is common and worth treating carefully. A price per item is an hourly rate with a throughput assumption baked in and hidden. Derive the assumption before agreeing the price: time a realistic sample of hard cases rather than easy ones, and check whether the resulting hourly is one a competent person would accept. If it is not, the quality problem is already in the contract.

How the eligible pool shrinks and the rate rises as a requirement narrows from any competent reviewer to a native speaker in one market.
How the eligible pool shrinks and the rate rises as a requirement narrows from any competent reviewer to a native speaker in one market.

The costs that get left out of comparisons are the ones that decide the total. Policy authorship and revision. Quality sampling. Appeal handling. Time zone coverage at the edges. A vendor quote covering only the review labour will look cheaper than an honest budget and will not be.

Evidence, and why "we reviewed it" is not enough

When the reviewer is not an employee at a desk you control, an assertion that the work happened is not adequate. The decision needs an artifact attached that somebody else can inspect afterwards — the structured response, the reasoning against the specific policy clause, the timestamp, the item state before and after.

This is the part that makes distribution workable rather than risky. What counts as evidence and how to specify it before the work starts covers the formats and where each is appropriate, and the sequencing matters: a proof requirement decided after the work has been done is both unfair and useless, because the bar can be raised retroactively.

  1. Fix the requirement before posting

    What the reviewer must return is set while the brief is written, not negotiated after somebody has already spent the time.

  2. Make it checkable by a third party

    Evidence nobody can verify without specialist tooling is evidence nobody will ever actually check.

  3. Mix in known answers

    A proportion of items with an established correct decision, scored and fed back, is the only thing that catches drift before an audit does.

  4. Keep the appeal independent

    The person deciding an appeal must not be the person who made the original call. Anything else is a complaints process, not an appeal.

Language and market coverage

This is where most moderation programmes are genuinely weak, and where the weakness is invisible until it is expensive.

A platform operating in fifteen countries typically staffs three languages properly, machine-translates the rest into English for review, and accepts the resulting error rate because nobody has quantified it. Translation strips the register, the local reference and the implied context, which are precisely the things the judgment depended on.

Hiring is the wrong instrument for this, because the requirement is not a skill you can train. There is nobody to hire who is simultaneously a native speaker of eleven languages. What you need is access to a pool, with eligibility set per item rather than per employee. How each layer of targeting changes both the pool and the rate sets out where that curve gets steep and where a filter costs more than it returns.

Machine-translating content before review does not moderate fifteen markets. It moderates one market's reading of fifteen.

Questions worth asking any supplier

Six, and the answers describe the operation more accurately than any capability deck.

What is the throughput target and how was it derived? An answer that does not involve timing real hard cases means the target came from a spreadsheet.

Who writes the policy? If the supplier writes it, you have outsourced your enforcement position along with the labour.

What happens when a reviewer disagrees with the guideline? If disagreement is slow, expensive or ignored, it stops happening within weeks and you lose the early warning that a policy is wrong.

How is quality measured, and against what? "Client satisfaction" is not a measurement. A gold set with known answers, scored and reported, is.

What support exists for people reviewing distressing material, and is it real? Ask what proportion of staff used it last quarter.

And what will you refuse to take? A supplier that accepts every queue including the ones that need a legal protocol has told you how carefully they think about this.

Where distribution is the wrong answer

Being direct about this is more useful than widening the pitch.

A large, steady, always-on queue in one language is a managed vendor's problem and they will do it better and cheaper. Sustained exposure to the worst categories needs employment and clinical support. Policy authorship, escalation and appeal decisions belong inside your organisation, permanently.

What distribution handles is the rest: market-specific judgment, surge capacity when volume spikes, coverage in the eleven languages nobody staffed, and structured review where the value comes from who is looking rather than how fast they clear items. How a brief becomes reserved capacity and verified evidence describes that mechanism end to end, including the parts that constrain what we will accept.

For the view from the other side of the queue — what this work is like to do, and what it costs the people doing it — the honest account of remote moderation work is worth reading before designing a programme around it.

Common questions

What are content moderation services?

Outsourced human review of user-submitted content against a platform's policy. In practice it covers several different jobs — pre-publication review, reactive queues from user reports, appeals, and sampling for quality measurement — which have different staffing needs and are frequently bought as though they were one thing.

How much does content moderation cost?

Roughly $10 to $30 an hour for general review in English, $15 to $35 where a specific language or market is required, and more for anything needing a legal or clinical background. Per-item pricing is common and hides the throughput assumption everything else depends on.

Should moderation be outsourced or kept in-house?

Keep the policy, the escalation path and the appeal decision in-house. The volume review underneath can be outsourced. Platforms that outsource the judgment rather than the labour end up unable to explain their own decisions.

What content should never be sent to a distributed pool?

Anything involving illegal material, particularly child sexual abuse material, which has mandatory reporting and evidence-handling requirements that a distributed workforce cannot meet. That queue belongs with a specialist provider working under a defined legal protocol.

Can AI replace human moderators?

It replaces most of the volume and none of the hard cases. Automation removes the routine items first, which concentrates the human queue in exactly the ambiguous material that takes longest and is most likely to be appealed.

What should you ask a moderation supplier?

What the throughput target is and how it was derived, who writes the policy, what happens when a reviewer disagrees, how quality is measured against a known answer, and what support exists for people reviewing distressing material.

Need review capacity in markets you do not staff?

Tell us the queue, the languages and the evidence you need back. If distribution is the wrong shape for it, we would rather say so than sell you a pilot.

Discuss a pilot