What changed

Meta published a set of child-safety measures on 7 October, and the one that matters technically is a change in what an ad review looks at.

Until now the question an automated ad review answers is whether the ad itself breaks a rule. Meta says it is adding a destination-based assessment: the systems now evaluate where an ad sends a person, so a non-compliant destination can be blocked and the accounts behind it acted against.

Alongside it the company has put a large language model on what it calls signposting — ads that look entirely ordinary but are strongly suspected of steering users towards illegal content or harmful activity somewhere else online. Meta says bad actors adopted the tactic specifically to get past content-only review.

The numbers Meta gave

For the first half of 2026, Meta says it took action on 33.2 million pieces of child sexual exploitation content across Facebook and Instagram, and that over 97% of it was found and addressed before anyone reported it. For India it gives 5.3 million pieces with over 98% found proactively.

A person working at two computer monitors in an office
The company reports acting on 33.2 million pieces of exploitation content in six months. Illustrative image. Kampus Production · pexels · Pexels License

Those are enforcement counts, not prevalence estimates. A rising number can mean more abuse or better detection, and the post does not separate the two. Meta gives no figures at all for ad-related enforcement, which is the thing the new measures are about.

The rest of the package

Meta also describes additional AI-driven sweeps aimed at material earlier systems missed, strengthened detection of removed users returning under new accounts, and a red-teaming AI agent that probes Meta’s own defences to, in the company’s words, identify new adversarial tactics before they scale.

The red-teaming agent is the interesting one. An automated adversary that finds the gaps in a classifier before an attacker does is the same technique security teams use on code, pointed at a moderation stack. Meta publishes no results from it.

A hand holding a smartphone showing a grid of apps
A red-teaming agent probes Meta's own defences, the company says. Illustrative image. Brian Ramirez · pexels · Pexels License

The context it arrives in

Meta has been under sustained legal and regulatory pressure on child safety. In August it agreed to pay up to $18 billion to settle a lawsuit brought by 29 US states over harms to children. Earlier in the year it added parental controls to Meta AI, parent-linked WhatsApp accounts for preteens, and Instagram alerts when a teenager searches for suicide or self-harm content.

The post sits in Meta’s India newsroom, and names a September commitment to report child-safety cases directly to India’s national cybercrime portal, run by the Indian Cybercrime Coordination Centre.

What to watch

Whether Meta publishes any measure of how often the destination check fires, and what share of flagged destinations turn out to be violating. A classifier pointed at link targets is also a classifier that can be wrong about a legitimate advertiser’s landing page, and neither error rate has a number next to it yet.