Meta Board Orders 2 Deepfakes Removed, Citing 9 Gaps in AI Safeguards
Updated
Updated · The Guardian · Sep 17
Meta Board Orders 2 Deepfakes Removed, Citing 9 Gaps in AI Safeguards
3 articles · Updated · The Guardian · Sep 17
Summary
Two Facebook deepfakes — one targeting a Scottish Labour councillor, another mocking a young Muslim campaign volunteer — must be removed under binding Meta Oversight Board rulings.
The board said Meta’s AI defenses were “consistently and fundamentally inadequate,” faulting the company for leaving up the councillor clip even after review and for failing to apply a “high risk AI” label.
Nine policy changes were recommended, including broader use of high-risk labels, algorithmic demotion of labeled posts, stronger penalties for repeat sharers, and warning screens before users can view AI-generated content.
The councillor video was ruled hateful conduct because it falsely portrayed refugees as sexual predators, while the second case was found to breach bullying and harassment rules.
The board said AI deepfakes are increasingly used to harass and silence women in public life, with manipulated videos of the Muslim volunteer drawing tens of millions of views online.
Will Meta enforce these strict deepfake rules, or are they merely a PR shield?
Can any social platform truly protect you from becoming the next deepfake victim?
Meta’s Deepfake Crisis: Why 48-Hour Review Windows and Labeling Fail to Protect Victims in the Age of AI Abuse and Global Takedown Laws
Overview
This report reveals how Meta’s outdated and unclear policies failed to protect victims from AI-generated deepfakes, as shown by two high-profile cases involving explicit images of public figures. When Meta’s automated systems closed user reports and appeals without review, harmful content stayed online, causing lasting damage. Public backlash intensified after Meta’s AI tool allowed users to create synthetic images of anyone, forcing the company to disable the feature. Meanwhile, technical flaws in Meta’s watermarking let abusers easily bypass detection, leaving victims exposed. New global laws now demand rapid removal of such content, pushing platforms to overhaul their moderation and victim support systems.