When the Bots Purge the Archive: Reddit’s AI Moderation Erased a Decade of r/AskHistorians Work

when the bots purge the archive reddits ai moderation erased a decade of raskhistorians work Over the course of April, dozens of posts and comments stretching back a full decade disappeared from r/AskHistorians. They weren't edited. They weren't quietly queued for review. They were removed automatically — from a subreddit whose entire reason for existing is that its answers remain readable years after they're written.

Over the course of April, dozens of posts and comments stretching back a full decade disappeared from r/AskHistorians. They weren’t edited. They weren’t quietly queued for review. They were removed automatically — from a subreddit whose entire reason for existing is that its answers remain readable years after they’re written.

“And there was nothing we or the experts [who posted the deleted content] could do about it,” said Dr. Sarah Gilbert, one of the community’s moderators.

This is the side of AI moderation that never makes it into a press release. It’s also the right place to begin, because in that very same month Reddit was publicly touting that its AI had pushed enforcement actions against hate and violent content up by more than 200 percent.

What the mods actually saw

The Slack channel where the r/AskHistorians mod team receives links to modmail messages is normally active anyway. On that day, it was drowning in alerts.

Once the team managed to recover the text of some deleted posts, one moderator spotted the common thread: every single casualty had linked to Rare Historical Photos, a site devoted to sharing historical imagery. The mods’ working theory is that Reddit tagged the domain as spam, and any post using its images as explanatory illustrations was swept away alongside it.

The moderators attribute the purge to Reddit’s recently overhauled AI moderation tooling. Reddit has not responded to a request for comment.

What elevates this above an ordinary false positive is the nature of the loss. According to Gilbert, some contributors devote hours — “sometimes over the course of days” — to researching and drafting a single answer, and the subreddit’s readers treat the place as an archive. Wiping a 10-year-old comment isn’t an afternoon’s inconvenience; it destroys work that was still actively teaching people.

The numbers that don’t mean what they look like

Reddit’s pitch is that AI delivers “faster, higher volume enforcement.” This month it stated that AI has “helped reduce exposure to potentially harmful content by more than 40 percent.” It also says its tools have “revoked nearly [2 million] fake votes daily,” and that large language models are deployed “to catch the highly subtle, coordinated patterns of fake behavior and artificial hype that older systems once missed.”

Each of those is a volume metric. Not one of them separates a deleted neo-Nazi rant from a decade-old citation pointing at a photo archive. And if the AskHistorians removals were tallied as enforcement actions, they helped inflate that 200 percent figure.

Gilbert is skeptical of the accounting, and she can articulate exactly why. “Back when there was more transparency in the system, we would routinely report hate and get an automated response that it wasn’t actually in violation of Reddit’s rules, prompting us to start an appeals process,” she said. “So it’s hard to trust the numbers because it’s hard to trust the ‘judgment’ of Reddit’s systems.”

On Reddit, she said, false positives are a “huge problem.”

Discord banned 8,400 people over chessboards

For the sharpest demonstration of how this collapses when no human is involved, look at Discord. The company acknowledged that its AI moderation system wrongfully banned roughly 8,400 accounts between May and early July.

The reason: the AI read images containing square grids as CSAM. Chessboards. Spreadsheets. Post one, receive a permanent ban. Discord says every affected account has since been restored.

The company’s own explanation is the detail worth dwelling on. Discord said its AI moderation was never designed to operate unsupervised — a human employee is meant to review anything the AI flags before action is taken. A bug allowed the AI to bypass that human step and issue bans on its own.

The safeguard, in other words, existed on paper. It took one software defect and about two months for 8,400 accounts to discover it wasn’t actually running.

The problem AI created, and is now sold as the fix for

Generative AI didn’t only break moderation at the enforcement end. It swamped the input end first.

Because large language models are engineered to mimic authentic human voices, they “have made spam detection a lot harder,” Gilbert said. “Over the last two to three months, we’ve been absolutely flooded by LLM-powered spambots,” she said.

Part of that surge has a commercial motive behind it. Marketing agencies are now churning out social posts engineered to get brands name-checked by chatbots. Inauthentic posting for visibility is nothing new; targeting chatbot outputs with it is. The startup ReachLLM is built specifically around marketing via chatbots, and its representatives have created and moderate subreddits on Reddit.

Who gets caught in the net

Conventional moderation stacks rely on machine learning classifiers that scan posts and flag rule-breaking material. Sarcasm, satire and slang happen to be precisely what classifiers handle worst.

That failure isn’t evenly distributed. Research indicates that marginalized groups bear a disproportionate share of the harm from AI moderation. Gilbert, who also serves as research director of Cornell’s Citizens and Technology Lab, said “marginalized and vulnerable populations are among those who experience the highest rates of moderation, and that typically this is a result of ‘false-positives,'” often set off by counter-speech, language reclamation and “responses to hateful content.”

“False positives are an equity issue. They mean that groups that are already marginalized are further silenced and censored,” she said.

Hold that up against the 200 percent enforcement increase. Systems marketed as shielding vulnerable communities from hate end up punishing those communities for pushing back against it.

Meta and Tumblr have the same bug

Facebook and Instagram users have been complaining since 2025 about waves of bans they blame on AI moderation. Meta has declined to say whether AI is responsible. What the company has done is shift further toward generative-AI-based moderation and away from human reviewers — a pivot that some observers, Meta employees among them, argue is happening too quickly.

Having nobody on the other end is a harm in itself. Users who insist they broke no rules have found no route to an actual Meta employee who could explain what happened or how to be reinstated.

Tumblr has its own variant of the story. In March, Chenda Ngak, head of communications at Tumblr’s parent company Automattic, said Tumblr’s automated systems wrongfully banned “sub-200” accounts over a single afternoon. During 2025, users reported that automatic moderation was misflagging content as “mature,” throttling its reach. Tumblr has never confirmed AI caused either issue, though it has acknowledged using “a mix of machine-learning classification and human moderation.”

The quieter cost: mods lose the call

One consequence never surfaces in any transparency report. Certain subreddit moderators would prefer to ban a user outright over hateful or violent rhetoric. But if Reddit’s AI strips the content before any human mod lays eyes on it, that mod has no basis for deciding whether a ban is justified.

In the logs, automated removal reads as a victory. On the ground, it deprives the community of the context it needs to govern itself.

Reddit, to its credit, is making moves here. This week it broadened testing of Rules Hub, which allows human mods to “choose which rules should be automatically enforced, decide what happens when a rule is triggered (send to queue, filter, or remove), preview the experience before enabling it, and review logs and insights.” Reddit anticipates it will eventually supersede Automod, which depends on exact keyword matching.

That’s the correct trajectory: handing more control to the people who understand the community, not less.

What would actually fix this

AI moderation cuts platform costs and takes harmful material down faster than any human team could. Both claims hold up. But a system incapable of distinguishing a checkerboard from CSAM hasn’t earned the authority to act by itself, and Discord’s own post-mortem is the strongest case yet for keeping a person in the loop.

Moderators consistently identify the generative AI boom as the driver behind the spike in rule-breaking content. For platforms whose product is whatever users contribute, that isn’t a support ticket — it’s existential.

The solution isn’t less AI. It’s detection at machine scale coupled with human judgment empowered to overrule it, plus enough transparency for mods to audit the machine’s work. Cutting human input while low-effort AI content keeps piling up is a deliberate step in the wrong direction.

Gilbert’s team spent April salvaging text from posts that had taken contributors days to produce. Nobody at Reddit had to approve deleting them. That’s the design flaw.

Advance Publications is the largest shareholder in Reddit.