Any intervention strong enough to reduce misinformation on a social platform is also strong enough to be called censorship. We designed moderation mechanisms for X, whose audience is primed to read moderation as attack, and treated that trap as the central design constraint: each mechanism had to specify not just a screen but a policy, and both had to survive contact with the people they correct.
We anchored testing on one real, debunked claim: that the National Guard had been called in on the spring 2024 protests at UT Austin. The claim was charged without being party-line partisan, believable enough that several posts carried it, and clearly debunked by the Associated Press, so every session could end with the record.
Working as a three-person team, we designed, specced, and tested three mechanisms, each intervening at a different point in a rumor's life:
- Correct, after belief. Community Notes Notifications tells a user when a post they've seen later receives community context.
- Contain, at the moment of spread. Quarantine pulls a flagged post out of algorithmic circulation while moderators review it.
- Contextualize, before exposure. The Developing Story Tracker opens a trending story on a journalist-curated frame beside the raw feed.
Findings
We walked five participants through all three mechanisms in think-aloud sessions, evaluating each against five design values: free speech, reducing bias, transparency, societal impact, and user effort. Each section below opens with the mechanism as we specced it, surface and policy; the quotes are the evidence; the revision is what the evidence forced.
- 5 think-aloud sessions
- 3 intervention concepts
- 5 design values
Correct · Community Notes Notifications: corrections work until they flood
- Surface
- A notification when a post you viewed later receives a community note; a category label above the corrected post (Misleading · Factual Error); a history view for posts corrected more than once.
- Policy
- A trigger threshold decides which viewers get notified; an update standard decides when a changed note notifies again.
"I would appreciate being notified if something I read is factually incorrect or has other errors because oftentimes I bring this up in discussion and it would be embarrassing to be wrong."
The trigger narrowed from viewed to interacted with: liked, reposted, or replied to. Note updates now require evolving guidance from verifiable, authoritative sources rather than evolving rumors. And the notification links straight to X's documentation of how Community Notes works.
Participants received the notification as an extension of how they already use X: keeping up with stories that change, and not repeating things they later learn are wrong. The threat was volume. If merely viewing a post could trigger a correction, people expected a flooded inbox, and a correction users learn to ignore is no correction at all. So the threshold is the design: tying notification to likes, reposts, and replies keeps it pointed at posts a person actually amplified. The documentation link answers a second finding: participants carried wildly different mental models of Community Notes, and those who didn't understand the ranking system distrusted its output regardless of content.
Step 1 of 5:A charged claim about the UT Austin protests sits in the timeline, reading as ordinary news.
Contain · Quarantine: understood, but it concentrated power
- Surface
- A quarantined post is visibly marked in the feed; liking, sharing, and commenting are disabled while review runs; any user can opt into a notification of the outcome.
- Policy
- Flagged posts enter quarantine during moderation review; they leave the algorithmic For You feed and search results but remain visible in Following; review ends one of three ways: removed, context added, or cleared.
"I don't repost much – I'm a lurker. So I guess I want them [X] to just clean up my feed from misinfo garbage and make sure that others don't repost it."
"I think this could be good but yea…screenshots. I feel like it's almost a challenge for them like 'oh you don't want me to share this I'm gonna prove you wrong and share it'…they really think the radical Left is out to get them."
Initiation moved to the moderation team alone: users still flag, but a flag no longer quarantines. Scope narrowed to content with potential for immediate harm. And removal became inspectable for followers: a moderation view shows the removed post and the guideline it broke.
Participants grasped the appeal immediately: quarantine works on the infrastructure of spread and asks nothing of the reader. The same property made it the least trusted of the three. It concentrates authority in the platform, and participants predicted the failure modes in detail: quarantine notices screenshotted and recirculated as proof of persecution, and report flows weaponized against queer creators, artists, activists, and small businesses. The revision follows from that evidence. What people feared was not review itself; it was who could trigger it.
Step 1 of 5:A flag from the moderation team, and only the moderation team, holds the post: still visible to followers, frozen for interaction while review runs.
Contextualize · Developing Story Tracker: the mechanism that fit existing behavior
- Surface
- Entry points from the feed and trending; a story opens on a journalist-curated Analysis tab beside an uncurated Latest tab; the engagement-ranked Top tab is gone.
- Policy
- A participating journalist curates credible posts and adds framing commentary as the story develops; the platform decides which journalists cover which stories.
"I would use this as a reference when debating my family during stuff in group texts. My parents are smart but sometimes they share and believe crazy sh*t. This could automate it for me to send them here."
The entry label changed: "Follow story" collided with what follow already means on X, so the final screens prompt with view-coverage language. A share button was added so the curated frame itself can be sent onward. And we recommended a journalist-selection policy that accounts for outlet lean, pairing outlets across the spectrum where possible.
The tracker drew the warmest reception because it formalizes something people already do: open X to find out what is happening right now. The curated tab beside the raw one read as transparency rather than substitution, and the share button turned the tracker into a personal debunking tool. The contested decision was the default. Opening on Analysis struck one participant as the platform "trying to narrate rather than having me research for myself," though they admitted they would read it anyway; the journalist's frame is the mechanism's power and its bias risk in a single design decision.
Step 1 of 5:A post in a developing story carries an entry point. The prompt says view coverage: the original label, Follow story, collided with what follow already means on X.
Three principles
Across the three mechanisms, three principles held, and they are what we would carry into any moderation work:
Agency distribution determines reception. Quarantine was the most technically aggressive mechanism and drew the most resistance, not because participants disagreed with the goal but because it concentrated authority in the platform. The mechanisms that distributed agency across communities, journalists, and readers were trusted. Interventions that augment user agency beat interventions that replace it.
Notification design is misinformation design. Even participants who valued staying corrected feared overload, and a correction users learn to ignore is no correction at all. The threshold policy, not the notification screen, is where that fight is won.
Mental models shape trust as much as mechanisms do. People who understood how Community Notes works trusted its corrections; people who didn't were skeptical regardless of content. Shipping a moderation mechanism also means shipping comprehension of it; the two can't be separated.
Process
We worked through a six-stage, values-centered process: literature review, success criteria, ideation, medium-fidelity mockups, user testing, and refinement. The work was collective at every stage; I built this with Nina Lutz and Ben Yamron in HCDE 598 (Designing for Trust), taught by Kate Starbird at the University of Washington.
From three criteria to five values
The literature review produced three success criteria: user experience and affect, trust and facts, adaptability and scalability. We translated those into five single-issue values participants could actually argue with: free speech, reducing bias, transparency, societal impact, and user effort. The same five values scored our concepts, structured the test protocol, and framed the spec revisions above.
From six ideas to three mechanisms
We sketched six ideas against the criteria: a Community Notes extension, more descriptive note labels, an interstitial manipulation warning, share-blocking for posts under fact-check review, badges rewarding nuanced posts, and journalist-curated story tracking. Scoring merged the first two and carried three forward, one for each point in the rumor's life: correct, contain, contextualize.
Testing with five people who use X differently
We recruited five participants across usage patterns: a software engineer who mostly lurks, a medium-to-high-usage participant from Texas, an urban planner in Washington DC, a medical researcher in Seattle, and a Seattle PhD student. Each walked through all three prototypes in a think-aloud session, answering the five value questions and revealing, along the way, their mental models of how moderation works.
Reflection
The biggest limitation is sample diversity. Five educated, mostly urban participants, tested on a claim that was politically complex but not explicitly partisan, leaves real uncertainty about how these mechanisms would land with people who hold strong partisan identities or place free speech above accuracy: the population most likely to resist moderation, and a meaningful share of X's base. Future rounds should recruit across political identity first.
The prototypes also limit ecological validity. Medium-fidelity, non-interactive screens examined attentively in a session are not a feed scrolled mindlessly; a quarantine notice met mid-scroll may provoke reactions our sessions never surfaced. Higher-fidelity testing inside a simulated feed is the obvious next build.
What held up: the values framework structured ideation, testing, and revision without collapsing into vibes; the real, debunked claim kept sessions concrete; and the triad (correct, contain, contextualize) gave us a vocabulary for when to deploy which mechanism, which is the kind of language a platform team can actually plan around.