Work index / Case study

CS·01 Research + Design · University of Washington · 2024

Intervention Design for Misinformation

Three moderation mechanisms for X, specced surface to policy and tested against five design values: correct after belief, contain at the moment of spread, contextualize before exposure.

Any intervention strong enough to reduce misinformation on a social platform is also strong enough to be called censorship. We designed moderation mechanisms for X, whose audience is primed to read moderation as attack, and treated that trap as the central design constraint: each mechanism had to specify not just a screen but a policy, and both had to survive contact with the people they correct.

We anchored testing on one real, debunked claim: that the National Guard had been called in on the spring 2024 protests at UT Austin. The claim was charged without being party-line partisan, believable enough that several posts carried it, and clearly debunked by the Associated Press, so every session could end with the record.

Working as a three-person team, we designed, specced, and tested three mechanisms, each intervening at a different point in a rumor's life:

Findings

We walked five participants through all three mechanisms in think-aloud sessions, evaluating each against five design values: free speech, reducing bias, transparency, societal impact, and user effort. Each section below opens with the mechanism as we specced it, surface and policy; the quotes are the evidence; the revision is what the evidence forced.

Contextualize · before exposure Contain · at spread Correct · after belief
A rumor's path from publication to belief, with each mechanism at its intervention point: the tracker frames the story before exposure, quarantine holds it at the moment of spread, notifications correct it after belief. Red is the reader the mechanisms protect.

Correct · Community Notes Notifications: corrections work until they flood

Surface
A notification when a post you viewed later receives a community note; a category label above the corrected post (Misleading · Factual Error); a history view for posts corrected more than once.
Policy
A trigger threshold decides which viewers get notified; an update standard decides when a changed note notifies again.

"I would appreciate being notified if something I read is factually incorrect or has other errors because oftentimes I bring this up in discussion and it would be embarrassing to be wrong."

Spec revision

The trigger narrowed from viewed to interacted with: liked, reposted, or replied to. Note updates now require evolving guidance from verifiable, authoritative sources rather than evolving rumors. And the notification links straight to X's documentation of how Community Notes works.

Participants received the notification as an extension of how they already use X: keeping up with stories that change, and not repeating things they later learn are wrong. The threat was volume. If merely viewing a post could trigger a correction, people expected a flooded inbox, and a correction users learn to ignore is no correction at all. So the threshold is the design: tying notification to likes, reposts, and replies keeps it pointed at posts a person actually amplified. The documentation link answers a second finding: participants carried wildly different mental models of Community Notes, and those who didn't understand the ranking system distrusted its output regardless of content.

Step 1 of 5:A charged claim about the UT Austin protests sits in the timeline, reading as ordinary news.

Community Notes Notifications, stepped: interaction arms the correction, context arrives after the reader moves on, and the note reaches belief.
Revised notification screen reading 'A post you interacted with has been given additional context,' with a link to learn about Community Notes
The revised notification: an interaction-based threshold and a link to how Community Notes works.

Contain · Quarantine: understood, but it concentrated power

Surface
A quarantined post is visibly marked in the feed; liking, sharing, and commenting are disabled while review runs; any user can opt into a notification of the outcome.
Policy
Flagged posts enter quarantine during moderation review; they leave the algorithmic For You feed and search results but remain visible in Following; review ends one of three ways: removed, context added, or cleared.

"I don't repost much – I'm a lurker. So I guess I want them [X] to just clean up my feed from misinfo garbage and make sure that others don't repost it."

"I think this could be good but yea…screenshots. I feel like it's almost a challenge for them like 'oh you don't want me to share this I'm gonna prove you wrong and share it'…they really think the radical Left is out to get them."

Spec revision

Initiation moved to the moderation team alone: users still flag, but a flag no longer quarantines. Scope narrowed to content with potential for immediate harm. And removal became inspectable for followers: a moderation view shows the removed post and the guideline it broke.

Participants grasped the appeal immediately: quarantine works on the infrastructure of spread and asks nothing of the reader. The same property made it the least trusted of the three. It concentrates authority in the platform, and participants predicted the failure modes in detail: quarantine notices screenshotted and recirculated as proof of persecution, and report flows weaponized against queer creators, artists, activists, and small businesses. The revision follows from that evidence. What people feared was not review itself; it was who could trigger it.

Step 1 of 5:A flag from the moderation team, and only the moderation team, holds the post: still visible to followers, frozen for interaction while review runs.

Quarantine, stepped: the post is frozen for followers, held out of circulation, reviewed, and returned with its outcome.
Three notification screens showing the possible quarantine outcomes: post removed, context added, or post cleared
The three outcomes after review: removed, context added, or cleared.

Contextualize · Developing Story Tracker: the mechanism that fit existing behavior

Surface
Entry points from the feed and trending; a story opens on a journalist-curated Analysis tab beside an uncurated Latest tab; the engagement-ranked Top tab is gone.
Policy
A participating journalist curates credible posts and adds framing commentary as the story develops; the platform decides which journalists cover which stories.

"I would use this as a reference when debating my family during stuff in group texts. My parents are smart but sometimes they share and believe crazy sh*t. This could automate it for me to send them here."

Spec revision

The entry label changed: "Follow story" collided with what follow already means on X, so the final screens prompt with view-coverage language. A share button was added so the curated frame itself can be sent onward. And we recommended a journalist-selection policy that accounts for outlet lean, pairing outlets across the spectrum where possible.

The tracker drew the warmest reception because it formalizes something people already do: open X to find out what is happening right now. The curated tab beside the raw one read as transparency rather than substitution, and the share button turned the tracker into a personal debunking tool. The contested decision was the default. Opening on Analysis struck one participant as the platform "trying to narrate rather than having me research for myself," though they admitted they would read it anyway; the journalist's frame is the mechanism's power and its bias risk in a single design decision.

Step 1 of 5:A post in a developing story carries an entry point. The prompt says view coverage: the original label, Follow story, collided with what follow already means on X.

The Developing Story Tracker, stepped: entry from feed or trending, the journalist's frame first, the raw stream one tap away.
Revised Story Tracker screens with clarified view-coverage entry language and an added share button for sending curated stories to others
The revised tracker: view-coverage entry language and a share button for sending the curated frame onward.

Three principles

Across the three mechanisms, three principles held, and they are what we would carry into any moderation work:

Agency distribution determines reception. Quarantine was the most technically aggressive mechanism and drew the most resistance, not because participants disagreed with the goal but because it concentrated authority in the platform. The mechanisms that distributed agency across communities, journalists, and readers were trusted. Interventions that augment user agency beat interventions that replace it.

Notification design is misinformation design. Even participants who valued staying corrected feared overload, and a correction users learn to ignore is no correction at all. The threshold policy, not the notification screen, is where that fight is won.

Mental models shape trust as much as mechanisms do. People who understood how Community Notes works trusted its corrections; people who didn't were skeptical regardless of content. Shipping a moderation mechanism also means shipping comprehension of it; the two can't be separated.

Process

We worked through a six-stage, values-centered process: literature review, success criteria, ideation, medium-fidelity mockups, user testing, and refinement. The work was collective at every stage; I built this with Nina Lutz and Ben Yamron in HCDE 598 (Designing for Trust), taught by Kate Starbird at the University of Washington.

From three criteria to five values

The literature review produced three success criteria: user experience and affect, trust and facts, adaptability and scalability. We translated those into five single-issue values participants could actually argue with: free speech, reducing bias, transparency, societal impact, and user effort. The same five values scored our concepts, structured the test protocol, and framed the spec revisions above.

From six ideas to three mechanisms

We sketched six ideas against the criteria: a Community Notes extension, more descriptive note labels, an interstitial manipulation warning, share-blocking for posts under fact-check review, badges rewarding nuanced posts, and journalist-curated story tracking. Scoring merged the first two and carried three forward, one for each point in the rumor's life: correct, contain, contextualize.

Ideation board showing brainstormed concepts scored against the success criteria, with each team member's contributions labeled
Ideation: six concepts sketched and scored against the success criteria.

Testing with five people who use X differently

We recruited five participants across usage patterns: a software engineer who mostly lurks, a medium-to-high-usage participant from Texas, an urban planner in Washington DC, a medical researcher in Seattle, and a Seattle PhD student. Each walked through all three prototypes in a think-aloud session, answering the five value questions and revealing, along the way, their mental models of how moderation works.

Reflection

The biggest limitation is sample diversity. Five educated, mostly urban participants, tested on a claim that was politically complex but not explicitly partisan, leaves real uncertainty about how these mechanisms would land with people who hold strong partisan identities or place free speech above accuracy: the population most likely to resist moderation, and a meaningful share of X's base. Future rounds should recruit across political identity first.

The prototypes also limit ecological validity. Medium-fidelity, non-interactive screens examined attentively in a session are not a feed scrolled mindlessly; a quarantine notice met mid-scroll may provoke reactions our sessions never surfaced. Higher-fidelity testing inside a simulated feed is the obvious next build.

What held up: the values framework structured ideation, testing, and revision without collapsing into vibes; the real, debunked claim kept sessions concrete; and the triad (correct, contain, contextualize) gave us a vocabulary for when to deploy which mechanism, which is the kind of language a platform team can actually plan around.