One platform, many communities — UX Research Case Study

UX Research · PhD Dissertation

One platform, many communities

Is Wikipedia one community, or 312?

Twenty years of research on how people collaborate online had been built almost entirely on the English edition. I tested whether any of it held in French and Spanish.

Role: Doctoral researcher, end to end Methods: Content analysis · qualitative coding · interviews Context: PhD, Human Centered Design & Engineering, University of Washington, 2021
Three languages 34 editors interviewed Research inside the community Replication and validation

If you read one thing

The three that matter

The rest of this page is the evidence. These are the parts that travel.

01

Same behaviors, different amounts

Every language edition does the same things, in noticeably different proportions. A model built on the biggest community doesn’t transfer, and an average across all of them describes none of them.

02

I coded the smallest community first

Spanish set the categories, not English. Whichever group you analyze first becomes the baseline everyone else gets measured against. Choosing differently costs nothing.

03

Language is its own process

Translation didn’t just change the words. It created new coordination work, new roles, and new failure points that neither the social nor the technical view could account for.

The short version

What this found, in five sentences

Wikipedia runs in more than 300 languages, but nearly everything researchers believed about how Wikipedians work together had been studied on the English edition alone. I ran three studies across the English, French and Spanish editions to test whether the English model actually described the others.

The behaviors turned out to be universal. The proportions did not, and authority sat somewhere completely different in each community. The practical conclusion is that a single platform serving several language communities is not serving one population, and any design or measurement that treats it as one will quietly be built around the largest group.

Context

The blind spot

For nearly twenty years, researchers analyzed how people collaborate on Wikipedia almost entirely on the English edition. Wikipedia runs in more than 300 language editions, yet nearly every assumption about how Wikipedians work together rested on an English-language view.

The bias was built into the platform. Wikipedia was founded in English by two Americans. Its technical foundation, its five pillars, and even its idea of a reliable source are English- and US-based. Because those foundations are shared across every edition, it was easy to assume one universal model and never test whether it held anywhere else.

Prior work had already shown that encyclopedic knowledge itself is not consistent across languages and cultures. If the collaboration models don’t generalize either, then designing or studying any edition from English-based assumptions is bad design.

Is Wikipedia a standardized platform with one common model of collaboration, or 312 active language editions, each with its own?

48talk pages coded for behavior
~2,050posts coded for claims to authority
34editors interviewed across three languages

My role

Inside the community, not observing from outside

I designed and led the whole dissertation: framing the questions, choosing the method for each stage, collecting and analyzing data in three languages, and synthesizing three studies into one comparative model.

I also participated. I took part in edit-a-thons, edited on my own time, met with editors virtually, and presented at Wikimania. That mattered for access and for reading nuance, and it meant the community got something back rather than being a data source.

I’m a native English speaker who learned French and Spanish. That is not the same as reading nuance in either, so I recruited and trained native-speaker research teams for the French and Spanish data, and we met weekly to reconcile coding and translation until the findings held up in each language on their own terms. Running research in languages you didn’t grow up in is possible. Pretending you can do it alone is how the nuance goes missing.

The studies

Three angles on the same question

Rather than study one edition and generalize, I asked the same question three ways: what people do, how they claim control, and how they understand it themselves. Two of the three deliberately replicated well-known English-Wikipedia models, because testing whether a known model travels is how you find out if it generalizes.

Three studies Study one asked what talk pages are used for, using content analysis of forty-eight talk pages. Study two asked how editors claim control, using qualitative coding of about two thousand and fifty posts. Study three asked where editors think power lives, using interviews with thirty-four editors. BEHAVIOR What people do Content analysis 48 talk pages POWER How they argue Qualitative coding ~2,050 posts PERCEPTION What they think Interviews 34 editors
English, French and Spanish editions in every study, so the comparison held across all three.

Finding 01

Same behaviors, different amounts

I replicated an eleven-category model of what editors post on talk pages and applied it across all three editions. Every category appeared in every language. How often people did each one was not the same anywhere.

The English edition’s supposed norm, a one-to-two ratio of article edits to talk-page edits, no longer held, and probably never held for French or Spanish. Talk pages had also quietly shifted from coordinating edits to debating the subject of the article itself. References to vandalism, for example, collapsed from 8.5% of posts to 0.9% as bots took over that work.

In the second study I replicated a model of “power plays”, the moves editors use to claim authority through the language of policy. No new ones appeared in any language, so the original model held up as a complete list. But the frequencies differed: French leaned hardest on article scope, Spanish on power of interpretation.

Why this matters outside Wikipedia

A model built on the behavior of the largest community can’t be assumed to transfer to the others. And an average across all of them describes none of them, which means an aggregate metric can improve while one community’s experience gets worse.

Finding 02

Authority lives somewhere different in each community

I interviewed 34 experienced editors: 13 English, 13 Spanish and 8 French. They described strikingly similar conversations, mostly about sourcing, translation and defusing conflict. They diverged sharply on who gets to decide.

Where authority sits in each edition In Spanish, authority sits with administrators, roughly sixty-eight of them, after the edition dissolved its arbitration committee in 2009. In English, it is spread across the most experienced editors and formal process. In French, it sits in the strength of the argument and its sources rather than in who is speaking. SAME SOFTWARE. SAME RULES. THREE DIFFERENT ANSWERS. SPANISH A small group Roughly 68 administrators, after the edition dissolved its arbitration committee ENGLISH Experience The longest-serving editors, working through formal process and appeals FRENCH The argument Who you are matters less than how well you argue and what you cite
Every edition runs the same software under the same founding principles. Who actually decides is different in each one.

Why this matters outside Wikipedia

Designing one process for one model of authority will fail quietly in the community that runs on a different one. Nobody complains, because from the inside it just looks like the system isn’t for them.

Finding 03

Translation created its own work

In French and Spanish, importing content from larger editions spawned entire coordination projects that had no equivalent in English. Language didn’t just change the content. It generated new social processes and new labour that somebody had to organize.

Existing analysis treated Wikipedia as a socio-technical system, meaning the product of social processes and technical ones. That framing had nowhere to put what I kept finding: translation, sentence structure, clarity and verbosity were actively shaping how people argued and how they reached agreement. So I added language as a first-class process alongside the other two.

Language added as a third process Collaboration was previously analyzed through two processes: social, covering norms, roles and authority, and technical, covering software, tools and bots. This work added a third, language, covering translation, structure, clarity and verbosity. THE USUAL VIEW Social norms, roles, who decides Technical software, tools, bots WHAT I ADDED Language translation, structure, clarity, verbosity, and the work each of those creates
A lens for any system where one platform serves several language communities at once.

Why this matters outside Wikipedia

Translation is usually treated as a finishing step applied to a finished design. It isn’t. It generates its own coordination work, its own roles and its own failure modes, and all of that stays invisible if language is treated as a property of the content rather than a process in its own right.

The decision I’d make again

I coded the smallest community first

The conventional order builds the coding scheme on English, the largest and best-studied edition, then applies it to the others and notes the differences. I started with Spanish, the smallest of the three, so its patterns shaped the categories rather than being scored against a scheme built somewhere else.

Two orders for building a coding scheme The usual order builds the scheme on the largest community, then measures the others against it, so smaller communities appear only as deviations. The order used here builds the scheme on the smallest community first, then tests it against the larger ones, so every community's patterns end up inside the scheme. THE USUAL ORDER Biggest Everyone else Smaller groups can only appear as deviations from someone else’s normal. WHAT I DID Smallest Everyone else Every group’s patterns end up inside the scheme, not in the margins.
A sequencing choice, not a methodological invention. It costs nothing and it decides whose normal the analysis is built around.

Why this matters outside Wikipedia

The same logic applies anywhere one product serves several language communities. A design built for the majority and then translated is a different product from one designed for both, and the difference usually shows up first in the things the second group never complains about, because they already stopped trying.

Contributions

What the work established

  • Empirical. An evidence-based account of how three language editions actually collaborate, showing that findings from English Wikipedia do not automatically generalize.
  • Methodological. A worked example of multi-site, cross-language research, and a framework extended to treat language as a first-class process.
  • Validation. A systematic replication of two canonical English-Wikipedia models across three languages, a direct response to the replication gap in this field.
  • Reframed the default. Made the case that designing or studying any edition from English-based assumptions is bad design.
  • Extended beyond Romance languages. A follow-on replication in the Farsi and Chinese editions, using non-Latin scripts and different language families, was accepted to CSCW 2021.

Where this applies

What I take into product work

This is a dissertation, so it ends in theory rather than a shipped decision. What it left me with is a set of questions I now ask about any system serving more than one community, and they change what gets measured.

  • Whose normal is this built around? Usually the largest group, usually without anyone deciding to.
  • Does the aggregate number describe anyone? If two communities behave differently, the average describes neither, and it will look like progress either way.
  • What work does translation create? Not what it costs to translate, but what new coordination, roles and failure points appear because two languages are in play.
  • Who holds authority in each community? It usually isn’t the same person, and one process for one model of authority fails quietly in the other.
  • Which community do we look at first? The cheapest way to avoid building around the majority by default is to start somewhere else.

Reflection

What I’d do differently

  • Scale the language expertise further. Native speakers gathered and reconciled data weekly, but most analysis still ran through one researcher.
  • Acknowledge the sample. The models draw on a predominantly male editor population. Consistent with Wikipedia’s demographics, and still an unbalanced base to generalize from.
  • Go further from Romance languages. English, French and Spanish share a script and a language family, which is why I extended the work to Farsi and Chinese.
  • Map the policy regimes. Each edition organizes policy differently. A dedicated study of how editors value and use policy across languages is the obvious next step.

About this dissertation. Comparing Language Communities: Characterizing Collaboration in the English, French and Spanish Language Editions of Wikipedia. PhD dissertation, University of Washington, 2021. Committee co-chaired by David W. McDonald and Mark Zachry.

Taryn Bipat · UX Research · PhD, Human Centered Design & Engineering, University of Washington