Zotero at Twenty

6 minute read

As we approached midnight on October 5, 2006, a small team at the Center for History and New Media put the final touches on the public launch of Zotero. Lead developer Dan Stillman and I chatted with each other the way everyone did back then: on AOL Instant Messenger. We had high hopes for what was to come the next day: our last discussion was about how to count Zotero downloads once we flipped the switch.

From a distance of twenty years it might be hard to remember just how revolutionary Zotero promised to be: unlike everything before it, Zotero could recognize what you were reading and automatically save a reference; it could generate accurate citations formatted for any publication’s style guide in a brilliantly designed new style language; it operated entirely inside your own web browser, a space where researchers had only recently begun to conduct the bulk of their research; and it did all of this for the price of… nothing.

Although we had tested Zotero privately for weeks, that October we essentially started with a blank slate. I became user 3, and I still am (admin is 1; Dan Cohen is user 2). Today Zotero has over twenty million other people using it. The story of how we got here is less one of continuity – Zotero hasn’t merely perdured – than one of innovation and transformation, as our team has repeatedly taken a great idea far beyond the limits of the technical and institutional infrastructure that launched it. From the beginning Zotero promised to be open and free, forever; keeping that promise has meant continuous experimentation and adaptation.

Technical experimentation

Zotero’s two key innovations – automatic recognition of research objects (a book, a newspaper story, a journal article, a conference paper) and automatic citation generation – each became established not only as core features of Zotero but also of research infrastructure writ large. Wikipedia repurposed Zotero’s recognition code to build citations into articles, and ProQuest copied it for its own reference manager. Zotero’s citation generation code has had an even more significant impact. Zotero automatically generates citations using Citation Style Language, a new standard which geographer Bruce D’Arcus designed to describe seemingly arbitrary citation styles in a logical and easily editable format. Zotero’s implementation of CSL was the first to show that the specification could actually work in practice and at scale; Zotero employees and volunteers produced nearly all of CSL’s initial styles. CSL today powers not just Zotero but most of its competitors, and our shared open repository now holds over ten thousand styles covering the full range of scholarly literature.

Because Zotero is free and open-source software, its innovations quickly became everyone’s; remaining free and open required us to continue to innovate. Rather than coast on its original design, Zotero has adapted to and reshaped the research ecosystem: while it has always worked with commercial publishers’ databases, in 2018 Zotero began pointing researchers to free and legal open-access copies of paywalled publications; in 2019 Zotero began automatically flagging retracted works. Both improvements brought valuable insights about research publications – where free copies, including preprints and self-archived versions, lived; when publications had been identified as problematic – to all researchers, and at no cost. While any number of software platforms promised to allow PDF annotation or collaborative annotation, Zotero was the first to deliver both in a custom interface designed around research practices. Most recently, Zotero has brought natural-sounding read-aloud voices to its users, enabling researchers to listen to their sources instead of reading them, and they can even annotate while they listen.

In contrast to flashy new features, Zotero’s most significant technical transformations have often been its least visible ones. Massive rewrites to accommodate underlying platform changes went unnoticed outside the developer community. Relocating Zotero’s cloud infrastructure to Amazon Web Services initially involved zero functional change on the outside, but it easily proved to be Zotero’s most significant technical move in two decades. When we launched Zotero in 2006, we stood up a small constellation of related web services: support forums, a documentation wiki, a blog, and a homepage. All ran on a small physical server running at George Mason University’s Center for History and New Media. Today Zotero’s twenty million users are served by a sprawling range of services hosted by AWS, only a few miles away from CHNM but unimaginably complex in comparison, facilitating collaboration in millions of public and private collaborative research groups, instantly converting text to speech in dozens of languages, and silently synchronizing billions of research items around the globe.

Institutional adaptation

Zotero’s technical successes also drew the attention of competitors. In 2008 Thomson Reuters, the owner of EndNote, sued George Mason University over Zotero. Even if ultimately nugatory, the lawsuit underscored the need for an institutional home with no daylight between its objectives and those of Zotero’s users. The following year we founded the nonprofit Corporation for Digital Scholarship to handle Zotero’s business operations. By 2014 Zotero’s core operations were funded entirely by revenue derived from its own services, and in that same year we transferred Zotero’s intellectual property away from the university. Along with the transition to AWS, these moves set the conditions for leaving behind the university environment which had nurtured Zotero in its early days but which was not built for the technological and administrative demands of a worldwide research platform.

Today known as Digital Scholar, our company consists of software teams responsible not just for Zotero, but for allied open-source software projects Omeka, Tropy, Sourcery, and PressForward. Zotero’s twelve full-time developers and designers work from seven countries around the world. Powered by revenue from paid services associated with Zotero and Omeka (and soon, Tropy), Digital Scholar can do things a university simply can’t: purchase equipment and services quickly, pay competitive salaries to technical staff, hire independent experts for legal and financial advice, and, most important, serve the interests of Zotero’s vast user community first and foremost.

To the credit not just of Digital Scholar but of George Mason University, these transitions passed unnoticed by Zotero users; today most probably have no idea where Zotero happens. This uprooting was also eased by remarkable consistency on Zotero’s human side: not only is Stillman Zotero’s lead developer today, another dozen former employees of the Roy Rosenzweig Center for History and New Media are current staff and board members at Digital Scholar. And the vast community of volunteer contributors to Zotero support forums, user documentation, codebase, and plugins has only continued to grow.

Twenty years young

Today at twenty, Zotero’s community is far younger than the project. Few are likely to remember AOL Instant Messenger, let alone Zotero’s cramped pane inside the Firefox browser, the only place Zotero once ran. More than half of Zotero’s twenty million users joined in the last three years. For them, Zotero is a recent tool, and the only Zotero they know is the one built in its second decade, not its first. This is the Zotero of iOS and Android devices, of Google Docs and collaborative annotation, of Read Aloud and accessibility. And it’s the Zotero of a nimble nonprofit corporation which has evolved and grown to serve its communities through mindful adaptation of new technology to research needs. Looking ahead, we renew the promise we made twenty years ago, that Zotero will be open and free, knowing that keeping it will take us in directions hard to imagine today. We’ve got some great stuff on deck already.

Leave a comment