From the archive · August 5, 2026
The login wall
On the difference between saving the July revolution and preserving it.
By the July Archive team
Not long ago, we sent a browser out to save eight Facebook posts from the July uprising. Not a screenshot tool: a real browser, headless, driven by archival crawler software, loading each page the way your phone loads it and writing down everything the server said. It came back with eight records. One contained zero characters of text. The other seven were identical copies of the first. Facebook had served our crawler its login wall eight times, and the crawler had preserved that wall faithfully, in a tamper-evident package, with checksums.
We shut the job off and wrote down what happened. The posts are presumably still up there behind the wall, for whatever a platform’s “still” is worth. Capturing them will mean a curator logging in as a person and saving what they actually see, which is slower. That’s fine. Archiving is mostly made of slow.
We tell this story first because it’s embarrassing. Preserving the record of the জুলাই গণঅভ্যুত্থান, the July mass uprising, is harder than it looks, and the hard parts don’t show until you hit them.
In the two years since the uprising, preservation sites have appeared the way flowers appear at a shrine. We’ve counted more than a dozen. Three of them, ours included, share the very name July Archive; we don’t mind. Students scan posters. Families upload photographs of the dead. Somebody keeps a playlist of the protest rap. The impulse behind all of it is right, and worth honoring: people watched history happen on their phones and understood, correctly, that the phones would not keep it.
But most of these sites are websites about the revolution, not archives of it. That sounds like pedantry. It’s the whole subject.
Uploading is not preserving.
There is an international standard for what an archive is: ISO 14721, the OAIS reference model, drafted by the space-data community, people who need a probe’s telemetry readable in fifty years, and adopted since by libraries everywhere. Its definition is bracing. An archive is an organization that has formally accepted responsibility for preserving information and keeping it usable. An organization. Not a folder. A set of promises. Provenance is the promise that you know where a thing came from, who captured it, and when. Fixity is the promise that it hasn’t been altered or quietly rotted. Persistence is the promise that its name and address will outlive any particular server, domain, or team. Add rights and credit, and, least glamorous of all, a plan for what happens when you burn out, lose your funding, or die. Preservationists keep a self-assessment grid, the NDSA Levels, whose tiers read like a catechism: know your content, protect it, monitor it, sustain it.
None of this takes exotic technology. It takes paperwork, and the paperwork has formats. Let us explain them the way we’d explain them over cha.
When the Wayback Machine saves a page, it doesn’t take a picture. It writes a WARC file (for Web ARChive, an ISO standard since 2009), which records the whole transaction: the request the crawler sent, the headers the server answered with, the exact timestamp, and the raw bytes of the page, plus cryptographic digests so tampering after capture is detectable. A screenshot is a picture somebody took, and pictures can be edited. A WARC is the original bytes plus the paperwork saying where and when they came from, and anyone with free software can replay it and check the math. Its younger sibling, WACZ, packs WARCs into one portable file with a manifest of SHA-256 hashes; change a single byte and the manifest breaks. If the WARC is the evidence, the WACZ is the sealed evidence bag.
The case for it fits inside another July. On July 17, 2014, Igor Girkin, a separatist commander in eastern Ukraine, boasted on social media about the downing of a military plane, shortly before news broke that the plane was Malaysia Airlines Flight 17, a civilian airliner carrying 298 people. The post was deleted within about two hours. The Wayback Machine had already captured it, and the capture became evidence cited around the world. A page that lived two hours has outlived a decade.
Checksums answer a quieter threat. Storage fails silently: disks and controllers flip bits and report nothing. In 2007, physicists at CERN checked 8.7 terabytes of their own data, 33,700 files, and found 22 silently corrupted, about one in fifteen hundred, with no error raised anywhere. The file opens fine and is quietly wrong. At archive scale you don’t ask whether that will happen. You ask whether your audit will catch it. So archives record each file’s fingerprint on arrival, re-check on a schedule, and keep a second copy to repair from. Librarians compressed the doctrine into an acronym, LOCKSS: lots of copies keep stuff safe.
Persistence has machinery too. An ARK, an Archival Resource Key, is a permanent name for a digital object, issued under an institutional number registered with the California Digital Library; more than 1,700 organizations hold one, and the registry is itself mirrored at national libraries in the United States and France. The point of an ARK is that it’s attached to an institution rather than a hostname. Domains die of redesigns and unpaid invoices. When a website moves, the resolver gets re-pointed and every identifier keeps working. Interoperability has machinery as well: OAI-PMH, a plain six-command protocol from 2002, lets any library on earth harvest a catalog’s metadata by machine. It’s the plumbing under Europeana and the Digital Public Library of America. An archive that speaks it stops being an island and becomes a node other institutions can copy.
Why does July need all this ceremony? Because the July record is contested, and it is disappearing while people argue over it.
The disappearing is measurable. Pew Research Center sampled a decade of the web and found 38 percent of the pages that existed in 2013 unreachable by October 2023; even 8 percent of pages that existed in 2023 were already dead by that October. A Harvard team checked the web links cited in United States Supreme Court opinions and found half no longer led to what the justices had cited. Platforms delete fastest of all. When Human Rights Watch audited the Syrian Archive’s collection in 2020, 21 percent of the 1.75 million conflict videos it had preserved from YouTube were no longer on YouTube. Moderation software built to purge extremist propaganda cannot tell it from evidence. Nobody has to be malicious. Deletion is what platforms do.
The contest is measurable too, because the counts refuse to agree. Tallies circulating in public run from a draft government list of around seven hundred dead to one memorial site’s claim of more than eight thousand martyrs. The United Nations human-rights office’s fact-finding report of February 2025 said as many as 1,400, and that is the number our homepage carries; the report itself sits in our harvest inventory as a captured source. The stat tile still owes its reader a citation right beside the number. We know. Under those numbers are names, and eventually courts and historians will need material that holds up. A screenshot proves little; it can be doctored over lunch. A hashed WARC with a capture record and a chain of custody is a different order of thing. There is an international standard for exactly this, the Berkeley Protocol on Digital Open Source Investigations, published by the UN human-rights office with UC Berkeley: hash content the moment you acquire it, preserve the metadata, document every step. The precedent arrived in August 2017, when the International Criminal Court issued its first arrest warrant built substantially on social-media video, against a Libyan commander accused of thirty-three murders. Some of the videos had been posted by his own brigade.
An archive of a contested revolution carries one more obligation, stranger than the rest: it has to be usable by people who distrust each other. This is why the Historian, our question-answering interface, attributes every claim to a source, surfaces contradictions without resolving them, and declines to answer who was right. When it quotes a Bengali source in an English answer, it gives the original words and a labeled translation, never a silent paraphrase. This isn’t squeamishness. The evidence has to work for everyone, including people we disagree with, or it stops being evidence and becomes advocacy with footnotes.
Hold the upload sites up to the promises and you can see, gently, what’s missing. We say this as allies. The material on those sites is precisely what we want to survive; the criticism is of method, never of motive. Most run on consumer publishing tools, so an “item” is a blog post with an upload date at best: no capture date, no original URL, no chain of custody, often no rights statement, and a photographer’s credit surviving, when it survives, inside a filename. Consumer platforms recompress each image and strip its metadata on the way in, and the favorite intake channels are web forms and Facebook pages, which means material rescued from platform deletion is being stored back onto the platforms. No checksums, so silent corruption will never be noticed. No export path, no API, so no other institution can mirror the collection. One server, usually one founder. Some sites render entirely in JavaScript from an empty HTML shell, so the Wayback Machine sees nothing when it visits. The archives themselves cannot be archived.
Two years out, the decay is under way. As we write this, one project’s domain already redirects to a parked host. A government museum’s chronology page, still indexed by Google, returns a 404. And when our crawler tried the official portals from outside Bangladesh, it could not reach them at all, which leaves the national record invisible to the diaspora and to international investigators. The sites built to fight forgetting are being forgotten first.
A preservation site with no preservation plan is link rot with a mission statement.
This has all happened before. GeoCities held tens of millions of hand-built homepages; Yahoo bought it for about $3.6 billion in 1999 and switched it off in October 2009; a volunteer brigade spent six frantic months crawling what it could before the lights went out. Nobody plans to be GeoCities. That is what the plan is for.
So, concretely, what we do. A working archivist would find the list unremarkable, which is the highest compliment we’re chasing. Every automated crawl we run writes a WARC. JavaScript-heavy pages are captured by a real browser into WACZ packages, and curators can capture at-risk pages by hand with ArchiveWeb.page and upload the file; either way the capture is kept byte for byte, and readers can replay it in their own browsers using software we host ourselves, because an archive that may someday be read under hostile network conditions shouldn’t depend on somebody’s CDN. Crawling has taught us humility: early on, two-thirds of what we’d gathered turned out to be duplicates of itself, and a domain-scoped crawl once published roughly 471 unrelated articles before we learned to fence a crawl to its paths. We wrote both mistakes down. Every original and every WARC gets a SHA-256 checksum on arrival, and every Monday at 05:00 UTC a scheduled job re-downloads a sample of up to two hundred stored objects, re-hashes them, and raises an alarm on any mismatch. Every published source carries an ARK under our registered number, 19160, minted so the identifier survives moves and redesigns; append ?info to any of them and a machine answers with the record, titles in both languages. The whole published catalog is harvestable over OAI-PMH, in Dublin Core, by any library that wants a copy. And every night the entire store is copied to a second bucket under a separate account, with deletions quarantined rather than propagated, so a compromise of the primary cannot reach the backup.
And the human part. We’ll be precise, since precision is the habit we’re arguing for: harvests from seed lists we’ve already vetted can publish automatically, so we won’t claim a human reviews everything. The claim we can defend is that every publication decision, human or automated, leaves an auditable record in an append-only trail, with explicit neutrality and rights sign-offs on every human publication, and that thin captures wait for human eyes. Attribution is mandatory on every source, free to use with due credit, and the credit follows the material into search results, cited answers, the cite-this widget on every source page, and the machine-readable record behind each ARK. There is a takedown path that requires reasons in writing, writes its audit record before it deletes, and then really deletes, down to the archived WARC when nothing else references it and, by runbook, the copy in the replica. A rights regime without a route to erasure isn’t one.
As of this writing, that machinery guards 68 published sources and a timeline of 14 entries. Modest numbers, deliberately. We left eight days out of the chronology because the sourcing was too thin, and of the sixty-nine protest-rap tracks on our list, Hannan’s “Awaaz Utha” among them, we’ve published the eight whose significance we could document; the lyrics stay in the recordings, where a curator decided they belong. We would rather be small and checkable than large and vague. The promise that worries us most is the last one, the plan for outliving ourselves. That work never finishes.
Brewster Kahle, who founded the Internet Archive, used to say the average life of a web page was about a hundred days. The figure was folklore even when he said it. The direction was not.
The people of Bangladesh carried this history through an internet shutdown. Losing it now to a lapsed domain renewal would be the bitterest joke of all.
So here is the invitation. If you run one of the July collections, we’re glad you exist, and we want your material to outlive us both. Write WARCs instead of screenshots; the tools are free. Hash your files. Register a persistent-identifier namespace; it costs nothing. Put your metadata where machines can harvest it. And if that is more infrastructure than one exhausted volunteer can carry, bring us your material. Bring us your dead links, your unlabeled photographs, your folder of screenshots nobody can verify anymore. We will preserve what you collected with its provenance intact and your name on it, and it stays yours to take elsewhere, because an archive you can’t leave is just another wall.
Somewhere behind a login wall, eight posts from a Bangladeshi July are still there. For now.