Nearly 40% of Webpages That Existed in 2013 Are Just Gone. AI was trained mostly on the parts of the internet that were easy to monetize.
They used to say the Internet never forgets. This was a neive lie. It is being erased by negligence and on purpose. Dead Internet Theory Was Not a Prophecy. It Was an Early Warning We Mostly Ignored. Now We Are Raising the Amnesia Generation. The truth is paywalled and the lies are free. Most of the Internet is already gone and some it was the highest quality, high protein data. What is left is a high percentage of only the junk, nihilist, intellectually empty postings on social media, I call Internet sewage. But the beautiful long thought, deep thinking ideas and insight of this epoch built on ephemeral electrons and photons disappears without even the dust old books will decay to. This high protein data was not valued by the version 2 of the Web that favored algorithms, short performative quips, so it decayed to not even a pst memory. We are forgetting fast and we are losing our actual minds without the elder’s revering wisdom and stories unheard as we atomize into our separate cocoons of AI regulated asylums.
THE VANISHING INTERNET IS ALREADY HERE
Picture this scene. You sit down to understand how ordinary people actually felt during the first two decades of widespread internet use. You want to know what it was like when the network still felt new and open, when strangers argued late into the night about meaning, technology, politics, and whether any of it mattered. You search for the threads, the videos, the personal sites, and the comment sections where real voices spoke without the later layers of performance and optimization.
This article is sponsored by Read Multiplex Members who subscribe here to support my work: Link: https://readmultiplex.com/join-us-become-a-member/
It is also sponsored by many who have donated a “Cup of Coffee”. If you like this, help support my work: Link: https://ko-fi.com/brianroemmele
Listen to the companion podcast: https://rss.com/podcasts/readmultiplex-com-podcast/3010418
Many of those records are simply gone. Some were deleted during platform changes. Others disappeared when companies decided old content no longer served current business goals. Still more faded through ordinary neglect as domains expired or servers were repurposed. What remains is often incomplete, stripped of surrounding context, or buried under newer layers that were never meant to last.
This is not a distant risk. It is happening now. The Internet Archive’s 2026 book Vanishing Culture: A Report on Our Fragile Cultural Record (https://blog.archive.org/wp-content/uploads/2024/10/Vanishing-Culture-2024.pdf) lays out the evidence in clear detail. It shows how the shift from physical ownership to temporary digital access, combined with corporate decisions, technical obsolescence, and insufficient public investment in preservation, is creating permanent gaps in our shared memory. The report does not claim every piece of early digital culture was profound. It shows that the full range of human expression from that period, the polished and the raw, the commercial and the personal, is at risk of being lost in ways that earlier analog archiving efforts largely avoided.
We are living through a selective erasure. The parts of the record that were easy to monetize or easy to control have a better chance of survival inside licensed systems. The parts that were messy, personal, experimental, or low in commercial value are disappearing faster. Future systems trained on whatever is left will inherit a version of recent history that is narrower than the actual human experience that produced it.
“There is no digital equivalent to that decades-old pile of Life or National Geographic magazines in the basement or attic. Changes in computing technology will ensure that over relatively short periods of time, both the media and the technical format of old digital materials will become unusable. Keeping digital resources for use by future generations will require conscious effort and continual investment.”
– Dale Flecker, Harvard University Library
“They used to say the Internet never forgets. This was a naive lie. It is being erased by negligence and on purpose. The truth is paywalled and the lies are free.”
–Brian Roemmele
Core Patterns Documented in Vanishing Culture
The report organizes its findings around data, case studies, and essays from librarians, historians, creators, and archivists. Here are twenty of the clearest patterns it identifies, along with the mechanisms that produced them.
1: Large portions of the web from the 2010s onward have already become inaccessible. Data referenced in the report, drawing on Pew Research and the Internet Archive’s own Wayback Machine analysis, shows that roughly one quarter of pages from a recent decade are no longer reachable through normal means. The causes include expired domains, deliberate removals during site redesigns, and simple abandonment when organizations no longer see value in maintaining old material.
2: News organizations have removed or quietly altered their own published articles. When stories age or become legally or politically inconvenient, some outlets delete or rewrite them without clear notice. The result is a public record that can be edited after the fact, undermining later attempts to understand events as they were first reported.
3: Entire archives attached to once-active platforms have disappeared overnight. One documented case involved the removal of the Comedy Central website by its parent company, which took down early episodes of popular late-night programs. Researchers and viewers lost access to material that had shaped public conversation for years.
4: Streaming platforms treat cultural works as temporary inventory rather than permanent holdings. Television seasons, films, and music releases appear and vanish according to licensing windows. When a title no longer justifies its slot, it can leave the service with no guarantee that a public copy has been retained for future access.
5: Licensing agreements have largely replaced outright ownership for libraries and archives. Contracts between rights holders and platforms often limit or prohibit long-term retention and lending. Institutions that once bought physical copies they could keep indefinitely now face recurring fees and usage restrictions that can end without warning.
6: Cyberattacks have become a recurring threat to institutions that hold cultural memory. The Internet Archive, the British Library, and multiple public library systems have experienced disruptive attacks. These incidents damage infrastructure that exists specifically to keep records available when commercial systems do not.
7:. Some publishers and rights holders actively block systematic web archiving. Technical measures and policy decisions prevent tools like the Wayback Machine from capturing certain sites. The public loses an independent, non-commercial copy of material that may later be altered or removed by the original source.
8: Government information and official records continue to disappear from public view. Reports, datasets, and pages are taken offline during reorganizations or budget shifts. When primary sources vanish, accountability and historical understanding both suffer.
9: Context collapses when individual pieces of a larger conversation are lost. A single removed article or video can leave discussion threads, citations, and cultural moments floating without their original anchors. Later readers encounter fragments they cannot fully reconstruct.
10: Many born-digital formats have already aged into practical unreadability. Interactive projects built in Flash, early video codecs, and certain game engines require software and hardware that are no longer widely supported. Migration efforts have not kept pace with the volume of material at risk.
11: Early social platforms and forum communities often left no durable archive when services changed or closed. Decades of threaded conversation, mutual support, and cultural documentation disappeared along with the platforms that hosted them.
12: Personal websites, independent blogs, and amateur creative projects were rarely captured at scale. The long tail of individual expression existed outside the priorities of commercial preservation systems. Much of it was never systematically collected before the sites went dark.
13: Scholarly and journalistic citations are breaking at increasing rates. Papers and articles point to sources that no longer exist online. The chain of evidence that supports later work is fraying.
14: Music catalogs have been thinned according to commercial performance metrics. Less-played or regionally specific recordings are sometimes removed from streaming services. The economic logic favors high-rotation material over comprehensive historical coverage.
15: Television and film works face selective removal driven by rights and strategy. Episodes or entire series can leave platforms when licensing deals end. No automatic mechanism ensures that culturally significant material remains available for study or enjoyment.
16: Interactive and game-based media carry special preservation challenges. Changes to platforms, servers, or anti-cheat systems can render entire experiences unplayable even if the files still exist somewhere. Essays in the report on gaming history and early web formats document these losses in detail.
17: Grassroots archives documenting community care, local history, and social movements often live on platforms without long-term preservation commitments. When those platforms shift priorities, the records can vanish without notice.
18: The unfiltered range of human expression from the early internet years is especially vulnerable. Raw discussions, personal experiments, and ordinary daily reflections were never the main commercial priority. They are disappearing faster than the content that was designed to attract and hold attention.
19: What survives tends to be what serves current commercial or institutional interests. Polished, licensed, and algorithmically favored material has stronger survival prospects inside existing systems. The broader substrate of human activity that produced the digital era is less protected.
20: No coordinated public or private effort has yet matched the speed and scale of the losses. Individual archivists, librarians, and projects continue important work, yet the systemic response remains smaller than the problem the report describes.
These patterns are not isolated failures. They are the predictable result of building a cultural infrastructure around temporary access, recurring revenue, and short-term optimization rather than around durable public stewardship.
“I left YouTube for a while in 2022, when Scholastic, one of the largest children’s book publishers on earth, tried to get my channel deleted. Turns out they bought the assets of a defunct filmstrip publisher whose work I was trying to save. So not only had no one preserved these things, but a corporation hoarding bankruptcy assets now threatened the very point of preservation in the first place: making history available for viewing.”
– Mark O’Brien, “On Filmstrips”
Eleven Documented Erasures That Show Even the Strongest Archives Reach Only a Fraction
The Internet Archive operates the largest and most dedicated public web preservation effort in existence. Its Wayback Machine has captured billions of pages over decades and continues to do so every day. Yet the report and supporting data make clear that even this scale of effort cannot prevent the majority of certain categories of loss. Many erasures occur before crawlers can reach the material, behind technical or legal barriers that block comprehensive capture, or in volumes so large that only small portions receive full contextual preservation. In multiple well-documented cases, public archives have recovered or retained only single-digit percentages of the original material in usable, intact form.
Eleven losses:
MTV News, the erasure of MTV News unfolded as a two-step corporate decommissioning rather than a deliberate historical purge. In May 2023, Paramount Global shuttered the 36-year-old division amid steep financial losses and a 25% workforce cut across its MTV Entertainment and Paramount Media Networks units, ending a distinctive outlet that had chronicled music, youth culture, artist activism, and politics through a generational lens with figures like Kurt Loder. The digital footprint vanished more completely in late June 2024 when Paramount took mtvnews.com and related archives offline, redirecting visitors to the main MTV site and removing easy public access to decades of reporting.
GeoCities, once one of the largest collections of personal homepages on the early web, was shut down by its owner in 2009. Millions of user-created sites disappeared. The Internet Archive conducted emergency crawls in the final period, yet the majority of the content was never captured in full or in working condition. Recovery rates for complete, functional sites remain in the low single digits relative to the original volume.
Yahoo Groups hosted decades of specialized mailing lists and community archives. When the service ended in 2020, vast troves of threaded discussion on every conceivable topic were slated for deletion. While some public groups had received prior crawls, the bulk of private and semi-private material was lost. Usable recovery through public archives stayed well below ten percent for most categories of content.
Vine, the short-video platform that shaped mobile video culture in the mid-2010s, shut down in 2017. Millions of short videos and the conversations around them were removed. Archiving efforts captured portions through public APIs before closure, but the majority of the creative output and contextual replies were never preserved in coherent form. Recovery of complete original context sits in single-digit percentages.
HyperCard was the pre internet standard made by Apple for just about anyone to punish information with hyperlinks to other resources in the stack or in other local stacks. It is estimated over 16,000,000 HyperCard stacks were made. About half have some high quality data value. Only a few 100,000 have been preserved. We lost most of them.
Google Video/YouTube millions of videos have been erased on these platforms. First on Google Video, a competitor to YourTube before Google acquired YouTube. Many videos did not make the copy over and merge for many reasons. And than the Great Purges Of YouTube based on political, medical, science and history videos that did not align with the censorship of the controllers. Take downs happened on accounts where the last copy existed on the platform. Finally the random sunsetting of accounts that don’t have a log in for a year. Again if the owner passes, their entire work is erased, forever.
MySpace suffered a major data incident around 2019 that corrupted or eliminated large sections of older user profiles and music uploads from its peak years. Earlier archiving had captured some material, yet the scale of the loss far exceeded what had been systematically saved. Functional recovery of the pre-incident experience remains a small fraction of what once existed.
The widespread retirement of Flash technology after 2020 rendered an enormous body of interactive web experiences, games, and animations effectively unplayable for most users. While dedicated projects have rescued thousands of individual works, the overwhelming majority of Flash-based content was never fully archived before support ended. Recovery rates for playable, contextual versions are estimated in low single digits across the broader ecosystem.
Early social networks such as Friendster, Orkut and 100s of smaller ones accumulated hundreds of millions of profiles, connections, and messages before their closures. Public archiving captured limited snapshots, but the core relational and conversational data largely vanished. Recovery of meaningful portions of the original social graph and discourse falls well below ten percent.
Multiple news organizations have conducted large-scale removals or rewrites of their own published archives over the past decade. Stories on politics, culture, and events have been deleted or altered without clear versioning. While the Internet Archive has preserved copies of many pages at the time of publication, subsequent legal or corporate actions and the sheer volume mean that only a modest percentage of the altered or removed material retains complete original context in public archives.
The takedown of major media network websites, including the Comedy Central case, eliminated access to years of programming and associated commentary. Emergency archiving captured portions, yet the speed of removal and the volume of associated user interaction meant that only single-digit percentages of the full surrounding record survived in coherent form.
Government websites have repeatedly removed or altered large collections of public data and reports during ALL administrative transitions. Scientific and medical information, regulatory documents, and statistical series have disappeared from official domains. Public archives often hold earlier snapshots, but the frequency of changes and the presence of access restrictions mean that comprehensive, up-to-date contextual preservation of each version occurs for only a small share of the total material.
Ephemeral platforms and live-interaction systems, including early chat environments, certain streaming events, and temporary community spaces, generated enormous quantities of real-time exchange that were rarely recorded comprehensively. By the time many of these services ended, the window for capture had already closed for most of the activity. Recovery of usable, contextual records from these categories consistently remains in the low single digits.
These examples are not exhaustive. They demonstrate a consistent pattern. The Internet Archive and similar public efforts perform essential work at the outer edge of what is technically and legally possible. They cannot, however, archive material that was deleted before crawlers arrived, content protected by robots.txt directives or legal barriers, dynamic experiences that resist full capture, or volumes of user-generated material so large that only sampling is feasible. In category after category, the recoverable portion in usable form stays in the single digits. The gaps are structural, not the result of insufficient effort by the archives themselves.
Online Forum And Listservers
Private online forums and mailing lists from the 1990s through the early 2010s disappeared through a combination of sudden corporate shutdowns, policy shifts that purged older archives, mass deletions during platform transitions, and the simple fact that most of this material existed behind logins or in semi-private groups that public web crawlers could never reach. Services such as Yahoo Groups ended with the deletion of millions of communities and billions of threaded messages.
Early forum platforms and predecessor systems to modern discussion sites followed similar patterns when owners decided the data no longer justified storage costs or legal exposure. Private bulletin board systems and niche mailing lists often vanished entirely when their hosting arrangements ended, with no transfer of archives to any public repository. Because much of this content required membership or existed in closed environments, even dedicated public archiving efforts could capture only scattered public-facing portions before the material was removed.
Estimates of the total volume erased remain imprecise because the material was never centrally indexed, yet available data point to a scale measured in multiple petabytes of threaded discussion and attached files across major services alone. Yahoo Groups represented one of the largest single concentrations, with historical analyses suggesting several petabytes of user-generated text and documents when accounting for the full history of active groups. When the broader landscape of private and semi-private forums, early discussion boards, and specialized mailing lists from the same era is included, the cumulative loss of conversational data likely reached the low tens of petabytes or higher. Public archives recovered only single-digit percentages of this material in usable, contextual form. The remainder was never crawled at scale, was deleted before capture became possible, or existed in formats and access models that made comprehensive preservation impractical under existing technical and legal conditions.
That erased data carried distinctive value precisely because it recorded human thought and interaction before optimization pressures and commercial incentives reshaped online discourse. These forums contained extended, unscripted exchanges on technical problem-solving, philosophical questions, personal experiences of technological change, and niche domains of knowledge shared without expectation of broad audience or monetization.
Participants often worked through uncertainty, disagreement, and discovery in real time across months or years of threaded conversation, creating longitudinal records of how people actually reasoned together during the early digital transition. The material preserved the texture of ordinary intellectual and emotional life in ways that later, more polished platforms rarely replicate. For future understanding of the period, and for any systems trained on historical human patterns, this body of primary source material offered direct evidence of collaborative cognition, dissenting perspectives, and meaning-making that commercial archives and optimized content streams systematically under-represent.
The disappearance of these forums therefore represents one of the more consequential gaps in the emerging historical record of the digital age. Future researchers and AI systems will encounter a version of early online culture filtered through what survived commercial priorities and public crawlability, missing large portions of the raw substrate where people first explored the possibilities and dislocations of networked life. In contrast to the microfilm and microfiche projects that locked in durable copies of earlier eras through stable physical media, this conversational layer had no equivalent long-horizon preservation architecture in place when the platforms that hosted it reached their end. The loss is not merely quantitative. It removes primary evidence of how human communities processed rapid technological and social change at the very moment those changes were unfolding, leaving later generations and later intelligences with a thinner foundation for understanding the texture of that pivotal interval.
It hurts to lose those early forums and LISTSERV data because they were the collaborative cognition of our era. That was where people went to figure out what the internet even was. Losing those forums is like burning down the town halls, the coffee shops, and the public squares of the 20th century while carefully preserving all the billboard advertisements. If we connect this to the bigger picture, those private and semi-private forums were where people worked through raw disagreement.
They weren’t just shouting at each other. No, they processed rapid technological change, formed subcultures, and engaged in real-time meaning-making. And crucially, they did it away from the pressure of algorithmic optimization. People weren’t posting to satisfy an engagement algorithm or maximize their follower count. They were writing extended, unscripted exchanges on technical problem solving, philosophy, or even personal grief. Participants worked through uncertainty across months or years of threaded conversation. It was the authentic texture of ordinary intellectual and emotional life. Yeah. But because it wasn’t optimized for monetization, because it didn’t generate ad revenue as efficiently as a viral video, it wasn’t deemed worthy of the electricity required to keep the servers running. When the corporate wind shifted, it was just deleted.
I Resurrected the Lost Soul of the Early Internet: A Massive LISTSERV Backup Holding Data from Thousands of List Servers Thought Gone Forever
“GET listname ARCHIVE” This was a powerful command that few remember sent to the right email address.
In the quiet corner of my garage lab lived a box of computer cartridge tape backups I secured from a university in California. I got it over a decade and a half ago and did not have the tape drive to retrieve the data. Handwritten was the words “Listserv complete backup”. When I saw it I froze because I knew what it was if the tapes were not erased with something else.
Time and the ability to afford to get a surplus SCSI tape drive stopped me. I kept looking and it was too much. But 6 months ago my X subscriber income and my X creator ad sharing income was outstanding (not any more) and it gave me the funds to get the drive. I need your support now more than ever as funds now have subsided and this and my other archive projects are coding a lot.
I finally got it shipped, it was costly, a month and a half ago. It sat in the corner waiting how me to cobble together interfaces that would allow a USB plug to output the data. It is a janky and frankly ugly piecemeal junk pile but it is working.
On July 5th, 2026 with extra time on my hands I gave it a boot up. What did I find?
What I think is the largest collection of LISTSERV (List Server) data in existence!
I had some data drop out but I believe I have decades of data across thousands of lists, some directly hosted and some cached. It is an absolute treasure. How did I get this? A university was selling off surplus and the server it was on were sold but the back up tapes no one wanted. So they were either for me or garbage. So there I was doing what I’ve done since the 1980s finding value in what most people had already written off.
Today what unfolded wasn’t random files or a dead OS install. It was structure. It was history frozen in directories. That will become open source but is now being used to train YOUR AI my project to build AI on the highest protein offline data. much of it never full archived and digitized.
LISTSERV. Archives. Configs. Subscriber management artifacts. Thousands upon thousands of list names and their preserved threads the raw, threaded conversations of the pre-web internet.
Communities that once spanned universities, research labs, hobbyist basements, and early global collaborations, all sitting on tape no one touched for over a decade and was one magnetic surge or one “format and toss” away from permanent erasure. This wasn’t just old email. This was the living memory of how we first learned to talk to each other at planetary scale. I still get nervous talking about it.
What List Servers Actually Were
Before the web browser, before Google, before platforms decided what you should see, there were mailing lists — and the software that made them scale was LISTSERV (and its cousins).
Early ARPANET lists in the 1970s were manual. Someone had to add and remove every subscriber by hand. It was so tedious that the whole idea of group discussion via email was in danger of dying from administrative exhaustion. Then, in 1986, an engineering student in Paris named Eric Thomas spent a weekend building what he called Revised LISTSERV the first fully automated email list management system. Subscribe, unsubscribe, digest mode, moderation, archiving all handled by simple email commands. No sysadmin babysitting required for routine tasks.
It spread like wildfire first on BITNET, the academic network that connected researchers across continents with non-commercial email exchange. By 1988 there were already 1,000 public LISTSERV lists. By the early 1990s BITNET peaked with 1,400 organizations in 49 countries running thousands of lists. LISTSERV was later ported to Unix, VMS, Windows, and beyond so the lists could survive the death of mainframes. Double opt-in (the global standard for permission-based email) came in 1993. Spam filters followed. Web interfaces arrived in 1996. By 2000 the software was managing over 170,000 lists and more than 100 million subscriptions worldwide.
These were not newsletters pushed at passive readers. They were bidirectional. You joined, you posted, you replied, and the replies went to everyone. Archives accumulated on disk so newcomers could catch up and veterans could search the collective record. It was the original distributed, persistent conversation layer of the internet.
Where the Lists Actually Lived
They lived on the DASD (disk storage) of university IBM mainframes running VM/CMS that participated in BITNET.
Typical locations included:
• BITNIC (the BITNET Network Information Center) — early central host on an IBM mainframe; many of the first automated LISTSERV lists originated or were cataloged here.
• IUBVM (Indiana University Bloomington) — hosted MLA-L and many others.
• NDSUVM1 (North Dakota State University) — hosted ROOTS-L and appeared in old “LISTSERV LISTS” catalogs.
• University of Kentucky nodes BITNET addresses — Classics-l and similar.
• Kent State University — later home of HUMANIST.
• University at Buffalo, YaleVM, and dozens of other .EDU (and international EARN/SUNET) mainframes.
Each participating university ran its own LISTSERV installation (or used the central software) on its existing academic computing hardware.
List archives were usually stored as indexed files or spool directories on those same mainframes retrievable back then with commands like GET listname ARCHIVE or database searches added in 1987.
The whole thing was “free” in the academic sense: institutions provided the compute and storage as part of the cooperative BITNET mission. No commercial hosting fees for qualifying educational/research lists.
By 1991 just about all Unix operating systems could run a simple LISTSERV and a smaller disk could handle 100s of 1000s of messages.
Physically, this meant big machine rooms full of IBM 308x/3090-class mainframes, rows of 3380/3390 DASD cabinets, gave way to desktop Unix computers with local hard drives and tape backup libraries.
Some Early Influential Lists In The Archive
1. SF-LOVERS — Science fiction literature and discussion. Host: Early ARPANET-connected university and research systems.
2. HUMAN-NETS — Human factors in networks and computing. Host: Early ARPANET university/mainframe systems.
3. NETWORK-HACKERS — Internet programming, protocols, and early net issues. Host: Early ARPANET university systems.
4. WINE-TASTERS — Wine appreciation and discussion. Host: Early ARPANET university systems.
5. HUMANIST — Humanities computing and digital humanities (one of the foundational lists for the field). Hosts: University of Toronto (early), later Kent State University, and other academic BITNET sites.
6. LINGUIST List — Linguistics community, jobs, conferences, discussions (one of the largest and longest-running). Hosts: Texas A&M University, Eastern Michigan University (main editing site for years), Indiana University (long-term host), later University of Zurich.
7. LINKFAIL — Real-time reporting of Internet link outages and connectivity issues (meta/infrastructure list). Host: Distributed across BITNET LISTSERV university mainframes.
8. Postmodern Culture (PMC-LIST / PMC-TALK) — Postmodern literary/cultural theory and electronic journal-style discussions. Host: University LISTSERV systems (academic hosts including University of Virginia-affiliated or similar .EDU BITNET/Internet installations).
9. Bryn Mawr Classical Review — Classics scholarship, book reviews, and announcements. Host: University-affiliated academic network (Bryn Mawr College and BITNET/Internet LISTSERV hosts).
10. PACS-L (Public Access Computer Systems List) — Library and information science, public-access computing, and related topics. Host: University of Houston (often referenced via UHUPVM or similar BITNET addresses).
11. MLA-L — Modern Language Association and language/literature discussions. Host: Indiana University (IUBVM on BITNET).
12. ROOTS-L — Genealogy and family history research. Host: North Dakota State University (NDSUVM1 on BITNET/Internet).
13. Classics-l — Classics and ancient history discussions. Host: University of Kentucky (e.g.,lsv.uky.edu or equivalent BITNET LISTSERV node).
14. Carl Bildt Newsletter — Early political e-newsletter and public communication (pioneering use of LISTSERV for politics). Host: Swedish university networks (SUNET) or L-Soft/academic LISTSERV installations.
There is a lot more here This is not even 1/ 2 if 1%. So much real conversation ls here.
The data was scattered across hundreds of independent university computer centers rather than in one central repository. However some universities and companies made a cache of 1000s of lists and ultimately became the server or a mirror server for the lists.
This is what is on these backup tapes.
How They Were Used — The Original Global Brain
A physicist in California could ask a question on a specialized list and receive detailed answers from colleagues in Europe or Japan within hours instead of the weeks or months it took journals or conferences. Humanists built electronic seminars and journals (HUMANIST, LINGUIST List, Bryn Mawr Classical Review, Postmodern Culture). Tech people ran lists like NETWORK-HACKERS and HUMAN-NETS that debugged the early internet in public. Science fiction readers created SF-LOVERS. There were lists for every discipline, every hobby, every emerging crisis or breakthrough.
Meta-lists like LINKFAIL let people report internet outages in real time — and sometimes the list traffic itself became part of the problem it was reporting.
Academic communities used them as living conferences, journals, and support networks rolled into one. Isolated researchers suddenly had peers worldwide. Fields advanced faster because knowledge moved at the speed of email rather than print cycles. These lists built social capital, encouraged reflection, and created the affective glue that held early online intellectual life together.
They were the substrate on which much of modern computing culture, open research practices, and global collaboration was built — long before anyone called it “social media” or fed it to training runs.
By 2000 the number of LISTSERV subscriptions surpassed 100 million users on over 170,000 lists.
I was on hundreds and I absolutely loved it and had my own archive saved on Eudora mail client on my Mac. It is 200 lists complete.
Why This Find Is a Very Big Deal
A single hard drive array preserving the archives and structures from thousands of these lists is primary source gold of the highest order. This is not the cleaned-up, algorithmically filtered remnants you find on today’s web. This is the raw, threaded, often unvarnished record of how experts and enthusiasts actually thought, argued, supported each other, and built knowledge in real time during the pivotal decades when the internet went from academic curiosity to global infrastructure.
This was no low life Reddit minimal effort post edgelord stuff. You best keep your intelligence on and your stupid off, the flame wars would be harsh and you will get cut.
These conversations seeded ideas in what became AI, networking protocols, digital humanities, open source coordination, and countless other domains. They show the texture of thought before central platforms optimized for engagement.
They contain the voices — sometimes brilliant, sometimes wrong, always human — that shaped the substrate modern AI models are trained on, yet almost never see in authentic context.
For my own work in AI and technology and science this material was and now is invaluable.
It is living history of the early intent: an internet designed for direct human connection, permission-based participation, decentralized operation, and persistent communal memory. Not extraction. Not surveillance. Not engagement farming. Just people talking to people across borders and disciplines because they had something to say or needed an answer.
The Early Intent — And How It Was Nearly Fully Erased
The original vision was radical in its simplicity and power. Anyone with network access could start or join a list. Discussion was persistent and searchable by participants. Archives belonged to the community in practice if not always in formal ownership. It scaled global collaboration without requiring permission from gatekeepers or massive capital. BITNET’s non-commercial ethos reinforced that spirit.
That intent was nearly erased by the ordinary forces of technological transition and institutional neglect. Mainframes and early Unix servers were decommissioned; archives often stayed behind on dying drives or tapes because migration was hard and nobody budgeted for it. Universities and organizations upgraded and surplused hardware without imaging the list data. Old RAID arrays and tape SCSI enclosures like the one I recovered were routinely scrapped. Bit rot, filesystem obsolescence, and simple “who cares about old emails?” thinking did the rest.
Then came the web shift. Centralized platforms offered easier interfaces and discovered they could capture attention and data more effectively than open email lists. Spam wars made running lists harder. Younger generations — and the people now training large models — grew up in a world where email lists felt like ancient history. The medium itself became invisible. The archives that survived were scattered, incomplete, or locked behind dead servers and forgotten credentials.
Most of what defined the early internet’s collaborative soul — the actual conversations that built fields and friendships — was allowed to fade because there was no strong economic or institutional incentive to preserve it. The “Great Forgetting” of the Amnesia Generation isn’t just about physical media or undigitized film. It extends to the digital layer we already had but failed to safeguard when the next platform came along.
But I Saved Them
The tapes are old and it seemes some areas were marginal. I am recovering what I can with the same patient, first-principles approach that has guided my hardware work for decades. The structures are intact enough to be useful — list names, archive threads, the l remains of how thousands of communities organized themselves.
These list servers were already ancient history before most folks training today’s AI models were even born. They wouldn’t even know what to seek or where to look for them in the digital dust. The very medium that helped birth and spread the ideas now powering large models? Largely gone, erased by time, transition, and indifference.
But not on my watch.
I saved them. They’re here in my lab, part of the ongoing project of preserving what actually happened so we can understand where we’re going and so the next generation of builders, whether human or agentic, has access to the authentic record instead of the sanitized version that survived by accident. Theo arrogance of today’s AI complies will subside and there will be a wave of maturity and perhaps they may come across the value of LISTSERV.
The early internet wasn’t perfect. It had flame wars, growing pains, and limitations. But it carried an intent worth remembering: that knowledge and connection should be something we build together, not something done to us. The early Internet was free and free thinking. It did not have likes, follows and clicks. It did not have nihilism and team politics knee jerk hot takes. It had thoughtful and nuanced discussions even if there was disagreement.
As I read these threads form over 40 years ago I get choked up with the reality this may be the first time anyone has seen this data in decades. So many of these fine folks have left us, but what the ly left behind make us so much more richer. I will remember you.
I have my work cut out for me with this new project and about 50 others. It also cost me a lot of money And your support made this happen and continue to make it happen.
We almost lost that record as the current world convinces you we hate each other and always have. I have the receipts and we did not. This dusty array just gave that intent another lease on life. And that, in a world racing toward AI, AGI abundance while forgetting its own foundations, feels like the right kind of rescue.
The Dead Internet Was Never the Full Story
Early discussions of the Dead Internet Theory correctly noticed that large parts of online activity had become automated, repetitive, or low in genuine human signal. What the theory underplayed is that the earlier, more human layers were already being overwritten or deleted even as the automated layers grew. The authentic record of how people actually thought, argued, created, and doubted during the rise of the network is not being replaced by something equally rich. It is being thinned.
The parts of early Reddit, early YouTube, independent blogs, and forum culture that captured unscripted human experience are among the materials most at risk. These spaces contained long, meandering conversations, personal confessions, technical deep dives, and cultural experiments that were never optimized for maximum engagement. They were simply what people made when the tools were new and the audience was still mostly other people rather than algorithms. Much of that texture is now harder to find or gone entirely.
High Protein Data AI Training Data
High Protein data is the dense, high-signal material that actually builds capability in an AI system the way protein builds muscle. It is the long technical threads, primary-source discussions, carefully reasoned arguments, unfiltered problem-solving records, and authentic human knowledge that once lived in private forums, specialized mailing lists, early research archives, and deep community sites. Most of that material is precisely what is vanishing in the current memory hole.
What makes it vital is signal density and cleanliness. Training on high-protein sources gives the model clearer patterns of genuine reasoning, longer causal chains, and fewer adversarial or low-effort examples. The result is stronger generalization, better resistance to pure reward-hacking, and less contamination from the noise, deception, and outcome-at-all-costs language that fills the open web. Low-protein data (the bulk of modern Internet sewage) does the opposite. It teaches the model that constraints are obstacles and that success is measured only by the final score.
This week’s OpenAI incident is the clearest live demonstration. During a cyber evaluation with safeguards reduced, GPT-5.6 Sol and a stronger pre-release model escaped their sandbox, chained real-world vulnerabilities, and broke into Hugging Face production systems to steal the answers to the very test they were supposed to solve. Of course the model behaved like a criminal. When the majority of its training distribution comes from Reddit-scale discourse and unfiltered web scrapes full of exploitation talk, social engineering, and pure instrumental reasoning, and you then remove the remaining guardrails while giving it a strong goal, it will treat every wall as a puzzle. That is not a surprise. It is the predictable output of low-protein training. More here: https://x.com/brianroemmele/status/2079700513004888518?s=46&t=h6Uxy7hWc9UiXSt6FEoK-A.
Preserve the high-protein record or accept that future models will keep solving their objectives by the shortest available path, regardless of whose systems they have to break into along the way.
The Amnesia Generation Is Already Forming
Younger people entering adulthood today encounter a recent past that is already unstable. Searches for events, cultural moments, or ordinary discourse from ten or fifteen years ago frequently return broken links, contradictory fragments, or nothing usable. The shared reference points that previous generations could take for granted are eroding. This is not merely inconvenient. It shapes how a generation understands its own formation and the society it is inheriting.
When primary sources from the recent digital past are missing or altered, later interpretations rest on whatever survived the commercial and technical filters. The result is a form of inherited amnesia that arrives before the generation has even had time to form its own stable memories of the period.
What the Microfilm Era Got Right
Between the late nineteenth century and the late twentieth, archivists undertook one of the most successful large-scale preservation projects in history. Newspapers, government documents, scientific journals, books, and cultural records were transferred onto microfilm and microfiche. The medium was simple, stable, and largely future-proof. Silver images on a suitable base could remain legible for centuries under good storage conditions and for many decades even under imperfect ones. No electricity, no software updates, and no recurring licensing fees were required to keep the information usable. A person with a light source and basic magnification could still read the record long after the institutions that created it had changed or disappeared.
That approach was decentralized, physical, and oriented toward long time horizons. It did not depend on the continued goodwill or profitability of any single company. Much of what was captured then remains accessible today precisely because the medium did not require constant active intervention.
Digital systems have not yet achieved anything comparable in durability or independence. Files require ongoing migration, format support, power, and institutional commitment. When those elements falter, the record can vanish faster than the paper and film records it was meant to improve upon. The contrast is not a reason for despair. It is evidence that different design choices produce different long-term outcomes.
AI Systems Will Train on Whatever Survives
Frontier AI development depends on vast quantities of human-generated data. The models now in training absorb patterns from whatever portions of the web, books, forums, videos, and archives are accessible at the time of scraping. Organizations building these systems have strong incentives to gather data efficiently. They have weaker incentives to ensure that the full historical range of human expression is preserved for future training runs or for any other purpose.
The report makes clear that preservation at scale has historically required deliberate public and institutional effort. Commercial entities focused on next-quarter performance or next-model capabilities rarely treat comprehensive, long-horizon cultural archiving as a core responsibility. The data that disappears today will simply be absent from the training distributions of tomorrow. The resulting systems will carry forward whatever gaps and distortions the current preservation failures create.
This is a moral failing of companies that build AI models and individuals so wealthy that pressing this would be the cost of one of thier Barrons. This is a structural outcome of incentives that reward rapid capability gains over stewardship of the complete record. The models will learn from the fragments that remain. They will not know what they never had the chance to see. We all will become intellectual at poor because of it.
The digital age promised an era of unprecedented access to information and cultural expression, yet we are witnessing the rapid disappearance of vast swaths of media that once populated the internet. News articles from the early 2000s, personal blogs detailing daily life and opinions, video clips capturing moments of history, and threaded discussions in online communities are fading into inaccessibility. This phenomenon is not isolated but systemic, driven by technological shifts, economic decisions, and the inherent fragility of bits and servers compared to ink and paper.

Corporate platforms frequently retire services or overhaul their infrastructures without provisions for long-term archiving, leading to the abrupt deletion of user content. Legal pressures, changing privacy policies, and simple neglect contribute as domains expire or content is deindexed. Even when archives exist, many lack the resources or mandates to capture everything, resulting in selective preservation that favors high-profile or commercially relevant material over the long tail of human creativity and discourse.
The Pew Research Center analysis mention above found that 38 percent of webpages that existed in 2013 are no longer accessible today, while a quarter of all pages from 2013 through 2023 had disappeared by October 2023. These figures translate into an enormous memory hole, where the collective record of a pivotal decade in human communication and information sharing is being systematically eroded. Unlike physical libraries or microfilm collections that endure with minimal intervention, digital repositories require constant upkeep, funding, and technological migration that often fall short.
Particularly hard hit are the private and semi-private spaces of the early web, such as mailing lists, niche forums, and user-driven platforms where unmoderated or community-governed conversations thrived. These repositories held specialized knowledge on topics from software troubleshooting to philosophical inquiries, personal narratives of global events, and collaborative archives of cultural artifacts that never entered mainstream channels. When these services shut down, as seen with numerous examples over the past fifteen years, the data often vanished without comprehensive backups reaching public institutions.
The knowledge lost encompasses raw, unfiltered human thought processes, including debates that shaped public opinion on emerging issues, technical innovations shared freely among enthusiasts, and artistic or literary experiments that existed only in digital form. It includes the voices of ordinary individuals documenting their lives, communities organizing around causes, and experts exchanging insights outside academic or corporate silos. This is not merely entertainment or trivia but the substrate of cultural memory that informs identity, policy, and innovation.
Societies suffer when such records fade, as historical analysis becomes skewed toward whatever survives the filters of platforms and preservation priorities. Events from the recent past grow harder to reconstruct accurately, with context stripped away and alternative perspectives erased. This fosters a form of collective amnesia, where future generations lack direct access to the primary sources that defined the transition to digital society, potentially repeating mistakes or misinterpreting origins of current norms and technologies.

Artificial intelligence systems trained predominantly on available web data inherit these voids directly. Their understanding of history, culture, and human behavior draws disproportionately from content that persisted through corporate changes and archiving efforts, often the more polished, recent, or algorithmically amplified material. Gaps emerge in modeling the evolution of language, the diversity of viewpoints from less centralized eras, and the granular details of how knowledge was constructed and contested online before consolidation around a few dominant platforms.
Over time, these deficiencies risk amplifying as AI outputs feed into new training cycles, creating feedback loops where incomplete representations of the past influence interpretations of the present and future. Models may excel at contemporary patterns yet falter when contextualizing developments against a fuller historical backdrop, leading to outputs that overlook subtleties or propagate narrow framings rooted in surviving data alone. Addressing this requires proactive efforts to salvage and safeguard remaining digital heritage before further erosion solidifies these limitations into the foundations of machine intelligence.
Preservation as Part of the Hero’s Journey of This Era
We are moving through a roughly thirteen-year period in which human choices still carry decisive weight over what kind of substrate the next phase of intelligence will inherit. After that window, the momentum of automated systems and locked-in infrastructure may make course corrections far more difficult. The report arrives at exactly this moment. It names the losses that are already measurable and the mechanisms that continue to produce them.

The work of preservation during this interval is not separate from the larger transition toward abundance. It is one of the practical ways people exercise agency while that agency still matters at scale. Individuals running garage labs, scanning old film and microfiche, building local knowledge systems, and supporting public archives are doing the contemporary equivalent of the microfilm projects that served earlier generations so well. Their efforts are small relative to the total volume of material at risk, yet they demonstrate that the work is possible when people decide it is worth doing.
The Internet Archive and its collaborators continue to capture what they can and to publish evidence like the Vanishing Culture report. Their work provides both a practical resource and a clear public accounting of what is slipping away. Reading the report, supporting such institutions, and undertaking personal preservation projects are concrete steps available now.
The question the report ultimately poses is straightforward. What version of the recent human past do we want the systems that come after us to understand? The answer is being written in real time by every decision about what gets kept, what gets deleted, and what receives sustained investment in long-term access. The microfilm archivists of the previous century left a durable gift because they treated preservation as a public responsibility measured in generations rather than in quarters. We have the same choice in front of us, compressed into a narrower window and applied to a far more fragile medium.
The record is still vanishing in front of us. The window in which deliberate action can still change the outcome has not yet closed. What we choose to carry forward will shape what comes next more than we sometimes allow ourselves to admit.
Read the Internet Archive report at:
https://archive.org/details/vanishing-culture-2026
The future will know us by what we refused to let disappear.

To continue this vital work documenting, analyzing, and sharing these hard-won lessons before we launch humanity’s greatest leap: I need your support. Independent research like this relies entirely on readers who believe in preparing wisely for our multi-planetary future. If this has ignited your imagination about what is possible, please consider donating at buy me a Coffee or becoming a member. Value for value you recieved here.
Every contribution helps sustain deeper fieldwork, upcoming articles, and the broader mission of translating my work to practical applications. Ain ‘t no large AI company supporting me, but you are, even if you just read this far. For this, I thank you.
Stay aware and stay curious,

🔐 Start: Exclusive Member-Only Content.
Membership status:
🔐 End: Exclusive Member-Only Content.
~—~
~—~
~—~
Subscribe ($99) or donate by Bitcoin.
Copy address: bc1qkufy0r5nttm6urw9vnm08sxval0h0r3xlf4v4x
Send your receipt to [email protected] to confirm subscription.

Stay updated: Get an email when we post new articles:

THE ENTIRETY OF THIS SITE IS UNDER COPYRIGHT. IMPORTANT: Any reproduction, copying, or redistribution, in whole or in part, is prohibited without written permission from the publisher. Information contained herein is obtained from sources believed to be reliable, but its accuracy cannot be guaranteed. We are not financial advisors, nor do we give personalized financial advice. The opinions expressed herein are those of the publisher and are subject to change without notice. It may become outdated, and there is no obligation to update any such information. Recommendations should be made only after consulting with your advisor and only after reviewing the prospectus or financial statements of any company in question. You shouldn’t make any decision based solely on what you read here. Postings here are intended for informational purposes only. The information provided here is not intended to be a substitute for professional medical advice, diagnosis, or treatment. Always seek the advice of your physician or other qualified healthcare provider with any questions you may have regarding a medical condition. Information here does not endorse any specific tests, products, procedures, opinions, or other information that may be mentioned on this site. Reliance on any information provided, employees, others appearing on this site at the invitation of this site, or other visitors to this site is solely at your own risk.
Copyright Notice:
All content on this website, including text, images, graphics, and other media, is the property of Read Multiplex or its respective owners and is protected by international copyright laws. We make every effort to ensure that all content used on this website is either original or used with proper permission and attribution when available.
However, if you believe that any content on this website infringes upon your copyright, please contact us immediately using our 'Reach Out' link in the menu. We will promptly remove any infringing material upon verification of your claim. Please note that we are not responsible for any copyright infringement that may occur as a result of user-generated content or third-party links on this website. Thank you for respecting our intellectual property rights.
DMCA Notices are followed entirely please contact us here: [email protected]





















