Login
Sign Up
Woofun AI reports that the auction of Spirit Airlines' bankruptcy assets concluded with Google securing a $10 million bid for the airline's entire corporate data archive, marking a definitive shift in how digital legacies are valued during corporate dissolution. This transaction, which saw competitor Mercor outbid by $2.5 million with a $7.5 million offer, establishes a new precedent where internal employee communications are no longer treated as liabilities to be scrubbed but as high-value training assets for artificial intelligence companies. The sale encompasses a vast repository of organizational memory, transforming what was once considered ephemeral digital waste into a critical resource for the next generation of enterprise software.
The auction details reveal a granular valuation of corporate communication history, with the total asset package divided into three distinct categories. The first category comprised approximately 100 million employee emails, representing years of formal internal correspondence. The second category consisted of 500 million messages from Microsoft Teams, bringing the total volume of direct communications to 600 million messages. The third category included a broader array of operational artifacts: calendars, spreadsheets, financial databases, project files, operational records, and various internal software tools.
Notably, passenger records and frequent flyer information were explicitly excluded from this sale, ensuring that customer-facing data remained separate from the internal workflow data being acquired. To contextualize the sheer volume, if a person speaks 100 sentences per day, it would take over 16,000 years to articulate all 600 million messages. At the final price of $10 million, each message was valued at approximately 1.67 cents. In American colloquial terms, where a one-cent coin is called a "penny" and often ignored if dropped on the ground, a single Teams message from a Spirit Airlines employee was worth just 1.5 pennies. This micro-valuation highlights the commoditization of individual data points when aggregated at scale.
Spirit Airlines' trajectory from industry benchmark to bankruptcy provides the backdrop for this data sale. Founded in 1980 in Detroit as a trucking company named Charter One, the entity transitioned to air travel and was renamed Spirit in 1992, a name chosen to signify "soul." Over the subsequent 34 years, the airline established itself as a leader in the ultra-low-cost carrier segment. At its peak, Spirit operated a fleet of 205 Airbus aircraft, executing around 300 flights per day and generating approximately $5 billion in revenue in 2024.
However, this operational scale was accompanied by significant financial distress, with a net loss of around $1.2 billion in 2024 and a debt burden of about $9 billion. The collapse accelerated in May 2026, when the company announced the suspension of all flights one late night. The following day, approximately 17,000 employees learned of their job losses through news reports rather than direct notification. The company entered Chapter 11 bankruptcy, undergoing liquidation without a traditional bankruptcy trustee taking over; instead, the company itself sold off remaining assets under court supervision.
Sensitive data, including employee emails and Teams messages, required processing by an independent agency to remove names, email addresses, and other identifying information. This agency was selected by Google, which also covered the associated costs. As far as public reports indicate, there was no prior precedent for bankrupt companies selling internal data to AI companies, making this deal potentially the first of its kind in history. American tech media outlet Gizmodo captured the sentiment of the era with the headline: "Spirit is Dead, But Its Data Will Haunt Google's Servers for Generations."
The emergence of this market is driven by two primary structural shifts in the technology sector. First, the rate of corporate failure has increased significantly, while capital has become increasingly concentrated. In the first quarter of 2024, the failure rate of U.S. startups rose by 58% compared to the previous year, while the number of active venture capital firms dropped by 62% from peak levels. Capital availability has not decreased overall but has shifted decisively toward the AI sector. In 2024, U.S. AI startups raised $97 billion in funding, setting a new record.
This concentration means that companies outside the AI sphere, which might have previously survived on extended funding rounds, are hitting liquidity walls earlier. A notable example is Tally, a fintech company backed by Andreessen Horowitz, which announced its shutdown in August 2024. Tally had raised a total of $172 million and reached a valuation of $855 million by its Series D round, yet it failed to secure further funding. This pattern of high-valuation failures creates a supply side of distressed assets, including valuable internal data, that becomes available for acquisition.
The demand side of this equation is fueled by data scarcity and the rise of enterprise Agents. The appetite for training data in large language models has reached unprecedented levels. GPT-4's training dataset consists of around 13 trillion tokens. For comparison, Google Books scanned approximately 40 million books over more than 40 years, translating to about 4 trillion tokens. This means the data used to train GPT-4 once is equivalent to over three times the entire corpus of Google Books. Epoch AI calculated that high-quality language data from books, news, and wikis would be exhausted around 2026.
With high-quality public internet text rapidly depleting, synthetic data is accounting for an increasing share of training inputs. Consequently, AI companies are forced to look beyond the public internet for new data sources. The market for AI training datasets is projected to reach $9.7 billion by 2030, with the total scale potentially reaching $67.5 billion when including various types of approvals. Gartner predicts that by 2026, 40% of enterprise applications will incorporate task-oriented AI Agents, up from less than 5% the previous year. This surge in Agent adoption creates a specific demand for data that reflects real-world work processes, which are abundant in the internal communications of defunct companies.
Woofun AI data shows that a new class of market participants has emerged to facilitate this trade: dissolution service providers. SimpleClosure, founded in 2023, specializes in helping startups "die with dignity." The company raised $1.5 million in its pre-seed round and then $15 million in its Series A round in May 2025, led by TTV Capital. Even Carta, a platform that manages equity and corporate affairs for many U.S.
startups, ceased offering its own dissolution services and instead invested in SimpleClosure, delegating these client needs to the specialist. By October 2025, SimpleClosure had handled the 'funerals' of over a thousand companies, earning the moniker "A Better Way To Fail" from Crunchbase. This reflects a cultural shift among American entrepreneurs who advocate for "failing fast," but SimpleClosure argues that speed alone is insufficient; dignity and asset recovery are also necessary.
The company even offers a pricing calculator on its website to estimate the cost of corporate dissolution. In April 2026, SimpleClosure launched Asset Hub, a platform dedicated to handling intangible assets left behind after company closures. This hub lists brands, software, customer lists, and, for the first time, internal work data such as Slack records, emails, and Jira work orders. This indicates that the market for internal data from failed companies was already forming before the Spirit Airlines deal, albeit limited to startups and smaller transactions.
The infrastructure for trading this data is being built by specialized platforms and legal entities. SimpleClosure delegated the data sales process for one case to Protege, a data trading platform focused on licensing AI training data. Protege received $30 million in funding led by Andreessen Horowitz in January 2026, with founder Bobby Samuels at the helm. Initially focused on medical imaging, where it gathered millions of images within 30 days for pre-training, Protege is now expanding into corporate data. In a real-world example, when transcription and captioning company cielo24 shut down, it used SimpleClosure to sell Slack messages, internal emails, and Jira work orders accumulated over 13 years.
CEO Shanna Johnson later told Forbes that these datasets were eventually sold for hundreds of thousands of dollars. For a company already decided to close, data that would have been deleted became a recoverable asset. This process involves bankruptcy courts and liquidation lawyers, who have traditionally inventoried physical assets like aircraft and furniture. Under U.S. bankruptcy law, internal data can now be included as intangible assets in the bankruptcy estate and sold under court supervision. Law firms like Redgrave LLP are establishing specialized teams to handle data preservation, collection, and organization in bankruptcy cases.
Additionally, e-discovery service providers such as KLDiscovery, Epiq, and Consilio handle the technical execution, collecting, organizing, storing, and reviewing corporate data from emails, Microsoft Teams, and SharePoint. The industry has a mature pricing structure, with EDRM regularly releasing pricing surveys that include billing units for data collection per GB, storage per GB per month, and document review fees.
The value of this data lies in its ability to train enterprise Agents in complex workflows. Unlike standard large models that learn from public web pages providing knowledge and finished content, enterprise Agents need to understand judgment, collaboration, error correction, and execution processes. Public data cannot replicate the nuances of how a request is made, how multiple people discuss it, how tasks are assigned, and how issues are resolved. These details are preserved in internal company emails, chat records, work orders, and project documents.
Training data is shifting from simple answers to complete task execution records. In the past, a single code training sample might have contained only a few hundred tokens, representing changes to a few lines of code. Now, a single Agent training sample often includes the entire process of understanding requirements, locating files, modifying code, and testing validation. For Agents, the final result is important, but the judgments, operations, and feedback left throughout the task completion process are more valuable for training.
This shift explains why internal communications, which capture the "how" of business operations, are becoming premium assets.
Pricing benchmarks for this emerging market are still forming, but comparisons can be drawn from existing data trades. Transactions handled by SimpleClosure and Protege for data from closed startups typically range from $10,000 to $100,000 per deal. Mercor has offered bids as high as $300,000 for employee chat records and emails from acquired startups. The Spirit Airlines deal pushed the price to tens of millions, with Mercor bidding $7.5 million and Google offering $10 million for around 600 million internal communications.
Calculated solely on the 600 million messages, the price was about 1.67 cents per message. While this seems low, Reuters reported in 2024 that Photobucket negotiated licensing rights with AI companies for around 13 billion photos and videos, with prices ranging from about 5 cents to $1 per photo and over $1 per video. In the B2B data market, a piece of contact information can sell for a few cents to several dollars depending on completeness and accuracy.
However, these prices do not form a unified standard. The value of internal data from closed companies depends largely on individual negotiations, influenced by factors such as data volume, industry, time span, completeness, uniqueness, and buyer plans. Google's $10 million bid is a rare example of a large-scale deal that has been made public, highlighting the potential scale of this market.
The standardization of this market remains a significant challenge. Compared to mature data licensing markets, such as Reddit's agreement to license user posts and comments to Google for around $60 million per year, the bankruptcy data market is in its infancy. The gap in pricing transparency and regulatory clarity is substantial. As more companies fail and AI demand for proprietary workflow data grows, the tension between creditor recovery and data privacy will likely intensify. The Spirit Airlines case serves as a bellwether, demonstrating that digital legacies can be monetized at scale, but it also underscores the lack of established norms for valuing and transferring such sensitive assets. This marks the beginning of a new era where the internal life of a company, once private and ephemeral, becomes a tradable commodity in the global AI economy.