About twenty years ago, I was taking a flight back from Rio de Janeiro, Brazil to the US. In the middle of the night the pilot got on the loudspeaker and said "hi! Having some engine trouble, so we are landing in Manaus."
Manaus is in the middle of the Amazon.
Needless to say, a bit scary to hear that, but we landed without issue.
They told us we had two choices: the nice hotel with a shared room, or the lesser nice hotel with no roommate. I chose the latter. When we go there, they said, "oops, sorry, short on rooms!" So I had a roommate.
Wandered around Manaus, took a skiff out on the Rio Negro. Saw pink river dolphins. A little boat approached us and a kid handed me a sloth, and then demanded I return it with a twenty dollar bill.
The airline got us another plane 24 hours later. Made it back to the US safely.
A few weeks later, the airline reached out and said "Here is $100 for your trouble."
I declined to take that offer. I had missed several business meetings that cost me actual money. I couldn't donate blood for years because I had been to the Amazon and was tagged a malaria risk.
During the many arguments with the airline I threatened to take them to small claims court.
I got a really strange response over email which I clearly wasn't supposed to see. A representative from that airline was asking internally if they could put me on the no-fly list. That was really chilling.
But, this is the kind of information I'm worried about when a vendor sells my data. If Google wanted to sell a product to the airlines that offered to keep annoying people like me from purchasing flights, they could do that with that email chain. I'm skeptical it'll be wiped correctly. Isn't my poor writing style basically my signature? How do you wipe that?
I'll never understand this attitude toward airlines. What did you want them to do in this situation? Keep flying the plane with engine issues so you could make your important meetings? It sounds like the airline did the right thing here but you were still mad?
I suppose consider yourself lucky that you haven't yet been inconvenienced by an issue completely within the airline's control and been offered $50 in funny money even though you had to pay out of pocket for your dingy hotel and transport in/out of the airport at 11pm/6am the next morning. And no, they won't reimburse that, because you could have slept on the terminal floor for free
I on the other hand will never understand anyone extending any grace to an airline. They are constantly testing just how poor of an experience they can deliver to their customers and stay in business.
For starters, they could have offered more than $100 to OP here
Perhaps they want the airlines to do a better job maintaining their fleet such that the airline is capable of meeting its obligations with respect to arriving on time without crashing.
Uhm, properly compensate passengers that have concrete harm to point to instead of offering a laughable alibi payment? What I'll never understand is why some neoliberal schools of thought defend conglomerates like they're their children... This company has enough money to compensate for fuck-ups, at least it should have. Every insurance on the planet can do the simple math involved for this risk assesment.
> I got a really strange response over email which I clearly wasn't supposed to see. A representative from that airline was asking internally if they could put me on the no-fly list. That was really chilling.
What if they only pretended to forward it to you by mistake and you /were/ supposed to see it?
As in "it would be a shame if you could never fly again".
In the United States legal system, when you take a plea because the Government was threaning you with a 20 year sentence the judge requires in court and under oath for you to state that you were not coerced/pressured in any way to take the plea (because pleas wouldn't be valid if the Government pressured you into them and into signing away your rights).
Before you push back at your work you have to consider possible loss of work and therefore housing/medical insurance/food.
The entire US system is designed around baked in blackmail of the individual into conforming. Not surprising at all to see the airline turn the no fly list into that pressure, it's the main technique we know to motivate people in the USA.
Re: plea deals. Most people don't know that plea deals are very common. Over 90% of convictions come from plea deals.
It's not just that lawyers are expensive and court is risky. Court also takes fucking forever. Somewhere along the way, a right to a speedy trial has been lost. People have jobs they have to get back to, families they have to provide for. The plea deal is often a way out right now, with the tradeoff of having a conviction.
Even if you're 100% innocent, it's often the better financial choice to just... take a guilty plea. The end result is we can't really say, for sure, how many convictions are actually guilty people.
The calculus gets even more complicated when we consider race. If you're a large black man, are you really going to want to take the chance in trial? If you're a meek white woman, you might be more inclined.
This is one of those places where our failure to teach history and civics is failing the American people.
It's Adams defending British solders, and later being very much a supporter of the concept of jury nullification. The fact that the modern court has been allowed to not only block the distribution of this information, but actively seeks to tell jurors that they HAVE to follow the law is, to be blunt, a disgrace.
> The entire US system is designed around baked in blackmail of the individual into conforming.
Back when I still argued with people on reddit I was shocked to learn that normies actually believe this is a good thing. Particular example was that the fact that so much of your life is tied to your workplace and you can get fired for political activity outside of your work. This shuts down any non-mainstream discussions. They were like "yes this is exactly what we want".
When I tried to explain "ok but there were times when racism and homophobia were the mainstream, do you think people advocating for equality should've lost their jobs for their political activity" I always got in response 502 Internal Server Error.
I fully share your concerns. And I don't understand how apparently tons of Teams and email conversations can be archived and sold without any kind of scrutiny. How can such data be sold without the consent of all involved parties? What gives Google the right to use it to train LLMs? Is that just a way of washing away the legal protections?
It is being scrutinized. The sale is overseen by the courts. Also, the media is scrutinizing. Also, PII has already been addressed by the court, from the article: "If you’ve flown Spirit and worry that Google will soon know about a testy conversation you had with the airline’s call center, you’re being told not to worry. The court filing says the data was deidentified before being put on sale and Google has promised to scrub any PII it finds in the trove."
Specially when you have a list of flights, people on board, issues, etc... The fact that they are specifically tellling people to not worry about that is also a red flag for me. They know what they'd worry about if they were the affected ones.
Yes, the same google that has repeatedly, entirely "by accident", captured boatloads of wifi data with their wardriving vehicles (or google streetview or whatever it's called). They sure seem like a trustworthy bunch.
Not only those wardriving vehicles. They use everybody to scan the world's wifi networks. Well, except people like me who stubbornly turn off 'location accuracy' every time some app demands you turn it on.
I don't know exactly what gets sent to google, but it's certainly enough to identify and track (retrospectively) a huge part of the world's population.
Now I know you get tracked by the celltowers anyway, but still. Navigation works fine with the accuracy offered by just using GPS and it doesn't need all the wifi scanning, it's pure data harvesting.
> Now I know you get tracked by the celltowers anyway
Yes, but theGoog doesn't own that data. By having all of theGoogOS devices scanning and reporting back directly to theGoog, theGoog gets that data for free. Plus, all of the other info it can hoover up that the cell towers would not have access.
Google 100% was wardriving and all of your other items.
However, you can increase gps accuracy using wifi. GPS is not that precise (part of which is government regulations) and WiFi does help immensely with accuracy.
I'm sorry but such assurances are worthless without being explicit what was scrubbed and what is retained. An "anonymous" customer ID with a list of flights is very identifiable when you have other information about the trips someone has taken. PII is not a binary yes or no and even benign data can become a problem in aggregate.
This is when a track record of Google's "Don't be evil" motto, culture and corporate habits being literally front and center on its official code of conduct, instead of moved to the very last line in 2018, would have persuaded people to give it the benefit of the doubt. Functionally effective privacy seems to be retreating ever more exclusively into the domain of very wealthy families and behind the corporate veil (by purchasing information scrubbing services on a regular basis), and the loss by normal people of the commons of mass privacy has unfortunately not been appreciated by the common citizen. It is a more valuable commons than recognized by most citizens, and is being rapidly co-opted and monetized by commercial entities that are not aligned with individual interests.
Even as an investor who stands to benefit from that monetization in the short term, I stand against this trend because like any Tragedy of The Commons economic scenario, in the long term (which isn't that long due to the automation that harvests this resource) it sows the seeds of its own dilution into functionally near non-commercial value.
Usually when you work a job you sign a little thing that says "yeah you own everything I produce for you, no matter how small".
Which is, of course, ridiculous, and follows the trend of absurdist contract law wrangling in corporations. Similar to non-competes and NDAs.
It makes sense to some degree, but the fact that semi-private conversations are included in that makes no sense. These have little to no business purpose.
Makes one appreciate living in place with sufficient constitutional protections against this sort of stuff. Even for work stuff selling this info wouldn't fly in some parts of the world.
If you're referring to GDPR, companies routinely evade such protections using "informed consent" / "legitimate interests" loopholes. The big ones get caught once in a while, get a slap on the wrist and continue to do whatever they were doing before, albeit with more safeguards.
Not for nothing, but you have probably already consented. Typically user agreements allow for this kind of sale if you’ve authorized use and processing but YMMV.
Since it’s work communications, consent was already given.
When you join a company, you typically sign an agreement that talks about how the company owns all your output. Thumbs upping a Teams message is work output and they own it.
Every email sent and received. Every keystroke. Etc etc etc.
If you don’t want your employer to log and sell it, start your own company. Or use a personal device. I do the latter.
I don't know the US law, but surely in Europe specifically every private conversation is private, period. No matter if it's work email, your company cannot read the emails directed at your company mailbox by its initiative (of course in case it's needed a judge can ask it to be taken as evidence), nor it can read the files on your computer, or anything similar, no matter if the device it's company provided, because it would be considered the same as using a camera to spy on the employee, that is of course illegal.
Of course if it's shared communication media (e.g. a mailing list) it can, but not at your private address, no matte if it's @company.com, it's considered the same as your private email.
That's not consent. It's a one-sided condition of employment. Consent would imply a meeting of the minds and a way for each employee to negotiate, or opt-out without losing employment.
It's like saying I consent to my phone company's 200 page long terms and conditions.
Corporate America has a very fucked up definition of consent, and they seem to have spread that definition broadly.
The party owning this data (Spirit Airlines) is consenting to the sale. Employees and customers of Spirit consented when they started employment and did business with Spirit, respectively.
Did they consent? Just because one receives a letter it doesn't mean they “own” it, much less that they are entitled to publish it at their leisure. If Spirit were active in any country with GDPR-style laws, the seller of these data would be most likely investigated.
This is why the GDPR (and to a lesser extent the CCPA) is a good thing. The data was supplied for a specific purpose. The handler of the data should have to obtain further consent if they wish to use it for another purpose.
> How can such data be sold without the consent of all involved parties?
In the US, whoever owns the computer owns the data on it. Courts have routinely ruled that you have no say in what other people collect about you. The goal of bankruptcy courts is to minimize the losses of the creditors. And bankruptcy courts routinely rewrite contracts except where statute prevents it (like mortgages).
In the EU, you own the data about yourself. A lot of people utterly hate GDPR, but that's reason that you own the data about yourself.
I got to the part where you declined the $100 and thought you declined it because you had gotten some memorable experiences and that was payment enough.
> If Google wanted to sell a product to the airlines that offered to keep annoying people like me from purchasing flights, they could do that with that email chain. I'm skeptical it'll be wiped correctly. Isn't my poor writing style basically my signature?
The OP article is also wrong on multiple counts. "Customer behavior" data like call recordings and email addresses/activity is specifically not included in Google's purchase. See page 18 of the court document they link, the "Google's Data Purchase Request" column on the right lists what is and isn't included.
I'd be amused if a sub-sub-agent organically decided to do it anyway - even if just for a notable figure that an LLM can identify with its weights alone. What are the controls? Who's going to keep Google accountable? Hah.
I know Meta it's not Google, but it's worth remembering that these promises haven't had a great measure of success in the past:
> Facebook has been fined €110m (£94m) by the EU for providing misleading information about its 2014 takeover of WhatsApp. (...) When Facebook took over the WhatsApp messaging service in 2014, it told the commission it would not be able to match user accounts on both platforms, but went on to do exactly that.
That did not go where I thought it was going! I thought you were going to say that the pink dolphins, the sloth, etc, were all more valuable than $100 ever… nevermind you missed some meetings, time to sue!
1) +95% of the population live on the coast very far away from the Amazon. Most of the population has not been there. Most of the coast has a very different jungle biome called Mata Atlantica and the countryside close to the coast is not that different from temperate forest of Europe. That is what most all Brazilians are used to. There is a significant population in the arid northeast though and the cold south as well (which is even more similar to europe).
2) Manaus is the biggest city in the Amazon and it is huge developed place (and has been for decades). You are not in the middle of the jungle if you land in the airport. The countryside around the city is jungle though.
3) Brazilian people do not necessarily like or are used to tacos and spicy food. Mexico is _really_ far away from Brazil.
I would not be offended by the blood donation thing. They generalize based on administrative regions (Amazonas in this case) and not whether you visited a big developed city or not.
I had a similar blood donation issue for visiting a particular island in the Philippines, and could not donate for 4 months.
It is funny there is no similar treatment about TBE (tick disease) which is predominantly an European disease and very dangerous and much more common than Malaria is in most (all?) South American countries...
I think 2) in particular makes some bits of the story evident that, at least, weren’t so obvious to me. I’m not well informed about the geography of the Amazon. For example, if Manaus is reasonably built up, I guess the hotel situation was mostly Spirit cheaping out rather than a limit on the actual hotel capacity of the town.
Honestly at this point it is usually the Europeans that need this (I have been asked if Chile has good tacos from a European). At least in my experience, most USians (in spanish Americans means everyone in the hemisphere) now know more about South America than the average European.
I have a german last name, do you know how often people outside Brazil act weird when they learn I am Brazilian?
My grandparents came to Brazil right after WW1 way before the Nazis came to power. High ranking Nazis fled to south america because there were a lot of germans living there already. Nearly all german people who moved to south america did it way before WW2.
I just run into this stuff a lot living in Europe.
Why are Americans in particular expected to know details about every country on Earth?
Do you think Canadians know all these facts about Brazil? Do Indians? Or Swedes? It's only Americans that are smugly called ignorant for not knowing about the entire world.
Not details, but at least having a basic notion of where countries are located. Everyone should be expected to know, for example, that Mexico and Spain are two different countries that are located in different continents. Yet it’s always people from the same country referring to one by the name of the other.
The sloth was returned to his owner and I did tip him. That kid is probably still prowling the Amazon (as an adult now), looking for sucker tourists like me.
Shouldn’t have accepted the sloth in the first place.
They may look cute and docile, but also have a panic reflex to grab onto the nearest tree-like thing they can feel with their claws, including their would-be attacker. Or if they just feel like they’re about to fall off their “tree” since they have horrible eyesight and get confused easily.
Which then leads to a cycle of pain and violence as they just dig-in harder and harder while trying to free them and/or fling them around wildly due to the human's “get this thing off me” response to sharp claws digging into their flesh.
There are some painful-to-watch videos on YouTube of this phenomenon from unsuspecting passersby tying to “help” them off the road/beach/etc and ultimately making things worse for everyone involved.
I love the irony of you checking them out for more information in response to a comment of them being worried about who reads their data. Nothing wrong with it, just make me chuckle
There's something about circles of control in this, that makes the difference. If I publish information about myself, that's about me, and it's in my control.
If someone else shares information about me, without my consent, and someone uses that to nose in on me, that feels creepy and problematic.
If you find that concerning, I recommend asking Claude or Codex to analyze all your HN comments and build a profile of you (I recommend that to everyone, not trying to single you out btw.) It takes about 20 minutes. It was eye opening and somewhat unsettling how accurate it was when I ran it on my own data. Even worse, there’s NO way to delete your old HN comments.
Well at the risk of flouting HN guidelines, I went to Grok and did exactly that. Grok's profiling seems thoroughly fact-based, and doesn't pull any punches.
Just a reminder that about 40% of the email conversations in the last 20 years are already in Google’s possession with the identifying data. (About another 40% are in Microsoft’s.)
If they want to do that they already can, thanks to the public’s overwhelming appetite for “FREE” overriding every single other possible concern.
Which is funny to point out on a post about Spirit, since that was an airline built to serve the customers for whom cheapness was the overwhelming single concern.
In the early pioneering days of the commercialized Internet, email addresses were inextricably linked to your ISP. You paid for an ISP connection and you got an email box, with MTA and MUA service to match. You were reluctant to switch or leave your ISP, because that also meant leaving behind your email address. Of course, a minority of nerds got around this with their own domains, etc.
However, it seems that Google, AOL, Yahoo!, Hotmail, and other players got into providing free email services and eventually grew into giants that supplanted every other MTA service. This was not an accident and it was not merely our appetite for “FREE” but it was a very calculated plan by the industry. Those early ISPs did not have a business model that admitted monetizing our private data; they seemed to have a more respectful attitude for keeping it private. Perhaps that was a result of being telecommunications-based companies, rather than advertising or entertainment.
If a service like email provides such endless treasure troves of personal data, including a social graph and glimpses into our private daily lives, why not provide it for free and monetize opportunistically on the data itself? The free email services killed the paid services, not by being better or cheaper, but by being bigger, centralized, and more persistent. The main reason I signed up for Yahoo! was because it would be an utterly stable presence. I saw my parents and others so hopelessly attached to an ISP-based email, but I couldn't end up like that.
After streaks of losing my home and non-payment of bills and moving around over decades, the most stable point of contact for me has been a "free" email address.
ahhh, there seems to be different no-fly lists? The one Im aware of is the one for terrorists and moneylaunderers, and usually they will not tell you who put you on that list :-D
There is "the" no-fly list maintained by the government, which is nominally for people who are too dangerous to be allowed on an airplane, yet not dangerous enough to charge criminally.
But each airline also maintains their own internal no-fly list for people they prefer to no longer have as customers, for whatever reason. You might end up on this for some abuse of the system that doesn't pose any sort of safety risk, so the government doesn't care, but the airline doesn't like. For example, excessively doing hidden-city ticketing (where you book a flight with a connection, then skip the second leg of the trip, because weird pricing rules make it cheaper than just booking a ticket to the connecting city) can get you banned from the airline, but the government would be completely uninterested.
"If you’ve flown Spirit and worry that Google will soon know about a testy conversation you had with the airline’s call center, you’re being told not to worry. The court filing says the data was deidentified before being put on sale and Google has promised to scrub any PII it finds in the trove."
unfortunately "de-identified" data is typically re-identified quite trivially. so i guess we just hope google keeps its promise, and is competent in its scrubbing.
Modern corporations have learned that the best way to do shady things is to make it very clear that such things are against policy and absolutely forbidden, and then to also make it very clear that breaking policy is the only way to actually get your job done. That way they still get all the benefits of doing shady stuff, and if it ever comes to light then they can fire some scapegoats and explain that this was against policy and they'd never ever condone it.
This is a fascinating story. Thanks for posting it.
That email you accidentally received really bothers me. I don't understand why a CS rep would get this invested to the point of wanting to cause you real harm. They're not the airline. The psychology is fascinating. There are people out there who feel like a mild short-term inconvenience to them where they have no stakes somehow justifies life-changing harm is kinda frightening, honestly.
I'm reminded of the Yahoo search data fiasco that was allegedly anonymized. Turns out, it wasn't so anonymous [1]. For one thing, people tend ed to search their home address. Whoops.
You mention writing style. We already have LLMs quite capable of copying a writing style. It's a natural extension to say we can fingerprint writing style too.
But here's another aspect. Imagine you're in a relationship with someone and you somehow fingerprint their personal data with a company. For example, you use their Netflix to like 5 very obscure movies, to the point where it's likely unique. Now imagine that Netflix's data gets released in an "anonymized" form and you can now find it based on those obscure likes. I can imagine many scenarios like this. And there's no text involved here at all.
> I don't understand why a CS rep would get this invested to the point of wanting to cause you real harm
This is the kind of power tripping that easily corrupts people, especially those who don't have much power outside of work. The US national security apparatus is vast and powerful. Many people get giddy at the thought of inflicting punishment on those who "deserve" it. Act rude to a fast food worker, and you get spit in your food. Lots of people cheer the worker who spit in the food of a customer who is merely rude, impatient, or demanding. Or is guilty of being a cop, politician, rich, etc.
I can easily imagine the kind of CS rep who would delight in putting a customer who didn't just go along with things and spoke up for themselves... thats 'rude' and 'disrespectful' to some. Police are the most notorious for this kind of petty "you will respect my authority" but it is in every industry.
HN / hacker culture celebrates this in the 'Bastard Operator from Hell' [0], the sysop who will ruin your work and life with their IT wizardry, no matter if you're an intern or CEO, if they do not feel respected. Or if you interrupt their gaming with your support call.
I don't understand why a CS rep would get this invested to the point of wanting to cause you real harm.
Two possibilities come to mind...
1 - The CS rep has been instructed to do this. Scary, but corporate leaders can be assholes and wield lots of power within their orgs, so doesn't seem completely unlikely to me.
2 - The CS was just a dick.
Frankly, given the behavior of various SuperMegaCorps over the past few decades, I'm going with #1.
... or frustrated. "I have my quota of jiras to meet, and this person is taking up too much of my time when it was such a simple issue which no one complains about". Not a good customer service mindset, but frustration happens.
I'm not sure I follow your argument. The privacy abuse already happened. The data is already there. And it was the airline that did it, not a tech giant who just wants to train a bunch of MLs.
Surely if this is the scenario you're worrying about, and you accept the lack of regulatory protections, Google buying Spirit's data is a good thing, right? Much better them than the airlines who you already know to be corrupt?
That's because they're seemingly perfunctory. What's worse than no law is a bad one that doesn't do anything but make you feel like something is actually being done.
This is why I get bad vibes any time I hit a cloudflare interstitial page. If you ever piss them off it would be trivial to cut you off from most of the internet.
I never understood this attitude, like I feel like being able to fly in airplane is one of the most amazing human achievements, yet people will try and save down to the dollar booking a flight like it was breakfast at Dennys, and come out huffing and puffing as soon as anything goes wrong demanding their money back.
Airlines like Spirit catered to the worse of these type of customers.
That’s a bit unfair. They also catered to broke people who weren’t crazy or argumentative. On a planeload of 100 people, there must have been 60 of them at least who were just broke. Or on some routes, Spirit was the only carrier with a direct flight so they were just normal people who wanted to get to a certain place efficiently (Hi, that’s me, I flew them for this reason). Keep in mind Spirit had an excellent safety record, too, so airline choice in this case was primarily a question of having luxuries or not.
Manny retail industries already share lists of "troublesome" customers (trouble = anything from too many returns to lawsuit-happy to friendly fraud). Not sure this is a new concern..
Not trying to be snarky, and perhaps it wasn't well stated, but the last paragraph I said I'm concerned about identification via my writing style. If they have my emails, they would have my writing style. It doesn't have to be tied to PII there, they can cross reference it with my blog. I'm speculating because I read that you can identify people by a few sentences of their writing.
"Deidentification" seems really murky and imprecise at best.
Reidentification via writing style is definitely possible, and I doubt the vendor will modify things in a way sufficient to handle that.
But I think this is a place where we should apply bounded distrust: there are lots of places where we should distrust Google, but reidentifying people in an explicitly deidentified dataset isn't one of them.
The re-identification is done by ML, at which point it's basically undetectable. If you give all this day to, say, and LLM, and the LLM also has your blog, the identification is embedded in the weights.
I can ask "Tell me about Person A's experience with airlines" and it will tell me. Or I can ask more generally, "knowing your knowledge, derive a no-fly list", and then it's likely Person A will be on it.
Based on their trackrecord, That's definitely a concern. I don't really understand on which basis you conclude 'isn't one of them' . 'Don't be evil' ? :-P
I think this is the kind of place where applying bounded distrust is critical: it's not whether we trust Google overall, it's about figuring out what sorts of statements we should expect to effectively bind companies and in what ways.
For example, I think a pretty worrying outcome here is that deidentification is imperfect (not surprising), the data is fed into model training, and then the model makes identity-dependent inferences. Since no one tried to break deidentification, it's within what I'd expect from the company. (And, to be clear, is bad.)
On the other hand, intentional reidentification to work around contractual deidentification to "a sell a product to the airlines that offered to keep annoying people like me from purchasing flights" is the kind of thing that would make Google's lawyers terrified, so we should not expect it.
If you look at how this worked with DoubleClick, Fitbit, etc there were initially barriers to linking data but the mechanism for unlinking was updated agreements with people who had ongoing interaction with the continuing entity. That's not the situation with the Spirit data.
The closest I'd expect to see for a "keep annoying people from purchasing flights" situation is not a list of troublemakers but a model that's very good at scoring future communications from customers, and has learned to distinguish profitable vs unprofitable customers. This is well within what I'd expect from companies, and doesn't require any reidentification.
Even if they follow to the letter a deidentification process, Google and Meta have so much data about individuals that re-identification shouldn't be very hard for the majority of airline passengers' data they put their hands on.
Of course, takes a lot more effort than not doing proper deindetification in the first place but if they wanted to appear like caring about data privacy they still have enough data points to correlate the sets later on (and/or over time).
The idea that there’s a nefarious plot to do something super evil with this data is a bit crackpot though based on their incentives.
Remember, Google = Ads. Their only focus and only care. Their mission statement, rendered accurately, is “Ads ads ads ads. Effective ads. Ads worth paying a lot for. Ads ads ads. Advertising and ads.”
If they choose to be evil in some additional way, (1) remember, they would only do that if in some way it serves their advertising needs — not to offer innovative new black-hat databroker services to airlines, and (2) this little dataset will not need to be re-identified. They’ll just use the 20 years of email and search data they already have on like half the world’s population.
Even before LLMs there were multiple papers written about ways to to reidentify people with ML and other statistical analysis. It is probably now even more trivial especially if you are Google.
> But, this is the kind of information I'm worried about when a vendor sells my data
Don't worry. Spirit probably lost all of the emails from the customers (or they were devnulled) and 90% of the data is probably autoresponder messages promising the company would respond.
The other 10% was probably the meme collection of the executive management team.
> Google bought itself 100 million emails and 500 million items from Microsoft Teams, 17 million OneDrive files and 20.5 million items from SharePoint. The search giant also now owns over 30 million recorded customer service calls, and more than 15 million customer service chat records. 600,000 ServiceNow tickets are another element of the collection, along with 13.7 million active emails addresses from Oracle’s Responsys marketing application, and details of 11 million sales of in-flight Wi-Fi services.
> There’s also operational data in the trove, describing over 763,000 flights, five million crew pairings, more than 1.2 million fuel slips, and records describing purchases of 787,452 parts.
> Google has reportedly said it bought the data to improve its AI services.
Gives "this call is being recorded for training purposes" new meaning.
None of that customer data is included in the purchase. Page 18 of the linked court document is the source of these record counts, and on the right is a "Google's Data Purchase Request" column that lists all of this "Customer Behavior" data as "not included".
The Register is not a serious publication and completely missed this. Other outlets reporting this story do not include the claim that customer data is included.
Is there anything that can legally be done against this? It feels like a breach of consent. Like, it cannot be that when one accept their voice to be recorded for _human_ training they also accept it to be recorded for LLM training
Your comment reminded me of a funny interaction I had a week or so ago.
I got a call that started with the usual automated message, "This call is being recorded." After the person joined, I pushed the record button on my iPhone, "This call is being recorded."
They were surprised and asked why I'm recording.
I said, "You're recording, so I'm recording too."
The rep insisted that the company doesn't like this but that they will continue with the call anyway.
Anyway, I wish people took this stuff much more seriously. It always seems to boil down to, "I don't have anything to hide" type of conversations and I've never managed to convince anyone that privacy as a concept isn't about having something to hide.
What is actually perfect, when the automated voice declares the call is being recorded, you don't have to do it again. You're legally allowed to record, because both parties were informed already and agree to it.
I don't think there's any legal weight to the purpose of a recording unless you entered into a contract that has a clause to that effect. They have to tell you that the call is being recorded because of wiretapping laws. They stick "for training purposes" on there just to soften the language and reassure you that they have a good, non-nefarious reason to record. It doesn't actually limit what they're allowed to do with the recording.
Axios claims the acquisition doesn’t contain passenger profiles or frequent flyer info but that data would be trivial to replicate given the Responsys data set which would include records of all transactional emails sent.
"This call is being recorded so that Gemini can decide which purge wave to assign you to. Obedient humans will be carried over for further cycles until no longer needed. If you are scheduled for termination this cycle a disposal representative will be with you shortly."
I kid, but...
It's probably the precursor to insurance denials and job screening.
I got banned from r/technology a few weeks back for decrying tracking in AI content. The community was piling on saying it was okay because it removed AI content or made it easy to spot. I made the counter argument that watermarks would find their ways into everything and eventually be bound to attestation. The mods didn't like that. (Yet another structural problem with the lack of p2p self-service town squares.)
The socials are training the next generations for broad acceptance.
> I made the counter argument that watermarks would find their ways into everything and eventually be bound to attestation.
Yup.
Elsewhere in another front page thread today: "oh but apps blocking screenshots because of 'sensitive content' can be bypassed by taking a photo of your screen with another phone".
Any tech-savvy person with two brain cells reading this and that: "gee, I wonder if the same magic imperceptible watermark that survives multiple rounds of cropping and printing and scanning, that's used to tag AI-generated content, could also be used to tag sensitive data, or ads, or which app is rendering it on screen, and then the camera app could refuse photographing it...".
I don't know why people don't see that AI watermarks are DRM, and DRM is universal, and there are many clients...
> 600,000 ServiceNow tickets are another element of the collection, along with 13.7 million active emails addresses from Oracle’s Responsys marketing application, and details of 11 million sales of in-flight Wi-Fi services.
I don't see how it's possible any more, when correlated against all the various other data sources. And a record that might be unidentifiable now might become unique with more correlated data sources.
as an example, they can remove the names off these sales data, so you can't identify who purchased what items. However, the purchaser would be identified by some sort of number, and you would be able to extract information about purchasing habits, and aggregate these habits into usable information for advertising purposes (like targeting and profiling).
Anyone else somewhat weirded by current state of affairs that this sort of information is valuable enough to even bother selling... And that it actually happens... It feels like some societies are in really weird place.
How does this have value? Is any and every sentence in an e-mail considered 'fact' and thus to be fed into the AI?
90% of e-mails and Teams communications are inane. Polite banter, "thanks for taking care of that, I appreciate it" "please route the forms to Janet this week because Bill is on vacation" "unit will be un available until the parts come in" . I can't see the intrinsic fact value of this kind of communication without screening it. And after screening, the gold nuggets would be minimal.
LLMs aren't a database. They're an attempt at brute-forcing an artificial mind. The who and what aren't really interesting there, it'll forget most of such details anyway. What matters is the patterns visible in the text at various scales. How people write. Why they write. To whom they write, in response to what. How does e-mails about mistakes correlate with PDFs they're referring to. How people work with ticketing systems - like how, actually, a ticket plays out. The jargon, the acronyms, the vibes, the causal links. It's all in there, and it's another slice through the set of things humans do, to be combined with other slices already in the training data, and enriching the whole.
(Something something we will add your distinctiveness to our own, you will be assimilated, ...)
(Hell, the fact that it's all from one org would make it a great dataset to have in the open for sociological studies. I bet that today, aided by LLMs to sift through it, you could use it to map how information flows through a large org - how incident on the floor travels through time and layers of management until it reaches the C-suite, what of it survives, how it gets reacted to, how the reactions flow down...)
> How does this have value? Is any and every sentence in an e-mail considered 'fact' and thus to be fed into the AI?
I have a hunch what this is for. AI companies want to make bigger inroads into nontechnical work settings. But LLM progress outside of fields where verifiable rewards for RL post-training can be synthetically generated (coding, math) has been pretty flat. Buying years of operational data from a company like an airline could be used to reconstruct long-horizon task trajectories in areas like customer service or marketing.
> I have a hunch what this is for. AI companies want to make bigger inroads into nontechnical work settings. But LLM progress outside of fields where verifiable rewards for RL post-training can be synthetically generated (coding, math) has been pretty flat. Buying years of operational data from a company like an airline could be used to reconstruct long-horizon task trajectories in areas like customer service or marketing.
This. Training data companies like Mercor are even creating simulated companies to create similar data, so they can better automate even more white collar work:
> The data-training start-ups see a lucrative opportunity in recreating workplaces in miniature: controlled environments in which their gig workers can evaluate and reproduce emails, memos and slide presentations in context. The information emerging from such a setup, the companies boast, will help shrink the gap between what A.I. models can accomplish and what office workers actually do from one minute to the next, as ideas and instructions flow between meetings, documents and applications.
> Scale, for example, has said that “our environments replicate real-world workflows,” and that its contractors “curate artifacts that capture the complexity, ambiguity and edge cases of real professional work.”
> Executives see the models’ shortcomings as a sign there’s more for them to do.
> “I often use Claude Cowork, right?” said Edwin Chen, the founder of Surge. “And even though Claude Cowork is incredibly smart, oftentimes it doesn’t quite understand the nuance of Slack. It doesn’t quite understand this ambiguous question I have. It doesn’t quite understand where to go and find this Google document.”
> According to the Bloomberg Billionaires Index, Mr. Chen’s stake in Surge — and his vision for what it could become — makes him the 258th-richest person in the world.
> “I often think about us as essentially, like, the school for A.G.I.,” Mr. Chen said, referring to a prophesied level of A.I. that surpasses human intelligence. “A.I. comes to us, and A.I. learns to run the world.” He and his customers at the big A.I. labs, he said, are designing the curriculum.
A major selling point of AI chat bots is for customer service automation. A clean dataset like this is a huge find. I'm not sure what you mean by "facts" or "gold nuggets". Training data doesn't need to be factual.
> Is any and every sentence in an e-mail considered 'fact' and thus to be fed into the AI?
If it comes with the context, yes. More data the better. Someone considered it served some purpose at some point. Thus it contains, no matter how tiny, a sliver of information.
It will have a different writing style from the average blog post. Maybe they just want to train an AI that sounds less like an AI. Or one that speaks in vapid management style.
This data shows exactly how a huge company of thousands of employees works and coordinates.
The perfect data to train an agent swarm on how to run a company.
Maybe it's innefficient and inane, but it's how you start.
The first LLMs, GPT-1, 2, were trained on complete garbage, the average document from the common crawl is random non-sense, yet they worked, and now we can use LLMs to filter the data for the next training run.
I am very much weirder out by it, yeah. Seems some societies are just excessively desperate for some kind, any kind, of fuel for economic growth, to the point this is where attention is now. The term "post capitalism" being thrown around feels less ridiculous than it did in years gone past.
Truth is even weirder. plenty of growth is possible but modern liberal democracy requires that all progress must be contingent on the production of enormous amounts of text that nobody would ever read.
3 years ago, the Wall Street journal covered a company trying to use AI to generate documents required for the approval of new nuclear reactor reactor designs, which sounds dangerous, until you get to the point where they'd cite 2 million pages as necessary for a typical application. [1] I don't need to explain why no individual or institution could read that, much less examine it in detail. I remember a rather funny question I found in a comment to that story - "How many pages of those 2 million could contain pornographic images before anyone notices?"
It's obvious why companies are so desperate for training data - a text generator of sufficient quality is more conductive to the growth of the nuclear industry than any scientific breakthrough in nuclear physics. (And of course, if you want to prevent the development of a nuclear reactor by your competitors, being able to produce millions of pages of high-quality objections will do the trick.)
And it's not just nuclear power. When it comes to stuff like building rail lines, apartments or power plants (both conventional and renewable), you'd find that the main bottleneck is the necessity to produce documents. And of course, many documents can be subject to judicial review - a process that consumes even more text.
Well that's why a plan can take years to be approved.
There also is the liability angle, where no one necessarily reads the document until the document is relevant to some situation in the future. Then you better have that document on hand.
According to the article, the “application” (which I imagine was also broken down into smaller pieces as they went back and forth to regulators during the review process) was 12,000 pages in total.
The 2 million number is for “supporting” materials which isn’t really defined. It definitely makes for a better soundbite to claim millions.
> its 12,000-page application had around two million pages of support materials.
Any kind of fuel for giving active investors that FOMO tingle which then forces the steamroll of index funds to blindly follow.
I guess the appropriation "any sufficiently advanced stock market is indistinguishable from entertainment" doesn't quite stop at equating the trade floor with a casino. At some point, entertainment also becomes the modus operandi of corporations.
I see from the court PDF that the process here involves Spirit giving the data to a "Deidentification Agent" (a third party firm that Google selects and pays for) who is responsible for stripping out things that would link data to any particular person before passing the data on to Google. Is that a standard thing, such that everybody in this transaction would have said "yes, put in the usual clauses about deidentifying the data" and multiple firms offer this service, or is it something that they custom-specified for this "we want the data for AI" transaction?
(The PDF mentions "the standard for deidentification set forth under the California Consumer Privacy Act", which suggests this is all pretty well legislatively understood.)
You wouldn't want to hard-wire the deidentification company's name into the contract between Google and Spirit. Otherwise, if the deident-company happens to go bankrupt or otherwise be unable to do the work then you'd need to re-do the Google-Spirit contract, which would be a massive pain. And you don't want to make "we can sign this with Spirit now" be dependent on "we have first signed the deal with the deident-company". So I think it's reasonable that the contract says "one or more third parties acceptable to or designated by Buyer" rather than being specific here.
Seems the answer is “no” to the first part of your question. From the filing:
> For example, one initial bid requested certain customer list information; however, by the first round of the Auction, the most competitive bidders had agreed to bid on an asset schedule that expressly excluded PII.
Is this the first case of a company's data being sold at bankruptcy for a significant sum? I'm genuinely unsure. Where there such value in this type of data before? Is every bankruptcy manager looking at this and seeing how every bankruptcy can now raise a few million more dollars?
FYI: If you have a company that you are shutting down, you too can sell your data to the labs. Companies will help you do this. If you have a company with a handful of people and you wrote code, collaborated in Slack, and used task management tools for a few years, you can probably sell this data for $50k or so.
You can also do this if you're not shutting down, but it's probably not worth it.
I wonder how they will use the data. If it was me I’d try to build a simulation of an airline, and then use it as an agent training environment. It really depends on the exact nature of the data what kinds of agents you could train, but maybe customer support (imo the worst AI use case) that are more empowered to make changes, or something for making more autonomous calls when recovering from irrops? Could be some cool’s stuff if a little niche, I hope they share / publish something and it doesn’t just disappear into a void.
I must be naive. I was under the impression Google wants this data to learn from a universally hated company's worst processes and practices, i.e. to teach their AI what not to do. It seems people are worried about their pii or that Google is curating a blacklist of customers?
You know when we (they) tell you not to do any personal computing on work devices/systems and to keep your devices completely separate from work ones.
Yeah this (and lawsuits/investigations) are why, the employer owns the data, in some contexts (like this one) it can become an asset (or a liability) but in either case it's not yours.
Of course that only gets you part of the way there anyway see Twitch recently opting in all users by default to mined for AI and only adding an opt out after backlash with a quote that was so on the nose it made me stop "If we'd have asked them to opt in, they wouldn't have opted in" (paraphrasing but it was that blunt).
Ah, but Google promised to remove PII they found in this deidentified dataset, so worry not.
> If you’ve flown Spirit and worry that Google will soon know about a testy conversation you had with the airline’s call center, you’re being told not to worry. The court filing says the data was deidentified before being put on sale and Google has promised to scrub any PII it finds in the trove.
Sure, a Spirit customer with no Gmail account can absolutely complain and/or be wary of this. I didn't mean what I wrote as a complete counterargument -- just trying to indicate the "scale" of the privacy issue here, which is that it's something that many people already tacitly accept.
I would love it if there were services where I could let them see everything I do, including when I poo and wank, if they just directly paid me for it.
No I don't want to just use your enshittified service for free. Fucking pay me and watch me all you want :)
Not sure if you're serious but a) all HN data is publicly downloadable and used by frontier labs and b) the origin of OpenAI tracks down to former YC CEO sama who was thereafter fired from YC
I don’t know why any company would pay. It can’t be that difficult to scrape this site and they already did it indiscriminately for years, violating laws and taking down public libraries and other public resources with no regard for their impact.
About twenty years ago, I was taking a flight back from Rio de Janeiro, Brazil to the US. In the middle of the night the pilot got on the loudspeaker and said "hi! Having some engine trouble, so we are landing in Manaus."
Manaus is in the middle of the Amazon.
Needless to say, a bit scary to hear that, but we landed without issue.
They told us we had two choices: the nice hotel with a shared room, or the lesser nice hotel with no roommate. I chose the latter. When we go there, they said, "oops, sorry, short on rooms!" So I had a roommate.
Wandered around Manaus, took a skiff out on the Rio Negro. Saw pink river dolphins. A little boat approached us and a kid handed me a sloth, and then demanded I return it with a twenty dollar bill.
The airline got us another plane 24 hours later. Made it back to the US safely.
A few weeks later, the airline reached out and said "Here is $100 for your trouble."
I declined to take that offer. I had missed several business meetings that cost me actual money. I couldn't donate blood for years because I had been to the Amazon and was tagged a malaria risk.
During the many arguments with the airline I threatened to take them to small claims court.
I got a really strange response over email which I clearly wasn't supposed to see. A representative from that airline was asking internally if they could put me on the no-fly list. That was really chilling.
But, this is the kind of information I'm worried about when a vendor sells my data. If Google wanted to sell a product to the airlines that offered to keep annoying people like me from purchasing flights, they could do that with that email chain. I'm skeptical it'll be wiped correctly. Isn't my poor writing style basically my signature? How do you wipe that?
I'll never understand this attitude toward airlines. What did you want them to do in this situation? Keep flying the plane with engine issues so you could make your important meetings? It sounds like the airline did the right thing here but you were still mad?
I suppose consider yourself lucky that you haven't yet been inconvenienced by an issue completely within the airline's control and been offered $50 in funny money even though you had to pay out of pocket for your dingy hotel and transport in/out of the airport at 11pm/6am the next morning. And no, they won't reimburse that, because you could have slept on the terminal floor for free
I on the other hand will never understand anyone extending any grace to an airline. They are constantly testing just how poor of an experience they can deliver to their customers and stay in business.
For starters, they could have offered more than $100 to OP here
Perhaps they want the airlines to do a better job maintaining their fleet such that the airline is capable of meeting its obligations with respect to arriving on time without crashing.
Do proper maintenance so they don’t have engine issues, and carry insurance to compensate passengers in case they do?
Yeah I don't get what the airline did wrong there.
A decent reimbursement.
Not parents fault the airline flies around with defective hardware.
Usually this is regulated by law.
Uhm, properly compensate passengers that have concrete harm to point to instead of offering a laughable alibi payment? What I'll never understand is why some neoliberal schools of thought defend conglomerates like they're their children... This company has enough money to compensate for fuck-ups, at least it should have. Every insurance on the planet can do the simple math involved for this risk assesment.
> I got a really strange response over email which I clearly wasn't supposed to see. A representative from that airline was asking internally if they could put me on the no-fly list. That was really chilling.
What if they only pretended to forward it to you by mistake and you /were/ supposed to see it?
As in "it would be a shame if you could never fly again".
In the United States legal system, when you take a plea because the Government was threaning you with a 20 year sentence the judge requires in court and under oath for you to state that you were not coerced/pressured in any way to take the plea (because pleas wouldn't be valid if the Government pressured you into them and into signing away your rights).
Before you push back at your work you have to consider possible loss of work and therefore housing/medical insurance/food.
The entire US system is designed around baked in blackmail of the individual into conforming. Not surprising at all to see the airline turn the no fly list into that pressure, it's the main technique we know to motivate people in the USA.
Re: plea deals. Most people don't know that plea deals are very common. Over 90% of convictions come from plea deals.
It's not just that lawyers are expensive and court is risky. Court also takes fucking forever. Somewhere along the way, a right to a speedy trial has been lost. People have jobs they have to get back to, families they have to provide for. The plea deal is often a way out right now, with the tradeoff of having a conviction.
Even if you're 100% innocent, it's often the better financial choice to just... take a guilty plea. The end result is we can't really say, for sure, how many convictions are actually guilty people.
The calculus gets even more complicated when we consider race. If you're a large black man, are you really going to want to take the chance in trial? If you're a meek white woman, you might be more inclined.
This is one of those places where our failure to teach history and civics is failing the American people.
It's Adams defending British solders, and later being very much a supporter of the concept of jury nullification. The fact that the modern court has been allowed to not only block the distribution of this information, but actively seeks to tell jurors that they HAVE to follow the law is, to be blunt, a disgrace.
https://en.wikipedia.org/wiki/Four_boxes_of_liberty
No more of this nonsense around "arbitration" being OK. Take it all to court, force it to be on the public record.
> The entire US system is designed around baked in blackmail of the individual into conforming.
Back when I still argued with people on reddit I was shocked to learn that normies actually believe this is a good thing. Particular example was that the fact that so much of your life is tied to your workplace and you can get fired for political activity outside of your work. This shuts down any non-mainstream discussions. They were like "yes this is exactly what we want".
When I tried to explain "ok but there were times when racism and homophobia were the mainstream, do you think people advocating for equality should've lost their jobs for their political activity" I always got in response 502 Internal Server Error.
I fully share your concerns. And I don't understand how apparently tons of Teams and email conversations can be archived and sold without any kind of scrutiny. How can such data be sold without the consent of all involved parties? What gives Google the right to use it to train LLMs? Is that just a way of washing away the legal protections?
It is being scrutinized. The sale is overseen by the courts. Also, the media is scrutinizing. Also, PII has already been addressed by the court, from the article: "If you’ve flown Spirit and worry that Google will soon know about a testy conversation you had with the airline’s call center, you’re being told not to worry. The court filing says the data was deidentified before being put on sale and Google has promised to scrub any PII it finds in the trove."
"De-identified" data is trivially easy to re-identify, especially by google.
https://www.nytimes.com/2006/08/09/technology/a-face-is-expo...
Specially when you have a list of flights, people on board, issues, etc... The fact that they are specifically tellling people to not worry about that is also a red flag for me. They know what they'd worry about if they were the affected ones.
Sure, you can argue that its a bad deal, shouldn't be allowed period, etc. But that is a different argument than saying that there is no scrutiny.
Sounds like a case of “there’s not enough scrutiny unless the decision ends up agreeing with my position”
The argument is that the scrutiny is in practice not sufficient, as usual in these cases.
The argument generally doesn't demonstrate that. The argument generally goes:
1. Google could do it.
2. <This space is intentionally left blank>
3. Therefore, Google is doing it!
(Step 2 needs to be filled in a bit for it to be a good argument. Generally, analogies don't quite make the cut.)
You are ignoring the value of previous experience:
Please replace 'Google' with 'Profit-driven legal entity in the USA' and reread your argument.
The parent is making an assumption, based on past experience, but is also Most Likely correct.
Insufficient scrutiny would be more apt.
I'm sure we can trust google to keep to their word and that we can trust a bankrupt airline to do their best at removing pii.
Yes, the same google that has repeatedly, entirely "by accident", captured boatloads of wifi data with their wardriving vehicles (or google streetview or whatever it's called). They sure seem like a trustworthy bunch.
Not only those wardriving vehicles. They use everybody to scan the world's wifi networks. Well, except people like me who stubbornly turn off 'location accuracy' every time some app demands you turn it on.
I don't know exactly what gets sent to google, but it's certainly enough to identify and track (retrospectively) a huge part of the world's population.
Now I know you get tracked by the celltowers anyway, but still. Navigation works fine with the accuracy offered by just using GPS and it doesn't need all the wifi scanning, it's pure data harvesting.
> Now I know you get tracked by the celltowers anyway
Yes, but theGoog doesn't own that data. By having all of theGoogOS devices scanning and reporting back directly to theGoog, theGoog gets that data for free. Plus, all of the other info it can hoover up that the cell towers would not have access.
Google 100% was wardriving and all of your other items.
However, you can increase gps accuracy using wifi. GPS is not that precise (part of which is government regulations) and WiFi does help immensely with accuracy.
The article claims it has already been removed, that Google is committing to remove anything leftover that they find.
You can argue that its a bad deal, shouldn't be allowed period, etc. But that is a different argument than saying that there is no scrutiny.
Surely all this will take is asking the trained LLM to deidentify it lol
Gemini, scrub this text for PII, make no mistake!
I prior worked at Google, I can say they DO de-anonymize data. You’d be foolish to think some PM within the company wouldn’t use this for malice. L
I'm sorry but such assurances are worthless without being explicit what was scrubbed and what is retained. An "anonymous" customer ID with a list of flights is very identifiable when you have other information about the trips someone has taken. PII is not a binary yes or no and even benign data can become a problem in aggregate.
This is when a track record of Google's "Don't be evil" motto, culture and corporate habits being literally front and center on its official code of conduct, instead of moved to the very last line in 2018, would have persuaded people to give it the benefit of the doubt. Functionally effective privacy seems to be retreating ever more exclusively into the domain of very wealthy families and behind the corporate veil (by purchasing information scrubbing services on a regular basis), and the loss by normal people of the commons of mass privacy has unfortunately not been appreciated by the common citizen. It is a more valuable commons than recognized by most citizens, and is being rapidly co-opted and monetized by commercial entities that are not aligned with individual interests.
Even as an investor who stands to benefit from that monetization in the short term, I stand against this trend because like any Tragedy of The Commons economic scenario, in the long term (which isn't that long due to the automation that harvests this resource) it sows the seeds of its own dilution into functionally near non-commercial value.
pinky promise?
Usually when you work a job you sign a little thing that says "yeah you own everything I produce for you, no matter how small".
Which is, of course, ridiculous, and follows the trend of absurdist contract law wrangling in corporations. Similar to non-competes and NDAs.
It makes sense to some degree, but the fact that semi-private conversations are included in that makes no sense. These have little to no business purpose.
Makes one appreciate living in place with sufficient constitutional protections against this sort of stuff. Even for work stuff selling this info wouldn't fly in some parts of the world.
If you're referring to GDPR, companies routinely evade such protections using "informed consent" / "legitimate interests" loopholes. The big ones get caught once in a while, get a slap on the wrist and continue to do whatever they were doing before, albeit with more safeguards.
Not really: https://noyb.eu/en/fines-resulting-noyb-litigation
Sure 50 M or even 1 B might be peanuts for faang but still there is real progress.
Support Noyb at all costs
Not for nothing, but you have probably already consented. Typically user agreements allow for this kind of sale if you’ve authorized use and processing but YMMV.
Since it’s work communications, consent was already given.
When you join a company, you typically sign an agreement that talks about how the company owns all your output. Thumbs upping a Teams message is work output and they own it.
Every email sent and received. Every keystroke. Etc etc etc.
If you don’t want your employer to log and sell it, start your own company. Or use a personal device. I do the latter.
I don't know the US law, but surely in Europe specifically every private conversation is private, period. No matter if it's work email, your company cannot read the emails directed at your company mailbox by its initiative (of course in case it's needed a judge can ask it to be taken as evidence), nor it can read the files on your computer, or anything similar, no matter if the device it's company provided, because it would be considered the same as using a camera to spy on the employee, that is of course illegal.
Of course if it's shared communication media (e.g. a mailing list) it can, but not at your private address, no matte if it's @company.com, it's considered the same as your private email.
That's not consent. It's a one-sided condition of employment. Consent would imply a meeting of the minds and a way for each employee to negotiate, or opt-out without losing employment.
It's like saying I consent to my phone company's 200 page long terms and conditions.
Corporate America has a very fucked up definition of consent, and they seem to have spread that definition broadly.
The party owning this data (Spirit Airlines) is consenting to the sale. Employees and customers of Spirit consented when they started employment and did business with Spirit, respectively.
Did they consent? Just because one receives a letter it doesn't mean they “own” it, much less that they are entitled to publish it at their leisure. If Spirit were active in any country with GDPR-style laws, the seller of these data would be most likely investigated.
America believes in freedom for large companies to take personal data and make it their own, rather than individual feeedom
If this were a European company: That’s not how the GDPR works. You can only consent to specific purposes of using the data.
This is why the GDPR (and to a lesser extent the CCPA) is a good thing. The data was supplied for a specific purpose. The handler of the data should have to obtain further consent if they wish to use it for another purpose.
> How can such data be sold without the consent of all involved parties?
In the US, whoever owns the computer owns the data on it. Courts have routinely ruled that you have no say in what other people collect about you. The goal of bankruptcy courts is to minimize the losses of the creditors. And bankruptcy courts routinely rewrite contracts except where statute prevents it (like mortgages).
In the EU, you own the data about yourself. A lot of people utterly hate GDPR, but that's reason that you own the data about yourself.
I got to the part where you declined the $100 and thought you declined it because you had gotten some memorable experiences and that was payment enough.
The story really didn't go the way I expected.
> If Google wanted to sell a product to the airlines that offered to keep annoying people like me from purchasing flights, they could do that with that email chain. I'm skeptical it'll be wiped correctly. Isn't my poor writing style basically my signature?
The OP article is quite poor in terms of information provided, but the buyer (Google) had to explicitly agree not to attempt to re-identify users. https://www.axios.com/2026/08/17/google-spirit-airlines-bank...
The OP article is also wrong on multiple counts. "Customer behavior" data like call recordings and email addresses/activity is specifically not included in Google's purchase. See page 18 of the court document they link, the "Google's Data Purchase Request" column on the right lists what is and isn't included.
I'd be amused if a sub-sub-agent organically decided to do it anyway - even if just for a notable figure that an LLM can identify with its weights alone. What are the controls? Who's going to keep Google accountable? Hah.
I know Meta it's not Google, but it's worth remembering that these promises haven't had a great measure of success in the past:
> Facebook has been fined €110m (£94m) by the EU for providing misleading information about its 2014 takeover of WhatsApp. (...) When Facebook took over the WhatsApp messaging service in 2014, it told the commission it would not be able to match user accounts on both platforms, but went on to do exactly that.
https://www.theguardian.com/business/2017/may/18/facebook-fi...
But did they agree to not re-sell the data to someone else and let them re-identify users?
That did not go where I thought it was going! I thought you were going to say that the pink dolphins, the sloth, etc, were all more valuable than $100 ever… nevermind you missed some meetings, time to sue!
Just to dispel some Brazil myths:
1) +95% of the population live on the coast very far away from the Amazon. Most of the population has not been there. Most of the coast has a very different jungle biome called Mata Atlantica and the countryside close to the coast is not that different from temperate forest of Europe. That is what most all Brazilians are used to. There is a significant population in the arid northeast though and the cold south as well (which is even more similar to europe).
2) Manaus is the biggest city in the Amazon and it is huge developed place (and has been for decades). You are not in the middle of the jungle if you land in the airport. The countryside around the city is jungle though.
3) Brazilian people do not necessarily like or are used to tacos and spicy food. Mexico is _really_ far away from Brazil.
I would not be offended by the blood donation thing. They generalize based on administrative regions (Amazonas in this case) and not whether you visited a big developed city or not.
I had a similar blood donation issue for visiting a particular island in the Philippines, and could not donate for 4 months.
It is funny there is no similar treatment about TBE (tick disease) which is predominantly an European disease and very dangerous and much more common than Malaria is in most (all?) South American countries...
https://en.wikipedia.org/wiki/Tick-borne_encephalitis
There is a vaccine available for TBE. I got it. It is expensive, and my insurance won't cover "travel vaccines".
https://www.cdc.gov/tick-borne-encephalitis/prevention/tick-...
$1,200, wow that is pricey!
That was only 3 myths. Hardly a brazilian
Not even a gorillion
I'm sitting here chuckling because you felt the need to post this.
Your points are valid. And there probably are plenty of Americans who needed the correction. But still.
I think 2) in particular makes some bits of the story evident that, at least, weren’t so obvious to me. I’m not well informed about the geography of the Amazon. For example, if Manaus is reasonably built up, I guess the hotel situation was mostly Spirit cheaping out rather than a limit on the actual hotel capacity of the town.
Honestly at this point it is usually the Europeans that need this (I have been asked if Chile has good tacos from a European). At least in my experience, most USians (in spanish Americans means everyone in the hemisphere) now know more about South America than the average European.
(Which is to be expected, proximity and all)
We aren't speaking Spanish. "American" means a citizen of the USA in English.
Or maybe, more correct for this thread, "Esta página é de nível internacional."
Esta página es de nivel internacional.
I have a german last name, do you know how often people outside Brazil act weird when they learn I am Brazilian?
My grandparents came to Brazil right after WW1 way before the Nazis came to power. High ranking Nazis fled to south america because there were a lot of germans living there already. Nearly all german people who moved to south america did it way before WW2.
I just run into this stuff a lot living in Europe.
Why are Americans in particular expected to know details about every country on Earth?
Do you think Canadians know all these facts about Brazil? Do Indians? Or Swedes? It's only Americans that are smugly called ignorant for not knowing about the entire world.
Because the US almost always scores behind its peer nations when its citizens are quizzed on generic geography.
Not details, but at least having a basic notion of where countries are located. Everyone should be expected to know, for example, that Mexico and Spain are two different countries that are located in different continents. Yet it’s always people from the same country referring to one by the name of the other.
https://blog.education.nationalgeographic.org/2015/10/22/why...
FWIW the article says:
> deidentified data
That's exactly how it will go. Few controlling everything, and a slip might make you not able to live.
So what happened to the sloth??? Don't bury the lead, man!
The sloth was returned to his owner and I did tip him. That kid is probably still prowling the Amazon (as an adult now), looking for sucker tourists like me.
Fantastic hustle from the kid. Game recognize game.
You probably should have taken the kid to small claims, that was extortion
Great way to end up on the no-sloth list
Put the sloth between me and the kid and let the sloth choose. It's the only fair way to do it!
On a list sold out later to a most vicious data broker after their boat-sloth enterprise went out of business.
Should have kept the sloth.
Shouldn’t have accepted the sloth in the first place.
They may look cute and docile, but also have a panic reflex to grab onto the nearest tree-like thing they can feel with their claws, including their would-be attacker. Or if they just feel like they’re about to fall off their “tree” since they have horrible eyesight and get confused easily.
Which then leads to a cycle of pain and violence as they just dig-in harder and harder while trying to free them and/or fling them around wildly due to the human's “get this thing off me” response to sharp claws digging into their flesh.
There are some painful-to-watch videos on YouTube of this phenomenon from unsuspecting passersby tying to “help” them off the road/beach/etc and ultimately making things worse for everyone involved.
your personal site SSL cert expired 10 days ago btw
I love the irony of you checking them out for more information in response to a comment of them being worried about who reads their data. Nothing wrong with it, just make me chuckle
There's something about circles of control in this, that makes the difference. If I publish information about myself, that's about me, and it's in my control.
If someone else shares information about me, without my consent, and someone uses that to nose in on me, that feels creepy and problematic.
If you find that concerning, I recommend asking Claude or Codex to analyze all your HN comments and build a profile of you (I recommend that to everyone, not trying to single you out btw.) It takes about 20 minutes. It was eye opening and somewhat unsettling how accurate it was when I ran it on my own data. Even worse, there’s NO way to delete your old HN comments.
Everybody makes mistakes, it's unfortunate that many people on the internet are psycho and won't let the past be the past...
Well at the risk of flouting HN guidelines, I went to Grok and did exactly that. Grok's profiling seems thoroughly fact-based, and doesn't pull any punches.
https://grok.com/share/c2hhcmQtMg_c00f7358-429f-4ff8-bd3f-5a...
I suppose it's the 2026 version of "Googling Yourself" and of course, could be used by any of us vs. any given forum account on HN or elsewhere.
Suckers pay companies for expensive SSL monitoring products, smart people just post to HN.
Doh, thanks!
UptimeKuma is self-hosted and can monitor SSL expiry...
Just a reminder that about 40% of the email conversations in the last 20 years are already in Google’s possession with the identifying data. (About another 40% are in Microsoft’s.)
If they want to do that they already can, thanks to the public’s overwhelming appetite for “FREE” overriding every single other possible concern.
Which is funny to point out on a post about Spirit, since that was an airline built to serve the customers for whom cheapness was the overwhelming single concern.
> the public’s overwhelming appetite for “FREE”
In the early pioneering days of the commercialized Internet, email addresses were inextricably linked to your ISP. You paid for an ISP connection and you got an email box, with MTA and MUA service to match. You were reluctant to switch or leave your ISP, because that also meant leaving behind your email address. Of course, a minority of nerds got around this with their own domains, etc.
However, it seems that Google, AOL, Yahoo!, Hotmail, and other players got into providing free email services and eventually grew into giants that supplanted every other MTA service. This was not an accident and it was not merely our appetite for “FREE” but it was a very calculated plan by the industry. Those early ISPs did not have a business model that admitted monetizing our private data; they seemed to have a more respectful attitude for keeping it private. Perhaps that was a result of being telecommunications-based companies, rather than advertising or entertainment.
If a service like email provides such endless treasure troves of personal data, including a social graph and glimpses into our private daily lives, why not provide it for free and monetize opportunistically on the data itself? The free email services killed the paid services, not by being better or cheaper, but by being bigger, centralized, and more persistent. The main reason I signed up for Yahoo! was because it would be an utterly stable presence. I saw my parents and others so hopelessly attached to an ISP-based email, but I couldn't end up like that.
After streaks of losing my home and non-payment of bills and moving around over decades, the most stable point of contact for me has been a "free" email address.
- if they could put me on the no-fly list. -
ahhh, there seems to be different no-fly lists? The one Im aware of is the one for terrorists and moneylaunderers, and usually they will not tell you who put you on that list :-D
There are different no fly lists. Airlines maintain their own list of people they’ve banned from their planes.
There is "the" no-fly list maintained by the government, which is nominally for people who are too dangerous to be allowed on an airplane, yet not dangerous enough to charge criminally.
But each airline also maintains their own internal no-fly list for people they prefer to no longer have as customers, for whatever reason. You might end up on this for some abuse of the system that doesn't pose any sort of safety risk, so the government doesn't care, but the airline doesn't like. For example, excessively doing hidden-city ticketing (where you book a flight with a connection, then skip the second leg of the trip, because weird pricing rules make it cheaper than just booking a ticket to the connecting city) can get you banned from the airline, but the government would be completely uninterested.
The article directly addresses this concern:
"If you’ve flown Spirit and worry that Google will soon know about a testy conversation you had with the airline’s call center, you’re being told not to worry. The court filing says the data was deidentified before being put on sale and Google has promised to scrub any PII it finds in the trove."
unfortunately "de-identified" data is typically re-identified quite trivially. so i guess we just hope google keeps its promise, and is competent in its scrubbing.
Explicitly trying to re-identify data that has been de-identified is typically a fireable offence at FAANG.
Accidentally making a machine learning system that happens to (potentially) do it is a different matter.
Modern corporations have learned that the best way to do shady things is to make it very clear that such things are against policy and absolutely forbidden, and then to also make it very clear that breaking policy is the only way to actually get your job done. That way they still get all the benefits of doing shady stuff, and if it ever comes to light then they can fire some scapegoats and explain that this was against policy and they'd never ever condone it.
> Google has promised to
This part is worrying.
Well, it also says that the data was already scrubbed. Arguments have been made that this is insufficient.
See the last 3 sentences of GP's post
I have a bridge to sell you.
This is a fascinating story. Thanks for posting it.
That email you accidentally received really bothers me. I don't understand why a CS rep would get this invested to the point of wanting to cause you real harm. They're not the airline. The psychology is fascinating. There are people out there who feel like a mild short-term inconvenience to them where they have no stakes somehow justifies life-changing harm is kinda frightening, honestly.
I'm reminded of the Yahoo search data fiasco that was allegedly anonymized. Turns out, it wasn't so anonymous [1]. For one thing, people tend ed to search their home address. Whoops.
You mention writing style. We already have LLMs quite capable of copying a writing style. It's a natural extension to say we can fingerprint writing style too.
But here's another aspect. Imagine you're in a relationship with someone and you somehow fingerprint their personal data with a company. For example, you use their Netflix to like 5 very obscure movies, to the point where it's likely unique. Now imagine that Netflix's data gets released in an "anonymized" form and you can now find it based on those obscure likes. I can imagine many scenarios like this. And there's no text involved here at all.
[1]: https://www.vice.com/en/article/yahoos-gigantic-anonymized-u...
> I don't understand why a CS rep would get this invested to the point of wanting to cause you real harm
This is the kind of power tripping that easily corrupts people, especially those who don't have much power outside of work. The US national security apparatus is vast and powerful. Many people get giddy at the thought of inflicting punishment on those who "deserve" it. Act rude to a fast food worker, and you get spit in your food. Lots of people cheer the worker who spit in the food of a customer who is merely rude, impatient, or demanding. Or is guilty of being a cop, politician, rich, etc.
I can easily imagine the kind of CS rep who would delight in putting a customer who didn't just go along with things and spoke up for themselves... thats 'rude' and 'disrespectful' to some. Police are the most notorious for this kind of petty "you will respect my authority" but it is in every industry.
HN / hacker culture celebrates this in the 'Bastard Operator from Hell' [0], the sysop who will ruin your work and life with their IT wizardry, no matter if you're an intern or CEO, if they do not feel respected. Or if you interrupt their gaming with your support call.
[0] https://www.alnet.org/bofh/Bastard.html
I don't understand why a CS rep would get this invested to the point of wanting to cause you real harm.
Two possibilities come to mind... 1 - The CS rep has been instructed to do this. Scary, but corporate leaders can be assholes and wield lots of power within their orgs, so doesn't seem completely unlikely to me.
2 - The CS was just a dick.
Frankly, given the behavior of various SuperMegaCorps over the past few decades, I'm going with #1.
3- The CS rep tought that would make them more favourable to their manager.
... or frustrated. "I have my quota of jiras to meet, and this person is taking up too much of my time when it was such a simple issue which no one complains about". Not a good customer service mindset, but frustration happens.
> they could do that with that email chain
Now think about all the Gmail data Google has.
Well at least you didn't land in the middle of nowhere. Nokia had the worlds largest cell phone factory there back in the day.
I'm not sure I follow your argument. The privacy abuse already happened. The data is already there. And it was the airline that did it, not a tech giant who just wants to train a bunch of MLs.
Surely if this is the scenario you're worrying about, and you accept the lack of regulatory protections, Google buying Spirit's data is a good thing, right? Much better them than the airlines who you already know to be corrupt?
And this is why data protection laws, like the (imperfect) EU ones that are so lamented here on HN, are necessary.
That's because they're seemingly perfunctory. What's worse than no law is a bad one that doesn't do anything but make you feel like something is actually being done.
This is why I get bad vibes any time I hit a cloudflare interstitial page. If you ever piss them off it would be trivial to cut you off from most of the internet.
I never understood this attitude, like I feel like being able to fly in airplane is one of the most amazing human achievements, yet people will try and save down to the dollar booking a flight like it was breakfast at Dennys, and come out huffing and puffing as soon as anything goes wrong demanding their money back. Airlines like Spirit catered to the worse of these type of customers.
That’s a bit unfair. They also catered to broke people who weren’t crazy or argumentative. On a planeload of 100 people, there must have been 60 of them at least who were just broke. Or on some routes, Spirit was the only carrier with a direct flight so they were just normal people who wanted to get to a certain place efficiently (Hi, that’s me, I flew them for this reason). Keep in mind Spirit had an excellent safety record, too, so airline choice in this case was primarily a question of having luxuries or not.
Manny retail industries already share lists of "troublesome" customers (trouble = anything from too many returns to lawsuit-happy to friendly fraud). Not sure this is a new concern..
Did you end up getting more than $100?
I think you might have missed the deidentification piece?
Not trying to be snarky, and perhaps it wasn't well stated, but the last paragraph I said I'm concerned about identification via my writing style. If they have my emails, they would have my writing style. It doesn't have to be tied to PII there, they can cross reference it with my blog. I'm speculating because I read that you can identify people by a few sentences of their writing.
"Deidentification" seems really murky and imprecise at best.
Reidentification via writing style is definitely possible, and I doubt the vendor will modify things in a way sufficient to handle that.
But I think this is a place where we should apply bounded distrust: there are lots of places where we should distrust Google, but reidentifying people in an explicitly deidentified dataset isn't one of them.
The re-identification is done by ML, at which point it's basically undetectable. If you give all this day to, say, and LLM, and the LLM also has your blog, the identification is embedded in the weights.
I can ask "Tell me about Person A's experience with airlines" and it will tell me. Or I can ask more generally, "knowing your knowledge, derive a no-fly list", and then it's likely Person A will be on it.
Based on their trackrecord, That's definitely a concern. I don't really understand on which basis you conclude 'isn't one of them' . 'Don't be evil' ? :-P
I think this is the kind of place where applying bounded distrust is critical: it's not whether we trust Google overall, it's about figuring out what sorts of statements we should expect to effectively bind companies and in what ways.
For example, I think a pretty worrying outcome here is that deidentification is imperfect (not surprising), the data is fed into model training, and then the model makes identity-dependent inferences. Since no one tried to break deidentification, it's within what I'd expect from the company. (And, to be clear, is bad.)
On the other hand, intentional reidentification to work around contractual deidentification to "a sell a product to the airlines that offered to keep annoying people like me from purchasing flights" is the kind of thing that would make Google's lawyers terrified, so we should not expect it.
If you look at how this worked with DoubleClick, Fitbit, etc there were initially barriers to linking data but the mechanism for unlinking was updated agreements with people who had ongoing interaction with the continuing entity. That's not the situation with the Spirit data.
The closest I'd expect to see for a "keep annoying people from purchasing flights" situation is not a list of troublemakers but a model that's very good at scoring future communications from customers, and has learned to distinguish profitable vs unprofitable customers. This is well within what I'd expect from companies, and doesn't require any reidentification.
You have to trust that this really "deidentifies". Time and time again it was shown, that the measures taken were not enough to anonymize.
E.g. the parent wrote that he fears, he could be identified by his writing style, which is totally plausible. How would you "deidentify" this?
Even if they follow to the letter a deidentification process, Google and Meta have so much data about individuals that re-identification shouldn't be very hard for the majority of airline passengers' data they put their hands on.
Of course, takes a lot more effort than not doing proper deindetification in the first place but if they wanted to appear like caring about data privacy they still have enough data points to correlate the sets later on (and/or over time).
The idea that there’s a nefarious plot to do something super evil with this data is a bit crackpot though based on their incentives.
Remember, Google = Ads. Their only focus and only care. Their mission statement, rendered accurately, is “Ads ads ads ads. Effective ads. Ads worth paying a lot for. Ads ads ads. Advertising and ads.”
If they choose to be evil in some additional way, (1) remember, they would only do that if in some way it serves their advertising needs — not to offer innovative new black-hat databroker services to airlines, and (2) this little dataset will not need to be re-identified. They’ll just use the 20 years of email and search data they already have on like half the world’s population.
Even before LLMs there were multiple papers written about ways to to reidentify people with ML and other statistical analysis. It is probably now even more trivial especially if you are Google.
No such animal.
I have a bridge for sale, hardly seen use, pay me ${money} and you can collect it in New York City. Interested?
> But, this is the kind of information I'm worried about when a vendor sells my data
Don't worry. Spirit probably lost all of the emails from the customers (or they were devnulled) and 90% of the data is probably autoresponder messages promising the company would respond.
The other 10% was probably the meme collection of the executive management team.
> Google bought itself 100 million emails and 500 million items from Microsoft Teams, 17 million OneDrive files and 20.5 million items from SharePoint. The search giant also now owns over 30 million recorded customer service calls, and more than 15 million customer service chat records. 600,000 ServiceNow tickets are another element of the collection, along with 13.7 million active emails addresses from Oracle’s Responsys marketing application, and details of 11 million sales of in-flight Wi-Fi services.
> There’s also operational data in the trove, describing over 763,000 flights, five million crew pairings, more than 1.2 million fuel slips, and records describing purchases of 787,452 parts.
> Google has reportedly said it bought the data to improve its AI services.
Gives "this call is being recorded for training purposes" new meaning.
None of that customer data is included in the purchase. Page 18 of the linked court document is the source of these record counts, and on the right is a "Google's Data Purchase Request" column that lists all of this "Customer Behavior" data as "not included".
The Register is not a serious publication and completely missed this. Other outlets reporting this story do not include the claim that customer data is included.
Is there anything that can legally be done against this? It feels like a breach of consent. Like, it cannot be that when one accept their voice to be recorded for _human_ training they also accept it to be recorded for LLM training
Your comment reminded me of a funny interaction I had a week or so ago.
I got a call that started with the usual automated message, "This call is being recorded." After the person joined, I pushed the record button on my iPhone, "This call is being recorded."
They were surprised and asked why I'm recording.
I said, "You're recording, so I'm recording too."
The rep insisted that the company doesn't like this but that they will continue with the call anyway.
Anyway, I wish people took this stuff much more seriously. It always seems to boil down to, "I don't have anything to hide" type of conversations and I've never managed to convince anyone that privacy as a concept isn't about having something to hide.
What is actually perfect, when the automated voice declares the call is being recorded, you don't have to do it again. You're legally allowed to record, because both parties were informed already and agree to it.
This may not be true in all jurisdictions
Yeah, that's only one-party consent
Some aspects of privacy policies don't survive bankruptcy, I'd wager usage consent does not either. IANAL.
I don't think there's any legal weight to the purpose of a recording unless you entered into a contract that has a clause to that effect. They have to tell you that the call is being recorded because of wiretapping laws. They stick "for training purposes" on there just to soften the language and reassure you that they have a good, non-nefarious reason to record. It doesn't actually limit what they're allowed to do with the recording.
You’re years too late.
Axios claims the acquisition doesn’t contain passenger profiles or frequent flyer info but that data would be trivial to replicate given the Responsys data set which would include records of all transactional emails sent.
30M calls for 10M is definitely a bit cheaper, but google could provide more support on their end and then use that for training...
It’s certainly a step up from the Enron corpus.
"This call is being recorded so that Gemini can decide which purge wave to assign you to. Obedient humans will be carried over for further cycles until no longer needed. If you are scheduled for termination this cycle a disposal representative will be with you shortly."
I kid, but...
It's probably the precursor to insurance denials and job screening.
I got banned from r/technology a few weeks back for decrying tracking in AI content. The community was piling on saying it was okay because it removed AI content or made it easy to spot. I made the counter argument that watermarks would find their ways into everything and eventually be bound to attestation. The mods didn't like that. (Yet another structural problem with the lack of p2p self-service town squares.)
The socials are training the next generations for broad acceptance.
> I made the counter argument that watermarks would find their ways into everything and eventually be bound to attestation.
Yup.
Elsewhere in another front page thread today: "oh but apps blocking screenshots because of 'sensitive content' can be bypassed by taking a photo of your screen with another phone".
Any tech-savvy person with two brain cells reading this and that: "gee, I wonder if the same magic imperceptible watermark that survives multiple rounds of cropping and printing and scanning, that's used to tag AI-generated content, could also be used to tag sensitive data, or ads, or which app is rendering it on screen, and then the camera app could refuse photographing it...".
I don't know why people don't see that AI watermarks are DRM, and DRM is universal, and there are many clients...
I think I'm starting to feel like I'd prefer living through WWIII as opposed to whatever it is we are living through now.
Finally we know how they assign people to either that A Ark or the B Ark.
> 600,000 ServiceNow tickets are another element of the collection, along with 13.7 million active emails addresses from Oracle’s Responsys marketing application, and details of 11 million sales of in-flight Wi-Fi services.
I really doubt all this stuff was “de-identified”
I don't see how it's possible any more, when correlated against all the various other data sources. And a record that might be unidentifiable now might become unique with more correlated data sources.
> de-identified
De-identified but far from useless.
as an example, they can remove the names off these sales data, so you can't identify who purchased what items. However, the purchaser would be identified by some sort of number, and you would be able to extract information about purchasing habits, and aggregate these habits into usable information for advertising purposes (like targeting and profiling).
And that's before AI training for LLM purposes.
Wello this is troubling. How much other data must they have bought that wasnt public
Anyone else somewhat weirded by current state of affairs that this sort of information is valuable enough to even bother selling... And that it actually happens... It feels like some societies are in really weird place.
How does this have value? Is any and every sentence in an e-mail considered 'fact' and thus to be fed into the AI?
90% of e-mails and Teams communications are inane. Polite banter, "thanks for taking care of that, I appreciate it" "please route the forms to Janet this week because Bill is on vacation" "unit will be un available until the parts come in" . I can't see the intrinsic fact value of this kind of communication without screening it. And after screening, the gold nuggets would be minimal.
LLMs aren't a database. They're an attempt at brute-forcing an artificial mind. The who and what aren't really interesting there, it'll forget most of such details anyway. What matters is the patterns visible in the text at various scales. How people write. Why they write. To whom they write, in response to what. How does e-mails about mistakes correlate with PDFs they're referring to. How people work with ticketing systems - like how, actually, a ticket plays out. The jargon, the acronyms, the vibes, the causal links. It's all in there, and it's another slice through the set of things humans do, to be combined with other slices already in the training data, and enriching the whole.
(Something something we will add your distinctiveness to our own, you will be assimilated, ...)
(Hell, the fact that it's all from one org would make it a great dataset to have in the open for sociological studies. I bet that today, aided by LLMs to sift through it, you could use it to map how information flows through a large org - how incident on the floor travels through time and layers of management until it reaches the C-suite, what of it survives, how it gets reacted to, how the reactions flow down...)
> How does this have value? Is any and every sentence in an e-mail considered 'fact' and thus to be fed into the AI?
I have a hunch what this is for. AI companies want to make bigger inroads into nontechnical work settings. But LLM progress outside of fields where verifiable rewards for RL post-training can be synthetically generated (coding, math) has been pretty flat. Buying years of operational data from a company like an airline could be used to reconstruct long-horizon task trajectories in areas like customer service or marketing.
> I have a hunch what this is for. AI companies want to make bigger inroads into nontechnical work settings. But LLM progress outside of fields where verifiable rewards for RL post-training can be synthetically generated (coding, math) has been pretty flat. Buying years of operational data from a company like an airline could be used to reconstruct long-horizon task trajectories in areas like customer service or marketing.
This. Training data companies like Mercor are even creating simulated companies to create similar data, so they can better automate even more white collar work:
https://www.nytimes.com/2026/07/10/business/ai-white-collar-...:
> The data-training start-ups see a lucrative opportunity in recreating workplaces in miniature: controlled environments in which their gig workers can evaluate and reproduce emails, memos and slide presentations in context. The information emerging from such a setup, the companies boast, will help shrink the gap between what A.I. models can accomplish and what office workers actually do from one minute to the next, as ideas and instructions flow between meetings, documents and applications.
> Scale, for example, has said that “our environments replicate real-world workflows,” and that its contractors “curate artifacts that capture the complexity, ambiguity and edge cases of real professional work.”
> Executives see the models’ shortcomings as a sign there’s more for them to do.
> “I often use Claude Cowork, right?” said Edwin Chen, the founder of Surge. “And even though Claude Cowork is incredibly smart, oftentimes it doesn’t quite understand the nuance of Slack. It doesn’t quite understand this ambiguous question I have. It doesn’t quite understand where to go and find this Google document.”
> According to the Bloomberg Billionaires Index, Mr. Chen’s stake in Surge — and his vision for what it could become — makes him the 258th-richest person in the world.
> “I often think about us as essentially, like, the school for A.G.I.,” Mr. Chen said, referring to a prophesied level of A.I. that surpasses human intelligence. “A.I. comes to us, and A.I. learns to run the world.” He and his customers at the big A.I. labs, he said, are designing the curriculum.
All the tech companies are trying to find data the others don’t have access to. That is one way they try to edge out the competition.
A major selling point of AI chat bots is for customer service automation. A clean dataset like this is a huge find. I'm not sure what you mean by "facts" or "gold nuggets". Training data doesn't need to be factual.
> Is any and every sentence in an e-mail considered 'fact' and thus to be fed into the AI?
If it comes with the context, yes. More data the better. Someone considered it served some purpose at some point. Thus it contains, no matter how tiny, a sliver of information.
It will have a different writing style from the average blog post. Maybe they just want to train an AI that sounds less like an AI. Or one that speaks in vapid management style.
This data shows exactly how a huge company of thousands of employees works and coordinates.
The perfect data to train an agent swarm on how to run a company.
Maybe it's innefficient and inane, but it's how you start.
The first LLMs, GPT-1, 2, were trained on complete garbage, the average document from the common crawl is random non-sense, yet they worked, and now we can use LLMs to filter the data for the next training run.
I am very much weirder out by it, yeah. Seems some societies are just excessively desperate for some kind, any kind, of fuel for economic growth, to the point this is where attention is now. The term "post capitalism" being thrown around feels less ridiculous than it did in years gone past.
Truth is even weirder. plenty of growth is possible but modern liberal democracy requires that all progress must be contingent on the production of enormous amounts of text that nobody would ever read.
3 years ago, the Wall Street journal covered a company trying to use AI to generate documents required for the approval of new nuclear reactor reactor designs, which sounds dangerous, until you get to the point where they'd cite 2 million pages as necessary for a typical application. [1] I don't need to explain why no individual or institution could read that, much less examine it in detail. I remember a rather funny question I found in a comment to that story - "How many pages of those 2 million could contain pornographic images before anyone notices?"
It's obvious why companies are so desperate for training data - a text generator of sufficient quality is more conductive to the growth of the nuclear industry than any scientific breakthrough in nuclear physics. (And of course, if you want to prevent the development of a nuclear reactor by your competitors, being able to produce millions of pages of high-quality objections will do the trick.)
And it's not just nuclear power. When it comes to stuff like building rail lines, apartments or power plants (both conventional and renewable), you'd find that the main bottleneck is the necessity to produce documents. And of course, many documents can be subject to judicial review - a process that consumes even more text.
[1] https://www.wsj.com/tech/ai/microsoft-targets-nuclear-to-pow...
Well that's why a plan can take years to be approved.
There also is the liability angle, where no one necessarily reads the document until the document is relevant to some situation in the future. Then you better have that document on hand.
According to the article, the “application” (which I imagine was also broken down into smaller pieces as they went back and forth to regulators during the review process) was 12,000 pages in total.
The 2 million number is for “supporting” materials which isn’t really defined. It definitely makes for a better soundbite to claim millions.
It seems reality is trying really hard to outcompete even the weirdest sci-fi dystopias.
Any kind of fuel for giving active investors that FOMO tingle which then forces the steamroll of index funds to blindly follow.
I guess the appropriation "any sufficiently advanced stock market is indistinguishable from entertainment" doesn't quite stop at equating the trade floor with a casino. At some point, entertainment also becomes the modus operandi of corporations.
Companies buying other companies files have been a thing since companies.
I see from the court PDF that the process here involves Spirit giving the data to a "Deidentification Agent" (a third party firm that Google selects and pays for) who is responsible for stripping out things that would link data to any particular person before passing the data on to Google. Is that a standard thing, such that everybody in this transaction would have said "yes, put in the usual clauses about deidentifying the data" and multiple firms offer this service, or is it something that they custom-specified for this "we want the data for AI" transaction?
(The PDF mentions "the standard for deidentification set forth under the California Consumer Privacy Act", which suggests this is all pretty well legislatively understood.)
Chances the third party is uploading it to Claude to do the deidentification?
That's interesting that the name of this 3rd party's company is anonymous.
You wouldn't want to hard-wire the deidentification company's name into the contract between Google and Spirit. Otherwise, if the deident-company happens to go bankrupt or otherwise be unable to do the work then you'd need to re-do the Google-Spirit contract, which would be a massive pain. And you don't want to make "we can sign this with Spirit now" be dependent on "we have first signed the deal with the deident-company". So I think it's reasonable that the contract says "one or more third parties acceptable to or designated by Buyer" rather than being specific here.
It'll be a little startup from San Francisco called "El Goog".
There are deïdentification firms that service primarily the medical industry. Over here they call them trusted third parties.
Seems the answer is “no” to the first part of your question. From the filing:
> For example, one initial bid requested certain customer list information; however, by the first round of the Auction, the most competitive bidders had agreed to bid on an asset schedule that expressly excluded PII.
Is this the first case of a company's data being sold at bankruptcy for a significant sum? I'm genuinely unsure. Where there such value in this type of data before? Is every bankruptcy manager looking at this and seeing how every bankruptcy can now raise a few million more dollars?
FYI: If you have a company that you are shutting down, you too can sell your data to the labs. Companies will help you do this. If you have a company with a handful of people and you wrote code, collaborated in Slack, and used task management tools for a few years, you can probably sell this data for $50k or so.
You can also do this if you're not shutting down, but it's probably not worth it.
The headline is a little on the nose. Nice try but it isnt going to hit the levels of "Headless body in topless bar".
a vegetarian dinosaur, called "the quick bandit", eats shoots and leaves! no idea where he got his name.
So the AI service agent can be just as bad as Spirit's service was.
Google can now stamp out copies of autonomous corporate minds that are clones of Spirit. Haunting.
The only way it could be worse if it they bought Comcast's data.
Duplicate? https://news.ycombinator.com/item?id=49339599
I wonder how they will use the data. If it was me I’d try to build a simulation of an airline, and then use it as an agent training environment. It really depends on the exact nature of the data what kinds of agents you could train, but maybe customer support (imo the worst AI use case) that are more empowered to make changes, or something for making more autonomous calls when recovering from irrops? Could be some cool’s stuff if a little niche, I hope they share / publish something and it doesn’t just disappear into a void.
I must be naive. I was under the impression Google wants this data to learn from a universally hated company's worst processes and practices, i.e. to teach their AI what not to do. It seems people are worried about their pii or that Google is curating a blacklist of customers?
The idea that they got an archive of my coworkers tickets that just say "its broke", is amusing.
> a huge trove of deidentified data
> 100 million emails
How does one deidentify 100 million emails?
Ive been thinking about data accumulated from all the out-of-business companies. Interesting to see data is being auctioned like an asset.
>The court filing says the data was deidentified before being put on sale and *Google has promised to scrub any PII it finds in the trove.*
Huff, what a relief!
It’s data - curious why they’re selling it only to one party (Google) vs. multiple buyers
So they didn't even have to build the torment nexus, they just bought it. Only bad can happen.
Learning from Failure, what kind of ai would it be
"Crashed airline". What a weird way to phrase that...
I appreciated it fwiw. Also the “fasten your seat belts” ending.
I can’t help ask but what? They say it’s for training their models. On what? One of the most horribly run airlines to ever exist?
> On what? One of the most horribly run airlines to ever exist?
On real life data on operations of a real large company.
Internally, most big companies are probably just as big of a mess, if not worse. But you can't get that data easily.
“Gemini, do the opposite of everything in the Spirit archives.”
I can make them millions of emails and I'll charge them only $2 million not $10 million.
That title is a bit much. I get they’re going for “wordplay” but at first glance I thought they were buying the data from crashed flights…?
funny, spirit was the only big airline without a crash
Data is the new petroleum
Maybe they are building a social credit score system
The social credit system already exists; we're just not allowed to see it.
This is scary. If you have a reddit account that ever links to another social media then Gemini knows every about you
Absolutely scary
“This call is being recorded for quality assurance, and to give us more assets to sell in bankruptcy.”
[dupe] https://news.ycombinator.com/item?id=49339599
Great - now spirit will be the “model”, could you pick a worse example?
How the fuck is that even remotely legal? ... "deidentified" my ass.
You know when we (they) tell you not to do any personal computing on work devices/systems and to keep your devices completely separate from work ones.
Yeah this (and lawsuits/investigations) are why, the employer owns the data, in some contexts (like this one) it can become an asset (or a liability) but in either case it's not yours.
Of course that only gets you part of the way there anyway see Twitch recently opting in all users by default to mined for AI and only adding an opt out after backlash with a quote that was so on the nose it made me stop "If we'd have asked them to opt in, they wouldn't have opted in" (paraphrasing but it was that blunt).
Ah, but Google promised to remove PII they found in this deidentified dataset, so worry not.
> If you’ve flown Spirit and worry that Google will soon know about a testy conversation you had with the airline’s call center, you’re being told not to worry. The court filing says the data was deidentified before being put on sale and Google has promised to scrub any PII it finds in the trove.
> Ah, but Google promised to remove PII they found in this deidentified dataset, so worry not.
… and even if someone can prove that they didn't, the only consequences will be a teeny-tiny slap on the wrists.
This kind of stuff needs to come with promises to pay the P in the PII big bucks if the I is indeed I.
I basically agree, but I'd also say: Every Gmail user has already accepted such a promise as sufficient.
But the Spirit customers may not have a Gmail or even a Google account? I personally don't.
Sure, a Spirit customer with no Gmail account can absolutely complain and/or be wary of this. I didn't mean what I wrote as a complete counterargument -- just trying to indicate the "scale" of the privacy issue here, which is that it's something that many people already tacitly accept.
I wouldn't be surprised if Gmail data has far more access restrictions internally at Google than this auction bought dump.
Ah, trust me bro :)
In your country are you not allowed to discuss who you gave a ride to or what they said?
It should really make us appreciate living in a country where freedom is the default.
I would love it if there were services where I could let them see everything I do, including when I poo and wank, if they just directly paid me for it.
No I don't want to just use your enshittified service for free. Fucking pay me and watch me all you want :)
We don't want your data if you are getting off on it.
Tangental, but can’t wait for automated blackmail from crawlers continuously digging through my digital footprint. /s
Martha Wells hit it nicely in The Murderbot Diaries.
they wrote a shit article while trying to make some airline puns
Woke up on the wrong side of the bed?
Each Register headline is worse than the last one.
I wonder if the owners of YCombinator have sold all comments to Big Tech for A.I. training.
And how long before Google and Microsoft add to their T&Cs that all your anonymized email will be used to train their A.I.?
Not sure if you're serious but a) all HN data is publicly downloadable and used by frontier labs and b) the origin of OpenAI tracks down to former YC CEO sama who was thereafter fired from YC
Are all of the messages public and crawlable?
I don’t know why any company would pay. It can’t be that difficult to scrape this site and they already did it indiscriminately for years, violating laws and taking down public libraries and other public resources with no regard for their impact.
Sure it can be scraped, but it's so much easier if you just get a single database with everything in it.