Share:

Thousands of people are selling their identities to train AI — but at what cost?

Human identity has always held value, but it has never been so literally negotiable as it is right now. With the explosion of artificial intelligence, a new kind of market has emerged almost silently: platforms that pay ordinary people to share their voices, faces, conversations, and even the ambient sounds around them — all to feed the models that are reshaping the digital world.

We are not talking about science fiction or something confined to Silicon Valley. This reality is already part of everyday life for thousands of people spread across every continent, from college students in India to welding apprentices in the United States. And the most striking part is that this market has grown so fast that most of the people involved still do not fully understand what is at stake.

Real stories from the people feeding the machine

Jacobus Louw, 27, lives in Cape Town, South Africa, and filmed his morning walks feeding seagulls for a task called Urban Navigation on the Kled AI app. A single video of his feet and the view as he walked along the sidewalk earned him 14 dollars — roughly ten times the country’s daily minimum wage. In two weeks, he racked up 50 dollars just by uploading photos and videos from his daily routine. Louw had struggled with a nerve disorder for years and could not hold down a formal job, but the money earned through AI data marketplaces allowed him to save up for a 500-dollar massage therapy course.

Thousands of miles away, in Ranchi, India, Sahil Tigga, a 22-year-old student, earns money regularly by letting the Silencio app access his phone’s microphone to capture urban noise — like the sounds inside a restaurant or traffic at a busy intersection. He also records his own voice and travels to capture unique environments, such as hotel lobbies that have not yet been documented on the app’s map. Doing this, he pulls in more than 100 dollars a month, enough to cover all of his food expenses.

And in Chicago, Ramelio Hill, an 18-year-old welding apprentice, made a few hundred dollars selling his private phone calls with friends and family to Neon Mobile, a conversational AI training platform that paid 0.50 dollars per minute. For Hill, the logic was simple: he figured tech companies were already capturing most of his personal data anyway, so why not at least get a cut of the profit? 🤔

These three profiles are not outliers. They are a snapshot of a gig economy that is growing fast, fueled by big tech’s hunger for quality human data and the very real need of thousands of people around the world to find new sources of income. But there is a side to this equation that rarely shows up in the terms of service — and that is exactly the side worth understanding before hitting any accept button. 👀

Why AI companies are desperate for human data

Behind every language model like ChatGPT or Gemini, every facial recognition system or voice assistant, there is an endless need: training data. And not just any data — real human data, rich in context, diversity, and naturalness. That is what separates a mediocre model from one that actually seems to understand what you are saying.

The problem is that the most commonly used training sources — datasets like C4, RefinedWeb, and Dolma, which account for roughly a quarter of the highest-quality datasets available on the web — are increasingly restricting access for generative AI companies. Researchers estimate that AI companies could run out of fresh, high-quality text for training their models as early as 2026. Some labs have tried to work around this by feeding models synthetic data generated by the AI itself, but this recursive process can cause systems to produce content riddled with errors and distortions, seriously compromising quality.

Receive the best innovation content in your email.

All the news, tips, trends, and resources you're looking for, delivered to your inbox.

By subscribing to the newsletter, you agree to receive communications from Método Viral. We are committed to always protecting and respecting your privacy.

This is exactly the gap where platforms like Kled AI, Silencio, and Neon Mobile step in. On these data marketplaces, millions of people are monetizing their identities to feed and train AI models. Beyond those three, there are several other options: Luel AI, backed by the famous Y Combinator incubator, buys multilingual conversations for about 0.15 dollars per minute. ElevenLabs lets you digitally clone your voice and authorize anyone to use it for a base rate of 0.02 dollars per minute.

Bouke Klein Teeselink, a professor of economics at King’s College London, has said that AI gig training is a new emerging category of work and one that will grow substantially. According to him, AI companies know that paying people to license their data helps avoid the copyright disputes they would face if they relied solely on content scraped from the web. Veniamin Veselovsky, an AI researcher, added that these companies also need high-quality data to model new behaviors in their systems. In his view, human data is, for now, the gold standard for out-of-distribution sampling. 💡

The gig economy has a new face

The gig economy is nothing new. Rideshare drivers, delivery workers, and freelancers are well acquainted with the on-demand work model — no formal employment contract, variable pay. But what is happening now with the AI data market represents a significant evolution of that model, because the product being delivered is no longer a physical service or a specific professional skill. It is the person’s own identity — their voice, their image, their behavioral patterns, the way they express themselves. That completely changes the nature of the relationship between worker and platform and raises questions that traditional labor law still does not know how to answer properly.

The humans feeding these machines, especially those in developing countries, frequently need the money and have few other ways to earn it. For many AI gig trainers, doing this work is a pragmatic response to economic disparity. In countries with high unemployment and devalued currencies, earning in U.S. dollars is often more stable and rewarding than local jobs. As Louw himself put it: as a South African, being paid in dollars is worth more than people realize.

Even in wealthier nations, rising cost of living has turned selling personal data into a logical financial decision for a lot of people. Some of these workers struggle to land entry-level jobs and turn to AI training out of necessity. The model works in a fairly straightforward way: platforms post tasks with specific instructions — record your voice reading certain phrases, film your hands performing common movements, or capture ambient sounds at specific locations. The worker completes the task, submits the data, and gets paid, usually via PayPal or credits that can be converted to cash.

It sounds simple, and for many people it works well as a side income. The problems start to surface when you read the fine print in these platforms’ terms of service. 😬

Blanket permissions and the invisible risks

There is an important difference between passively sharing data — like what happens when you use social media or navigation apps — and actively selling data as a form of work. In the second case, the person is making a conscious decision, but not always with all the information needed to assess the long-term consequences.

In most cases, the transfer of rights over the data is total, irrevocable, and permanent. When AI trainers share their data on platforms like Neon Mobile and Kled AI, they are granting a blanket license — worldwide, exclusive, irrevocable, transferable, and royalty-free — to sell, use, publicly display, and store their likeness, and even create derivative works from it. That means a 20-minute voice recording made today could power a customer service bot for years to come, without the trainer ever seeing another cent.

Avi Patel, founder of Kled AI, has stated that his company’s data agreements limit usage to AI training and research purposes. According to him, the entire business depends on user trust, and the company vets buyers before selling datasets, avoiding working with companies that have questionable intentions — such as pornography — or government agencies that might use the data in ways that conflict with that trust.

Neon Mobile, on the other hand, did not respond to requests for comment. And perhaps that says plenty about how these platforms approach transparency.

Jennifer King, a data privacy researcher at the Stanford Institute for Human-Centered Artificial Intelligence, points out that the most troubling aspect is the lack of clarity about how and where users’ data will be used. Without negotiating or knowing their rights, consumers risk having their data repurposed in ways they do not like, did not understand, or did not foresee — and they will have little legal recourse.

Enrico Bonadio, a law professor at City St George’s, University of London, goes further: according to him, the terms of these agreements allow platforms and their clients to do virtually anything with this material, forever, without any additional payment and without any realistic way for the contributor to withdraw consent or renegotiate. More serious risks include the use of data for deepfakes and impersonation. Even though data marketplaces claim to strip identifying information like names and locations before selling the data, biometric patterns are inherently difficult to anonymize in any robust way.

When a bargain turns costly: stories of regret

Ramelio Hill’s case illustrates well how things can go wrong. For about 11 hours of phone calls, he earned 200 dollars from Neon Mobile. But the app frequently went down and delayed payments. In September, just weeks after launch, the platform went offline after TechCrunch uncovered a security flaw that allowed anyone to access users’ phone numbers, call recordings, and transcripts. Hill said Neon Mobile never informed him about it, and now he worries about how his voice could be misused online.

Even more telling is the case of Adam Coy, a New York actor who sold his likeness in 2024 for 1,000 dollars to Captions, an AI-powered video editor now called Mirage. His contract included more detailed protections: his identity could not be used for political purposes, nor to sell alcohol, tobacco, or pornography, and the license would expire in one year.

Even with those safeguards, it did not take long before friends started forwarding him videos found online featuring his face and voice racking up millions of views. In one of those videos, an Instagram reel, Adam’s AI replica introduced itself as a gynecologist and promoted unverified medical supplements for pregnant and postpartum women.

It was embarrassing having to explain that to people, Coy said. He admitted that the decision to sell his likeness came from a logic similar to Hill’s: if most models were going to scrape the internet for data and images anyway, at least he might as well get something for it. Since then, however, Coy has not signed up for any other AI data tasks. He would only consider doing it again if a company offered truly meaningful compensation.

Tools we use daily

The future for AI data workers

Mark Graham, a professor of internet geography at the University of Oxford and author of the book Feeding the Machine, has acknowledged that for people in developing countries the money can be significant in the short term. But he warned that, structurally, this work is precarious, non-progressive, and effectively a dead end.

According to Graham, AI marketplaces depend on a race to the bottom on wages and on a temporary demand for human data. When that demand shifts — and it will shift — workers will be left without protections, without transferable skills, and without a safety net. The only winner that emerges, in his view, is the global-north platforms that capture all of the lasting value.

The conversation around regulating this market is advancing in some regions, but still in a fairly fragmented way. The European Union has been the most active in this space, with the AI Act establishing some guidelines on how personal data can be used in training artificial intelligence models. In the United States, the debate is more decentralized, with some states pushing forward specific legislation while the federal government still seeks consensus. In Brazil, the LGPD provides an important legal foundation, but its practical application in the context of the AI data market is still a work in progress, with plenty of gray areas that platforms know how to exploit very well.

The result is that, in practice, anyone selling their data today rarely has clear guarantees about how it will be used, for how long, or with whom it will be shared.

Why this conversation matters right now

What makes this discussion even more relevant is that it is not happening in some distant future. It is happening right now, with real people making real choices about their identity and their data every single day. The AI data market already moves billions of dollars globally, and demand is only set to grow as models become more sophisticated and need larger and more diverse volumes of human data to keep evolving.

The new AI training gig economy presents a fairly clear Faustian bargain: in exchange for a few dollars, its trainers are feeding an industry that may eventually render their own skills obsolete, while leaving them vulnerable to a future of deepfakes, identity theft, and digital exploitation that they are only beginning to comprehend.

Understanding how this market works, who benefits from it, and what risks are involved is not a technical question reserved for experts. It is a question that concerns anyone who has a voice, a face, and a story — in other words, everyone. 🌐

Picture of Rafael

Rafael

Operations

I transform internal processes into delivery machines — ensuring that every Viral Method client receives premium service and real results.

Fill out the form and our team will contact you within 24 hours.

Related publications

Performance and Growth: Nvidia, AI Agents, and Data Centers

Nvidia accelerates revenue with data centers, GB300 NVL72, and Rubin; efficiency and AI Agents demand drive record growth and profit.

AI and Copyright: Supreme Court Denies Copyright Protection for Artistic Creation

Supreme Court rejected the AI-generated art case; in the US only humans can hold authorship — a direct impact on

AI Reveals the Identity of Anonymous Social Media Users

Vulnerable anonymity: how modern AI unmasks social media profiles and why this threatens your online privacy.

Receba o melhor conteúdo de inovação em seu e-mail

Todas as notícias, dicas, tendências e recursos que você procura entregues na sua caixa de entrada.

Ao assinar a newsletter, você concorda em receber comunicações da Método Viral. A gente se compromete a sempre proteger e respeitar sua privacidade.

Rafael

Online

Atendimento

Website Pricing Calculator

Find out how much the ideal website for your business costs

Website Pages

How many pages do you need?

Drag to select from 1 to 20 pages

In just 2 minutes, automatically find out how much a custom website for your business costs

More than 0+ companies have already calculated their quote

Fale com um consultor

Preencha o formulário e nossa equipe entrará em contato.