The Last 24
✓ Verified publishers only
Tuesday, 23 June 2026

✓ Verified source: The Times of India · 16:07 IST · Tuesday, 23 June 2026

'AI eating itself': Trainers use chatbots to generate training data, experts warn of model collapse risk

'AI eating itself': Trainers use chatbots to generate training data, experts warn of model collapse risk
Photo: geralt via Pexels (free license)

Key facts

  • AI model trainers admit using chatbots to generate training data
  • Experts warn the practice risks 'model collapse' over time
  • AI-generated training data can amplify errors across model generations

People hired to train the next generation of AI models are openly admitting they are simply using existing chatbots to do the job — feeding AI-generated text back into AI training pipelines rather than producing the carefully crafted human-generated data the process is supposed to rely on.

Experts have reacted with alarm. The practice, which has been described as a form of 'AI eating itself', creates a dangerous feedback loop. When models are trained on AI-generated data that was itself produced by models with errors, biases, and hallucinations, those flaws can amplify across successive training generations. Researchers refer to the worst-case outcome as 'model collapse' — a degradation of the model's ability to represent the real world accurately.

The problem is partly economic. Generating high-quality human training data is expensive, time-consuming, and requires genuine expertise. Chatbots can produce plausible-looking text at a fraction of the cost and time. For AI companies under pressure to ship models quickly and cheaply, the temptation to use AI-generated data is significant — and apparently, already widespread.

For India, where a booming AI services industry and a wave of AI startups are building and fine-tuning models for local languages, legal documents, medical records, and financial services, the warning has direct operational significance. Models trained on contaminated or AI-generated data could produce unreliable outputs in high-stakes domains where errors carry real consequences.

The revelations are likely to intensify calls for mandatory disclosure of training data provenance and for regulatory standards around what counts as acceptable data for AI model development.

Read the full story at The Times of India →

This is a summary brief. Original reporting and all facts: The Times of India.

Ad space