UPSC Darpan

Science & TechnologyGS329 September 2026

AI Image Models Reproduce Caste Bias as India’s AI Rules Lean on Voluntary Codes and Self-Certification

Open in the app — quiz, notes, Mistake Vault हिंदी में पढ़ें

The news

New Delhi. Writing in The Hindu, workers’ rights advocate Rejimon Kuttappan argues that India’s craze for AI-generated ‘1980s’ images carries a hidden caste bias, and that the government’s light-touch approach to AI regulation lets it pass unexamined. Generative image models are software that produce pictures from a text prompt after learning patterns from millions of existing images. Retro-Bollywood portraits became the signature trend, and India became the No. 1 country for Google’s Nano Banana image-generation model. The author asks where this aesthetic comes from. The source material, he says, is film stills, magazine spreads, studio portraits and family albums, and in the India of the 1980s all four belonged to households with money and standing, overwhelmingly savarna (upper-caste), urban or landed. He cites the Planning Commission’s estimate that in 1983, 44.5% of Indians, some 323 million people, lived below the poverty line. Dalit and Adivasi lives were photographed, but mostly by others: the state for welfare files, activists after atrocities, and anthropologists. He opens with the Karamchedu massacre of July 17, 1985 in coastal Andhra Pradesh, where a dispute at a public drinking-water tank ended with six Madiga (Dalit) men killed, and which gave birth to the Andhra Pradesh Dalit Mahasabha. A model trained on such an archive, he writes, “reproduces the silence, and hardens it”. Tests by MIT Technology Review found that GPT-5 chose the stereotypical answer in 80 of 105 sentences. A study at the ACM FAccT conference this year analysed 1,536 images from Gemini’s image model, the engine behind Nano Banana, prompted with Indian names and no caste labels; caste surfaced anyway, through food, neighbourhood, work and worship. On policy, he notes that the Ministry of Electronics and Information Technology’s (MeitY) AI Governance Guidelines name bias and discrimination as risks but lean on voluntary codes and self-certification, and that the Centre has told the Rajya Sabha that no new “horizontal” AI law, meaning one law for all uses of AI, is needed at this stage. The ₹10,371-crore IndiaAI Mission subsidises “sovereign” models, and Union Minister Ashwini Vaishnaw has promised Indian-trained models free of “the bias of other models”. A committee is reportedly drafting firmer rules. He asks MeitY and the IndiaAI Safety Institute to say publicly what is in the training data, who labelled it, whether a caste evaluation was done, and when it will be published. In The Indian Express, former ambassador to China Ashok K. Kantha argues that India’s weight in the US-China AI contest will rest on compute, chips, models, datasets, talent and the ability to test frontier systems independently. The Economic Times reports that OpenAI lists India among countries whose national AI safety institutes could contribute to global standards. The syllabus link is GS3 on IT and emerging technology, and GS1 on caste.

The chain in one line: Photography in 1980s India is largely a privilege of savarna, urban and landed households → the surviving visual archive shows Dalit and Adivasi lives mostly through welfare files and atrocity records → generative models trained on that archive learn the privileged view as ‘normal’ and link Dalit identity with dirt and menial work → a viral retro-image trend spreads the stereotype at mass scale → India’s voluntary, self-certification approach to AI governance leaves no mandatory test for caste bias

Static syllabus linkage

  1. Generative models learn statistical patterns, so they inherit the skew of their training data. A generative AI model is trained on a very large collection of examples, such as images with captions, and learns which features tend to occur together; when prompted, it produces the most probable output given what it has seen. It has no understanding of fairness, so if a category of people appears rarely, or only in certain roles, the model reproduces that pattern as the default. Researchers distinguish historical bias (the data faithfully records past discrimination), representation bias (some groups are under-sampled) and label bias (the people who tag the data bring their own assumptions). Algorithmic bias is therefore not a coding error that can be patched once; it has to be measured through audits and evaluations designed for the specific society in which the model is used.
  2. Articles 15 and 17 and the Atrocities Act make caste discrimination a constitutional wrong, not a matter of taste. Article 15(1) bars the State from discriminating on grounds only of religion, race, caste, sex or place of birth. Article 15(2) goes further and forbids any disability or restriction on these grounds in access to shops, public restaurants, hotels, places of public entertainment, and the use of wells, tanks, bathing ghats and roads maintained by the State or dedicated to public use, which is why the Karamchedu water-tank dispute is a constitutional story. Article 17 abolishes untouchability and forbids its practice in any form, and is enforced through the Protection of Civil Rights Act, 1955. The Scheduled Castes and the Scheduled Tribes (Prevention of Atrocities) Act, 1989 created specific offences and special courts for atrocities against SCs and STs.
  3. India funds AI capacity through the IndiaAI Mission and has chosen guidelines over a new statute. The Union Cabinet approved the IndiaAI Mission in March 2024 with an outlay of about ₹10,371.92 crore over five years, implemented by MeitY. Its pillars include subsidised access to computing capacity (GPUs), development of indigenous foundation models, a platform for non-personal datasets, support for AI start-ups and applications, skilling, and a ‘Safe and Trusted AI’ component. Under that component, MeitY announced an IndiaAI Safety Institute in January 2025 to work on standards, testing and risk assessment. The India AI Governance Guidelines released by MeitY in November 2025 took the position that existing laws such as the Information Technology Act, 2000 and the Digital Personal Data Protection Act, 2023 can address most AI harms for now, backed by voluntary commitments and institutional coordination rather than a separate AI law.
  4. UNESCO’s 2021 Recommendation and the EU AI Act are the two global reference points on AI fairness. The UNESCO Recommendation on the Ethics of Artificial Intelligence, adopted by its General Conference in November 2021, was the first global standard-setting instrument on AI ethics; it names fairness and non-discrimination, transparency and human oversight among its core principles and asks states to carry out ethical impact assessments. It is not binding. The European Union’s Artificial Intelligence Act, which entered into force in 2024, is binding and risk-based: it bans a few ‘unacceptable-risk’ uses, imposes strict duties on ‘high-risk’ systems, including data governance to examine training data for possible biases, and requires transparency such as labelling of AI-generated content. India’s approach sits between the two: stated principles, but mostly voluntary compliance.

Why UPSC loves this

  1. GS3 asks about IT, computers and the social effects of new technology. The GS3 syllabus covers awareness in the fields of IT and computers and issues relating to intellectual property, and UPSC has repeatedly asked about the promise and risks of artificial intelligence in governance, health and security. The expected answer has moved from listing uses to weighing risks such as bias, privacy and accountability. This story gives a concrete Indian example of bias, which is exactly what most answers lack.
  2. GS1 on caste and GS4 on fairness make this a cross-paper case study. GS1 asks about the salient features of Indian society, social empowerment and the persistence of caste in new forms. GS4 case studies increasingly involve technology and fairness. A candidate who can explain how an old social exclusion re-enters society through an algorithm can use the same example in an essay on technology and equality.
  3. Prelims tests India’s AI institutions and constitutional anchors. Questions on government missions and their ministries, and on Articles 15 and 17, are standard fare. The IndiaAI Mission, the IndiaAI Safety Institute and the UNESCO Recommendation are the natural Prelims targets from this story, alongside the difference between binding law and guidelines.

Prelims nuggets

  • Article 15(2) of the Constitution prohibits any disability or restriction on grounds only of religion, race, caste, sex or place of birth in access to shops, public restaurants and places of public entertainment, and in the use of wells, tanks, bathing ghats and roads maintained by the State or dedicated to public use.
  • Article 17 abolishes untouchability and forbids its practice in any form; the Protection of Civil Rights Act, 1955 is the law that enforces it.
  • The Scheduled Castes and the Scheduled Tribes (Prevention of Atrocities) Act was enacted in 1989.
  • The IndiaAI Mission, approved by the Union Cabinet in March 2024, is implemented by the Ministry of Electronics and Information Technology and includes a ‘Safe and Trusted AI’ pillar.
  • The UNESCO Recommendation on the Ethics of Artificial Intelligence, adopted in November 2021, is a non-binding global standard that includes fairness and non-discrimination among its principles.
  • The European Union’s Artificial Intelligence Act follows a risk-based approach, prohibiting certain AI practices and imposing stricter obligations on high-risk systems.
  • Algorithmic bias refers to systematic and repeatable errors in an automated system that produce unfair outcomes for particular groups, often because of skewed or unrepresentative training data.

Analysis

  1. This is bias by omission, and ordinary safety filters are not built to catch it. Most AI safety work in content moderation looks for toxic output: slurs, violence, explicit material. The problem the author describes is quieter. When a prompt for a 1980s family returns only upper-caste settings, nothing offensive has been produced, yet a whole population has been left out of the picture of the past. Filters cannot flag an absence. Catching it needs representational audits, in which a model is tested with prompts across names, regions and occupations and its outputs are compared systematically, as the FAccT study did. That is a measurement task a public institution such as the Safety Institute is well placed to standardise.
  2. ‘Sovereign’ models will not be bias-free by default; they may be more caste-aware in the wrong way. The Minister’s promise that Indian-trained models will not carry the bias of other models assumes the problem is foreign data. But an Indian model trained on Indian films, matrimonial advertisements, news archives and social media may learn caste signals more precisely, because those sources are full of them. Local training is an opportunity only if the data is deliberately curated to include under-represented languages and communities, and if results are tested. Without a published evaluation, the claim cannot be checked, which is the author’s point when he asks which bias and tested how.
  3. Disclosure mandates are a middle path between a heavy AI law and pure self-certification. The government’s case for guidelines is reasonable: a sweeping horizontal law could burden start-ups, lag behind the technology and duplicate the IT Act and data protection law. But self-certification means the firm marks its own homework. The author’s four questions point to a lighter tool: mandatory public disclosure of what data a model was trained on, who labelled it and what fairness tests were run, for models used at scale or by the government. Disclosure does not dictate design, yet it lets researchers, courts and users hold firms to account. The counter-view is that training data is commercially sensitive, which can be handled by disclosure to a regulator rather than to the public.
  4. The invisible labeller is part of the problem and part of the solution. Data labelling, often done by workers paid per task, decides what a model considers ‘Indian enough’ or ‘normal’. If labellers have never seen a basti or a colony, they cannot recognise when a model misrepresents it. This connects AI fairness to labour policy: the composition, training and pay of the labelling workforce shape output quality. Requiring diverse evaluation panels, including Dalit and Adivasi reviewers, for models deployed in public services would be a practical step. It also creates skilled work in regions and communities that the AI economy has so far bypassed.
  5. Kantha’s case for building at home is also a case for controlling whose values are built in. Kantha argues from geopolitics: India should avoid dependence on either the American or the Chinese AI stack and build compute, models, datasets and independent testing capacity. The caste-bias debate supplies a social reason for the same conclusion. A country that only imports models also imports the evaluation choices of their makers, who have no reason to test for caste. Independent testing capacity is therefore not only a security asset but a guarantee that Indian constitutional values, including Articles 15 and 17, can be checked against the machines Indians use.

Possible Mains question

“Artificial intelligence does not invent social prejudice; it inherits it and scales it.” Examine this statement with reference to caste bias in generative AI models in India. Suggest a regulatory approach that protects equality without stifling innovation. (15 marks, 250 words)

Model approach

  1. Introduction. Open with the retro-1980s AI image trend and the finding, cited in The Hindu, that caste surfaced in 1,536 images from Gemini’s image model even without caste labels, while GPT-5 chose stereotypes in 80 of 105 sentences in MIT Technology Review tests.
  2. Body — how bias is inherited. Explain training data in plain terms and the three sources of bias: a historical archive dominated by privileged households, under-representation of Dalit and Adivasi lives, and labellers’ assumptions. Use the Karamchedu example and Article 15(2) to show why this touches constitutional equality.
  3. Body — current Indian approach. Describe the IndiaAI Mission (₹10,371 crore), the IndiaAI Safety Institute and MeitY’s guidelines relying on voluntary codes and self-certification. Give the government’s rationale: avoiding premature regulation and using existing laws.
  4. Body — a balanced regulatory design. Propose mandatory disclosure of training data sources and fairness evaluations for large or government-used models, India-specific bias benchmarks set by the Safety Institute, diverse labelling and evaluation panels, and funding for community archives as data. Contrast with the EU’s risk-based law and UNESCO’s non-binding principles.
  5. Conclusion. Conclude that India does not need a heavy AI statute to address caste bias, but it does need binding transparency and public testing, because a promise of fairness that cannot be checked is not a safeguard.

Administrator's brainstorm

You head a State’s e-governance department, which plans to use a generative AI tool to create images and text for welfare campaigns. What checks would you put in place?

Before deployment I would ask the vendor to disclose the model’s training data sources and any fairness evaluations it has run. I would have a diverse internal panel, including officers from the SC/ST welfare department, test the tool with prompts about different communities and occupations. All public-facing output would pass human review before release. The procurement contract would carry a clause allowing termination if the tool is found to produce discriminatory content.

As an officer at the IndiaAI Safety Institute, how would you design a caste-bias evaluation?

I would build a benchmark of prompts using names, surnames, regions, occupations and everyday scenes, and measure how outputs differ across groups, following the method of the published academic studies. Community organisations and social scientists would help design and validate it. The results for models used by government would be published. Repeating the test on every major update matters, because models change quickly.

An interview board asks: is it fair to blame a machine for a society’s prejudice?

The machine is not morally responsible, but the people who build and deploy it are, because they choose the data and decide whether to test it. A biased human decision affects one case; a biased model repeats the same error millions of times. So the standard for deployment should be higher, not lower. The remedy is to treat bias testing as a routine part of building AI, not as an afterthought.