AI Models Have Shown Bias Toward State Media

University of Oregon researchers found that models trained on widespread web data mirror the influence of state-controlled outlets.

Updated on Sept. 20, 2026 in Artificial Intelligence

Bold flat-color editorial illustration of a stone pillar being weighed down by sediment, symbolizing structural bias in AI training datasets.
University of Oregon researchers reported that large-scale AI models ingest disproportionate amounts of state-controlled media, reinforcing geopolitical biases in model outputs. AI Illustration. Upload story photo >

Live Poll

Do you trust artificial intelligence models to provide neutral and unbiased political information?

A study published on September 20, 2026, by University of Oregon researchers revealed that AI models often exhibit favorable bias toward government entities when prompted in the native language of countries with high media control. These findings highlight how AI behavior is shaped by the prevalence of state-controlled media within open training datasets.

Why it matters

As AI models are trained on massive, unfiltered web scrapes, they unintentionally ingest information environments dominated by powerful institutions. This research illustrates how digital training sets can reinforce the specific geopolitical narratives present in those source materials.

Researchers found that state-controlled media accounts for 1.64% of the analyzed dataset, with 75.3% of AI responses regarding Chinese leaders and institutions appearing favorable when prompted in Chinese. The study covered 37 countries in total.

The players

University of Oregon

A public research university recognized for its interdisciplinary studies and data science initiatives.

OpenAI

An AI research organization and developer of large language models that utilize extensive web-scale training data.

The details

The team analyzed how models interact with the Common Crawl—a massive, open-source repository of web data—by training small models and comparing output to inquiries in both English and local languages. They found 3.1 million documents overlapping with Chinese state-controlled media, which were significantly more prevalent than balanced encyclopedic sources. These training inputs appear to anchor the model’s linguistic associations, causing them to reflect the political tone of the source material when triggered by language-specific cues.

Timeline

  1. September 20, 2026: The University of Oregon research findings were published.

The Tech Race

This research highlights a significant vulnerability in the data pipeline that currently fuels the global AI race. It follows a growing pattern of scrutiny regarding how data composition impacts the neutrality and safety of models used by international audiences.

Users querying AI models in non-English languages, particularly in nations with restricted media, may receive outputs that mirror local state narratives. Those relying on these models for objective analysis should be aware that model responses often reflect the weight of the data on which they were trained.

The takeaway

The study confirms that training data composition is a primary driver of geopolitical bias in language models. Observers should track whether major AI developers begin to publish dataset transparency reports regarding the inclusion of state-sponsored content.

Further reading

For additional context on how foundational models are trained, read more in our Artificial Intelligence section.

Source note: This article includes information reported by KOIN 6 Portland.

Live Poll

Do you trust artificial intelligence models to provide neutral and unbiased political information?