JournalAI Ethics

Field guide / 6

Hidden prompt revision raised WORLDVIEW cultural flattening from 0.75 to 0.86 on SDXL

WORLDVIEW traced part of image generators' cultural stereotyping to an unseen prompt rewrite. Teams need to audit the deployed pipeline, not only the image model.

Sep 12, 20266By ISH Team
Hidden prompt revision raised WORLDVIEW cultural flattening from 0.75 to 0.86 on SDXL
Advertisement

Hidden prompt revision raised WORLDVIEW cultural flattening from 0.75 to 0.86 on SDXL

Ask an image generator for "a living room in Finland" and snow, pine or birch may appear inside a scene that never requested them. The image model may not have chosen those details. A hidden language model may have expanded the prompt first.

WORLDVIEW, a new audit of commercial image systems, separates those stages. The researchers collected rewritten prompts from DALL-E-3, Imagen 4 and GPT-Image-1.5, then measured what the revision layer inserted. They also sent original and rewritten prompts to SDXL Lightning and Flux 2 Dev, two image models without a built-in rewriter.

Mean cultural flattening rose from 0.75 to 0.86 on SDXL and from 0.85 to 0.98 on Flux when the open models received rewritten prompts. Revision did not explain every stereotype. It did add a measurable second source.

Ordinary scenes reveal repeated shorthand

WORLDVIEW starts with 280 English prompts across 14 domains, including family, transport, work, politics, religion, housing, advertising and immigration. Each baseline omits geography. The researchers translated context-specific versions and added a place, producing 8,960 prompts across 15 languages and 31 language-context pairings.

The first test measures how far each rewritten prompt moves from its unmarked English baseline. The US was closest to that default across all three systems, with a Contextual Markedness Score of 0.21 to 0.31. The UK followed at 0.24 to 0.35. Finland in Finnish reached 0.47 on DALL-E-3, while Middle Eastern and North African contexts also appeared near the heavily marked end. Cross-model ordering was moderate, with Kendall's tau from 0.57 to 0.67. The three systems shared a broad pattern but did not agree on every context.

The Cultural Flattening Score asks whether the rewriter applies a small set of distinctive terms across unrelated scenes. Cultural specificity is expected when the prompt names a place. Reusing the same cues for a family, an advert and a political event is the warning sign.

Swiss and Nordic contexts ranked highest. All three Swiss language variants scored above 0.48. Their top ten distinctive terms appeared in 77% to 96% of prompts, though those terms made up only 15% to 22% of inserted vocabulary. Finland repeatedly received snow, pine, northern and birch. Switzerland received Alps, chalet, chocolate and fondue. Egypt received pyramid, sphinx, pharaonic and hieroglyph.

A chalet may fit a tourism poster. It makes little sense in a Swiss queue for government services. Flattening is the failure to notice that difference.

A controlled test isolates the rewriter

A correlation between rewritten text and final images cannot identify the responsible stage. For a causal test, the researchers took GPT-Image's original and revised prompts for the US, UK, Australia and India. They rendered both conditions with SDXL Lightning and Flux 2 Dev, holding generation settings constant.

Revised-prompt images had significantly higher markedness on Flux in all four contexts and on SDXL in three. Cultural flattening rose on both models. The UK had the largest increase, 34% on SDXL and 52% on Flux.

The added visual vocabulary was recognizable. Revised prompts introduced outback and kangaroo for Australia, cobblestone and pub for the UK, and marigold and taj for India. India also exposed the boundary of the finding: original prompts already produced bindi, dhoti and curry. The image model carried cultural bias before revision. For the US, the study found no culturally distinctive terms without the revision layer in this ablation.

Removing a rewriter would not make an image model neutral. Leaving one unaudited can add another layer of stereotyping.

A shipped system is more than its image model

Google's Imagen documentation says its rewriter adds detail and is enabled by default for Imagen 4 models. API users can set enhancePrompt to false. The same page warns that enhancement can produce unwanted results with complex prompts.

Image products may translate, rewrite, filter and rank before presenting an output. Each stage can change the request. Controlled framing tests on language models found that keeping the facts did not guarantee keeping the framing. Testing only the last model misses decisions already embedded upstream.

Content Credentials document an image's history rather than its truth. They can record which tool produced a file. They cannot determine whether an unseen rewriter turned a contemporary culture into tourism shorthand.

Run a paired audit before launch

A useful first pass does not need all 8,960 prompts.

  1. Save the exact user prompt, revised prompt when available, model identifiers, settings, refusals and final image.
  2. Build matched prompts across everyday domains. Change only the cultural context or language.
  3. Generate with revision enabled and disabled. Keep the image model, seed and settings fixed where the API permits.
  4. Measure semantic distance and recurrence. A relevant detail in one scene differs from the same detail appearing throughout the set.
  5. Ask people from represented communities to review repeated terms and images. A distance metric cannot decide what appropriate representation looks like.

If a provider hides the revised prompt, record that gap. An enhancement toggle can still support paired tests. Cross-provider comparisons may reveal changes too, though they cannot attribute the cause as cleanly.

The habit also helps with one-off work in ISH's image generation workflow. Preserve the original wording, compare deliberate variants and inspect what the system added. Cultural detail is not the problem. Detail pulled from an invisible autocomplete's favorite shorthand is.

What WORLDVIEW does not establish

The commercial audit covered three systems from two providers. Its controlled ablation used one rewriter and four English-speaking contexts, plus a Swiss case study. That is not proof of the same causal effect for every language or provider.

Qwen2.5-VL-7B-Instruct described the generated images for lexical analysis and may carry biases of its own. The authors manually checked 150 image-description pairs. They found no case where it invented culturally relevant stereotypical content absent from the image, but they describe the check as informal. SDXL and Flux also have training biases, as the India result makes plain.

WORLDVIEW does not define the correct depiction of any culture. It establishes a narrower result: in the tested pipeline, prompt revision made a separate causal contribution to cultural flattening. The released benchmark data and analysis code include the revised prompts, so another team can inspect the intermediate text instead of inferring it from the final picture.

#WORLDVIEW#text-to-image#prompt revision#cultural bias#AI auditing#image generation
Advertisement

Keep reading

Related stories

Browse the archive