AI has a data problem. Not just a copyright problem. Not just a bias problem. Something messier. A lot of what AI “knows” about people comes from images, captions, websites, and datasets that communities never approved, never reviewed, and sometimes never even saw. Then those same systems generate images that can flatten real lives into stereotypes.
Microsoft is now testing a different approach. The company’s research team has introduced the Community Library Creator, a platform designed to let communities help define how they appear in AI-generated imagery. The idea is simple enough: instead of letting scraped internet data decide representation, communities build their own visual libraries and explain what those images actually mean. Microsoft published the project on July 22, 2026.
Community Library Creator Gives Groups a Say in AI Representation
Microsoft says the Community Library Creator helps groups with shared experiences build collections of images or videos that reflect how they want to be represented. Each image is paired with descriptions, so AI systems can learn not only what something looks like, but why it matters from the community’s point of view.
That second part matters.
An image of someone at work, at home, with family, or in public can carry context that a generic AI model may miss. Without that context, AI tools often lean on whatever patterns already exist online. Those patterns are sometimes incomplete. In other cases, they are weird. At worst, they can be harmful.
Microsoft researcher Anja Thieme put it bluntly: there is no single ground truth for how people should be represented. Communities have to define and negotiate that themselves.
Why AI Image Data Needs a Different Model
AI image models are trained on huge collections of pictures and text. That sounds powerful, but it also means the models can repeat whatever gaps, distortions, and lazy assumptions exist in the data.
Microsoft gives a sharp example. Online images of mystical dwarves can influence AI systems to generate people with dwarfism using fantasy-like features, such as pointy ears. People with limb differences may also appear mainly in medical or athletic contexts because that is what the available data overrepresents. Real life is wider than that.
People with disabilities go to school. Many work, drink coffee with friends, and raise families. Meetings are part of their lives too. Boredom happens like it does for anyone else. They live ordinary lives that AI often fails to show because the training data does not carry enough ordinary context. That is where community libraries start to feel important.
The Data Comes From the Community, Not Just the Open Internet
The Community Library Creator gives advocacy groups a structured process for collecting and annotating images. Community members choose meaningful images, explain why they matter, then organize larger collections around themes such as family life, health, work, care, or everyday routines.
Microsoft says each library aims to gather about 400 real-world images from community members. That number is not huge compared with the enormous datasets used in AI training, but that is partly the point. This is more about quality, consent, and context than scraping everything in sight.
The images can then support training and evaluation. Prompts generated from the library are used to create AI images, and community members review whether those outputs match the representation they want. Over time, that feedback helps define what “good” representation looks like from the community’s own perspective.
Ownership Is the Real Shift
The most interesting part is not only the data. It is the control. Microsoft says the community, through the advocacy organization that built the library, owns the library and decides how the data is shared. That could include making it available to researchers or developers through a platform such as Hugging Face, or limiting its use.
That is very different from the usual AI data pipeline. Much of AI today still depends on massive online scraping, unclear provenance, and limited visibility for the people represented in the data. Microsoft’s model gives communities a way to say yes, no, or not like that.
Consent also remains part of the process. If someone later wants their data removed, the community can make that change. Microsoft says this kind of control is not typical in AI today and required legal, engineering, and internal process work to make possible.
This Is Still Early, Not a Finished Fix
This should not be mistaken for a full answer to AI bias. Microsoft is not saying every model will suddenly represent everyone correctly because a few libraries exist.
For now, the Community Library Creator is being used in a controlled way with selected advocacy organizations. Microsoft researchers are still working on the engineering, safety, and legal systems needed to expand the project to more communities. So yes, it is early.
Still, the direction is worth watching. AI companies often talk about responsible AI in broad language. This project is more concrete. It asks who gets to build the data, who owns it, who reviews it, and who decides whether the output feels accurate. Those are uncomfortable questions. Good.
What This Means for the Future of AI Data
The Community Library Creator points to a possible future where AI datasets are not only extracted from the web, but built with communities directly involved. That would change the power balance a little.
Instead of engineers guessing what fair representation means, communities could help define it. Models would also learn from images with lived context attached, not only from noisy internet patterns. Most importantly, data would no longer be treated only as something companies collect and own. It could become something communities govern. Will this scale easily? Probably not.
Community-led data takes time. Trust has to come first. Consent matters just as much. Responsible organizations also need to manage the libraries properly. Compared with scraping the web, this approach moves more slowly. But maybe that is exactly why it matters. Fast data built AI into what it is today. Better data may decide whether people actually trust what comes next.

