Anthropic is bringing outside AI evaluators closer to the development of its most advanced models. The Claude maker has partnered with Accenture to create an embedded evaluation team that will work alongside Anthropic’s researchers and engineers.
The initiative is designed to test frontier models, examine safeguards and identify risks while AI systems are still being developed. Anthropic and Accenture each expect to invest at least $1 billion over five years in the effort, bringing their combined planned investment to at least $2 billion.
Accenture’s Faculty Will Work Inside Anthropic
Accenture will carry out the work through Faculty, its specialist AI business. The team will work closely with Anthropic while maintaining an independent evaluation role. Its work will include red teaming, alignment assessments and tests of safeguards built around Claude and future Anthropic models.
This arrangement gives evaluators more visibility than a typical outside model assessment. Rather than receiving a finished AI system and testing it afterward, the team can examine risks closer to the development process. That could help uncover issues that standard benchmarks or external testing may miss.
What Embedded AI Evaluation Actually Means
Embedded evaluation places outside specialists much closer to the teams building frontier AI. Anthropic says evaluators could receive access similar to employees. That may include observing models during training, speaking with staff and reviewing decisions related to development and deployment.
The idea is to give evaluators enough access to understand how an AI system is being built, not simply how it performs after release. They could also examine whether internal safeguards and safety commitments are working as intended.
Anthropic Remains Responsible for Its AI Models
The presence of an outside evaluator does not shift responsibility away from Anthropic. The company says it remains responsible for the safety and behavior of the models it develops and releases.
Instead, embedded evaluation is intended to add another layer of scrutiny. Independent specialists can challenge assumptions, test safeguards and identify weaknesses. Anthropic still decides how those findings affect training, deployment and future model releases.
Independence Remains a Difficult Question
Embedded evaluation creates an obvious challenge. Anthropic is helping fund the organization evaluating its own systems.
The company acknowledges that the model is still developing. There are no widely accepted standards governing how embedded evaluators should operate, what information they should receive or how findings should be published.
Anthropic has suggested that future evaluation programs could use pooled funding or government-backed mechanisms. Such approaches could reduce the financial dependence of evaluators on individual AI companies.
Anthropic Is Looking Beyond Accenture
The Accenture agreement is not exclusive. Anthropic says it is also discussing embedded evaluation with METR and other nonprofit organizations.
Those groups could test different versions of the approach using their own funding. Anthropic expects to announce additional evaluation partners as the program develops.
Working with several organizations could also prevent one evaluator from becoming the sole outside voice assessing Anthropic’s models. It may eventually provide different perspectives on safety, alignment and emerging capabilities.
AI Evaluation Is Moving Earlier in Development
Most AI evaluation becomes visible when a model is close to launch or already available. Embedded evaluation changes that timeline.
Outside specialists could examine systems while developers are still training and testing them. That creates an opportunity to identify problems before they become part of a widely deployed product.
The approach may become more important as AI models gain stronger reasoning and agentic capabilities. Systems that can use tools, write code or take actions create different risks from traditional chatbots. Testing them only after development may provide an incomplete picture.
Anthropic and Accenture Already Have a Bigger AI Partnership
The two companies were already working together before the embedded evaluation initiative. In December 2025, they announced a multi-year partnership focused on enterprise AI adoption.
That agreement created the Accenture Anthropic Business Group. Accenture also announced plans to train about 30,000 professionals to help organizations deploy Claude and related AI technologies.
The embedded evaluation project adds a different dimension to that relationship. Accenture will help organizations use Anthropic technology while Faculty examines the safety of Anthropic’s frontier systems.
Embedded Evaluators Could Change Frontier AI Oversight
The experiment is still at an early stage. Questions remain around independence, funding, access and how much information evaluators can eventually make public.
Still, the size of the planned investment makes this more than a small research project. Anthropic and Accenture each expect to commit at least $1 billion over five years.
If embedded evaluation proves useful, other frontier AI developers could explore similar arrangements. Independent testing may then start much earlier in the development cycle rather than appearing mainly before or after a major model release.
That would change the role of outside AI evaluators. Instead of simply testing the finished product, they could watch parts of the development process unfold.
Sources
Anthropic — Partnering with Accenture on Embedded Evaluation
https://www.anthropic.com/news/accenture-embedded-evaluation
Accenture — Accenture and Anthropic Partner to Build Team of Embedded Evaluators at Anthropic
https://newsroom.accenture.com/news/2026/accenture-and-anthropic-partner-to-build-team-of-embedded-evaluators-at-anthropic
Anthropic — Accenture and Anthropic Launch Multi-Year Partnership
https://www.anthropic.com/news/anthropic-accenture-partnership

