Close Menu
  • Home
  • Events
    • Machine Can Think Summit 2026
    • Step Dubai Conference 2026
  • Technology & Innovation
  • Industry Applications
  • Business & Marketing
  • Trends & Insights
  • World AI Awards 2026
    • Aerospace, Aviation & Space
    • Agriculture & Food Systems
    • AI Infrastructure, Compute & Cloud
    • AI Leadership
    • Climate, Environment & Sustainability
    • Construction, Real Estate & PropTech
    • Consumer, Lifestyle & Accessibility
    • Core AI, Models & Algorithms
    • Customer Experience & Service
    • Cybersecurity & Digital Trust
    • Data, Analytics & Intelligence
    • Defense, Security & Public Safety
    • Destinations & Sustainable Tourism
    • Education & Learning
    • Energy, Utilities & Resources
    • Enterprise Productivity & Collaboration
    • Events, Sports & Live Experiences
    • Exhibitions, Conferences & Business Events
    • Financial Services & Insurance
    • Food, Beverage & Dining
    • Government & Public Sector
    • Healthcare & Life Sciences
    • Hotels, Accommodation & Stays
    • HR, Talent & Workforce
    • Legal, Governance, Risk & Compliance
    • Manufacturing, Industrial & Industry 4.0
    • Marketing, Advertising & Brand Experience
    • Media, Entertainment, Gaming & Culture
    • Retail, E-commerce & Consumer Commerce
    • Robotics, Autonomy & Drones
    • Sales, Revenue & Growth
    • Science, Research & Discovery
    • Software, Platforms & IT Operations
    • Supply Chain, Logistics & Procurement
    • Telecommunications, Networks & Connectivity
    • Tours, Activities & Attractions
    • Transport & Mobility
    • Travel Commerce, Booking & Platforms
    • Travel Infrastructure, Risk & Intelligence
    • Wellness, Spa & Retreats
What's Hot
Business & Marketing

Samsung Profit Surges Nearly Eightfold as AI Memory Boom Drives $80 Billion Q3 Forecast

By Art RyanOctober 9, 20260

Samsung Electronics is getting a massive boost from the global AI infrastructure buildout, with the…

Biohub Expands Virtual Biology Initiative With $1.8 Billion AI Push

October 9, 2026

Google Cloud Unveils Gemini Agent as AI Work Race Accelerates

October 9, 2026

Council of Europe and Microsoft Sign AI Cooperation Agreement Focused on Human Rights

October 9, 2026
Facebook X (Twitter) Instagram
Facebook X (Twitter) Instagram
Breaking AI News
Friday, October 9
  • Home
  • Events
    • Machine Can Think Summit 2026
    • Step Dubai Conference 2026
  • Technology & Innovation

    Samsung Profit Surges Nearly Eightfold as AI Memory Boom Drives $80 Billion Q3 Forecast

    October 9, 2026

    Biohub Expands Virtual Biology Initiative With $1.8 Billion AI Push

    October 9, 2026

    Google Cloud Unveils Gemini Agent as AI Work Race Accelerates

    October 9, 2026

    Council of Europe and Microsoft Sign AI Cooperation Agreement Focused on Human Rights

    October 9, 2026

    OrbitronAI and Aramco Digital Target Billions in Industrial AI

    October 9, 2026
  • Industry Applications

    Biohub Expands Virtual Biology Initiative With $1.8 Billion AI Push

    October 9, 2026

    OrbitronAI and Aramco Digital Target Billions in Industrial AI

    October 9, 2026

    SPARK Launches Sharjah Advanced AI Industrial Accelerator

    October 8, 2026

    Dubai Launches Agentic AI Accelerator to Redesign Government Services

    October 7, 2026

    UK Backs New Blueprint for AI-Enabled Medical Devices as Healthcare Regulation Shifts

    October 7, 2026
  • Business & Marketing

    Samsung Profit Surges Nearly Eightfold as AI Memory Boom Drives $80 Billion Q3 Forecast

    October 9, 2026

    Google Cloud Unveils Gemini Agent as AI Work Race Accelerates

    October 9, 2026

    AI Everything Abu Dhabi Puts AI to Work Across Business and Government

    October 8, 2026

    Dubai Launches AED2.5 Million Create AI Agents Championship

    October 8, 2026

    Schneider Electric Makes $23.7 Billion PTC Bet on Industrial Software and AI

    October 6, 2026
  • Trends & Insights

    Google Unveils Gemini 4 Argon With 1 Million-Token Output Capacity

    October 2, 2026

    OpenAI Unveils ‘Always-On’ Dots AI Agents as Altman Pushes a More Personal AI Future

    October 1, 2026

    Trump and AI Leaders Sign Voluntary Safety Accord as Washington Bets on Industry Self-Policing

    October 1, 2026

    Nvidia Launches Record $150 Billion Share Buyback as AI Profits Surge

    September 30, 2026

    OpenAI Targets $30 Billion Funding Round at $1.4 Trillion Valuation, Report Says

    September 30, 2026
  • World AI Awards 2026
    • Aerospace, Aviation & Space
    • Agriculture & Food Systems
    • AI Infrastructure, Compute & Cloud
    • AI Leadership
    • Climate, Environment & Sustainability
    • Construction, Real Estate & PropTech
    • Consumer, Lifestyle & Accessibility
    • Core AI, Models & Algorithms
    • Customer Experience & Service
    • Cybersecurity & Digital Trust
    • Data, Analytics & Intelligence
    • Defense, Security & Public Safety
    • Destinations & Sustainable Tourism
    • Education & Learning
    • Energy, Utilities & Resources
    • Enterprise Productivity & Collaboration
    • Events, Sports & Live Experiences
    • Exhibitions, Conferences & Business Events
    • Financial Services & Insurance
    • Food, Beverage & Dining
    • Government & Public Sector
    • Healthcare & Life Sciences
    • Hotels, Accommodation & Stays
    • HR, Talent & Workforce
    • Legal, Governance, Risk & Compliance
    • Manufacturing, Industrial & Industry 4.0
    • Marketing, Advertising & Brand Experience
    • Media, Entertainment, Gaming & Culture
    • Retail, E-commerce & Consumer Commerce
    • Robotics, Autonomy & Drones
    • Sales, Revenue & Growth
    • Science, Research & Discovery
    • Software, Platforms & IT Operations
    • Supply Chain, Logistics & Procurement
    • Telecommunications, Networks & Connectivity
    • Tours, Activities & Attractions
    • Transport & Mobility
    • Travel Commerce, Booking & Platforms
    • Travel Infrastructure, Risk & Intelligence
    • Wellness, Spa & Retreats
Breaking AI News
Home » Perplexity WANDR Benchmark Shows AI Research Agents Are Still Not Ready for Full Research Work
Technology & Innovation

Perplexity WANDR Benchmark Shows AI Research Agents Are Still Not Ready for Full Research Work

Art RyanBy Art RyanJuly 16, 2026No Comments6 Mins Read
Facebook Twitter Pinterest LinkedIn Tumblr Email
Share
Facebook Twitter LinkedIn Pinterest Email

Perplexity has released a new benchmark called WANDR, and it quietly says something many people using AI research tools already know. These agents can be impressive. They can move fast. They can find sources, summarize pages, compare companies, build lists, and produce something that looks useful.

Then the hard part arrives. Can they do it again? And again? Across dozens of companies, people, filings, dates, sources, and claims, without dropping pieces along the way? That is where Perplexity’s WANDR benchmark starts asking uncomfortable questions.

What Is Perplexity WANDR?

WANDR stands for Wide ANd Deep Research. It is an open benchmark built to test AI research agents on large research tasks, not simple one-answer prompts. Most AI benchmarks feel neat. A model answers a question. The answer is checked. Done.

WANDR is messier because real research is messier. A business analyst may need a full competitor map. A due diligence team may need dozens of companies with executives, ownership details, funding history, and proof for each claim. A recruiter may need a long list of qualified candidates, with evidence attached to every entry.

That is not one answer. That is structured research at scale. Perplexity says WANDR includes 500 realistic research tasks based on high-volume, evidence-heavy knowledge work. The benchmark asks agents to search widely, go deep enough on each record, and support every claim with specific sources.

The Result Is Not Exactly Comforting

The strongest system in Perplexity’s evaluation reached only 0.363 soft F1 and 0.133 hard F1. That sounds technical, but the meaning is simple enough. Even the best-performing system struggled to complete the full job properly. It could make partial progress, sometimes good progress, but full coverage was rare.

That is the part worth paying attention to. AI research agents are often marketed like they can take over multi-step research work. WANDR suggests they are better viewed as assistants that can gather useful leads, not finished-work machines that should be trusted without review. There is a difference. A big one.

Why Wide Research Breaks AI Agents

A single polished answer can hide a lot of gaps. Wide research does not let those gaps stay hidden for long. If an agent needs to find 70 companies and provide proof for each one, the weak spots become obvious. Maybe it finds 40 good records. Maybe it repeats companies. Maybe it uses a source that looks relevant but does not prove the actual claim. Maybe one branch of the research is complete, while another is missing.

This is where WANDR becomes useful. It does not only ask whether an answer sounds good. It checks whether each submitted record is supported by evidence. Perplexity’s own explanation says WANDR grades claims against cited evidence instead of relying on a fixed answer key. That matters because many research questions change over time. A fixed list can go stale quickly.

So WANDR checks the page, the claim, the excerpt, and whether the evidence actually supports what the agent says. That is stricter. Also more realistic.

Perplexity’s Search as Code Leads, But the Win Has Limits

Perplexity’s Search as Code system led the benchmark, while Anthropic came second in the reported results. Other systems scored much lower. Still, this is not a victory lap moment.

Perplexity’s own numbers show that nobody is close to solving the problem. The leading system still had a low hard F1 score, which means complete, fully supported research paths remain difficult for current AI agents. That is probably the most honest takeaway here. Yes, Perplexity performed best in its own benchmark. But the larger message is not “Perplexity wins.” It is more like: even the best current systems still miss too much when the task becomes broad, layered, and evidence-heavy.

And yes, because Perplexity built and released the benchmark, readers should treat the leaderboard with some caution. Vendor benchmarks can still be useful, but they should not be treated like neutral ground without checking the methodology.

The Real Problem Is Not Finding a Page

One interesting detail from the WANDR results is that finding a page is not always the hardest part. The bigger issue is proving the claim. An AI agent may find a source that looks close enough. But does the source clearly support the exact requirement? Does the excerpt include the right evidence? Does the page prove the date, role, company, location, or status being claimed?

That is where many systems lose accuracy. This is also where human researchers still matter. People do not just collect links. They judge whether a source actually proves something. AI agents are getting better at the first part. The second part is still shaky.

Why This Matters for AI Research Tools

WANDR lands at a time when AI companies are pushing agents into workplace research, market intelligence, finance, legal support, recruiting, and business operations. The pitch is attractive. Give the agent a task. Let it search. Let it organize the results. Let it return something useful while you do other work.

But WANDR shows the current ceiling more clearly. Agents can help speed up research, especially early-stage discovery. They can surface options, build starting lists, and reduce blank-page work. What they should not do, at least not yet, is replace final verification.

For businesses using AI research agents, the safe workflow is probably not “agent does the research.” It is closer to “agent creates the first draft of the research trail, then a human checks the important parts.” Less glamorous. More realistic.

WANDR Could Help Build Better Agents

The useful part of WANDR is not only the score. It is the diagnosis. Perplexity says the benchmark can show where an agent fails: discovery, enrichment, identity handling, page qualification, or evidence extraction. That gives developers something more specific to improve.

Instead of saying “the model hallucinated,” teams can ask better questions. Did it fail to find enough entities? Did it stop too early? Did it attach weak evidence? Did it confuse two similar companies? Did it cite a page that looked authoritative but did not prove the exact claim?

That kind of breakdown is valuable. It also points toward the next stage of AI agents: systems that know when their own research is incomplete before they confidently hand it back to the user.

The Bottom Line

Perplexity’s WANDR benchmark is a reality check for AI research agents. They are not useless. Far from it. They can already help with competitive research, due diligence, literature review, product comparisons, and sourcing work. But at scale, the cracks show fast.

The future of AI research may not be one agent producing a perfect report on demand. It may be a more careful system where agents search, structure, cite, self-check, and flag weak evidence before a human makes the final call. Not as magical as the sales pitch. Probably more useful.

Sources

  • Times of AI: Perplexity’s WANDR Benchmark Shows AI Research Agents Fail at Scale
  • Perplexity Research: WANDR Benchmark — Evaluating Research Agents That Must Search Wide and Deep
Share. Facebook Twitter Pinterest LinkedIn Tumblr Email
Art Ryan

Related Posts

Samsung Profit Surges Nearly Eightfold as AI Memory Boom Drives $80 Billion Q3 Forecast

October 9, 2026

Biohub Expands Virtual Biology Initiative With $1.8 Billion AI Push

October 9, 2026

Google Cloud Unveils Gemini Agent as AI Work Race Accelerates

October 9, 2026

Comments are closed.

Latest News

Samsung Profit Surges Nearly Eightfold as AI Memory Boom Drives $80 Billion Q3 Forecast

October 9, 2026

Biohub Expands Virtual Biology Initiative With $1.8 Billion AI Push

October 9, 2026

Google Cloud Unveils Gemini Agent as AI Work Race Accelerates

October 9, 2026

Council of Europe and Microsoft Sign AI Cooperation Agreement Focused on Human Rights

October 9, 2026
Facebook X (Twitter) Pinterest Vimeo WhatsApp TikTok Instagram LinkedIn YouTube Spotify Reddit Snapchat Threads

AI University

  • Global Universities
  • Universities in Africa
  • Universities in Asia
  • Universities in Europe
  • Universities in Latin America
  • Universities in Middle East
  • Universities in North America
  • Universities in Oceania

AI Tools & Apps Directory

  • AI Productivity Tools
  • AI Coding Tools
  • AI Voice Tools
  • AI Video Tools
  • AI Image Generators
  • AI Writing Tools

Info

  • Home
  • About Us
  • AI Organizations & Associations
  • Contact Us
  • Cookie Policy
  • Copyright Policy
  • Disclaimer
  • Editorial Policy
  • Terms and Conditions

Subscribe to Updates

Get the latest creative news from FooBar about art, design and business.

© 2026 Breaking AI News.
  • Privacy Policy

Type above and press Enter to search. Press Esc to cancel.

Sign Up

Want to stay ahead In Artificial Intelligence?

 Sign up now and get exclusive breaking AI news and special updates—FREE!