Perplexity AI Releases WANDR: An Open Benchmark Evaluating Research Agents That Must Search Wide And Deep
Perplexity's WANDR is an open benchmark and evaluation harness with 500 evidence-heavy tasks. It tests whether research agents can discover many qualifying entities and back each one with cited, re-verifiable evidence. Perplexity Search as Code leads at 0.363 soft F1 and 0.133 hard F1. The post Perplexity AI Releases WANDR: An Open Benchmark Evaluating Research Agents That Must Search Wide And…
This is a summary curated by AIFuture. Read the complete article at the original source:
Read the full story on MarkTechPost