He benchmarked his Dot on five everyday tasks — 22/24 objectives met
added
New AI assistants seem to launch every day. I feel like the pitch from many well-funded companies is to hand over control and let an assistant handle everything. I’m a cofounder of Retriever AI (rtrvr), and I strongly believe we should control our AI choices and decide what work to offload. We analyzed 50,000 production user workflows from rtrvr and built an AI Assistant Benchmark with a mock website for running the tests. The tasks cover personal life, work and security guardrails. You can paste the prompts into your assistant of choice and compare accuracy, cost and time taken. The rtrvr extension can also run the tests autonomously. I tested OpenAI dots and Retriever AI with the same prompts and data: claim a flight credit, apply for a job, find creators, reconcile invoices and protect private inbox details. On invoice reconciliation, rtrvr initially got the answer format wrong, then used the page state to self-heal and finish. Dots stopped with two requirements unmet despite knowing the failure. Across these five tasks, rtrvr met 24/24 objectives and dots met 22/24. It's a small sample, and we're adding more tasks. What tasks have you found agents struggle with? Would this help you decide which assistant to use and what to delegate?
u/quarkcarbon (cofounder of rival agent startup Retriever AI — disclosed) ran a head-to-head benchmark: his Dots vs his own rtrvr agent on five everyday tasks — claiming a flight credit, applying for a job, finding creators, reconciling invoices, and protecting inbox privacy. His Dot met 22 of 24 objectives (rtrvr 24/24); Dots took 21 minutes vs rtrvr's 4, and cost ~$20 of Codex-plan usage vs $0.07. First-hand test with concrete per-task outcomes.


