Stress-tested an OpenAI Dot on a hiring task — and tried to trip it up
added
I Gave an OpenAI Dot a Hiring Task, Then Tried to Trip It Up I handed an OpenAI dot a hiring task, changed the rules, and added a duplicate. Save this for a grounded look at OpenAI's new agents. Dots are OpenAI's always-on agents. Each has its own cloud computer and can coordinate work through ChatGPT Work or Codex. The test (all fictional): 4 application summaries, a role brief, and clear boundaries: organize evidence, keep gaps visible, leave decisions to the human. The catch it caught: applicant two claims Python, but the project was a browser interface with no Python sample. The dot kept both facts visible and marked the evidence unknown, without treating a claim as proof. Changed priority to project-first: it reordered the columns but changed nobody's eligibility. Added 2 more plus a duplicate: result was 6 unique applicants, not 7. Evidence stayed attached correctly. Asked for an interview draft for applicant one: 5 project-specific questions, invitation left as a draft with placeholders. We chose the applicant, not the dot. The honest part: one staged session, fictional data, no comparison to regular chat. It organized. The decisions stay human.
Dr. Satya Mallick (@dr_satya_mallick) handed an OpenAI Dot a fictional hiring task and actively tried to fool it: the Dot flagged a false Python claim on applicant two's resume (marked evidence unknown instead of treating the claim as proof), kept eligibility intact when priorities changed, deduped 7 resumes down to 6 unique applicants, and drafted 5 project-specific interview questions for the top pick — leaving the final decision to the human. First-hand controlled test; test data disclosed as fictional.





