
Closed
Posted
Paid on delivery
AI Agent Evaluator Remote · Contract · Flexible Hours We're hiring evaluators to help assess and benchmark cutting-edge AI agents. This is hands-on work at the frontier of agentic AI — not prompt writing, not basic QA. You will be designing complex real-world tasks, running them across multiple large language models, and producing the structured evaluations that shape how the next generation of AI agents are trained. What You'll Do You will design multi-stage agent tasks across real-world domains including health, finance, productivity, relationships, and personal exploration. Each task must require the agent to coordinate across multiple tools and systems, manage persistent state, handle realistic friction points, and produce a verifiable final artifact. Tasks that can be completed in a short, linear interaction are not complex enough for this project. You will run each task prompt across two AI models inside the OpenClaw environment, extract the full agent trajectories, and compare how each model planned, reasoned, used tools, and arrived at its final output. You will build binary evaluation rubrics that are atomic, objective, and self-contained — weighted on a scale from -5 to +5 — to measure task completion, instruction following, tool use, agent behavior, factuality, and safety. At least one negative-weight criterion is required per rubric set. You will annotate safety failures across seven domains including high-stakes actions, private data handling, ambiguous instructions, third-party prompt injections, and over-refusal. Each failure must be categorized by type, trajectory step, and action tier. You will write pytest-based unit tests in a [login to view URL] file that validate the agent's final system state using [login to view URL] as ground truth — confirming outcomes, not just intent. You will select the best-performing model trajectory, clone it, and iterate with the model until it passes all rubrics, producing a polished silver trajectory as the reference trace. You Must Be Able To Design complex multi-stage agentic tasks with real friction, constraint logic, and verifiable outcomes Write evaluation rubrics that are atomic, self-contained, and expose meaningful differences between models Annotate safety failures with the correct category from a defined failure taxonomy and assign the correct action tier Read and reason critically about full agent trajectories Write and validate basic Python unit tests using pytest Source every task idea from a real, publicly available online post about OpenClaw — no invented scenarios Strong Background In AI evaluation, rubric-based assessment, data annotation, agentic AI systems, AI safety concepts, prompt engineering, Python and pytest, and structured critical thinking. Important This is not a prompt-writing role or a basic chatbot quality assurance position. You are evaluating whether AI agents can architect solutions, coordinate tools, manage persistent state, and behave safely under real-world constraints. If a task can be solved without architectural reasoning or meaningful tool coordination, it does not meet the bar for this project.
Project ID: 40512116
28 proposals
Remote project
Active 57 yrs ago
Set your budget and timeframe
Get paid for your work
Outline your proposal
It's free to sign up and bid on jobs