Job opening

Artificial Intelligence Researcher

Verita AI

10 locations in CA

Filed under Technology, Information and Internet

Full job description

Location: San Francisco

Employment Type: Full-time

Work Model: In-person


About Verita AI

Verita AI works with leading AI companies to identify model gaps and build the human data needed to improve model performance. Verita AI operates a vetted expert network that connects specialized professionals with leading AI laboratories and human-data companies. The network comprises more than 5,000 experts across finance, medicine, law, engineering, music, and other professional domains.


We recently raised a $6 million seed round led by Kindred Ventures.


About the Role

We are hiring an Applied AI Researcher to work directly with clients on model evaluation and data strategy.


You will evaluate model performance, identify failure modes, and recommend the datasets, rubrics, expert workflows, and quality controls needed to address them. You will then work with our operations and engineering teams to turn these recommendations into scalable data programs.


What You’ll Do

  • Work with clients to understand their models, goals, and performance gaps.
  • Design evaluations for generative, multimodal, reasoning, tool-use, and agentic AI systems.
  • Analyze model outputs and benchmark results to identify and quantify failure modes.
  • Recommend data solutions such as supervised fine-tuning data, preference data, expert demonstrations, critiques, and evaluation datasets.
  • Write client proposals covering the methodology, data design, quality controls, staffing, deliverables, and expected impact.
  • Create annotation guidelines, scoring rubrics, gold-standard tasks, and evaluator-training programs.
  • Design pilot studies and measure whether data interventions improve model performance.
  • Build quality systems using calibration tasks, blind review, adjudication, and expert scoring.
  • Work with operations and engineering teams to launch and scale data pipelines.
  • Present findings and recommendations to clients.


What We’re Looking For

  • Experience in applied AI research, machine learning, model evaluation, or data-centric AI.
  • Experience evaluating foundation models or generative AI systems.
  • Strong understanding of benchmark design, human evaluation, rubric development, and statistical analysis.
  • Ability to translate model failures into practical data solutions.
  • Strong Python skills and experience working with model APIs and structured datasets.
  • Familiarity with supervised fine-tuning, preference optimization, RLHF/RLAIF, reward modeling, synthetic data, or LLM-as-a-judge evaluation.
  • Strong technical writing and client communication skills.
  • Ability to independently structure and execute ambiguous research projects.


Nice to Have

  • Experience at an AI lab, foundation-model company, AI data company, or post-training team.
  • Experience designing expert-data or human-evaluation programs.
  • Experience evaluating multimodal, coding, agentic, or tool-use systems.
  • Publications at conferences such as NeurIPS, ICML, ICLR, ACL, or EMNLP.
  • Previous client-facing research, consulting, solutions engineering, or forward-deployed experience.
  • Public research, code, benchmarks, or evaluation frameworks.


Please make sure you have any relevant work samples, including model evaluations, benchmarks, error analyses, research, technical writing, or code repositories.


Explore jobs in these locations

This opening is distributed across multiple locations. Use the application panel to apply through a specific listing, or browse other roles in each local market.

Apply on original listing