Data Scientist

Actively hiring Posted 3 months ago 2 min read

Role overview

We're seeking a data-driven analyst to conduct comprehensive failure analysis on AI agent performance across finance-sector tasks. You'll identify patterns, root causes, and systemic issues in our evaluation framework by analyzing task performance across multiple dimensions (task types, file types, criteria, etc.).

What you'll work on

Statistical Failure Analysis: Identify patterns in AI agent failures across task components (prompts, rubrics, templates, file types, tags)
Root Cause Analysis: Determine whether failures stem from task design, rubric clarity, file complexity, or agent limitations
Dimension Analysis: Analyze performance variations across finance sub-domains, file types, and task categories
Reporting & Visualization: Create dashboards and reports highlighting failure clusters, edge cases, and improvement opportunities
Quality Framework: Recommend improvements to task design, rubric structure, and evaluation criteria based on statistical findings
Stakeholder Communication: Present insights to data labeling experts and technical teams

What we're looking for

Experience with AI/ML model evaluation or quality assurance
Background in finance or willingness to learn finance domain concepts
Experience with multi-dimensional failure analysis
Familiarity with benchmark datasets and evaluation frameworks
2-4 years of relevant experience

We consider all qualified applicants without regard to legally protected characteristics and provide reasonable accommodations upon request.

Tags & focus areas

Used for matching and alerts on DevFound

Contract Remote Ai Data Science

Data Scientist

Role overview

What you'll work on

What we're looking for

Tags & focus areas

Ready to Join the Team?