James NiuSr. Staff AI EngineerOpen to AI engineering opportunitiesSan Francisco · Remote USEMAILLINKEDIN

I build agentic AI systems, then make them hold up under real traffic.

6 months became 5 days. At Basis Technologies, I rebuilt media planning from ideation through the right audience, right time, and right channels, reaching 98% accuracy across 200 turns.

6 months to 5 daysEND-TO-END MEDIA PLANNING WORKFLOW
98%ACCURACY ACROSS 200 TURNS
1.7M tokensCONTEXT COMPACTION, 8X LONGER SESSIONS

I do the whole loop: the agent, the tools it calls, and the evals that decide whether it ships. If you are putting AI in front of real users, that is the part I own.

Media planning workflow cut from 6 months to 5 days 98% accuracy across 200 turns $175M+ media spend, 3B+ impressions 40 stakeholder recommendations on LinkedIn AI copilot serving 120+ enterprise advertisers 10 years building production AI systems

01 / STAKEHOLDER PROOF

Trusted across the customer, product, engineering, and quality table.

“He focuses on what happens when the agent is wrong, who is accountable, and how you would know before your customer does.”
Daniel JimenezFounder and CEO at Nexus LegalExternal client
“When an agent workflow behaved unexpectedly, James was the person I trusted to diagnose it. Any team building production AI agents would be stronger with him on it.”
Robin LeeSoftware Engineer at Basis TechnologiesSame team
“He understood edge cases, welcomed QA feedback, and jumped on bugs early. That commitment to quality showed up directly in our releases.”
Inna KoganQA Manager at BasisDifferent team
“An extremely forward thinker as AI capabilities emerged. A growth mindset with the ability to learn fast.”
Julianna JavorMicrosoft AI Strategy and TransformationSenior colleague

ABOUT

Measure the model, not the demo.

Ten years building production systems, the last several on AI agents. I started in quant finance and operations research, so I treat model behavior as a system to measure, not a demo to admire.

At Basis I built Compass, an AI media-strategy assistant over Claude, GPT, and Gemini with LangGraph and MCP, and I owned the eval harness that gated its releases: 6 months of media planning down to 5 days, 98% accuracy across 200 turns.

I work the full stack: Python and TypeScript services, agent orchestration, MCP and A2A tools, and the CI gates that keep agents honest in production.

Based in San Francisco. Open to full-time AI engineering roles, here or remote US.

02 / WORK EXPERIENCE

A decade of measurable AI and business impact.

JAN 2026 TO PRESENT

Basis Technologies

Sr. Staff AI Engineer

Built Compass, an AI media-strategy assistant over Claude, GPT, and Gemini via LangGraph and MCP, and the eval harness that gated its releases.

  • 6 months to 5 days for end-to-end media planning.
  • 98% accuracy across 200 turns.
  • 0.03 false positives across 190+ evals.
AUG 2023 TO DEC 2025

Advantage Solutions

Principal AI Solutions Architect

AI systems for retail analytics and workforce operations.

  • 40% staffing accuracy.
  • 20% campaign ROI.
  • 70% lower reporting cost and runtime.
DEC 2022 TO AUG 2023

Leadership Circle

AI Solutions Engineer

Customer-facing AI support systems.

  • 50% fewer support tickets.
  • 99% reliability.
  • 30% lower infrastructure cost.
JAN 2020 TO DEC 2022

Medlytix LLC

Data and AI Architect

Data and AI for healthcare claims and revenue cycle.

  • 25% higher recovery efficiency.
  • 25% fewer denials.
  • 20% better early detection.
NOV 2018 TO DEC 2019

GeoAlliance Consultants Pte Ltd

Forward Deployed Engineer (FDE)

Deployed ML on construction projects: contract review and worksite operations.

  • 60% faster contract review.
  • 15+ late projects prevented.
  • 15% higher worksite efficiency.
OCT 2017 TO OCT 2018

M Capital Group (MCG)

Machine Learning (ML) Engineer

ML for portfolio forecasting and risk.

  • 30% higher forecast accuracy.
  • 20% lower risk.
  • 50% less model-ops friction.

EDUCATION

2017 TO 2019Columbia University in the City of New YorkMS in Operations Research

2014 TO 2017London School of Economics and Political ScienceBS in Economics, First-Class Honors

ORCHESTRATION · EVALUATION · RELEASESeven systems, with evals and failure paths.

04 / SELECTED WORK

Production AI agents, evaluation systems, and inference services.

01 / VOICE AI

Realtime Voice Agent

Turn-taking, interruption recovery, and failover for realtime voice. A voice agent that never talks over your customer.
TWILIOELEVENLABSWEBSOCKETEVALS
02 / REINFORCEMENT LEARNING

Multi-Agent Drone Navigation

Coordination under motion, drift, and conflict. RL policies that survive contact with other agents.
PYTORCHGYMNASIUMPPOFASTAPI
03 / LLM EVALUATION

Self-Hosted Evals Lab

Prompt strategies measured with confidence intervals. Stop arguing about prompts; measure them.
OLLAMALM-EVALSTATISTICSLOAD TESTING
04 / FINE-TUNING

Turn Detection Model

The model waits when the speaker is not done. 33.1 ms p95, because in voice, latency is a feature.
DISTILBERTHUGGING FACEONNXGOLD SET
05 / ENTERPRISE MCP

Atlassian MCP Server

Enterprise tools with identity and audit built in. MCP the way a security team will actually approve.
FASTMCPAUTH0BITBUCKETAUDIT LOGS
06 / COMPUTER VISION

Form Recognition Harness

Uncertain form reads route to human review. Vision that knows when to ask a human.
OPENCVPYTORCHONNXHUMAN REVIEW

05 / HOW I WORK

The checks between a demo and production traffic.

01

Trace every tool call.

02

Fail the build on bad evals.

03

Route around slow or failing models.

04

Stop risky actions at the gate.

06 / TECHNICAL SKILLS

What actually ran behind these systems.

AGENTS
Agentic LLM SystemsLangGraphModel Context Protocol (MCP)Agent-to-Agent (A2A)Retrieval-Augmented Generation (RAG)Function CallingTool CallingMulti-Model OrchestrationPrompt Engineering
EVALS
LLM EvaluationLLM-as-JudgeGolden DatasetsCI-Gated RegressionGuardrailsObservabilityTracing
APPS
PythonTypeScriptFastAPIReactSQLAlchemyWebSocketspandas
DATA
PostgreSQLRedisSQLMultimodal AIComputer VisionSnowflakeDatabricksVector Databases
QUANT
Operations ResearchOptimizationQuantitative EconomicsEconometricsQuantitative FinanceStatisticsStochastic ModelingSimulationDecision Science
INFRA
DockerKubernetesAWSCI/CDLangfuseMLOpsLLMOpsMicrosoft Azure
MODELS
ClaudeGPTGeminiPyTorchHugging FaceONNXTensorFlowTransformersFine-TuningElevenLabs

07 / NEXT

Need an AI agent that survives production traffic?

Full-time, individual contributor: Staff or Principal AI engineer, or forward-deployed AI work where agents meet real users. San Francisco Bay Area or remote US.