All roles

AI Quality Engineer

Own the quality of what our AI actually says: grade answers, tune prompts, onboard brands.

Product team·Full-time·Remote-first

This role has been filled.

1Who we are

Captain is a US company based in Abu Dhabi, building AI for car dealers across the GCC. We are working with Mercedes-Benz, Porsche, Volvo and many other brands.

We're a small, experienced team: one of our previous startups, Fura (a digital freight brokerage), reached $80M in GMV per year in the US. We don't romanticize this business — it's tough sales and hard work.

We're AI-native to the core — the agents are the product, and AI is how we build them. The bet: we're not selling tools for dealers to operate — we're replacing entire functions inside the dealership. Every agent climbs the same ladder — copilot → hybrid → autopilot — until it owns a job end to end: reception, qualification, follow-up.

2Who we're looking for

An AI Quality Engineer — you own the quality of what our AI actually says. Not checklist manual QA. You judge the model's answers — facts, language, brand tone, hallucinations, off-topic, flakiness — and close the loop: hunch → tweak the prompt or knowledge base → re-check → compare.

You're fine digging through configs (JSON/Markdown) and the command line, you actually read the session analytics, and when an answer's off you can pin down where it broke — the prompt, the knowledge base, the catalog logic, the product, or the test itself.

Day to day:

  • Test the product end-to-end — real user flows and the UX around them (empty states, answer length, handing off to a human).
  • Grade what the AI says in chat and voice: is it correct, on-brand, in the right language, not hallucinating or drifting off-topic.
  • Write and tweak prompts (or run an LLM-assisted loop) — small edits, then check they actually helped.
  • Onboard new brands — their voice, their rules, their knowledge base, test scenarios.
  • Run QA passes and dig into what broke, tighten the test cases, and write reports people can actually read.

You'll thrive here if:

  • you judge an answer on correctness, not on how nice it sounds
  • you'd rather make a small, verified change than rewrite everything
  • you can pull requirements out of the business and turn them into rules and tests
  • configs (JSON/Markdown), the command line, and basic Git don't scare you
  • you write confidently in Russian and English

It's probably not for you if you need a rigid checklist to work from, you stop at "I like it / I don't," or a fast change → check → compare rhythm wears you out.

Nice to have:

  • Playwright / E2E or other browser automation
  • the auto-dealer domain — inventory, models, test drive, multi/monobrand
  • LLM evals, adversarial testing, LLM-as-judge
  • CRM & admin panels, product / session analytics
  • a bit of JS/TS — enough to tweak a test or a config, not to ship services

3What we do

We're building one AI workforce for the entire auto-dealer funnel — voice and chat agents that answer in 60 seconds, 24/7, in 50+ languages, qualify leads, book test drives, and win back cold customers. Your job is to keep what they say correct, on-brand, and reliable.

What you'll work with:

scope · quality across the funnel
├──productweb UI/UX · admin panels · session analytics
├──aiprompt engineering · grounding · language policy · brand voice
├──contentknowledge base · brand onboarding · test scenarios
└──toolingterminal / CLI · JSON + Markdown · Git

4Next steps

This role is closed — we've already found someone. Thanks for the interest.

We're still hiring for other roles — take a look at the careers page.