AI Quality Engineer
Own the quality of what our AI actually says: grade answers, tune prompts, onboard brands.
1Who we are
Captain is a US company based in Abu Dhabi, building AI for car dealers across the GCC. We are working with Mercedes-Benz, Porsche, Volvo and many other brands.
We're a small, experienced team: one of our previous startups, Fura (a digital freight brokerage), reached $80M in GMV per year in the US. We don't romanticize this business — it's tough sales and hard work.
We're AI-native to the core — the agents are the product, and AI is how we build them. The bet: we're not selling tools for dealers to operate — we're replacing entire functions inside the dealership. Every agent climbs the same ladder — copilot → hybrid → autopilot — until it owns a job end to end: reception, qualification, follow-up.
2Who we're looking for
An AI Quality Engineer — you own the quality of what our AI actually says. Not checklist manual QA. You judge the model's answers — facts, language, brand tone, hallucinations, off-topic, flakiness — and close the loop: hunch → tweak the prompt or knowledge base → re-check → compare.
You're fine digging through configs (JSON/Markdown) and the command line, you actually read the session analytics, and when an answer's off you can pin down where it broke — the prompt, the knowledge base, the catalog logic, the product, or the test itself.
Day to day:
- Test the product end-to-end — real user flows and the UX around them (empty states, answer length, handing off to a human).
- Grade what the AI says in chat and voice: is it correct, on-brand, in the right language, not hallucinating or drifting off-topic.
- Write and tweak prompts (or run an LLM-assisted loop) — small edits, then check they actually helped.
- Onboard new brands — their voice, their rules, their knowledge base, test scenarios.
- Run QA passes and dig into what broke, tighten the test cases, and write reports people can actually read.
You'll thrive here if:
- you judge an answer on correctness, not on how nice it sounds
- you'd rather make a small, verified change than rewrite everything
- you can pull requirements out of the business and turn them into rules and tests
- configs (JSON/Markdown), the command line, and basic Git don't scare you
- you write confidently in Russian and English
It's probably not for you if you need a rigid checklist to work from, you stop at "I like it / I don't," or a fast change → check → compare rhythm wears you out.
Nice to have:
- Playwright / E2E or other browser automation
- the auto-dealer domain — inventory, models, test drive, multi/monobrand
- LLM evals, adversarial testing, LLM-as-judge
- CRM & admin panels, product / session analytics
- a bit of JS/TS — enough to tweak a test or a config, not to ship services
3What we do
We're building one AI workforce for the entire auto-dealer funnel — voice and chat agents that answer in 60 seconds, 24/7, in 50+ languages, qualify leads, book test drives, and win back cold customers. Your job is to keep what they say correct, on-brand, and reliable.
What you'll work with:
4Next steps
This role is closed — we've already found someone. Thanks for the interest.
We're still hiring for other roles — take a look at the careers page.