> AI + backend engineering studio

Your AI feature works in the demo.We make it work in production.

We build RAG agents, LLM features and automation for teams who need them to hold up in front of real customers. Every build ships with an eval harness and a measured hallucination rate — so you know the number before your users find it.

Every build
Ships with an eval harness
Every quote
Fixed scope, fixed price
Agency work
White-label by default

Three ways to start

Fixed scope. Fixed price.
Named before you commit.

All offers →
  • 01 / auditRAG Eval Audit

    We measure the AI you already shipped, and tell you exactly where it is wrong.

    Fixed scope · one to two weeks
  • 02 / buildProduction RAG Build

    A retrieval agent grounded in your data, delivered with the eval harness that proves it works.

    Fixed scope · four to eight weeks
  • 03 / retainerFractional AI Lead

    Senior AI engineering judgment on your team, without the hire.

    Monthly retainer · ongoing

How we work

Measured,
not demoed.

A demo shows you the best case. It tells you nothing about how often the system is wrong. This is the loop we run instead, on every build.

  1. 01

    Agree the bar first

    Before any code, we write down what the system has to get right and how often. A threshold you sign off on, not one we grade ourselves against later.

  2. 02

    Build the eval, then the feature

    The eval set comes from your real data and your real questions — not a public benchmark that proves nothing about your domain.

  3. 03

    Ship with a number attached

    You get a measured accuracy and hallucination rate for your system, and the ranked list of what still fails. No launch until it clears the bar.

  4. 04

    Hand over the harness

    The eval suite is yours. Re-run it after every model change, prompt tweak or vendor swap, long after we have gone.

What we build

Six things, done properly.

All services →
For agencies

Your clients are asking for AI. We build it under your name.

Mutual NDA and non-solicitation before scoping. Nothing we deliver carries our name unless you ask it to.

How it works →

Products we run

  • Iris

    Conversational commerce for f-commerce merchants. Messenger-first, Bangla and English.

  • Nous

    The channel-agnostic conversational commerce engine. Iris is built on it.

Questions

Before you
ask us.

  • Every AI build we ship comes with an eval harness and a stated accuracy and hallucination rate, measured against your real data. Most AI work is sold on a demo, which shows you the best case and hides the failure rate. We agree a quality threshold before the build, measure against it, and hand you the harness so you can keep checking after we leave.

  • Most audits start within five business days. A RAG eval audit runs one to two weeks; a full production build is typically four to eight weeks with a short scoping phase first.

  • Yes — white-label, under your brand, with a mutual NDA and non-solicitation signed before scoping. If you have the client relationship but not the AI team, that is a large part of what we do.

  • No. RAG agents, eval and reliability systems, LLM features inside existing software, workflow automation, MCP servers, and the backend work that makes any of it hold up in production.

  • Dhaka gives us a deep senior engineering pool and lets us deliver founder-led, without agency layers. You work directly with the engineers building your system. UTC+6 means full-day overlap with the UK, Europe, the Middle East, Australia and Asia.

  • They keep us honest. Iris and Nous run on the same retrieval, agent and reliability patterns we sell — so when we claim something works in production, we are describing something we operate ourselves.

Tell us what it has to get right.
We'll tell you the number.