Skip to content

AI implementation & agent engineering

AI agents that survive contact with your business.

A working demo is about twenty percent of the job. Codicx builds the other eighty — the integrations, the evaluation, the guardrails, and the human review paths that make automation something you can hand real work to.

Founded
2021
Systems in production
40+
Team
28
Client retention
91%
Diagram of a production AI agent runInbound work items on the left feed a controller, which retrieves context, calls a bounded set of permitted tools on the right, and routes low-confidence cases to human review.INBOUNDCarrier emailEDI 214Rate confirmation~4,000 / WEEKControllerPLAN → RETRIEVE → ACTcontext18 chunksconfidence0.94THRESHOLDPERMITTED TOOLSTMS recordREADContract termsREADRate tableREADLedger writeWRITELEAST PRIVILEGE, SCOPEDPER RUNBELOW THRESHOLDHuman reviewevidence cited inlinecorrection → eval setFEEDBACKAUDIT TRAILrun.start+00.00sretrieve+00.34stool.read+01.02stool.write+02.41srun.commit+02.88sFIG. 01 — EXCEPTION RESOLUTION, MERIDIAN FREIGHT

Operators in logistics, lending, insurance, and healthcare administration

  • Meridian
  • NORTHBROOK
  • Atlas Health
  • CORVEL
  • Halden & Roe
  • PEMBERTON
  • Vantage Mutual
  • KESTREL

The problem we exist for

The pilot was never the hard part.

Almost every company we meet has already proved a model can do the task. What stopped them was everything the demo did not touch: nobody had defined what correct means, the integration was faked with exported spreadsheets, there was no path for the cases the system should refuse, and no team had agreed to be paged when quality dropped at two in the morning.

None of that is a modelling problem, which is why waiting for a better model does not fix it. It is implementation work — unglamorous, measurable, and the entire reason a system either reaches production or spends a year being nearly finished.

We do the eighty percent that decides whether the twenty percent was worth building.

Selected work

Engagements, with the numbers attached.

Every case below states what changed, how it was measured, and how the result is verified on an ongoing basis.

How we work

Five phases. None of them produces a slide deck.

The order matters more than the labels. Defining correctness before building is the single strongest predictor of whether a programme reaches production.
  1. 011–3 weeks

    Discovery

    Understand the work as it is actually done, not as it is documented.

  2. 021–2 weeks

    Definition

    Decide what correct means — with the people who will have to live with the answer.

  3. 036–14 weeks

    Build

    Two-week cycles, working software at the end of each one, evaluated continuously.

  4. 042–4 weeks

    Deploy

    Shadow first, then a narrow slice, then scale — with a rollback at every step.

  5. 05Ongoing

    Operate

    Quality decays quietly. Somebody has to be watching, and it should be measurable.

From the engagement

The part I did not expect was how much we learned about our own process. Half the value was discovering that three of my best dispatchers had been handling the same charge three different ways for years.

Dana WhitcombeVP Operations, Meridian Freight Group

Measured result

Median exception handling time
9 min → 40 sMedian exception handling time
Exceptions resolved without human touch
71%Exceptions resolved without human touchUp from 0%
Annualised capacity recovered
$1.9MAnnualised capacity recovered
Read the full engagement →

Next step

Tell us what is stuck.

Most useful first conversations start with a process that is not working rather than a technology you want to buy. Thirty minutes, no deck, and a straight answer about whether we are the right firm for it.