Skip to content
Subscribe
AI and Automation

Loan officer AI: how to tell a shipped agent from a press release

Dark cover card reading Shipped, or a press release, with four demo checkpoints marked along a track
Nine tests, all of them checkable inside one demo.
In short

Nine tests that separate a working loan officer AI from a demo. What to ask, what to watch during the walkthrough, and what breaks after the pilot.

Most mortgage AI announcements describe software that does not exist yet. Nine specific things separate a press release from a shipped agent. Every one of them is checkable inside a single demo. This page is the rubric the desk scores products against on the mortgage automation software board. Use it before you sign, not after the pilot stalls.

Three figures: 32 products scored, 19 scoring below 3.5, and 9 tests in the rubric
The AI and Automation board, 12 August 2026.

How do I tell a shipped AI agent from a demo?

Nine tests. A product that fails four or more is a roadmap, not a product.

  1. It runs on your file. The vendor loads a loan file you brought and runs it live. A curated sample file means the model works on curated sample files.
  2. It shows its work. Every output traces back to the document and the page it came from. No trace means no audit, and no audit means no examiner defence.
  3. It states an accuracy rate and a denominator. "High accuracy" is not a number. Ask what percentage of what population, measured over what period.
  4. It has an exception path. Ask what happens to the files it cannot handle, and who touches them. Every real system has a queue behind it.
  5. It names a production lender. Not a pilot, not a design partner. A lender running volume through it today, who will take your call.
  6. It survives a bad document. Hand it a phone photo of a paystub, upside down. Watch what it does. Watch what it says it did.
  7. It integrates without a services engagement. If the answer starts with a scoping call, the integration is custom work priced as software.
  8. It prices on volume you can predict. Per-file pricing you can forecast beats a platform fee plus consumption you find out about in month four.
  9. It has a version history. Ask what shipped in the last two quarters. A product with no changelog has no engineering behind it.

Score one point per test. Seven or higher is a product. Four to six is early and worth a pilot with an exit. Under four is a press release with a login screen.

Bar chart of 32 mortgage AI products grouped into score bands, with 19 falling below 3.5
Most of this category is not ready. That is the finding, not an opinion.

Is the AI doing the work, or routing it to a human?

This is the question that separates the category. Ask it directly and ask for the number.

Plenty of mortgage AI runs a human review step on every output. That is a fine architecture. It is also a different cost structure, a different turn time, and a different thing to buy. Vendors rarely volunteer it because the offshore team sits behind the interface, not in the deck.

The tell is turn time under load. A model returns in seconds. A human queue returns in hours, and the hours stretch at month end. Ask for turn time at peak volume, not average turn time. Then ask what peak volume they have actually handled.

Ocrolus publishes more about its review layer than most of the category. MOZAIQ sits at the top of the board on integration depth. Both answer this question without being cornered, which is itself a signal.

Horizontal bar chart of the top twelve products on the AI and Automation board by score
Amber marks the two products named in this section.

A model returns in seconds. A human queue returns in hours, and the hours
stretch at month end.

The test that separates the category

What happens when the AI is wrong?

Ask three things and write down the answers.

Who owns the error. Read the contract language on accuracy. Most mortgage AI contracts place the entire burden on the lender. That is defensible. Know it before an examiner asks, not during.

How you find out. A system that fails silently is worse than no system. The condition that was cleared incorrectly surfaces at closing, or it surfaces in a repurchase demand two years later.

What the correction does to the model. If your correction never reaches the vendor, you are paying to train nothing.

Fair lending exposure sits in this section too. A model that touches credit decisions needs an adverse action path. The reason codes have to be real reason codes. Zest AI built its position on exactly that problem. Underwriting AI and document AI should never be scored with the same questionnaire. The governance questions are covered separately in AI and fair lending.

What to watch during the demo itself

The demo tells you more than the questionnaire, because nobody rehearses the parts that break.

Watch the latency. Watch whether the presenter clicks through a screen quickly. Watch whether they load the file or whether it was loaded before the call started. Ask them to do one thing that was not on the agenda.

Two vendor-published questionnaires are worth reading before the call. They show what the category itself considers a fair question. Wilqo published a lender questionnaire and Mortgage Brain published a supplier list. Neither scores products against the answers. That gap is the reason this board exists.

The category moves fast enough that a demo from six months ago is stale. ICE announced AI voice and chat agents for servicing at ICE Experience 2026. Better and ElevenLabs have published on running a loan agent at scale. Re-demo anything you evaluated before this year.

What the nine tests do not cover

Two things sit outside the rubric and kill more deployments than accuracy does.

Adoption. An agent nobody uses scores the same as software nobody bought. The pattern is identical to the one that kills CRM rollouts, and it has the same cause. Work lands on the desk in a different shape than the desk expects. The CRM adoption post covers the fix, and it transfers directly.

Data exit. Every AI vendor holds a copy of your loan files. Ask what leaves with you and in what format. Ask whether your corrections leave with you too. A vendor who refuses an export clause in the first draft has told you what the renewal conversation looks like.

Neither shows up in a demo. Both belong in the contract.

What this means for your desk

Bring your own loan file to every AI demo. One real file, with a bad scan in it.

Score the nine tests during the call, not after. Send the score to the vendor and ask them to correct anything you got wrong. The ones that engage are the ones worth a pilot, and the ones that go quiet have answered the question.

Then price the pilot with an exit at 90 days and a data export clause. Both belong in the first draft, not the redline.

Frequently asked questions

What is LO AI? Software that handles part of a loan officer's workflow without a person doing it. In practice that covers lead follow-up, document collection, condition clearing, and pre-underwriting. The term covers products at very different levels of maturity.

How do I evaluate a mortgage AI vendor? Run the nine tests above during a live demo on your own loan file. Score each one. Seven or higher is a shipped product. Anything under four is a roadmap being sold as software.

Is the AI doing the work or routing it to a human? Ask for turn time at peak volume and for the size of the review team. A model returns in seconds. A human queue returns in hours and stretches at month end. Both architectures work. They cost different amounts and they scale differently.

Can I see it run on my own loan file? Yes, and a vendor who refuses has told you something. Bring a file with a poor scan in it. The handling of a bad document separates shipped products from demos faster than any question on a questionnaire.

What accuracy rate should I expect? No credible single number exists across the category, because vendors measure different populations. Ask for the percentage, the population, and the measurement window together. A rate with no denominator is marketing.

Products covered in this piece

Ocrolus
MOZAIQ
Zest AI

More analysis

Aug 20, 2026 Your Verification Vendor Just Changed Owners. Price the Renewal Off a Unit, Not a Percentage. Aug 11, 2026 What mortgage technology is actually worth in 2026 Aug 10, 2026 CRM adoption dies at the loan officer’s desk
Who stands behind this review

MortgageTechReview

This score rests on evidence anyone can check. It also rests on the vendor's own documentation, pricing, integration pages, and ownership records. We do not claim to run every product ourselves. Nobody can. The rubric was published before this review existed. The vendor did not write this, and no vendor can buy a word of it. Every product in this category is weighted the same way.

How this was scored · Who publishes this · Dispute this score · Disclosure

The Stack Memo · free · one email a month

One email a month: what's actually worth demoing.

New reviews, category shake-ups, pricing changes we've spotted. No vendor spam, unsubscribe anytime.

No vendor spam·We never sell your address·Unsubscribe in one click

Or read the buying guides →

317 products · 15 categories · one rubric

Every mortgage tool, scored the same way.

No pay-for-play, no vendor-written listicles, no gate. Start from the category you are actually buying in.

Compare →