AI CRE Underwriting: 4 Auditable Options Compared
Why “where did this number come from?” is the new underwriting question
The question on a lender call used to be what’s the going-in cap rate? Increasingly it’s how did you get to that in-place NOI?
That expectation has a verifiable root. U.S. bank model-risk expectations trace to the interagency Supervisory Guidance on Model Risk Management — issued as SR 11-7 by the Federal Reserve and OCC Bulletin 2011-12 by the Office of the Comptroller of the Currency — which is built around documentation, independent validation, and explicitly knowing a model’s limitations. You are not the bank, and that guidance does not bind a brokerage. But it shapes the vocabulary and the habits of the credit team reading your package, and in our view that is why AI-assisted numbers now draw follow-up questions that hand-built ones did not. Confirm what your specific lender requires with the lender, not with an article.
The practical translation for a brokerage: an AI-assisted underwriting file needs provenance, not just output.
Standards to hold yourself to
These are targets we think are worth adopting, not measured benchmarks:
What “auditable” actually means in a deal file
Before comparing tools, define the bar. In our view, an auditable AI-assisted underwriting package has five properties:
- Every input is traceable to a source document, page, and (for rent rolls) line item.
- Arithmetic is deterministic. Language models are unreliable calculators; the sum of 214 lease lines should be computed by a formula, not generated.
- Assumptions are labeled as assumptions — vacancy, credit loss, reserves, mark-to-market — with who set them and why.
- Changes are versioned. If the rent roll was restated, the file shows both versions.
- A human signed off, by name, on the numbers that went to the lender.
Every option below — including the one with no AI in it — should be judged against those five, not against how impressive the demo looked.
Option 1: the analyst-built model with a source column
No AI. An analyst keys the rent roll into the model and adds a column citing the source page for anything non-obvious.
Scored on the five properties: traceability is as good as the analyst’s discipline and, critically, unverifiable after the fact if that discipline slipped; math is deterministic by default; labeled assumptions depend entirely on your template; versioning is usually the weakest link, since restatements get saved as file copies with no diff; sign-off is strong, because a named person did the work.
Cost profile: highest per-deal labor cost, lowest tooling and governance cost. That trade is the honest baseline every AI option has to beat.
Weaknesses: it is the slowest path, and re-keying introduces its own errors — transposed suite numbers, missed escalations, a percentage-rent clause nobody read. Manual is not automatically accurate; it’s just accountable.
Best fit: one-off complex assets, small teams, or any deal where the document set is messy enough that extraction tools will need heavy correction anyway.
Option 2: AI features inside the platforms you already own
CRE data platforms, brokerage marketing systems, valuation tools, and general document-AI products increasingly ship extraction and summarization features. The capability that matters for auditability is click-back-to-source: a figure in the output links to the highlighted region of the source PDF. Some products offer it, some don’t, and the feature set changes fast enough that naming vendors here would be stale within a quarter. Check the vendor’s own current documentation, then test it on your own documents.
Strengths: zero integration work, usually inside a security perimeter you’ve already approved, and the vendor maintains the extraction models as document formats drift.
Weaknesses: you get the vendor’s schema, not yours. If your model needs mark-to-market by suite type or a specific reimbursement treatment, off-the-shelf output often lands close but not aligned, and reconciliation eats the time savings.
Best fit: teams already standardized on one platform, doing high volume of conventional office, retail, or industrial rent rolls. This is also the honest default — most brokerages should exhaust it before building anything.
Option 3: a general AI assistant plus a written skill
Here you use Claude, ChatGPT, or Copilot with a skill — a packaged, reusable set of instructions that teaches the assistant to do one job the same way every time. For underwriting intake, the skill specifies: output this exact table, one row per lease line, with a source_page field for every row; flag anything ambiguous instead of guessing; never compute totals.
Strengths: cheap to start, and the skill is where your firm’s underwriting conventions get written down for the first time. That’s a real byproduct.
Weaknesses: there’s no system log. If a lender asks in March what the assistant did in January, you have a chat transcript at best. Context limits also bite on large data rooms, and file-by-file work reintroduces manual handling.
Best fit: a proving ground. Run it for a quarter, see whether the workflow is worth hardening.
Option 4: a custom MCP agent over your data room and CRM
MCP — the Model Context Protocol — is an open standard for giving an AI assistant governed access to your data and tools. A custom MCP server exposes your systems (data room, CRM, comps database, model templates) as callable tools with permissions you define.
For underwriting, the agent gets tools like list_data_room_documents, extract_rent_roll_lines(document_id), fetch_comps(submarket, asset_type), and write_model_tab(deal_id, rows). Two design choices do the audit work:
- The server, not the model, does the math. The agent hands over structured rows; your code computes in-place NOI. Same inputs, same output, every time.
- Every tool call is logged with timestamp, deal, document ID, and user. That log is the audit trail, and it can be exported as an evidence appendix behind the lender package.
Weaknesses: it’s a build. Someone owns it when a data room vendor changes its API, and a half-maintained agent is worse than no agent. It also creates a data-governance exposure the other options mostly avoid: confidential data room contents — leases, rent rolls, sponsor financials, sometimes tenant-identifying information — flow through an external model provider, so retention settings, training-use terms, and your data processing agreement need review before the first document moves, and confidentiality obligations in your listing and NDA paperwork should be checked with counsel.
Where AI genuinely struggles in underwriting
Be specific about failure modes, because vendors rarely are:
- Scanned and handwritten rent rolls. Faxed 1990s leases and marginalia are still the hard case.
- Ambiguous lease language. Reimbursement structures, caps on controllable CAM, and co-tenancy clauses require judgment; extraction tools tend to pick the most common reading.
- Arithmetic and rounding. Treat any model-generated total as unverified until a formula reproduces it.
- Silent gaps. The dangerous failure isn’t a wrong number, it’s a lease line quietly dropped. Always reconcile line count and total base rent against the source document before anything else.
The point of an audit trail isn’t to prove the AI was right. It’s to make it fast to find out when it wasn’t.
Two rules of thumb not to hand an agent
The 2% rule — that a rental should generate monthly rent of roughly 2% of purchase price — is a residential rental screening heuristic, not a commercial underwriting standard. Institutional CRE runs on cap rate, debt yield, DSCR, and unlevered IRR against actual documented cash flow. Never encode a consumer heuristic into an underwriting skill.
The so-called 30% rule in AI is the other one. There’s no formal standard by that name in any technical literature we can point to; it circulates informally to mean something like “let AI do the first chunk, humans do the rest.” Don’t build policy on a folk rule. Set an explicit review threshold instead: which figures require a human to reconcile against source, and who signs.
What this actually changes about brokerage work
The realistic answer on the underwriting side is narrower than the headlines: AI compresses document intake and first-pass modeling, and it raises the evidentiary bar on everything downstream. Judgment on submarket risk, sponsor quality, and pricing strategy isn’t going anywhere.
To size it for your shop, don’t borrow anyone’s percentage. Time your own baseline: hours per deal on rent roll intake and model build, multiplied by loaded analyst rate, multiplied by deals per year. Then subtract review time — which goes up with AI, not down — plus tooling and maintenance. Our broker-ops payback framework walks the arithmetic. If the remainder is small, the honest answer may be that your existing platform’s AI features are enough.
-
Write the standard before the tool
Document what a complete, auditable underwriting file contains at your firm. One page. This is the spec every option gets measured against. -
Test on three closed deals
Run each option against deals where you know the right answer. Score line-count accuracy, base-rent reconciliation, and whether you can trace ten random figures to a page. -
Exhaust what you own
Check current AI features in your CRE platform and document tools against the vendor’s documentation. If they clear the bar, stop here. -
Codify conventions as a skill
Even if you never build a custom system, write the skill. It forces your underwriting standards into text and makes any future build far cheaper. -
Build only where the schema gap is real
If off-the-shelf output requires meaningful rework every deal, that’s the case for a custom MCP server — and name an internal owner, plus a data-handling policy, before the first line of code.
Related: building a custom MCP server for CRE data.
Not sure where to start?
Get a free automation audit: we map your deal pipeline, marketing, and back-office workflows and show you what's worth automating — before you spend a dollar.
Get a free automation audit