Kaaj raises $3.8M in seed funding to power the future of small business lending 🎉Read more
Loan Document Organization

How Kaaj Classifies Loan Documents with Generative AI

Utsav Shah·July 06, 2025·5 min read
How Kaaj Classifies Loan Documents with Generative AI
Table of contents

About the author

Utsav Shah

AI and decision-systems operator with experience building large-scale systems at Uber and Cruise.

Loan packages often arrive as inconsistently named PDFs, scans, and images. Kaaj uses vision-language models, large language models, and traditional machine-learning methods to classify those files, map them to a lender-readable taxonomy, and route low-confidence documents for review. The objective is a cleaner, traceable package before underwriting judgment begins.

What Is Document Chaos and How Does It Slow Underwriting?

Manual triage is the hidden tax on every deal—yet few credit teams track it. Let’s unpack the pain.

Volume: a Monday-morning fire-hose

Consider an illustrative Monday-morning file: a deal arrives in the submissions inbox as “Acme_Corp_Small_Business_Loan.zip.” Inside sit 28 PDFs plus a handful of stray JPEGs. Some are four-page bank statements; others are ID scans, tax returns, invoices, insurance certs, and the obligatory “Final-revised-actual-statement-v3(1).pdf.” Even opening the ZIP can take minutes on a remote desktop. Multiply that by dozens of deals per week and the hours vanish fast.

Variety: layouts that change weekly

Templates drift: last quarter the logo was top-left; now it’s a floating watermark. A file named “BankStatement-Feb.pdf” might actually be a pay stub because a busy borrower grabbed the wrong attachment. Analysts open—and reopen—files just to be sure.

Human drag: the silent cost center

Renaming, re-ordering, and routing files consumes analyst attention before credit work begins. The exact cost varies by package volume, document quality, staffing model, and current workflow, so lenders should measure it against their own queue rather than rely on a universal benchmark.

Error risk: the audit-time nightmare

A misfiled document can create a gap that surfaces later in review. If a required tax return or identity document is classified as “other,” the downstream check may not start when expected. A reviewable workflow should preserve the source file, classification, confidence, and exception path.

The net result: poor classification creates rework and makes the package harder to review consistently.


Why Accurate Classification Matters

Accurate labels aren’t vanity; they power every downstream decision. Let’s zoom in.

1. Fraud and risk controls

Each document type can route to a different configured check. A bank statement may move to cash-flow and document-integrity review; an identity document may move to identity-verification review. A wrong or uncertain label should remain visible so a person can correct it before downstream use.

2. Review and audit trail

A clear record of source document → classification → review state → outcome makes it easier to reconstruct which workflow ran and where a person intervened.

3. Downstream automation

Modern underwriting resembles an assembly line:

Classification → Tailored extraction → Structured data → Review workflow → Human underwriting decision

If the first classification is wrong, downstream extraction and review can start from the wrong document context. Correct tags help the appropriate extraction and verification workflow run, while low-confidence or conflicting files should be routed to a person before they affect underwriting.

4. Borrower experience & brand equity

Repeated requests for a document that was already submitted create avoidable friction. Classification can reduce that back-and-forth when the system preserves the source file, reports confidence, and makes exceptions visible to the reviewer.

Bottom line: classification is preparation infrastructure. It should organize evidence and route uncertainty, not make the lending decision.

Limitations of Traditional Approaches

ChallengeWhy Rule-Based or Classic ML Falls Short
Layout driftSmall template tweaks break brittle regex rules.
Low-text scansOCR alone can miss logos, stamps, or handwriting.
Edge casesThe “other” bucket balloons, forcing manual review.
Scaling new typesAdding a label may require reviewed examples, taxonomy changes, and a new exception path.
 

See how this preparation layer connects to the broader workflow in the Kaaj underwriting FAQs.


Inside Kaaj AI Pipeline

We’re English-first on purpose—our lenders serve U.S. markets, so every research hour goes into squeezing maximum accuracy from English-language files.

StageWhat happensWhy it matters
1. Visual checkpointA vision pass looks for layout, logo, and other visual cues.Provides an initial classification and confidence state.
2. Contextual readA language model reads headings and key phrases.Adds document context and helps identify ambiguous cases.
3. Multimodal reviewA multimodal model considers image, text, and layout together when configured.Supports low-quality scans and uncommon formats; uncertain results still route to review.
 

Smart routing lets high-confidence files continue through the configured workflow while ambiguous, low-quality, or unfamiliar documents move to review. Accuracy, latency, and unknown rates should be measured on the lender's own document mix.


Kaaj’s Lending Document Taxonomy

Document familyExample labelsTypical next workflow
Financial statementsBank statement, financial statement, financial projectionsBank analysis or financial-spreading review
Credit and risk reportsCredit report, PayNet report, business web-presence reportCredit and risk review
Applications and packagesLoan application, submission packageIntake, completeness, and entity review
Revenue and equipment documentsInvoice, quote, equipment scheduleVendor, asset, and amount verification
Identity documentsPassport, Social Security card, commercial driver's licenseIdentity-verification review
Government and taxIRS SS-4 EIN letter, tax returnEntity and tax review
Business formationArticles of incorporation, operating agreement, Secretary-of-State certificateKYB and ownership review
Other or low confidenceUnrecognized, ambiguous, or low-quality filesHuman taxonomy review before downstream use

Names such as “P&L” and “Profit and Loss” can map to the same configured category, while unfamiliar or low-confidence files remain visible for review.


What a Reviewable Classification Output Should Include

  • A lender-readable label tied to the original source file.

  • A confidence or review state that makes uncertain classifications visible.

  • A downstream route to the appropriate extraction, verification, or underwriting workflow.

  • An exception path for unfamiliar, conflicting, or low-quality documents.

API and workflow connections depend on the lender's existing systems and implementation scope. See how classification fits into SMB underwriting automation and document-integrity review.


Burning Questions (FAQ)

Q1. What is a vision-language model (VLM)?
A neural network trained to understand images and the text inside them at the same time—ideal for low-quality scans or documents with important visual cues (e.g., watermarks, stamps).

Q2. How should lenders evaluate document-classification accuracy?
Test the system on the lender's own document mix and review accuracy by class, especially confusing pairs and the low-confidence queue. A single aggregate percentage can hide the errors that create the most downstream work.

Q3. Does it replace my existing OCR or extraction tool?
Not necessarily. Classification can sit before an existing extraction or review workflow so each file reaches the appropriate process. The integration pattern depends on the lender's systems and document stack.

Q4. Can I add a new document type?
Document taxonomies can be configured around the lender's workflow. New types should be introduced with reviewed examples, a destination workflow, and an explicit fallback for ambiguous files.

Q5. Is my data secure?
Kaaj maintains SOC 2 Type II compliance. Review the current controls and security information on the Kaaj Security page.


Ready to See It in Action?

Bring a representative credit package and review how Kaaj classifies files, surfaces uncertainty, and routes documents into the configured workflow. Book a live demo and turn document chaos into a competitive edge.

Ready to see Kaaj in action?

Book a demo and walk through a live deal with our team — from intake to credit memo.

Book a demo

Related articles

Loan Document OrganizationDocument Intelligence for SMB Lenders: Comparing the Top PlatformsAugust 22, 2026 · 11 min readLoan Document OrganizationA $100k Equipment Loan Shouldn't Take 3 Days to get approved and fundedMay 6, 2026 · 4 min readLoan Document OrganizationHow Kaaj Uses AI to Auto-Rename Loan FilesJuly 06, 2025 · 5 min read