Skip to article
Skip to main content

Land small, prove fast, expand - Explore the Squirro AI Agent Catalog – Download Now

Blog

How to Build a Generative AI Roadmap for Finance Transformation

	Abstract layered structure in which each tier carries a visible reference line back to the tier beneath it, suggesting traceable sources

Imagine somebody in your Frankfurt office is booking a flight to Singapore and wants to know whether they can fly business class. It's in the policy. The policy runs forty pages, it has a regional addendum, and the last time they looked something up in it they got it wrong and the claim came back rejected.

So they do what people do. They message someone in Finance.

That question, and many other questions like it, is what an expense and travel policy agent is for. It’s accessible directly via Microsoft Teams, reads the policy documents you already have, and answers in plain language. A good answer looks like this:

Q: Can I expense business class for a 7-hour flight?
 
A: Yes. Grade 6 and above, EMEA, on flights over 6 hours.
Source: T&E Policy §4.2
Approval: Line manager

It's made up of four parts: the verdict, the conditions behind it, the clause it came from, and who signs it off. About three seconds, no ticket, nobody interrupted.

Producing something that looks like that is the easy part. Any general-purpose assistant will format an answer that way if you ask it to, and it will sound every bit as certain. The difficulty is in being sure that §4.2 exists, that it says what the answer claims it says, and that it applies to the person who asked..

Which is roughly where finance teams are stuck. Gartner found that 63% of finance organizations saw slower-than-expected AI implementation in 2025, and Deloitte's Q2 2026 CFO Signals survey found 59% of CFOs naming the pressure to deploy AI quickly while managing its risks as a top governance challenge. The appetite is clearly there. What's missing, in our read of the market, is a first deployment where the risk of being wrong feels small enough but that is still visible enough to prove its worth and build trust within the organization.

Expense policy turns out to be particularly good for that. The rest of this piece works through why, starting with the easy questions and moving toward the ones your policy doesn't answer.

What a Grounded Answer Actually Contains

Start with the easy version, where the policy genuinely settles the question. A grounded answer to that kind of question carries four things:

  • the verdict,
  • the conditions that produced it,
  • a citation to the document and section it came from,
  • and the approval route.

The citation is what lets anyone other than the person who asked make use of the answer.

An expense and travel policy agent needs to be able to search an approved set of documents: your expense policy, the travel guidelines, the approval matrices, the relevant ERP documentation rather than simply providing an answer that sounds plausible. It has to actually read your documents and report what they say.

Then it footnotes where it read it so that an employee can open §4.2 and confirm that the claim in front of them matches the rule. Six months later, an auditor looking at a disputed expense has an explicit reference to the relevant document, not just somebody's vague memory.

Why the Same Question Has Two Correct Answers

Now let’s make the question slightly harder by changing who's asking.

In many large organizations, cabin class depends on seniority and geography. Your Frankfurt colleague at Grade 6 and a colleague at Grade 4 in Singapore are subject to different rules, so both can ask the same question and get a different, correct answer:

Q: Can I expense business class for a 7-hour flight?
(Grade 4, APAC)
 
A: No. Economy applies at Grade 4 for flights under 10 hours. Premium economy is available above 10 hours.
Source: T&E Policy §4.2, APAC Addendum §2.1
Approval: Line manager

See how the answer cites two documents?  That's because the addendum modifies the base policy for that region.

The agent can do the resolution step because it knows who is asking. How? Because it's running inside Teams, where identity is already established, along with grade and region. It can then work out which clause to follow before it answers, instead of handing the employee the full rule set.

The thresholds themselves stay with Finance. Limits, grade bands, and approval routes are owned, versioned, and attributed, so changing a number changes what the agent says and the record shows who changed it and when. That versioning pays off in an audit, as a limit that was £200 in March and £250 in September needs both answers to have been correct at the moment they were given, and only a dated record of the change can show that.

What Happens When the Answer Isn't in the Policy

Alright, so let’s raise the difficulty by another notch: Some questions have no clause to cite because the information they need isn't in your policy documents.

Budget questions are the clearest example. No policy file holds how much of your team's Q3 travel budget is left. Answering that requires live figures from the ERP or the planning system, so where that data is connected, the agent returns something structurally different:

Q: How much of our Q3 travel budget is left?
 
A: $12,500 remaining of $80,000.
Source: Cost center 4410, actuals as of 17 Sep 2026
Last posted: $2,000 on 15 Sep 2026

The underlying principle carries over from the clause citation. For a document, provenance is a section number. For a number, it's a cost center and an as-of time. In both cases the person receiving the answer can work out how much weight to put on it without going and asking somebody else.

Pointing your agent at policy documents is a two-week job, because the documents already exist and nothing needs to be built. Connecting live budget actuals requires a more complex scoping conversation with whoever owns the ERP, and the answer depends on the type of system ane its configuration.

What Should an Agent Do When Policy Doesn't Settle It?

The hardest version of the question is the one your policy can't answer.

Take a client dinner in Zurich for six people, two of whom are public officials. Your expense policy has a per-head entertainment limit. So far so good. But the public-official dimension is governed by a different policy. A wrong answer there becomes a regulatory problem rather than an expense problem.

Q: Can I expense a €400 client dinner for six in Zurich, including two attendees who are public officials?
 
A: Partially answered.
Entertainment limit at Grade 6, EMEA is €80 per head.
Source: T&E Policy §6.1
 
Attendance by public officials is governed by the Anti-Bribery Policy and sits outside the scope of this policy.
 
→ Case #4417 opened for Finance.
Required before filing: attendee list, business purpose, prior contact with the named officials.

Look closely at the response and you'll see that there are three things happening there. The agent answers the part it can and cites it. Next, it states plainly which part it can't answer and why. And finally, it opens a case rather than producing a confident answer to a question it lacks the context to settle.

How Do You Know What It Told People?

Every answered question is logged and attributed pseudonymously, and each entry expands to show what the agent retrieved in order to produce that answer.

Why is that important? Deloitte's Q2 2026 survey found 43% of CFOs citing insufficient visibility into AI tools and their use as a top governance challenge. A queryable record with unambiguous provenance directly addresses that. It gives you a direct answer to a question like what did this tell people about per diems in July.

The same log turns out to be useful for a second reason: Escalations cluster. When forty people in a quarter ask a question the agent can't settle, the pattern is telling you something about your documents. Ranking those clusters tells you which policy to rewrite first.

What It Takes to Set This Up

At two to four week’s set up time, this is an easy to manage first project to get your Finance department's AI transformation rolling. All you need to supply are three things: the policy documents you already have, a channel where people will actually ask, and a Finance owner for thresholds and routes. You don't need to restructure your policy, build a taxonomy, or finish a data warehouse project first.

We estimate ticket deflection at roughly half of routine policy queries, and time recovered at 15 to 20 minutes per query, most of which is context-switching that doesn’t get tracked. And rapid deployment builds trust while people are still paying attention. Expense policy is an unusually cheap place to run that test.

Why This Is a Smart First Move in a Finance Transformation Roadmap

The practical reason is that all the infrastructure the expense and budget policy agent establishes can be reused by subsequent agents. The policy corpus, the permissions model, the thresholds that Finance owns, and the audit log all carry forward unchanged.

Consider the following roadmap:

Stage

What It Does

What It Newly Requires

What Carries Over

1. Expense & Budget Policy Agent

Source-cited answers on expense, travel, and approval questions, in Teams

Your existing policy documents, a channel, a Finance owner for thresholds

Starting point

2. Proactive Budget Alerting

Real-time budget monitoring, anomaly alerts, and forecast summaries for cost-center owners

Live actuals from the ERP or planning system, and a Finance definition of what counts as an anomaly

Corpus, permissions, thresholds, audit log

3. Intelligent Approval Routing

Routes exceptions to the correct authority with policy context and precedent attached

The approval matrix as structured data, integration with the approval system, and a review gate before anything routes unseen

Everything above

4. Full Finance Shared Services

Extends to procurement, AP/AR, vendor onboarding, and intercompany guidance

Additional corpora per domain. Taxonomy and classification begin to pay here, where routing crosses domains

Everything above

Stage three is where the agent stops answering questions and starts acting on them, so it's the stage where the review gate does the most work.

Stage four is where a taxonomy earns its place. Classifying and routing across procurement, AP, AR, and vendor records is a genuinely different problem from retrieving within a single policy estate. A structured knowledge layer smooths out any differences in the terminology used across these departments.

What to Check Before You Commit

Ask vendors you reach out to to show you an answer with a visible source, then open the source and read the clause it points at. Ask what happens the policy isn’t enough to settle a question, and look at what the resulting case contains before a human has touched it. Ask who owns the thresholds and how a change to one gets recorded. Then ask for the log, and try to answer a question like what did this tell people about per diems during the month of July.

The first of those is easy to demonstrate and the other three aren't. More on separating a working deployment from a good demo sits in six lessons from enterprise AI deployments.

Just remember: none of this tells you whether your policy is any good. What it does is show you, week after week, are the specific places where it isn't, which is more useful in the long run.

Are you ready to start your finance transformation with an easily manageable project that pays for itself, while building trust and a solid foundation for more comprehensive workflow automation? Explore the click-through demo or talk to our team to learn more.

Frequently Asked Questions.

How do you stop an AI assistant from inventing expense rules?
By restricting what it can read and requiring a citation on every answer. A policy agent retrieves only from an approved corpus: your expense policy, travel guidelines, approval matrices, and ERP documentation. Each answer footnotes the document and section behind it. A plausible answer and a correct one look identical on screen, so the citation is the only practical way to tell them apart.
Do you need to restructure your expense policy before deploying an AI agent?
No. Retrieval runs over the documents in the state they are in, including long PDFs, regional addenda, and approval matrices that several people have edited over several years. What you supply is a channel where employees will ask and a Finance owner for thresholds and routes. A taxonomy sharpens classification once scope widens across procurement and AP, as an enhancement rather than a prerequisite.
What happens when an expense question falls outside company policy?
The agent answers the part policy covers, cites it, and opens a case for the rest. A client dinner including public officials has a per-head entertainment limit that policy settles and an anti-bribery dimension that it does not. The case reaches Finance with the policy position assembled and cited and the intake answers attached, so nothing routes past a person on the agent's confidence.
How do you prove to an auditor what an AI agent told employees?
Through the answer log. Every answered policy question is persisted and logged, attributed pseudonymously, and each entry expands to show what the agent retrieved to produce it. That turns a question like 'what did this tell people about per diems in July' into a lookup. Deloitte's Q2 2026 CFO Signals survey found 43% of CFOs cite insufficient visibility into AI tools and their use as a top governance challenge.
Can an AI agent tell employees how much budget is left?
Yes, where budget data is connected. Policy questions need only documents. Budget questions need live actuals from the ERP or planning system, so the answer arrives with a cost center and an as-of timestamp in place of a policy clause. This is the one capability that depends on your environment rather than on the agent, so scope it in week one.
Why start a finance transformation roadmap with expense policy?
Because a wrong answer gets caught within minutes by the person who asked. Larger business cases such as invoice matching or forecasting hide errors inside numbers nobody re-derives for weeks. Expense policy also deploys in two to four weeks, and the corpus, permissions model, and audit log it leaves behind carry into budget alerting, approval routing, and shared services.