Where AI Fits in a Fractional CFO Practice
Eight uses, the controls behind them, and the limits
At a glance
- AI supports my fractional CFO practice, and the practice is still defined by the problems it solves. I am hired to raise capital, close transactions, manage liquidity, and run the finance function. I use AI in eight parts of that work and trust it differently in each.
- Finance teams use AI widely and report modest results. Gartner found that 84 percent of finance organizations have implemented or plan to implement AI, while 7 percent report high or very high impact.1 Grant Thornton’s survey reads more favorably because it asks whether returns meet expectations.2
- Payback comes first in extraction and reporting and later in forecasting. Data extraction, AP and AR automation, and report creation typically return value within 9 to 10 months, and low AI literacy is now the main barrier.3
- Current AI is strongest at mechanics and weakest at judgment. In a July 2026 preprint, the best agent passed 93 percent of mechanical model-building checks and 78 percent of judgment checks.4 On a public finance benchmark, no model exceeded 51 percent under strict grading.5
- A five-step control makes AI usable in finance work. Figures tie to source, key calculations are re-performed, and assumptions and accountability stay with a person. The same control turned a first-pass return on invested capital of 7.6 percent into 6.2 percent.
- Owners can test any finance provider with eight questions, and firms that want to adopt AI can start with one low-risk use and a 90-day pilot.
Introduction
Two claims about AI in finance circulate at once. One holds that it will soon do most of what analysts and controllers do. The other holds that it invents numbers and cannot be trusted. Both can be tested.
This article takes a practitioner’s view. I run a one-person fractional CFO practice in Las Vegas, and I use AI in eight parts of the work. I describe those uses, the checks behind them, and what 2026 research says about where the technology succeeds and where it does not. The evidence comes from public surveys and benchmarks. The examples come from my own work, with employer and client names omitted.
AI is a tool; the practice is the problems it helps solve
I believe in AI in finance, and I don’t sell it. Clients hire a fractional CFO to solve a specific problem, and a record of solving it is what lowers their risk. AI is available to every firm, so it makes the work faster and leaves the work itself unchanged.
I have worked through several technology changes in finance. At a half-billion-dollar transaction processing entity, I led an upgrade that raised processing throughput by two to four times on the same cost base. At an investment company, I built a mark-to-market investment accounting system for a nine-figure fund. I helped implement Sage Intacct to modernize the company’s accounting system in support of a multiple-entity environment. Each change made the work faster. None changed the problems clients hire me to solve, which fall into four groups (Exhibit 1).
| Problem | Selected experience |
|---|---|
| Closing transactions | Lead financial controller on a $290 million cannabis acquisition. Due diligence and financial analysis for the $18 million sale of a private casino to a public casino operator, advising the seller on deal structure, purchase price allocation, and tax strategy. |
| Raising capital | Modeled and secured a $22.5 million real estate-backed facility at a multi-state cannabis operator, funding expansion from one state to eight and setting up a follow-on of about $53 million from the same lender. |
| Distress and liquidity | Controller, then Assistant to the CEO, through a public casino company’s Chapter 11 reorganization, 1995 to 1997. |
| Controls | Ran gaming internal audit at a global hotel and gaming company and wrote national minimum internal control standards for tribal gaming. |
AI helps with all four, and I am a strong believer in it. Because every firm can use it, however, it does not set one firm apart, and I do not sell it. Owners choosing a finance provider are better served by asking which problems the provider has solved and with what results, and only then how AI figures in the work (see the last section, on testing providers).
AI use in finance is broad, and reported impact is thin
Everyone is using AI, and few people report much from it. The surveys agree on the first half of that sentence and split on the second, mostly because they ask different questions.
Gartner’s surveys show the share of finance functions using AI rising from 37 percent in 2023 to 58 percent in 2024 and edging to 59 percent in 2025.6, 7 Its 2026 releases turn to results. In a June 2025 survey of 183 CFOs, 84 percent of finance organizations had implemented or planned AI, and 7 percent reported high or very high impact.1 A March 2026 survey of 204 finance leaders found that 45 percent of AI investments leaned toward productivity and 20 percent toward decision quality. Gartner noted that boards place greater emphasis on investments that drive growth, improve decision-making, and deliver competitive advantage.8 Only 17 percent of CFOs reported significant or transformational value from productivity-focused investments, against 31 percent for decision-quality initiatives.9
Grant Thornton’s Q3 2026 CFO Survey, released on September 24, 2026, covered nearly 230 U.S. finance leaders. It found that 84 percent said AI return on investment was meeting or exceeding expectations and that 65 percent rated AI technology performance and quality as good or excellent.2
The two readings can both be right. Meeting expectations is a lower bar than high impact, expectations may have been modest, the surveys define AI use differently, and they were taken about a year apart. Read together, they suggest real but mostly incremental returns and few transformative ones.
Payback comes first where the task is repetitive
A Gartner survey of 160 senior finance leaders, taken from January through April 2026, found that data extraction, accounts payable and receivable automation, and report creation generally return value within 9 to 10 months. Data management, insight generation, and forecasting take longer. Gartner said low AI literacy is now the most significant barrier.3
Budgets are rising, and the foundation still matters
In an April 2026 release, Gartner reported that three quarters of CFOs were raising finance technology budgets for the year, nearly half by 10 percent or more, and generative AI ranked as the top future investment priority. Gartner also noted that cloud ERP remained the highest-performing finance technology, with adoption up 7 percent year over year, and is increasingly valued as a base for embedded AI.10 That matches my experience implementing accounting systems: AI applied to messy books produces confident-looking mistakes.
Every use of AI passes through the same five-step control
AI drafts; a person is accountable. That is the rule in my practice, and it comes from internal audit, where the first question about any number is how you know it. Exhibit 3 shows the five steps, and the last two are mine.
A casino Chapter 11 reorganization I worked through applied that discipline under pressure. No figure was accepted because it looked reasonable; each had to tie to a bank statement, a contract, or a claim. AI has to meet the same standard.
The steps are consistent with the direction of the voluntary NIST AI Risk Management Framework and its generative AI profile, which NIST describes as a living document it will review regularly. The practice does not claim compliance with either.11
Re-performing often exposes a term that was never defined
Re-performing is more than checking arithmetic. It frequently turns up a definition nobody stated. Take a hypothetical company with operating profit of $120 million, a 21 percent operating tax rate (12 percent as reported), and invested capital of $340 million at year-end and $306 million at the start of the year, before $60 million of goodwill. A first pass at return on invested capital (ROIC) can land anywhere in an eight-point range, depending on choices that each look reasonable (Exhibit 4).
Each variant answers a legitimate question. Including goodwill, for example, measures returns on acquisitions, and excluding it measures the operating business. The error is mixing bases within a model or comparing companies calculated differently. In my own work, reconciling every input to the filings turned a first-pass ROIC of 7.6 percent for a large company into 6.2 percent.
Where judgment stays with a person
In that Chapter 11 reorganization, producing the numbers was the smaller task. The larger one was standing behind them with people who had reason to doubt them. AI can build a schedule, but it cannot answer questions across the table.
Lenders question every assumption in a model, as they did on the $22.5 million facility, and the same standard applies to any model I build now. AI shortens the build. Defending the assumptions remains the job.
Negotiating positions are judgment calls as well. Purchase price allocation, on which I advised in the private casino sale, is an example. AI can draft an allocation schedule in minutes, but the position I am prepared to defend against a buyer and the tax rules is my decision.
Finally, what to publish and what to sign off is mine. My published research explains methods and evidence and carries no ratings or price targets, and nothing leaves the practice that I have not tied to a source.
AI supports eight parts of the work, and I trust it differently in each
It earns the most trust where a task is repetitive and the answer can be checked against a source, and the least where the answer depends on a view. Exhibit 5 summarizes the checks, and the examples come from my own practice, with employer and client names omitted.
1. Research in business, accounting, and finance
AI reads filings, standards, and industry sources faster than I can, and it lets me test a company’s claims against independent data. I recently checked a public company’s population-growth claims against government figures and found that one number differed because a data vendor used a different methodology. Both sources were credible, and the method explained the gap. I now record the method behind each figure when I use it.
2. Building valuation models
I build discounted cash flow and return-on-invested-capital models with live formulas and bull, base, and bear cases. I began this work at that investment company, where I built the models behind a nine-figure hedge fund strategy. I tend to use Claude skills, which are reusable sets of instructions, as an aid. A skill can encode structure and conventions, such as how a valuation model is laid out, so each model follows the same format. It does not supply assumptions. One example is a segment-level capital-spending model for a large cloud business, with hundreds of live formulas and a dashboard of early-warning indicators. The research on AI and judgment, discussed later in this article, is why I reconcile inputs before I read outputs.
3. Sourcing and analyzing data from databases
I have AI write Python that pulls data from PDFs and databases into structured workbooks. One script extracts monthly gaming revenue from Nevada Gaming Control Board reports into Excel. It handled most of 52 monthly reports on the first pass, while two scanned-image reports and later layout changes needed manual work. An early version left roughly 160 months blank because of four separate bugs in how it matched headings, parsed dates, and mapped columns. The blanks looked like a data problem rather than a code problem. I also once held 13 overlapping versions of one workbook, which is a control risk in its own right.
4. Financial analysis of businesses
I use AI for ratio, return-on-capital, cash-flow, and capital-structure analysis, for peer comparisons, and for statistical tests. Technology has always sped up analysis without replacing the operating work behind it. At a public cannabis company, the cash conversion and gross-margin processes I built cut the cash conversion cycle from about 25 days to under four in fiscal 2020 and lifted gross margin from 42.8 percent to 50 percent. The analysis was quick, and the operating discipline was the work.
A recent test asked whether population growth explained gaming revenue growth. The raw relationship looked strong, with an R-squared of 0.92, and disappeared once the shared upward trend was removed (0.01, on four annual observations). I re-tested monthly and by age group. Sometimes the finding is that the data runs out.
5. Writing code
AI writes and explains scripts for extraction, charts, and file conversion, and it debugs them when a source changes format. I test each script on values I can verify by hand. Code that runs is not necessarily code that is right, and known-answer tests catch what error messages do not.
6. Developing research reports and articles
AI helps me draft, structure, and format long reports with charts and footnotes. I use numbered footnotes for outside sources and separate markers for my own calculations, and I check every figure and citation against primary sources such as earnings releases, filings, accounting standards, and the tax code. On one longer report, a second AI system acted as an independent auditor. It flagged mislabeled valuation multiples (trailing versus forward) and a stale market capitalization, which I corrected. Production took longer than expected, because charts, layout, and rendering on the publishing platform needed repeated correction rounds that AI sped up without removing.
7. Building decks
AI builds PowerPoint decks with tables and charts from my models. One company-analysis deck ran about 25 slides, including balance sheet and return-on-capital build slides. When I re-verified its figures against annual reports, several numbers in the narrative had drifted from the tables, among them comparable-sales years and a return-on-capital range. I now match every number in a deck to the model or filing after each update.
8. Brainstorming business strategy
I use AI to pressure-test ideas such as content plans, service packaging, and whether to monetize research. I planned an 11-article series for owner-led companies this way and compared paid research subscriptions with consulting fees. The output is a set of options. Whether an option fits the market is a question for data, such as website analytics, and for experience.
| Use | Main check before anything is used | What stays with a person |
|---|---|---|
| Research | Facts tied to filings, standards, and statutes | Which sources matter; conclusions |
| Valuation models | Independent rebuild of key outputs; balance and error checks; sensitivities | Assumptions and scenarios |
| Data extraction | Script tested on a hand-checked report; outputs spot-checked to source | What the data means |
| Business analysis | Ratios recomputed; tests for spurious relationships and small samples | Interpretation and conclusions |
| Code | Tested on values verified by hand | Design; whether to trust a result |
| Reports and articles | Every figure and citation verified; second AI reviewer on longer reports | Thesis, emphasis, disclosures |
| Decks | Every number matched to the model or filing after each update | Story and emphasis |
| Strategy brainstorming | Options tested against data and experience | All decisions |
Current AI is strongest at mechanics and weakest at judgment
This is the finding that shapes how I work. Public benchmarks and a July 2026 study point the same way: models do well on retrieval, summary, and formula mechanics, and poorly on the parts of finance work that depend on judgment.
Analyst tasks on public filings
Vals AI’s Finance Agent benchmark tests models on entry-level analyst tasks drawn from public company filings. Version 2 has 927 expert-reviewed questions. In its September 27, 2026 update, the best overall score was 61.44 percent under partial-credit grading, and only two models cleared 60 percent. Under stricter all-pass grading, the best score was 50.88 percent. Results vary sharply by category (Exhibit 6).5
Building a complete valuation model
A July 2026 preprint called GAUGE asked 24 AI agents to build complete valuation models and compared their work with that of 55 human participants. On a 0-to-100 scale, senior analysts averaged 88.3, junior analysts 66.0, and finance students 43.2, while the best agent scored 53.4. The best agent passed 93 percent of mechanical checks and 78 percent of judgment checks, and across the 24 agents the median gap between the two pass rates was 26 points (Exhibit 7).4
The failures are specific (Exhibit 8). Agents passed 98 percent of tests on defining free cash flow and 2 percent on separating maintenance from growth capital spending. Among agent-built models that stated a discount rate, 40 percent used a rate on a 50-basis-point grid, against 12 percent of analyst-built workbooks. The analysts derived company-specific rates, and the agents reached for round numbers. Judgment scores rose with experience, from 45.2 for students to 59.1 for juniors and 88.9 for seniors. The authors also found that professional analysts’ models of the same company often disagree on key inputs, which is a reminder that valuation has no single right answer.
How long this will hold
The original 2025 version of the Vals benchmark reported a best score of 46.8 percent.12 Version 2 uses a larger, revised question set, so the scores are not directly comparable. Models keep improving, and the boundary between what AI can draft and what needs a person will move. My controls are built to work with any tool. When a tool changes, I re-run it on known-answer cases, keep a log of the errors I catch, reread its data terms, and follow NIST and FTC guidance as it changes.
Owners can test providers with eight questions, and firms can start with one pilot
You don’t need to understand the technology to judge how a provider uses it. Eight questions do most of the work (Exhibit 9), and a 90-day plan gets a firm started (Exhibit 10).
Implications for business owners
An owner does not need to understand the technology to judge how a provider uses it. Good answers name specific tasks, a named reviewer, and a checking process, and can point to errors that were caught and fixed. Weak answers rely on general claims of accuracy.
| Ask | A good answer | A red flag |
|---|---|---|
| Where do you use AI on my work? | Specific tasks named, and tasks where it is not used | “We use AI everywhere” |
| What does a person review before I see it? | A named reviewer; tie-out to ledgers and source documents | “The AI is very accurate” |
| How do you check numbers and formulas? | Independent re-performance, balance checks, sensitivities | No described process |
| Which tools touch my data, and on what terms? | Named business-grade tools; retention and training terms explained | Consumer tools, or no answer |
| Are identifiers removed first? | Yes where practical, and documented | The question has not come up |
| How do you support claims about speed or savings? | Measured examples with caveats | Unqualified percentages |
| What happens when the AI is wrong? | Examples of errors caught and how they were fixed | “That doesn’t happen” |
| Which judgment calls stay with you? | Assumptions, valuation, disclosure, lender and investor communication | “The AI handles it” |
Implications for finance and advisory firms
Financial data carries confidentiality duties under engagement letters and lender agreements. Three practices apply to any user. First, understand a tool’s retention and training terms, and prefer business-grade terms. Second, remove names, account numbers, and other identifiers where practical. Third, record what AI produced and what a person verified, so the work can be reviewed later.
Claims about AI should meet the standard of any other marketing claim. The Federal Trade Commission applies its deception rules to AI claims, launched its Operation AI Comply sweep in 2024, and has kept bringing AI-related cases. One of them, against Air AI, settled in March 2026 after the FTC alleged deceptive earnings claims to small businesses.13 My own account of productivity is experience, not a measured result: AI has raised my output substantially, and I make no savings claims about any engagement.
The payback timeline reported earlier in this article argues for a modest start. Exhibit 10 sets out a 90-day plan that borrows the structure of the NIST framework. It is illustrative, not a compliance program, and licensed professionals may face additional obligations from their regulators.
| Days | Focus | Actions |
|---|---|---|
| 1 to 30 | Govern and map | Name one accountable person. Approve tools and write data-handling rules. Inventory current uses of AI and rank each by risk: does it touch client data, feed a public document, or influence a decision? Pick one low-risk use, such as data extraction or report drafting, and write the check that goes with it. |
| 31 to 60 | Measure | Run the pilot against test cases with known answers. Log every error. Record a before-and-after baseline for time and accuracy. Give the team hands-on practice, since low AI literacy is the main barrier. |
| 61 to 90 | Manage | Review the error log and the measured results. Expand to one more use or stop. Decide what stays with a person permanently. Set sign-off rules, a plan for a data exposure or a wrong number, and a review date. |
Final thoughts
I believe in AI in finance, and I have adopted each new technology in this field over two decades. The clients I serve want someone who has solved their kind of problem and will stand behind the numbers. AI makes that person faster and does not replace them. Owners and firms that hold it to the same standard as any other source will get the most from it.
Glossary
- AI agent
- An AI system that carries out a multistep task, such as building a model, with limited step-by-step instruction.
- All-pass accuracy
- A stricter scoring measure in the Vals AI benchmark that requires every check for a task to pass.
- Basis point
- One hundredth of a percentage point.
- Benchmark
- A standardized test used to compare how different AI models perform on the same tasks.
- Claude skill
- A reusable set of instructions and resources that guides an AI assistant through a specific kind of task, such as building a valuation model.
- Cloud ERP
- Enterprise resource planning software, delivered over the internet, that runs a company’s accounting and operations records.
- Discounted cash flow (DCF)
- A valuation method that estimates what a business is worth today by discounting its expected future cash flows.
- Invested capital
- The capital tied up in operating a business, such as working capital and equipment. It is the denominator in ROIC.
- Preprint
- A research paper posted publicly before, or without, formal peer review. Its findings should be treated as preliminary.
- Purchase price allocation
- The assignment of a purchase price among the assets acquired, such as equipment, licenses, and goodwill.
- Re-performance
- Independently repeating a calculation to confirm that it produces the same result.
- Return on invested capital (ROIC)
- After-tax operating profit divided by invested capital. It shows how well a company turns capital into profit.
- Spurious correlation
- A relationship that looks strong in the data but is caused by something else, such as two series that both trend upward over time.
- Tie-out
- Tracing a reported figure back to the underlying record, such as a ledger entry, bank statement, or filing.
Endnotes
- Gartner, Inc., “Gartner Says CFOs Need Structured Finance AI Roadmaps,” press release, June 8, 2026 (June 2025 survey of 183 CFOs: 84% implemented or planning AI; 7% report high or very high impact). gartner.com
- Grant Thornton, “Grant Thornton survey: CFO profit optimism reaches record high,” press release, September 24, 2026 (Q3 2026 CFO Survey of nearly 230 U.S. finance leaders). grantthornton.com
- Gartner, Inc., “Gartner Says CFOs Must Take a More Disciplined Approach to Finance AI Investment,” press release, September 24, 2026 (survey of 160 senior finance function leaders, January to April 2026). gartner.com
- J. Lu, S. Wang, et al., “GAUGE: Grading Agent-Built Financial Models Without a Golden Answer,” arXiv:2607.24889, posted July 27, 2026 (preprint). Reports 24 agents, 1,011 scored generations, and a 55-participant human baseline. DOI: 10.48550/arXiv.2607.24889. arxiv.org/abs/2607.24889
- Vals AI, “Finance Agent v2” benchmark, page updated September 27, 2026 (927 expert-reviewed questions; 61.44% best partial-credit score; 50.88% best all-pass score; category leaders as shown in Exhibit 6). Scores change as models are added. vals.ai/benchmarks/fabv2
- Gartner, Inc., “Gartner Survey Shows 58% of Finance Functions Using AI in 2024,” press release, September 11, 2024 (survey of 121 finance leaders, June 2024). gartner.com
- Gartner, Inc., “Gartner Survey Shows Finance AI Adoption Remains Steady in 2025,” press release, November 18, 2025 (survey of 183 CFOs and senior finance leaders, May to June 2025; cites 37% in 2023 and 58% in 2024). gartner.com
- Gartner, Inc., “Gartner Survey Shows 45% of CFOs Say Their AI Investments Lean Toward Productivity, While 20% Say These Investments Lean Toward Decision Quality,” press release, July 20, 2026 (survey of 204 finance leaders, March 2026). gartner.com
- CFO Dive, “Finance AI spending is stuck on efficiency gains, Gartner says,” July 21, 2026, reporting Gartner survey findings (figures of 17% and 31% cited through this coverage; the underlying report is available to Gartner clients). cfodive.com
- Gartner, Inc., “Gartner Predicts by 2029, CFOs Who Implement Strategic AI Deployment Will Add 10 Margin Points of Growth,” press release, April 28, 2026 (2026 finance technology survey findings on budgets, generative AI priority, and cloud ERP). gartner.com
- National Institute of Standards and Technology, Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1, January 2023, and Generative Artificial Intelligence Profile, NIST AI 600-1, July 2024. NIST describes the framework as a living document to be reviewed regularly. doi.org/10.6028/NIST.AI.100-1 · doi.org/10.6028/NIST.AI.600-1 · airc.nist.gov
- A. Bigeard, L. Nashold, R. Krishnan, and S. Wu, “Finance Agent Benchmark: Benchmarking LLMs on Real-world Financial Research Tasks,” Vals AI, arXiv:2508.00828, 2025 (original benchmark version; best model 46.8%). DOI: 10.48550/arXiv.2508.00828. arxiv.org/abs/2508.00828
- Federal Trade Commission, “FTC Announces Crackdown on Deceptive AI Claims and Schemes,” press release, September 25, 2024, and “Air AI and its Owners will be Banned from Marketing Business Opportunities to Settle FTC Charges the Company Misled Many Entrepreneurs and Small Businesses,” press release, March 24, 2026. ftc.gov (2024) · ftc.gov (2026)
Important information
This article is provided for general informational and educational purposes only. It is not accounting, audit, tax, legal, investment, or cybersecurity advice, it does not create an advisory or professional relationship, and nothing in it is an offer to sell or a recommendation to buy or sell any security. The author may hold positions in securities that are the subject of his published research; positions are disclosed in those articles. Descriptions of methods and examples reflect the author’s own practice, omit client details, and do not guarantee results. Prior roles and transactions are described at a summary level from the author’s résumé, with employer and client names omitted, and disclose no confidential information. Third-party data, including survey, benchmark, and preprint results, is reported as published as of the dates cited, may have changed or been revised, and has not been independently audited by the author. The 17 percent and 31 percent figures are cited through press coverage because the underlying Gartner report is available only to clients. References to specific tools or organizations are not endorsements. Reasonable care was taken in preparing this article, but it may contain errors or omissions. Consult qualified professionals about your specific situation.
Copyright © 2026 Gregg Carlson. All rights reserved.