What an AI Treasury Pilot Actually Tests

An AI treasury pilot is a controlled test of whether software can improve cash visibility, forecasting, liquidity decisions, or routine treasury work without creating unacceptable operational, security, or compliance risk. It is not simply an experiment with a chatbot, and successful pilot demonstrations do not prove that a system can run safely in production. The test should connect to a defined process, use representative data, and compare performance against a credible baseline such as current forecasts, spreadsheet models, bank portals, or human review.

Also worth reading: What Are Treasury AI Controls, and How Should APAC Finance Teams Implement Them in 2026? · How Is AI Cash Flow and Treasury Intelligence Reshaping Corporate Finance in 2026? · What Is the Best Way to Run an APAC Treasury AI Pilot in 2026?

For Asia-Pacific operators, the most defensible first use cases are usually cash-position forecasting, payment and receivable prioritisation, variance analysis, scenario preparation, and reconciliation assistance. A 2026 pilot should avoid beginning with autonomous payment execution unless the organisation already has strong controls, reliable master data, and a mature treasury operating model. AI may predict or recommend, but people and existing approval rules should retain authority over bank instructions, counterparty changes, and material cash movements.

A useful pilot runs for 12 to 16 weeks, although complex integrations can require six months. It should cover at least one complete business cycle and, where relevant, one month-end or quarter-end close. The minimum evidence should include forecast error, time saved, exception detection, adoption, control performance, and total operating cost. A vendor claiming a “30% productivity improvement” without a documented baseline is not enough; the organisation must know whether that figure excludes integration, data preparation, review time, and model maintenance.

Why Treasury Teams Are Piloting AI Now

Treasury is dealing with more frequent cash updates, fragmented banking portals, multiple entities, currencies, payment rails, and local reporting requirements. A mid-sized business operating in six markets may have dozens of bank accounts, yet still lack a dependable consolidated cash position. Manual work then shifts staff from analysis toward collecting data, cleaning spreadsheets, chasing confirmations, and reconciling differences. These are expensive activities because they consume scarce specialist time and make late decisions more likely.

The economic case is strongest where errors are visible and frequent. Reducing a daily cash-update task from 60 to 35 minutes saves about seven hours per week for one employee, or roughly 364 hours annually, before considering quality gains. Those hours are not automatically removable as cost, however; they may simply be redirected to counterparty analysis, bank relationship work, and better liquidity planning. A pilot should therefore track capacity released as well as task time, rather than presenting automation hours as guaranteed savings.

Research from finance and enterprise-ai studies supports the need for discipline after a successful demonstration. Pilot projects often stall because organisations treat the prototype as the product, overlook data ownership, or fail to redesign the surrounding process. A 2025 PwC treasury survey and BCG’s finance-function AI work both point toward the importance of governance, workflow redesign, and measurable returns rather than technology adoption alone. These sources are directional rather than a guarantee for any individual company, so finance leaders should request evidence from vendors, existing customers, and independent evaluations.

Choosing a High-Value, Low-Risk First Use Case

The best first project has frequent volume, accessible data, a repeatable workflow, and an outcome that can be measured within 90 days. Daily cash consolidation and short-horizon forecasting are often suitable because the business cycle is short and users can compare AI output with actual results. Variance explanations can also work well when the system receives bank statements, ledger data, invoice information, and a controlled list of business events. By contrast, autonomously selecting counterparties or releasing payments introduces risks that are rarely justified during an initial pilot.

A practical scoring model can assign 30% of the decision weight to forecast or decision value, 20% to data readiness, 20% to measurability, 15% to frequency, and 15% to control risk. The final score out of 100 should also consider integration effort, but no highly material control risk should be waived simply because the model scores well. For example, a forecast that improves 10-minute cash visibility by five business days could receive a higher priority than an ambitious reconciliation agent that requires access to sensitive banking credentials.

Start with a “recommendation mode” rather than an “execution mode.” The system may identify likely shortfalls, unusual transactions, or receivable delays, but a treasury analyst should review the reason codes and approve action. Every recommendation should be logged with its source data, model version, user, decision, and outcome. Over time, this record supports model monitoring, audit work, and retraining. It also gives the team a way to determine whether the software is genuinely useful or merely producing explanations that staff already knew.

How to Run the Pilot in Four Stages

The first stage is baseline design, normally taking two to three weeks. Treasury should document the current process, system users, data sources, decision rights, failure points, and existing service levels. Record measures such as daily cash availability, absolute percentage forecast error, manual touches, late-payment incidents, and analyst hours. Baselines should use at least eight to twelve representative weeks, and twelve months is preferable when the business is seasonal or affected by payroll, tax, or customer collection cycles.

The second stage is data preparation and configuration, usually taking four to six weeks. Connect read-only sources first, including ERP accounts receivable and payable, general ledgers, bank statements, enterprise-resource-planning forecasts, and approved counterparty master data. The organisation should remove credentials from prompts, apply role-based access, and test whether currencies, time zones, and entity structures are mapped correctly. Data quality thresholds should be explicit; for instance, 98% of scheduled bank balances may be accepted for a visibility pilot, while lower coverage should generate a warning rather than a fabricated completion.

The third stage is shadow operation, lasting four to eight weeks. The AI system produces forecasts or recommendations while treasury staff continue using the established process. Analysts compare the two methods without penalties that encourage blind trust, and every material difference is classified as data error, business-event omission, model error, or acceptable uncertainty. A weekly review should include error magnitude, false alarms, turnaround time, security events, and user feedback. The target should not be zero variance; it should be economically and operationally better than the baseline.

The fourth stage is controlled production, lasting at least four weeks. Management approves a narrow production release, such as daily recommendations to two treasury analysts, while payment authority remains unchanged. Stop conditions should include unauthorised access, unreconciled model output, material unexplained forecast drift, missed regulatory reporting, or a critical workflow failure. Pilot success should then be confirmed over one additional month rather than declared on demonstration day.

Metrics That Survive Scrutiny

A balanced scorecard combines financial, operational, control, and adoption measures. Forecast accuracy should be reported by currency, entity, and horizon because a favourable group average can conceal poor performance in a smaller but volatile account. Cash-flow models are often evaluated with mean absolute error, root mean square error, or mean absolute percentage error, although percentage error becomes unstable when actual balances are close to zero. Teams should therefore use absolute currency error alongside percentage measures and explain the baseline used in every claim.

Operational measures include time to produce a consolidated position, time to investigate exceptions, manual touches, and the percentage of recommendations accepted after review. Control measures include access exceptions, stale master data, unsupported explanations, model overrides, and incidents. Adoption measures include weekly active users, response time, and the share of recommendations receiving a documented disposition. As a screening rule, at least 80% weekly active use during the final eight weeks is stronger than a one-time training attendance rate, but no universal target guarantees value.

A pre-agreed business case might require at least a 10% reduction in forecast error, a 20% reduction in manual processing time, zero material control breaches, and payback within 18 to 24 months. Those figures are decision thresholds, not industry facts. The appropriate threshold depends on the value at risk, project cost, and quality of the current process. If an AI forecast is only 3% better but integrates cleanly with a workflow that saves substantial analyst time, it may still merit controlled expansion; if it is 20% more accurate but requires expensive data work, the calculation may be less attractive.

Pilot featureForecast and analytics optionTreasury workflow agentFull autonomous execution
Typical scopeCash visibility, variance analysis, scenario forecastsReconciliation, collections, exception routing, payment preparationBank instruction and payment execution
Human controlReviews outputs and assumptionsReviews cases and prepared actionsApplies policy within delegated limits
Data accessUsually broad but read-onlyBank, ERP, and master data with workflow accessPrivileged payment and banking access
Pilot duration8–12 weeks12–20 weeks6–12 months in many mature programmes
Main riskForecast error or poor data qualityHallucinated actions, workflow failure, access misuseFraud, payment error, compliance breach
Suitable first decisionControlled shadow runRecommendation modeGenerally not a first pilot
## Cost, Vendor Selection, and Commercial Terms

Pricing varies because vendors may charge by entity, account, user, transaction volume, data volume, forecast horizon, or enterprise contract. A narrow departmental pilot can cost roughly US$10,000 to US$50,000 over three to six months, while a cross-entity production programme may range from US$100,000 to more than US$500,000 in the first year. These are budgeting ranges, not market-wide quoted prices. Integration, security review, internal labour, and process redesign can exceed the subscription fee, especially in organisations with many banks or legacy systems.

Request an all-in proposal separating subscription, implementation, bank connectivity, data migration, support, model usage, and premium security features. Ask whether the pilot converts automatically to an annual contract and whether prices rise after the test. Commercial terms should include a defined data-export process, service-level credits, audit rights, breach-notification periods, model-change notice, termination assistance, and protection against unlimited per-seat expansion. A pilot fee should not be disguised as a discount that creates a mandatory multi-year commitment.

Security and procurement evaluation should cover data residency, encryption, subprocessors, retention, model-training use, privileged access, and incident response. References should be checked with treasury users rather than only procurement champions. A credible vendor should permit a proof of concept using synthetic or masked data, provide a documented evaluation plan, and explain how its product differs from a general-purpose chatbot wrapped around treasury terminology. The 2026 buying decision should be based on demonstrated control and integration performance, not expected feature breadth alone.

Common Mistakes and Failure Conditions

The most common mistake is automating a broken process. If entity mappings, payment terms, opening balances, or ownership are unclear, AI may produce polished output from unreliable inputs. Another error is declaring victory from a polished demonstration using selected historical examples. A demonstration does not reveal latency during bank outages, poor performance on unusual transactions, user workarounds, or the cost of reviewing incorrect recommendations.

Teams also underestimate workflow adoption. Users may ignore the tool if it arrives as another dashboard, produces alerts without prioritisation, or does not save steps in existing systems. Overly aggressive alert rules create fatigue, while overly permissive rules make the system appear reliable without providing actionable information. “Human in the loop” is not a complete control by itself: reviewers need sufficient time, understandable evidence, clear escalation paths, and authority to reject a recommendation.

A third mistake is measuring only forecast accuracy. A more accurate model can still be poor if it arrives after the treasury decision, requires too much manual checking, or omits known events such as tax payments. Conversely, a simpler rules-based forecast may outperform a costly AI model for certain recurring transactions. The final recommendation should compare AI with spreadsheets, existing treasury-management-system functions, deterministic rules, and managed-service alternatives. If a conventional rule handles 85% of cases reliably, the pilot may direct AI toward the remaining 15% rather than applying it to everything.

When to Expand, Pause, or Stop

Expansion should occur only after the controlled pilot reaches its pre-agreed financial, operational, control, and adoption thresholds. A sensible sequence is to expand the number of users, then entities or banks, and only later more sensitive workflows. Between each stage, management should require another four to eight weeks of evidence and a refreshed risk review. Expansion is not justified merely because a vendor is supportive, the pilot is visually impressive, or a deadline is approaching.

Pause when forecast performance deteriorates, source integrations are unreliable, or the organisation cannot maintain master data. A temporary decline may be acceptable during an ERP migration, but the team should define a recovery date and materiality threshold. Stop if the vendor cannot meet agreed security requirements, cannot explain material recommendations, or if total cost exceeds the verified value for two consecutive review periods. Early termination is a valid control when the business case fails, not an admission that treasury has failed.

By 30 September 2026, the practical standard is not whether a company calls itself an “AI treasury company.” The standard is whether it can show current, traceable evidence that its system improves a defined treasury decision under realistic Asia-Pacific operating conditions. A focused 12-week pilot, supported by reliable data and human approval, offers a better route to evidence than an open-ended transformation. The strongest outcome may be a smaller deployment with clear economics; the weakest is a broad rollout justified by claims that cannot be reproduced.