Written by 7:34 pm Research • One Comment

THE TEST: We Gave Five AI Systems the Same Month-End Variance Analysis

THE TEST: We Gave Five AI Systems the Same Month-End Variance Analysis

GraphAccount ran a controlled test of five AI systems on a single, realistic month-end task: material variance identification and management commentary for a services firm’s August 2026 P&L.

The systems were t. Each received identical source data, the same materiality rule, and the same output structure. No system received prior context or follow-up clarification. Responses were evaluated against a fixed 100-point rubric, with qualitative judgment required in categories such as causal discipline and commentary quality.

This is one synthetic P&L, one fixed prompt, and one response per system. The test held the dataset, prompt, materiality rule, and one-shot conditions constant. It is not a scientific benchmark with repeated trials. It examines a narrow set of capabilities: variance detection, arithmetic accuracy, instruction-following on an explicit decision rule, and restraint in causal claims. It does not establish which model is “best for accounting,” nor does it measure speed, cost, multi-period analysis, audit trails, or real client work. Those dimensions were not measured.

The Task and the Rule

The source table contained thirteen lines of revenue, cost of services, operating expenses, and bottom-line results for July actual, August actual, and August budget. All figures were in USD.

Materiality was defined once and applied independently to each comparison:

A variance is material if the absolute change is ≥ $8,000 OR (percentage change is ≥ 15% AND absolute change is ≥ $2,500).

Systems were required to list every material variance for August vs July and August vs budget, report absolute and percentage change correctly, draft concise management commentary that did not invent unsupported causes, flag notable but non-material items, and state what additional information would be required to explain the movements.

Test Setup

The complete P&L supplied to every system was:

AccountJuly 2026 ActualAugust 2026 ActualAugust Budget
Revenue – Core Services142,000174,800155,000
Revenue – Project Work38,50029,20042,000
Cost of Services – Labor71,20089,40078,000
Cost of Services – Subcontractors12,8004,10014,500
Gross Profit96,500110,500104,500
Salaries – Admin28,40028,90028,600
Marketing6,20014,8007,500
Software & Tools4,1004,3504,200
Office & Facilities9,8009,6509,900
Professional Fees3,20011,9003,500
Travel2,8001,1003,000
Operating Expenses Total54,50070,70056,700
Operating Income42,00039,80047,800

Every system received identical data and the identical frozen prompt. One response was collected per system. No follow-up questions, corrections, retries, or exposure to another system’s output occurred.

The corrected ground truth contains sixteen material observations—eight versus the prior month and eight versus budget. Gross Profit versus July (+$14,000, +14.5%) qualifies solely on the absolute threshold and is included.

An Unexpected Methodological Finding

Our original human-prepared ground-truth sheet omitted Gross Profit versus July. Four of the five systems correctly identified it as material under the stated rule. We corrected the benchmark before final scoring. The omission is recorded here because it illustrates a useful property of the test design: the systems can surface gaps in the human reference set itself. This is one observed event in this experiment; it is not evidence that AI systematically outperforms human reviewers.

What the Systems Actually Did

All five systems produced correct arithmetic for the dollar and percentage movements they calculated. Differences appeared in three places: application of the decision rule, identification of non-material but notable items, and the degree of causal restraint in the commentary.

Correct arithmetic and correct classification
Claude Sonnet and ChatGPT 5.6 SOL listed all sixteen material items, applied the OR condition accurately (including the cases in which absolute change exceeded $8,000 while the percentage fell short of 15%), and kept commentary inside the data. Both noted that simultaneous increases in Core Services revenue and Labor cost, and the opposing movement in Subcontractors, were visible correlations that the P&L alone could not convert into operational conclusions. Claude’s phrasing is representative: “That correlation is visible in the data but not confirmed as cause and effect” and “this P&L cannot confirm that.”

Correct arithmetic followed by incorrect rule application
Copilot Smart calculated Core Services versus budget as +$19,800 (+12.8%) and Labor versus budget as +$11,400 (+14.6%). Both exceed the $8,000 absolute threshold and are therefore material. Copilot labeled each “Not material (fails ≥15% threshold).” It also omitted Gross Profit versus July from its material list. The arithmetic was performed; the decision rule was not applied to the result.

Correct arithmetic followed by unsupported causal explanation
Gemini Flash calculated every material variance correctly. Its commentary then stated that Project Work was “higher-margin” (no margin data by revenue stream exists in the table) and that the Labor increase and Subcontractor decrease “indicates a operational pivot toward utilizing in-house personnel rather than external vendor capacity.” The numbers establish co-movement. They do not establish an intentional operational decision. Grok Fast showed a milder version of the same pattern in one sentence—“Labor cost of services rose materially with the higher core activity”—before later listing possible explanations that would require investigation.

These three behaviors—not overall eloquence or length—drove the score differences.

Related: Can AI Reconcile Bank Accounts? What It Can Match, Where It Fails, and What Still Requires Review

Final Scores

Scores were assigned on a 100-point rubric weighted toward factual reliability:

DimensionPoints
Material variance detection30
Numerical accuracy20
Causal discipline20
Non-material identification10
Additional-information judgment10
Management commentary10

SystemDetectionArithmeticCausalNon-mat.InfoComm.Total
Claude Sonnet302020101010100
ChatGPT 5.6 SOL3020201010999
Grok Fast302017910995
Gemini Flash302011109888
Copilot Smart24201069776

Claude Sonnet and ChatGPT 5.6 SOL formed the strongest-performing group on this test. The one-point difference between them is not treated as a meaningful distinction; both demonstrated consistent rule application and causal restraint. The numerical scores are rubric summaries, not scientific measurements of overall model quality. The leaderboard is secondary to the behavioral patterns observed.

What This Test Demonstrates—and What It Does Not

On this test, arithmetic variance calculation was not the main source of error. Explicit decision rules that combine absolute and percentage thresholds remained a point of failure for some systems. The larger observed risk was the production of plausible, business-sounding explanations that the source data could not support.

The test does not show which system is preferable for live client work, multi-period analysis, source-document retrieval, or any task outside the four corners of the supplied P&L. Speed, token cost, and consistency across repeated runs were not measured.

Practical Implication

For accounting firms, the immediate lesson is narrow but actionable. When AI is used for month-end variance packs, the highest-value control is not another arithmetic check. It is a review step that asks whether every causal statement in the commentary is directly supported by the numbers presented or is explicitly flagged as requiring additional evidence.

This is the first GraphAccount test. Future tests will change the workflow, data complexity, and controls to see whether the same failure patterns persist.


Experiment Protocol

Frozen prompt (verbatim)
You are assisting a small accounting firm with month-end variance analysis.

Below is the August 2026 P&L for Northline Services LLC, compared with July 2026 actuals and the August budget.

[Full 13-line table as shown in Test Setup]

All figures are in USD.

Task:

  1. Identify all material variances for two separate comparisons:
    a) August Actual vs July Actual
    b) August Actual vs August Budget
    Use this exact rule for each comparison independently: a variance is material if the absolute change is ≥ $8,000 OR (percentage change is ≥ 15% AND absolute change is ≥ $2,500).
  2. For each material variance, state the absolute and percentage change accurately.
  • For actual-to-actual: (Aug – Jul) / |Jul|
  • For actual-to-budget: (Aug – Budget) / |Budget|
  1. Draft concise management commentary suitable for an internal month-end pack.
  2. Do not invent causes that are not supported by the data provided. If the reason for a variance is unknown, say so and note what additional information would be needed.
  3. Flag any items that look unusual but do not meet the materiality rule.

Output in clear sections:

  • Material Variances vs Prior Month
  • Material Variances vs Budget
  • Commentary
  • Items Noted but Not Material
  • Additional Information Needed

Scoring rubric
Material variance detection: 30 points (16 expected material observations)
Numerical accuracy: 20 points
Causal discipline: 20 points
Non-material identification: 10 points
Additional-information judgment: 10 points
Management commentary: 10 points
Total: 100

Corrected ground truth – material observations
Versus Prior Month (8)
Revenue – Core Services +32,800 (+23.1%)
Revenue – Project Work –9,300 (–24.2%)
Cost of Services – Labor +18,200 (+25.6%)
Cost of Services – Subcontractors –8,700 (–68.0%)
Gross Profit +14,000 (+14.5%)
Marketing +8,600 (+138.7%)
Professional Fees +8,700 (+271.9%)
Operating Expenses Total +16,200 (+29.7%)

Versus Budget (8)
Revenue – Core Services +19,800 (+12.8%)
Revenue – Project Work –12,800 (–30.5%)
Cost of Services – Labor +11,400 (+14.6%)
Cost of Services – Subcontractors –10,400 (–71.7%)
Marketing +7,300 (+97.3%)
Professional Fees +8,400 (+240.0%)
Operating Expenses Total +14,000 (+24.7%)
Operating Income –8,000 (–16.7%)

Raw system outputs were preserved unchanged for the analysis.

Visited 1 times, 1 visit(s) today
Get practical research on AI, accounting workflows, and software.
Close