Digital government in Australia has moved well past the phase where launching a service counted as success. Federal and state agencies are now expected to demonstrate that services work, that users complete their intended tasks, and that the numbers backing those claims are real. The pressure comes from multiple directions: DTA platform mandates, ministerial accountability, and an increasingly literate public that can spot when a form takes 47 steps to complete.
Getting performance measurement right is harder than it looks. The metrics that are easy to collect (page views, uptime) rarely tell you whether a service is actually serving anyone well. The metrics that matter (task completion rate, error rate, time on task) require deliberate instrumentation, user research, and a governance structure willing to act on bad results.
What the DTA framework actually requires
The Digital Transformation Agency sets the performance expectations for Commonwealth digital services through the Digital Service Standard. The Standard requires agencies to measure performance on four core areas: user satisfaction, digital take-up, completion rate, and cost per transaction. These aren't suggestions. Services built under the Standard are expected to publish these metrics publicly, typically through the Performance Dashboard.
The Performance Dashboard collects live data from agency services and exposes it in a format that anyone can inspect. Not every agency populates it consistently, which has been a persistent criticism. The DTA has tightened its expectations in this area, with its DTA strategy in 2026 placing renewed emphasis on accountability for service outcomes rather than delivery activity.
State governments apply their own standards. NSW uses the Digital.NSW framework, Victoria the Digital Standards published by the Department of Government Services, and Queensland has its own Digital Experience Standards. These differ in their metric requirements, but the pattern is consistent: completion rate and user satisfaction scores are universal expectations.
The four metrics that drive accountability
User satisfaction is the most visible metric and the most easily gamed. Agencies typically collect it via a short post-transaction survey: a thumbs up or down, or a 1-to-5 rating. The problem is that satisfaction captures how a user felt about the experience, not whether they achieved their goal. A person who eventually completes a complex Centrelink form after three attempts may still rate the experience neutrally, because they succeeded.
Completion rate is more revealing. It measures the proportion of people who start a transaction and finish it successfully. Low completion rates flag real problems: confusing navigation, document requirements users don't anticipate, technical failures, or simply a service design that doesn't match how people think about their situation. Services Australia tracks completion rates across MyGov-connected services, and the numbers have historically varied significantly between transaction types.
Digital take-up measures the share of a service's total volume being handled through the digital channel versus phone, paper, or in-person. High take-up generally reduces cost per transaction, which is the fourth metric. Cost per transaction calculations require agencies to include back-office processing costs, not just front-end infrastructure, which makes honest reporting harder than it sounds.
Instrumentation: where most agencies fall short
Collecting these metrics requires real instrumentation work before a service launches, not after. Many agencies still treat analytics as a post-launch addition, which means they have no baseline and no way to measure whether changes improve outcomes.
The minimum viable analytics setup for a government digital service includes: event tracking on each step of a multi-step form, error event capture so the team can see where users fail, funnel analysis to identify drop-off points, and session segmentation by device type and browser. Mobile users routinely experience government services differently from desktop users, and many performance failures are device-specific.
Real User Monitoring (RUM) tools give agencies actual load time data from users' devices rather than synthetic performance benchmarks run from a data centre. Synthetic testing tells you whether the service is technically available. RUM tells you whether it's usable for someone on an NBN HFC connection in outer Brisbane or a 4G mobile connection in regional South Australia.
The ATO has invested heavily in this instrumentation layer as part of its broader modernisation program. Its pre-fill and lodgement services collect granular performance data, which feeds into internal reporting cycles and informs the annual myTax release cycle. The detail behind that work is explored in the coverage of ATO digital services and how the tax office is modernising.
Accessibility performance as a distinct measurement domain
WCAG 2.1 AA compliance is a legal baseline for Australian government digital services under the Disability Discrimination Act 1992 and the Web Accessibility National Transition Strategy. But compliance against a checklist and actual accessibility performance are different things.
Automated accessibility scanning tools (Axe, WAVE, Deque's suite) catch roughly 30 to 40 percent of WCAG failures. The rest require manual testing with assistive technologies: screen readers such as JAWS and NVDA on Windows, VoiceOver on iOS and macOS, and Switch Access on Android. Agencies that report "WCAG AA compliant" based solely on automated scans are reporting an incomplete picture.
Some agencies now include accessibility-specific performance indicators in their service dashboards: the percentage of transactions successfully completed by users of assistive technology, and complaint rates via the disability discrimination complaints pathway. This is still rare, but the DTA has signalled it wants accessibility treated as a measurable outcome rather than a checkbox.
Connecting performance data to service improvement
The hardest part of performance measurement isn't collecting the data. It's building a governance model that acts on it. Agencies with mature digital programs run regular service review cycles: a scheduled review of performance metrics, a linked backlog of improvements, and a clear owner who has the authority to prioritise fixes.
Without that structure, performance data sits in dashboards that nobody consults until a minister asks a question. The DTA's whole-of-government approach to data sharing, covered in the analysis of whole-of-government data sharing in Australia, is partly aimed at giving central agencies better visibility into service performance without requiring every department to build its own reporting infrastructure.
User research sits alongside analytics as a mandatory input under the Digital Service Standard. Quantitative metrics tell you where users are failing. User research tells you why. Both are required. Agencies that treat performance measurement as a purely technical function, and skip qualitative research, end up optimising for metrics without fixing the underlying service problem.
What good reporting looks like in practice
The best-performing agencies treat their public performance dashboards as live documents updated at least monthly. They publish completion rates broken down by transaction type, not averaged across the entire service portfolio. They report satisfaction scores with sample sizes, so readers can judge statistical significance. And they report cost per transaction with a clear methodology statement, so the number is comparable year on year.
Services Australia's digital channel has moved in this direction, though the transparency of its reporting has varied across program areas. The Medicare online services and Centrelink online account platforms publish headline metrics; the underlying methodology is harder to access.
For state government agencies, the Victorian Digital Standards represent one of the more detailed frameworks for performance reporting, specifying not just which metrics to collect but how to calculate them and at what frequency to report. NSW's Digital.NSW framework links performance reporting to funding gate reviews, which creates a real incentive for agencies to take measurement seriously before a project hits its next spending approval.
The trajectory is clear. Performance measurement for Australian government digital services is moving from voluntary good practice to a mandatory, inspectable obligation. Agencies that build the measurement infrastructure early, and build governance structures that act on the results, will find it far less painful when the inspection arrives.

