‘Productivity’ is too broad to be a useful project target on its own. A business needs to identify the task, define the outcome that matters and measure the whole workflow, including review and correction, before it can decide whether AI has made the work better.
What to take into your next decision
- Define productivity for one workflow in terms the people doing the work can observe.
- Measure the current process before introducing the tool, including review, waiting and rework.
- Track quality and failure modes alongside time or throughput.
- Use the evidence to scale, adjust or stop; do not treat adoption itself as success.
Replace the broad promise with a testable question
Ask whether a particular workflow improves for a defined group of people under stated conditions. For example: can the team prepare an accurate first draft from approved material with less total effort, without increasing corrections or risk?
The question should name the outcome and the boundary. Faster drafting is not useful if approval takes longer, important exceptions are omitted or the work simply moves to a more senior reviewer.
Our recommendation is to measure the task in your own setting. Record who does the work, how the tool is introduced and what people must check, so the comparison describes your workflow rather than a headline from elsewhere.
Build a baseline before the pilot
Observe the current workflow before changing it. Capture enough representative work to understand normal variation, difficult cases and the difference between active effort and waiting time.
Include the steps that are easy to overlook: finding source material, asking a colleague, checking the final output, correcting errors and recovering when information is missing.
A baseline does not need to become a large analytics program. It needs to be consistent enough that the organisation can compare like with like and explain how the measurement was made.
- Total effort across everyone involved
- Elapsed time where delay matters
- Output quality against a shared rubric
- Corrections, rework and escalations
- Common failure and exception types
- User confidence and ability to challenge the output
Measure the whole assisted workflow
Do not stop the clock when the model produces an answer. Include prompt preparation, source gathering, output review, edits, approvals and any follow-up caused by an error.
Keep quality visible. A useful rubric might assess factual support, completeness, adherence to instructions, appropriate tone and correct handling of uncertainty. The criteria should reflect the actual business task rather than generic ideas of good writing.
Track serious failures separately from average quality. A workflow can look acceptable overall while still producing an unacceptable type of error that should prevent wider use.
Use a fair comparison
Compare similar work where possible. A set of easy assisted tasks should not be compared with a set of difficult unassisted tasks. Record the user, task type, source quality and level of review so important differences are visible.
Allow for learning without hiding it. Early use may involve setup and practice; later use may become faster. At the same time, familiarity can make reviewers less attentive, so monitoring should continue after the initial pilot.
When a formal experiment is not practical, use a transparent before-and-after or matched-task comparison and state its limitations. Honest imperfect evidence is more useful than a precise-looking number with an unclear method.
Look for distribution, not only an average
An average can hide who benefits and who carries new work. Examine whether the workflow affects newer and experienced staff differently, whether review shifts to managers, and whether some task types consistently fail.
Ask what people do with any capacity that is released. Faster completion does not automatically become a business benefit; the team needs a useful way to apply the time while maintaining service and quality.
Also watch for behavioural effects. People may over-rely on plausible output, avoid the tool after a poor experience or create unofficial workarounds when the approved process is awkward.
Turn findings into a decision
Before the pilot, agree which findings would support expansion, require a redesign or stop the work. This protects the decision from becoming a defence of money or enthusiasm already invested.
A useful result may be narrower than expected. The system might help with a first draft but not classification, or work well with complete inputs while creating too much review when information is missing.
Document the decision and the evidence behind it. If the workflow continues, keep a small set of meaningful measures and review failure patterns over time. Measurement becomes part of operating the system, not a one-off proof point.
Sources and limitations
The practical method above is TechGuider’s suggested approach. These references explain the underlying guidance; they do not endorse TechGuider or prove an outcome for your business.
- NIST: AI Risk Management Framework
A voluntary framework for identifying and managing AI risks.
- Making the most of the AI opportunity: productivity, regulation and data access
Australian analysis of AI uptake, productivity and the complementary skills, process changes, trust and infrastructure involved.
Suitability depends on your task, information and review process. A source-check date records when the references were checked; it is not professional certification.
How to request a correction ↗
