AI and traditionally human tasks

Think of all the tasks that AI is doing in a company that were previously carried out by humans. These range from mundane, repetitive tasks through to writing meeting minutes and summaries, drafting reports, analysing documents, building presentations, as well as answering customer questions, making bookings, and completing transactions. Then there’s marketing and sales activities as well as HR activities; qualifying leads, creating content, managing campaigns, generating quotes, preparing proposals, screening CVs, onboarding employees and identifying training needs to name a few. 

AI effectiveness and value

When AI takes over human activities, its effectiveness and value should be measured the same as an organisation would measure employee effectiveness. This requires a change in the organisational mindset. AI should not be seen simply as a ‘technological automation’ capability. It needs to be viewed as a ‘knowledge worker’ assessed on the tangible business value and outcomes it delivers, which is the way an organisation assesses a human employee. 

In the same way that departments are reorganising, restructuring and redefining roles as AI becomes embedded in the operating model, the way AI value and the return on investment is measured needs to evolve. 

High performing AI 

Organisations talk about high performing employees. Simply, they are employees who deliver exceptional results. Isn’t that what organisations should want from AI? 

It is important to note, however, that AI systems are built on software infrastructure and so traditional software metrics such as uptime, security, latency, and reliability must still be monitored and measured. 

To really know how your AI is performing a hybrid approach to measurement needs to be adopted. This would involve traditional software-based metrics as well as employee-like metrics such as whether it achieves business goals and produces valuable work. 

Judging the efficacy of AI using both sets of metrics aligns with how generative AI is being used in real-world situations. It ensures organisations know how the AI is performing from a software perspective in terms of its operational reliability and as a digital worker that must deliver tangible outcomes; using the same method used to evaluate human workers. 

Afterall, a chatbot or voicebot that boasts 99.9% uptime but only resolves 50% of customer problems isn’t successful, just as a human agent wouldn’t be seen as successful resolving just 50% of enquiries. 

Employee-like metrics for AI

AI measurement needs to be approached the way people management is. When AI is doing the work of a human employee then it should be measured the same way which is outcomes focused and not just whether it is online. 

Every organisation has different employee metrics on which performance is assessed. In terms of AI performance measurements there are a number that most organisations adopt as a baseline. These are:

  • Time to competency
    • Reaching useful performance in two weeks is more valuable than one requiring three months. Metrics include the number of training examples needed, length of time to acceptable performance, amount of prompt engineering required and how much human oversight initially needed
  • First contact resolution
    • It is of no value if AI has a fast response time but can’t or doesn’t resolve the problem. Measurement metrics should include percentage of conversations resolved, escalation rate, resolution quality
  • Customer satisfaction
    • Customers care most about whether their issue was resolved and they felt understood. Measurement should include CSAT, Net Promoter Score, user confidence, and conversation ratings 
  • Productivity
    • Employees are often measured in terms of productivity and so should AI. Metrics could include tasks completed, cases handled, documents processed, revenue influenced – and the associated cost of the AI
  • Policy adherence
    • AI should be held to the same standard as employees when it comes to company rules. AI governance must align with a company’s operational controls. Measurement should include: did the AI follow company policy in relation to compliance, security, brand, privacy, and escalation procedures? 
  • Learning and/or improvement
    • AI needs to be constantly improving, month on month. Just as great employees improve, so should AI. Measurement metrics should include reduction in repeated errors, accuracy improvements after feedback, speed of adapting new policies, knowledge retention 
  • Quality of decisions
    • In addition to measuring accuracy, AI should be measured on whether recommendations were accepted, the outcomes were positive and whether decisions reduced business risk
  • Reliability
    • This is not about uptime but about consistent performance at a high standard, which requires AI to produce stable, repeatable results. AI should be measured on delivering consistent answers, behaving predictably, and trustworthiness 
  • Judgement
    • AI needs to demonstrate maturity which means knowing both its capabilities and limits. Without it, trust is compromised. Measurement for AI would include things such as knowing when to escalate, identifying exceptions, recognising uncertainty, and saying ‘I don’t know’
  • Collaboration
    • In the same way that employees don’t work in isolation neither does AI. It is required to collaborate with people, and measurement on its collaborating capability should include successful handoffs, human acceptance rate, team productivity improvements, and time saved for employees

Investing in AI

Every organisation invests in people. They conduct performance reviews for employees, measuring the contribution to business goals. People are employed to make a positive impact via the work they do. The same should apply to AI. 

The questions a manager would ask an employee during a performance review are the right questions that managers should be asking of conversational AI.