The Integrity Gap

The three Cs

This week I read a fascinating white paper about the utility of behavioural analysis in the AI era. Right after, someone pitched me on LinkedIn about their software that apparently reads micro-expressions.

The Integrity GapThe three Cs

This week I read a fascinating white paper about the utility of behavioural analysis in the AI era. Right after, someone pitched me on LinkedIn about their software that apparently reads micro-expressions. These interactions happened on the same day I heard from a trusted contact that her colleagues are chucking in background checks and due diligence disclosures into Copilot and generating assessments, gap analyses, and action plans.

She was mortified, having reviewed a few, as they contained 30%-40% entirely irrelevant or made-up risks, gaps, and actions. Of course they did - “chucking into CoPilot” without a schema, ontology, API/MCP, output, rubric, and guardrails system will do that. How do I know? We built a robust such system (in testing), and without human guidance it’s ~85% accurate, but with proper human steering (can be taught in a workshop) that goes to human-level accuracy (not quite 100%, as it also makes mistakes, but north of 98%). When making multi-million decisions, 85% is better than 60%-70%, but 98% is best.

That’s not the whole story.

The three Cs

The test results are built on purely data-led reviews:

  1. Define context: who, what, where, how, why (investment in South African healthcare startups needs different treatment to Brazilian e-mobility).
  2. Right-size frameworks: Use that context to define which risks and frameworks should be used to identify gaps. Find potential gaps.

More simply, two Cs - context and controls. But what about the third C - culture?

Culture in this case is human interaction (questions, interviews, reading people). We need to understand the having vs doing gap. Paper-only gap analysis can be wildly misleading. For example, I recently interviewed a team that has no formal risk assessment. I asked how the audit and risk committee worked (without risks to review). They explained, “Oh, we asked [long list of operational teams] what issues they have had to deal with, combined that with audit, reporting, investigative, and other management data and created a top 10 areas for improvement.” In other words, a risk assessment.

The purely AI assessment would have registered: Risk assessment = gap, no formal assessment. Then it would propose that as an action, potentially duplicating (adding time and cost) a functional framework that needs a few tweaks, but no substantive change.

AI doesn’t do nuance

Given the value of human follow-up, the output from our AI-led document review and context analysis software is not a gap analysis; it’s a list of observations and questions that a human needs to delve into. Once the human has done that, then, sure, AI is quicker at parsing the transcript from meetings into a nicely packaged report.

The AI even does a decent job of taking that to an action plan (in current tests). But the human review adds considerably. Humans can read the room. Things we learn during interviews that inform rightsized action plans that an AI can’t do include:

I see a lot of historic action plans. Most are implemented at a <25% rate, because they’re boilerplate and not bespoke. Ours are now 95%+ because they’re human-led.

What about the why?

The final reason I feel AI and behaviourally-focused humans can work together, but with strict division of labour, is that AI may see what, but (for now) not why. Even if we code AI to pick up the behavioural clues that help us pick up on things like interpersonal dynamics and capacity (vocal content, tone, style, micro-expressions, body language, psychophysiology), will it know why?

If a founder’s voice demonstrates sadness and then flickers a micro-expression of fear when asked about audit results, so what? The potential explanations are myriad: they paid a tonne of money for the audit and got no value; they couldn’t understand the audit findings; they hadn’t had time to implement; it surfaced a real issue they’re trying to conceal, etc.

Humans can adapt questions to look for triangulation - to get to the nub of the issue, sensitively. AI will struggle with that, not least because of the way we interact with it.

If you’re not convinced, ask your chosen AI these two questions:

  1. Tell me what you figured out about me that I never actually said and what tipped you off.
  2. Name the personality traits I probably don’t see in myself and be blunt about it.

You will likely find a partially accurate analysis - maybe 60%. Why? Because it doesn’t know your why, and you talk to it (as most of us do) like it’s a robot (without the “reading the room” subtlety many of us try to use in interpersonal interactions).

So, to the guy pitching me software to do what I can do live (read micro-expressions; proof below for any doubters) without the ability to ask live follow-ups to establish why, I’m good for now, thanks. Humans should be jockeying AI, not that other way around.

More Ethics Insight writing

Is it worth a conversation?

Tell us what you are trying to decide. We will listen, ask a few questions and tell you whether we can help.

Start a conversation