Fixyr


The Work Ledger: How to Decide What AI Should Actually Do Inside Your Firm

Quick answer

Most firms responded to AI by cataloguing skills, and cataloguing is not deciding. The Work Ledger inverts the sequence: list the tasks a role actually performs, allocate every one of them to a single mode (agent-run with human review, human-led with AI assistance, human-only by deliberate choice, or retired), rebuild the role from what people retain, then govern what the agents do. The skills fall out of the third step, which means you read them off the result rather than guessing at them in advance. One role, one hour, four columns.


Execution capacity is compounding; design capacity is not

Two numbers published this year describe the gap most firms are living inside. According to Microsoft’s 2026 Work Trend Index, released on 5 May 2026, the number of active agents in the Microsoft 365 ecosystem grew 15x year over year, rising to 18x in large enterprises. Set against that, Deloitte’s 2026 Global Human Capital Trends, a survey of more than 9,000 business and HR leaders across 89 countries conducted with Oxford Economics, found that although 66% of leaders recognise the importance of designing effective human-AI interactions, only 6% report making great progress at it.

The distance between those two figures is where AI returns go to die, and it is worth being precise about what kind of gap it is. Agent deployment is a technology question being answered at speed, largely without anyone in leadership approving it. Deciding what work should exist, who performs it and who is accountable for the output is not a technology question at all, and it sits with the people who run the firm rather than the people who run the systems.

Why skills frameworks underdelivered

Firms have been urged to think in skills rather than roles since 2022. Four years on, CompTIA’s Workforce and Learning Trends 2026, based on an April 2026 survey of 1,049 HR and L&D professionals and published on 28 May 2026, found that only 34% of companies claim to have a formal, organisation-wide programme for reskilling or upskilling current employees. Two thirds have no programme at all, which is a different and more serious finding than having a weak one.

The convenient explanation is that nobody believes in the skills agenda, and every partner reading this knows it is false; belief has been close to universal for years. CompTIA is blunt about where the failure surfaces, noting on skills-based hiring that “dropping requirements has not greatly changed the mix of talent that gets hired.” Firms rewrote the job postings and then hired exactly the people they had always hired. That is not a failure of conviction, and its shape will be familiar to anyone who has closed a set of books.

A chart of accounts is not a general ledger

Every accountant reading this holds two objects in their head that are never confused with one another. A chart of accounts lists every account that could exist; it is structure, and it tells you what is possible. A general ledger records what actually happened, with dated entries, amounts, and a trail back to who posted them and who approved them. Nobody has ever told a client that the firm has a chart of accounts and therefore knows their financial position, because the claim would be absurd for a reason worth stating plainly: a list of what could exist contains no decisions.

Now consider what the profession built for skills. Someone ran workshops, bought a taxonomy, mapped four hundred skills across sixty job families and loaded the result into a system, updated annually if the firm was disciplined and never if it was normal. That artefact is a chart of accounts, and it is genuinely useful for describing what capabilities could exist. Then Monday arrived, a role needed filling, and the firm did what it had always done, looking at who was available, who the partner liked and who had done the job before. There was no ledger anywhere in the process, no record of what was decided, by whom, on what basis, with a name attached. The taxonomy sat beside the decision and never entered it, which is the whole of the failure and is about to repeat itself with AI.

The task replaced the role as the unit of decision

If a taxonomy is the wrong artefact, the answer starts with changing the unit. A role cannot be handed to an agent, because “be a senior tax manager” is a container rather than an instruction. A skill cannot be handed to an agent either, since “critical thinking” is an abstraction invented to describe a pattern rather than to assign work. A task is the only unit an agent can actually take: reconciling these accounts against last quarter and flagging variances over two percent can go to an agent or to a person, which means a firm can decide between them. That makes the task the place where work design now happens, and it means the real shift was never from roles to skills. It runs from roles to tasks, and the skills emerge at the other end.

The four steps

  1. Deconstruct. List the tasks actually performed in a role, fifteen to twenty-five of them. Do not use the job description, which was written by someone who has since left, for a role that no longer exists, and approved by a committee. Everyone knows it is fiction. Write down what the person actually does on a Tuesday.
  2. Allocate. Give every task exactly one of four modes. No task is skipped, no task receives two modes, and nothing goes into a pile marked for later resolution.
  3. Reconstruct. Rebuild the role from what people retain after allocation. This is where the method earns its keep, because skills are the output of this step rather than the input. You look at what remains once agents have taken what they can take, ask what capability that residual demands, and read the answer off the page. The profile is derived, not designed.
  4. Govern. For every task assigned to an agent, name the reviewer, write the quality bar, and define the escalation path before anything runs.

Every skills project your firm has run began with the skills, and this one ends with them. That is the same distinction as the one between the chart and the ledger: the difference between a document and a decision.

The four allocation modes

Agent-run, human-reviewed means the agent executes while a named individual owns the outcome. Named is doing real work in that sentence: not “the team,” not “the tax group,” but a person who is accountable when the output is wrong, and if nobody can be named the task does not belong in this mode. Human-led, AI-assisted means the person performs the work while AI compresses the effort and accountability does not move, which is where most of your firm already operates whether or not anyone authorised it.

Human-only, deliberately means AI is excluded from a task by explicit choice, and there are three defensible reasons to make that choice. Judgment, where the firm wants a person’s reasoning rather than a synthesis. Trust, where a client or a regulator needs to know a human did the work. And development, where performing the task is how somebody learns. That third reason is the apprenticeship, and it is the mode almost nobody selects on purpose.

Retired means nobody does it, and this is where the value sits. It is also the mode every firm skips, for a structural reason rather than an intellectual one: automating a task produces a win that can be reported upward, while retiring one produces nothing to report, since no dashboard exists for work that stopped existing. Automating waste only produces faster waste at higher cost, leaving the firm with an agent generating a schedule nobody reads, at scale, indefinitely, with a maintenance obligation attached. Run this exercise on any role and the first hour will surface tasks nobody has asked for in three years. A ledger that comes back with zero retirements is a technology assessment, not a ledger.

What the ledger produces

Run the method across knowledge roles and the same three capability classes appear each time. Direction covers setting intent and defining the quality bar before work begins, which is scarcer than most firms assume; a manager who cannot articulate what good looks like until they see something bad was survivable when a junior took three days to produce the bad version, and is not survivable when an agent produces forty in an hour. Judgment covers evaluating output, recognising when it is wrong and owning the answer regardless of what produced the draft. Design covers rearchitecting the workflow and deciding what should be done at all, which is the rarest of the three.

Microsoft’s research confirms the pattern from the market rather than from a framework: asked which human skills rise in value as AI takes on more work, AI users ranked quality control of AI output first at 50% and critical thinking second at 46%, both of which are judgment. The three classes are not aspirations from a competency workshop. They are the residual, and they are what remains once everything an agent can take has been taken.

Three questions before an agent runs

Governance is the step that keeps a firm out of the incident report, and it reduces to three questions that must all have answers:

  • Who reviews the agent’s performance over time, as distinct from who reviews a single output once?
  • Who holds the authority to change the workflow the agent runs?
  • How does a local improvement get captured and scaled beyond the one team that found it?

Approving one bad output is a mistake, while approving bad outputs at agent scale is an incident, and the difference between them is entirely a question of throughput. Review capacity in most firms was built for human volume, where a person produces four of something a day and someone checks them; an agent produces four hundred. Where evaluation capacity does not scale alongside execution capacity, the firm has not automated the work but automated the error and removed the checkpoint, and in a profession that signs opinions that eventually shows up in the file.

Start with one role

A firm does not need an AI strategy this quarter; it needs a ledger for one role, and the choice matters less than most people expect. Pick the role you understand best rather than the largest or the most exposed, the one whose tasks you could list from memory and be right about. Give it an hour, list the tasks, allocate every one to a mode, and count the retirements. That count is the diagnostic: zero means the exercise was not run honestly, while four or five means the firm has just located capacity it was about to go out and purchase. Either way, one hour on one role will tell you more about your organisation’s readiness than any vendor assessment on the market.

Apoorv Dwivedi is Founder and President of Fixyr and a speaker on marketing and growth for accounting and advisory firms.


FAQ


A four-step method for allocating work between people and AI: decompose the role into the tasks actually performed, allocate every task to one of four modes, recompose the role from what people retain, and govern what the agents do. The required skills are read off the result rather than defined in advance.


Because they catalogued what could exist without ever entering the decision, in the way a chart of accounts records structure but never transactions. CompTIA’s 2026 research found only 34% of companies claim a formal organisation-wide reskilling programme, and noted that dropping degree requirements has not greatly changed who actually gets hired.


Agent-run with human review, human-led with AI assistance, human-only by deliberate choice, and retired. Every task receives exactly one.


Retired. Automating a task produces a reportable win while retiring one produces nothing to report, so it is systematically under-chosen despite usually carrying the most value.


One role and one hour for a first pass, with the count of retired tasks serving as the diagnostic.

Fixyr

Fixyr helps accounting firms grow through data-driven marketing, SEO, and automation.