When capability outruns the organization

How to build an institution that can learn fast enough to use what becomes possible

In this essay

In Interstellar, the crew returns from a planet where a few hours have cost more than twenty years aboard the ship. Romilly has lived through the interval they experienced as a short expedition. In the screenplay, Brand struggles to reconcile knowing what would happen with encountering its consequences. “I thought I was prepared. I knew all the theory. Reality’s different.”1

Organizations can understand a technological change and still make decisions on a timetable that leaves them unprepared for it. Leaders approve an investigation, wait for a planning cycle, commission training, and eventually assess a pilot against assumptions established months earlier. Meanwhile, the capability being investigated changes, as does what customers can obtain elsewhere. Every step may be defensible within the organization’s calendar. The cumulative delay can make the result much less valuable.

I think this is becoming one of the central leadership problems of the AI transition. A company can have talented people, extensive use of AI tools, and impressive demonstrations while remaining slow at turning a newly useful capability into a reliable way of working. The delay matters because implementing the technology is itself a source of learning. Teams discover which information is missing, which decisions are poorly specified, and which parts of the organization must change before the capability creates value. Deferring that work defers the understanding needed for the next advance as well.

There is a much more promising possibility. An organization could become able to investigate opportunities that previously required a dedicated team, give more people access to specialist knowledge, and make smaller customer problems economically worthwhile to solve. As the cost of producing an initial answer falls, it could invest more effort in finding the right question and following a useful answer through to a result. That would be a substantial change in what the institution can attempt.

Getting there requires decisions about people, resources, and authority that installing a tool will never make for us. The organization we should build is one that can repeatedly change how work gets done when the evidence warrants it, without requiring each change to become a campaign against its own structure.

1. What waiting costs

The pace of AI progress is uneven, but it is already sufficient to make a static account of its usefulness unreliable. METR’s work on task-completion time horizons tracks an expanding range of software tasks that agents can complete at a specified success rate. The measure concerns tasks indexed by how long they take people, rather than a guarantee that an agent can run unattended for that duration. Its relevance to an organization is that a class of work dismissed as impractical can become worth testing again.2

The capability need not improve everywhere for the competitive implications to be large. A system that becomes reliable at one expensive part of a workflow can change the economics of the service around it. If preparing a customer-specific configuration becomes much cheaper, a business might serve accounts it previously rejected as too small. If investigating a service failure becomes faster, it might offer a level of support competitors cannot provide profitably. Those opportunities depend on the rest of the workflow being able to absorb the change.

Adoption statistics provide only a partial view of that readiness. In a 2026 nationally representative survey, Alexander Bick and colleagues find generative AI use spread across many occupations and tasks, with substantial differences in adoption among people doing similar work. Their account describes adoption as widespread but shallow. A capability being relevant to a job does not establish that the person performing it has incorporated the capability into the job.3

I would distinguish three delays inside that apparent gap. The first is the time needed to recognize that something has become possible. The second is the time needed to test whether it helps in the organization’s circumstances. The third is the time needed to make a successful method available beyond the person who discovered it. They have different causes. A technology briefing might reduce the first while doing almost nothing about a team’s lack of data access or a promotion system that gives no credit for teaching others.

The last delay is particularly consequential. Suppose two organizations obtain the same models at the same prices. One tests a narrowly defined workflow, records the failures, and changes its data and operating procedures accordingly. The other waits for the next generation of models. When that generation arrives, both can buy it, but only the first has learned which parts of the problem the new model would need to solve. Access to the latest technology has become equal again; the accumulated ability to use it has not.

This does not make the earliest adopter the inevitable winner. An early investment can harden around an approach that later becomes unnecessary, and a company can exhaust its people by changing tools faster than anyone can learn them. The useful distinction is between postponing a consequential deployment and postponing inexpensive learning. A bank might reasonably withhold authority from an untested agent while still examining its proposals against completed work. A software company can test a different development process without migrating its entire production system. The option to wait on the irreversible decision becomes more valuable when the organization uses the interval to reduce uncertainty.

The question for a leadership team is therefore quite specific. Which important capability have we declined to investigate because we lack evidence, and what have we done to obtain that evidence? If the answer has remained unchanged through several planning cycles, the delay itself deserves an owner and a decision.

2. Follow the work across the organization

Individual productivity can improve without changing the product or service a customer receives. A randomized study of 7,137 knowledge workers across 66 firms found that access to generative AI reduced time spent on email, but did not detect changes in the quantity or composition of tasks performed. The study followed workplace-tool access over six months. It provides a useful reason to examine the difference between saving time within a task and reorganizing the work that follows it.4

Consider a hypothetical business onboarding a new customer. Sales records a commitment, an operations team reconstructs what was promised, finance checks the commercial terms, and specialists determine whether the requested configuration can be supported. Someone then reconciles the answers before the customer can begin. AI could make each team’s summary faster while leaving the customer waiting through the same sequence.

The first investigation should follow one onboarding all the way through. How much time is spent acquiring information, resolving contradictory information, deciding what is acceptable, or simply waiting for the next person? Which questions are genuinely different, and which are repetitions caused by teams keeping separate accounts of the same request? A process map becomes useful when it explains the delay well enough to suggest a change.

One possible redesign would assemble the customer’s requirements, commercial commitments, and applicable constraints into a shared record at the outset. An agent could identify missing information before the request enters a specialist queue. Checks that do not depend on one another could proceed together. Routine commitments within agreed limits could advance automatically, while a novel configuration reaches someone who has both the expertise and authority to decide it.

The benefit comes from changing the work that must pass between people. A faster summary is helpful; eliminating the need to reconstruct the same commitment five times changes the process more substantially. It also reveals questions that the old process may have left conveniently ambiguous. Who is entitled to promise an exception? Which record prevails when a sales note conflicts with a contract? What happens when one department benefits from a decision whose cost falls on another?

Those are leadership decisions. Connecting an agent to several systems can expose them more quickly, but cannot settle them by making an unauthorized choice easy to execute. The organization needs an agreed result, a way to resolve conflicting objectives, and an accountable owner for the whole journey. Otherwise, each function can report a successful AI initiative while the customer’s elapsed time remains unchanged.

This also changes how we should interpret an investment period. Brynjolfsson, Rock, and Syverson’s work on the productivity J-curve explains why general-purpose technologies require complementary investments in processes, skills, and other intangible assets before their benefits are fully reflected in measured productivity.5 The appropriate response is to examine what a transformation program is actually accumulating. Reliable data access, reusable tests, and a working handoff have prospective value. Another presentation describing a future operating model may not.

For the onboarding example, progress should eventually appear in completed, usable customer setups, the time required to achieve them, and the cost of fixing what went wrong. The team should also investigate the slower and more difficult cases. A reduction in average onboarding time obtained by quietly excluding complex customers would represent a different commercial choice, which leaders would need to acknowledge.

The broader implication is that AI adoption should be organized around work with an observable destination. Department boundaries remain relevant to expertise, staffing, and professional responsibilities. They should not prevent us from discovering that the customer’s problem crosses all of them.

3. Expertise after the handoffs change

When people can carry work farther with AI, the question of how to organize expertise becomes more interesting. Some handoffs exist because one person lacks a capability another possesses. Others exist because someone needs an independent examination, a distinct authorization, or knowledge that is expensive to maintain. Removing the first kind can accelerate learning. Removing the others without understanding their purpose can make a small team dependent on abilities it does not actually have.

The Cybernetic Teammate field experiment offers evidence that some capability boundaries are already permeable. In product-innovation tasks involving 791 professionals, individuals using AI matched the performance of teams without it, and the technology helped participants combine technical and commercial perspectives. The experiment concerned newly formed teams and bounded innovation work; it did not follow those proposals through implementation and operation.6 It nevertheless makes an important organizational assumption testable: perhaps some work requires fewer transfers between functions than we have become accustomed to making.

I would begin with small groups responsible for an outcome, supported by specialists who can engage before the difficult choices have been made. A group improving customer onboarding might include operations and engineering capability, with finance and security involved in defining the commitments the service may make. It should be able to investigate, build, test, and use a limited version of the process without obtaining a new sponsor for each stage.

Specialist involvement can take several forms. A capability needed throughout the work may belong inside the group. Rare but consequential questions may be better served by a shared expert team. A repeated specialist decision might eventually become an explicit test or a bounded automated procedure. These arrangements should be selected according to the work, and revisited when the work changes. Making every team self-sufficient can duplicate expensive expertise; routing every decision through a central group can recreate the delay we intended to remove.

The center still has an important job. Common identity and access controls, reliable connections to operational records, versioned evaluation tools, and a way to account for AI costs are difficult for every team to build well independently. Providing them can let teams experiment more safely and move between models without reconstructing their entire workflow. But that infrastructure should grow from real use. A platform initiative with no team able to complete a valuable task through it has yet to demonstrate its organizational purpose.

Shared context deserves particular attention. An agent searching the company’s documents may find an obsolete policy, an unapproved proposal, and a current commitment expressed in nearly identical language. A larger context window does not establish which one governs the request. The organization needs to make the status of consequential information legible, with an owner, a source, and a way to correct it. Once that exists, it helps people as well as machines.

Independent examination also needs an explicit purpose. A meta-analysis of 106 experiments found that human–AI combinations were, on average, better than humans alone but worse than the stronger of the human or AI acting alone. Results differed substantially across tasks, and the underlying studies ran through mid-2023.7 That is a reason to test the arrangement rather than presume that adding a human approval step improves it. The reviewer needs information that can reveal a consequential error, sufficient time to investigate it, and authority to change the result.

The organization should be able to explain what each retained boundary contributes. It should be equally willing to remove a boundary whose purpose has disappeared. That is a more demanding approach than deciding in advance that everyone becomes a generalist, or that every existing specialty must preserve the same territory.

4. Give people a credible reason to change

Organizational resistance is sometimes described as a failure of imagination among employees. I would first examine what the organization is asking them to risk.

A specialist invited to automate part of their own work may be unsure whether success will earn them greater scope, reduce their standing, or remove their job. A manager may support experimentation in principle while knowing that their performance will be assessed against commitments made before the experiment began. A novice may use AI secretly because the approved route is too slow, while an experienced colleague avoids it because correcting its output creates work that no one has allowed for. These responses can be understandable under the incentives people face.

Leaders need to make the terms of participation explicit. What time is available for learning? Which current commitments will be reduced? How will a useful method be credited when it helps other teams? Where efficiency genuinely changes staffing needs, what support and transition choices will be available? Promising that every existing role will remain unchanged would make little sense in an essay arguing for organizational change. Allowing people to discover the consequences only after they have supplied the knowledge would be equally shortsighted.

There is an opportunity to distribute expertise more widely. In Generative AI at Work, a study of 5,172 customer-support agents, access to an AI assistant increased issues resolved per hour by 15% on average, with the largest gains among less experienced and lower-skilled workers. Effects on experienced workers were much smaller and included modest quality declines.8 The practical question is what allows a person to benefit: useful contextual assistance, knowledge they can apply, and a job in which the improvement has somewhere to go.

Assistance can also remove the practice through which competence develops. Shen and Tamkin’s randomized study of developers learning an unfamiliar Python library found lower immediate knowledge-assessment scores among those with AI assistance. Their short experiment does not establish a long-term effect on expertise, but it makes skill development part of the design problem rather than an automatic by-product of task completion.9

Training should therefore include unfamiliar work that the participant must understand well enough to change. Someone who can generate a prototype should also investigate why it failed when an input changed, explain the assumptions its behavior depends on, and make a correction. Someone using AI to reconcile commercial records should be able to identify which source supports a discrepancy and what evidence would settle it. The appropriate standard differs by role, but each needs an observable demonstration of understanding.

This changes the manager’s work. Coaching becomes more concerned with how a person frames a task, directs the investigation, recognizes a weak result, and handles the next case without continuous intervention. There is still a place for foundational practice without assistance, especially where understanding is being assessed. There is also a place for ambitious AI-assisted work beyond the person’s former scope. A good development plan gives them both.

Access needs to be practical. A training session cannot help an employee whose workstation cannot use the tools, whose data is inaccessible, or whose manager regards experimentation as time taken away from the “real” job. Nor should fluency be inferred from confidence in a demonstration. People need different routes into the work, including written explanations, opportunities to ask elementary questions, and assistance at the moment they become blocked.

The individual has responsibilities as well. Given reasonable access and support, they should engage with how the work is changing, test their assumptions, and make the result understandable to others. Expertise deserves respect; it should also be exposed to new evidence about what can now be done. The goal is a person whose capabilities expand, rather than someone who must preserve a particular task to preserve their value.

5. A bias for action that produces evidence

An organization cannot learn its way into a new operating model through planning alone. It needs a short path from an important question to a test that can change a decision.

For an initial effort, I would choose a workflow that matters, has a visible outcome, and can be bounded without removing the difficulty that makes it worth improving. Leadership should name an owner with enough cross-functional authority to move it, provide the relevant access, and reserve time from the people whose knowledge is needed. The first review date should be tied to a working comparison, with the current process serving as a baseline.

For a workflow with manageable consequences, thirty days could be enough to inspect representative cases, establish the current performance, and put a limited alternative into the hands of the people who would use it. A high-risk process may require a longer path to deployment, but the first month should still produce something more informative than another statement of intent. If the team cannot obtain the necessary records or make the test environment work, that is already evidence about the institutional constraint. The next leadership decision should address that constraint, rather than reset the pilot’s clock and ask for patience.

The question is what would be different at that review. An operations team might show that a request can pass from intake through resolution without repeated reconstruction. An engineer might demonstrate that a proposed change survives integration and recovery tests. A finance partner might establish whether the apparent saving survives reconciliation and exception handling. A skeptical specialist might identify why the process should remain limited to a narrower class of work. Each result should affect the next allocation of effort.

A pilot also needs a route out of the pilot stage. If it succeeds, who can authorize broader use, what must be maintained, and which duplicated process can stop? Organizations often create an enthusiastic experimental layer above an unchanged operating system, leaving employees to perform both the new process and the old reporting required to prove it happened. That burden can make an effective intervention look unattractive in practice.

Measurement must cover the whole workflow. I would examine elapsed time, work waiting between stages, quality of completed outcomes, correction effort, and the cost of models and people together. Time saved by one person is less valuable if another has to spend it repairing the result. A comparison should also record which model and method were actually used. If the deployed workflow frequently falls back to the old process, its outcomes cannot tell us much about a capability that customers rarely received.

Even measurement methods need to evolve. METR’s February 2026 update on AI-assisted software development explains how changing participant selection, task selection, and parallel agent use made later speed estimates harder to interpret than its earlier experiment. The researchers considered a growing speed benefit plausible, while cautioning against a confident estimate from the newer data.10 An organization should be equally willing to revise a measurement process that no longer describes how its people work.

To accommodate a moving technical frontier, I would keep exploration and dependable operation on different schedules. Teams need a protected route for trying new models and methods against representative work. Changes to a consequential production process need evidence that the new version preserves what already works and improves something that matters. This lets the organization investigate quickly without forcing every model release directly into a customer commitment.

What gets learned should remain useful after the experiment. WikiSkill, a research framework for agent learning, separates execution experience, accumulated knowledge, and reusable procedures; a rejected procedure can leave knowledge behind that helps a later attempt. Its evaluations concern bounded benchmark tasks.11 The organizational opportunity is to make experiments cumulative in a similarly practical way, retaining the problem, the conditions, the failure, and the reason a proposed change was accepted or rejected.

A useful shared method needs maintenance as much as discovery. Someone has to notice when the source data changes or the procedure becomes obsolete. Leadership should fund that responsibility, because otherwise a library of once-successful prompts can become another store of unexamined assumptions.

The questions should reach each level of the institution. Senior leaders should ask which commitment they are prepared to change if the evidence contradicts the current plan. Teams should ask which step in their own workflow would disappear if they were designing it with today’s capabilities. Individuals should be able to identify a part of their work they can now carry farther, together with a limitation they encountered and how they checked it. The answers provide a more useful account of readiness than confidence about AI in the abstract.

The strongest evidence of bias for action is a decision that follows the result. Expand the method that works. Repair or narrow the one that partly works. Stop the investment whose premise failed. Make the time and resources released by those choices available for the next worthwhile question.

6. The organization after the next advance

There is a tempting way to make this transition feel stable. We can assign execution to AI and reserve strategy, creativity, and judgment for people. That may describe a useful arrangement for some work today, but it would be a fragile foundation for the institution’s future.

Suppose models become substantially better at diagnosing why a workflow fails, designing experiments, or proposing an allocation of resources. Their recommendations would deserve examination on the merits, including when they challenge an experienced leader’s account. The organization would need a way to evaluate and act on that contribution. Treating the present division of work as permanent would exclude some of the very capabilities it was being redesigned to use.

The relevant unit of change could eventually become a whole operating method. A system might identify a recurring source of delay, propose a different sequence of work, test it in a controlled environment, and assemble evidence for broader use. People and other systems could challenge that evidence, define the commitments it must preserve, or reject the proposed objective. The institution’s ability to inspect and authorize change would then matter as much as its ability to execute the current process.

This future also changes what scale can mean. A larger organization might use its accumulated experience to make small teams exceptionally capable, rather than requiring each new service to reproduce a large hierarchy. A smaller organization might provide careful, customized service to customers who previously could obtain only a generic product. Some forms of expertise could become more concentrated, while their useful application becomes more widely available.

Whether those possibilities improve working lives depends on the choices made along the way. Saved time can become more volume, fewer employees, more attention to difficult cases, or a new service the company could not previously afford to provide. There will be circumstances in which each has a commercial argument. Leadership should make the allocation explicit and evaluate its consequences rather than let the easiest number to measure decide it by default.

We do not need to know the final shape of that organization to begin building its capacity to adapt. We need to know which meaningful question we will test next, who has the power to act on the result, and what the institution will change if the answer differs from its expectations.

Brand understood the difference between the clocks before returning to the ship. Preparation would have meant organizing the mission around its consequences. For our organizations, the opportunity is still in front of us: give capable people the means to investigate new possibilities, make their findings consequential, and let useful learning change the work before the old timetable makes the decision for us.


Notes

Footnotes

  1. Jonathan Nolan and Christopher Nolan, Interstellar: The Complete Screenplay with Selected Storyboards (2014), printed p. 76. Brand’s words are quoted from the screenplay; the scene follows the crew’s return from Miller’s planet. Screenplay. Back to reference 1

  2. METR, Time Horizon 1.1 (29 January 2026) and the task-completion time-horizon methodology. These measurements concern a specified task distribution and success rate. The organizational argument here does not extrapolate a doubling time or an autonomous-work deadline from the curve. Back to reference 2

  3. Alexander Bick, Adam Blandin, David J. Deming, and Tyler R. Schumacher, What Work Does Generative AI Do?, NBER Working Paper 35677 (August 2026). The cited finding is from the authors’ public abstract describing their representative occupation- and task-linked survey. Back to reference 3

  4. Eleanor Wiske Dillon, Sonia Jaffe, Nicole Immorlica, and Christopher T. Stanton, Shifting Work Patterns with Generative AI, version 4 (13 November 2025). The six-month randomized study covers 7,137 workers in 66 firms. Workplace-tool access is distinct from an experiment that redesigns the whole organization. Back to reference 4

  5. Erik Brynjolfsson, Daniel Rock, and Chad Syverson, The Productivity J-Curve: How Intangibles Complement General Purpose Technologies, American Economic Journal: Macroeconomics 13(1), 333–372 (2021). Complementary investment can precede measured gains; the framework does not establish that an unproductive initiative will eventually succeed. Back to reference 5

  6. Fabrizio Dell’Acqua and colleagues, The Cybernetic Teammate: A Field Experiment on Generative AI and Teamwork, Organization Science, published online 12 June 2026. The published study reports 791 professionals. Its task scope does not establish whole-lifecycle team replacement. Back to reference 6

  7. Michelle Vaccaro, Abdullah Almaatouq, and Thomas Malone, When combinations of humans and AI are useful: A systematic review and meta-analysis, Nature Human Behaviour 8, 2293–2303 (2024). The 106 experiments were published from January 2020 through June 2023; heterogeneous tasks and systems limit universal prescriptions. Back to reference 7

  8. Erik Brynjolfsson, Danielle Li, and Lindsey R. Raymond, Generative AI at Work, The Quarterly Journal of Economics 140(2), 889–942 (2025). Public paper record. The figures used here follow the revised 5,172-agent study, rather than earlier working-paper figures. Back to reference 8

  9. Judy Hanwen Shen and Alex Tamkin, How AI Impacts Skill Formation (2026). The randomized experiment involves 52 developers learning an unfamiliar library, followed by an immediate assessment. Long-term expertise and workplace career outcomes were not measured. Back to reference 9

  10. METR, We are Changing our Developer Productivity Experiment Design (24 February 2026). The account explains selection and measurement difficulties in follow-up work. The discussion here concerns those methodological difficulties, not a universal estimate of developer speed. Back to reference 10

  11. Liyan Tang and colleagues, WikiSkill: Compiling Agent Experience into Persistent Knowledge for Skill Evolution (2026). The proposed organizational application extends the distinction between experience, knowledge, and procedures beyond the paper’s benchmark evaluations. Back to reference 11

Back to top

IN THIS SERIES

Organizations and products in the age of AI

  1. When capability outruns the organization — You are here
  2. Product management with the means to build

THE INDEX

Find a thread.

Loading the index…