Companies & AI
Microsoft “Dogfooding” Copilot Was Not The Best Way To Enhance Adoption
I’m reading the second article in my journey of deep diving on the science behind technology and AI adoption. One of the key challenges with research, experiments and the resulting published articles on the topic of AI is of course the crazy speed of developments. Meaning that experiments and results from four years ago are with a completely different level and quality of AI than we are seeing today.
Still I think it’s important to sift through these articles to see what insights we can get from them. Even if sometimes it’s just to question methodologies used since it can enhance future research approaches and insights (yes, a small spoiler alert of one of my criticisms).
Anyway, let me give you some context about the article and dive in.
The article describes three experiments run at Microsoft, Accenture and an anonymous company, where developers were invited to start using Copilot (back then just a tool to autocomplete coding and give suggestions). The goal was to figure out if the use of generative AI, in the form of copilot, would enhance productivity.
My core criticism
Now we immediately get to my core criticism. Productivity. As a concept, it’s literally how productive developers are. Defined in the article as the number of so-called Pull Requests (when you have built code and you are ready to publish it to your application you create a Pull Request (PR) and after review, manually or automated, that code is pushed to production and part of the live software product).
In my own AI building and coding journey as a solopreneur, I run a gazillion PR’s. Am I the one defining the size and timing of the PR’s? Nope. Is AI often suggesting to split tasks into multiple PR’s. Yep. So I was very surprised to read they used this as their key indicator because in my view it’s telling me nothing about productivity (let alone quality).
Granted: this was 3-4 years ago, so there were no agents running code and suggesting PR’s. And these are large corporations and not solopreneurs, so there were manual code review procedures in place that could assess the quality of the PRs. And the authors assess this shortcoming themselves in the article and state that the quality of code is very hard to measure.
Zero surveys
The three experiments show some significant productivity gains (26%) but the data is pretty noisy as they reveal in their summary. That’s fine. It’s such a new field, that the research will often be noisy and challenging. What I’m mostly critical about, is their complete lack of human experience measurement. They only use hard data. Zero surveys toward the almost 5.000 developers that participated in the pilots.
If they would have asked them, they could have assessed:
- If autocompletion and suggestions were a helpful innovation
- If no training on the use of Copilot was costing them more time to figure it out than to just write the code themselves (only Accenture had done training and saw better adoption)
- If at home they were using better AI than Copilot and if this was frustrating them
- If they felt they were building better solutions for customers
- If they felt they were building faster and what they were doing with the time gained
That’s only in hindsight. Valuable, but it could have been even more valuable.
If they had done a survey beforehand, they could have tested even more impact. They could have discovered the daily actual pains and struggles of the developers and connected Copilot use cases to those pains. That for sure would have enhanced adoption and the impact the tool would have on their work.
In the Microsoft case, the developers in the test group received this email below.
I’m sure you can understand why only 8% of developers initially responded to join.
Tap to open the full size image
Design for the experience of your employees
So often companies are overly focused on productivity without quality of the output.
Are you still tracking hours worked of your employees or the speed and quality of the output they are delivering for your organisation?
Same issue.
Let alone take into account the experience of both employees as well as customers.
Remember that super shitty customer service chatbot you recently experienced?
That there was no way to actually talk to a human, despite “Botty” being super nice?
Or all these employees who need to use an AI solution in their organisation that is so much worse than what they can use at home?
I know, a corporation can not pivot and use tools as fast as entrepreneurs can.
But they can at least be honest and acknowledge the challenges and not blame the employees for not using the AI while the only “adoption” approach was the centralized e-learning on the risks of sharing confidential information in the corporate AI.
You have to design for the experience of your employees.
Their daily work, their pains, their level of knowledge of using AI, etc.
This is the part where application of technology becomes transformation of the organisation.
Article reviewed
- Cui, Z. K., Demirer, M., Jaffe, S., Musolff, L., Peng, S. & Salz, T. (2026). The Effects of Generative AI on High-Skilled Work: Evidence from Three Field Experiments with Software Developers. Management Science. In the library