The Efficiency-Throughput Gap with GitHub Copilot

(cacm.acm.org)

7 points | by fscaramuzza 1 hour ago

4 comments

  • aurareturn 52 minutes ago
    Note the timelines of this study:

    4/2024: Baseline metrics

    8/2024: Participants given Github Copilot licenses

    11/2024: Conducted surveys, 261 people invited and 97 responded in survey

    9/2026: Study published

    • erikgahner 46 minutes ago
      Setting aside issues with the sample size and response rate, the current progress within the field shows how difficult it is to study the impact of AI on productivity.

      Even if the survey was conducted one year later, I would not find it useful to make inferences about the use of AI tools in September 2026.

      • aurareturn 28 minutes ago
        What's very weird is that this study was published by two people working professionally at Okta. Yes, the auth tech company.

        This is the title they chose for the study:

          Beyond the Hype: The Efficiency-Throughput Gap with GitHub Copilot
        
        How can you possibly have a title like that when the study was done in 2024? I'm guessing even their own engineers at Okta would roll the eyes at this study.
  • thevinter 53 minutes ago
    GitHub Copilot is an incredibly limited tool compared to any real harness, and I'm so tired about all of these studies that claim a lack of efficiency while at the same time doing everything in their power to shoot themselves in the foot.

    Also, it's insane to me that this whole article has been written without specifying anywhere the models being used.

    Like, of course you're not gonna be productive if all you have is Sonnet 4.6???

    • jbjbjbjb 38 minutes ago
      I think you’re missing that it’s not only about coding. That’s just one aspect of the work.
      • thevinter 28 minutes ago
        Claude Cowork is also a better harness than GitHub Copilot. Many such cases. I never mentioned coding being the only goal.
    • njaa 38 minutes ago
      Considering the timeline, it would have been Sonnet 3.5
  • sublimefire 33 minutes ago
    > we found no immediate increase in key engineering metrics such as monthly pull requests and lines of code

    > To establish a before-Copilot baseline, we used data from April, May, and June 2024. After-Copilot data was represented by the period of September, October, and November 2024

    > GitHub Copilot usage varied significantly among engineers, the tool demonstrably fostered positive changes in perceived engineer value, reduced time spent on various engineering activities, and boosted motivation and perceived skills

    > subsequent monitoring of PRs and LOC for participants from December 2024 to May 2025 showed no statistical improvements

    > the implications of more advanced capabilities, such as retrieval-augmented generation (RAG) over enterprise codebases or deeper engineering workflow integrations, warrant separate investigation

    I think it is just out of date, habits have changed as well. I have not seen much gain personally at that period except in the last 12 months. Also, models not named, token counts not shown. Not to mention it was the older autocomplete + chat that were in use, these days it is much more advanced with RAG, cli use, MCPs, etc.

  • henrydoughty 37 minutes ago
    [dead]