Can gzip be a language model?

(nathan.rs)

73 points | by networked 2 hours ago

9 comments

  • GodelNumbering 0 minutes ago
    3blue1brown did a series on this topic: https://www.youtube.com/watch?v=l6DKRf-fAAM https://www.youtube.com/watch?v=GlYgs6v2YfU (i think one more is yet to release)
  • Culonavirus 55 minutes ago
    This tracks perfectly with Winrar being more profitable than OpenAI... coincidence? I think not!
    • wolfi1 41 minutes ago
      winrar is profitable? sure? well, on the other hand, they sure don't make losses
      • shezi 12 minutes ago
        They are a German GmbH and must publicly state their financials: https://www.northdata.de/win%C2%B7rar%20GmbH,%20Berlin/Amtsg...

        Looks pretty profitable to me.

        • amiga386 2 minutes ago
          They're one of the few companies that actually manage to sell "boxed software" (i.e. has not changed much in years but new customers keep buying it)

          That said, Windows users should use 7-Zip. Better compression format, unpacks more kinds of archives

        • jurgenburgen 2 minutes ago
          That’s surprising. Seems there is a niche for everything.
  • montebicyclelo 7 minutes ago
    This is fun, but historically people have gone a bit overboard with saying that models like this, or n-gram language models, are anywhere close to large neural network models. There is certainly a connection though.
  • tromp 25 minutes ago
    I'm more interested in the converse question: how well does an LLM perform as a compressor, compared to gzip (ignoring its insanely lower speed)?
    • gkbrk 21 minutes ago
      Top contestant in the Hutter Prize uses a neural network for compression. So fair to say, LLMs would perform pretty well compared to gzip.
    • asdfsa32 21 minutes ago
      lossless vs lossy is the question.
  • Tornhoof 29 minutes ago
    Previous discussions of that specific page https://news.ycombinator.com/item?id=48557691
  • mentalgear 9 minutes ago
    Interesting approach, I wonder how this could be used as a classifier. :)
    • Sesse__ 2 minutes ago
      I've used LZO as a spam classifier on chat. Spam tends to be very content-less and repetitive...
  • mg 53 minutes ago

        give it a normal text prompt, and it
        continues that prompt by searching
        for the byte sequences that compress
        best.
    
    One moment, how are we supposed to know how well that search was done? There is no way to search a meaningful part of the search space.

    So the result only gives us some lower bound of how well gzip works as a "plausibility tester" of a continuation of a text. The space of possible sequences is many orders of magnitude larger than what was searched. So there might be sequences in there that compress much better.

    The text mentions beamsearch, but I don't see a discussion about how well beamsearch performs in finding the global optima when it comes to gzip compressibility of a text?

  • 0x20cowboy 8 minutes ago
    .
  • bob1029 39 minutes ago
    Not without attention or something approximating it.

    The fact that gzip is relatively fast should be your first clue that something important is missing.

    Gzip is great at predicting the next token for one very specific narrative. LLMs can predict next tokens for entire universes of narratives. Searching for the correct next token across this space scales ~quadratically with the input size. Gzip scales linearly. I can gzip a one terabyte file. Imagine feeding that much into an LLM. These are wildly different animals that happen to overlap in a very small way. Equating compression to intelligence looks increasingly silly to me.

    If we must compare language models to compression, they are much more like jpeg and mp3 than they are gzip and flac. I can go fuck with a jpeg file pretty severely at the bitstream level and still have something resembling performance on the other side. Gzip cannot remotely approach this.

    • fedeb95 4 minutes ago
      I agree, but also equating LLMs with intelligence is wrong.
    • Retr0id 28 minutes ago
      > Gzip scales linearly. I can gzip a one terabyte file.

      In part because gzip only has a 32KiB window size, and I think it'd be at least quadratic within that window if you were going for optimal compression.

      • bob1029 7 minutes ago
        I'll concede the window part, but Gzip runs within the physical confines of a single cpu core and is typically entirely resident in local caches. The point is not just the quadratic scaling but also what it scales with.

        Show me an LLM that can run at 300 megabytes per second. Even dedicated ASICs with weights burned in will never move this fast.

      • Sesse__ 21 minutes ago
        Match-finding does not need to be quadratic. However, truly optimal gzip block splitting is very slow, indeed.
    • amelius 36 minutes ago
      Perhaps a better question is if LLMs are used as compressors, how well is that expected to work.
      • magicalhippo 26 minutes ago
        > if LLMs are used as compressors, how well is that expected to work

        Quite well. This project[1], by Fabrice Bellard of ffmpeg fame, is quite old in AI years and uses an ancient LLM, but still beats xz by a solid margin.

        [1]: https://bellard.org/ts_zip/

      • Retr0id 27 minutes ago
        Extremely well, aside from speed.
        • bob1029 23 minutes ago
          > aside from speed.

          And energy consumption.