• 0 Posts
  • 11 Comments
Joined 1 month ago
cake
Cake day: August 10th, 2026

help-circle

  • I think it’s a little disingenuous to imply that deepseek is a small model. Even flash requires 160gb of vram and it’s important to note that’s a size aimed at a constraint, the vram density of hbm equipment from two generations ago. (And the density allowed by using last generations consumer equipment plugged into ten amp circuits with doubled vram see all those 48gb 4090s floating around!)

    They didn’t arbitrarily decide that they’d use a smaller model, they were targeting a constraint. Which of course was really juicy because of all the demand for text inference and the limited new hardware to fulfill it.

    The paper you cite is true, I haven’t read it but it based on the title it can’t really be wrong unless they just get wildly over their skis with their claims…

    You’re right that everything’s gotta walk through the layers and more layers means a longer walk. That’s limited by memory and interconnect bandwidth though which is still doubling or close to it every year or close to it.

    Which means it would have to be twice as fast to use a smaller model on cutting edge hardware to be a real “wall”.

    We are not near the end of the memory bandwidth road yet, quantum tunneling isn’t rearing its head again as the workaday wafers that ferry serialized streams from place to place get faster and more numerous. Because they don’t need to get smaller really. The port of New York can grow and sprawl and sprout heretofore unseen support structures like coolers and voltage regulators and coolers for its voltage regulators.

    I’ve seen in action what you describe though. On a, sensible chuckle, tiny card like a 3050 slotted into a system with pcie3 very small models run faster than those who can barely fit in the vram with their little bitty context, but that difference shrinks significantly when the card is a 5060 or something that still has limited vram but can load faster due to a faster pcie interconnect.

    Now that may seem like apples and oranges because it’s two completely different things but the point of the comparison is to show that the difference between performance measures in time to first, tps or whatever other measurement might be in vogue at the moment when two models are compared shrinks when the interconnect and memory bandwidth gets faster.

    To butcher a car metaphor, you can turbocharge an ls and get more power but you’re not escaping engine wear=crankshaft rotations. That might not be as butchered as it first seemed even though the domains are all shifted.


  • You said there’s no reason to think you can just keep making the model bigger and keep getting improved capability.

    there’s the structure of the neural network itself. Fundamentally, adding nodes and layers increases the ability of the model to handle more complex input.

    Then there’s the actual models we see in use. They are literally as large as the hardware allows. The only reason to use smaller models are to fit some constraint.

    So both by the book and in practice bigger is always better.

    Now we can’t always go big. I can’t afford to purchase a dgx or even upgrade my wiring to power it, let alone pay the power bill it would rack up or all the other utilities alone when my wife leaves me because of the sound.

    My computer can only fit so many expansion cards and pcie is so slow compared to hbm that I’m better off running a small model quickly that fits on one card as opposed to a larger one slowly across several cards.

    But those are all constraints. When I replace my motherboard with supermicro gpu host fabric I no longer am limited by the pcie bandwidth and can quickly use models that fit across several cards.

    I do agree with you that the future is smaller models, not because of the fundamental nature of the concepts involved but because of the complex constraints that are coming into play.






  • Hey, I think ai can be a useful tool and I use both cloud hosted and local models.

    The ecological impact is god awful.

    Datacenters run on fossil fuel turbines to meet demand when there isn’t enough capacity on the grid or just 24/7 because the residents of the nearby municipality said they don’t want brownouts.

    Fundamentally ai datacenters transform expensive to transmit electricity into cheap to transmit data, financializing what was previously a relatively insulated commodity whose pricing and production reflect local demand into a globally traded commodity whose pricing and production reflect global demand.

    If you don’t have a background in the history of industry and environment: that’s bad news for the environment (and anyone who relies on stable power rates).

    Production of ai computing components requires mining rare earth minerals and phenomenal amounts of power. This is the same process with nearly the same inputs as the production of general purpose computing parts (if you leave off the increased rare earths needed to mitigate electromigration in single digit nanometer processes!) but the increased demand has already brought more fabs up to meet it meaning more use of resources and energy in production before the first kilowatt of inference is run on those chips.

    Turbine blades for datacenter generators are also limited by production and are currently years out for delivery. These are components produced by advanced aerospace manufacturing shops which use massive amounts of energy to shape alloys of metals that have to be yanked out of the ground and smelted into some esoteric type of steel or aluminum first.

    So even without the water concerns datacenters have tremendous impact on the environment.

    But the water concerns are legitimate. Swamp chillers divert flowing water into the air where it is blown anywhere but where it would have flowed. Which is fine if you aren’t downstream of one I guess.

    All this is to say: if you don’t care about the impacts you can rationalize that however you like, but they exist.