How is this "Flash"?

#10
by MasterYoba - opened

This is full size of GLM-4.7 or half of 5.2. Calling it "flash" is a very weird marketing ploy. It is nonsensical with your own naming convention (even calling it "Air" would be a stretch). It is not for consumer hardware. Don't advertise it as such. Call it GLM-5.3-Turbo or something (similar to how DiT models come in Base and Turbo versions that are usually 2x smaller).
People with consumer hardware waited for a GLM-4.7-Flash successor, not for this.

I think the naming is a bit confusing too. GLM-4.7-FLASH is a 30B-A3B model, while GLM-4.5-AIR or GLM-4.6V is a 106B-A12B model. That's way too big to be called "flash"
If you want to call it GLM-5.3-AIR, that wouldn't be wrong either, because 106B is roughly one-third of 352B, and 321B is also about one-third of 744B - so we can all just consider it one-third for now.

Yeah, not a bad flash =))) it’s like Mistral small 4=)))

Deepseek 284B is also called 'flash', chill out.

Deepseek-V4-Flash is 6 times smaller than V4-Pro. For GLM that would be 120B, not 320B.

Yes! We need a true successor to the GLM 4.7 Flash, not a fake "Flash" model that's 10x bigger and impossible to run on local hardware.
It's even 3x bigger than the Air! How are we supposed to run it when Flash was supposed to be a fast model for GPU poor local hardware?

Sorry Z.AI, but we don't have 300GB of VRAM at home. At least the new Qwen Flash is 3x smaller, you can offload it to RAM and it'll still run at a decent speed because of engrams that GLM also doesn't have...

So now a model doesn’t deserve to be called “Flash” unless it can run on consumer hardware? Lmao.

The funniest part is that the entire meltdown from people like you comes from your little fantasy not becoming reality: you want ultra-low VRAM usage AND Opus 4.8-tier performance at the same time.

In 2026, what GPU or LLM tech has actually achieved both? Seriously, name one.

Or is your argument just “I can’t run it at home, therefore the model is bad”?

And apparently Z.AI somehow owes you a model that’s insanely capable, tiny enough to fit on your consumer GPU, and open-sourced for free?

That’s not a technical expectation. That’s just entitlement.

always somebody complaining...

So now a model doesn’t deserve to be called “Flash” unless it can run on consumer hardware? Lmao.

The funniest part is that the entire meltdown from people like you comes from your little fantasy not becoming reality: you want ultra-low VRAM usage AND Opus 4.8-tier performance at the same time.

In 2026, what GPU or LLM tech has actually achieved both? Seriously, name one.

Or is your argument just “I can’t run it at home, therefore the model is bad”?

And apparently Z.AI somehow owes you a model that’s insanely capable, tiny enough to fit on your consumer GPU, and open-sourced for free?

That’s not a technical expectation. That’s just entitlement.

Everyone is saying “z.ai, thank you!” But just a part of the community and fans were expecting a smaller model and its derivative distillation. Like it was with 4.7 flash, but in the end they got a monster that can barely run on an ultra‑top‑tier user config. Essentially, people were expecting an analogue of Qwen 3.8 27B 35B‑A3B and were deeply disappointed. You also need to hear these people... No one is saying that the model is bad or that z.ai did the wrong thing by releasing it... but perhaps it’s possible to create a glm 5.3 mini 20-36b, since flash models now look like giants compared to a year ago. Still, user capacities are not growing as fast as server racks under Vera Rubin.

The funniest part is that the entire meltdown from people like you comes from your little fantasy not becoming reality: you want ultra-low VRAM usage AND Opus 4.8-tier performance at the same time.

No, nobody said that. We want SOMETHING modern that can run on consumer hardware, like GLM-4.7-Flash or Qwen-3.5 35B-a3b or Gemma. At this moment only Zai can make a good model of that size to rival Qwen.

So now a model doesn’t deserve to be called “Flash” unless it can run on consumer hardware? Lmao.

The GLM naming scheme is simple:

  • GLM x.x for new SOTA models.
  • GLM x.x Air for models around 100b
  • GLM x.x Flash for models that runs on computer hardware (GLM 4.7 Flash for example is 30b-a3b MoE)
  • GLM x.x Turbo is their "Faster than SOTA, but slower and bigger than Air".

GLM 5.3 Flash is not a Flash, but Turbo.
Our frustration is that probably this means that we won't get a small model from them again, because their models series for consumer hardware is WAY out of range of any consumer hardware.

The funniest part is that the entire meltdown from people like you comes from your little fantasy not becoming reality: you want ultra-low VRAM usage AND Opus 4.8-tier performance at the same time.

Nobody said that. Maybe your LLM has allucinated when you gave the prompt "Answer them with an argument that is impossible for them to win. Make it perfect, with no mistakes."

We just want a comparable model to qwen 3.6 35b-a3b, because they are alone in that range from months now and only Z.ai can make a great model to compete. More competition is always better, but Qwen stands alone at the top of the local hardware model tier.

Or is your argument just “I can’t run it at home, therefore the model is bad”?

Your model hallucinated again. We never said it was bad.

And apparently Z.AI somehow owes you a model that’s insanely capable, tiny enough to fit on your consumer GPU, and open-sourced for free?

They don't owe us anything, but labeling their model as one thing and not delivering it is just... messed up. They could have named it GLM 5.3 Turbo and that would be okay, but naming it the same as their consumer hardware models (their lowest-parameter models) is the same as saying that this is their new lowest-parameter model from now on, nothing below these parameters. Or nothing much below them (like a 30B MoE).

That’s not a technical expectation. That’s just entitlement.

Yeah, we already know you're using LLM to write your answers.

I, for one, am grateful for the release. Thank you Z.ai!

So now a model doesn’t deserve to be called “Flash” unless it can run on consumer hardware? Lmao.

The funniest part is that the entire meltdown from people like you comes from your little fantasy not becoming reality: you want ultra-low VRAM usage AND Opus 4.8-tier performance at the same time.

In 2026, what GPU or LLM tech has actually achieved both? Seriously, name one.

Or is your argument just “I can’t run it at home, therefore the model is bad”?

And apparently Z.AI somehow owes you a model that’s insanely capable, tiny enough to fit on your consumer GPU, and open-sourced for free?

That’s not a technical expectation. That’s just entitlement.

Thank you for this.... The level of entitlement is just outrageous. It shows how delusional the human nature can be especially given the fact that before these open source models came out, they willingly paid frontier models for lesser intelligent models.

It's small enough to run on two Sparks in nvfp4 and it's allegedly smarter than GLM 5.2. I don't see if there's an MTP/dspark/dflash drafter available but even if not, A18B means it'll be reasonably performant. If there IS a drafter this thing will fly.

This is insanely small for frontier-level AI. If you want a tiny model Qwen 3.8 was just released like a week ago and now the preview for Qwen 4 is already out.

This is insanely small for frontier-level AI.

And it's still a frontier-level AI, just like GLM-4.7 was. I'm not saying it's bad in any way. It's just not that much smaller than the base model to be designated as Flash.

I think Z.AI doesn't commit to multiple model family product lines—whether a model is labeled "Flash" or not purely depends on its relative size or position within the current series.

The GLM naming scheme is simple:

And still they seem to have another view on this than you. The truth is that they probably fokus on the best model they can build. Hardware makes it possible for them to create bigger and bigger models, when that is created they try to make a efficient destil. That is the smaller model and they choose to call it flash. I am pretty sure no one ever took time to create naming rules for their models. Just understand this generation GLM was maybe not for your hardware. Look at Qwen 3.8 Next Flash or some smaller model. Dont start arguing the names.....

Flagship, Flash, Air - seems a lot easier to follow than some other labs. I mean wtf is Luna, Terra, Sol - not intuitive. We all know Mythos, Fable, Opus, Sonnet, Haiku ONLY because the names have been around for a while. They're not intuitive in the least.

Gemini Flash? How big do you think that is? I'd bet it's easily as big as this model. No one ever said it doesn't deserve a Flash tag. I guess the main point here is: Doesn't fit on my hardware then it's not flash.

I think appropriateness of the Flash tag should be considered relative to the lab's lineup. If Moonshot dropped a Kimi-K3-Flash 500B model, that would make sense and yet still be bigger than this model.

GLM-4.7-Flash was the big thing when it released, it was the absolute best small model of it's time (before Qwen3.5), and a lot of people with consumer hardware could access practically useful local AI for the first time. That is the reason a lot of people waited for GLM-5.x-Flash and were excited for this to finally drop.

I am pretty sure no one ever took time to create naming rules for their models.

What about Claude Haiku, Sonnet, Opus and Fable? It doesn't mean anything to you? Or GPT Luna, Terra and Sol? Maybe even Mistral, Mixtral and Ministral? Gemini Flash and Pro models? DeepSeek Flash and Pro?
They all have rules about what they're supposed to be.

And also answering the other guy:

Gemini Flash? How big do you think that is? I'd bet it's easily as big as this model. No one ever said it doesn't deserve a Flash tag. I guess the main point here is: Doesn't fit on my hardware then it's not flash.

No? We never said it. If you stop and start to think, you will understand what really is the problem here.
Some years ago Claude still updated the Haiku and it was the cheapest model you could get out there. There was no chinese cheap model, DeepSeek r1 was not even released. Chinese and other AI labs cheap models were genuine trash. Garbage.
So when Haiku 3.5 was announced with almost Sonnet capabilities, the hype was REAL. IT would be like to get the power of the sun, now fitting in our wallets. But then it released.
It was 400% MORE EXPENSIVE than the previous generation. The cheap model of the Claude family wasn't cheap anymore. People hated it, I hated it, everyone hated it.
Why? Nobody was saying "Leave my multibillionare AI Lab alone, you are just sad because you're poor and can't afford it" as some users are saying here in the comments, everyone knew it was supposed to be the cheap model and if the cheap model is expensive, it would be the end of us.

The family of Claude Haiku is just like Flash for GLM. But instead of being accessible for our wallets, it's supposed to be the accessible for our consumer hardware.
And note, we're not raging out, calling them out of something and boycotting them. No. We're just saying that it shouldn't be named "Flash", but "Turbo" or stretching a bit, "Air".

If the Flash model FOR GLM is supposed to be the model that can be run in the consumer hardware, as GLM 4.7-Flash was. Then it isn't Flash.
As I wrote some comments ago:

"The GLM naming scheme is simple:

  • GLM x.x - For new SOTA models.

  • GLM x.x Air - For models around 100b

  • GLM x.x Flash - For models that runs on computer hardware (GLM 4.7 Flash for example is 30b-a3b MoE)

  • GLM x.x Turbo - For models that are 'Faster than SOTA, but slower and bigger than Air'.

GLM 5.3 Flash is not a Flash, but Turbo."

Edit: Spelling correction

just give us a modern GLM 4.7 flash-KINDA model: 4.7 flash was 80% of 4.7 in every benchamrk in days when model size mattered more than reasoning quality and RL-training, if GLM does it now - it surely will score >85% of full model, which will make it the absolute DESTROYER of Qwen 3.6 35B A3B. example: GLM 5.3 DeepSWE 1.1 score = 66.9%. Then GLM 5.3 air (let's name it air) = 55% (let the multiplier be 0,82 perhaps). which is Opus 4.8 tier, which scores 59%. So we have 32B A4B (perhaps) which runs on 5070 + 32GB ram and rivals 120-200 USD API models.

Sign up or log in to comment