Absolutely not you can check my own models Ivme-Conversate-S-v1-Base and S-v2-Instruct if you want I hate to do unrelated mentioning but
They literally do what you say
Low training tokens, different architecture (for v1), high quality data and it flopped so bad I didn't put v2 to leaderboards.
Second you said SmolLM2 uses curated data
Well
I guess what do we use
Take a wild guess it's FineWeb-Edu!
Why?
BECAUSE IT'S HUGGINGFACE
THAT'S THE POINT OF HUGGINGFACE
Also you said things about water usage I can tell you BananaMind's model is definitely more efficient in the amount of compute spent on what hardware.