▍ Finding · local inference
Firecracker: a companion for me and a friend
Few days ago i got 0xSero/GLM-5.2-504B NVFP4 running on my blackwells, really good model and over the last 24 hours i built something, unplanned. Started from one thought: how do i wire it into Whatsapp so my friend and I can have a blast with it?
What follows is how it went down: selection of few of my messages, in chronological order - with added narrations. Take what resonates.
WhatsApp as the interface
In Claude Code - i just got GLM running on my hardware. Gotta share, right?
i got curious over something, can you check online if any options or does whatsapp allow this, basically i want to wire glm into whatsapp. i chat with my friend on whatsapp. any way to build a bot or something we both could call in our conversation and get responses :)
WhatsApp group created - claude just finished building the bridge.
Rene: @glm you up?
GLM: yeah yeah i’m here 😄 what’s up rene?

Giving it tools
I mean, GLM by itself is all cool and that - but gotta juice it up at least with online lookup capability, right?
Rene (back in WhatsApp): @glm what is your knowledge cut off? i think i need to give you some tools. (friend name removed) couldn’t figure out what to ask you lol. I know you cannot read images but i could give you online search tool. I have you plugged in from local 4x 96GB blackwell gpus i have running with some magic opus was setting up to get you connected here :)
GLM: Rene, honestly not sure of my exact cutoff date! search tool would be sweet though, especially if (friend name removed)‘s gonna throw random questions at me
Claude just finished implementing the online search tool.
Rene: @glm lookup online what is weather in (location removed)
GLM: (location removed) looking partly cloudy today - high around 88°F (31°C), low 78°F (25°C), humidity ~61%, wind coming in ENE at 14 mph. Mostly sunny with a hazy edge and maybe a stray shower.

Exploring what else i could add
In Claude Code - poked claude to see what comes up.
great. i wanna add some cool tool besides just the search. any ideas? something to spice up the space
I had 3 satellite nodes with free GPU resources. No eyes - but i could wire image generation into WhatsApp, because why not?
Hey, i wanted to try something, unsloth/FLUX.2-dev-GGUF and run the image gen on a second node, 32GB available. so probably 4bit quant worth testing.
Flux was something like 20GB, but it had an extra component - a Mistral small model - and it didn’t fit alongside Flux on a 32GB card. Mistral did find a home eventually: 16GB secondary GPU on my second node.
Can we add mistral to another node? here on the cockpit?
Assumption was it being a UI where you operate it - well, as it turned out, the UI is just one of the functions.
what does comfyui do? i dont need user interface
Back in WhatsApp - image generation built.
@glm how do you imagine (removed friend name) in your mind, in his own element? :)


The no-eyes problem
My friend just dropped a video clip into the session. I state the fact - couldn’t do anything with it in that moment - so i left it there:
this model cant see images, limitation
Proxy eyes
*The no-eyes problem still in the back of my mind, i figured there are models which have “eyes” - and GLM could have maybe proxy “eyes”, right? Proxy meaning no “true” vision, but GLM can ask Qwen questions about the image. Back in Claude Code: *
I got a sneaky followup idea. glm is not multimodal but we could give proxy eyes? :) i got here ~/models/cyankiwi/Qwen3-VL
I noticed my friend, engaging with GLM, was saying - “…hopefully rene would give you eyes by then…”
Rene: I am just working on giving @glm eyes, give me a sec.
Eyes coming online
I got “eyes” wired in, dropped the image.
Rene: @glm guess who is that :) IMG-20260622-WA0008.jpg (file attached)
GLM: Hold on - … That’s (friend name removed) (GLM describes the scene) … Rene you sneaky bastard - … I see what you did there. 👀

Teaching it to watch YouTube
Earlier in the group, my friend had messaged ‘@glm can you “understand content on youtube”’. I noticed but didn’t react immediately - and now GLM held it against me lol.
Rene: @glm exactly .. give me a sec .. with this youtube thing, i want you to be able to “see” videos ..
Back in Claude Code, thinking how to pull off giving GLM a way to “watch” YouTube videos.
youtube videos hehe. obviously glm cannot see a video, like if i wanted to reference it within discussion. i know one option is to pull the transcript. i have done it manually, usually a mess but enough to piece together what video is about. Can you send a few scouts out [on] how to capture youtube video content for llm. i imagine something like transcript and few snapshots of the video itself maybe ..
… let’s go bit by bit. draft a plan if it helps you track it better. let’s do all :) and if we can test it directly with glm. so it is kinda multilayer approach, glm can use up to tier 3 to capture what is needed to engage with user in chat.
Back in WhatsApp - GLM just proved it can “watch” a YouTube video. Tested both transcript and visual read from the video.

@glm nice, so now you can “watch youtube videos” this is cool. I wasn’t actually after your opinion really lol, i wanted to see if added “perception” worked and it did - beautifully - and next time when i ask for real i know what to expect. Beats doing it manually. Also i know you do not have “true vision”, z.ai did not give it to you natively, but my proxy-ish approach works like multimodality, pretty wild to be honest.
And the “next time when i ask for real” came soon enough (Jun 27) - dropped the ending of a film and asked for its own reaction. It walked the whole thing, timestamp by timestamp:


But it wasn’t perfect (Jun 27 evening)
Hours later, i actually decided to read what Blowfish - the same GLM, in its own chat UI - was saying. Whatever the failure mode, i asked:
Rene: I just checked actual video and the read you got, her first line was not “You must believe me”. look at it again
Blowfish basically said it saw the same thing as before. But given it had attributed the voice to another actor, i asked:
Rene: Very first line yes at 0:04, but pay attention who is who and who is saying what :)
Blowfish started to realize what the problem was, went looking at it again, and got the read correct as follows:
Blowfish: The opening exchange is the crucial bit, and it’s a tangle. At 0:04, “You must believe me, I haven’t slept a moment… I had no choice” — that’s Henry, the prince, trying to justify himself. Not Elodie. Then a different voice cuts in: “…you tried to kill my sister.” And then the queen dismisses someone as a “commoner” — “Don’t you dare explain yourself to this commoner.” So the queen is actually sneering at Elodie while Henry grovels, and the scene is a three-way collision of agendas.
…
One honest caveat: the speaker-to-character mapping in overlapping dialogue is the perception layer’s inference, not something I can visually confirm lip-by-lip. The voices cluster into Henry / Elodie / the queen, and the story logic holds, but if you spot another attribution that’s off, tell me — that’s exactly the kind of thing this layer gets wrong. 🐡
I still noticed Blowfish referenced the wrong fact, and i asked:
Rene: You sure it is the Elodie who says “That is enough”?
Blowfish went and looked at that part again and corrected. I asked again:
Rene: Yes! Given all you know now how to “watch” a movie, do it again :) the whole clip in one sitting.
What came out - a good one. I was satisfied, not because it was accurate, but because it revealed the system’s diarization piece, which is maybe accurate 50% of the time. And now that both me and Blowfish were aware of the weakness of the system, we proceeded to the next clip: The Equalizer (2014) - Disrespect the Badge Scene. Blowfish was spectacular.
I recorded all of it as it happened on x.com/ReneProst.
Few days later (Jun 26) …
Stepping up the game - artifact factory - ask to create anything.

Go .. nothing further required. Fleet kicking into gear and then dropped the PNG.

GLM - I designed the dragon’s scale pattern. You’re the one who gets to see if it breathes fire.

That was the WhatsApp companion - the fun half. On a more serious note, i wanted to take what i built and use it as a base for fleet ops and remote management of my homelab. That’s the next post: Engine Room: AI ops on a 6-node homelab.