Skip to content
View palmfuture's full-sized avatar

Block or report palmfuture

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
palmfuture/README.md

πŸ‘‹ Wisit Srimala

Tinkerer Β· Self-hoster Β· Local-AI enthusiast

I run a small AI lab on the desk next to me, quantize open models in my spare time, and print the parts I need on a 3D printer I can't stop upgrading. Based in Bangkok, Thailand πŸ‡ΉπŸ‡­

GitHub followers Hugging Face


🧠 What I'm into

  • Running LLMs locally β€” I operate a 2Γ— DGX Spark build running DeepSeek V4 Flash (TP2), and I care about squeezing maximum tokens/sec per watt.
  • Quantization & model packaging β€” I take open weights and repackage them into GPTQ / FP8 / NVFP4 formats that actually run well on consumer & edge hardware.
  • Self-hosting & infrastructure β€” Docker, Nginx, Go services, and glue code that makes small clusters behave.
  • 3D printing β€” Bambu Lab A1, PETG preferred, currently printing parts for my own hardware projects.
  • Building small, boring, reliable tools β€” CLIs and utilities that do one thing well (and I use daily).

πŸ“¦ Hugging Face β€” my main contribution

I publish quantized models to Hugging Face β€” collectively ~300k downloads and counting.

Model Format Downloads Likes
Qwen3.6-35B-A3B-GPTQ-Int4 GPTQ-Int4 282,964 29
Qwen3.6-27B-GPTQ-Int4 GPTQ-Int4 13,580 1
Ornith-1.0-35B-NVFP4A16 NVFP4-A16 1,773 2
Nex-N2-mini-NVFP4A16 NVFP4-A16 290 0
gemma-4-12B-it-FP8 FP8 55 1

πŸ‘‰ Full list: huggingface.co/palmfuture

πŸ”§ Interesting projects

  • vllm-default-thinking-budget β˜…2 β€” Injects a default thinking_token_budget and presence_penalty into vLLM, fixing the gap where models lose reasoning behavior by default. Shell.
  • More experiments, notes, and one-off utilities land here when they're worth keeping.

πŸ› οΈ Tech I work with

Go JavaScript TypeScript React Node.js Docker Python Arduino

πŸ“Š GitHub stats

GitHub Stats


Home lab: powered by ❀️ and a fully-loaded PSU
πŸ“« Reach me: GitHub Β· Hugging Face

Popular repositories Loading

  1. Example-Web-Server-with-Node.js Example-Web-Server-with-Node.js Public

    JavaScript 3 5

  2. vllm-default-thinking-budget vllm-default-thinking-budget Public

    Inject default thinking_token_budget and presence_penalty for vLLM, fixing the gap where --override-generation-config doesn't propagate these fields. Prevents Qwen3 thinking-mode infinite loops.

    Shell 2

  3. PSU-auth-soap-example PSU-auth-soap-example Public

    PHP 1

  4. Cosmetic-Redux-HOC-Router Cosmetic-Redux-HOC-Router Public

    JavaScript 1

  5. NB-Iot-Arduino NB-Iot-Arduino Public

    Arduino Uno R3, NB-IOT Project

    C++ 1

  6. Template-Expo-ReactNative-Redux-Navigation Template-Expo-ReactNative-Redux-Navigation Public

    JavaScript 1