Instructions to use bottlecapai/ThinkingCap-Qwen3.8-27B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use bottlecapai/ThinkingCap-Qwen3.8-27B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="bottlecapai/ThinkingCap-Qwen3.8-27B") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("bottlecapai/ThinkingCap-Qwen3.8-27B") model = AutoModelForMultimodalLM.from_pretrained("bottlecapai/ThinkingCap-Qwen3.8-27B", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use bottlecapai/ThinkingCap-Qwen3.8-27B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "bottlecapai/ThinkingCap-Qwen3.8-27B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "bottlecapai/ThinkingCap-Qwen3.8-27B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/bottlecapai/ThinkingCap-Qwen3.8-27B
- SGLang
How to use bottlecapai/ThinkingCap-Qwen3.8-27B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "bottlecapai/ThinkingCap-Qwen3.8-27B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "bottlecapai/ThinkingCap-Qwen3.8-27B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "bottlecapai/ThinkingCap-Qwen3.8-27B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "bottlecapai/ThinkingCap-Qwen3.8-27B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use bottlecapai/ThinkingCap-Qwen3.8-27B with Docker Model Runner:
docker model run hf.co/bottlecapai/ThinkingCap-Qwen3.8-27B
GSQ-RCO quant for ThinkingCap-Qwen3.8-27B?
Hi bottlecapai,
Could you release a GSQ-RCO quant of ThinkingCap-Qwen3.8-27B, or coordinate with ISTA-DASLab to do it? Their GSQ-RCO method keeps near-BF16 quality at ~2–3 bits, cutting the model to ~1/5 size. For low-VRAM users, that’s the difference between needing 60GB+ / multi-GPU and running on 12–16GB consumer GPUs. Combined with ThinkingCap’s ~46% fewer thinking tokens, it’d be much faster and locally usable. Your model already quantizes well, so it’s a natural fit.
Thanks!
Hi @Exo87 ,
it would be great to combine. I will investigate it but no promises 😇
Out of curiosity, what is your use case? How you used models from ISTA-DASLab ?
Aren't you limited by slow response time anyway despite ternary quantizations?
Hi @oplatek , thanks for looking into it!
My setup: RTX 3060 12GB VRAM. So yes, I'm definitely VRAM-limited, but the bigger problem is actually speed and token burn.
Use case: I take existing GitHub projects and modify them to fit my specific needs. The base Qwen3.8-27B thinking model is the issue, it burns so many thinking tokens per task that it either runs out of context before finishing the project, or takes forever. So I'm stuck choosing between slow-and-incomplete local, or fast-and-complete but cloud (DeepSeek), which I'd rather avoid.
On ISTA-DASLab: I haven't run their quants myself yet, I found their GSQ-RCO work and it looked like the missing piece. Ternary/low-bit quants usually hurt quality, but their results suggest you can get the VRAM savings without the accuracy hit, which is exactly what a 12GB card needs.
Why your model matters here: ThinkingCap's ~46% token reduction directly attacks my real bottleneck (thinking length / context exhaustion), and GSQ-RCO attacks the VRAM side. Together they'd be the first combo that actually fits my workflow locally, smaller and smarter about tokens. That's why I'm pushing for it.
No pressure on the promise, just wanted to plant the flag. Thanks for the great work either way 🙏