Tagged "multi-gpu-inference"
- llama.cpp b10549: Tensor Parallelism Support for LFM2/LFM2MOE Models
- Building a Dual V100 AI Workstation for Local LLMs
- K3 Model Achieves 20 Tokens/Second on 80x RTX 5090 Cluster
- RTX 5080 and RTX 3090 Setup Achieves 80 Tok/s on Qwen 3.6 27B Q8
- Pluggable's TBT5-AI: First Thunderbolt Dock Explicitly Targeting Local LLM Workstations
- This External GPU Enclosure Tries to Break Cloud Dependence for Local AI Inference
- Running Qwen3.5-27B Across Multiple GPUs Over LAN Achieves Practical Speed for Local Inference
- Qwen3.5-397B Achieves 282 tok/s on 4x RTX PRO 6000 Blackwell Through Custom CUTLASS Kernel
- Llama.cpp Celebrates Major Milestone: From Leak to Industry Standard
- Qwen3.5 122B Achieves 25 tok/s on 72GB VRAM Setup
- Qwen3-Next 80B MoE Achieves 39 Tokens/Second on RTX 5070/5060 Ti Dual-GPU Setup