Tagged "memory-bandwidth"
- llama.cpp Improves CUDA Performance with Kernel Fusion
- DeepSeek V4 Flash Optimized for Single AMD MI300X GPU
- Kioxia UFS 5.0 Embedded Flash Memory Enables On-Device AI with Advanced Storage Architecture
- Kioxia's UFS 5.0 Embedded Flash Enables Practical On-Device AI
- SK hynix 3D-Stacked DRAM-on-Logic Architecture Could Solve On-Device AI Memory Constraints
- AI Inference is Rewriting the GPU Buying Playbook
- On-Device AI Ignites WAIC 2026: How Compute-in-Memory Chips Are Stuffing 100-Billion-Parameter LLMs Into Your Pocket
- AMD Ryzen 7 7700X3D Linux Performance Review
- Apple's MacBook Lineup Overhaul Features M7 Chip for Enhanced Local AI
- On-Device AI Technology Emerges as Key Growth Driver for Hardware Makers
- Theoretical Bottlenecks for Scaling LLM Inference to Achieve Higher Token per Second
- Apple's M7 Chip Delivers 56% Memory Bandwidth Increase for On-Device AI
- NVIDIA DFlash Block Diffusion Accelerates Autoregressive LLM Inference
- Boost Inference Performance up to 15x on NVIDIA Blackwell Using DFlash Speculative Decoding
- Samsung Unveils UFS 5.0 Storage Optimized for On-Device AI Applications
- Samsung's UFS 5.0 Addresses Critical Memory Bandwidth Bottleneck in Mobile AI Inference
- Longsys Redefines On-Device AI with Groundbreaking Edge Memory Solutions
- Tether AI Upgrades QVAC SDK With TurboQuant for Data Center-Sized Memory on Everyday Devices
- Samsung's Exynos 2800 Brings HBM Memory to Mobile AI, Enabling Faster Local Model Inference
- M5 Max MacBook Runs Local Large Language Models Efficiently
- Samsung's Exynos 2800 Could Be the First Mobile Chip to Use HBM for Powerful On-Device AI
- Samsung's Exynos 2800 Brings Significant On-Device AI Capabilities
- Unweight: Lossless MLP Weight Compression for LLM Inference
- OpenUMA – Apple-Style Unified Memory for x86 AI Inference
- Google's TurboQuant Shows Memory Constraints Remain Critical for Local LLM Inference
- Samsung Galaxy Book6 Brings Consumer-Grade On-Device AI Hardware to Market
- M5 Max Delivers 1.7x Faster Inference Than M3 Max on Qwen 3.5 Models
- Snapdragon 8 Elite Gen 5 Hands the Galaxy S26 the AI Upgrade We've Been Waiting For
- Cutile.jl Brings Nvidia CUDA Tile-Based Programming to Julia
- SK Hynix Completes Qualification for LPDDR6 Memory Optimized for AI Inference
- M5 Max and M5 Ultra Chipsets Demonstrate Significant Bandwidth Improvements for Local LLM Inference
- SK Hynix Develops 1c LPDDR6 DRAM to Boost On-Device AI Performance in Mobile Devices
- The Emerging Role of SRAM-Centric Chips in AI Inference
- Apple Unveils MacBook Pro with M5 Pro and M5 Max Featuring On-Device AI
- Apple Unveils MacBook Pro With M5 Pro and M5 Max for On-Device AI
- Snapdragon 8 Elite Gen 5 for Galaxy Official: 5 Key Improvements that Push the Boundaries
- DeepSeek Releases DualPath: Addressing Storage Bandwidth Bottlenecks in Agentic Inference
- DeepSeek Paper – DualPath: Breaking the Bandwidth Bottleneck in LLM Inference
- Nvidia Could Launch Its First Laptops With Its Own Processors
- Same INT8 Model Shows 93% to 71% Accuracy Variance Across Snapdragon Chipsets
- High Bandwidth Flash Memory Could Alleviate VRAM Constraints in Local LLM Inference
- Carmack Proposes Using Long Fiber Lines as L2 Cache for Streaming AI Data