J’ai testé une IA géante chinoise sur Mac mini, PC gamer et mini PC : voici le grand gagnant sur Bonsai 2

In a significant leap towards democratizing artificial intelligence, American startup Prism ML has successfully compressed Alibaba’s formidable Qwen3.8-27B, a large language model boasting 27 billion parameters, into an astonishingly compact 5.95 GB file. This breakthrough, dubbed "Bonsai 2," allows the sophisticated AI model to run locally on a wide array of consumer-grade computers, from Apple’s Mac mini M6 to Nvidia-powered gaming PCs and various Intel and AMD mini PCs. While the model demonstrates universal compatibility, its operational speed varies dramatically across different machines, highlighting a crucial insight: hardware specifications are vital, but two simple software adjustments can profoundly alter performance dynamics.

The Dawn of Accessible Local AI: A Paradigm Shift

The ability to run advanced AI models directly on personal devices, rather than relying on cloud-based services, represents a pivotal shift in the AI landscape. This "local AI" approach offers numerous benefits, including enhanced privacy (data remains on the user’s device), reduced latency, lower operational costs (no subscription fees), and the capability to function offline. Historically, deploying large language models (LLMs) locally has been a formidable challenge due to their immense size and computational demands. Models like Qwen3.8-27B, in their original high-precision formats, can easily exceed 50 gigabytes, making them incompatible with the limited video RAM (VRAM) of most consumer graphics cards or the unified memory of integrated systems.

Alibaba’s Qwen series, developed by the Alibaba Cloud intelligence team, has emerged as a significant player in the global open-source AI community. Known for its robust performance across various benchmarks, the 27-billion-parameter version, Qwen3.8-27B, is a testament to the rapid advancements in AI research originating from diverse technological hubs worldwide. Its original size, however, meant it was largely confined to powerful data centers or high-end professional workstations equipped with specialized hardware.

J’ai testé une IA géante chinoise sur Mac mini, PC gamer et mini PC : voici le grand gagnant sur Bonsai 2

Prism ML, a startup with roots in Caltech research, has consistently pushed the boundaries of AI model compression. Their previous endeavors include the remarkable feat of enabling a 54 GB model to run efficiently on an iPhone, demonstrating their expertise in optimizing large models for resource-constrained environments without sacrificing core capabilities. Bonsai 2, released on September 17, 2026, under an Apache 2.0 license, further solidifies Prism ML’s position as a leader in making cutting-edge AI more accessible. While the model weights are freely downloadable from Hugging Face, running Bonsai 2 currently requires Prism ML’s proprietary runtime, preventing a direct "one-click" launch in standard AI deployment tools like LM Studio or Ollama. This strategic decision by Prism ML aims to ensure optimal performance and compatibility with their highly specialized compression format.

The Technical Alchemy: Quantization and Data Reduction

At its core, an AI model is a vast collection of numerical values, known as "weights," which encapsulate the knowledge acquired during its training phase. In its standard FP16 (16-bit floating-point) representation, each of Qwen3.8-27B’s 27 billion parameters occupies 16 bits of memory, culminating in an imposing file size of approximately 54 GB. This gargantuan footprint makes it impossible to fully load into the 16 GB VRAM of a typical gaming graphics card, such as the Nvidia RTX 4070 Ti SUPER used in the tests.

Prism ML’s innovation lies in its radical approach to quantization—a process that reduces the precision of these weights. Instead of traditional methods that might scale weights down to 4 or 8 bits (Q4 or Q8), Prism ML employs a ternary representation (PTQ1_0). In this method, each weight is constrained to one of only three possible values: -1, 0, or +1. This extreme post-training quantification technique achieves an average of just 1.75 bits per weight, shrinking the model to an astonishing 5.95 GB GGUF file. For context, a more common Q4 quantization, used for comparative testing, still results in a 17.1 GB file, requiring about 5 bits per weight. This dramatic reduction is the primary enabler for Bonsai 2’s local deployment on consumer hardware.

The immediate benefit of this aggressive compression is evident: a 17 GB model cannot fully reside within a 16 GB VRAM graphics card, leading to performance bottlenecks as data is constantly swapped with slower system memory. However, when the model size is reduced to a mere 6 GB, it can comfortably fit within the VRAM, unlocking the full computational power of the GPU.

J’ai testé une IA géante chinoise sur Mac mini, PC gamer et mini PC : voici le grand gagnant sur Bonsai 2

A Multi-Platform Gauntlet: Testing Across Diverse Hardware

To thoroughly evaluate Bonsai 2’s performance and the impact of underlying hardware and software, a diverse array of machines was assembled for rigorous testing:

  1. Nvidia Gaming PC: Equipped with an RTX 4070 Ti SUPER graphics card featuring 16 GB of VRAM. This represents a high-performance, dedicated GPU setup commonly found in gaming and enthusiast machines.
  2. Apple Mac mini M6: Configured with 24 GB of unified memory. Apple’s M-series chips are renowned for their integrated architecture, where CPU and GPU share a single, high-bandwidth memory pool, theoretically offering efficient data access for AI workloads.
  3. Intel Mini PC: A GMKtec EVO-T2S, featuring an Intel Core Ultra processor with an integrated Arc graphics chip and 64 GB of shared system memory. This represents a compact, general-purpose mini PC with a modern integrated GPU.
  4. AMD Mini PC: A GMKtec EVO-X3, powered by an AMD Ryzen AI Max+ 395 processor with 128 GB of unified memory, of which 64 GB was specifically reserved for the integrated graphics processing unit (GPU). This high-end mini PC showcases AMD’s latest AI acceleration capabilities.

A key architectural detail shared by the Mac mini and both mini PCs is their unified or shared memory design, where the CPU and GPU access the same pool of RAM. This can be advantageous for AI tasks by minimizing data transfer overheads between different memory types.

The testing protocol was standardized to ensure fair comparisons: the same version of Prism ML’s runtime was used across all machines, alongside the 5.95 GB Bonsai 2 model file and, for comparative purposes, a 17.1 GB Qwen3.8 model quantized to Q4. Performance was measured using the "llama-bench" tool, focusing on "hot generation" throughput. Each test involved processing a 512-token prompt followed by the generation of 128 tokens, with five repetitions to establish an average. A separate, real-world probability problem was also posed to the AI on each machine to verify functional correctness, ensuring no system was "delivering gibberish" due to computational errors, though this specific test wasn’t designed to assess overall model quality.

Performance Unveiled: Startling Disparities and Unexpected Turnarounds

J’ai testé une IA géante chinoise sur Mac mini, PC gamer et mini PC : voici le grand gagnant sur Bonsai 2

The benchmark results revealed stark performance differences, underscoring the complex interplay between model format, hardware, and software optimization:

Machine / Model Bonsai 2 (tokens/s) Qwen3.8 Q4 (tokens/s)
Nvidia RTX 4070 Ti SUPER 69.0 17.3
Apple Mac mini M6 18.1 9.47
AMD Mini PC (ROCm) 21.57 12.45
AMD Mini PC (Vulkan) 3.95 12.67
Intel Mini PC (Vulkan) 1.75 4.14

Two insights immediately stand out. First, the Nvidia RTX 4070 Ti SUPER unequivocally dominated, demonstrating a generation speed 39.3 times faster than the Intel Arc GPU on Bonsai 2, and 3.8 times faster than the Mac mini M6. It’s crucial to contextualize this "39.3x" figure as a specific measurement of generation throughput under these precise test conditions, with this particular weight format and backend.

Second, the relative performance of Bonsai 2 versus the standard Q4 model shifted dramatically depending on the machine and, more importantly, the software backend. Bonsai 2 outperformed the Q4 classic on CUDA (Nvidia’s backend) and Metal (Apple’s backend). However, the Q4 model surprisingly took the lead on the Intel mini PC using Vulkan. The AMD mini PC results proved to be the most illustrative of this phenomenon.

Prism ML’s own projections had suggested up to 143 tokens per second on an RTX 5090 and 46.8 tokens per second on an M5 Max chip. The observed 69 tokens per second on the RTX 4070 Ti SUPER and 18 tokens per second on the Mac mini M6 align logically, given the use of more modest hardware in these tests.

The Unseen Hand: Software Backends and Memory Management

J’ai testé une IA géante chinoise sur Mac mini, PC gamer et mini PC : voici le grand gagnant sur Bonsai 2

The dramatic performance variations are not solely attributable to raw hardware power; software backends play an equally critical role. Backends like Vulkan, Metal, CUDA, and ROCm are the low-level software layers that enable the application to communicate efficiently with the graphics processing unit (GPU). Vulkan, initially designed as a cross-platform graphics API for gaming (and now adopted by titles like Minecraft), is widely supported. However, its efficiency for AI workloads, especially with novel formats like Prism ML’s ternary quantization, hinges on the availability of highly optimized "kernels"—small, specialized programs that execute specific computational tasks on the GPU. The ternary format, it appears, is not yet optimally supported by all Vulkan kernels, leading to performance bottlenecks.

The RTX Paradox: Less Work, More Speed

One of the most counter-intuitive findings emerged from the Nvidia RTX testing. The Q4 version of Qwen3.8, weighing 17.1 GB, could not entirely fit into the RTX 4070 Ti SUPER’s 16 GB of VRAM. Initially, the runtime reported that all 66 layers were sent to the GPU, and the test completed at 11.0 tokens per second. However, a deeper analysis of the system logs revealed a critical issue: the VRAM was saturated, with 682 MB of model weights spilling over into the much slower system RAM. This constant swapping of data between VRAM and system RAM severely throttled the GPU’s performance.

Through careful experimentation, the optimal configuration involved offloading just three of the 66 layers to the CPU, which then calculated them using the PC’s 32 GB of DDR5 system memory. By relieving the VRAM of this minor burden, the GPU gained crucial memory headroom, and its generation rate surged to 17.3 tokens per second—a remarkable 57% improvement. This demonstrates a key principle: sometimes, strategically reducing the workload on a saturated component, even by redirecting it to a theoretically slower one, can significantly boost overall system performance by optimizing resource allocation. Had a GPU with more VRAM (e.g., 24 GB on an RTX 4090 or 32 GB on an RTX 5090) been used, the Q4 model would have fit entirely, likely resulting in even higher throughput without the need for CPU offloading.

The Backend Revelation: A Software Switch Transforms Performance

J’ai testé une IA géante chinoise sur Mac mini, PC gamer et mini PC : voici le grand gagnant sur Bonsai 2

The AMD mini PC provided the most compelling evidence for the profound impact of software backends. The same tests, using identical files and settings, were run twice, with the only variable being the backend: Vulkan versus ROCm. Under Vulkan, Bonsai 2 achieved a modest 3.95 tokens per second. Switching to ROCm (AMD’s open-source compute platform) propelled Bonsai 2’s performance to an impressive 21.57 tokens per second—a staggering 5.45-fold increase, all without any hardware changes. In contrast, the Q4 model’s performance remained relatively stable, shifting marginally from 12.67 to 12.45 tokens per second.

This backend switch effectively inverted the performance hierarchy on the AMD machine. Under Vulkan, Q4 was 3.2 times faster than Bonsai 2. Under ROCm, Bonsai 2 became 1.73 times faster than Q4. This striking result underscores that the speed of an AI model is not solely a function of hardware power or model size, but critically depends on the optimization of software kernels for a specific weight format on a particular hardware architecture. While ROCm proved superior for Bonsai 2 in this instance, it’s not a universal truth; other models, like Google’s Gemma 4, have shown faster performance under Vulkan in different tests, highlighting the complexity and model-specific nature of backend optimization.

The Quality Conundrum: Compression vs. Fidelity

While the performance gains are undeniable, the aggressive compression employed by Prism ML naturally raises questions about model quality and accuracy. Prism ML asserts that Bonsai 2 retains an impressive 98.2% of the original FP16 version’s performance, achieving an average score of 84.78 across 14 reasoning benchmarks compared to the FP16’s 86.32.

However, real-world testing conducted by MindStudio offered a more nuanced perspective. In their evaluation, Bonsai 2 reportedly failed to detect a deliberately hidden bug within an application and entered a loop during a Tamil translation task. This suggests that while core reasoning capabilities are largely preserved, highly aggressive quantization can, in certain edge cases, subtly impact the model’s robustness, nuance, or ability to handle specific linguistic or logical complexities. Users must weigh the significant performance and accessibility benefits against these potential, albeit infrequent, trade-offs in qualitative performance.

J’ai testé une IA géante chinoise sur Mac mini, PC gamer et mini PC : voici le grand gagnant sur Bonsai 2

Choosing Your Local AI Champion: A Multi-faceted Decision

The question of which machine is "best" for local AI is complex, with no single universal winner. The choice depends heavily on user priorities, budget, and specific AI workloads.

  • Nvidia RTX (Gaming PC): Offers unparalleled raw speed for Bonsai 2 when the model fits entirely within its VRAM. It is the top choice for enthusiasts prioritizing maximum throughput for well-fitting models. However, its performance can be severely hampered when models exceed VRAM capacity, necessitating careful memory management and potential CPU offloading.
  • Apple Mac mini M6 (24GB+): Presents an excellent balance of performance, energy efficiency, and silent operation. Its unified memory architecture provides a coherent and effective platform for local AI, achieving a respectable 18 tokens per second with Bonsai 2. For those seeking a compact and reliable AI workstation, the 24 GB configuration (or higher) is strongly recommended, as the 16 GB version can quickly become memory-constrained, especially with growing AI context in longer conversations.
  • AMD Ryzen AI Max+ Mini PC: This high-end mini PC, with its substantial 128 GB of unified memory (64 GB dedicated to the GPU in this test), is a powerhouse for handling large AI models without memory constraints. Its performance, however, is highly contingent on the software backend, as demonstrated by the dramatic 5.45x speedup when switching from Vulkan to ROCm for Bonsai 2. It offers significant potential for advanced users willing to optimize their software stack.
  • Intel Core Ultra Mini PC: While the slowest performer in these specific tests for Bonsai 2, it still represents a functional, albeit less optimized, platform. Future advancements in Intel’s AI acceleration software and Vulkan kernel optimizations could potentially improve its standing.

In terms of value among the compact machines tested, the Mac mini M6 (priced around €1,269 at the time of writing) offered a compelling price-to-performance ratio for running local AI, significantly outperforming the Intel mini PC (around €1,800) by a factor of ten on Bonsai 2, and even competing favorably with the much more expensive AMD mini PC (€3,499) in certain scenarios.

Crucially, the performance metrics observed in these tests represent "single-prompt" scenarios. Real-world AI usage often involves complex, multi-turn conversations, where the "context window" (the AI’s memory of the ongoing discussion) can rapidly expand, consuming significant amounts of memory. Therefore, machines with ample VRAM or unified memory are paramount for sustained, high-performance local AI interaction.

The Evolving Landscape: Software’s Enduring Impact

J’ai testé une IA géante chinoise sur Mac mini, PC gamer et mini PC : voici le grand gagnant sur Bonsai 2

This comprehensive evaluation vividly illustrates the dynamic nature of AI performance benchmarks. The fact that a mere software adjustment—changing a backend—could multiply Bonsai 2’s throughput by 5.45 times on a single machine underscores the immense, often underestimated, power of software optimization. A future update to Vulkan’s kernels, for instance, could easily redefine the performance hierarchy, highlighting that these benchmarks are merely snapshots of a rapidly evolving technological landscape.

Prism ML’s work is part of a broader industry trend to democratize AI. From previously enabling a 27-billion-parameter model to run on an iPhone without cloud connectivity, to the ongoing efforts by chip manufacturers like Qualcomm to bring local AI capabilities to even budget-friendly smartphones, the push for on-device intelligence is relentless. Tools are emerging that allow users to quickly ascertain their machine’s local AI compatibility, further facilitating this transition. As compression techniques become more sophisticated, hardware-software co-design improves, and integrated GPUs gain more power, local AI is poised to become an increasingly ubiquitous and essential part of our digital lives, offering unprecedented privacy, control, and accessibility.

Related Posts

Tado° X Wireless Starter Kit (2nd Generation) Offers Advanced Smart Heating Control with Significant Discount at Darty

The Tado° X wireless starter kit, representing the second generation of the company’s premium connected heating solutions, is currently available at a notable discount on Darty. Priced at €149.99, down…

Microsoft Cedes Low-Cost Laptop Market to OEMs, Declining to Challenge Apple’s MacBook Neo

In a significant strategic pronouncement that redefines its hardware ambitions, Microsoft has confirmed it will not launch a low-cost Surface device designed to directly compete with Apple’s recently introduced MacBook…

Leave a Reply

Your email address will not be published. Required fields are marked *

You Missed

Hasbro’s Little Miss No-Name: The Unsettling 1965 Doll That Divided Generations and Became a Cultural Artifact

Hasbro’s Little Miss No-Name: The Unsettling 1965 Doll That Divided Generations and Became a Cultural Artifact

Kingdom Come Deliverance 2 Developer Hopes Grand Theft Auto 6 Will Standardize Eighty Dollar Game Pricing to Ensure Industry Sustainability

Kingdom Come Deliverance 2 Developer Hopes Grand Theft Auto 6 Will Standardize Eighty Dollar Game Pricing to Ensure Industry Sustainability

Sean Parker Returns to Music Industry, Championing AI with Stability AI and Major Label Backing

Sean Parker Returns to Music Industry, Championing AI with Stability AI and Major Label Backing

Building the Next Generation of AI Giants: Blackstone’s Jas Khaira to Share Insights on Sustainable AI Growth at TechCrunch Disrupt 2026

Building the Next Generation of AI Giants: Blackstone’s Jas Khaira to Share Insights on Sustainable AI Growth at TechCrunch Disrupt 2026

GitLab Issues Urgent Patch for Critical AI Gateway Vulnerability Enabling Arbitrary Code Execution

GitLab Issues Urgent Patch for Critical AI Gateway Vulnerability Enabling Arbitrary Code Execution

Apple Acknowledges AT&T Network Bug on iPhone 18 Pro Max

Apple Acknowledges AT&T Network Bug on iPhone 18 Pro Max