Cactus Compute Released Needle 2 for Raspberry Pi 5

The 45-million-parameter model enables local natural-language function calling on edge hardware.

Updated on Sept. 24, 2026 in Artificial Intelligence

Isometric editorial illustration of a compact circuit board surrounded by floating metallic geometric nodes, representing local edge AI computation.
Cactus Compute launched Needle 2, a 45-million-parameter AI model designed to perform local natural-language function calling on Raspberry Pi 5 hardware. AI Illustration. Upload story photo >

Live Poll

Would you prefer using AI tools that function entirely offline on your personal devices?

Cactus Compute has released Needle 2, an open-source artificial intelligence model designed to run locally on the Raspberry Pi 5. The model enables offline, natural-language-to-Python function calling without relying on external cloud APIs.

Why it matters

By facilitating local function execution, the model removes the need for network connectivity and cloud-based API costs for edge computing workflows. This release underscores the growing viability of small, specialized language models for constrained hardware environments.

Needle 2 utilizes two-bit quantisation—a process of reducing the precision of model weights—to reach a 14MB footprint. The model maintains a 256-token sliding context window and achieves a decoding throughput of 248-314 tokens per second.

The players

Cactus Compute

A developer of compact, edge-focused artificial intelligence models that prioritize local execution and low memory consumption.

The details

The model employs an attention network architecture that incorporates grouped-query attention, a method that reduces memory usage by sharing key and value heads across multiple query heads. Developers integrate the model by applying the @needle.tool decorator to Python functions, which allows the system to interpret natural-language requests and select up to five relevant tools for execution. The entire process requires 28MB for the native session and peaks at 46.4MB for the full Python process, ensuring stability on the Raspberry Pi 5.

Timeline

  1. September 24, 2026: Needle 2 was publicly released.

The Tech Race

This development follows a broader industry trend of migrating complex AI tasks from centralized cloud clusters to local, edge-based execution. Needle 2 competes against larger, cloud-dependent model infrastructures by offering a lightweight alternative that sacrifices broad generality for high-speed, offline function calling.

Developers can implement this tool immediately for offline automation tasks, such as triggering hardware signals with measured latencies of 78ms for LED switching. The system is available now under an Apache 2.0 license, requiring no subscription or internet-dependent API keys.

The takeaway

Needle 2 demonstrates that high-performance function calling is now achievable on low-power devices. Developers should monitor Cactus Compute's repository for future updates to model accuracy or expansions to the tool-retrieval limit.

Further reading

For broader trends in local model deployment, see Artificial Intelligence.

Source note: This article includes information reported by Electronics For You - Official Site ElectronicsForU.com.

Live Poll

Would you prefer using AI tools that function entirely offline on your personal devices?