Cactus Compute Released Needle 2 for Raspberry Pi 5
The 45-million-parameter model enables local natural-language function calling on edge hardware.
Updated on Sept. 24, 2026 in Artificial Intelligence

Live Poll
Would you prefer using AI tools that function entirely offline on your personal devices?
Cactus Compute has released Needle 2, an open-source artificial intelligence model designed to run locally on the Raspberry Pi 5. The model enables offline, natural-language-to-Python function calling without relying on external cloud APIs.
Why it matters
By facilitating local function execution, the model removes the need for network connectivity and cloud-based API costs for edge computing workflows. This release underscores the growing viability of small, specialized language models for constrained hardware environments.
Needle 2 utilizes two-bit quantisation—a process of reducing the precision of model weights—to reach a 14MB footprint. The model maintains a 256-token sliding context window and achieves a decoding throughput of 248-314 tokens per second.
The players
Cactus Compute
A developer of compact, edge-focused artificial intelligence models that prioritize local execution and low memory consumption.
The details
The model employs an attention network architecture that incorporates grouped-query attention, a method that reduces memory usage by sharing key and value heads across multiple query heads. Developers integrate the model by applying the @needle.tool decorator to Python functions, which allows the system to interpret natural-language requests and select up to five relevant tools for execution. The entire process requires 28MB for the native session and peaks at 46.4MB for the full Python process, ensuring stability on the Raspberry Pi 5.
Timeline
September 24, 2026: Needle 2 was publicly released.
The Tech Race
This development follows a broader industry trend of migrating complex AI tasks from centralized cloud clusters to local, edge-based execution. Needle 2 competes against larger, cloud-dependent model infrastructures by offering a lightweight alternative that sacrifices broad generality for high-speed, offline function calling.
Developers can implement this tool immediately for offline automation tasks, such as triggering hardware signals with measured latencies of 78ms for LED switching. The system is available now under an Apache 2.0 license, requiring no subscription or internet-dependent API keys.
The takeaway
Needle 2 demonstrates that high-performance function calling is now achievable on low-power devices. Developers should monitor Cactus Compute's repository for future updates to model accuracy or expansions to the tool-retrieval limit.
Further reading
For broader trends in local model deployment, see Artificial Intelligence.
Source note: This article includes information reported by Electronics For You - Official Site ElectronicsForU.com.
Live Poll
Would you prefer using AI tools that function entirely offline on your personal devices?






