Ayush Pande is a PC hardware and gaming writer. When he's not working on a new article, you can find him with his head stuck inside a PC or tinkering with a server operating system. Besides computing, his interests include spending hours in long RPGs, yelling at his friends in co-op games, and practicing guitar.
I’ve documented all sorts of wacky AI experiments involving weak hardware here on XDA, with my Raspberry Pi 5 being a common guinea pig for a handful of these projects. To its credit, my tiny tinkering companion manages to deliver decent results with smaller LLMs, with Gemma 4 E2B being my preferred model for simple productivity tasks involving this SBC.
However, LLMs with chat interfaces aren’t the only AI models I can run on my weak hardware. Take the Cactus Compute Needle 3, for example. While it needs some tinkering with Python scripts, this slim model barely needs a few MBs to turn my “mere” SBC into an automation beast.
Calling the Needle 3’s footprint tiny would be an understatement
But the trade-off is that it can only handle specialized tasks
If you’re wondering how the Needle 3 manages to consume a maximum of 29MB, it’s because this AI model is a lot different from conventional LLMs. Typical large language models require billions of parameters to respond to queries in a chat-like interface. When you combine that with an underpowered system like the Raspberry Pi, it’s no wonder that even something as light as 3-4B models run under 10 tokens/second.
Meanwhile, the Cactus Compute Needle 3 ditches the general chat capabilities altogether. Instead, it focuses on machine level tasks such as calling tools for different services, converting unstructured data into its JSON equivalents, and generating vector spaces from text inputs – and these three jobs let it perform typical AI tasks without the extra computational overhead of typical chatbots. Combine that with the Needle 3’s unique sliceable ladder architecture that lets you choose the subnetworks from 2 to 20 layers depending on the underlying system, and you can see why I love tinkering with this edge LLM on my Raspberry Pi.
It meshes really well with Raspberry Pi SBCs
The setup process was surprisingly straightforward
For reference, the SBC I’m testing it on is the Raspberry Pi 5, specifically the 8GB variant. OS-wise, I wanted to use Raspberry Pi OS Lite, but since DietPi has an even smaller footprint, I decided to roll with the latter for this experiment. I initially tried running Needle 3 on my first-gen Raspberry Pi Zero, but since the AI model doesn’t support the outdated ARMv6 architecture, it’d throw errors every time building it on my ancient SBC.
Fortunately, setting everything up on my Raspberry Pi 5 was relatively simple. First, I ran the sudo apt install -y python3-venv python3-pip git command to grab the necessary packages. Then, I executed python3 -m venv ~/needle-env to create an isolated Python environment for Needle 3 and switched to it with source ~/needle-env/bin/activate. I also pulled the runtime environment with the pip install cactus-needle command. After a little bit of trial and error, I realized I had to run pip install numpy sentencepiece jax flax cactus-needle requests python-dotenv as well.
With the rest of the packages installed, the needle build --platform linux-arm64 --layers 20 --out ./pi succeeded without any more problems. For the --layers flag, I decided to go all out and assigned all 20 layers, but I could’ve gone as low as 2 layers if I was working with, say, a weak microcontroller (which I plan to do sometime in the future). But I still had to work out a way to assign some tools to this tiny AI model…
I even managed to pair it with Home Assistant via a custom Python file
Since I was particularly interested in the tool-calling aspect of Needle 3, I wanted to test it with some self-hosted services. Local smart home automation is one of the use cases mentioned in the official documentation, so I figured I could toss my Home Assistant server into the mix. But before I could write Python code for the tools, I had to pair my Home Assistant instance with Needle 3, which I did by creating a .needle-ha.env file with the following syntax:
sudo tee /root/.needle-ha.env >/dev/null <<'EOF'
HA_URL=http://192.168.0.24:8123
HA_TOKEN=private-ha-token
NEEDLE_TELEMETRY=0
EOF
If you’re wondering, the HA_TOKEN value is a long-lived access token that I created on Home Assistant. As for the Python tool-calling file, I used this specific snippet to control a smart_plug entity on my HASS instance:
import os
import requests
import needle
from dotenv import load_dotenv
load_dotenv("/root/.needle-ha.env")
HA_URL = os.environ["HA_URL"].rstrip("/")
HA_TOKEN = os.environ["HA_TOKEN"]
HEADERS = {
"Authorization": f"Bearer {HA_TOKEN}",
"Content-Type": "application/json",
}
def ha_service(domain: str, service: str, entity_id: str, data: dict | None = None):
payload = {"entity_id": entity_id}
if data:
payload.update(data)
response = requests.post(
f"{HA_URL}/api/services/{domain}/{service}",
headers=HEADERS,
json=payload,
timeout=10,
)
response.raise_for_status()
return {"ok": True, "entity_id": entity_id, "service": f"{domain}.{service}"}
def set_smart_plug(on: bool):
"""Turn the smart plug on or off."""
entity_id = "switch.smart_plug"
service = "turn_on" if on else "turn_off"
return ha_service("switch", service, entity_id)
agent = needle.Needle(
tools=[
set_smart_plug,
]
)
if __name__ == "__main__":
prompt = input("Command: ")
result = agent.run(prompt)
print(result)
Executing this Python script would immediately cause Needle 3 to ask for a prompt, so I wrote a simple command to turn on the smart_plug entity. Sure enough, Needle 3 had no issues understanding my query and the device came back online immediately. As for the performance, my tiny SBC + Needle 3 combo was able to handle the prompt at a decode rate of 280.2 tokens/s and a prefill rate of 996.4 tokens/s, which is amazing for edge AI tasks.
The best part? The tools themselves are extremely customizable
So far, my Needle 3 can only perform a simple action. But the real fun begins once I start adding more tools to the file. For example, I could add different tools for conventional smart devices, sensor modules, DIY ESP32 IoT products, and even the random FOSS services that Home Assistant recognizes as entities. Heck, I could even ditch the Home Assistant aspect and have it perform analysis tasks on documents, emails, and spreadsheets. Or, I could go into the Docker direction and expose every aspect of a self-hosted workstation to my edge model.
Sure, it’d take some time to create the necessary code for all the tools I want to pair with Needle 3. But the incredible performance more than makes up for the extra effort.
Raspberry Pi 5
- Arm Cortex-A76 (quad-core, 2.4GHz)
- Up to 8GB LPDDR4X SDRAM
- Raspberry Pi OS (official)
- 2× USB 3.0, 2× USB 2.0, Ethernet, 2x micro HDMI, 2× 4-lane MIPI transceivers, PCIe Gen 2.0 interface, USB-C, 40-pin GPIO header
- VideoCore VII
- $60





