Close Menu
WithO2WithO2

    Subscribe to Updates

    Get the latest AI News Tools Updates in your Inbox

    What's Hot

    Nvidia Nemotron 3.5 Lightning Cuts AI Agent Costs 74% — What Enterprises Get

    August 13, 2026

    Manus AI Splits From Meta After China Blocks $2B Deal — Back Up Your Data Now

    August 13, 2026

    Needle 2: The 14MB AI Agent That Runs on Any Phone or Device

    August 13, 2026

    Subscribe to Updates

    Get the latest creative news from FooBar about art, design and business.

    Facebook X (Twitter) Instagram
    Facebook X (Twitter) Instagram
    WithO2WithO2
    • AI
    • Blog
    • Business Software
    • Trending News
    • Stories
    WithO2WithO2
    Home » Trending News
    Trending News

    Needle 2: The 14MB AI Agent That Runs on Any Phone or Device

    By Amitabh SarkarAugust 13, 20264 Mins Read1
    Facebook Twitter Pinterest LinkedIn Tumblr Email
    Smartphone and smartwatch running an on-device AI agent neural network
    Needle 2 runs tool calling and device control entirely on the handset, with no cloud connection.
    Share
    Facebook Twitter LinkedIn Pinterest Email

    Published: August 11, 2026

    Cactus Compute released Needle 2 on August 11, 2026, an open 45-million-parameter AI agent model that ships as a single 14MB binary and runs a full session in 28MB of RAM. Needle 2 handles tool calling, device control, and structured data extraction on phones, wearables, smart home devices, and small robots, with no cloud connection and no API calls. The model runs as one dependency-free C++ binary across Cortex-M microcontrollers, ARM, x86, and WebAssembly. On tool-call and mobile device-use benchmarks, Needle 2 trades wins with models 5x to 70x larger, including FunctionGemma 270M, LFM2.5 230M, and Apple FM, while running at 2-bit quantization against their full-precision versions. The Show HN post announcing the release passed 368 points on August 11, 2026.

    Table of Contents

    Toggle
    • What Needle 2 Actually Does
    • Why Edge AI Agents Change the Cost Math
    • How Needle 2 Compares to Other Small Models
    • For Context: Our Coverage of AI Agent Economics

    What Needle 2 Actually Does

    Needle 2 is an agentic LLM built for three jobs: calling tools, operating devices, and extracting structured data. It is not a general chat assistant, and Cactus Compute does not present it as one. The narrow scope is the design: agentic workflows spend most of their compute on repetitive “read input, choose a function, emit structured output” loops, and those loops do not require the general reasoning capacity of a frontier model.

    The architecture is a Simple Attention Network compressed to CQ2-bit precision using Cactus Quants, the company’s own quantization method. The 28MB session footprint means Needle 2 runs alongside existing applications on low-end hardware rather than demanding a dedicated device. Needle 1, the previous version, was a 26M-parameter model distilled from Gemini tool-calling behaviour.

    “Every architectural choice was benchmarked on the target hardware before it earned its parameters.” — Cactus Compute team, cactuscompute.com/needle

    Why Edge AI Agents Change the Cost Math

    An on-device agent removes per-query API cost entirely: the business pays for the deployment once, and every subsequent tool call runs free on hardware it already owns. Cloud-hosted agents invert that — each function call, sensor read, and structured query bills at the provider’s token rate and adds network latency.

    Four business consequences follow from running the model on the device. Costs stop scaling with query volume, which matters most for high-frequency loops such as sensor polling, inventory scanning, and field-device automation. The agent works offline, which makes it viable for IoT fleets, remote sites, and air-gapped environments. Data never leaves the hardware, which removes a category of vendor-processing questions for regulated workloads. And the 28MB footprint fits devices that cannot host a larger model at all, such as wearables, smart home hubs, and Cortex-M controllers.

    For teams already cutting AI agent costs at the infrastructure layer, an on-device model attacks the same bill from the opposite direction — eliminating the per-call charge rather than reducing it.

    How Needle 2 Compares to Other Small Models

    Needle 2 is open and cross-platform, which separates it from the two closest alternatives. Apple FM runs on-device but stays closed and Apple-hardware-only. LFM2.5 from Liquid AI is efficient at 2.6B parameters, but that is roughly 57 times larger than Needle 2’s 45M. Cactus Compute publishes the Needle 2 weights publicly on GitHub.

    The benchmark claim needs precision: Needle 2 trades wins with larger models on tool-call and device-use tests. It does not beat them outright, and it does not replace general-purpose models such as Claude or GPT-4o for reasoning, writing, or open-ended conversation. Businesses evaluating AI agents for business tasks should read Needle 2 as a component for the execution layer, not a swap for the model that does the thinking.

    For Context: Our Coverage of AI Agent Economics

    Needle 2 lands in an ongoing shift in how businesses pay for and control agent behaviour.

    • Cloudflare Wallets lets AI agents spend money with limits you set — the spending-control layer that emerged as agents started transacting autonomously.
    • Humans miss 1 in 3 AI agent threats — the 40,000-run study on how much oversight agent deployments actually get.

    Our Take: Needle 2 is the first credible answer to a question nobody was asking loudly enough — do your AI agents really need the cloud? Cactus Compute is a small team and this is a research release, not enterprise software. But for any deployment where per-query cost scales with volume, a 14MB model that trades benchmark wins with models 50 times its size changes the calculation, and the architectural ideas here will sit in most edge AI stacks within 18 months.

    Share. Facebook Twitter Pinterest LinkedIn Tumblr Email
    Amitabh Sarkar
    • Website

    I am a software engineer, I have a passion for working with cutting-edge technologies and staying up-to-date with the latest developments in the field. In my articles, I share my knowledge and insights on a range of topics, including business software, how to set up tools, and the latest trends in the tech industry.

    Related Posts

    Nvidia Nemotron 3.5 Lightning Cuts AI Agent Costs 74% — What Enterprises Get

    August 13, 2026

    Manus AI Splits From Meta After China Blocks $2B Deal — Back Up Your Data Now

    August 13, 2026

    Salesforce Slackbot Can Now Pull CRM Data — No App-Switching

    August 12, 2026

    Comments are closed.

    Don't Miss
    Trending News

    Nvidia Nemotron 3.5 Lightning Cuts AI Agent Costs 74% — What Enterprises Get

    By Amitabh SarkarAugust 13, 2026

    Published: August 12, 2026 Nvidia released Nemotron 3.5 Lightning and NeMo Switchyard on August 11,…

    Manus AI Splits From Meta After China Blocks $2B Deal — Back Up Your Data Now

    August 13, 2026

    Salesforce Slackbot Can Now Pull CRM Data — No App-Switching

    August 12, 2026

    OpenAI Buys NextSlide: Slide Decks Are Now a ChatGPT Feature

    August 12, 2026

    Subscribe to Updates

    Get the latest creative news from SmartMag about art & design.

    Stay In Touch
    • Facebook
    • Twitter
    • Pinterest
    • Instagram
    Our Picks

    12 Best Project Management Software Tools in 2026

    August 1, 2026

    9 Best Ecommerce Platforms Compared (2026)

    July 30, 2026

    Rippling vs Gusto vs BambooHR: Full HRMS Comparison 2026

    July 15, 2026

    Best Ecommerce Platform 2026: Top 10 Options Compared

    July 5, 2026
    Editors Picks

    Nvidia Nemotron 3.5 Lightning Cuts AI Agent Costs 74% — What Enterprises Get

    August 13, 2026

    Manus AI Splits From Meta After China Blocks $2B Deal — Back Up Your Data Now

    August 13, 2026

    Salesforce Slackbot Can Now Pull CRM Data — No App-Switching

    August 12, 2026

    OpenAI Buys NextSlide: Slide Decks Are Now a ChatGPT Feature

    August 12, 2026
    About Us
    About Us

    Your Source for Innovation: Discover in-depth guides, solutions, and tools tailored to modern business challenges.

    Links
    • Blog
    • Privacy Policy
    • Contact WithO2.com
    • Terms and Conditions
    Facebook X (Twitter) Instagram Pinterest
    • About
    • Editorial Policy
    • Contact
    • Privacy Policy
    • Terms
    © 2026 WITHO2.COM

    Type above and press Enter to search. Press Esc to cancel.