Close Menu
WithO2WithO2

    Subscribe to Updates

    Get the latest AI News Tools Updates in your Inbox

    What's Hot

    Qwen-UI-Agent Beats GPT-5.6 at Clicking Through Business Apps

    August 25, 2026

    Fable 5 vs DeepSeek: Why Enterprises Are Switching AI Models

    August 24, 2026

    Anthropic’s Claude Academy Opens Free With 22 AI Courses and Completion Badges

    August 24, 2026

    Subscribe to Updates

    Get the latest creative news from FooBar about art, design and business.

    Facebook X (Twitter) Instagram
    Facebook X (Twitter) Instagram
    WithO2WithO2
    • AI
    • Blog
    • Business Software
    • Trending News
    • Stories
    WithO2WithO2
    Home » Trending News
    Trending News

    Qwen-UI-Agent Beats GPT-5.6 at Clicking Through Business Apps

    By Amitabh SarkarAugust 25, 20263 Mins Read0
    Facebook Twitter Pinterest LinkedIn Tumblr Email
    AI cursor detecting and clicking on-screen elements in a business software interface
    GUI agents control business software by recognising on-screen elements instead of calling an API.
    Share
    Facebook Twitter LinkedIn Pinterest Email

    Published: August 24, 2026

    Alibaba’s Tongyi-MAI team released Qwen-UI-Agent on August 20, 2026, a foundation GUI agent that scored 92.2% on the MobileWorld-Real benchmark and outperformed GPT-5.6 Sol by 12.0 percentage points and Claude Opus 4.8 by 14.6 percentage points on MobileWorld. The model operates phones, PCs, and web apps by reading on-screen elements rather than calling APIs.

    Qwen-UI-Agent posted state-of-the-art results on all 5 benchmarks its developers tested: 97.5% on AndroidDaily, 92.2% on MobileWorld-Real, 82.1% on MobileWorld, 81.5% on ScreenSpot-Pro for GUI grounding, and 73.6% on WebArena for web browsing and deep search. The MobileWorld-Real result, measured on physical devices, beat Gemini 3.1 Pro, Claude Opus 4.8, and GPT-5.6 Sol. Alibaba published the code at github.com/Tongyi-MAI/MAI-UI and a technical report on arXiv. Neither source published pricing or confirmed that model weights are downloadable.

    Table of Contents

    Toggle
    • What a GUI Agent Does That an API Agent Cannot
    • What the Benchmark Scores Do and Do Not Prove
    • For Context: Our GUI and Agent Automation Coverage

    What a GUI Agent Does That an API Agent Cannot

    A GUI agent is an AI system that controls software through the visual interface — recognising on-screen elements, clicking, and typing — instead of calling a programmatic endpoint. API agents require a documented interface; GUI agents require only a screen, which extends automation to software that exposes no API at all.

    That distinction decides which internal workflows are automatable. Legacy systems, such as older ERP installations, custom internal tools, government filing portals, and CRMs predating public APIs, hold large volumes of manual clicking work that API-based agents cannot reach. Robotic process automation vendors, such as UiPath, have served this gap with scripted selectors that break whenever a screen layout changes. A model that recognises elements visually degrades differently: it adapts to layout changes but fails on ambiguity. Teams mapping which of their workflows suit which approach can start from our roundup of the best AI agents for business tasks.

    What the Benchmark Scores Do and Do Not Prove

    ScreenSpot-Pro measures GUI grounding — whether the model correctly identifies which on-screen element matches an instruction — and Qwen-UI-Agent scored 81.5%. Grounding accuracy sets the ceiling for everything downstream, because a multi-step task fails at the first misidentified button.

    Read the 97.5% AndroidDaily result against that 81.5% grounding figure before planning a rollout. AndroidDaily covers routine mobile app sequences; ScreenSpot-Pro covers professional desktop interfaces with dense controls, which is where business workflows actually run. A 1-in-5 grounding miss on complex screens means unattended automation still needs a verification step, and the practical deployment pattern remains human-approved execution rather than fire-and-forget. That gap between pilot scores and shipped systems is the same one documented in reporting that 99% of companies plan agents while 10% ship them.

    For Context: Our GUI and Agent Automation Coverage

    • 15 AI agent examples across industries — the workflow categories GUI control now opens up.
    • Cloudflare Kitesurf’s agent browser — the browser-side approach to the same screen-automation problem.
    • Qwen3.8-27B open weights — Alibaba’s prior release and its self-hosting terms.
    • Alibaba’s agent-native cloud — the infrastructure this model plugs into.
    Our Take
    Benchmark leadership on real devices is the first credible signal that GUI agents have moved from demo to tool, and it arrives from Alibaba rather than from OpenAI or Anthropic — worth noting for anyone assuming the frontier labs own every agent category. Do not budget for an RPA replacement yet. Pick one high-volume, low-risk screen workflow, such as extracting a weekly report from a system with no export API, run it with a human approving each submission, and measure the error rate against your own interface rather than against AndroidDaily. Confirm licensing terms in the GitHub repository before any production plan, since neither pricing nor weight availability is documented.
    Share. Facebook Twitter Pinterest LinkedIn Tumblr Email
    Amitabh Sarkar
    • Website

    I am a software engineer, I have a passion for working with cutting-edge technologies and staying up-to-date with the latest developments in the field. In my articles, I share my knowledge and insights on a range of topics, including business software, how to set up tools, and the latest trends in the tech industry.

    Related Posts

    Fable 5 vs DeepSeek: Why Enterprises Are Switching AI Models

    August 24, 2026

    Anthropic’s Claude Academy Opens Free With 22 AI Courses and Completion Badges

    August 24, 2026

    Model Context Protocol Roadmap Names 5 Priorities for AI Agents After Stateless Rewrite

    August 24, 2026

    Comments are closed.

    Don't Miss
    Trending News

    Fable 5 vs DeepSeek: Why Enterprises Are Switching AI Models

    By Amitabh SarkarAugust 24, 2026

    Published: August 24, 2026 Anthropic’s flagship Fable 5 model accounts for roughly 11% of what…

    Anthropic’s Claude Academy Opens Free With 22 AI Courses and Completion Badges

    August 24, 2026

    Model Context Protocol Roadmap Names 5 Priorities for AI Agents After Stateless Rewrite

    August 24, 2026

    HubSpot Makes Its AI-First CRM Redesign Mandatory for All Users

    August 23, 2026

    Subscribe to Updates

    Get the latest creative news from SmartMag about art & design.

    Stay In Touch
    • Facebook
    • Twitter
    • Pinterest
    • Instagram
    Our Picks

    12 Best Project Management Software Tools in 2026

    August 1, 2026

    9 Best Ecommerce Platforms Compared (2026)

    July 30, 2026

    Rippling vs Gusto vs BambooHR: Full HRMS Comparison 2026

    July 15, 2026

    Best Ecommerce Platform 2026: Top 10 Options Compared

    July 5, 2026
    Editors Picks

    Fable 5 vs DeepSeek: Why Enterprises Are Switching AI Models

    August 24, 2026

    Anthropic’s Claude Academy Opens Free With 22 AI Courses and Completion Badges

    August 24, 2026

    Model Context Protocol Roadmap Names 5 Priorities for AI Agents After Stateless Rewrite

    August 24, 2026

    HubSpot Makes Its AI-First CRM Redesign Mandatory for All Users

    August 23, 2026
    About Us
    About Us

    Your Source for Innovation: Discover in-depth guides, solutions, and tools tailored to modern business challenges.

    Links
    • Blog
    • Privacy Policy
    • Contact WithO2.com
    • Terms and Conditions
    Facebook X (Twitter) Instagram Pinterest
    • About
    • Editorial Policy
    • Contact
    • Privacy Policy
    • Terms
    © 2026 WITHO2.COM

    Type above and press Enter to search. Press Esc to cancel.