Skip to content
← Tags

#Multi-agent (13)

AI CodingDevelopmentTopC

Qoder: Alibaba agentic platform for real work - nine product lines on one knowledge engine

Qoder is the agentic coding platform from Alibaba, positioned officially as an agentic platform for real work: not an editor but an end-to-end loop - understand the task and context, plan, call tools, verify results, iterate toward the deliverable - resting on three stated principles (context engineering, agent autonomy, goal-directed loops), with nine product lines sharing one knowledge engine (desktop Qoder and Qoder IDE coexisting rather than replacing each other, Editor and Quest forms, a JetBrains plugin, Qoder CLI, cloud agents and more); session history and memory are stored separately but can be imported from the IDE. Repo Wiki is generated locally by multiple agents, never uploads the codebase, is off by default and supports Auto Update, Auto Export and Citation back to source locations. Quest has four drives - Agent, Experts, Goal and Spec (convertible to scheduled tasks): Spec runs requirement clarification (multiple choice, with Recommend / Continue / Skip), a structured Spec covering requirements, design, task breakdown and acceptance criteria, human review, execution, then Review/Commit/Push, while Goal takes only the desired outcome and evaluates progress at the end of every round, continuing automatically until met. Two scaled cases: building Qoder with Qoder (10 people, 3 weeks, 500,000 lines of agent code merged into a 4-million-line legacy system, 99% agent-generated, still in production at v1.4.0 with zero incidents; the method is a cognitive base plus Ultra Spec plus Experts cross-review plus a verifier agent filtering hallucinated issues, with humans only deciding SLO definitions and irreversible operations, and each person driving 20-plus Experts tasks a day); and AutoSDK for AMap in-car systems across 20-plus repositories and over a million lines, where the strict first-pass rate went from 37.3% to 61.5% (problem framing cites KoCo-Bench / arXiv:2601.13240v3: general coding reaches 90% Pass@1 while domain code generation reaches only 8.9%). Boundaries: the client and knowledge engine are closed; every scaled number comes from official cases and vendor self-reporting, and self-evidence from a product about itself carries methodological self-interest, none of it independently reproduced; Experts cost per unit is clearly above a single agent (median about 75 versus about 50 Credits) and Credits reset each cycle rather than accumulating. Confidence C (vendor-claim).

500,000 行 · 99% agent 生成10 人 3 周并入 400 万行遗留系统(厂商案例)Vendor Claim · 2026
ProductAlibabaSiteRepo
Qoder: Alibaba agentic platform for real work - nine product lines on one knowledge engine
World Action Agent: Harnessing VLMs for Robot Manipulation via World Action Rehearsal

World Action Agent: Harnessing VLMs for Robot Manipulation via World Action Rehearsal

HKUST(GZ), CUHK and Knowin AI present World Action Agent (WAA), a multi-agent harness that changes what the VLM sees and how its decisions take effect — rather than the VLM itself — so a general-purpose VLM can pilot a robot with basic tools (point, drag, preview, execute), making every decision inside one visual action workspace. Three properties: Contact views, whose camera parameters are solved from interaction-region visibility under occlusion, framing compactness, view redundancy and stability, with the feasible set constraining the two views to orthogonal horizontal projections so any alignment error can be read along two independent directions — requiring only a base-frame point cloud, so the perception source (simulation, fused RGB-D, or VGGT reconstruction) is irrelevant; action rehearsal, where each action is an editable proposal planned by cuRobo, overlaid as a translucent robot in every view with a feasibility report, refined by an Imagination Agent in a separate context while the physical scene stays unchanged, and only proposals with executable plans become motion; and in-view correction, where the agent drags from a reference point to the desired location in the Contact view where the error is observed — the reference can be a visible point on the held object, so no conversion into an absolute gripper pose is needed. Skills evolve from expert videos and human Canvas-GUI teaching through Learner, Editor and Reviewer roles under evidence-citing and independent-review constraints. On LIBERO-Pro with Gemini 3.7 Flash, using skills evolved only from LIBERO-90 and frozen before evaluation, WAA reaches a state-of-the-art 75.6% average success, above ASPIRE (72.0%), end-to-end VLAs and the same-backbone Show-Harness (6.7%), scoring 80.0 / 73.3 on the two Spatial splits; an episode costs just 31 model calls, 150 s and $0.1996 versus 120 calls, 874 s and $0.5021 for Show-Harness. The frozen skills transfer to robosuite at 100.0%, and LoRA fine-tuning Qwen3.5-9B on 112 trajectories (1,774 decision steps) as the main agent only raises out-of-domain success from 1.7% to 43.3%.

Yehang Zhang, Haojian Huang, Yifan ChangSep 24, 2026
World Action AgentWAAVLMSep 24, 2026
Dreamer-CPC: Message Learning with World Models for Decentralized Multi-agent Reinforcement Learning

Dreamer-CPC: Message Learning with World Models for Decentralized Multi-agent Reinforcement Learning

In multi-agent reinforcement learning (MARL), inter-agent communication is effective for improving performance under partial observability. Representation learning-based approaches enable decentralized agents to learn messages grounded in their own observations, but they rely only on current observations and cannot convey information accumulated over time. We propose Dreamer-CPC, a decentralized model-based MARL method that integrates message learning based on Collective Predictive Coding (CPC) into the world model of DreamerV3. Each agent independently maintains a world model and a message module, and infers and exchanges messages from the latent states of the world model that reflect the history of past observations and actions. We evaluated Dreamer-CPC in two environments: Observer, a non-cooperative information-sharing task, and CatchApple, a newly introduced task in which task-relevant observations are temporarily missing. In both environments, Dreamer-CPC outperformed IPPO-CPC, an existing CPC-based method that generates messages from current observations, as well as no-communication baselines. In particular, in CatchApple, Dreamer-CPC achieved 4 to 5 times the episode return of IPPO-CPC, demonstrating effective coordination where other methods fail due to missing observations. These results suggest that communication grounded in the latent dynamics of world models can support decentralized decision-making when current observations alone are insufficient.

Taisuke Takayama, Naoto Yoshida, Tadahiro TaniguchiJul 22, 2026
DreamerWorld ModelsMulti-agentJul 22, 2026