0:40As an Amazon Associate, we earn from qualifying purchases.
As an Amazon Associate, we earn from qualifying purchases.

Convai Innovations' open answer to TypeSafe Jev: a 421M non-autoregressive decision model trained with RLCD proper scoring rules, plus the README's honest limits and real Jev comparisons.

A humanoid robot is not an assembly of seven modules but a stack of physics equations that set each other's boundary conditions. This article computes the whole-machine stack layer by layer: joint motor modules (declared torque versus real quasi-static CoP demand — knee margins across three vendors converge to 2.23-2.42x while BHL's knee has only 1.51x) -> IMU (lever-arm pseudo-acceleration is 21,752x the sensor noise floor, so mounting position matters four orders of magnitude more than the datasheet) -> materials and structure (three BOM revisions of AgiBot's X1 as a load-path history: every part entering the closed-chain drivetrain upgraded to 7075-T6 / TC4 / 17-4PH) -> sensors (fix the observation space before the shopping list) -> battery and BMS ('all joints at peak simultaneously' is physically impossible: G1's 46,062 W against a 421 Wh pack is 109 C) -> software control and the CAN-level low-side boards (22 nodes at 500 Hz on one bus is 130.9% load, so it must be split into four) -> simulation training and sim-to-real: domain randomization, sim2sim, zero calibration (ATOM01's 2.093 rad waist-yaw assembly offset, the |q| < 1e-2 rad acceptance gate, and write_motor_flash() being a no-op in three of the four motor drivers), plus 10 of 13 real failure modes being hardware calibration rather than simulation fidelity. Every figure comes from programmatic parsing of the five machines' public model files, deployment and calibration source, plus official vendor specifications, and is recomputable.

A 98-second build shows the open-source Hunter V2 (EC H130-V2) humanoid assembled from arm and leg joints, through torso integration, to indoor and outdoor walking.

Figure AI gave its retired second-generation humanoid F.02 a Terminator-style send-off: with the newer F.03 fleet arrived, maintaining old units no longer made sense, and dismantling each machine would eat engineering time needed for the next-gen F.04 while risking exposure of proprietary hardware and IP. After multiple steel mills declined, only a foundry in Imatra, Finland took the job. The team rigged a large air cushion at its San Jose HQ, captured stunt-performer motion data, and trained a new AI motion model in simulation so the robots could **autonomously leap from a second floor and land precisely inside a 75-ton electric arc furnace they had never surveyed** — then be melted into steel. The company cut a 3+ minute video homaging the T-800's thumbs-up furnace scene in Terminator 2 (complete with the thumbs-up), but blogger xiaohu finds it "a bit tragic and cruel — I worry about the robots of the future seeing this." Context: before retirement the F.02 fleet spent 11 months at BMW Spartanburg — 30,000+ X3s assisted, 90,000+ sheet-metal parts loaded (84 s cycles / 37 s loads / 99%+ accuracy), 1,250+ runtime hours, ~200 miles walked, zero injuries — with those battle scars feeding directly into Figure 03's redesign (no more wrist dynamic cabling or distribution boards). Only a few F.02 units remain at HQ for display.

Code2Skill turns 19,769 GitHub repos into 1,006,822 verified skill records; retrieved skills lift macro-average 11.7% and beat trajectory-derived banks on all seven shared benchmarks.

Dyna Robotics releases Dyna-2.1, the first Physical Agent achieving reliable super long-horizon whole-body autonomy, pairing brand-new semi-humanoid hardware with an agentic system built around Dyna-2. An uncut video shows it completing an entire hour-long laundry room workflow just like a human. The hardware is an upper torso with two agile arms (parallel-jaw grippers or dexterous hands) on four steerable wheels; the core is a vision-language orchestrator plus a whole-body controller that coherently manages driving, reaching, bending and lifting, running on Dyna own world-action model (Dyna-2 was pretrained on over 1M hours of egocentric human video, roughly 170 years of continuous waking time). Capabilities include autonomously operating washers and dryers, reaching deep into dryers for towels, folding and stacking by size, placing stacks at any height, and self-correcting errors such as picking up dropped items or retrying failed grasps. Downtime is measured by Mean Time Between Interventions (MTBI).

RLCDAlignBench spans 44 benchmarks and 7,193 instances. One generic question reaches median AUROC 0.886; soft probabilities beat argmax.

LLM agents increasingly answer questions over dynamic raw-document collections, where files may change before preprocessing, and relevant evidence (spans, sections, pages, or tables) is query-dependent. Existing retrieval-augmented approaches pre-materialize evidence via fixed chunking, embeddings, or persistent indexes: effective for lookup, yet costly, stale-prone, and committed to a granularity before the query is known. We formulate in-context search as Budgeted Evidence Localization over a latent evidence space induced by dynamic raw documents and propose LENS (Latent Evidence Exploration and Search), an index-free framework. Instead of pre-materializing the evidence space, LENS maintains a query-conditioned belief over candidate units, iteratively selecting candidates via complementary lexical, local, and exploratory proposal policies, updating the belief via an LLM relevance oracle, and narrowing toward high-posterior regions under a controllable budget. Evidence is consolidated into compact, source-grounded regions of interest and compressed into self-organizing knowledge clusters reused across related queries. On a controlled 500-question evaluation with matched corpus snapshots, LENS reaches 62.4% exact match and 84.8% evidence recall vs. 65.2% exact match but 50.4% evidence recall for a ReAct-style baseline. Across scales, LENS gives the strongest supporting-fact localization and answer grounding. On a fixed 150-question fullwiki subset over the raw Wikipedia dump with zero indexing, LENS and ReAct are nearly tied in official answer quality (43.3% vs. 42.7% EM), with LENS grounding more answers in retrieved evidence (84.0% vs. 70.7%). A no-retrieval Closed-Book reference highlights the contribution of model memory. LENS is query-ready after corpus changes, needs no preprocessing or persistent index, and preserves source-grounded evidence localization throughout.

Wildminder (@wildmindai) demos RunningHub's Qwen-Image-2.1 Consistent Lighting LoRA: blends an object seamlessly into the scene — no reflections or glare, enhanced specular highlights, natural shadows. Trigger word "pengyu", 160MiB weights, available via HuggingFace and ComfyUI (RunningHubAI/rh-qwen-image-2.1-lora-2104918997757157378).

@xiangxiang103 on Shanghai AI Lab's Shusheng DuanYan platform: sign-up grants a monthly base pack of 10 ink points (1 point covers up to 50M tokens — up to 500M total); in-house Intern-S2, Atria-Dawn-Preview and Agents-A1 are free for a limited time (0 points); plus 7 third-party models including DeepSeek-V4, GLM-5.3 and Kimi-K2.6; OpenAI- and Anthropic-compatible, so Claude Code and Codex work by swapping the base URL. The author tested it via API with workbuddy. Entry: discovery.intern-ai.org.cn/token-plan/home.
A coding-agent skill that shows three style previews for you to pick from, then builds a zero-dependency single-file HTML deck — and converts existing .pptx files with images and notes intact.
An open-source, project-based course (14 lectures, 8 projects, 15 languages) on building the instructions/state/verification/scope/lifecycle subsystems that make frontier models reliable inside real repos — includes reverse-engineered harness breakdowns of Pi, Claude Code, Codex and DeepSeek, plus loop- and graph-engineering advanced lectures. MIT licensed, ships a harness-creator skill and a zero-dependency audit script.
Foundation models: pre- and post-training, architecture and long context, inference serving and cost per token.
2026-09-22MiMo-V2.6 Deep Read: Six Days of Live RL, an AA Index of 46 and a Fully Open Self-Improvement RunXiaomi open-sources MiMo-V2.6 Pro/Flash: 30 live RL steps, ~750k trajectories at $850k/$2.62M; AA Index 46 tops open models, DeepSWE v1.1 gains +17/+14 out of sample; Vibe World, CUA, science and content demos plus 7k+ RL environments released.
2026-08-10Stealing Reasoning Traces from Proprietary LLM APIsLeading large language model providers now conceal their models' step-by-step reasoning, or chain-of-thought, to protect intellectual property and limit information leakage. Rather than storing these traces server-side, providers return them to the client as blocks of encrypted text, which the client passes back with each subsequent request. Building on prior research, we identify an architectural vulnerability: these encrypted blocks are fully compatible and interchangeable across different sessions, users, and models within a provider's ecosystem. We exploit this compatibility to develop a scalable decryption jailbreak. By injecting an encrypted reasoning trace from a given model into a weaker, and less safeguarded model from the same provider, we force it to decode and output the trace verbatim in plaintext, without ever jailbreaking the more capable model directly. This vulnerability enables four distinct attack vectors. First, it circumvents anti-distillation mechanisms, allowing adversaries to extract a proprietary model's reasoning, as we demonstrate across Anthropic, OpenAI, and Google. Second, it allows for large-scale private data extraction. Developers frequently share session logs publicly, unaware of contents of the encrypted blocks. By decoding 315,320 reasoning blocks scraped from public repositories, we recovered 367 Personally Identifiable Information (PII) artifacts and 182 credentials. Third, it inadvertently reveals hazardous information hidden within the reasoning process, even in cases where the model's final, visible output safely rejects a malicious request. Fourth, attackers can leverage this flaw to execute invisible prompt injections, embedding malicious payloads entirely within encrypted blocks to poison public agentic rollouts. Following responsible disclosure, we propose concrete cryptographic and system-level mitigations to secure client-side reasoning.Agent runtimes: plugins and sandboxes, tool surfaces, skill systems, long-term memory and self-improvement.
2026-09-30Caveman: The Token-Cutting Skill, Proxy and Middleware for Coding AgentsA MIT rule file makes agents talk like cavemen (code and errors never shortened), a local Go proxy compresses what agents read before each call with byte-exact originals recoverable, and middleware brings the same to your own app. JetBrains' 86-task A/B: -8.5% output tokens, quality flat; the repo's pinned 54-run suite: -33.2% input tokens, 18/18 checks passed.2026-09-23MiMo Code: Xiaomi terminal coding agent betting on memory and self-evolutionXiaomi open-sourced terminal-native AI coding assistant, TypeScript built with bun. The source is MIT but usage is additionally bound by USE_RESTRICTIONS.md, the MiMo terms of service and the trademark policy - read those before treating it as plain MIT. The README states it is a fork of OpenCode: it keeps the multi-provider, TUI, LSP, MCP and plugin core and adds persistent memory (SQLite FTS5 full-text search across four kinds - project MEMORY.md, session checkpoints, scratch notes and task progress - injected automatically on session resume), intelligent context management (near the limit it rebuilds from the latest checkpoint plus project memory plus task progress plus retained recent messages, ranked by importance against a token budget), goals and stop conditions (/goal sets the condition, and when the agent wants to stop a separate judge model assesses whether it was truly met, which targets optimistic early quitting), and deterministic JS workflows in a sandbox (compose splits independent tasks into isolated git worktrees with per-task TDD; also deep-research, fact-check with three-reviewer adversarial voting, and research-experiment with an anti-metric-gaming audit). Twenty built-in skills (arxiv, claude-code, codex, docx, pdf, pptx, xlsx, html-to-video, product-design and more), compatible with four skill roots - .agents/skills, .claude/skills, .codex/skills, .opencode/skills - where a user skill of the same name overrides the built-in. /dream and /distill are its signature: the first distils recent session trajectories into project memory and prunes stale entries, the second finds your repeated manual routines and packages high-confidence candidates into reusable skills. Model agnostic - the Xiaomi platform, Codex/ChatGPT OAuth, or any OpenAI-compatible endpoint. 13.4k stars. Not installed or run yet; memory-restore accuracy, judge-model effectiveness and the vendor-stated cache hit rates are unverified, so this is graded as pending reproduction.2026-09-23DeerFlow 2.0: ByteDance's super agent harnessByteDance's open-source super agent harness (82k+ stars). 2.0 is a ground-up rewrite sharing no code with 1.x: sub-agents, extensible skills, sandbox and file system, long-term memory, session goals, plus Manual Context Compaction that hands context management back to the operator. Positioning lifted from a deep-research framework to a runtime for any task. Not independently benchmarked; graded needs-reproduction.2026-09-23Hermes Agent: the open-source agent that turns self-improvement into a closed learning loopThe self-improving agent from Nous Research (248k+ stars). Its learning loop has five mechanisms: periodic memory nudges, autonomous skill creation after tasks, skills that self-improve in use, FTS5 session search with LLM summarization, and Honcho user modeling. One TUI, seven terminal backends (local/Docker/SSH/Singularity/Modal/Daytona/Vercel Sandbox), six chat platforms from a single gateway, no model lock-in. It doubles as a trajectory data-production apparatus. Self-improvement is falsifiable and no controlled experiment has been run, so this is graded as needing reproduction.Code generation, repair, refactoring, software engineering tasks
2026-10-01Inside Alibaba's 68-page AI Native R&D Handbook: coding is <1% of the pipeline — the real battlefield lies elsewhere2026 Handbook (AIDC Agent ) Agent 。 <1%(1 vs 3 ) 70% 90%+ Scrum Session→Commit→Change→Workitem Guardrail (UNKNOWN PASS) 。2026-09-30UI UX Pro Max: a 132k-star AI coding skill that generates complete design systems from 192 industry reasoning rulesAn open-source AI skill that injects design intelligence into coding agents: one request generates a complete design system (style + colors + typography + anti-patterns) from 192 industry reasoning rules, 79 searchable UI styles, 192 color palettes, 74 font pairings, 25 chart types and 22 tech stacks (React, Next.js, Astro, SwiftUI, Flutter, Three.js and more). The v2.0 flagship Design System Generator runs five parallel searches (product type, style, palette, landing pattern, typography) through a BM25-ranked reasoning engine to output pattern, style, colors, typography, effects, anti-patterns and a pre-delivery checklist (4.5:1 contrast, no emoji icons, cursor-pointer, prefers-reduced-motion, four responsive breakpoints). MIT licensed, one-line npm CLI install, works with Claude Code and other coding agents.Editor's takeIts premise is that models lack process discipline, not coding ability. The single best idea to steal is 'Rulings, not stalls': decide by default, and stop for a human only when an action is irreversible, touches a security-sensitive surface, leaks outside the worktree, or the plan is already too broken to continue.
v6.4.1 (2026-09-18), MIT, 289k stars — the largest single skill repository of its kind. It chains brainstorming, spec writing, git-worktree isolation, planning, subagent execution, red-green TDD, two-way code review and verification-before-completion into one trunk workflow, and the README lists 16+ harnesses: Claude Code, Codex, Cursor, Gemini CLI, OpenCode, Kimi Code, Devin CLI and more.
/plugin install superpowers@claude-plugins-officialEditor's takeIt sells parts rather than a whole process: code-review, diagnosing-bugs and resolving-merge-conflicts each stand alone. grill-me, grilling and wait-what turn the agent around to interrogate your requirements, and nothing else in this batch does that. CONTEXT.md at the repo root decouples project vocabulary from generic skills — the cleanest version of that idea in this batch.
The personal working set of the author of Total TypeScript; the repo describes itself as real engineering, not vibe coding. 38 SKILL.md files, primary language Shell, grouped as engineering (18), productivity (7), misc (4) and in-progress (9). It explicitly rejects takeover-style frameworks such as GSD, BMAD and Spec-Kit in favour of small, easy to adapt, composable.
claude plugins install mattpocock-skillsEditor's takeIf you want to know how a SKILL.md is actually supposed to be written, read this instead of any second-hand tutorial. Two things to know first: docx, pdf, pptx and xlsx are source-available rather than open source (the other 15 are Apache-2.0), and the repo carries its own disclaimer — these skills are demos, so test them in your environment before trusting them with anything critical.
Anthropic's own implementation, 177k stars: 19 skills under ./skills plus ./spec (the Agent Skills specification, now at agentskills.io) and ./template as a scaffold. Its definition of a skill is folders of instructions, scripts and resources loaded dynamically, and progressive disclosure comes straight from here.
/plugin marketplace add anthropics/skillsEditor's takeIts sharpest self-description is that models can write CSS but lack a taste database. Ten rule categories are ordered by severity, with Accessibility and Touch & Interaction first as CRITICAL and Charts last — that ordering is the stance. The Design System Generator runs five parallel searches, and the repo states plainly: do not persist unverified output, so a generated design system has to be checked before it lands in your repo.
v2.13.0, MIT, 129k stars, site uupm.cc. Instead of lecturing the model it ships structured catalogues the agent queries before writing code: 79 UI styles, 192 palettes with reasoning, 74 type pairings, 119 UX rules, 105 icon suggestions, 17 GSAP presets, 25 chart types, 22 stacks and 34 landing-page patterns.
npx ui-ux-pro-max-cli init --ai claudeEditor's takeThe design that matters most: every edge is labelled EXTRACTED or INFERRED, so you can tell what the source explicitly contains from what the tool inferred — which makes an agent's conclusions auditable, a hard requirement in production. It is not a vector index: no embeddings, no vector store, a real graph you can explain, path and query. In its own benchmark, building the graph costs zero LLM spend and ingest is an order of magnitude cheaper, a structural win from local parsing.
YC S26, Apache-2.0, 120k stars, Python 3.10+. /graphify . maps code, docs, PDFs, images and videos into a graph and emits a clickable graph.html, a human-readable GRAPH_REPORT.md and a machine-readable graph.json. Code goes through tree-sitter AST parsing (about 40 languages): deterministic, no LLM calls, nothing leaves the machine.
uv tool install graphifyy && graphify installEditor's takeThere is no one-liner that installs this repo; use it as a discovery layer. A list answers what exists, not what is worth running. The part you cannot get elsewhere is connect-apps-plugin: a single MCP endpoint to 1000+ integrations with authentication, team-level ACLs and audit logs, which is exactly the governance problem enterprises hit when agents touch SaaS.
A list maintained by Composio at 75k stars with an Apache-2.0 badge in the README. About 31 first-party skill directories physically exist here while the index references 864 SKILL.md files, and most of the difference is links out to external repositories. Coverage explicitly extends beyond Claude.ai and Claude Code to Codex, Cursor, Gemini CLI and Antigravity.
claude --plugin-dir ./connect-apps-pluginSource · artificialanalysis.aiofficial Intelligence Index v4.3synced Oct 5, 2026
Text-to-image, image editing, consistency and controllability
Video and speech/audio: physical and temporal coherence for text/image-to-video, TTS and voice cloning, music generation and ASR.
2026-09-28WB-WAM: Heterogeneous Body-Hand Pre-training for Humanoid Loco-ManipulationWB-WAM from Tsinghua IIIS, Xiong'an Institute of AI and University of Melbourne (Hang Zhao group) injects explicit whole-body action supervision into generative video pre-training: a shared 72-D physical action space (body 29 + root 3 + two hands 20+20) unifies partial annotations from 1,880.2 hours of nine-source heterogeneous video and motion data via masked flow matching, followed by PICO egocentric mid-training (22 h / 73 tasks, GMR retargeting + constrained IK + MINT) and real-robot post-training with a forward-kinematics loss on Unitree G1 with Wuji hands. It wins all seven HumanoidArena tasks (81.9% mean), reaches 84.0% on five real tasks vs OpenWAM's 80.0% (ACT 20%, Fast-WAM 6%), and PICO mid-training lets 30 real demos hit 73.8% — beating 100-demo direct training at 65.0%, a 70% cut in robot data.
2026-09-24Rolling-WAM: World Action Models with Rolling ImaginationWorld Action Models (WAMs) couple action generation with future visual prediction for robotic manipulation. However, completing the joint video-action denoising process at each replanning cycle incurs substantial latency, delaying action updates and limiting closed-loop responsiveness. We present Rolling-WAM, a formulation that distributes joint denoising across successive replanning cycles. Our method maintains a sliding window of video-action chunks at staggered noise levels. At each step, a rolling noise schedule fully denoises the imminent action chunk for execution, while partially refining farther-future chunks. As the window advances with new camera observations, the retained future chunks continue their denoising process. This distributes the computational cost over time while carrying an evolving visual-action context across chunk boundaries. Evaluations on LIBERO, RoboTwin, and a real-world Unitree G1 humanoid show that Rolling-WAM achieves competitive manipulation performance. By removing the need to denoise the entire prediction horizon from scratch, it delivers a 4.5x steady-state replanning speedup over standard joint WAMs.
2026-09-24BeyondRetarget: Learning Executable Humanoid Motions Directly from Monocular VideoBeyondRetarget removes the SMPL intermediate representation and maps monocular RGB video end-to-end to executable humanoid motions: 18 semantic keypoints with per-link non-uniform scaling form unified supervision, a shared representation decodes onto 8 humanoids, and contact-aware refinement yields 26.4mm error, zero collapse, and 193ms latency enabling real-time visual teleoperation.
2026-09-22MachEmbodied-U0: Unified Understanding and Generation Model for Embodied IntelligenceGeneral-purpose robot control requires models to understand task intent, identify where to interact, capture how the scene evolves, and generate precise actions. Vision-language-action models provide strong semantic priors but typically do not explicitly model scene dynamics, while world-action models couple visual prediction with control without necessarily exposing the task-relevant semantic and spatial structure needed for fine-grained manipulation. We present MachEmbodied-U0 (ME-U0), a unified embodied foundation model connecting understanding and generation experts through a Mixture-of-Transformers architecture. Subtask prediction and affordance grounding guide joint visual-dynamics and action generation via flow matching. Visual dynamics encompass future RGB, depth, surface normals, and optical flow, providing complementary supervision for appearance, geometry, and motion. Multi-rate Rotary Position Encoding (MRPE) aligns visual dynamics with fine-grained control. We pretrain ME-U0 on approximately 4,200 hours of curated demonstrations from robotic datasets and egocentric datasets. Using only the supervision natively available in each downstream benchmark, ME-U0 achieves an average score of 17.66 on the RoboDojo simulation benchmark and average success rates of 99.0\% and 82.5\% on LIBERO and LIBERO-Plus, respectively. We additionally validate ME-U0 on real-world robotic manipulation tasks, demonstrating its effectiveness beyond simulation. Without corresponding downstream supervision, ME-U0 further demonstrates zero-shot subtask prediction, affordance grounding, and visual dynamics on simulated and real-world observations. Overall, ME-U0 combines competitive downstream control performance with transferable task-grounding and visual-dynamics capabilities across simulation and the real world.3D asset generation, scene reconstruction, long-horizon world prediction
2026-10-04image-blaster: blast one image into a collidable 3D worldA set of Claude Code skills that turns one image into an explorable Gaussian-splat environment with a collider mesh and metric scale, plus interactable object meshes and sound effects, in under five minutes.
2026-10-02Introducing PROWL-2: Fidelity-Gated Dual Curricula — Jointly Learning Simulation and Decision-MakingOdyssey (with UCL AI Centre and the University of Basel) released PROWL-2 on Oct 1: the first framework coupling an agent curriculum and a world-model repair curriculum inside one continual training loop. Core insight: in imagination training, a high-learning-signal trajectory is ambiguous — a real policy weakness or just a wrong world-model prediction; prioritizing it indiscriminately reinforces the model's own hallucinations. A fidelity gate separates 'useful' from 'trustworthy': reliable high-potential imagined experience feeds the policy curriculum, unreliable rollouts — with their already-stored real continuations — go to a repair pool, fixed by a randomly-initialized, KL-anchored developer policy exploring the real environment, re-audited and readmitted once repaired. First place on all nine SMACv2 scenarios (+4-18% at 5v5, +20-91% at 10v10/10v11 over the backbone); big margins on hard MQE tasks — Gate-3 70.4% vs 29.8%, Shepherd-Hard 28.2% vs 7.3%, all other baselines at 0. Ablations show the curricula are not additive: repair alone is marginal, the ungated curriculum falls below the backbone in all nine scenarios, and the gate turns the same curriculum into the largest single gain. Architecture- and algorithm-agnostic; next steps point at humanoid coordination and long-horizon multi-agent games.Visual design, UI components, web and slide layout
Slides, Word, Excel, PDF generation and editing

A 10-step roadmap to a local LLM stack you use daily: memory sizing, model picking, Ollama install, quantization, hardware buying, serving as an API, small-model routing, and six ready-to-run workloads.

The same model behaves like two different models depending on the agent harness around it: GLM-5.2 scores 23% in one harness and 52% in another on SWE-bench Pro; in a controlled experiment swapping the harness moved scores 13 points while swapping the model moved only 2.5–5. A Hugging Face team (Adithya S Kolavi, Joel Niklaus, Lewis Tunstall, Leandro von Werra and colleagues, with Liquid AI) publishes an open fix: run agentic RL inside the real harnesses — Claude Code, Codex, OpenCode, Mini-SWE-Agent — with zero harness code changes. The stack: OpenEnv as the shared interface, a capture proxy that mints a session id per rollout and records exact token ids, per-token behavior logprobs and loss masks at the model endpoint (full-distribution sampling, top_p=1.0), Harbor for 40+ harness adapters and 26 sandbox backends, and TRL's Async GRPO consuming the masked training sequences. The numbers: LFM2.5-2.6B trained across four harnesses rises from 42.2% to 54.2% average pass@1 with gains under all four, and uses 31% fewer tool calls on tasks both it and the base model solved, while the OpenCode-only arm's gains stay mostly at home. The article also reports the SFT comparison (RL 54.6% vs SFT 47.5%), the full attribution of the earlier Qwen rise-and-decline (output-budget exhaustion, working without submitting, and a reward term that collapsed held-out accuracy from 0.740 to 0.178), and engineering boundaries like the proxy's concurrency ceiling.

Odyssey (with UCL AI Centre and the University of Basel) released PROWL-2 on Oct 1: the first framework coupling an agent curriculum and a world-model repair curriculum inside one continual training loop. Core insight: in imagination training, a high-learning-signal trajectory is ambiguous — a real policy weakness or just a wrong world-model prediction; prioritizing it indiscriminately reinforces the model's own hallucinations. A fidelity gate separates 'useful' from 'trustworthy': reliable high-potential imagined experience feeds the policy curriculum, unreliable rollouts — with their already-stored real continuations — go to a repair pool, fixed by a randomly-initialized, KL-anchored developer policy exploring the real environment, re-audited and readmitted once repaired. First place on all nine SMACv2 scenarios (+4-18% at 5v5, +20-91% at 10v10/10v11 over the backbone); big margins on hard MQE tasks — Gate-3 70.4% vs 29.8%, Shepherd-Hard 28.2% vs 7.3%, all other baselines at 0. Ablations show the curricula are not additive: repair alone is marginal, the ungated curriculum falls below the backbone in all nine scenarios, and the gate turns the same curriculum into the largest single gain. Architecture- and algorithm-agnostic; next steps point at humanoid coordination and long-horizon multi-agent games.
Under constructionThe games system is still under construction — everything listed starts in-page, and the library keeps growing.

A first-person match where every bot across from you is driven live by a policy this repo trained, the whole round stepping at a fixed 60 Hz: builtin, Rust/wasm or Rapier behind one physics interface, nothing scripted.
WASD move · mouse look · left-click fire · shift sprint · R reload · esc release

Unreal 5.5's first-person test map, running same-origin in the browser: the level, the materials, the lights and four weapon data assets are read out of the project's own .umap / .uasset bytes, six soldiers are driven by a four-state brain on one fixed 60 Hz step, and the sidebar lists every substitution the page had to make.
WASD move · mouse look · left-click fire · shift sprint · space jump · R reload · F pick up · esc release

Fifty-two professional tools, layers and masks, PSD files that open and save back - all in the browser, nothing to install.
retouchpi.com · 52 pro tools · PSD read and write · AI generate & edit

A 100 km 3D city that opens in seconds — crawling, blogs, papers, industry and investment analysis, all watchable in one digital twin.
100km city · opens in seconds · Live task stream · Agent-maintained

An SO-101 6-DoF arm running MuJoCo physics in your browser: joint teleoperation, IK end-effector dragging and gripper pick-and-place into a basket. No install.
MuJoCo WASM physics · SO-101 · 6 DoF · Contacts / telemetry HUD

Take a humanoid joint module apart layer by layer: brushless motor, magnetic encoder, planetary / harmonic / cycloidal drive and output flange. Switch architectures on one page — exploded view plus analytic kinematics.
Three gearbox types · Planetary / harmonic / cycloidal · Analytic kinematics · explode

Matcha-TTS mixed zh/en speech synthesis: server-side synthesis with sentence-streamed playback and full-article read-aloud for papers and blogs. Sign-in required.
Matcha-TTS · server-side · Mixed zh / en · Full-article read-aloud

A caring coding sprite on your desktop and a cockpit that never stops: the work resumes itself after a crash or a reboot. Every coding terminal, ssh session and local shell in one tree, every start, finish and permission request reported in time. Open source, one-line install.
Your caring coding desktop sprite · Resumes after crash & reboot · First-rate terminal / ssh / shell

Built for games and for physical AI: an ultra-realistic, ultra-high-performance simulation environment with ultra-low-cost procedural data; Unreal-grade visuals open and play in the browser, bridging the physical world and intelligent agents.
Games meet physical AI · Unreal-grade on the web · Open and play