0:40As an Amazon Associate, we earn from qualifying purchases.
As an Amazon Associate, we earn from qualifying purchases.

Odyssey (with UCL AI Centre and the University of Basel) released PROWL-2 on Oct 1: the first framework coupling an agent curriculum and a world-model repair curriculum inside one continual training loop. Core insight: in imagination training, a high-learning-signal trajectory is ambiguous — a real policy weakness or just a wrong world-model prediction; prioritizing it indiscriminately reinforces the model's own hallucinations. A fidelity gate separates 'useful' from 'trustworthy': reliable high-potential imagined experience feeds the policy curriculum, unreliable rollouts — with their already-stored real continuations — go to a repair pool, fixed by a randomly-initialized, KL-anchored developer policy exploring the real environment, re-audited and readmitted once repaired. First place on all nine SMACv2 scenarios (+4-18% at 5v5, +20-91% at 10v10/10v11 over the backbone); big margins on hard MQE tasks — Gate-3 70.4% vs 29.8%, Shepherd-Hard 28.2% vs 7.3%, all other baselines at 0. Ablations show the curricula are not additive: repair alone is marginal, the ungated curriculum falls below the backbone in all nine scenarios, and the gate turns the same curriculum into the largest single gain. Architecture- and algorithm-agnostic; next steps point at humanoid coordination and long-horizon multi-agent games.

A deep read of TypeSafe's first System One Model: three question primitives, the economics of parallel calls, confidence-gated routing, eval caveats, eight jagged edges, and the OpenJev local repro.

Xiaomi open-sources MiMo-V2.6 Pro/Flash: 30 live RL steps, ~750k trajectories at $850k/$2.62M; AA Index 46 tops open models, DeepSWE v1.1 gains +17/+14 out of sample; Vibe World, CUA, science and content demos plus 7k+ RL environments released.

First look at the LiberAI preview model Liber-0: fully autonomous dexterous manipulation of everyday objects, including recovering and retrying when execution goes wrong, as human experience scales.

Tsinghua University and Beihang University present BSC-Nav (Brain-inspired Spatial Cognition for Navigation). The paper argues that existing embodied agents — whether end-to-end RL or "MLLM plus modular pipeline" — are fundamentally reactive and stateless: they process an observation and discard it, lacking any durable internal model of space, which yields fragmented knowledge, short-sighted planning and poor generalization. The authors borrow the answer from neuroscience, where spatial knowledge consolidates into three interconnected forms: landmarks, route knowledge, and survey knowledge. BSC-Nav instantiates these computationally in three modules. Landmark memory stores 4-tuples (world coordinates, open-vocabulary category, detection confidence, GPT-4o contextual description) with a spatial-overlap set plus confidence-weighted fusion for deduplication. The cognitive map extracts DINO-v2 patch features, projects them through inverse perspective projection and cascaded coordinate transforms into a voxel grid, and adopts a free-energy-principle-inspired surprise-driven update: a new feature is written when its mean distance to features in the n-hop neighborhood exceeds a threshold, replacing the lowest-surprise entry when the buffer is full, preserving cross-viewpoint diversity while bounding storage. Working memory retrieves hierarchically by instruction complexity — simple targets use text-only GPT-4 reasoning over landmark memory (even inferring unrecorded targets from co-located landmarks), while complex targets first have descriptions refined by GPT-4o, then "imagine" the appearance via Stable Diffusion 3.5, encoded by DINO-v2 and center-distance weighted pooled to query the cognitive map (imagine-then-localize), with similarity-weighted DBSCAN yielding candidate coordinates. Rather than greedily taking the highest confidence, candidates are ordered by H_i = lambda*p_i + (1-lambda)(1 - d_i/d_max). Across 62 MP3D/HM3D scenes and 8,195 episodes: OGN reaches 78.5% SR on HM3D (24.0 points above SOTA UniGoal), OVON zero-shot beats supervised DAgRL, IIN reaches 71.4%; SPL gains are even more consistent (IIN 57.2% vs 23.7%). On A-EQA it achieves the highest LLM-Match of 54.6, still trailing humans by 27.5. Real-world deployment on a custom platform (AgileX Ranger-mini-3.0 chassis, Franka Research 3 arm, RealSense D435i) ran 75 episodes in a ~200 m² two-floor space, with IIN reaching 100% SR on 4 of 5 targets and reliable localization to semantically plausible regions even on failure, plus long-horizon navigation-plus-manipulation demos such as "make breakfast" over three open-vocabulary objects. The paper proposes an embodied Turing test for spatial cognition probing three dimensions: real-time construction of reusable spatial representations, abstraction from sparse partial observations, and translation of high-level goals into actionable spatial plans.

RLCDAlignBench spans 44 benchmarks and 7,193 instances. One generic question reaches median AUROC 0.886; soft probabilities beat argmax.

ModularRSI is a benchmark-disjoint, contrastive, and modular framework for generalizable agent-harness evolution. It diagnoses recurring failures from paired successful and failed trajectories, evolves five functional modules independently, and integrates validated changes into a unified harness. On Terminal-Bench 2.0 and SWE-Bench Verified, the evolved harness improves unseen in-domain and cross-domain tasks and transfers across foundation models.

Haocheng Xi (UC Berkeley) released FreeVideo (FlashML-org, Apache-2.0): a local inference engine for MiniMax H3 on consumer GPUs, running the 33B omni-modal video model with as little as 8GB VRAM and 16GB RAM. It builds on OpenVDN's 8-step VDN-H3 model — a hybrid-attention Video DeltaNet rework of H3 that replaces quadratic long-range attention with a linear Video Delta Attention branch. The core is hardware-adaptive execution planning: VRAM, system memory, and disk are scheduled as a single memory hierarchy; the engine picks native FP8 or FP8-storage-with-BF16-compute per GPU architecture and auto-probes available attention kernels (SageAttention, etc.); weight streaming, asynchronous prefetching, and chunked computation keep peak memory at 8GB. Shipped as a ComfyUI plugin: the Windows launcher installs and starts ComfyUI with a dedicated creative workspace supporting two-pass sampling, batch generation, and history; power users can switch to node view to add LoRAs or customize workflows; Linux also offers a CLI (./freevideo generate). Multimodal inputs cover text prompts, first/last frames, and image/video/audio references. Weights carry the MiniMax H3 Community License (territorial and acceptable-use restrictions); the code is open source

A set of Claude Code skills that turns one image into an explorable Gaussian-splat environment with a collider mesh and metric scale, plus interactable object meshes and sound effects, in under five minutes.

EPFL open-sources MuscleMimic, a JAX plus MuJoCo musculoskeletal benchmark: Hill-type muscle models actuate a 354-muscle full body, and a single generalist policy handles chained movements such as a golf swing zero-shot.
NVIDIA's inference optimisation library for LLMs and visual generation models, fully open source since 2025-03-22. Its 28 tech blogs form an auditable engineering ledger: Skip Softmax and sparse attention for long context, the three-part expert parallelism series with one-sided AlltoAll over NVLink, DWDP for NVL72, guided decoding cooperating with speculative decoding, inference-time compute, and evaluating agentic serving with trace replay and job-level metrics. Claims include Llama 4 Maverick above 1,000 TPS per user and 40,000+ tok/s on B200; Bing and NAVER Place run it in production.
CopilotKit open-sourced OpenDots (MIT) on Oct 1: a fully self-hostable template for always-on AI coworkers — its answer to OpenAI's closed Dots (powered by GPT-6 Astra). Each Dot is a specialist agent with a name, role, instructions and tool allowlist, plus its own computer via OpenBot's container supervisor (browser profile and workspace files persist across stop/start). Dots move seamlessly between text chat, realtime voice calls and Slack: calls pair a separate compute agent so long work keeps running mid-conversation, all sharing the same context and tool permissions. AG-UI carries streamed messages, tool calls and agent state between any agent harness and the UI; human-in-the-loop cards pause tool calls for approval; Spaces and Pages hold working documents. Built on TanStack AI + CopilotKit React SDK + Threads + Channels SDK. Clone and customize — enterprise-ready.
Foundation models: pre- and post-training, architecture and long context, inference serving and cost per token.
2026-08-17LENS: In-Context Search via Latent Evidence Exploration over Dynamic Raw DocumentsLLM agents increasingly answer questions over dynamic raw-document collections, where files may change before preprocessing, and relevant evidence (spans, sections, pages, or tables) is query-dependent. Existing retrieval-augmented approaches pre-materialize evidence via fixed chunking, embeddings, or persistent indexes: effective for lookup, yet costly, stale-prone, and committed to a granularity before the query is known. We formulate in-context search as Budgeted Evidence Localization over a latent evidence space induced by dynamic raw documents and propose LENS (Latent Evidence Exploration and Search), an index-free framework. Instead of pre-materializing the evidence space, LENS maintains a query-conditioned belief over candidate units, iteratively selecting candidates via complementary lexical, local, and exploratory proposal policies, updating the belief via an LLM relevance oracle, and narrowing toward high-posterior regions under a controllable budget. Evidence is consolidated into compact, source-grounded regions of interest and compressed into self-organizing knowledge clusters reused across related queries. On a controlled 500-question evaluation with matched corpus snapshots, LENS reaches 62.4% exact match and 84.8% evidence recall vs. 65.2% exact match but 50.4% evidence recall for a ReAct-style baseline. Across scales, LENS gives the strongest supporting-fact localization and answer grounding. On a fixed 150-question fullwiki subset over the raw Wikipedia dump with zero indexing, LENS and ReAct are nearly tied in official answer quality (43.3% vs. 42.7% EM), with LENS grounding more answers in retrieved evidence (84.0% vs. 70.7%). A no-retrieval Closed-Book reference highlights the contribution of model memory. LENS is query-ready after corpus changes, needs no preprocessing or persistent index, and preserves source-grounded evidence localization throughout.
2026-08-10Stealing Reasoning Traces from Proprietary LLM APIsLeading large language model providers now conceal their models' step-by-step reasoning, or chain-of-thought, to protect intellectual property and limit information leakage. Rather than storing these traces server-side, providers return them to the client as blocks of encrypted text, which the client passes back with each subsequent request. Building on prior research, we identify an architectural vulnerability: these encrypted blocks are fully compatible and interchangeable across different sessions, users, and models within a provider's ecosystem. We exploit this compatibility to develop a scalable decryption jailbreak. By injecting an encrypted reasoning trace from a given model into a weaker, and less safeguarded model from the same provider, we force it to decode and output the trace verbatim in plaintext, without ever jailbreaking the more capable model directly. This vulnerability enables four distinct attack vectors. First, it circumvents anti-distillation mechanisms, allowing adversaries to extract a proprietary model's reasoning, as we demonstrate across Anthropic, OpenAI, and Google. Second, it allows for large-scale private data extraction. Developers frequently share session logs publicly, unaware of contents of the encrypted blocks. By decoding 315,320 reasoning blocks scraped from public repositories, we recovered 367 Personally Identifiable Information (PII) artifacts and 182 credentials. Third, it inadvertently reveals hazardous information hidden within the reasoning process, even in cases where the model's final, visible output safely rejects a malicious request. Fourth, attackers can leverage this flaw to execute invisible prompt injections, embedding malicious payloads entirely within encrypted blocks to poison public agentic rollouts. Following responsible disclosure, we propose concrete cryptographic and system-level mitigations to secure client-side reasoning.Agent runtimes: plugins and sandboxes, tool surfaces, skill systems, long-term memory and self-improvement.
2026-09-30Caveman: The Token-Cutting Skill, Proxy and Middleware for Coding AgentsA MIT rule file makes agents talk like cavemen (code and errors never shortened), a local Go proxy compresses what agents read before each call with byte-exact originals recoverable, and middleware brings the same to your own app. JetBrains' 86-task A/B: -8.5% output tokens, quality flat; the repo's pinned 54-run suite: -33.2% input tokens, 18/18 checks passed.2026-09-23MiMo Code: Xiaomi terminal coding agent betting on memory and self-evolutionXiaomi open-sourced terminal-native AI coding assistant, TypeScript built with bun. The source is MIT but usage is additionally bound by USE_RESTRICTIONS.md, the MiMo terms of service and the trademark policy - read those before treating it as plain MIT. The README states it is a fork of OpenCode: it keeps the multi-provider, TUI, LSP, MCP and plugin core and adds persistent memory (SQLite FTS5 full-text search across four kinds - project MEMORY.md, session checkpoints, scratch notes and task progress - injected automatically on session resume), intelligent context management (near the limit it rebuilds from the latest checkpoint plus project memory plus task progress plus retained recent messages, ranked by importance against a token budget), goals and stop conditions (/goal sets the condition, and when the agent wants to stop a separate judge model assesses whether it was truly met, which targets optimistic early quitting), and deterministic JS workflows in a sandbox (compose splits independent tasks into isolated git worktrees with per-task TDD; also deep-research, fact-check with three-reviewer adversarial voting, and research-experiment with an anti-metric-gaming audit). Twenty built-in skills (arxiv, claude-code, codex, docx, pdf, pptx, xlsx, html-to-video, product-design and more), compatible with four skill roots - .agents/skills, .claude/skills, .codex/skills, .opencode/skills - where a user skill of the same name overrides the built-in. /dream and /distill are its signature: the first distils recent session trajectories into project memory and prunes stale entries, the second finds your repeated manual routines and packages high-confidence candidates into reusable skills. Model agnostic - the Xiaomi platform, Codex/ChatGPT OAuth, or any OpenAI-compatible endpoint. 13.4k stars. Not installed or run yet; memory-restore accuracy, judge-model effectiveness and the vendor-stated cache hit rates are unverified, so this is graded as pending reproduction.2026-09-23DeerFlow 2.0: ByteDance's super agent harnessByteDance's open-source super agent harness (82k+ stars). 2.0 is a ground-up rewrite sharing no code with 1.x: sub-agents, extensible skills, sandbox and file system, long-term memory, session goals, plus Manual Context Compaction that hands context management back to the operator. Positioning lifted from a deep-research framework to a runtime for any task. Not independently benchmarked; graded needs-reproduction.
2026-09-16OpenResearch: A Local-First Workspace That Turns Coding Agents into Research AgentsOpenResearch (orx) is the local-first research workspace open-sourced by alphaXiv: a single Rust binary that turns Claude Code, Codex, OpenCode, Cursor, or Google Antigravity into research agents that review literature, form hypotheses, run experiments, and produce artifacts. It organizes work as a git-native experiment tree under a fixed run contract, drives experiments through a repair/refill/promote/stop autoresearch loop, and treats run logs and manifests as the only evidence. Written from a close read of the v0.2.13 source and a hands-on run of the bundled nanochat demo.Code generation, repair, refactoring, software engineering tasks
2026-10-01Inside Alibaba's 68-page AI Native R&D Handbook: coding is <1% of the pipeline — the real battlefield lies elsewhere2026 Handbook (AIDC Agent ) Agent 。 <1%(1 vs 3 ) 70% 90%+ Scrum Session→Commit→Change→Workitem Guardrail (UNKNOWN PASS) 。2026-09-30UI UX Pro Max: a 132k-star AI coding skill that generates complete design systems from 192 industry reasoning rulesAn open-source AI skill that injects design intelligence into coding agents: one request generates a complete design system (style + colors + typography + anti-patterns) from 192 industry reasoning rules, 79 searchable UI styles, 192 color palettes, 74 font pairings, 25 chart types and 22 tech stacks (React, Next.js, Astro, SwiftUI, Flutter, Three.js and more). The v2.0 flagship Design System Generator runs five parallel searches (product type, style, palette, landing pattern, typography) through a BM25-ranked reasoning engine to output pattern, style, colors, typography, effects, anti-patterns and a pre-delivery checklist (4.5:1 contrast, no emoji icons, cursor-pointer, prefers-reduced-motion, four responsive breakpoints). MIT licensed, one-line npm CLI install, works with Claude Code and other coding agents.Editor's takeIts premise is that models lack process discipline, not coding ability. The single best idea to steal is 'Rulings, not stalls': decide by default, and stop for a human only when an action is irreversible, touches a security-sensitive surface, leaks outside the worktree, or the plan is already too broken to continue.
v6.4.1 (2026-09-18), MIT, 289k stars — the largest single skill repository of its kind. It chains brainstorming, spec writing, git-worktree isolation, planning, subagent execution, red-green TDD, two-way code review and verification-before-completion into one trunk workflow, and the README lists 16+ harnesses: Claude Code, Codex, Cursor, Gemini CLI, OpenCode, Kimi Code, Devin CLI and more.
/plugin install superpowers@claude-plugins-officialEditor's takeIt sells parts rather than a whole process: code-review, diagnosing-bugs and resolving-merge-conflicts each stand alone. grill-me, grilling and wait-what turn the agent around to interrogate your requirements, and nothing else in this batch does that. CONTEXT.md at the repo root decouples project vocabulary from generic skills — the cleanest version of that idea in this batch.
The personal working set of the author of Total TypeScript; the repo describes itself as real engineering, not vibe coding. 38 SKILL.md files, primary language Shell, grouped as engineering (18), productivity (7), misc (4) and in-progress (9). It explicitly rejects takeover-style frameworks such as GSD, BMAD and Spec-Kit in favour of small, easy to adapt, composable.
claude plugins install mattpocock-skillsEditor's takeIf you want to know how a SKILL.md is actually supposed to be written, read this instead of any second-hand tutorial. Two things to know first: docx, pdf, pptx and xlsx are source-available rather than open source (the other 15 are Apache-2.0), and the repo carries its own disclaimer — these skills are demos, so test them in your environment before trusting them with anything critical.
Anthropic's own implementation, 177k stars: 19 skills under ./skills plus ./spec (the Agent Skills specification, now at agentskills.io) and ./template as a scaffold. Its definition of a skill is folders of instructions, scripts and resources loaded dynamically, and progressive disclosure comes straight from here.
/plugin marketplace add anthropics/skillsEditor's takeIts sharpest self-description is that models can write CSS but lack a taste database. Ten rule categories are ordered by severity, with Accessibility and Touch & Interaction first as CRITICAL and Charts last — that ordering is the stance. The Design System Generator runs five parallel searches, and the repo states plainly: do not persist unverified output, so a generated design system has to be checked before it lands in your repo.
v2.13.0, MIT, 129k stars, site uupm.cc. Instead of lecturing the model it ships structured catalogues the agent queries before writing code: 79 UI styles, 192 palettes with reasoning, 74 type pairings, 119 UX rules, 105 icon suggestions, 17 GSAP presets, 25 chart types, 22 stacks and 34 landing-page patterns.
npx ui-ux-pro-max-cli init --ai claudeEditor's takeThe design that matters most: every edge is labelled EXTRACTED or INFERRED, so you can tell what the source explicitly contains from what the tool inferred — which makes an agent's conclusions auditable, a hard requirement in production. It is not a vector index: no embeddings, no vector store, a real graph you can explain, path and query. In its own benchmark, building the graph costs zero LLM spend and ingest is an order of magnitude cheaper, a structural win from local parsing.
YC S26, Apache-2.0, 120k stars, Python 3.10+. /graphify . maps code, docs, PDFs, images and videos into a graph and emits a clickable graph.html, a human-readable GRAPH_REPORT.md and a machine-readable graph.json. Code goes through tree-sitter AST parsing (about 40 languages): deterministic, no LLM calls, nothing leaves the machine.
uv tool install graphifyy && graphify installEditor's takeThere is no one-liner that installs this repo; use it as a discovery layer. A list answers what exists, not what is worth running. The part you cannot get elsewhere is connect-apps-plugin: a single MCP endpoint to 1000+ integrations with authentication, team-level ACLs and audit logs, which is exactly the governance problem enterprises hit when agents touch SaaS.
A list maintained by Composio at 75k stars with an Apache-2.0 badge in the README. About 31 first-party skill directories physically exist here while the index references 864 SKILL.md files, and most of the difference is links out to external repositories. Coverage explicitly extends beyond Claude.ai and Claude Code to Codex, Cursor, Gemini CLI and Antigravity.
claude --plugin-dir ./connect-apps-pluginSource · artificialanalysis.aiofficial Intelligence Index v4.3synced Oct 4, 2026
Text-to-image, image editing, consistency and controllability
Video and speech/audio: physical and temporal coherence for text/image-to-video, TTS and voice cloning, music generation and ASR.
2026-09-28WB-WAM: Heterogeneous Body-Hand Pre-training for Humanoid Loco-ManipulationWB-WAM from Tsinghua IIIS, Xiong'an Institute of AI and University of Melbourne (Hang Zhao group) injects explicit whole-body action supervision into generative video pre-training: a shared 72-D physical action space (body 29 + root 3 + two hands 20+20) unifies partial annotations from 1,880.2 hours of nine-source heterogeneous video and motion data via masked flow matching, followed by PICO egocentric mid-training (22 h / 73 tasks, GMR retargeting + constrained IK + MINT) and real-robot post-training with a forward-kinematics loss on Unitree G1 with Wuji hands. It wins all seven HumanoidArena tasks (81.9% mean), reaches 84.0% on five real tasks vs OpenWAM's 80.0% (ACT 20%, Fast-WAM 6%), and PICO mid-training lets 30 real demos hit 73.8% — beating 100-demo direct training at 65.0%, a 70% cut in robot data.
2026-09-24Rolling-WAM: World Action Models with Rolling ImaginationWorld Action Models (WAMs) couple action generation with future visual prediction for robotic manipulation. However, completing the joint video-action denoising process at each replanning cycle incurs substantial latency, delaying action updates and limiting closed-loop responsiveness. We present Rolling-WAM, a formulation that distributes joint denoising across successive replanning cycles. Our method maintains a sliding window of video-action chunks at staggered noise levels. At each step, a rolling noise schedule fully denoises the imminent action chunk for execution, while partially refining farther-future chunks. As the window advances with new camera observations, the retained future chunks continue their denoising process. This distributes the computational cost over time while carrying an evolving visual-action context across chunk boundaries. Evaluations on LIBERO, RoboTwin, and a real-world Unitree G1 humanoid show that Rolling-WAM achieves competitive manipulation performance. By removing the need to denoise the entire prediction horizon from scratch, it delivers a 4.5x steady-state replanning speedup over standard joint WAMs.
2026-09-24BeyondRetarget: Learning Executable Humanoid Motions Directly from Monocular VideoBeyondRetarget removes the SMPL intermediate representation and maps monocular RGB video end-to-end to executable humanoid motions: 18 semantic keypoints with per-link non-uniform scaling form unified supervision, a shared representation decodes onto 8 humanoids, and contact-aware refinement yields 26.4mm error, zero collapse, and 193ms latency enabling real-time visual teleoperation.
2026-09-22MachEmbodied-U0: Unified Understanding and Generation Model for Embodied IntelligenceGeneral-purpose robot control requires models to understand task intent, identify where to interact, capture how the scene evolves, and generate precise actions. Vision-language-action models provide strong semantic priors but typically do not explicitly model scene dynamics, while world-action models couple visual prediction with control without necessarily exposing the task-relevant semantic and spatial structure needed for fine-grained manipulation. We present MachEmbodied-U0 (ME-U0), a unified embodied foundation model connecting understanding and generation experts through a Mixture-of-Transformers architecture. Subtask prediction and affordance grounding guide joint visual-dynamics and action generation via flow matching. Visual dynamics encompass future RGB, depth, surface normals, and optical flow, providing complementary supervision for appearance, geometry, and motion. Multi-rate Rotary Position Encoding (MRPE) aligns visual dynamics with fine-grained control. We pretrain ME-U0 on approximately 4,200 hours of curated demonstrations from robotic datasets and egocentric datasets. Using only the supervision natively available in each downstream benchmark, ME-U0 achieves an average score of 17.66 on the RoboDojo simulation benchmark and average success rates of 99.0\% and 82.5\% on LIBERO and LIBERO-Plus, respectively. We additionally validate ME-U0 on real-world robotic manipulation tasks, demonstrating its effectiveness beyond simulation. Without corresponding downstream supervision, ME-U0 further demonstrates zero-shot subtask prediction, affordance grounding, and visual dynamics on simulated and real-world observations. Overall, ME-U0 combines competitive downstream control performance with transferable task-grounding and visual-dynamics capabilities across simulation and the real world.3D asset generation, scene reconstruction, long-horizon world prediction
Visual design, UI components, web and slide layout
Slides, Word, Excel, PDF generation and editing

The same model behaves like two different models depending on the agent harness around it: GLM-5.2 scores 23% in one harness and 52% in another on SWE-bench Pro; in a controlled experiment swapping the harness moved scores 13 points while swapping the model moved only 2.5–5. A Hugging Face team (Adithya S Kolavi, Joel Niklaus, Lewis Tunstall, Leandro von Werra and colleagues, with Liquid AI) publishes an open fix: run agentic RL inside the real harnesses — Claude Code, Codex, OpenCode, Mini-SWE-Agent — with zero harness code changes. The stack: OpenEnv as the shared interface, a capture proxy that mints a session id per rollout and records exact token ids, per-token behavior logprobs and loss masks at the model endpoint (full-distribution sampling, top_p=1.0), Harbor for 40+ harness adapters and 26 sandbox backends, and TRL's Async GRPO consuming the masked training sequences. The numbers: LFM2.5-2.6B trained across four harnesses rises from 42.2% to 54.2% average pass@1 with gains under all four, and uses 31% fewer tool calls on tasks both it and the base model solved, while the OpenCode-only arm's gains stay mostly at home. The article also reports the SFT comparison (RL 54.6% vs SFT 47.5%), the full attribution of the earlier Qwen rise-and-decline (output-budget exhaustion, working without submitting, and a reward term that collapsed held-out accuracy from 0.740 to 0.178), and engineering boundaries like the proxy's concurrency ceiling.

Odyssey (with UCL AI Centre and the University of Basel) released PROWL-2 on Oct 1: the first framework coupling an agent curriculum and a world-model repair curriculum inside one continual training loop. Core insight: in imagination training, a high-learning-signal trajectory is ambiguous — a real policy weakness or just a wrong world-model prediction; prioritizing it indiscriminately reinforces the model's own hallucinations. A fidelity gate separates 'useful' from 'trustworthy': reliable high-potential imagined experience feeds the policy curriculum, unreliable rollouts — with their already-stored real continuations — go to a repair pool, fixed by a randomly-initialized, KL-anchored developer policy exploring the real environment, re-audited and readmitted once repaired. First place on all nine SMACv2 scenarios (+4-18% at 5v5, +20-91% at 10v10/10v11 over the backbone); big margins on hard MQE tasks — Gate-3 70.4% vs 29.8%, Shepherd-Hard 28.2% vs 7.3%, all other baselines at 0. Ablations show the curricula are not additive: repair alone is marginal, the ungated curriculum falls below the backbone in all nine scenarios, and the gate turns the same curriculum into the largest single gain. Architecture- and algorithm-agnostic; next steps point at humanoid coordination and long-horizon multi-agent games.

A deep read of chapter 15 of wdfk-prog's ROS tutorial series: why you need robot_localization even when the driver already publishes /odom. Covers EKF vs UKF predict-correct over 15 states (31 sigma points), the real roles of P/Q/R covariances, the data-provenance thinking behind 15-bit _config switches, the semantic gap between differential (velocity) and relative (pose), TF ownership refactoring (/odom→/odom/raw, EKF takes over odom→base_link), world_frame and sensor_timeout time semantics, and the contract-first mindset: every observation must answer who measured it, in which frame, at which time, and how trustworthy it is.
Under constructionThe games system is still under construction — everything listed starts in-page, and the library keeps growing.

A first-person match where every bot across from you is driven live by a policy this repo trained, the whole round stepping at a fixed 60 Hz: builtin, Rust/wasm or Rapier behind one physics interface, nothing scripted.
WASD move · mouse look · left-click fire · shift sprint · R reload · esc release

Unreal 5.5's first-person test map, running same-origin in the browser: the level, the materials, the lights and four weapon data assets are read out of the project's own .umap / .uasset bytes, six soldiers are driven by a four-state brain on one fixed 60 Hz step, and the sidebar lists every substitution the page had to make.
WASD move · mouse look · left-click fire · shift sprint · space jump · R reload · F pick up · esc release

A 100 km 3D city that opens in seconds — crawling, blogs, papers, industry and investment analysis, all watchable in one digital twin.
100km city · opens in seconds · Live task stream · Agent-maintained

An SO-101 6-DoF arm running MuJoCo physics in your browser: joint teleoperation, IK end-effector dragging and gripper pick-and-place into a basket. No install.
MuJoCo WASM physics · SO-101 · 6 DoF · Contacts / telemetry HUD

Take a humanoid joint module apart layer by layer: brushless motor, magnetic encoder, planetary / harmonic / cycloidal drive and output flange. Switch architectures on one page — exploded view plus analytic kinematics.
Three gearbox types · Planetary / harmonic / cycloidal · Analytic kinematics · explode

Matcha-TTS mixed zh/en speech synthesis: server-side synthesis with sentence-streamed playback and full-article read-aloud for papers and blogs. Sign-in required.
Matcha-TTS · server-side · Mixed zh / en · Full-article read-aloud

A caring coding sprite on your desktop and a cockpit that never stops: the work resumes itself after a crash or a reboot. Every coding terminal, ssh session and local shell in one tree, every start, finish and permission request reported in time. Open source, one-line install.
Your caring coding desktop sprite · Resumes after crash & reboot · First-rate terminal / ssh / shell

Built for games and for physical AI: an ultra-realistic, ultra-high-performance simulation environment with ultra-low-cost procedural data; Unreal-grade visuals open and play in the browser, bridging the physical world and intelligent agents.
Games meet physical AI · Unreal-grade on the web · Open and play

Fifty-two professional tools, layers and masks, PSD files that open and save back - all in the browser, nothing to install.
retouchpi.com · 52 pro tools · PSD read and write · AI generate & edit