As an Amazon Associate, we earn from qualifying purchases.
As an Amazon Associate, we earn from qualifying purchases.






ScientistTwo from Google Cloud AI Research is a fully autonomous multi-agent research framework: it generates hypotheses targeting limitations of the human state of the art, validates subset-first before full-set scaling, refines ideas evolutionarily and via ablation-driven critique, and closes the loop with simulated peer review, rebuttal and meta-review to produce papers plus reproducible codebases end to end. Across 107 ICLR/ICML/NeurIPS-level challenges it completes 86 (80.4%) with 25.2% average relative gain; its papers earn 7.5/10 and 91.9% acceptance on ScholarPeer and 72.1% under an independent reviewer, at ~2.5 days and $3,765 per task.

Tsinghua University and Beihang University present BSC-Nav (Brain-inspired Spatial Cognition for Navigation). The paper argues that existing embodied agents — whether end-to-end RL or "MLLM plus modular pipeline" — are fundamentally reactive and stateless: they process an observation and discard it, lacking any durable internal model of space, which yields fragmented knowledge, short-sighted planning and poor generalization. The authors borrow the answer from neuroscience, where spatial knowledge consolidates into three interconnected forms: landmarks, route knowledge, and survey knowledge. BSC-Nav instantiates these computationally in three modules. Landmark memory stores 4-tuples (world coordinates, open-vocabulary category, detection confidence, GPT-4o contextual description) with a spatial-overlap set plus confidence-weighted fusion for deduplication. The cognitive map extracts DINO-v2 patch features, projects them through inverse perspective projection and cascaded coordinate transforms into a voxel grid, and adopts a free-energy-principle-inspired surprise-driven update: a new feature is written when its mean distance to features in the n-hop neighborhood exceeds a threshold, replacing the lowest-surprise entry when the buffer is full, preserving cross-viewpoint diversity while bounding storage. Working memory retrieves hierarchically by instruction complexity — simple targets use text-only GPT-4 reasoning over landmark memory (even inferring unrecorded targets from co-located landmarks), while complex targets first have descriptions refined by GPT-4o, then "imagine" the appearance via Stable Diffusion 3.5, encoded by DINO-v2 and center-distance weighted pooled to query the cognitive map (imagine-then-localize), with similarity-weighted DBSCAN yielding candidate coordinates. Rather than greedily taking the highest confidence, candidates are ordered by H_i = lambda*p_i + (1-lambda)(1 - d_i/d_max). Across 62 MP3D/HM3D scenes and 8,195 episodes: OGN reaches 78.5% SR on HM3D (24.0 points above SOTA UniGoal), OVON zero-shot beats supervised DAgRL, IIN reaches 71.4%; SPL gains are even more consistent (IIN 57.2% vs 23.7%). On A-EQA it achieves the highest LLM-Match of 54.6, still trailing humans by 27.5. Real-world deployment on a custom platform (AgileX Ranger-mini-3.0 chassis, Franka Research 3 arm, RealSense D435i) ran 75 episodes in a ~200 m² two-floor space, with IIN reaching 100% SR on 4 of 5 targets and reliable localization to semantically plausible regions even on failure, plus long-horizon navigation-plus-manipulation demos such as "make breakfast" over three open-vocabulary objects. The paper proposes an embodied Turing test for spatial cognition probing three dimensions: real-time construction of reusable spatial representations, abstraction from sparse partial observations, and translation of high-level goals into actionable spatial plans.

A text-to-CAD tool lands in the official Codex plugin directory. The demo designs a 3D-printable dual iPhone stand from plain language in about ten minutes.

The same model behaves like two different models depending on the agent harness around it: GLM-5.2 scores 23% in one harness and 52% in another on SWE-bench Pro; in a controlled experiment swapping the harness moved scores 13 points while swapping the model moved only 2.5–5. A Hugging Face team (Adithya S Kolavi, Joel Niklaus, Lewis Tunstall, Leandro von Werra and colleagues, with Liquid AI) publishes an open fix: run agentic RL inside the real harnesses — Claude Code, Codex, OpenCode, Mini-SWE-Agent — with zero harness code changes. The stack: OpenEnv as the shared interface, a capture proxy that mints a session id per rollout and records exact token ids, per-token behavior logprobs and loss masks at the model endpoint (full-distribution sampling, top_p=1.0), Harbor for 40+ harness adapters and 26 sandbox backends, and TRL's Async GRPO consuming the masked training sequences. The numbers: LFM2.5-2.6B trained across four harnesses rises from 42.2% to 54.2% average pass@1 with gains under all four, and uses 31% fewer tool calls on tasks both it and the base model solved, while the OpenCode-only arm's gains stay mostly at home. The article also reports the SFT comparison (RL 54.6% vs SFT 47.5%), the full attribution of the earlier Qwen rise-and-decline (output-budget exhaustion, working without submitting, and a reward term that collapsed held-out accuracy from 0.740 to 0.178), and engineering boundaries like the proxy's concurrency ceiling.

SANA is NVIDIA Labs open-source codebase for high-resolution image and video generation (9k+ stars, Apache-2.0, PyTorch). Its efficiency comes from two structural changes: linear attention replaces vanilla attention in the transformer, and a deep compression autoencoder compresses images 32x instead of the usual 8x. Official figures: the 0.6B model renders a 1024 image in 0.9 s, roughly one fortieth of FLUX-dev 23 s, while FID improves from 10.15 to 5.61. The series now spans one-step generation, video, world models and streaming editing, and runs 4K in 8 GB of VRAM under 4-bit quantization. No benchmarks reproduced; recorded as unverified.
v3.23.0, 50,461 stars, CC BY-NC 4.0. Five Claude Code skills connect deep research, paper writing, simulated peer review, a 10-stage orchestrator, and systematic-review screening, using a Material Passport, mandatory checkpoints, and integrity gates to keep the scholar in control.
Foundation models: pre- and post-training, architecture and long context, inference serving and cost per token.





Agent runtimes: plugins and sandboxes, tool surfaces, skill systems, long-term memory and self-improvement.




Code generation, repair, refactoring, software engineering tasks



Editor's takeIts premise is that models lack process discipline, not coding ability. The single best idea to steal is 'Rulings, not stalls': decide by default, and stop for a human only when an action is irreversible, touches a security-sensitive surface, leaks outside the worktree, or the plan is already too broken to continue.
v6.4.1 (2026-09-18), MIT, 289k stars — the largest single skill repository of its kind. It chains brainstorming, spec writing, git-worktree isolation, planning, subagent execution, red-green TDD, two-way code review and verification-before-completion into one trunk workflow, and the README lists 16+ harnesses: Claude Code, Codex, Cursor, Gemini CLI, OpenCode, Kimi Code, Devin CLI and more.
/plugin install superpowers@claude-plugins-officialEditor's takeIt sells parts rather than a whole process: code-review, diagnosing-bugs and resolving-merge-conflicts each stand alone. grill-me, grilling and wait-what turn the agent around to interrogate your requirements, and nothing else in this batch does that. CONTEXT.md at the repo root decouples project vocabulary from generic skills — the cleanest version of that idea in this batch.
The personal working set of the author of Total TypeScript; the repo describes itself as real engineering, not vibe coding. 38 SKILL.md files, primary language Shell, grouped as engineering (18), productivity (7), misc (4) and in-progress (9). It explicitly rejects takeover-style frameworks such as GSD, BMAD and Spec-Kit in favour of small, easy to adapt, composable.
claude plugins install mattpocock-skillsEditor's takeIf you want to know how a SKILL.md is actually supposed to be written, read this instead of any second-hand tutorial. Two things to know first: docx, pdf, pptx and xlsx are source-available rather than open source (the other 15 are Apache-2.0), and the repo carries its own disclaimer — these skills are demos, so test them in your environment before trusting them with anything critical.
Anthropic's own implementation, 177k stars: 19 skills under ./skills plus ./spec (the Agent Skills specification, now at agentskills.io) and ./template as a scaffold. Its definition of a skill is folders of instructions, scripts and resources loaded dynamically, and progressive disclosure comes straight from here.
/plugin marketplace add anthropics/skillsEditor's takeIts sharpest self-description is that models can write CSS but lack a taste database. Ten rule categories are ordered by severity, with Accessibility and Touch & Interaction first as CRITICAL and Charts last — that ordering is the stance. The Design System Generator runs five parallel searches, and the repo states plainly: do not persist unverified output, so a generated design system has to be checked before it lands in your repo.
v2.13.0, MIT, 129k stars, site uupm.cc. Instead of lecturing the model it ships structured catalogues the agent queries before writing code: 79 UI styles, 192 palettes with reasoning, 74 type pairings, 119 UX rules, 105 icon suggestions, 17 GSAP presets, 25 chart types, 22 stacks and 34 landing-page patterns.
npx ui-ux-pro-max-cli init --ai claudeEditor's takeThe design that matters most: every edge is labelled EXTRACTED or INFERRED, so you can tell what the source explicitly contains from what the tool inferred — which makes an agent's conclusions auditable, a hard requirement in production. It is not a vector index: no embeddings, no vector store, a real graph you can explain, path and query. In its own benchmark, building the graph costs zero LLM spend and ingest is an order of magnitude cheaper, a structural win from local parsing.
YC S26, Apache-2.0, 120k stars, Python 3.10+. /graphify . maps code, docs, PDFs, images and videos into a graph and emits a clickable graph.html, a human-readable GRAPH_REPORT.md and a machine-readable graph.json. Code goes through tree-sitter AST parsing (about 40 languages): deterministic, no LLM calls, nothing leaves the machine.
uv tool install graphifyy && graphify installEditor's takeThere is no one-liner that installs this repo; use it as a discovery layer. A list answers what exists, not what is worth running. The part you cannot get elsewhere is connect-apps-plugin: a single MCP endpoint to 1000+ integrations with authentication, team-level ACLs and audit logs, which is exactly the governance problem enterprises hit when agents touch SaaS.
A list maintained by Composio at 75k stars with an Apache-2.0 badge in the README. About 31 first-party skill directories physically exist here while the index references 864 SKILL.md files, and most of the difference is links out to external repositories. Coverage explicitly extends beyond Claude.ai and Claude Code to Codex, Cursor, Gemini CLI and Antigravity.
claude --plugin-dir ./connect-apps-pluginSource · artificialanalysis.aiofficial Intelligence Index v4.3synced Oct 6, 2026
Text-to-image, image editing, consistency and controllability



Video and speech/audio: physical and temporal coherence for text/image-to-video, TTS and voice cloning, music generation and ASR.






3D asset generation, scene reconstruction, long-horizon world prediction



Visual design, UI components, web and slide layout
Slides, Word, Excel, PDF generation and editing

A 10-step roadmap to a local LLM stack you use daily: memory sizing, model picking, Ollama install, quantization, hardware buying, serving as an API, small-model routing, and six ready-to-run workloads.

The same model behaves like two different models depending on the agent harness around it: GLM-5.2 scores 23% in one harness and 52% in another on SWE-bench Pro; in a controlled experiment swapping the harness moved scores 13 points while swapping the model moved only 2.5–5. A Hugging Face team (Adithya S Kolavi, Joel Niklaus, Lewis Tunstall, Leandro von Werra and colleagues, with Liquid AI) publishes an open fix: run agentic RL inside the real harnesses — Claude Code, Codex, OpenCode, Mini-SWE-Agent — with zero harness code changes. The stack: OpenEnv as the shared interface, a capture proxy that mints a session id per rollout and records exact token ids, per-token behavior logprobs and loss masks at the model endpoint (full-distribution sampling, top_p=1.0), Harbor for 40+ harness adapters and 26 sandbox backends, and TRL's Async GRPO consuming the masked training sequences. The numbers: LFM2.5-2.6B trained across four harnesses rises from 42.2% to 54.2% average pass@1 with gains under all four, and uses 31% fewer tool calls on tasks both it and the base model solved, while the OpenCode-only arm's gains stay mostly at home. The article also reports the SFT comparison (RL 54.6% vs SFT 47.5%), the full attribution of the earlier Qwen rise-and-decline (output-budget exhaustion, working without submitting, and a reward term that collapsed held-out accuracy from 0.740 to 0.178), and engineering boundaries like the proxy's concurrency ceiling.

Odyssey (with UCL AI Centre and the University of Basel) released PROWL-2 on Oct 1: the first framework coupling an agent curriculum and a world-model repair curriculum inside one continual training loop. Core insight: in imagination training, a high-learning-signal trajectory is ambiguous — a real policy weakness or just a wrong world-model prediction; prioritizing it indiscriminately reinforces the model's own hallucinations. A fidelity gate separates 'useful' from 'trustworthy': reliable high-potential imagined experience feeds the policy curriculum, unreliable rollouts — with their already-stored real continuations — go to a repair pool, fixed by a randomly-initialized, KL-anchored developer policy exploring the real environment, re-audited and readmitted once repaired. First place on all nine SMACv2 scenarios (+4-18% at 5v5, +20-91% at 10v10/10v11 over the backbone); big margins on hard MQE tasks — Gate-3 70.4% vs 29.8%, Shepherd-Hard 28.2% vs 7.3%, all other baselines at 0. Ablations show the curricula are not additive: repair alone is marginal, the ungated curriculum falls below the backbone in all nine scenarios, and the gate turns the same curriculum into the largest single gain. Architecture- and algorithm-agnostic; next steps point at humanoid coordination and long-horizon multi-agent games.
Under constructionThe games system is still under construction — everything listed starts in-page, and the library keeps growing.

A first-person match where every bot across from you is driven live by a policy this repo trained, the whole round stepping at a fixed 60 Hz: builtin, Rust/wasm or Rapier behind one physics interface, nothing scripted.
WASD move · mouse look · left-click fire · shift sprint · R reload · esc release

Unreal 5.5's first-person test map, running same-origin in the browser: the level, the materials, the lights and four weapon data assets are read out of the project's own .umap / .uasset bytes, six soldiers are driven by a four-state brain on one fixed 60 Hz step, and the sidebar lists every substitution the page had to make.
WASD move · mouse look · left-click fire · shift sprint · space jump · R reload · F pick up · esc release

Fifty-two professional tools, layers and masks, PSD files that open and save back - all in the browser, nothing to install.
retouchpi.com · 52 pro tools · PSD read and write · AI generate & edit

A 100 km 3D city that opens in seconds — crawling, blogs, papers, industry and investment analysis, all watchable in one digital twin.
100km city · opens in seconds · Live task stream · Agent-maintained

An SO-101 6-DoF arm running MuJoCo physics in your browser: joint teleoperation, IK end-effector dragging and gripper pick-and-place into a basket. No install.
MuJoCo WASM physics · SO-101 · 6 DoF · Contacts / telemetry HUD

Take a humanoid joint module apart layer by layer: brushless motor, magnetic encoder, planetary / harmonic / cycloidal drive and output flange. Switch architectures on one page — exploded view plus analytic kinematics.
Three gearbox types · Planetary / harmonic / cycloidal · Analytic kinematics · explode

Matcha-TTS mixed zh/en speech synthesis: server-side synthesis with sentence-streamed playback and full-article read-aloud for papers and blogs. Sign-in required.
Matcha-TTS · server-side · Mixed zh / en · Full-article read-aloud

A caring coding sprite on your desktop and a cockpit that never stops: the work resumes itself after a crash or a reboot. Every coding terminal, ssh session and local shell in one tree, every start, finish and permission request reported in time. Open source, one-line install.
Your caring coding desktop sprite · Resumes after crash & reboot · First-rate terminal / ssh / shell