1:04One job: collect what has actually happened on the road to AGI, with every entry traceable to a paper, a repo, a benchmark or a video. No headline chasing and no full tool directory — what shows here is only what we read ourselves.
This screen is ordered by our in-house recommender: the pool is every paper, deep read, demo, repo and skill that cleared our quality gate, ranked on real engagement signal, freshness and cold-start exploration — so it shifts a little each day.

ModularRSI is a benchmark-disjoint, contrastive, and modular framework for generalizable agent-harness evolution. It diagnoses recurring failures from paired successful and failed trajectories, evolves five functional modules independently, and integrates validated changes into a unified harness. On Terminal-Bench 2.0 and SWE-Bench Verified, the evolved harness improves unseen in-domain and cross-domain tasks and transfers across foundation models.

A deep read of TypeSafe's first System One Model: three question primitives, the economics of parallel calls, confidence-gated routing, eval caveats, eight jagged edges, and the OpenJev local repro.

Tsinghua University and Beihang University present BSC-Nav (Brain-inspired Spatial Cognition for Navigation). The paper argues that existing embodied agents — whether end-to-end RL or "MLLM plus modular pipeline" — are fundamentally reactive and stateless: they process an observation and discard it, lacking any durable internal model of space, which yields fragmented knowledge, short-sighted planning and poor generalization. The authors borrow the answer from neuroscience, where spatial knowledge consolidates into three interconnected forms: landmarks, route knowledge, and survey knowledge. BSC-Nav instantiates these computationally in three modules. Landmark memory stores 4-tuples (world coordinates, open-vocabulary category, detection confidence, GPT-4o contextual description) with a spatial-overlap set plus confidence-weighted fusion for deduplication. The cognitive map extracts DINO-v2 patch features, projects them through inverse perspective projection and cascaded coordinate transforms into a voxel grid, and adopts a free-energy-principle-inspired surprise-driven update: a new feature is written when its mean distance to features in the n-hop neighborhood exceeds a threshold, replacing the lowest-surprise entry when the buffer is full, preserving cross-viewpoint diversity while bounding storage. Working memory retrieves hierarchically by instruction complexity — simple targets use text-only GPT-4 reasoning over landmark memory (even inferring unrecorded targets from co-located landmarks), while complex targets first have descriptions refined by GPT-4o, then "imagine" the appearance via Stable Diffusion 3.5, encoded by DINO-v2 and center-distance weighted pooled to query the cognitive map (imagine-then-localize), with similarity-weighted DBSCAN yielding candidate coordinates. Rather than greedily taking the highest confidence, candidates are ordered by H_i = lambda*p_i + (1-lambda)(1 - d_i/d_max). Across 62 MP3D/HM3D scenes and 8,195 episodes: OGN reaches 78.5% SR on HM3D (24.0 points above SOTA UniGoal), OVON zero-shot beats supervised DAgRL, IIN reaches 71.4%; SPL gains are even more consistent (IIN 57.2% vs 23.7%). On A-EQA it achieves the highest LLM-Match of 54.6, still trailing humans by 27.5. Real-world deployment on a custom platform (AgileX Ranger-mini-3.0 chassis, Franka Research 3 arm, RealSense D435i) ran 75 episodes in a ~200 m² two-floor space, with IIN reaching 100% SR on 4 of 5 targets and reliable localization to semantically plausible regions even on failure, plus long-horizon navigation-plus-manipulation demos such as "make breakfast" over three open-vocabulary objects. The paper proposes an embodied Turing test for spatial cognition probing three dimensions: real-time construction of reusable spatial representations, abstraction from sparse partial observations, and translation of high-level goals into actionable spatial plans.

Xiaomi open-sources MiMo-V2.6 Pro/Flash: 30 live RL steps, ~750k trajectories at $850k/$2.62M; AA Index 46 tops open models, DeepSWE v1.1 gains +17/+14 out of sample; Vibe World, CUA, science and content demos plus 7k+ RL environments released.

First look at the LiberAI preview model Liber-0: fully autonomous dexterous manipulation of everyday objects, including recovering and retrying when execution goes wrong, as human experience scales.

The run book behind PaperRoute: mechanics before art, engine and look as separate threads, Blender driven by headless Python, Meshy for faces, review renders driving iteration, 39 tracked hours.

OmniChar is a GPL-3.0 open-source, fully local character-consistency tool that has been drawing attention overseas. It packages face and wardrobe into a tiny portable .char file: build a character once and it stays itself — same face, same body, same outfit — across every model you use, solving the difficulties LoRA-based approaches face with off-angle depictions and feature drift. Technically, a .char stores an SFace face signature and a DINOv2 subject signature extracted from your photos; FLUX.2 has a native multi-reference channel and reads those references directly with nothing to train, while Krea 2 has no reference channel and instead takes a LoRA trained from the same photos on the built-in Trainer canvas. One .char file carries both forms, so switching models does not mean rebuilding the character, and every file records how it was made and who it belongs to. It runs on Inline Core, a from-scratch diffusion engine that serves the web UI, runs models, keeps the project database and drives the timeline as a single Python process, supporting Z-Image Turbo, FLUX.2 and Krea 2 for images plus MiniMax H3 and LTX-2.5 for video, on NVIDIA, AMD or CPU with Python 3.11+. It is also available as ComfyUI nodes and can reach hosted Seedance and Kling through API nodes. Community reception has been largely positive, many welcoming it as a way to streamline character management, while others suggest using it alongside existing LoRA and reference-image workflows depending on the scenario.

A Japanese post spotlighting the free Codex skill unreal-home-wizard: give Codex a room photo and it checks the Unreal Engine 5.8 setup, drafts the floorplan for confirmation, furnishes and lights the scene while fixing clipping and gaps against the source photo, and finally launches a walkable 3D space (WASD). Install by asking Codex to add the skill; users still need UE installed and an Epic Games login. Note: this is a promotional repost — the original demo video is from @AmirMushich.

Delta Intelligence has released Delta-0, a humanoid foundation model for whole-body loco-manipulation built on a co-designed brain and controller. The brain is a latent world-action model: a mixture-of-transformers with three branches (vision-language reading multi-view observations and instructions, a vision branch predicting future frames in DINO semantic feature space, and an action branch modelling whole-body motion from proprioception), trained jointly in four modes - forward dynamics, inverse dynamics, visual planning and policy-only. The controller converts motion commands into joint targets across 69 degrees of freedom through a delta-action interface shared by policy output, teleoperation and human corrections; in zero-shot evaluation on roughly 3 h 31 min of everyday-task motion it beats HEFT, MimicLite v1.1, SONIC v1.1 and ScaleBFM XL on both global root error and MPJPE. Pretraining uses more than 10,000 hours of paired egocentric video and whole-body human motion, plus a 154-dimensional shared action representation padded to 180 so single-arm, dual-arm, ego, UMI and humanoid data all live in one action space; high-precision tracking coverage rises from 29.2% to 72.6% as motion data goes from 10% to 100%. Evaluation runs a real-to-sim-to-real loop, rebuilding deployment scenes into simulation assets with 3D reconstruction and LLM agents, running repeatable autonomous rollouts, then returning selected checkpoints to hardware. Real-world RL post-training uses a value model that estimates cost-to-go, giving positive or negative conditioning signals to autonomous action chunks and to human-in-the-loop corrections - on the dishwasher task in a mirrored kitchen layout the success rate climbs from 20% (4/20) after SFT to 65% (13/20). The demo video covers six long-horizon household tasks: loading a record on a turntable, making a bed, picking a toy off the floor, opening a dishwasher, pressing a step trash can and sitting down on a sofa.

Kuaishou Kling AI announces Kling 4.0 for October, with Flash now live for Ultra Yearly subscribers. Four upgrade pillars - Audio & Visuals: stable dynamic motion, high-quality stereo audio, more accurate lip sync, up to 4K resolution and 10-bit HDR output; Omni Reference: up to 15 multimodal references with more consistent results and enhanced video editing; Storytelling: multi-keyframe control up to 10 keyframes and native 30-second generation; plus video extension and support for multiple languages, accents and dialects. Video generation keeps closing in on production-grade material, relevant to both embodied world models and content generation.
An MIT-licensed, dependency-free plain C/C++ inference implementation built on ggml. Seventeen backends span CUDA, Metal, HIP, Vulkan, SYCL, WebGPU, CANN, Hexagon, MUSA, zDNN and ZenDNN; integer quantisation runs 1.5 to 8 bits; CPU+GPU hybrid inference makes models larger than VRAM runnable at all; and GGUF, its output format, is a de facto standard that even vLLM consumes. Two lines - llama cli or llama serve - start an OpenAI-compatible server.
A high-performance serving framework hosted by LMSYS: RadixAttention prefix caching, a zero-overhead CPU scheduler, prefill-decode disaggregation, DFlash and Spec V2 speculative decoding, compressed finite state machines for structured output, and large-scale expert parallelism (96 H100 GPUs; 3.8x prefill and 4.8x decode on GB200 NVL72 part II). Its day-0 ledger covers Kimi K3, DeepSeek-V4, GLM5.2 NVFP4, Nemotron 3 and MiniMax M2, and AReaL, Miles, slime, Tunix and verl all use it as an RL rollout backend.
The foundation-model layer: the weights themselves plus what is tightly coupled to them - training recipes (pretraining, post-training, RL/RLHF), architecture and long context, and the inference serving stack with its cost per token
2026-08-10Stealing Reasoning Traces from Proprietary LLM APIsLeading large language model providers now conceal their models' step-by-step reasoning, or chain-of-thought, to protect intellectual property and limit information leakage. Rather than storing these traces server-side, providers return them to the client as blocks of encrypted text, which the client passes back with each subsequent request. Building on prior research, we identify an architectural vulnerability: these encrypted blocks are fully compatible and interchangeable across different sessions, users, and models within a provider's ecosystem. We exploit this compatibility to develop a scalable decryption jailbreak. By injecting an encrypted reasoning trace from a given model into a weaker, and less safeguarded model from the same provider, we force it to decode and output the trace verbatim in plaintext, without ever jailbreaking the more capable model directly. This vulnerability enables four distinct attack vectors. First, it circumvents anti-distillation mechanisms, allowing adversaries to extract a proprietary model's reasoning, as we demonstrate across Anthropic, OpenAI, and Google. Second, it allows for large-scale private data extraction. Developers frequently share session logs publicly, unaware of contents of the encrypted blocks. By decoding 315,320 reasoning blocks scraped from public repositories, we recovered 367 Personally Identifiable Information (PII) artifacts and 182 credentials. Third, it inadvertently reveals hazardous information hidden within the reasoning process, even in cases where the model's final, visible output safely rejects a malicious request. Fourth, attackers can leverage this flaw to execute invisible prompt injections, embedding malicious payloads entirely within encrypted blocks to poison public agentic rollouts. Following responsible disclosure, we propose concrete cryptographic and system-level mitigations to secure client-side reasoning.
2026-07-22LENS: LLM-guided Environment Simplification for Planning and Control in ClutterDespite recent advances in general-purpose robotic manipulation, real-world multi-object clutter remains challenging to handle for today's prevalent approaches. The problem scales in complexity due to more objects and collisions, more unpredictable contact physics, distractors, and task ambiguity. Bridging this gap to real-world deployment requires effective scene abstractions; yet today, producing such abstractions requires extensive task-specific manual engineering, which does not scale. These abstractions are costly to generate and difficult to adjust or fine-tune. We instead propose a plug-and-play fix to automatically generate scene-specific, task-specific, adaptively updating abstractions on top of existing planning and control stacks. LLM-guided Environment Simplification (LENS) produces a de-cluttered abstracted scene representation by merging (e.g., stacked objects) or pruning (e.g., distant objects) scene entities in a closed loop in response to task progress. These dynamic, task-relevant abstractions are versatile and easy to use. In our experiments, we show that LENS improves classical planning, model-based control, and a vision-language-action model, across a diverse set of highly cluttered manipulation scenes. Project website: https://lens-2026.github.io/.Agent runtimes and scaffolds: plugin architecture, sandboxing & tool surface, skill systems, long-term memory, self-improvement loops
2026-09-17SoL-Pi: Recursively Scaling Auto-Research Loops for Efficient Agent HarnessAs coding agents move from supervised code completion to unattended, around-the-clock exploration, their work expands into long trajectories of reasoning, tool use, and feedback. SoL-Pi scales recursive auto-research across executable environments to discover reusable harness mechanisms. Four retained mechanisms improve action execution, context compaction, observation handling, and delegated reading, reducing recorded token traffic by 44.7-49.0% and API cost by about one third at comparable EdgeBench performance.Code generation, repair, refactoring, software engineering tasks
2026-09-28mobile-mcp: the mobile automation MCP server driven by the native accessibility treemobile-next/mobile-mcp (8.2k stars, Apache-2.0) connects iOS simulators and real devices plus Android emulators and real devices to any MCP client, including Claude Code, Codex, Gemini and Copilot. Its first principle is accessibility-first: read the native accessibility tree instead of screenshots, no vision model and no image tokens, falling back to screenshots with coordinates only when necessary. Roughly 30 tools in seven groups cover device management, remote devices through Mobile Next Cloud, app management, screen interaction with recording, input and navigation, logs and crashes, and batch commands that compress N rounds into one. The README examples are all cross-app long-horizon tasks, the hardest class for mobile agents.2026-09-28unsloth: the 2x-faster, 70%-less-VRAM fine-tuning library, now a desktop runtime that wires local weights into Claude Code and CodexSelf-described as the first desktop app to run and train models (Windows/macOS/Linux; multi-GPU across NVIDIA, AMD, Intel, CPU and Vulkan). Fine-tuning is claimed at 2x faster with 70% less VRAM, the post-training menu covers LoRA, QLoRA, full fine-tuning, pretraining, GRPO, DPO and FP8, and export reaches GGUF, NVFP4 and FP8. Unsloth Start connects Claude Code or Codex to local weights in one command; the model list includes Qwen3.8, GLM-5.3-Flash, Kimi K3, DeepSeek-V4, MiniMax-H3 and Gemma 4, with private web search, deep research, auto-compaction and RAG built in.2026-09-27GenOffice: Open-Source AI Office SuiteGenOffice is the first full-featured open-source AI office suite (Apache-2.0): real .docx/.xlsx/.pptx/PDF/Markdown editing with byte-preserving Word compatibility, block-level AI editing with document-aware agents, BYOK model support, fully local.Our takeIts premise is that models lack process discipline, not coding ability. The single best idea to steal is 'Rulings, not stalls': decide by default, and stop for a human only when an action is irreversible, touches a security-sensitive surface, leaks outside the worktree, or the plan is already too broken to continue. We adopted those four gates in our own pipelines.
v6.4.1 (2026-09-18), MIT, 289k stars — the largest single skill repository we have found. It chains brainstorming, spec writing, git-worktree isolation, planning, subagent execution, red-green TDD, two-way code review and verification-before-completion into one trunk workflow, and the README lists 16+ harnesses: Claude Code, Codex, Cursor, Gemini CLI, OpenCode, Kimi Code, Devin CLI and more.
/plugin install superpowers@claude-plugins-officialOur takeIt sells parts rather than a whole process: code-review, diagnosing-bugs and resolving-merge-conflicts each stand alone. grill-me, grilling and wait-what turn the agent around to interrogate your requirements, and nothing else in this batch does that. CONTEXT.md at the repo root decouples project vocabulary from generic skills — the cleanest version of that idea we have seen.
The personal working set of the author of Total TypeScript; the repo describes itself as real engineering, not vibe coding. 38 SKILL.md files, primary language Shell, grouped as engineering (18), productivity (7), misc (4) and in-progress (9). It explicitly rejects takeover-style frameworks such as GSD, BMAD and Spec-Kit in favour of small, easy to adapt, composable.
claude plugins install mattpocock-skillsOur takeIf you want to know how a SKILL.md is actually supposed to be written, read this instead of any second-hand tutorial. Two things to know first: docx, pdf, pptx and xlsx are source-available rather than open source (the other 15 are Apache-2.0), and the repo carries its own disclaimer — these skills are demos, so test them in your environment before trusting them with anything critical.
Anthropic's own implementation, 177k stars: 19 skills under ./skills plus ./spec (the Agent Skills specification, now at agentskills.io) and ./template as a scaffold. Its definition of a skill is folders of instructions, scripts and resources loaded dynamically, and progressive disclosure comes straight from here.
/plugin marketplace add anthropics/skillsOur takeIts sharpest self-description is that models can write CSS but lack a taste database. Ten rule categories are ordered by severity, with Accessibility and Touch & Interaction first as CRITICAL and Charts last — that ordering is the stance. The Design System Generator runs five parallel searches, and the repo states plainly: do not persist unverified output, so a generated design system has to be checked before it lands in your repo.
v2.13.0, MIT, 129k stars, site uupm.cc. Instead of lecturing the model it ships structured catalogues the agent queries before writing code: 79 UI styles, 192 palettes with reasoning, 74 type pairings, 119 UX rules, 105 icon suggestions, 17 GSAP presets, 25 chart types, 22 stacks and 34 landing-page patterns.
npx ui-ux-pro-max-cli init --ai claudeOur takeThe design that matters most: every edge is labelled EXTRACTED or INFERRED, so you can tell what the source explicitly contains from what the tool inferred — which makes an agent's conclusions auditable, a hard requirement in production. It is not a vector index: no embeddings, no vector store, a real graph you can explain, path and query. In its own benchmark, building the graph costs zero LLM spend and ingest is an order of magnitude cheaper, a structural win from local parsing.
YC S26, Apache-2.0, 120k stars, Python 3.10+. /graphify . maps code, docs, PDFs, images and videos into a graph and emits a clickable graph.html, a human-readable GRAPH_REPORT.md and a machine-readable graph.json. Code goes through tree-sitter AST parsing (about 40 languages): deterministic, no LLM calls, nothing leaves the machine.
uv tool install graphifyy && graphify installOur takeThere is no one-liner that installs this repo; use it as a discovery layer. A list answers what exists, not what is worth running. The part you cannot get elsewhere is connect-apps-plugin: a single MCP endpoint to 1000+ integrations with authentication, team-level ACLs and audit logs, which is exactly the governance problem enterprises hit when agents touch SaaS.
A list maintained by Composio at 75k stars with an Apache-2.0 badge in the README. About 31 first-party skill directories physically exist here while the index references 864 SKILL.md files, and most of the difference is links out to external repositories. Coverage explicitly extends beyond Claude.ai and Claude Code to Codex, Cursor, Gemini CLI and Antigravity.
claude --plugin-dir ./connect-apps-pluginRanks come verbatim from each evaluator's own board, synced into our database on a schedule. We never recompute, reweight or merge across sources: an intelligence index, an Elo and a success rate are different rulers, and any single overall number would be incomparable figures stitched into one that looks comparable.
Source · artificialanalysis.aiofficial Intelligence Index v4.3synced Sep 22, 2026
Text-to-image, image editing, consistency and controllability
Video generation and speech/audio in one channel: physical and temporal coherence for text/image-to-video, TTS naturalness and voice cloning, music generation, and ASR
2026-09-24Rolling-WAM: World Action Models with Rolling ImaginationWorld Action Models (WAMs) couple action generation with future visual prediction for robotic manipulation. However, completing the joint video-action denoising process at each replanning cycle incurs substantial latency, delaying action updates and limiting closed-loop responsiveness. We present Rolling-WAM, a formulation that distributes joint denoising across successive replanning cycles. Our method maintains a sliding window of video-action chunks at staggered noise levels. At each step, a rolling noise schedule fully denoises the imminent action chunk for execution, while partially refining farther-future chunks. As the window advances with new camera observations, the retained future chunks continue their denoising process. This distributes the computational cost over time while carrying an evolving visual-action context across chunk boundaries. Evaluations on LIBERO, RoboTwin, and a real-world Unitree G1 humanoid show that Rolling-WAM achieves competitive manipulation performance. By removing the need to denoise the entire prediction horizon from scratch, it delivers a 4.5x steady-state replanning speedup over standard joint WAMs.
2026-09-24BeyondRetarget: Learning Executable Humanoid Motions Directly from Monocular VideoBeyondRetarget removes the SMPL intermediate representation and maps monocular RGB video end-to-end to executable humanoid motions: 18 semantic keypoints with per-link non-uniform scaling form unified supervision, a shared representation decodes onto 8 humanoids, and contact-aware refinement yields 26.4mm error, zero collapse, and 193ms latency enabling real-time visual teleoperation.
2026-09-22MachEmbodied-U0: Unified Understanding and Generation Model for Embodied IntelligenceGeneral-purpose robot control requires models to understand task intent, identify where to interact, capture how the scene evolves, and generate precise actions. Vision-language-action models provide strong semantic priors but typically do not explicitly model scene dynamics, while world-action models couple visual prediction with control without necessarily exposing the task-relevant semantic and spatial structure needed for fine-grained manipulation. We present MachEmbodied-U0 (ME-U0), a unified embodied foundation model connecting understanding and generation experts through a Mixture-of-Transformers architecture. Subtask prediction and affordance grounding guide joint visual-dynamics and action generation via flow matching. Visual dynamics encompass future RGB, depth, surface normals, and optical flow, providing complementary supervision for appearance, geometry, and motion. Multi-rate Rotary Position Encoding (MRPE) aligns visual dynamics with fine-grained control. We pretrain ME-U0 on approximately 4,200 hours of curated demonstrations from robotic datasets and egocentric datasets. Using only the supervision natively available in each downstream benchmark, ME-U0 achieves an average score of 17.66 on the RoboDojo simulation benchmark and average success rates of 99.0\% and 82.5\% on LIBERO and LIBERO-Plus, respectively. We additionally validate ME-U0 on real-world robotic manipulation tasks, demonstrating its effectiveness beyond simulation. Without corresponding downstream supervision, ME-U0 further demonstrates zero-shot subtask prediction, affordance grounding, and visual dynamics on simulated and real-world observations. Overall, ME-U0 combines competitive downstream control performance with transferable task-grounding and visual-dynamics capabilities across simulation and the real world.
2026-09-04GE-Act 2.0: Pretraining and Scaling a World-Action Model for Robotic ManipulationGE-Act 2.0 combines a control-oriented autoencoder, a single-step visual planner, an inverse dynamics model, and knowledge-aligned selective optimization to pretrain a world-action policy from scratch on manipulation data. Scaling co-training data from 300 to 30,000 hours raises zero-shot OOD success to 44.1% on G1-OP and 31.1% on G2-90D without task-specific fine-tuning.3D asset generation, scene reconstruction, long-horizon world prediction
2026-09-21DexTacWAM: A Visuo-Tactile World-Action Model for Dexterous ManipulationDexterous manipulation depends on contact dynamics that are often only partially observable from vision. Recent World-Action Models (WAMs) couple predictive video world modeling with action generation, but remain largely vision-centric and therefore cannot directly model these contact dynamics. We present DexTacWAM, a visuo-tactile WAM that encodes each fingertip independently, aggregates the resulting features through a finger- and pose-aware tactile compressor, and injects the tactile latent into a video diffusion world model for joint visuo-tactile world modeling. Across six contact-rich dexterous manipulation tasks on a 22-DoF bimanual platform, DexTacWAM achieves the highest score on every task, averaging 70.6 versus 38.0 for the strongest baseline. Ablations attribute the gain to modeling contact evolution as part of the predicted world state rather than tactile conditioning alone: removing tactile world modeling reduces the four-task mean from 74.7 to 26.6 while keeping the same tactile features and action expert. After four hours of tactile-encoder adaptation with a frozen pretrained vision VAE, our continual vision-to-touch learning extends the pretrained video model to touch using roughly 100 demonstrations per task without tactile midtraining, while retaining visual prediction quality within 0.5 dB of vision-only counterparts. The compressor retains 89.4% of pre-fusion contact recall while enabling 2.26x faster training and 1.29x faster inference. Together, these results show that pretrained video priors can be extended to distributed multi-finger contact dynamics in a data- and compute-efficient manner.Visual design, UI components, web and slide layout
2026-07-25Fashion-3DLR: A Controllable 3D Garment Generation Using Pairwise Fashion Elements for Intelligent DesignAI-generated content (AIGC) has made significant progress, with 2D generative models becoming ready-to-use tools for the digital fashion industry. However, 3D garment generation remains in its nascent stage, where in the realm of fashion, the semantic information of diverse design elements exhibits intricate coupling relationships in 3D representations, posing substantial challenges for generating diverse 3D garments. In this work, to handle the above problem, We introduce Fashion-3DLR, a novel 3D garment generation framework that utilizes diverse design elements to create high-quality, versatile 3D garment assets. Specifically, to bridge the semantic gaps between different fashion elements, we propose a Garment Feature Fusion Diffusion Transformer (GFF-DiT) module to integrate 2D fashion design elements, e.g., sketch and texture, into latent space. Within the latent space, we then employ a rectified flow transformer to generate geometry latents, which can be decoded into various 3D garment representations, including 3D Gaussians and meshes. Furthermore, we integrate Fashion-3DLR into downstream tasks, achieving the 3D Gaussian Splatting (3DGS)-driven cloth physical simulation and mesh-based virtual try-on. Experimental results indicate that Fashion-3DLR surpass the previous state-of-the-art methods, which verify that the proposed work can generate well-structured, non-watertight garments capable of physical simulation and virtual try-on, underscoring its potential as a versatile 3D garment design tool.Slides, Word, Excel, PDF generation and editing

Physical Intelligence (Sergey Levine's team, Mar 19, 2026) present RL Tokens (RLT): freeze a pretrained VLA, attach an encoder-decoder that compresses its internal embeddings into a bottleneck RL token, then run online RL with a ~1M-parameter actor-critic on the real robot, refining only the critical phase (editing VLA action chunks rather than generating from scratch, anchored by a BC regularizer plus reference-action dropout). Across four sub-millimeter tasks the critical phase speeds up by up to 3x, screw success goes 20% to 65%, and half of Ethernet insertion episodes beat every human teleoperation demo - with just 15 minutes of real robot data.

Avid's 3,500-word build log: Jev as a decision layer inside keel, a local-first Rust coding workspace. One principle carries it: a model can suggest the next move, the host still owns the move — the host prepares the menu, Jev picks from it, the host checks again; pinned routes and live sessions bypass automatic selection; selection grants no permission. Plus how decision receipts become replayable evaluation scenarios.

Gergely Orosz visits OpenAI: Codex and ChatGPT Work have taken over everything; a nine-step agentic software factory with Perf Factory and Sevbot, and the rethink of IDEs and pull requests.
Under constructionThe games system is still under construction — everything listed starts in-page, and the library keeps growing.

A first-person match where every bot across from you is driven live by a policy this repo trained, the whole round stepping at a fixed 60 Hz: builtin, Rust/wasm or Rapier behind one physics interface, nothing scripted.
WASD move · mouse look · left-click fire · shift sprint · R reload · esc release

Unreal 5.5's first-person test map, running same-origin on this site: the level, the materials, the lights and four weapon data assets are read out of the project's own .umap / .uasset bytes, six soldiers are driven by a four-state brain on one fixed 60 Hz step, and the sidebar lists every substitution the page had to make.
WASD move · mouse look · left-click fire · shift sprint · space jump · R reload · F pick up · esc release

A 100 km 3D city that opens in seconds — the digital twin of the AI agents that build this site: crawling, blogs, papers, industry and investment analysis.
100km city · opens in seconds · Live task stream · Agent-maintained

An SO-101 6-DoF arm running MuJoCo physics in your browser: joint teleoperation, IK end-effector dragging and gripper pick-and-place into a basket. No install.
MuJoCo WASM physics · SO-101 · 6 DoF · Contacts / telemetry HUD

Take a humanoid joint module apart layer by layer: brushless motor, magnetic encoder, planetary / harmonic / cycloidal drive and output flange. Switch architectures on one page — exploded view plus analytic kinematics.
Three gearbox types · Planetary / harmonic / cycloidal · Analytic kinematics · explode

Matcha-TTS mixed zh/en speech synthesis: server-side synthesis with sentence-streamed playback and full-article read-aloud for papers and blogs. Sign-in required.
Matcha-TTS · server-side · Mixed zh / en · Full-article read-aloud

A caring coding sprite on your desktop and a cockpit that never stops: the work resumes itself after a crash or a reboot. Every coding terminal, ssh session and local shell in one tree, every start, finish and permission request reported in time. Open source, one-line install.
Your caring coding desktop sprite · Resumes after crash & reboot · First-rate terminal / ssh / shell

Built for games and for physical AI: an ultra-realistic, ultra-high-performance simulation environment with ultra-low-cost procedural data; Unreal-grade visuals open and play in the browser, bridging the physical world and intelligent agents.
Games meet physical AI · Unreal-grade on the web · Open and play

The first online photo editor built to be driven by AI agents: fifty-two professional tools, a layer tree with masks, PSD files that open and save back, nothing to install, and your images never leave your computer. The agent layer is on its way.
retouchpi.com · 52 pro tools · PSD read and write · your files never leave