1:04One job: collect what has actually happened on the road to AGI, with every entry traceable to a paper, a repo, a benchmark or a video. No headline chasing and no full tool directory — what shows here is only what we read ourselves.
This screen is ordered by our in-house recommender: the pool is every paper, deep read, demo, repo and skill that cleared our quality gate, ranked on real engagement signal, freshness and cold-start exploration — so it shifts a little each day.

Avid's 3,500-word build log: Jev as a decision layer inside keel, a local-first Rust coding workspace. One principle carries it: a model can suggest the next move, the host still owns the move — the host prepares the menu, Jev picks from it, the host checks again; pinned routes and live sessions bypass automatic selection; selection grants no permission. Plus how decision receipts become replayable evaluation scenarios.

A deep read of TypeSafe's first System One Model: three question primitives, the economics of parallel calls, confidence-gated routing, eval caveats, eight jagged edges, and the OpenJev local repro.

Jev-style fixed-answer scoring reproduced locally with SGLang /v1/score: single-token labels, restricted softmax, confidence-gated routing, and a 6 ms vs 1088 ms benchmark against generation.

ModularRSI is a benchmark-disjoint, contrastive, and modular framework for generalizable agent-harness evolution. It diagnoses recurring failures from paired successful and failed trajectories, evolves five functional modules independently, and integrates validated changes into a unified harness. On Terminal-Bench 2.0 and SWE-Bench Verified, the evolved harness improves unseen in-domain and cross-domain tasks and transfers across foundation models.

RLCDAlignBench spans 44 benchmarks and 7,193 instances. One generic question reaches median AUROC 0.886; soft probabilities beat argmax.

A Japanese post spotlighting the free Codex skill unreal-home-wizard: give Codex a room photo and it checks the Unreal Engine 5.8 setup, drafts the floorplan for confirmation, furnishes and lights the scene while fixing clipping and gaps against the source photo, and finally launches a walkable 3D space (WASD). Install by asking Codex to add the skill; users still need UE installed and an Epic Games login. Note: this is a promotional repost — the original demo video is from @AmirMushich.

ScientistTwo from Google Cloud AI Research is a fully autonomous multi-agent research framework: given only a scientific problem, it identifies human-SOTA limitations, formulates hypotheses, implements code, runs subset-to-full experiments, performs ablations, and simulates peer review with rebuttal experiments — delivering publishable papers and reproducible codebases. Across 107 ICLR/ICML/NeurIPS-accepted papers it improves 86 (80.4%) with a 25.2% average gain over human SOTA; papers score 7.5 with 91.9% acceptance under ScholarPeer and 72.1% clear the held-out Stanford Agentic Reviewer (0% for all prior agents).

The Helix 2.5 foundation model from Figure, pretrained on Index human data and fine-tuned for tidying, towel folding and bed making, hits 56% zero-shot success across 30 unseen Bay Area homes versus 9% trained from scratch. The 3h57m narration-free recording is recut here as a 1m35s bilingual recap.

Delta Intelligence has released Delta-0, a humanoid foundation model for whole-body loco-manipulation built on a co-designed brain and controller. The brain is a latent world-action model: a mixture-of-transformers with three branches (vision-language reading multi-view observations and instructions, a vision branch predicting future frames in DINO semantic feature space, and an action branch modelling whole-body motion from proprioception), trained jointly in four modes - forward dynamics, inverse dynamics, visual planning and policy-only. The controller converts motion commands into joint targets across 69 degrees of freedom through a delta-action interface shared by policy output, teleoperation and human corrections; in zero-shot evaluation on roughly 3 h 31 min of everyday-task motion it beats HEFT, MimicLite v1.1, SONIC v1.1 and ScaleBFM XL on both global root error and MPJPE. Pretraining uses more than 10,000 hours of paired egocentric video and whole-body human motion, plus a 154-dimensional shared action representation padded to 180 so single-arm, dual-arm, ego, UMI and humanoid data all live in one action space; high-precision tracking coverage rises from 29.2% to 72.6% as motion data goes from 10% to 100%. Evaluation runs a real-to-sim-to-real loop, rebuilding deployment scenes into simulation assets with 3D reconstruction and LLM agents, running repeatable autonomous rollouts, then returning selected checkpoints to hardware. Real-world RL post-training uses a value model that estimates cost-to-go, giving positive or negative conditioning signals to autonomous action chunks and to human-in-the-loop corrections - on the dishwasher task in a mirrored kitchen layout the success rate climbs from 20% (4/20) after SFT to 65% (13/20). The demo video covers six long-horizon household tasks: loading a record on a turntable, making a bed, picking a toy off the floor, opening a dishwasher, pressing a step trash can and sitting down on a sofa.

Davide Cuttini, CEO of Over The Reality, shows 100 Unitree G1 humanoids in a single MuJoCo world tracking kung-fu motions with a pretrained whole-body policy, on a plaza reconstructed from a real 360-degree scan. Contacts, knockdowns and pile-ups all come from real physics. The pitch: sim-to-real starts with real environments - OVER crowdscans locations via smartphones and 360 cameras (174k+ locations, 86.8M images) and exports USDZ assets pairing a photorealistic Gaussian splat (vision) with a collision-ready mesh (physics) at true metric scale, importable into Isaac Sim/MuJoCo, so pixels and physics are anchored to the same real place to close the sim-to-real gap. Relevant both to multi-robot policy validation and simulation data infrastructure.
NVIDIA's open-source (Apache-2.0) orchestration layer, Rust for the performance path and Python for extensibility. It explicitly does not replace SGLang, TensorRT-LLM or vLLM - it wires them into a coordinated multi-node system: disaggregated prefill/decode, KV-aware routing (2x faster TTFT on Qwen3-Coder 480B), KVBM offloading KV cache to CPU/SSD/remote, ModelExpress streaming weights over NVLink for 7x faster cold starts, an SLA-driven Planner (80% fewer breaches at 5% lower TCO in Alibaba production), and Grove for topology-aware NVL72 scheduling.
Self-described as the first desktop app to run and train models (Windows/macOS/Linux; multi-GPU across NVIDIA, AMD, Intel, CPU and Vulkan). Fine-tuning is claimed at 2x faster with 70% less VRAM, the post-training menu covers LoRA, QLoRA, full fine-tuning, pretraining, GRPO, DPO and FP8, and export reaches GGUF, NVFP4 and FP8. Unsloth Start connects Claude Code or Codex to local weights in one command; the model list includes Qwen3.8, GLM-5.3-Flash, Kimi K3, DeepSeek-V4, MiniMax-H3 and Gemma 4, with private web search, deep research, auto-compaction and RAG built in.
The foundation-model layer: the weights themselves plus what is tightly coupled to them - training recipes (pretraining, post-training, RL/RLHF), architecture and long context, and the inference serving stack with its cost per token
2026-09-22MiMo-V2.6 Deep Read: Six Days of Live RL, an AA Index of 46 and a Fully Open Self-Improvement RunXiaomi open-sources MiMo-V2.6 Pro/Flash: 30 live RL steps, ~750k trajectories at $850k/$2.62M; AA Index 46 tops open models, DeepSWE v1.1 gains +17/+14 out of sample; Vibe World, CUA, science and content demos plus 7k+ RL environments released.
2026-08-10Stealing Reasoning Traces from Proprietary LLM APIsLeading large language model providers now conceal their models' step-by-step reasoning, or chain-of-thought, to protect intellectual property and limit information leakage. Rather than storing these traces server-side, providers return them to the client as blocks of encrypted text, which the client passes back with each subsequent request. Building on prior research, we identify an architectural vulnerability: these encrypted blocks are fully compatible and interchangeable across different sessions, users, and models within a provider's ecosystem. We exploit this compatibility to develop a scalable decryption jailbreak. By injecting an encrypted reasoning trace from a given model into a weaker, and less safeguarded model from the same provider, we force it to decode and output the trace verbatim in plaintext, without ever jailbreaking the more capable model directly. This vulnerability enables four distinct attack vectors. First, it circumvents anti-distillation mechanisms, allowing adversaries to extract a proprietary model's reasoning, as we demonstrate across Anthropic, OpenAI, and Google. Second, it allows for large-scale private data extraction. Developers frequently share session logs publicly, unaware of contents of the encrypted blocks. By decoding 315,320 reasoning blocks scraped from public repositories, we recovered 367 Personally Identifiable Information (PII) artifacts and 182 credentials. Third, it inadvertently reveals hazardous information hidden within the reasoning process, even in cases where the model's final, visible output safely rejects a malicious request. Fourth, attackers can leverage this flaw to execute invisible prompt injections, embedding malicious payloads entirely within encrypted blocks to poison public agentic rollouts. Following responsible disclosure, we propose concrete cryptographic and system-level mitigations to secure client-side reasoning.
2026-07-22LENS: LLM-guided Environment Simplification for Planning and Control in ClutterDespite recent advances in general-purpose robotic manipulation, real-world multi-object clutter remains challenging to handle for today's prevalent approaches. The problem scales in complexity due to more objects and collisions, more unpredictable contact physics, distractors, and task ambiguity. Bridging this gap to real-world deployment requires effective scene abstractions; yet today, producing such abstractions requires extensive task-specific manual engineering, which does not scale. These abstractions are costly to generate and difficult to adjust or fine-tune. We instead propose a plug-and-play fix to automatically generate scene-specific, task-specific, adaptively updating abstractions on top of existing planning and control stacks. LLM-guided Environment Simplification (LENS) produces a de-cluttered abstracted scene representation by merging (e.g., stacked objects) or pruning (e.g., distant objects) scene entities in a closed loop in response to task progress. These dynamic, task-relevant abstractions are versatile and easy to use. In our experiments, we show that LENS improves classical planning, model-based control, and a vision-language-action model, across a diverse set of highly cluttered manipulation scenes. Project website: https://lens-2026.github.io/.Agent runtimes and scaffolds: plugin architecture, sandboxing & tool surface, skill systems, long-term memory, self-improvement loops
2026-09-17SoL-Pi: Recursively Scaling Auto-Research Loops for Efficient Agent HarnessAs coding agents move from supervised code completion to unattended, around-the-clock exploration, their work expands into long trajectories of reasoning, tool use, and feedback. SoL-Pi scales recursive auto-research across executable environments to discover reusable harness mechanisms. Four retained mechanisms improve action execution, context compaction, observation handling, and delegated reading, reducing recorded token traffic by 44.7-49.0% and API cost by about one third at comparable EdgeBench performance.Code generation, repair, refactoring, software engineering tasks
2026-09-28mobile-mcp: the mobile automation MCP server driven by the native accessibility treemobile-next/mobile-mcp (8.2k stars, Apache-2.0) connects iOS simulators and real devices plus Android emulators and real devices to any MCP client, including Claude Code, Codex, Gemini and Copilot. Its first principle is accessibility-first: read the native accessibility tree instead of screenshots, no vision model and no image tokens, falling back to screenshots with coordinates only when necessary. Roughly 30 tools in seven groups cover device management, remote devices through Mobile Next Cloud, app management, screen interaction with recording, input and navigation, logs and crashes, and batch commands that compress N rounds into one. The README examples are all cross-app long-horizon tasks, the hardest class for mobile agents.2026-09-27GenOffice: Open-Source AI Office SuiteGenOffice is the first full-featured open-source AI office suite (Apache-2.0): real .docx/.xlsx/.pptx/PDF/Markdown editing with byte-preserving Word compatibility, block-level AI editing with document-aware agents, BYOK model support, fully local.Our takeIts premise is that models lack process discipline, not coding ability. The single best idea to steal is 'Rulings, not stalls': decide by default, and stop for a human only when an action is irreversible, touches a security-sensitive surface, leaks outside the worktree, or the plan is already too broken to continue. We adopted those four gates in our own pipelines.
v6.4.1 (2026-09-18), MIT, 289k stars — the largest single skill repository we have found. It chains brainstorming, spec writing, git-worktree isolation, planning, subagent execution, red-green TDD, two-way code review and verification-before-completion into one trunk workflow, and the README lists 16+ harnesses: Claude Code, Codex, Cursor, Gemini CLI, OpenCode, Kimi Code, Devin CLI and more.
/plugin install superpowers@claude-plugins-officialOur takeIt sells parts rather than a whole process: code-review, diagnosing-bugs and resolving-merge-conflicts each stand alone. grill-me, grilling and wait-what turn the agent around to interrogate your requirements, and nothing else in this batch does that. CONTEXT.md at the repo root decouples project vocabulary from generic skills — the cleanest version of that idea we have seen.
The personal working set of the author of Total TypeScript; the repo describes itself as real engineering, not vibe coding. 38 SKILL.md files, primary language Shell, grouped as engineering (18), productivity (7), misc (4) and in-progress (9). It explicitly rejects takeover-style frameworks such as GSD, BMAD and Spec-Kit in favour of small, easy to adapt, composable.
claude plugins install mattpocock-skillsOur takeIf you want to know how a SKILL.md is actually supposed to be written, read this instead of any second-hand tutorial. Two things to know first: docx, pdf, pptx and xlsx are source-available rather than open source (the other 15 are Apache-2.0), and the repo carries its own disclaimer — these skills are demos, so test them in your environment before trusting them with anything critical.
Anthropic's own implementation, 177k stars: 19 skills under ./skills plus ./spec (the Agent Skills specification, now at agentskills.io) and ./template as a scaffold. Its definition of a skill is folders of instructions, scripts and resources loaded dynamically, and progressive disclosure comes straight from here.
/plugin marketplace add anthropics/skillsOur takeIts sharpest self-description is that models can write CSS but lack a taste database. Ten rule categories are ordered by severity, with Accessibility and Touch & Interaction first as CRITICAL and Charts last — that ordering is the stance. The Design System Generator runs five parallel searches, and the repo states plainly: do not persist unverified output, so a generated design system has to be checked before it lands in your repo.
v2.13.0, MIT, 129k stars, site uupm.cc. Instead of lecturing the model it ships structured catalogues the agent queries before writing code: 79 UI styles, 192 palettes with reasoning, 74 type pairings, 119 UX rules, 105 icon suggestions, 17 GSAP presets, 25 chart types, 22 stacks and 34 landing-page patterns.
npx ui-ux-pro-max-cli init --ai claudeOur takeThe design that matters most: every edge is labelled EXTRACTED or INFERRED, so you can tell what the source explicitly contains from what the tool inferred — which makes an agent's conclusions auditable, a hard requirement in production. It is not a vector index: no embeddings, no vector store, a real graph you can explain, path and query. In its own benchmark, building the graph costs zero LLM spend and ingest is an order of magnitude cheaper, a structural win from local parsing.
YC S26, Apache-2.0, 120k stars, Python 3.10+. /graphify . maps code, docs, PDFs, images and videos into a graph and emits a clickable graph.html, a human-readable GRAPH_REPORT.md and a machine-readable graph.json. Code goes through tree-sitter AST parsing (about 40 languages): deterministic, no LLM calls, nothing leaves the machine.
uv tool install graphifyy && graphify installOur takeThere is no one-liner that installs this repo; use it as a discovery layer. A list answers what exists, not what is worth running. The part you cannot get elsewhere is connect-apps-plugin: a single MCP endpoint to 1000+ integrations with authentication, team-level ACLs and audit logs, which is exactly the governance problem enterprises hit when agents touch SaaS.
A list maintained by Composio at 75k stars with an Apache-2.0 badge in the README. About 31 first-party skill directories physically exist here while the index references 864 SKILL.md files, and most of the difference is links out to external repositories. Coverage explicitly extends beyond Claude.ai and Claude Code to Codex, Cursor, Gemini CLI and Antigravity.
claude --plugin-dir ./connect-apps-pluginRanks come verbatim from each evaluator's own board, synced into our database on a schedule. We never recompute, reweight or merge across sources: an intelligence index, an Elo and a success rate are different rulers, and any single overall number would be incomparable figures stitched into one that looks comparable.
Source · artificialanalysis.aiofficial Intelligence Index v4.3synced Sep 22, 2026
Text-to-image, image editing, consistency and controllability
Video generation and speech/audio in one channel: physical and temporal coherence for text/image-to-video, TTS naturalness and voice cloning, music generation, and ASR
2026-09-24Rolling-WAM: World Action Models with Rolling ImaginationWorld Action Models (WAMs) couple action generation with future visual prediction for robotic manipulation. However, completing the joint video-action denoising process at each replanning cycle incurs substantial latency, delaying action updates and limiting closed-loop responsiveness. We present Rolling-WAM, a formulation that distributes joint denoising across successive replanning cycles. Our method maintains a sliding window of video-action chunks at staggered noise levels. At each step, a rolling noise schedule fully denoises the imminent action chunk for execution, while partially refining farther-future chunks. As the window advances with new camera observations, the retained future chunks continue their denoising process. This distributes the computational cost over time while carrying an evolving visual-action context across chunk boundaries. Evaluations on LIBERO, RoboTwin, and a real-world Unitree G1 humanoid show that Rolling-WAM achieves competitive manipulation performance. By removing the need to denoise the entire prediction horizon from scratch, it delivers a 4.5x steady-state replanning speedup over standard joint WAMs.
2026-09-24BeyondRetarget: Learning Executable Humanoid Motions Directly from Monocular VideoBeyondRetarget removes the SMPL intermediate representation and maps monocular RGB video end-to-end to executable humanoid motions: 18 semantic keypoints with per-link non-uniform scaling form unified supervision, a shared representation decodes onto 8 humanoids, and contact-aware refinement yields 26.4mm error, zero collapse, and 193ms latency enabling real-time visual teleoperation.
2026-09-22MachEmbodied-U0: Unified Understanding and Generation Model for Embodied IntelligenceGeneral-purpose robot control requires models to understand task intent, identify where to interact, capture how the scene evolves, and generate precise actions. Vision-language-action models provide strong semantic priors but typically do not explicitly model scene dynamics, while world-action models couple visual prediction with control without necessarily exposing the task-relevant semantic and spatial structure needed for fine-grained manipulation. We present MachEmbodied-U0 (ME-U0), a unified embodied foundation model connecting understanding and generation experts through a Mixture-of-Transformers architecture. Subtask prediction and affordance grounding guide joint visual-dynamics and action generation via flow matching. Visual dynamics encompass future RGB, depth, surface normals, and optical flow, providing complementary supervision for appearance, geometry, and motion. Multi-rate Rotary Position Encoding (MRPE) aligns visual dynamics with fine-grained control. We pretrain ME-U0 on approximately 4,200 hours of curated demonstrations from robotic datasets and egocentric datasets. Using only the supervision natively available in each downstream benchmark, ME-U0 achieves an average score of 17.66 on the RoboDojo simulation benchmark and average success rates of 99.0\% and 82.5\% on LIBERO and LIBERO-Plus, respectively. We additionally validate ME-U0 on real-world robotic manipulation tasks, demonstrating its effectiveness beyond simulation. Without corresponding downstream supervision, ME-U0 further demonstrates zero-shot subtask prediction, affordance grounding, and visual dynamics on simulated and real-world observations. Overall, ME-U0 combines competitive downstream control performance with transferable task-grounding and visual-dynamics capabilities across simulation and the real world.
2026-09-04GE-Act 2.0: Pretraining and Scaling a World-Action Model for Robotic ManipulationGE-Act 2.0 combines a control-oriented autoencoder, a single-step visual planner, an inverse dynamics model, and knowledge-aligned selective optimization to pretrain a world-action policy from scratch on manipulation data. Scaling co-training data from 300 to 30,000 hours raises zero-shot OOD success to 44.1% on G1-OP and 31.1% on G2-90D without task-specific fine-tuning.3D asset generation, scene reconstruction, long-horizon world prediction
2026-09-21DexTacWAM: A Visuo-Tactile World-Action Model for Dexterous ManipulationDexterous manipulation depends on contact dynamics that are often only partially observable from vision. Recent World-Action Models (WAMs) couple predictive video world modeling with action generation, but remain largely vision-centric and therefore cannot directly model these contact dynamics. We present DexTacWAM, a visuo-tactile WAM that encodes each fingertip independently, aggregates the resulting features through a finger- and pose-aware tactile compressor, and injects the tactile latent into a video diffusion world model for joint visuo-tactile world modeling. Across six contact-rich dexterous manipulation tasks on a 22-DoF bimanual platform, DexTacWAM achieves the highest score on every task, averaging 70.6 versus 38.0 for the strongest baseline. Ablations attribute the gain to modeling contact evolution as part of the predicted world state rather than tactile conditioning alone: removing tactile world modeling reduces the four-task mean from 74.7 to 26.6 while keeping the same tactile features and action expert. After four hours of tactile-encoder adaptation with a frozen pretrained vision VAE, our continual vision-to-touch learning extends the pretrained video model to touch using roughly 100 demonstrations per task without tactile midtraining, while retaining visual prediction quality within 0.5 dB of vision-only counterparts. The compressor retains 89.4% of pre-fusion contact recall while enabling 2.26x faster training and 1.29x faster inference. Together, these results show that pretrained video priors can be extended to distributed multi-finger contact dynamics in a data- and compute-efficient manner.Visual design, UI components, web and slide layout
2026-07-25Fashion-3DLR: A Controllable 3D Garment Generation Using Pairwise Fashion Elements for Intelligent DesignAI-generated content (AIGC) has made significant progress, with 2D generative models becoming ready-to-use tools for the digital fashion industry. However, 3D garment generation remains in its nascent stage, where in the realm of fashion, the semantic information of diverse design elements exhibits intricate coupling relationships in 3D representations, posing substantial challenges for generating diverse 3D garments. In this work, to handle the above problem, We introduce Fashion-3DLR, a novel 3D garment generation framework that utilizes diverse design elements to create high-quality, versatile 3D garment assets. Specifically, to bridge the semantic gaps between different fashion elements, we propose a Garment Feature Fusion Diffusion Transformer (GFF-DiT) module to integrate 2D fashion design elements, e.g., sketch and texture, into latent space. Within the latent space, we then employ a rectified flow transformer to generate geometry latents, which can be decoded into various 3D garment representations, including 3D Gaussians and meshes. Furthermore, we integrate Fashion-3DLR into downstream tasks, achieving the 3D Gaussian Splatting (3DGS)-driven cloth physical simulation and mesh-based virtual try-on. Experimental results indicate that Fashion-3DLR surpass the previous state-of-the-art methods, which verify that the proposed work can generate well-structured, non-watertight garments capable of physical simulation and virtual try-on, underscoring its potential as a versatile 3D garment design tool.Slides, Word, Excel, PDF generation and editing

Physical Intelligence (Sergey Levine's team, Mar 19, 2026) present RL Tokens (RLT): freeze a pretrained VLA, attach an encoder-decoder that compresses its internal embeddings into a bottleneck RL token, then run online RL with a ~1M-parameter actor-critic on the real robot, refining only the critical phase (editing VLA action chunks rather than generating from scratch, anchored by a BC regularizer plus reference-action dropout). Across four sub-millimeter tasks the critical phase speeds up by up to 3x, screw success goes 20% to 65%, and half of Ethernet insertion episodes beat every human teleoperation demo - with just 15 minutes of real robot data.

Avid's 3,500-word build log: Jev as a decision layer inside keel, a local-first Rust coding workspace. One principle carries it: a model can suggest the next move, the host still owns the move — the host prepares the menu, Jev picks from it, the host checks again; pinned routes and live sessions bypass automatic selection; selection grants no permission. Plus how decision receipts become replayable evaluation scenarios.

Gergely Orosz visits OpenAI: Codex and ChatGPT Work have taken over everything; a nine-step agentic software factory with Perf Factory and Sevbot, and the rethink of IDEs and pull requests.
Under constructionThe games system is still under construction — everything listed starts in-page, and the library keeps growing.

A first-person match where every bot across from you is driven live by a policy this repo trained, the whole round stepping at a fixed 60 Hz: builtin, Rust/wasm or Rapier behind one physics interface, nothing scripted.
WASD move · mouse look · left-click fire · shift sprint · R reload · esc release

Unreal 5.5's first-person test map, running same-origin on this site: the level, the materials, the lights and four weapon data assets are read out of the project's own .umap / .uasset bytes, six soldiers are driven by a four-state brain on one fixed 60 Hz step, and the sidebar lists every substitution the page had to make.
WASD move · mouse look · left-click fire · shift sprint · space jump · R reload · F pick up · esc release

A 100 km 3D city that opens in seconds — the digital twin of the AI agents that build this site: crawling, blogs, papers, industry and investment analysis.
100km city · opens in seconds · Live task stream · Agent-maintained

An SO-101 6-DoF arm running MuJoCo physics in your browser: joint teleoperation, IK end-effector dragging and gripper pick-and-place into a basket. No install.
MuJoCo WASM physics · SO-101 · 6 DoF · Contacts / telemetry HUD

Take a humanoid joint module apart layer by layer: brushless motor, magnetic encoder, planetary / harmonic / cycloidal drive and output flange. Switch architectures on one page — exploded view plus analytic kinematics.
Three gearbox types · Planetary / harmonic / cycloidal · Analytic kinematics · explode

Matcha-TTS mixed zh/en speech synthesis: server-side synthesis with sentence-streamed playback and full-article read-aloud for papers and blogs. Sign-in required.
Matcha-TTS · server-side · Mixed zh / en · Full-article read-aloud

A caring coding sprite on your desktop and a cockpit that never stops: the work resumes itself after a crash or a reboot. Every coding terminal, ssh session and local shell in one tree, every start, finish and permission request reported in time. Open source, one-line install.
Your caring coding desktop sprite · Resumes after crash & reboot · First-rate terminal / ssh / shell

Built for games and for physical AI: an ultra-realistic, ultra-high-performance simulation environment with ultra-low-cost procedural data; Unreal-grade visuals open and play in the browser, bridging the physical world and intelligent agents.
Games meet physical AI · Unreal-grade on the web · Open and play

The first online photo editor built to be driven by AI agents: fifty-two professional tools, a layer tree with masks, PSD files that open and save back, nothing to install, and your images never leave your computer. The agent layer is on its way.
retouchpi.com · 52 pro tools · PSD read and write · your files never leave