As an Amazon Associate, we earn from qualifying purchases.
As an Amazon Associate, we earn from qualifying purchases.






The ModelScope team at Alibaba formulates QA over dynamic raw documents as Budgeted Evidence Localization and proposes LENS, an index-free framework: a low-cost prior compresses the search domain, a budget-constrained propose–observe–update loop narrows candidate evidence windows via an LLM relevance oracle, and selected regions consolidate into a source-grounded evidence set. On a controlled 500-question evaluation, LENS ties ReAct on answer quality (62.4% vs 65.2% EM, p=0.1143) while dominating evidence recall (84.8% vs 50.4%, +34.4pp) and grounding (96.8% vs 71.8%); with zero indexing over 15,517 raw Wikipedia shards, answers tie while LENS grounds more answers (84.0% vs 70.7%). After corpus growth, stale-index systems collapse to 0–5.5% evidence recall while LENS drops only 0.8pp.

MIT presents GaussLite, the first 3DGS mapping system that takes a natural-language task and allocates representation capacity online by task relevance: a one-shot LLM parser extracts target and anchor objects, Grounding DINO + FastSAM ground them per frame into relevance masks that steer seeding density, gradient flow and initial scale. At matched Gaussian budget and real-time 4 Hz mapping, ROI PSNR improves by +2.72 dB on Replica and +2.23 dB on real campus scenes; multi-agent maps fuse via per-voxel voting on active-optimization counts, beating concatenation by +3.42 dB while sharing only 7.08% of the map.

A conversation-first, image-first PPT workflow skill: build a content basis, confirm style with real 16:9 previews, lock design_spec / slide_blueprint / spec_lock, then retouch through a bundled review shell before exporting PPTX. Page visuals are rendered by an image model, so native per-element editing is not promised.

NVIDIA open-sourced a frame-level autoregressive framework for real-time interactive long video. KV-recache swaps prompts mid-generation, short window attention plus a frame sink holds long-range coherence, and it sustains 20.7 FPS on a single H100 for clips up to 240 seconds.
Modly is an open-source desktop app from Lightning Pixel (Windows, Linux, Apple Silicon macOS; v0.4.3; MIT plus an attribution clause) that turns a photo into a usable 3D mesh with open models running on your own GPU, so inference never leaves the machine. An Electron shell supervises a Python FastAPI backend bound to 127.0.0.1:8765. The default model is TripoSR (about 2.4 GB); Hunyuan3D 2 Mini, TripoSG and Trellis2 GGUF arrive as extensions, where an extension is a GitHub repo with a manifest.json that can declare multi-repo model_sources, shared weight_groups and separately installable weight_variants carrying size_gb and vram_gb. Generation is expressed as a node-graph workflow (Image to Generate Mesh to Add to Scene), scenes follow the modly.scene-manifest.v1 contract, and export covers glb/stl/obj/ply with a dedicated stl/obj printing path that scales to millimetres for OrcaSlicer. Three agent surfaces ship in the repo: a stdlib-only JSON-first CLI with its own SKILL.md, an MCP server exposing nine modly_* tools, and an in-app tool-calling agent that can run on a local llama.cpp GGUF pool or an external model endpoint. AMD ROCm auto-detection, memory-budgeted Apple Silicon workflows and headless Jetson runs are documented with measured numbers.

LiveTalking connects LLMs, TTS and three lip-sync backends into a real-time digital-human pipeline with WebRTC, RTMP, virtual-camera output, interruption, recording and multi-session support.
Foundation models: pre- and post-training, architecture and long context, inference serving and cost per token.



Agent runtimes: plugins and sandboxes, tool surfaces, skill systems, long-term memory and self-improvement.




Code generation, repair, refactoring, software engineering tasks



Editor's takeIts premise is that models lack process discipline, not coding ability. The single best idea to steal is 'Rulings, not stalls': decide by default, and stop for a human only when an action is irreversible, touches a security-sensitive surface, leaks outside the worktree, or the plan is already too broken to continue.
v6.4.1 (2026-09-18), MIT, 289k stars — the largest single skill repository of its kind. It chains brainstorming, spec writing, git-worktree isolation, planning, subagent execution, red-green TDD, two-way code review and verification-before-completion into one trunk workflow, and the README lists 16+ harnesses: Claude Code, Codex, Cursor, Gemini CLI, OpenCode, Kimi Code, Devin CLI and more.
/plugin install superpowers@claude-plugins-officialEditor's takeIt sells parts rather than a whole process: code-review, diagnosing-bugs and resolving-merge-conflicts each stand alone. grill-me, grilling and wait-what turn the agent around to interrogate your requirements, and nothing else in this batch does that. CONTEXT.md at the repo root decouples project vocabulary from generic skills — the cleanest version of that idea in this batch.
The personal working set of the author of Total TypeScript; the repo describes itself as real engineering, not vibe coding. 38 SKILL.md files, primary language Shell, grouped as engineering (18), productivity (7), misc (4) and in-progress (9). It explicitly rejects takeover-style frameworks such as GSD, BMAD and Spec-Kit in favour of small, easy to adapt, composable.
claude plugins install mattpocock-skillsEditor's takeIf you want to know how a SKILL.md is actually supposed to be written, read this instead of any second-hand tutorial. Two things to know first: docx, pdf, pptx and xlsx are source-available rather than open source (the other 15 are Apache-2.0), and the repo carries its own disclaimer — these skills are demos, so test them in your environment before trusting them with anything critical.
Anthropic's own implementation, 177k stars: 19 skills under ./skills plus ./spec (the Agent Skills specification, now at agentskills.io) and ./template as a scaffold. Its definition of a skill is folders of instructions, scripts and resources loaded dynamically, and progressive disclosure comes straight from here.
/plugin marketplace add anthropics/skillsEditor's takeIts sharpest self-description is that models can write CSS but lack a taste database. Ten rule categories are ordered by severity, with Accessibility and Touch & Interaction first as CRITICAL and Charts last — that ordering is the stance. The Design System Generator runs five parallel searches, and the repo states plainly: do not persist unverified output, so a generated design system has to be checked before it lands in your repo.
v2.13.0, MIT, 129k stars, site uupm.cc. Instead of lecturing the model it ships structured catalogues the agent queries before writing code: 79 UI styles, 192 palettes with reasoning, 74 type pairings, 119 UX rules, 105 icon suggestions, 17 GSAP presets, 25 chart types, 22 stacks and 34 landing-page patterns.
npx ui-ux-pro-max-cli init --ai claudeEditor's takeThe design that matters most: every edge is labelled EXTRACTED or INFERRED, so you can tell what the source explicitly contains from what the tool inferred — which makes an agent's conclusions auditable, a hard requirement in production. It is not a vector index: no embeddings, no vector store, a real graph you can explain, path and query. In its own benchmark, building the graph costs zero LLM spend and ingest is an order of magnitude cheaper, a structural win from local parsing.
YC S26, Apache-2.0, 120k stars, Python 3.10+. /graphify . maps code, docs, PDFs, images and videos into a graph and emits a clickable graph.html, a human-readable GRAPH_REPORT.md and a machine-readable graph.json. Code goes through tree-sitter AST parsing (about 40 languages): deterministic, no LLM calls, nothing leaves the machine.
uv tool install graphifyy && graphify installEditor's takeThere is no one-liner that installs this repo; use it as a discovery layer. A list answers what exists, not what is worth running. The part you cannot get elsewhere is connect-apps-plugin: a single MCP endpoint to 1000+ integrations with authentication, team-level ACLs and audit logs, which is exactly the governance problem enterprises hit when agents touch SaaS.
A list maintained by Composio at 75k stars with an Apache-2.0 badge in the README. About 31 first-party skill directories physically exist here while the index references 864 SKILL.md files, and most of the difference is links out to external repositories. Coverage explicitly extends beyond Claude.ai and Claude Code to Codex, Cursor, Gemini CLI and Antigravity.
claude --plugin-dir ./connect-apps-pluginSource · artificialanalysis.aisynced Oct 11, 2026
Text-to-image, image editing, consistency and controllability


Video and speech/audio: physical and temporal coherence for text/image-to-video, TTS and voice cloning, music generation and ASR.






3D asset generation, scene reconstruction, long-horizon world prediction



Visual design, UI components, web and slide layout
Slides, Word, Excel, PDF generation and editing

Uber's Legal Redlining Agent lives inside Microsoft Word: four iterations (RAG, lawyer feedback, tone, agentic modifications) cut average review time over 20% with 91% decision accuracy.

NVIDIA ships 38 Isaac Sim Agent Skills with the distribution: procedures, handoff contracts and acceptance thresholds in SKILL.md teach coding agents to finish and validate robot sims like engineers.

Google's open EmbeddingGemma 2 maps text, code, images, audio and video into one 768-d space with 740M params (270M text-only). We unpack modular loading, MRL truncation and on-device RAG.
Under constructionThe games system is still under construction — everything listed starts in-page, and the library keeps growing.

A first-person match where every bot across from you is driven live by a policy this repo trained, the whole round stepping at a fixed 60 Hz: builtin, Rust/wasm or Rapier behind one physics interface, nothing scripted.
WASD move · mouse look · left-click fire · shift sprint · R reload · esc release

Unreal 5.5's first-person test map, running same-origin in the browser: the level, the materials, the lights and four weapon data assets are read out of the project's own .umap / .uasset bytes, six soldiers are driven by a four-state brain on one fixed 60 Hz step, and the sidebar lists every substitution the page had to make.
WASD move · mouse look · left-click fire · shift sprint · space jump · R reload · F pick up · esc release

Fifty-two professional tools, layers and masks, PSD files that open and save back - all in the browser, nothing to install.
retouchpi.com · 52 pro tools · PSD read and write · AI generate & edit

A 100 km 3D city that opens in seconds — crawling, blogs, papers, industry and investment analysis, all watchable in one digital twin.
100km city · opens in seconds · Live task stream · Agent-maintained

An SO-101 6-DoF arm running MuJoCo physics in your browser: joint teleoperation, IK end-effector dragging and gripper pick-and-place into a basket. No install.
MuJoCo WASM physics · SO-101 · 6 DoF · Contacts / telemetry HUD

Take a humanoid joint module apart layer by layer: brushless motor, magnetic encoder, planetary / harmonic / cycloidal drive and output flange. Switch architectures on one page — exploded view plus analytic kinematics.
Three gearbox types · Planetary / harmonic / cycloidal · Analytic kinematics · explode

Matcha-TTS mixed zh/en speech synthesis: server-side synthesis with sentence-streamed playback and full-article read-aloud for papers and blogs. Sign-in required.
Matcha-TTS · server-side · Mixed zh / en · Full-article read-aloud

A caring coding sprite on your desktop and a cockpit that never stops: the work resumes itself after a crash or a reboot. Every coding terminal, ssh session and local shell in one tree, every start, finish and permission request reported in time. Open source, one-line install.
Your caring coding desktop sprite · Resumes after crash & reboot · First-rate terminal / ssh / shell