HuggingFace 11个
Daily Papers
1 ActiveSaddler: Automated Curriculum Learning for Agent Harness Optimization
2 Argo-Bench: Evaluating Data Agents on Enterprise-Scale Workflows
3 Ego2Act: Evaluating Goal-Directed Manipulation in Egocentric Video Generation
4 Decoding Looped Transformers Better for (Almost) Free
5 KaliBench: A Fine-Grained Benchmark for Cybersecurity Tool Use on Kali Linux with Runtime-Free Verifiable Rewards
6 ScholarCatalyst: A Benchmark for Retrieving Papers That Inspire New Research
7 AgSpec: Pushing the Limits of Retrieval-Based Speculative Decoding in Coding Agent Pipelines
8 Where-OPD: Spatially Guided On-Policy Self-Distillation of MLLMs with Synthetic Scenes
9 InterEvolve: Test-Time Evolution of Reward Programs for Humanoid Loco-Manipulation
10 AutoGUIWorld: Image Generators as Visual World Models for GUI Agent
11 ROWBench: Do Video Models Render What the Program Specifies?
12 World Observer: Joint Actor-Observer Generation for Persistent World Modeling
13 Beyond Memory: Harnessing Long-Horizon Agents with Explicit Belief States
14 Sharpening Tax in Post-Training
15 Hierarchical Continuous Diffusion Language Models
16 RPTune: Learned Context Curation for LLM Catalog Search
17 Omni-Embed-Mini: Binding Modalities Without Forgetting via Dense Distillation
18 Scaling and Distilling Text Embeddings for Better Diffusibility
19 OmniSeek: Native Tool Integration for Multi-turn Audio-Visual Reasoning
20 Fewer Tokens, Better Action: GPT-6 Astra Robot Agents with 14% Higher Success Rate but 65% Fewer Tokens
21 VTR-Bench: A Systematic Benchmark for Evaluating Visual Text Rendering in Video Generation
22 SILSA: Sliding-Window Slice Latents for Topology-Preserving High-Resolution 3D Generation
23 Generalization Is Stability, Not Accuracy: Multi-Axis Evaluation of LLMs
24 4Director: Controlling Video World Models with Rigid 3D Geometry
25 It Takes Workflows to Evolve Better Workflows
26 Keyword Harnesses Fail Open: A Cheap Diagnostic Ladder for Tool-Use Claims in Small Language Models
27 Pay for the Fault, Not the Flow: Label-Free In-Flow Multi-Agent Workflow Optimization
28 Before It Fades: Reinforcing Temporal Representations at Inference Time in VideoLLMs
29 Latent-Foresight: End-to-End Learning Predictable Representations for Latent World Models
30 OneStreamer: Unifying Perception, Memory, and Proactive Response in Streaming Video Interaction
31 LOCI: Spatial Linear Memory for Streaming World Models
32 PhysVista: Benchmarking Physical Intelligence in VLMs via a Perception-Reasoning-Assessment Loop
33 Multimodal Flow: Unified Flow Modeling of Language and Vision in Embedding Spaces
34 JevSpawn: Adaptive Agentic Inference through Compositional Action Spaces
35 Memorizon: Training World Models Beyond Their Context Window
36 Align Then Reason: A Multimodal Lip-Sync Judge for Dubbing
37 Smaller Models, Better Rejects: Preference Distillation Scaling
38 Make Sparse Rewards Count: Density-Aware Reward Aggregation for Multi-Reward RL
39 A Missing Piece for Trustworthy AI Reviewers: From Benchmarking Rhetorical Robustness to SciCore Review
40 GraphForge: Training Working Agents with Graph-Anchored Workspace Synthesis
41 PixelDense: Dense Prediction as Representation Alignment for Pixel Diffusion
42 EgoTools: Towards Tool-Centric Reasoning in Real-World Egocentric Videos
43 Video Generation Models: A Survey of Post-Training and Alignment
44 Benchmarking and Enhancing Skill-Level Memory for Partially Observable Robotic Manipulation
45 Beyond the Current Scene: Event-Referential Grasping with Active View Selection
46 SemanTok: Predictable Semantic Tokens for Efficient Autoregressive Video Generation
47 Architect-Ant: Editable Automatic Furnishing of Architectural Floor Plans
48 Personalized Image Generation with Reasoning and Reflection
49 AutoDataBench: A Data-centric Testbed for Accelerating Auto Research
50 DexPolicy: Scheduled Exploration for Trajectory-Guided Dexterous Manipulation
51 Honeycomb: Constant-Size Scene Memory Representation for Video World Models
52 Better Supervision Is Nearby: Neighborhood On-Policy Self-Distillation
53 FlexRouter: Learning Complementary Model Sets for Flexible LLM Routing
54 MemFold: Learning Compact Soft Memory for Long-Context Personalization via On-Policy Optimization
55 Explore Broadly, Reason Sharply: Push Small Models toward the Frontier via Sampling
56 Prefill-Free Cross-Family KV Cache Transfer for Heterogeneous Multi-Agent LLMs
57 Prompt2Skill: Unsupervised Skill Optimization From Natural Language Instructions
58 Rules to Tools: Executable Checks for LLM Agents in Scientific Computing
59 Removing the NEEDLE in the Haystack: Backdoor Removal in LLMs via Weight Orthogonalisation
60 Predictive Credit: Measuring What Scientific Explanations Add to Experimental Forecasts
61 Pretrain Once, Route Anywhere: Towards a Foundation Model for LLM Routing
62 RLE-Bench: A Qualifying Exam for Coding Agents as Robot Learning Engineers
63 CorrGRPO: Correlation-Normalized GRPO for Multi-Reward Learning
64 Retrieval-Augmented Skill Optimization via Cross-Harness Adaptation
65 E-MoE: Enhanced Mixture-of-Experts for Non-Factorized Diffusion Language Models
66 Controlled Decoding Attacks on Black-Box LLMs
67 Devils in Question Relay: Source-Conditioned Relay Steering to Mitigate Hallucinations in Audio-visual Large Language Models
68 Transformers Stop Thinking Too Early, and a Tiny LoRA Fixes It
69 OTRetarget: Joint Robot and Object Motion Retargeting via Optimal Transport
70 Adaptive Reward Routing: Dynamic Multi-Reward Optimization for Joint Audio-Video Diffusion via Forward-Process RL
71 When Does Correction Become Repair? Mechanistic Auditing of Internal Interventions in Tool-Using LLMs
72 Learning What to Recall: Adaptive Multi-Cue Episodic Memory for World Models
73 Joint and Cross-Modal Video-Audio Generation and Editing: A Unified Formulation and Design Taxonomy
74 Does Native 3D Texture Generation Necessarily Require 3D Assets for Training?
75 Persona Dosing: Calibrated Activation Steering for Graded Trait Control
76 Agent Priors-guided Policy Learning
77 On-Policy or Off-Policy Learning? A Systematic Study of Distillation Dynamics
78 DataMagic: Authoring Data Videos through Declarative Multi-Agent Orchestration
79 Replacing Large Language Models with Jev Decision Models for Low-Latency Edge Service Orchestration
80 Do Audio LLMs Listen Before They Act? Diagnosing Acoustic-Context Gating in Voice Agents
81 OpenTumorBoard: A Real-World Benchmark of Multidisciplinary Tumor Board Discussion Trajectories
82 When Users Change Their Minds: Measuring and Repairing Intent Drift in LLM Agents
83 X-Tree: Tokenizing Reusable Experience for Efficient Agent Generalization
84 Decentralized Master-Mind: Joint Action Refinement through Iterative Intent Denoising in Multi-Agent Pathfinding
7分钟前
1 The Agent Said It Was Done. The Database Disagreed.
2 Open-sourcing AstaBrief, the fast report-generation model in Asta
3 AutoSynthData: Generating Training Data for Enterprise Agents
4 Open TTS Leaderboard: Scalable Evaluation for Multilingual Text-to-Speech and Voice Cloning
5 NVIDIA Kumo Tabular Sets a New Accuracy-Efficiency Frontier for Tabular Prediction
6 Getting the Source Right, Not Just the Fact: Source-Aware Verification for MCP Agents
7 Holo4: powering generalist computer-use agents
8 Accelerating vision-language models with LFM2.5-VL-DSpark
9 Jun Kim, oMLX creator and maintainer, joins Hugging Face to support the MLX community
10 Transformers now runs llama.cpp quants
11 How UK AISI and EvalEval Are Making Benchmark Results Reproducible
12 tokenizers v1: encode, decode and scaling, measured
13 Your Agent Aced the Task. Will It Do It Again?
14 Rebuilding AUTOMATIC1111 with Gradio Workflow
15 Async GRPO with LoRA across HF Jobs: a bucket, a proxy, and no NCCL
16 NeoMME: an efficient Multimodal-native and Multilingual Encoder
17 Training a coding model to paint watercolours with TRL and OpenEnv
18 Give Your Coding Agents a Memory You Own
19 Fine-tuning a 350M Model for Better Structured Outputs in 100 GRPO Steps
20 BenchMIRT: What are LLM benchmarks actually measuring?
21 Introducing @huggingface/kernels: 200+ WebGPU Kernels for Local AI
22 The Open ASR Leaderboard Adds Its First Global South Language
23 Training and Finetuning Multi-Vector Embedding Models with Sentence Transformers
24 Granite 4.2 LLMs: How They're Built
25 Quantization-Aware Healing: a compressed, 4-bit model that outperforms its full-precision original
26 Wire It, Run It, Deploy It: AI Workflows in Gradio
27 Measuring benchmark optimization in speech recognition
28 How Hugging Face Inference Endpoints, Jobs, and Buckets Power Search on Papers with Code
29 Up to 3.2x Faster Inference with LFM2.5-DSpark
30 How Much Memory Does Your Agent Actually Need?
31 Multi-Vector (Late Interaction) Embedding Models with Sentence Transformers
32 Same Cluster, 33 Points More Utilization: What Changed Was the Order
33 State of Open Models: Summer 2026 Observations
34 Record, train, and deploy from one place with Strands Agents, LeRobot, and Hugging Face Storage Buckets
35 What We Learned by Reproducing 2,200 papers from ICML
36 Introducing OlmoEarth embeddings: Custom embedding exports from OlmoEarth Studio for downstream analysis
37 Thinking of ACE? We Can Do It with Fewer Tokens
38 Build Low-Latency Multilingual Voice Agents: Open Weights & Full Deployment Control with NVIDIA Magpie TTS
39 Making Knowledge Distillation Cheap Enough to Run at Scale
40 Meta is back with Muse Glimmer: local, agentic, multimodal, and open source
41 Baseten on Hugging Face Inference Providers 🔥
42 GPU Management: Why Idle GPUs Are the New Grounded Aircraft
43 NVIDIA Cosmos-H-Dreams: Bringing Real-Time Generative Simulation to Surgical Robotics
44 Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 Incident
45 Bringing Nunchaku 4-bit Diffusion Inference to Diffusers
46 Grabette: an open system to record robot-manipulation data
47 Newer Models, Same Advantage
48 Security incident disclosure — July 2026
49 Model Routing Is Simple. Until It Isn’t.
50 Introducing Real World VoiceEQ: Measuring the human quality of voice AI
51 Welcome Inkling by Thinking Machines
52 Profiling in PyTorch (Part 3): Attention is all you profile
53 Native-speed vLLM transformers modeling backend
54 From Hugging Face to Amazon SageMaker Studio in one click
55 Hugging Face Models on Foundry Managed Compute
56 Run AI workloads on any cloud, store on Hugging Face: zero-egress storage with SkyPilot
57 LeRobot v0.6.0: Imagine, Evaluate, Improve
58 PRX Part 4: Our Data Strategy
59 🤗 Kernels: Major Updates
60 Hugging Face and Cerebras bring Gemma 4 to real-time voice AI
61 ScarfBench: Benchmarking AI Agents for Enterprise Java Framework Migration
62 Why Specialization Is Inevitable
63 Featuring Every Eval Ever Results on Hugging Face Model Pages
64 DiScoFormer: One transformer for density and score, across distributions
65 Run a vLLM Server on HF Jobs in One Command
66 Accelerating Transformers Fine-Tuning with NVIDIA NeMo AutoModel
67 Introducing the FFASR Leaderboard: Benchmarking ASR in the Real World
68 Experimenting with the proposed Cross-Origin Storage API in Transformers.js
69 Shipping huggingface_hub every week with AI, open tools, and a human in the loop
70 PP-OCRv6 on Hugging Face: 50-Language OCR from 1.5M to 34.5M Parameters
71 We got local models to triage the OpenClaw repo for FREE!*
72 MosaicLeaks: Can your research agent keep a secret?
73 Is it agentic enough? Benchmarking open models on your own tooling
74 Beyond LoRA: Can you beat the most popular fine-tuning technique?
75 From the Hugging Face Hub to robot hardware with Strands Agents and LeRobot
76 GLM-5.2: Built for Long-Horizon Tasks
77 Agentic Resource Discovery: Let agents search
78 Profiling in PyTorch (Part 2): From nn.Linear to a Fused MLP
79 How an Agent Built a 3D Paris Gallery by Chaining Two Hugging Face Spaces
80 Migrating Your GitHub CI to Hugging Face Jobs
81 The Open Source Community is backing OpenEnv for Agentic RL
82 Nemotron 3.5 Content Safety: Customizable Multimodal Safety for Global Enterprise AI
83 Designing the hf CLI as an agent-optimized way to work with the Hub
84 Direct Preference Optimization Beyond Chatbots
85 Adding MCP Tools to Reachy Mini
86 Holo3.1: Fast & Local Computer Use Agents
87 Introducing Mellum2: A 12B Mixture-of-Experts Model by JetBrains
88 Beyond LLMs: Why Scalable Enterprise AI Adoption Depends on Agent Logic
89 Profiling in PyTorch (Part 1): A Beginner's Guide to torch.profiler
90 Shipping a Trillion Parameters With a Hub Bucket: Delta Weight Sync in TRL
91 Reachy Mini goes fully local
92 Harness, Scaffold, and the AI Agent Terms Worth Getting Right
93 OlmoEarth v1.1: A more efficient family of Earth observation models
94 Introducing the Ettin Reranker Family
95 PaddleOCR 3.5: Running OCR and Document Parsing Tasks with a Transformers Backend
96 Granite Embedding Multilingual R2: Open Apache 2.0 Multilingual Embeddings with 32K Context — Best Sub-100M Retrieval Quality
97 Unlocking asynchronicity in continuous batching
98 Building Blocks for Foundation Model Training and Inference on AWS
99 vLLM V0 to V1: Correctness Before Corrections in RL
100 Adding Benchmaxxer Repellant to the Open ASR Leaderboard
7分钟前