NVIDIA's TensorRT Edge-LLM Runs Qwen3.6-27B 6.4x Faster in MLPerf Edge Agentic
NVIDIA says TensorRT Edge-LLM ran Qwen3.6-27B through the MLPerf Edge Agentic workload in 24 minutes 36 seconds on one Jetson AGX Thor. The 6.4x gap over llama.cpp puts long-context agent inference on edge.