-
Notifications
You must be signed in to change notification settings - Fork 250
Pull requests: SemiAnalysisAI/InferenceX
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
[Klaud Cold] minimaxm3-fp4-b300-vllm-agentic-mtp: long-prefill-token-threshold 512 / 设置 long-prefill-token-threshold 512
full-sweep-fail-fast
#2539
opened Aug 9, 2026 by
xinli-sw
Collaborator
Loading…
[Klaud Cold] minimaxm3-fp4-b200-vllm-agentic-mtp: long-prefill-token-threshold 512 / 设置 long-prefill-token-threshold 512
full-sweep-fail-fast
#2538
opened Aug 9, 2026 by
xinli-sw
Collaborator
Loading…
fix(evals): preserve GPQA reasoning responses / 修复 GPQA 推理响应截断
#2537
opened Aug 9, 2026 by
JordanNanos
Collaborator
Loading…
Retune DSV4 B300 AgentX MTP sweep / 调优 DSV4 B300 AgentX MTP 扫描
agentx
AgentX benchmarks, recipes, and infrastructure
full-sweep-fail-fast
#2536
opened Aug 8, 2026 by
ivanium
Collaborator
Loading…
Add dsv4f-fp8-b200-vllm: DeepSeek-V4-Flash-0731 FP8 B200 vLLM single-node recipe
#2535
opened Aug 8, 2026 by
stewtong
Loading…
[Klaud Cold] dsv4-fp4-b200-vllm: nightly-700d39b image, prefill-schedule-interval 16 at CONC≥256 / 更新 B200 DSv4 镜像并在高并发下启用 prefill 调度间隔
full-sweep-fail-fast
#2534
opened Aug 8, 2026 by
xinli-sw
Collaborator
Loading…
[tilert] Add glm5.1-fp8-b200-tilert 1k1k+8k1k: vLLM prefill + TileRT decode PD-disaggregation / 新增 glm5.1-fp8-b200-tilert 1k1k+8k1k:vLLM prefill + TileRT decode PD 分离
all-evals
Expand eval selection to every fixed-sequence config
non-canary-full-sweep-enabled
Run the full sweep without the canary gate (full search space, no trim)
#2533
opened Aug 8, 2026 by
Oseltamivir
Collaborator
Loading…
CollectiveX: b300 deepep-v2 LL EP8 (single-node HCA pin) + H-series LL EP16 bring-up findings
#2525
opened Aug 7, 2026 by
Oseltamivir
Collaborator
Loading…
perf(agentx): refresh dsv4-fp4-gb300-dynamo-sglang-agentic harness
full-sweep-enabled
#2520
opened Aug 7, 2026 by
cquil11
Collaborator
Loading…
CollectiveX: kv-transfer suite — NIXL + MoRI-IO KV-cache handoff benchmark (stacked on #2489)
#2510
opened Aug 6, 2026 by
Oseltamivir
Collaborator
Loading…
[AMD][AgentX] Add Kimi-K3 MXFP4 MI355X vLLM agentic MTP recipe
agentx
AgentX benchmarks, recipes, and infrastructure
AMD
full-sweep-enabled
#2508
opened Aug 6, 2026 by
seungrokj
Collaborator
Loading…
3 tasks
[Power] feat: extend dcgm energy lanes to gb dsv4/qwen3.5 fp4 / 扩展 dcgm 能耗采集到 gb dsv4 与 qwen3.5 fp4
#2507
opened Aug 5, 2026 by
edwingao28
Collaborator
Loading…
GLM-5.2-FP8 H200 llmd-vllm P/D disagg agentic benchmark
full-sweep-enabled
#2499
opened Aug 5, 2026 by
elvircrn
Collaborator
Loading…
3 tasks
[NV] llm-d-vllm: optimize DSv4-Pro GB200 recipe configs
full-sweep-enabled
#2498
opened Aug 5, 2026 by
ilmarkov
Collaborator
Loading…
[AMD][WIP][AGENTX]: Kimi-K3 DSpark on MI355X
agentx-fast
Run AgentX throughput with 1 warmup request per lane and a 20-minute profile; not reusable
AMD
full-sweep-fail-fast
#2496
opened Aug 5, 2026 by
haic0
Collaborator
Loading…
3 of 5 tasks
[AMD] [WIP] [AGENTX] GLM-5.2 FP4 MI355X SGLang Agentic MTP
agentx
AgentX benchmarks, recipes, and infrastructure
AMD
full-sweep-fail-fast
#2488
opened Aug 4, 2026 by
giovanniguastiamd
Collaborator
Loading…
[AMD] [WIP] [AGENTX] MiniMax-M3 Support on MI355X with MTP
agentx
AgentX benchmarks, recipes, and infrastructure
AMD
full-sweep-fail-fast
#2487
opened Aug 4, 2026 by
ajith-sirra-amd
Collaborator
Loading…
config: make aggregate multinode topology explicit
#2479
opened Aug 3, 2026 by
cquil11
Collaborator
•
2/2
Loading…
[AgentX]: B200 Kimi K3 DSpark with corrected harness
#2475
opened Aug 3, 2026 by
cquil11
Collaborator
Loading…
Refresh Qwen3.5 FP4 GB300 AgentX
full-sweep-enabled
#2474
opened Aug 3, 2026 by
cquil11
Collaborator
Loading…
Previous Next
ProTip!
Find all pull requests that aren't related to any open issues with -linked:issue.