Cptz's blog
Share my ideas, thoughts and experiences.
HOME
CATEGORIES
TAGS
ARCHIVES
ABOUT
Home
Tags
Tags
Cancel
Tags
abi
1
ada
1
ai infra
12
algbw
1
algorithm
1
all-gather
2
all-reduce
3
all-to-all
1
allocator
1
allreduce
2
ampere
1
architecture
2
async-copy
1
asynchronous
2
bandwidth
1
barrier
1
baseline
1
benchmark
2
blackwell
1
bootstrap
1
bucket
1
buffer
1
busbw
1
callgraph
1
canary
1
capstone
1
cdp2
1
channel
4
chunk
1
chunked prefill
1
cluster-launch-control
1
collective
1
collectives
1
collnet
1
communicator
2
connector
1
continuous batching
1
cooperative-groups
1
cp-async
1
cpu-affinity
1
cuda
37
cuda-graph
1
cuda-graphs
1
cuda-ipc
1
cuda-kernel
1
cuda-malloc-async
1
cuda-stream
1
cuda-visible-devices
1
cumem
1
ddp
3
debug
1
debugging
1
dependent-launch
1
device-runtime
1
diagnosis
1
diagnostics
1
direct3d
1
distributed system
1
distributed training
24
dmabuf
1
double-buffering
1
double-tree
1
driver-api
2
dynamic-linker
1
dynamic-loading
1
dynamic-parallelism
1
egm
1
elf
1
enqueue
1
event
1
experiment
1
expert parallel
1
failure
1
fence
1
flight-recorder
1
fsdp
2
governance
1
gpu
16
gpu architecture
1
gpudirect-rdma
1
grace-hopper
1
graph
2
green-context
1
group
1
hmm
1
hopper
3
image bed
1
infiniband
1
infra
4
interop
1
interview
35
iommu
1
ipc
2
kv cache
12
l2-cache
1
latency
1
launch-overhead
1
lazy-loading
1
ll
1
ll128
1
llm inference
24
logging
1
megatron
12
megatron core
12
memory
2
memory-model
1
memory-pool
1
memory-registration
1
minio
1
model parallel
12
module
1
moe
1
multi-gpu
1
multi-nic
1
multicast
1
nccl
51
nccl-tests
2
network
2
notes
1
nsight
1
nsight-systems
1
nsys
1
numa
4
nvidia
1
nvlink
5
nvlink-c2c
1
nvls
1
nvsci
1
nvswitch
1
nvtx
2
observability
1
opengl
1
optimizer
1
ordering
1
overlap
1
p2p
2
pcie
2
performance
3
performance-model
1
persisting-cache
1
pipeline
2
pipeline parallel
1
planner
1
plugin
2
primitives
1
process-group
2
processgroup
1
processgroupnccl
1
production
1
protocol
3
proxy
1
pxn
1
pytorch
7
qos
1
rail
1
rank
1
rdma
2
reduce-scatter
3
reducer
1
regression
2
resource-partition
1
ring
2
roce
1
root-cause
1
runtime
1
scheduler
12
scheduling
1
sglang
13
sharding
1
sharp
1
shm
2
simple
1
slice
1
socket
2
startup
1
stas
1
statistics
1
stream
1
synchronization
4
tensor core
1
tensor parallel
3
timeout
2
tma
2
tmem
1
topology
5
torchrun
1
transport
3
tree
2
troubleshooting
3
tuning
2
turing
1
typora
1
unified-memory
1
verbs
1
virtual-memory
1
vllm
13
vmm
2
volta
1
vulkan
1
watchdog
3
wgmma
1
work-stealing
1
worknccl
1
xml
1
zero
2
Trending Tags
nccl
cuda
interview
distributed training
llm inference
gpu
sglang
vllm
ai infra
kv cache
Trending Tags
nccl
cuda
interview
distributed training
llm inference
gpu
sglang
vllm
ai infra
kv cache
×
A new version of content is available.
Update