home
  • Notes
  • Software
  • Talks
  • Guestbook

Chai

I’m a fan of having chai every weekend I’m at home. I took some recipes online and after lots of trial and error got to a recipe that I really like.
Me

search Search Notes (⌘K)

Related Articles

  • Optimizing the LLM Inference Stack
    An end-to-end framework for optimizing LLM inference: map hardware, kernels, host execution, runtimes, and orchestration, define service demand, rank bottlenecks with capacity and queueing, and choose experiments that improve goodput. Includes case studies from Makora and Baseten.
  • The Pareto Frontier in LLM Inference Serving
    Why a single performance number tells only part of the story in LLM serving, and when you need the whole Pareto frontier instead. Covers the limits of scalarization on concave trade-offs, how Optuna and Google Vizier actually store trial values (the frontier is derived, not persisted), and how production tools like Prism, InferenceBench, and NVIDIA Dynamo present the frontier as the reportable artifact.
  • Benchmarking & Capacity Planning for Systems Engineers
    A systems engineering guide to load testing and capacity planning: applying the problem statement method, characterizing production workloads, eliminating Coordinated Omission, choosing between open and closed loops, and sizing infrastructure with Little’s Law.
  • Queuing Theory for Systems Engineers
    A practical, example-first guide to queuing theory in systems engineering: single-server M/M/1 response time and the 50% load paradox, the hockey stick latency curve, Pollaczek-Khinchine service variance, multi-server resource pooling (M/M/c), and capacity planning rules of thumb with interactive visualizers.
  • Performance Fundamentals
    Core concepts in systems performance engineering: formalizing latency, throughput, and resource utilization, Little’s Law, and analyzing multi-worker queue dynamics and diurnal traffic cycles with interactive simulators.

Recent Notes

  • Optimizing the LLM Inference Stack
    An end-to-end framework for optimizing LLM inference: map hardware, kernels, host execution, runtimes, and orchestration, define service demand, rank bottlenecks with capacity and queueing, and choose experiments that improve goodput. Includes case studies from Makora and Baseten.
  • The Pareto Frontier in LLM Inference Serving
    Why a single performance number tells only part of the story in LLM serving, and when you need the whole Pareto frontier instead. Covers the limits of scalarization on concave trade-offs, how Optuna and Google Vizier actually store trial values (the frontier is derived, not persisted), and how production tools like Prism, InferenceBench, and NVIDIA Dynamo present the frontier as the reportable artifact.
  • Benchmarking & Capacity Planning for Systems Engineers
    A systems engineering guide to load testing and capacity planning: applying the problem statement method, characterizing production workloads, eliminating Coordinated Omission, choosing between open and closed loops, and sizing infrastructure with Little’s Law.
  • Queuing Theory for Systems Engineers
    A practical, example-first guide to queuing theory in systems engineering: single-server M/M/1 response time and the 50% load paradox, the hockey stick latency curve, Pollaczek-Khinchine service variance, multi-server resource pooling (M/M/c), and capacity planning rules of thumb with interactive visualizers.
  • Performance Fundamentals
    Core concepts in systems performance engineering: formalizing latency, throughput, and resource utilization, Little’s Law, and analyzing multi-worker queue dynamics and diurnal traffic cycles with interactive simulators.
Browse all notes →
Plane by antonmoek - CC BY 4.0 Fullscreen experience Thank you for your visit, please sign the guestbook.