{"componentChunkName":"component---src-templates-tags-js","path":"/tags/speculative-decoding/","result":{"data":{"site":{"siteMetadata":{"title":"Bottlehs Tech Blog"}},"allMarkdownRemark":{"totalCount":1,"nodes":[{"excerpt":"LLM 추론 최적화 - KV Cache, Quantization, Speculative Decoding LLM 추론은 계산 집약적이고 메모리를 많이 사용한다. 프로덕션 환경에서 LLM…","fields":{"slug":"/ai/LLM-추론-최적화-KV-Cache-Quantization-Speculative-Decoding/"},"frontmatter":{"date":"2025년 11월 28일","title":"LLM 추론 최적화 - KV Cache, Quantization, Speculative Decoding","description":"LLM 추론 성능을 최적화하는 핵심 기술들을 실전 예제와 함께 정리합니다. KV Cache, Quantization, Speculative Decoding, 배치 처리, 모델 병렬화 등 모든 최적화 기법을 다룹니다."}}]}},"pageContext":{"tag":"Speculative Decoding"}},"staticQueryHashes":["213619243","3262363727"]}