LLM Systems in Production

Cloud-Native Patterns for AI Engineers

First Edition
Nearby Libraries

Only you can see this

Save Note
Last edited by Nabeel Khan
August 26, 2026 | History

LLM Systems in Production

Cloud-Native Patterns for AI Engineers

First Edition

Volume one of The Full-Stack AI Engineering Series, covering the infrastructure layer of a production large language model platform. Twelve chapters develop NexusCore, a fictional routing and observability gateway operated by a fictional regulated fintech called Nebula Financial, and treat inference and serving foundations, multi-cloud and hybrid topology, learned routing policies expressed as versioned artifacts, edge and on-device escalation, speculative decoding as a latency primitive, distributed and pipelined inference, observability, router lifecycle security, and compliance as evidence the system produces by design. A case study reconstructs a trading-desk outage and the router rollback that followed, and a closing chapter assembles the full reference architecture. Specifications throughout are conceptual, presenting schemas, policy objects, trace shapes and architecture diagrams rather than production code. Addressed to site reliability engineers, cloud architects and platform engineers who have been given responsibility for AI workloads.

Publish Date
Publisher
iSystematic Inc.
Language
English
Pages
260
Edition Availability
Cover of: LLM Systems in Production
LLM Systems in Production: Cloud-Native Patterns for AI Engineers
2026, iSystematic Inc.
in English - First Edition

Add another edition?

Book Details


First Sentence

"A gateway is not plumbing. It is the place where an institution decides what it is allowed to think, and how much that thought may cost."

Table of Contents

The Full-Stack AI Engineering Series (Series Overview)
A Note Where a Foreword Should Be
Preface
Acknowledgments
Introduction: How to Read This Book
Chapter 1. Why Fintech Needs an LLM Gateway
Chapter 2. Foundations of LLM Inference and Serving
Chapter 3. Multi-Cloud and Hybrid Topology for LLMs
Chapter 4. Designing the Routing Brain: From Rules to Learned Policies
Chapter 5. Edge and On-Device Routing: Confident or Seek Stronger
Chapter 6. Speculative Decoding as an SRE Primitive
Chapter 7. Distributed and Pipelined Inference in Practice
Chapter 8. Observability for LLM Routers and Models
Chapter 9. Router Lifecycle Security and Governance
Chapter 10. Compliance, Audit Trails, and AI Conformance
Chapter 11. Case Study: A Trading-Desk Outage and a Router Rollback
Chapter 12. The NexusCore Blueprint
Appendix A. NexusCore Artifact Catalog
Appendix B. Quick Reference Cards
Appendix C. Evaluation Metrics Reference
Appendix D. Tooling Ecosystem
Appendix E. Compliance Checklists
Appendix F. Troubleshooting Guide
Glossary
Bibliography
Index
Colophon

Edition Notes

Artificial intelligence

Machine learning

Natural language processing (Computer science)

Cloud computing

Software architecture

Computer network architectures

Systems engineering

Reliability (Engineering)

Electronic data processing--Distributed processing

Information technology--Management
Large language models

Site reliability engineering

Model serving

Inference optimization

MLOps

Published in
Canada
Series
The Full-Stack AI Engineering Series, volume 1
Copyright Date
2026

Edition Identifiers

Open Library
OL62487370M
ISBN 10
1067831711
ISBN 13
9781067831714
Amazon ID (ASIN)
1067831711

Work Identifiers

Work ID
OL45928739W

Community Reviews (0)

No community reviews have been submitted for this work.

Lists

Download catalog record: RDF / JSON / OPDS |

Wikipedia citation

Copy and paste this code into your Wikipedia page.

{{cite book|author=Nabeel Khan |date=2026 |title=LLM Systems in Production |publication-place=Canada |publisher=iSystematic Inc. |isbn=978-1-06-783171-4 |ol=62487370M}}