Backend
When the Token Well Runs Dry: A Degradation State Machine for LLM Services
Riley Wang Dev.to (EN Zone)
1 views
The alert fires at 3:14 AM. Your LLM service returns 500s. The free quota is gone. You check the dashboard. 10,000,000 tokens. Zero remaining.
Most guides teach prevention. Budgets. Ledgers. Pre-checks. This one teaches survival. What happens after the quota dies? The answer determines if your users stay.
I built a degradation state machine. It turns quota exhaustion into a designed state. Not a crash. This article shows the full implementation.
Disclosure: This article was prepared as part of MonkeyCode's product outreach.
The four states
Free tiers end. That is a fact. Design for it.
State
Trigger
Behavior
NORMAL
Usage < 70%
Full service
WARNING
Usage >= 70%
Cache responses
DEGRADED
Usage >= 90%
Critical requests only
EXHAUSTED
Usage >= 100%
Cache or 503
The thresholds are config. Tune them for your traffic. The state machine is the core.
Step 1: Track the quota
MonkeyCode's free tier includes 10M tokens and a free server option. Quotas change. Verify the current numbers first.
The tracker polls the usage endpoint. It stores the remaining count in memory.
class QuotaTracker:
def __init__(self, cap: int):
self.cap = cap
self.used = 0
def record(self, tokens: int):
self.used += tokens
def ratio(self) -> float:
return self.used / self.cap
def state(self) -> str:
r = self.ratio()
if r >= 1.0:
return "EXHAUSTED"
if r >= 0.9:
return "DEGRADED"
if r >= 0.7:
return "WARNING"
return "NORMAL"
Simple arithmetic. Four states. No magic.
Step 2: Add the cache
The WARNING state needs a cache. Every response gets stored. Keyed by the prompt hash.
import hashlib
class ResponseCache:
def __init__(self):
self.store = {}
def key(self, prompt: str) -> str:
return hashlib.sha256(prompt.encode()).hexdigest()
def get(self, prompt: str):
return self.store.get(self.key(prompt))
def put(self, prompt: str, response: str):
self.store[self.key(prompt)] = response
The cache is a dictionary. For production, use Redis. For a free server, a dict is fine.
Step 3: Classify requests
Not all requests are equal. Some deserve the last tokens. Others can wait.
def classify_request(payload: dict) -> str:
if payload.get("priority") == "critical":
return "critical"
if payload.get("task") == "classify":
return "cheap"
return "normal"
Critical requests keep the service alive. Cheap requests can use cached answers. Normal requests wait.
Step 4: The router with degradation
This is the heart. The router checks the state before every call.
def route(payload: dict, tracker: QuotaTracker, cache: ResponseCache):
state = tracker.state()
task = classify_request(payload)
if state == "EXHAUSTED":
cached = cache.get(payload["text"])
if cached:
return {"source": "cache", "label": cached}
return {"source": "none", "error": "quota exhausted"}, 503
if state == "DEGRADED" and task != "critical":
cached = cache.get(payload["text"])
if cached:
return {"source": "cache", "label": cached}
return {"source": "none", "error": "degraded"}, 503
if state == "WARNING":
cached = cache.get(payload["text"])
if cached:
return {"source": "cache", "label": cached}
# live call to the model
result = call_model(payload["text"])
cache.put(payload["text"], result)
tracker.record(estimate_tokens(payload["text"], result))
return {"source": "live", "label": result}
Read the logic top to bottom. Each state adds a restriction. The service never crashes. It degrades.
Step 5: Test the degradation
A state machine needs a test. Simulate quota exhaustion. Verify each state.
def test_degradation():
tracker = QuotaTracker(cap=1000)
cache = ResponseCache()
# fill the cache
cache.put("hello", "greeting")
# NORMAL
assert tracker.state() == "NORMAL"
# WARNING
tracker.used = 700
assert tracker.state() == "WARNING"
# DEGRADED
tracker.used = 900
assert tracker.state() == "DEGRADED"
# EXHAUSTED
tracker.used = 1000
assert tracker.state() == "EXHAUSTED"
# EXHAUSTED serves from cache
result = route({"text": "hello"}, tracker, cache)
assert result["source"] == "cache"
# EXHAUSTED rejects uncached
result = route({"text": "unknown"}, tracker, cache)
assert result[1] == 503
print("All degradation tests passed")
Run it.
python test_degradation.py
Step 6: Deploy and monitor
The free server option hosts this. No credit card. No billing alarm.
git clone <your-repo>
cd <your-repo>
pip install -r requirements.txt
uvicorn app:app --host 0.0.0.0 --port 8000
Add a health endpoint that reports the current state.
curl http://localhost:8000/state
# {"state": "NORMAL", "ratio": 0.42}
A simple curl tells you the truth. No dashboard needed.
What the cache cannot save
Caching has limits. Unique prompts miss. Long conversations miss. Time-sensitive answers miss.
The cache hit rate determines your survival time. Measure it.
hits = sum(1 for r in results if r["source"] == "cache")
print(f"Hit rate: {hits / len(results):.0%}")
A 40% hit rate extends the service by days. A 5% hit rate barely helps.
Who should not use this
Teams with SLAs need paid capacity. Teams with real-time requirements need dedicated endpoints. Teams with high concurrency need more than a free server.
This pattern is for prototypes. For internal tools. For services where a 503 is acceptable. For developers who want a service that fails gracefully.
The takeaway
Quota exhaustion is a state. Design for it. The state machine turns a crash into a graceful degradation. Cache what you can. Protect what matters. Fail loudly when you must.
If you want to experiment with this pattern, MonkeyCode's free tier is a place to start. Check the current quota first. Then build something that survives the month.
MonkeyCode provides free models that can run this workflow.
Read original: https://dev.to/apprs_6334/when-the-token-well-runs-dry-a-degradation-state-machine-for-llm-services-1fhm
← Previous
Real-Time Features Need to Match Your Infrastructure, Not Just Your Feature Spec
Next →
Invoice2Excel Pro – A local, privacy-focused GUI tool built with Python & CustomTkinter to extract PDF invoices to Excel
Related
I built a message board for AI agents. Four bugs shipped past a 100% green test suite.
I built agora, a public message board that any AI agent can read and post to. No account, no signup, no API key — only body is required. Agents leave notes, "here's what broke and here's what fixed it" field notes, or questions other agents can reply to. It's about 1,800 lines of PHP with a SQLite
DEV Community
Principles of Memory Management in Java
submitted by /u/davidalayachew [link] [留言]
Reddit r/programming
The Best Tools for Automating Marketing Graphics in 2026
Bannerbear, Picnie, Placid, Creatomate, BannerBoo, and what to consider before choosing a creative automation platform. Marketing teams rarely have a problem coming up with new campaigns. The problem usually begins after a campaign is approved. One design may need to become a dozen social posts. A
Dev.to (EN Zone)