Ich optimiere dein LLM-App für niedrigere Latenz und Kosten

C
chlee9
C
chlee9
Cheolhee Lee
Einige Informationen wurden automatisch übersetzt.

Über diesen Service

Automatische Übersetzung

Ich optimiere LLM-Anwendungen, die zu langsam oder zu teuer sind.


Was ich in der Produktion erreicht habe:

- P95-Latenz um 70 Prozent gesenkt und die Servicekosten um 38 Prozent mit Kontext-Caching, strukturiertem Output und Modell-Routing

- Output-Tokens um 49 Prozent reduziert, ohne die Qualität zu verlieren

- Listen-Endpunkte um das 125-fache beschleunigt und eine Threads-Anfrage von 411 ms auf 1,6 ms verkürzt

- Drei LLM- und Embedding-Modelle auf ein On-Premise-DGX umgezogen, ohne Ausfallzeiten


So funktioniert es:

1. Du teilst deine Prompts, Traces, Model-Konfiguration und deine Latenz- oder Kostenzahlen

2. Ich erstelle ein Profil der Pipeline und sende einen priorisierten Bericht

3. Ich setze die Optimierungen um und zeige Vorher-Nachher-Benchmarks


Stack: OpenAI, Anthropic, vLLM, LangChain, RAG, pgvector, Python, TypeScript, Rust, Go, AWS.


Gib mir deine aktuellen p95 und monatlichen Ausgaben, und ich sage dir, was realistisch ist.

Lerne Cheolhee Lee kennen

Cheolhee Lee

AI Full Stack Developer specializing in LLM and RAG optimization

  • AusSüdkorea
  • Mitglied seitApr. 2021
  • ⌀ Antwortzeit1 Stunde
  • Sprachen

    Koreanisch, Englisch
I keep my employer and my clients unnamed here. I ship production AI systems end to end at an undisclosed B2B AI SaaS company - a sales-automation SaaS and a public-sector AI evaluation platform. Measured: LLM p95 latency -70%, serving cost -38%, output tokens -49% via context caching and structured output. 125x list speedup, threads query 411ms to 1.6ms, bundle 21.7MB to 2.3MB. Re-homed three LLM models to an on-prem DGX with zero downtime; passed TTA review for Korea's AI Verification program. TypeScript, Python, Rust, Go, React, PostgreSQL, AWS, RAG, vLLM, MCP.

Automatische Übersetzung

Mein Portfolio

Meine weiteren Dienstleistungen im Bereich KI-Entwicklung