Skip to content
AImpact
IT EN
Models Beginner Also known as: Large Language Model · Modello linguistico di grandi dimensioni

LLM

/el-el-em/

An AI model trained on huge amounts of text to predict the next word and generate natural language responses.

ShareLinkedInX

In practice

It is the engine behind ChatGPT, Claude, Gemini. When you embed an LLM into your product you pay per token and get a service that reads and writes text. Quality depends heavily on the chosen model and the prompt you give it.

Related terms

Seen in the wild

62 entries mentioning it
  1. Mistral releases Devstral Small: 7B coding model for agentic tasks on consumer GPU
    Medium
  2. Figure AI releases Figure 02 autonomy update: fully autonomous warehouse picking without human teleoperation
    High
  3. Ollama 0.9: concurrent model serving, multi-GPU split, and REST API v2 for local AI
    Medium
  4. Local AI 2025: Ollama, MLX LM, Apple Foundation Models triple the speed
    Medium
  5. Private LLM: models up to 7B directly on iPhone and Mac, fully offline
    Medium
  6. vLLM v0.7: chunked prefill by default and a redesigned V1 engine
    Medium
  7. NVIDIA NIM 1.0: Containerized LLM Inference with OpenAI-Compatible API
    High
  8. WebLLM and LLM in WASM: browser-based LLM inference via WebGPU, no server needed
    Medium
  9. Continuous Batching for LLM Serving: survey and state of the art of Orca, vLLM, SGLang, TGI
    Medium
  10. DeepMind: 60+ cases of Specification Gaming in LLMs documented
    High
  11. FlashInfer 0.2: attention library for LLM serving with paged KV cache and RoPE fusion
    Medium
  12. Prefill/decode disaggregation: separate GPUs for low TTFT and high throughput
    High
  13. KV Cache Quantization FP8/INT8: Double User Density per GPU
    High
  14. AnythingLLM 1.0: the complete local RAG stack for enterprise use
    High
  15. LLM Compressor: unified toolkit for quantization and sparsity with native vLLM integration
    Medium
  16. CyberSecEval 2: Meta's LLM cybersecurity benchmark
    Medium
  17. Dify 0.7: visual agentic workflows with integrated RAG and 10+ LLMs
    Medium
  18. DrEureka: LLM automates simulation-to-real transfer without manual tuning
    Medium
  19. NeMo Guardrails 0.8: NVIDIA's framework for adding safety rails to any LLM
    Medium
  20. Microsoft RoboGen: generating robot tasks, skills and environments from text
    Medium
← All terms