Skip to content

GitLab

  • Menu
Projects Groups Snippets
    • Loading...
  • Help
    • Help
    • Support
    • Community forum
    • Submit feedback
    • Contribute to GitLab
  • Sign in
  • G GitLabMCP
  • Project information
    • Project information
    • Activity
    • Labels
    • Planning hierarchy
    • Members
  • Repository
    • Repository
    • Files
    • Commits
    • Branches
    • Tags
    • Contributors
    • Graph
    • Compare
  • Issues 0
    • Issues 0
    • List
    • Boards
    • Service Desk
    • Milestones
  • Merge requests 0
    • Merge requests 0
  • CI/CD
    • CI/CD
    • Pipelines
    • Jobs
    • Schedules
  • Deployments
    • Deployments
    • Environments
    • Releases
  • Monitor
    • Monitor
    • Metrics
    • Incidents
  • Packages & Registries
    • Packages & Registries
    • Package Registry
    • Container Registry
    • Infrastructure Registry
  • Analytics
    • Analytics
    • Value stream
    • CI/CD
    • Repository
  • Wiki
    • Wiki
  • Snippets
    • Snippets
  • Activity
  • Graph
  • Create a new issue
  • Jobs
  • Commits
  • Issue Boards
Collapse sidebar
  • Satyam Raj
  • GitLabMCP
  • Merge requests
  • !1

Merged
Created Aug 12, 2026 by Satyam Raj@satyamrajMaintainer

Make KB answer context size configurable

  • Overview 0
  • Commits 6
  • Pipelines 2
  • Changes 66

buildContext hardcoded a 2500-token cap on the source block sent to the LLM. On CPU inference the model reads every one of those tokens before emitting its first, which makes this the main latency knob after CHAT_MAX_TOKENS.

Expose it as CONTEXT_TOKENS (default 1800) and wire it through docker-compose.yml and .env.example.

Co-Authored-By: Claude Opus 5 noreply@anthropic.com

Assignee
Assign to
Reviewer
Request review from
Time tracking
Source branch: feature/knowledgeBase