@meta
  v: 1
  route: /run-locally/
  generated: 2026-09-29T01:38:42Z
  ttl: 1d

@intent
  purpose: Get Jebadiah (Jeb) running on your own machine and answering a first decision, then wire it into code
  audience: developer, ai-agent, coding-assistant
  capability: install, run, call, integrate

@state
  install: pip install jebadiah-decide
  requires: Python 3.10 or newer; installs tokenizers, jinja2, huggingface-hub; no torch, no transformers
  server: jeb serve on http://localhost:8100
  time_to_first_decision: 2 minutes 1 second on the Ollama path from a clean setup, most of it the 9.8 GB download
  default_model: hf.co/frontier-infra/jebadiah-9b-v2-GGUF:Q8_0 (9.8 GB); jeb serve --size 4b for 4.6 GB
  endpoints[3]{method,path,shape}:
    POST,/v1/systemone,"Jev wire (preferred): state plus questions of type choice, noul or score"
    POST,/v1/decide,AINode decide shape: question plus options
    GET,/health,200 when the server is up
  runtimes[6]{runtime,command,options_per_question,note}:
    Ollama,jeb serve,20,pulls the 9B Q8_0 on first run; Ollama 0.34 or newer
    LM Studio,jeb serve --backend lmstudio,20,load jebadiah-9b-v2 Q8_0 and start the server; LM Studio 0.4 or newer
    vLLM,jeb serve --backend vllm --url http://localhost:8000,20 by default,vllm serve frontier-infra/jebadiah-9b-v2 --max-model-len 4096 --language-model-only
    llama.cpp,jeb serve --backend llama-server,no limit,llama-server -m jebadiah-9b-v2-Q8_0.gguf -c 4096 -np 1 --port 8080; v0.5.0 or newer
    MLX on a Mac,jeb serve --backend mlx,no limit,install jebadiah-decide[mlx]; downloads the 8-bit 9B
    AINode,POST http://<node>:3000/api/models/load then /v1/systemone on the node,20,check with jeb doctor --backend ainode
  first_decision_request: curl localhost:8100/v1/systemone, route an invoice ticket to billing, support or sales
  first_decision_answer: billing 0.641142, support 0.036095, sales 0.322763 (Ollama; llama-server gives the same)
  scenarios[5]{scenario,question_type,answer}:
    Route a ticket or message,choice,billing 0.986
    Phishing or spam,choice yes/no/unknown,yes 0.939
    Does this reply answer the question,noul and score,answers_it 0.067; quality 0.40 of 2
    Gate an agent's tool call with JDE,noul plus JDE bands,rename act 0.964; delete escalate 0.021; flight confirm 0.680
    Content policy,noul per rule,threat 0.868; spam 0.030
  thresholds: act at 0.9 or more, confirm at 0.6 or more, escalate below
  integrate_languages[5]: curl, Python (requests), Python (jebadiah-decide package), TypeScript / Node, Go
  jde: JDE's default judge is a local Jeb at http://localhost:8100/v1/systemone, where jeb serve listens; ask({ decision, state, questions }) needs no key and no config
  jde_settings: JDE_JEB_ENDPOINT and JDE_JEB_MODEL move it (AINode: http://<node>:3000/v1/systemone with JDE_JEB_API_KEY); no silent fallback to a hosted service; TypeSafe's hosted Jev is opt in with judge jev and TYPESAFE_API_KEY
  judge_jeb: planned, not released; JDE moves to it with one JDE_JEB_MODEL setting
  jeb_commands[4]{command,does}:
    jeb serve,/v1/systemone and /v1/decide on localhost
    jeb doctor,"checks the runtime, model, tokenizer, prompt and a known answer"
    jeb ask,one answer from a question and a text file
    jeb request,"answers a {state, questions} file"
  models[3]{model,gguf_q8,same_pick_as_bf16_of_260}:
    Jebadiah 27B,29 GB,260
    Jebadiah 9B v2 (default),9.8 GB,257
    Jebadiah 4B v2,4.6 GB,256
  gotchas[6]: never use a runtime chat window or /api/chat, the prompt must be exact; jeb checks token counts, "20 options on Ollama, LM Studio, AINode and default vLLM", option order matters, stay on Q8_0 or 8-bit, "state over 2,048 prompt tokens is cut from the end with a warning"

@actions
  - id: view_human
    method: GET
    href: https://jebadiah.ai/run-locally/
    description: "The walkthrough page, with a runtime picker and copy buttons"
  - id: read_guide_markdown
    method: GET
    href: https://github.com/getainode/jebadiah/blob/main/docs/run-locally.md
    description: The same guide as markdown
  - id: read_llms_full
    method: GET
    href: https://jebadiah.ai/llms-full.txt
    description: The full guide as one text file
  - id: get_client
    method: GET
    href: https://github.com/getainode/jebadiah/tree/main/clients/python
    description: The jebadiah-decide client: the jeb CLI and Python API
  - id: get_ollama
    method: GET
    href: https://ollama.com/download
    description: Ollama download
  - id: get_lm_studio
    method: GET
    href: https://lmstudio.ai
    description: LM Studio
  - id: get_gguf_9b_v2
    method: GET
    href: https://huggingface.co/frontier-infra/jebadiah-9b-v2-GGUF
    description: "9B v2 GGUF, the default for local use"
  - id: get_jde
    method: GET
    href: https://github.com/Titanium-Devops/jde
    description: "JDE, the Jev Decision Engine"
  - id: get_ainode
    method: GET
    href: https://github.com/getainode/ainode
    description: AINode

@context
  > Jeb is an open decision model. Ask it a typed question about some text or JSON and it answers with a calibrated probability for every option. It never writes text.
  > Which runtime are you using? Every path ends the same way: jeb serve on http://localhost:8100, which JDE, your own code or a coding agent can call.
  > The page includes a copyable prompt that sets Jeb up in a project through a coding agent.

@nav
  self: /run-locally.agent
  parents: [/.agent]
  peers: [/.agent, /404.agent]
