rougail-ai/doc/WebLlm.md
2026-09-14 21:22:34 +02:00

3.6 KiB

WebLLM

liens

avec ses outils

webllmfit
webllmfit — can your machine run these models?
Hardware
  WebGPU.........: NOT available
  System RAM.....: unknown
  CPU threads....: 8
No WebGPU → no in-browser LLM can run in this browser.
Use a recent Chrome/Edge (≥ 113) or Safari 18+ (chrome://flags/#enable-unsafe-webgpu on Linux).


install webllm

Suivre l'install dans le readme.

sudo apt install -y nodejs npm

npm install @mlc-ai/web-llm

 python -m"http.server" 8000

ou bien, lancer le serveur nodejs :

npm install -g serve
serve

usage

La méthode principale : le System Prompt. Exemple de script:


import { CreateMLCEngine } from "@mlc-ai/web-llm";

// 1. Définir la persona (comme votre instruction SYSTEM dans Ollama)
const systemPrompt = "Tu es un assistant culinaire expert, spécialisé dans la cuisine française. Tu réponds toujours avec enthousiasme et proposes des alternatives aux ingrédients difficiles à trouver. Tu ne parles jamais de sujets hors de la cuisine.";

const messages = [
  // 2. Le message système définit le comportement pour toute la conversation
  { role: "system", content: systemPrompt },
  { role: "user", content: "Comment faire une bonne quiche lorraine ?" }
];

// 3. Envoyer la requête
const engine = await CreateMLCEngine("Llama-3.1-8B-Instruct-q4f32_1-MLC");
const reply = await engine.chat.completions.create({ messages });

console.log(reply.choices[0].message.content);

  • Aller dans le dossier webllm et lancer un serveur type python -m"http.server" dans ce dossier.
  • Utiliser un browser compatible (j'ai utilisé un-googled-chromium), le résultat du chat apparaît dans la console.

install the python mlc local chat utility

install

L'install doit se faire avec des version très précises sinon il y a des difficultés.

  • faire un environnement virtuel
  • installer de la manière suivante :
sudo apt-get install git-lfs

pip install --pre -f https://mlc.ai/wheels \
    mlc-ai-cpu==0.20.0 \
    mlc-llm-cpu==0.20.0.dev0


pip install psutil

# wget https://github.com/mlc-ai/package/releases/download/v0.9.dev0/mlc_llm_cpu-0.20.0.dev0-py3-none-manylinux_2_28_x86_64.whl
# pip install ./mlc_llm_cpu-0.20.0.dev0-py3-none-manylinux_2_28_x86_64.whl

pip install apache-tvm-ffi==0.1.11
# https://pypi.org/project/apache-tvm-ffi/#history

usage

python -m mlc_llm chat HF://mlc-ai/Llama-3-8B-Instruct-q4f16_1-MLC
You can use the following special commands:
  /help               print the special commands
  /exit               quit the cli
  /stats              print out stats of last request (token/sec)
  /metrics            print out full engine metrics
  /reset              restart a fresh chat
  /set [overrides]    override settings in the generation config. For example,
                      `/set temperature=0.5;top_p=0.8;seed=23;max_tokens=100;stop=str1,str2`
                      Note: Separate stop words in the `stop` option with commas (,).
  Multi-line input: Use escape+enter to start a new line.