Différences

Ci-dessous, les différences entre deux révisions de la page.

--- informatique:ai_lm [11/06/2026 09:23] – [ZML] cyrille
+++ informatique:ai_lm [23/06/2026 05:40] (Version actuelle) – [Compilation pour GPU] cyrille
@@ Ligne 265: / Ligne 265: @@
 # RTX 5060 : 120
-$ export CUDA_VERSION=12.9 && cmake -B build -DGGML_CUDA=ON \
+$ export CUDA_VERSION=12.9
+$ export CUDA_VERSION=13.3
+$ cmake -B build -DGGML_CUDA=ON \
  -DCMAKE_CUDA_ARCHITECTURES="86;120" \
  -DCMAKE_BUILD_WITH_INSTALL_RPATH=ON \
@@ Ligne 308: / Ligne 310: @@
 user	27m13.877s
 sys	1m24.687s
-</code>
-Avec CUDA 13.1 llama.cpp plante direct à la 1ère requête, mais sans message dans syslog : ce n'est donc pas le driver mais le logiciel llama.cpp qui ne support pas cette version de CUDA :
-<code>
-/home/cyrille/Code/bronx/AI_Coding/llama.cpp/ggml/src/ggml-cuda/ggml-cuda.cu:97: CUDA error
-CUDA error: invalid argument
-  current device: 0, in function ggml_cuda_mul_mat_q at /home/cyrille/Code/bronx/AI_Coding/llama.cpp/ggml/src/ggml-cuda/mmq.cu:179
 </code>
@@ Ligne 489: / Ligne 484: @@
 Headroom
+  * Compress tool outputs, logs, files, and RAG chunks before they reach the LLM. 60-95% fewer tokens, same answers. Library, proxy, MCP server.
+  * https://headroom-docs.vercel.app/docs
+  * https://github.com/chopratejas/headroom
   * https://www.lemondeinformatique.fr/actualites/lire-headroom-un-projet-open-source-pour-reduire-la-facture-des-tokens-100357.html
-  *
+RTK
+  * CLI proxy that reduces LLM token consumption by 60-90% on common dev commands. Single Rust binary, zero dependencies
+  * https://www.rtk-ai.app/
+  * https://github.com/rtk-ai/rtk
+Openwolf
+  * Sharper context. Fewer tokens. Open-source middleware for Claude Code.
+  * https://openwolf.com/
+  * https://github.com/cytostack/openwolf