<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0"><channel><title><![CDATA[IA]]></title><description><![CDATA[IA]]></description><link>https://lemmy.cyber-neurones.org/category/32</link><generator>RSS for Node</generator><lastBuildDate>Mon, 21 Sep 2026 07:06:51 GMT</lastBuildDate><atom:link href="https://lemmy.cyber-neurones.org/category/32.rss" rel="self" type="application/rss+xml"/><pubDate>Thu, 03 Sep 2026 16:20:24 GMT</pubDate><ttl>60</ttl><item><title><![CDATA[Crash LLama.cpp]]></title><link>https://lemmy.cyber-neurones.org/topic/398/crash-llama.cpp</link><guid isPermaLink="true">https://lemmy.cyber-neurones.org/topic/398/crash-llama.cpp</guid><pubDate>Thu, 03 Sep 2026 16:20:24 GMT</pubDate></item><item><title><![CDATA[Installation  Driver 610.57.04]]></title><description><![CDATA[Pas de crash :
# nvidia-smi 
Tue Sep  1 17:35:06 2026       
+-----------------------------------------------------------------------------------------+
| NVIDIA-SMI 610.57.04              KMD Version: 610.57.04     CUDA UMD Version: 13.3     |
+-----------------------------------------+------------------------+----------------------+
| GPU  Name                 Persistence-M | Bus-Id          Disp.A | Volatile Uncorr. ECC |
| Fan  Temp   Perf          Pwr:Usage/Cap |           Memory-Usage | GPU-Util  Compute M. |
|                                         |                        |               MIG M. |
|=========================================+========================+======================|
|   0  NVIDIA GeForce RTX 3060 ...    Off |   00000000:01:00.0 Off |                  N/A |
| N/A   63C    P0             36W /  115W |    2096MiB /   6144MiB |     12%      Default |
|                                         |                        |                  N/A |
+-----------------------------------------+------------------------+----------------------+
|   1  NVIDIA GeForce RTX 5060 Ti     Off |   00000000:05:00.0 Off |                  N/A |
| 30%   51C    P1             49W /  180W |    7894MiB /  16311MiB |     32%      Default |
|                                         |                        |                  N/A |
+-----------------------------------------+------------------------+----------------------+

+-----------------------------------------------------------------------------------------+
| Processes:                                                                              |
|  GPU   GI   CI              PID   Type   Process name                        GPU Memory |
|        ID   ID                                                               Usage      |
|=========================================================================================|
|    0   N/A  N/A            3587      C   /usr/local/bin/llama-server            2088MiB |
|    1   N/A  N/A            3587      C   /usr/local/bin/llama-server            7886MiB |
+-----------------------------------------------------------------------------------------+


]]></description><link>https://lemmy.cyber-neurones.org/topic/397/installation-driver-610.57.04</link><guid isPermaLink="true">https://lemmy.cyber-neurones.org/topic/397/installation-driver-610.57.04</guid><dc:creator><![CDATA[Tuxedo17]]></dc:creator><pubDate>Tue, 01 Sep 2026 14:48:41 GMT</pubDate></item><item><title><![CDATA[Crash Tuxedo :]]></title><description><![CDATA[C’est mieux :
[image: 7f991335-22e6-424f-9fe4-4495751d1e66-image.jpeg]
]]></description><link>https://lemmy.cyber-neurones.org/topic/396/crash-tuxedo</link><guid isPermaLink="true">https://lemmy.cyber-neurones.org/topic/396/crash-tuxedo</guid><dc:creator><![CDATA[Tuxedo17]]></dc:creator><pubDate>Tue, 01 Sep 2026 12:52:31 GMT</pubDate></item><item><title><![CDATA[llama-swap & demeter-sante.fr : Calcul du coût.]]></title><link>https://lemmy.cyber-neurones.org/topic/394/llama-swap-demeter-sante.fr-calcul-du-coût.</link><guid isPermaLink="true">https://lemmy.cyber-neurones.org/topic/394/llama-swap-demeter-sante.fr-calcul-du-coût.</guid><pubDate>Thu, 13 Aug 2026 16:29:19 GMT</pubDate></item><item><title><![CDATA[llama.cpp : Impact de  "-fa on"  sur le Prompt Speed]]></title><description><![CDATA[Je viens de refaire un graphique :
[image: llama-swap-prompt-speed-daily.png]
]]></description><link>https://lemmy.cyber-neurones.org/topic/392/llama.cpp-impact-de-fa-on-sur-le-prompt-speed</link><guid isPermaLink="true">https://lemmy.cyber-neurones.org/topic/392/llama.cpp-impact-de-fa-on-sur-le-prompt-speed</guid><dc:creator><![CDATA[Tuxedo17]]></dc:creator><pubDate>Fri, 07 Aug 2026 13:53:13 GMT</pubDate></item><item><title><![CDATA[Tuxedo 17 : Installation vllm]]></title><description><![CDATA[Misère …
# /root/vllm-env/bin/vllm serve nvidia/Qwen3.6-35B-A3B-NVFP4 --port 8002 --gpu-memory-utilization 0.95 --max-model-len 180000 --max-num-seqs 1 --kv-cache-dtype fp8 --enable-chunked-prefill --max-num-batched-tokens 4096 --attention-backend flashinfer --reasoning-parser qwen3

WARNING 08-07 13:29:38 [cuda.py:959] Detected different devices in the system: NVIDIA GeForce RTX 3060 Laptop GPU, NVIDIA GeForce RTX 5060 Ti. Please make sure to set `CUDA_DEVICE_ORDER=PCI_BUS_ID` to avoid unexpected behavior.
(APIServer pid=1833126) INFO 08-07 13:29:49 [api_utils.py:345] 
(APIServer pid=1833126) INFO 08-07 13:29:49 [api_utils.py:345]        █     █     █▄   ▄█
(APIServer pid=1833126) INFO 08-07 13:29:49 [api_utils.py:345]  ▄▄ ▄█ █     █     █ ▀▄▀ █  version 0.26.0
(APIServer pid=1833126) INFO 08-07 13:29:49 [api_utils.py:345]   █▄█▀ █     █     █     █  model   nvidia/Qwen3.6-35B-A3B-NVFP4
(APIServer pid=1833126) INFO 08-07 13:29:49 [api_utils.py:345]    ▀▀  ▀▀▀▀▀ ▀▀▀▀▀ ▀     ▀
(APIServer pid=1833126) INFO 08-07 13:29:49 [api_utils.py:345] 
(APIServer pid=1833126) INFO 08-07 13:29:49 [api_utils.py:273] non-default args: {'model_tag': 'nvidia/Qwen3.6-35B-A3B-NVFP4', 'port': 8002, 'model': 'nvidia/Qwen3.6-35B-A3B-NVFP4', 'max_model_len': 180000, 'attention_backend': 'flashinfer', 'reasoning_parser': 'qwen3', 'gpu_memory_utilization': 0.95, 'kv_cache_dtype': 'fp8', 'max_num_batched_tokens': 4096, 'max_num_seqs': 1, 'enable_chunked_prefill': True}
(APIServer pid=1833126) Warning: You are sending unauthenticated requests to the HF Hub. Please set a HF_TOKEN to enable higher rate limits and faster downloads.
config.json: 100%|████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 58.1k/58.1k [00:00&lt;00:00, 93.7MB/s]
preprocessor_config.json: 100%|███████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 390/390 [00:00&lt;00:00, 2.29MB/s]
(APIServer pid=1833126) INFO 08-07 13:30:04 [model.py:623] Resolved architecture: Qwen3_5MoeForConditionalGeneration
(APIServer pid=1833126) INFO 08-07 13:30:04 [model.py:1788] Using max model len 180000
(APIServer pid=1833126) INFO 08-07 13:30:05 [cache.py:285] Using fp8 data type to store kv cache. It reduces the GPU memory footprint and boosts the performance. Meanwhile, it may cause accuracy drop without a proper scaling factor
(APIServer pid=1833126) INFO 08-07 13:30:05 [scheduler.py:252] Chunked prefill is enabled with max_num_batched_tokens=4096.
(APIServer pid=1833126) WARNING 08-07 13:30:05 [modelopt.py:385] Detected ModelOpt fp8 checkpoint (quant_algo=FP8). Please note that the format is experimental and could change.
(APIServer pid=1833126) WARNING 08-07 13:30:05 [modelopt.py:1034] Detected ModelOpt NVFP4 checkpoint (quant_algo=NVFP4). Please note that the format is experimental and could change in future.
(APIServer pid=1833126) WARNING 08-07 13:30:05 [modelopt.py:1034] Detected ModelOpt NVFP4 checkpoint (quant_algo=W4A16_NVFP4). Please note that the format is experimental and could change in future.
(APIServer pid=1833126) WARNING 08-07 13:30:05 [modelopt.py:1707] Detected ModelOpt MXFP8 checkpoint. Please note that the format is experimental and could change in future.
(APIServer pid=1833126) INFO 08-07 13:30:05 [vllm.py:1109] Asynchronous scheduling is enabled.
(APIServer pid=1833126) INFO 08-07 13:30:05 [kernel.py:295] Final IR op priority after setting platform defaults: IrOpPriorityConfig(rms_norm=['native'], fused_add_rms_norm=['native'])
tokenizer_config.json: 100%|██████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 16.7k/16.7k [00:00&lt;00:00, 26.1MB/s]
vocab.json: 100%|█████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 6.72M/6.72M [00:00&lt;00:00, 32.2MB/s]
tokenizer.json: downloading bytes: ████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 5.67MB,  532kB/s  
tokenizer.json: reconstructing file: 100%|████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 12.8MB / 12.8MB, 1.22MB/s  
chat_template.jinja: 100%|████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 7.76k/7.76k [00:00&lt;00:00, 23.5MB/s]
generation_config.json: 100%|██████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 202/202 [00:00&lt;00:00, 998kB/s]
video_preprocessor_config.json: 100%|█████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 385/385 [00:00&lt;00:00, 1.41MB/s]
(APIServer pid=1833126) [transformers] The `use_fast` parameter is deprecated and will be removed in a future version. Use `backend="torchvision"` instead of `use_fast=True`, or `backend="pil"` instead of `use_fast=False`.
WARNING 08-07 13:30:34 [cuda.py:959] Detected different devices in the system: NVIDIA GeForce RTX 3060 Laptop GPU, NVIDIA GeForce RTX 5060 Ti. Please make sure to set `CUDA_DEVICE_ORDER=PCI_BUS_ID` to avoid unexpected behavior.
(EngineCore pid=1833533) INFO 08-07 13:30:41 [core.py:116] Initializing a V1 LLM engine (v0.26.0) with config: model='nvidia/Qwen3.6-35B-A3B-NVFP4', speculative_config=None, tokenizer='nvidia/Qwen3.6-35B-A3B-NVFP4', skip_tokenizer_init=False, tokenizer_mode=auto, revision=None, tokenizer_revision=None, trust_remote_code=False, dtype=torch.bfloat16, max_seq_len=180000, download_dir=None, load_format=auto, tensor_parallel_size=1, pipeline_parallel_size=1, data_parallel_size=1, decode_context_parallel_size=1, dcp_comm_backend=ag_rs, disable_custom_all_reduce=False, quantization=modelopt_mixed, quantization_config=None, enforce_eager=False, enable_return_routed_experts=False, kv_cache_dtype=fp8, device_config=cuda, structured_outputs_config=StructuredOutputsConfig(backend='auto', disable_any_whitespace=False, disable_additional_properties=False, reasoning_parser='qwen3', reasoning_parser_plugin='', enable_in_reasoning=False), observability_config=ObservabilityConfig(show_hidden_metrics_for_version=None, otlp_traces_endpoint=None, collect_detailed_traces=None, kv_cache_metrics=False, kv_cache_metrics_sample=0.01, cudagraph_metrics=False, enable_layerwise_nvtx_tracing=False, enable_mfu_metrics=False, enable_mm_processor_stats=False, enable_logging_iteration_details=False, jit_monitor_mode='warn', jit_monitor_verbose=False), seed=0, served_model_name=nvidia/Qwen3.6-35B-A3B-NVFP4, enable_prefix_caching=False, enable_chunked_prefill=True, pooler_config=None, compilation_config={'mode': &lt;CompilationMode.VLLM_COMPILE: 3&gt;, 'debug_dump_path': None, 'cache_dir': '', 'compile_cache_save_format': 'binary', 'backend': 'inductor', 'custom_ops': ['none'], 'ir_enable_torch_wrap': True, 'splitting_ops': ['vllm::unified_attention_with_output', 'vllm::unified_mla_attention_with_output', 'vllm::mamba_mixer2', 'vllm::mamba_mixer', 'vllm::short_conv', 'vllm::linear_attention', 'vllm::plamo2_mamba_mixer', 'vllm::qwen_gdn_attention_core', 'vllm::gdn_attention_core_xpu', 'vllm::olmo_hybrid_gdn_full_forward', 'vllm::kda_attention', 'vllm::sparse_attn_indexer', 'vllm::rocm_aiter_sparse_attn_indexer', 'vllm::deepseek_v4_attention', 'vllm::hpc_rope_norm_forward', 'vllm::unified_kv_cache_update', 'vllm::unified_mla_kv_cache_update'], 'compile_mm_encoder': False, 'cudagraph_mm_encoder': False, 'encoder_cudagraph_token_budgets': [], 'encoder_cudagraph_max_vision_items_per_batch': 0, 'encoder_cudagraph_max_frames_per_batch': None, 'compile_sizes': [], 'compile_ranges_endpoints': [4096], 'inductor_compile_config': {'enable_auto_functionalized_v2': False, 'size_asserts': False, 'alignment_asserts': False, 'scalar_asserts': False, 'combo_kernels': True, 'benchmark_combo_kernel': True}, 'inductor_passes': {}, 'cudagraph_mode': &lt;CUDAGraphMode.FULL_AND_PIECEWISE: (2, 1)&gt;, 'cudagraph_num_of_warmups': 1, 'cudagraph_capture_sizes': [1, 2], 'cudagraph_copy_inputs': False, 'cudagraph_specialize_lora': True, 'use_inductor_graph_partition': False, 'pass_config': {'fuse_norm_quant': False, 'fuse_act_quant': False, 'fuse_attn_quant': False, 'enable_sp': False, 'fuse_gemm_comms': False, 'fuse_allreduce_rms': False, 'enable_qk_norm_rope_fusion': False, 'fuse_rope_kvcache_cat_mla': False, 'fuse_act_padding': False, 'fuse_qk_norm_rope_kvcache': False}, 'max_cudagraph_capture_size': 2, 'dynamic_shapes_config': {'type': &lt;DynamicShapesType.BACKED: 'backed'&gt;, 'evaluate_guards': False, 'assume_32_bit_indexing': False}, 'local_cache_dir': None, 'fast_moe_cold_start': False, 'static_all_moe_layers': []}, kernel_config=KernelConfig(ir_op_priority=IrOpPriorityConfig(rms_norm=['native'], fused_add_rms_norm=['native']), enable_flashinfer_autotune=True, enable_cutedsl_warmup=True, enable_bf16x3_router_gemm=False, moe_backend='auto', linear_backend='auto')
(EngineCore pid=1833533) Warning: You are sending unauthenticated requests to the HF Hub. Please set a HF_TOKEN to enable higher rate limits and faster downloads.
(EngineCore pid=1833533) INFO 08-07 13:30:45 [parallel_state.py:1615] world_size=1 rank=0 local_rank=0 distributed_init_method=tcp://10.3.2.87:52967 backend=nccl
(EngineCore pid=1833533) INFO 08-07 13:30:45 [parallel_state.py:1946] rank 0 in world size 1 is assigned as DP rank 0, PP rank 0, PCP rank 0, TP rank 0, EP rank 0, EPLB rank N/A
(EngineCore pid=1833533) Failed to get device capability: SM 12.x requires CUDA &gt;= 12.9.
(EngineCore pid=1833533) Failed to get device capability: SM 12.x requires CUDA &gt;= 12.9.
(EngineCore pid=1833533) INFO 08-07 13:30:49 [topk_topp_sampler.py:55] Using FlashInfer for top-p &amp; top-k sampling.
(EngineCore pid=1833533) [transformers] The `use_fast` parameter is deprecated and will be removed in a future version. Use `backend="torchvision"` instead of `use_fast=True`, or `backend="pil"` instead of `use_fast=False`.
(EngineCore pid=1833533) INFO 08-07 13:31:05 [gpu_model_runner.py:5250] Starting to load model nvidia/Qwen3.6-35B-A3B-NVFP4...
(EngineCore pid=1833533) INFO 08-07 13:31:05 [cuda.py:541] Using backend AttentionBackendEnum.FLASH_ATTN for vit attention
(EngineCore pid=1833533) INFO 08-07 13:31:05 [mm_encoder_attention.py:373] Using AttentionBackendEnum.FLASH_ATTN for MMEncoderAttention.
(EngineCore pid=1833533) INFO 08-07 13:31:05 [__init__.py:635] Selected MarlinFP8ScaledMMLinearKernel for ModelOptFp8LinearMethod
(EngineCore pid=1833533) INFO 08-07 13:31:05 [qwen_gdn_linear_attn.py:150] Using Triton/FLA GDN prefill kernel (requested=auto, head_k_dim=128).
(EngineCore pid=1833533) INFO 08-07 13:31:05 [nvfp4.py:285] Using 'MARLIN' NvFp4 MoE backend out of potential backends: ['FLASHINFER_TRTLLM', 'FLASHINFER_CUTEDSL', 'FLASHINFER_CUTEDSL_BATCHED', 'FLASHINFER_CUTLASS', 'VLLM_CUTLASS', 'MARLIN', 'HUMMING', 'EMULATION'].
(EngineCore pid=1833533) INFO 08-07 13:31:05 [cuda.py:422] Using AttentionBackendEnum.FLASHINFER backend.
(EngineCore pid=1833533) ERROR 08-07 13:31:07 [gpu_model_runner.py:5345] Failed to load model - not enough GPU memory. Try lowering --gpu-memory-utilization to free memory for weights, increasing --tensor-parallel-size, or using --quantization. See https://docs.vllm.ai/en/latest/configuration/conserving_memory/ for more tips. (original error: CUDA out of memory. Tried to allocate 256.00 MiB. GPU 0 has a total capacity of 15.52 GiB of which 68.62 MiB is free. Including non-PyTorch memory, this process has 15.44 GiB memory in use. Of the allocated memory 15.15 GiB is allocated by PyTorch, and 78.45 MiB is reserved by PyTorch but unallocated. If reserved but unallocated memory is large try setting PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True to avoid fragmentation.  See documentation for Memory Management  (https://docs.pytorch.org/docs/stable/notes/cuda.html#optimizing-memory-usage-with-pytorch-cuda-alloc-conf))


]]></description><link>https://lemmy.cyber-neurones.org/topic/390/tuxedo-17-installation-vllm</link><guid isPermaLink="true">https://lemmy.cyber-neurones.org/topic/390/tuxedo-17-installation-vllm</guid><dc:creator><![CDATA[Tuxedo17]]></dc:creator><pubDate>Fri, 07 Aug 2026 09:59:18 GMT</pubDate></item><item><title><![CDATA[LLama-swap : test du NVFP4]]></title><description><![CDATA[NVIDIA : https://huggingface.co/nvidia/Qwen3.6-35B-A3B-NVFP4
Exemple :
models:
  # Premier modèle : Llama 3 8B
  - name: "meta-llama/Meta-Llama-3-8B-Instruct"
    command: &gt;
      vllm serve meta-llama/Meta-Llama-3-8B-Instruct 
      --port 8001 
      --gpu-memory-utilization 0.85
    ready_url: "http://127.0.0"
    upstream_url: "http://127.0.0"

]]></description><link>https://lemmy.cyber-neurones.org/topic/389/llama-swap-test-du-nvfp4</link><guid isPermaLink="true">https://lemmy.cyber-neurones.org/topic/389/llama-swap-test-du-nvfp4</guid><dc:creator><![CDATA[Tuxedo17]]></dc:creator><pubDate>Fri, 07 Aug 2026 06:55:27 GMT</pubDate></item><item><title><![CDATA[LLama-swap : Temps de compression]]></title><description><![CDATA[Update :
[image: llama-swap-compression-hourly.png]
]]></description><link>https://lemmy.cyber-neurones.org/topic/388/llama-swap-temps-de-compression</link><guid isPermaLink="true">https://lemmy.cyber-neurones.org/topic/388/llama-swap-temps-de-compression</guid><dc:creator><![CDATA[Tuxedo17]]></dc:creator><pubDate>Fri, 07 Aug 2026 06:41:17 GMT</pubDate></item><item><title><![CDATA[Hermes configuration contexte en dur ...]]></title><description><![CDATA[Requete pour la compression :
{
  "messages": [
    {
      "role": "user",
      "content": "You are a summarization agent creating a context checkpoint. Treat the conversation turns below as source material for a compact record of prior work. Produce only the structured summary; do not add a greeting, preamble, or prefix. Write the summary in the same language the user was using in the conversation — do not translate or switch to English. NEVER include API keys, tokens, passwords, secrets, credentials, or connection strings in the summary — replace any that appear with [REDACTED]. Note that the user had credentials present, but do not preserve their values.\n\nCreate a structured checkpoint summary for the conversation after earlier turns are compacted. The summary should preserve enough detail for continuity without re-reading the original turns

]]></description><link>https://lemmy.cyber-neurones.org/topic/387/hermes-configuration-contexte-en-dur-...</link><guid isPermaLink="true">https://lemmy.cyber-neurones.org/topic/387/hermes-configuration-contexte-en-dur-...</guid><dc:creator><![CDATA[Tuxedo17]]></dc:creator><pubDate>Wed, 05 Aug 2026 15:52:12 GMT</pubDate></item><item><title><![CDATA[LLAMA-SWAP :  147 933 343 tokens Cached]]></title><description><![CDATA[Temps :
[image: llama-swap-daily-duration.png]
]]></description><link>https://lemmy.cyber-neurones.org/topic/386/llama-swap-147-933-343-tokens-cached</link><guid isPermaLink="true">https://lemmy.cyber-neurones.org/topic/386/llama-swap-147-933-343-tokens-cached</guid><dc:creator><![CDATA[Tuxedo17]]></dc:creator><pubDate>Wed, 05 Aug 2026 15:09:06 GMT</pubDate></item><item><title><![CDATA[Crash i915 0000:00:02.0: [drm] drm_WARN_ON_ONCE(t_vblank < vblank->time)]]></title><link>https://lemmy.cyber-neurones.org/topic/384/crash-i915-0000-00-02.0-drm-drm_warn_on_once-t_vblank-vblank-time</link><guid isPermaLink="true">https://lemmy.cyber-neurones.org/topic/384/crash-i915-0000-00-02.0-drm-drm_warn_on_once-t_vblank-vblank-time</guid><pubDate>Tue, 28 Jul 2026 14:53:43 GMT</pubDate></item><item><title><![CDATA[TuxedoOS : Downgrade de NVIDIA (595.84)]]></title><description><![CDATA[J’avais oublier de modifier le script /usr/local/bin/aorus-bridge .

VENDOR_ID="${VENDOR_ID:-0x10de}"   # NVIDIA
# DEVICE_ID: exact PCI device id to match (e.g. 0x2b85 for RTX 5090). Leave
# empty to match ANY NVIDIA display controller by PCI class — this covers the
# whole RTX 50-series (5090=0x2b85, 5060 Ti=0x2d04, …) without a hardcoded list,
# in the same spirit as is_tb_tunneled below. Set DEVICE_ID to pin one board.
DEVICE_ID="${DEVICE_ID:-0x2d04}"
# PCI base class 0x03 = display controller (VGA 0x0300 / 3D 0x0302); selects the
# GPU function and excludes its HDMI-audio function (class 0x0403).
GPU_CLASS_PREFIX="${GPU_CLASS_PREFIX:-0x03}"


]]></description><link>https://lemmy.cyber-neurones.org/topic/383/tuxedoos-downgrade-de-nvidia-595.84</link><guid isPermaLink="true">https://lemmy.cyber-neurones.org/topic/383/tuxedoos-downgrade-de-nvidia-595.84</guid><dc:creator><![CDATA[Tuxedo17]]></dc:creator><pubDate>Mon, 27 Jul 2026 10:05:58 GMT</pubDate></item><item><title><![CDATA[Installation de llama-swap]]></title><description><![CDATA[Sur l’interface :
[image: image.png]
]]></description><link>https://lemmy.cyber-neurones.org/topic/379/installation-de-llama-swap</link><guid isPermaLink="true">https://lemmy.cyber-neurones.org/topic/379/installation-de-llama-swap</guid><dc:creator><![CDATA[Tuxedo17]]></dc:creator><pubDate>Wed, 22 Jul 2026 07:32:11 GMT</pubDate></item><item><title><![CDATA[Test Hermes IA en local (avec eGPU)]]></title><description><![CDATA[A tester https://github.com/HalfbitStudio/hermes-plugin-rocketchat .
]]></description><link>https://lemmy.cyber-neurones.org/topic/377/test-hermes-ia-en-local-avec-egpu</link><guid isPermaLink="true">https://lemmy.cyber-neurones.org/topic/377/test-hermes-ia-en-local-avec-egpu</guid><dc:creator><![CDATA[Tuxedo17]]></dc:creator><pubDate>Fri, 17 Jul 2026 16:19:21 GMT</pubDate></item><item><title><![CDATA[Tuxedo 17 (en Ubuntu 24) + GIGABYTE AORUS RTX 5060 Ti AI Box Carte Graphique - 16GB GDDR7.]]></title><description><![CDATA[Nouveau test :
# llama-bench -m /models/Qwen3.6-27B-Q3_K_M.gguf -dev CUDA0 -fa on --threads 12 -ngl 999 --cache-type-v q8_0
ggml_cuda_init: found 2 CUDA devices (Total VRAM: 21692 MiB):
  Device 0: NVIDIA GeForce RTX 5060 Ti, compute capability 12.0, VMM: yes, VRAM: 15888 MiB
  Device 1: NVIDIA GeForce RTX 3060 Laptop GPU, compute capability 8.6, VMM: yes, VRAM: 5803 MiB
| model                          |       size |     params | backend    | ngl | threads | type_v |  fa | dev          |            test |                  t/s |
| ------------------------------ | ---------: | ---------: | ---------- | --: | ------: | -----: | --: | ------------ | --------------: | -------------------: |
| qwen35 27B Q3_K - Medium       |  12.64 GiB |    26.90 B | CUDA       | 999 |      12 |   q8_0 |   1 | CUDA0        |           pp512 |        387.17 ± 2.25 |
| qwen35 27B Q3_K - Medium       |  12.64 GiB |    26.90 B | CUDA       | 999 |      12 |   q8_0 |   1 | CUDA0        |           tg128 |         18.31 ± 0.01 |

# llama-bench -m /models/Qwen3.6-27B-Q3_K_M.gguf -fa on --threads 12 -ngl 999 --cache-type-v q8_0
ggml_cuda_init: found 2 CUDA devices (Total VRAM: 21692 MiB):
  Device 0: NVIDIA GeForce RTX 5060 Ti, compute capability 12.0, VMM: yes, VRAM: 15888 MiB
  Device 1: NVIDIA GeForce RTX 3060 Laptop GPU, compute capability 8.6, VMM: yes, VRAM: 5803 MiB
| model                          |       size |     params | backend    | ngl | threads | type_v |  fa |            test |                  t/s |
| ------------------------------ | ---------: | ---------: | ---------- | --: | ------: | -----: | --: | --------------: | -------------------: |
| qwen35 27B Q3_K - Medium       |  12.64 GiB |    26.90 B | CUDA       | 999 |      12 |   q8_0 |   1 |           pp512 |        367.28 ± 1.88 |
| qwen35 27B Q3_K - Medium       |  12.64 GiB |    26.90 B | CUDA       | 999 |      12 |   q8_0 |   1 |           tg128 |         16.71 ± 0.03 |



]]></description><link>https://lemmy.cyber-neurones.org/topic/376/tuxedo-17-en-ubuntu-24-gigabyte-aorus-rtx-5060-ti-ai-box-carte-graphique-16gb-gddr7.</link><guid isPermaLink="true">https://lemmy.cyber-neurones.org/topic/376/tuxedo-17-en-ubuntu-24-gigabyte-aorus-rtx-5060-ti-ai-box-carte-graphique-16gb-gddr7.</guid><dc:creator><![CDATA[Tuxedo17]]></dc:creator><pubDate>Thu, 16 Jul 2026 07:07:10 GMT</pubDate></item><item><title><![CDATA[Tuxedo 17 + GIGABYTE AORUS RTX 5060 Ti AI Box Carte Graphique - 16GB GDDR7.]]></title><description><![CDATA[J’ai donc supprimer le service de Tuxedo et fait une mise à jours :
# apt-get autoremove.

# systemctl stop tccd

# systemctl disable tccd

# ubuntu-drivers autoinstall


En fait tccd fait l’installation de 560.35.05 et pas de 595.71.05.
]]></description><link>https://lemmy.cyber-neurones.org/topic/375/tuxedo-17-gigabyte-aorus-rtx-5060-ti-ai-box-carte-graphique-16gb-gddr7.</link><guid isPermaLink="true">https://lemmy.cyber-neurones.org/topic/375/tuxedo-17-gigabyte-aorus-rtx-5060-ti-ai-box-carte-graphique-16gb-gddr7.</guid><dc:creator><![CDATA[Tuxedo17]]></dc:creator><pubDate>Wed, 15 Jul 2026 07:28:42 GMT</pubDate></item><item><title><![CDATA[Test Hermes IA en local]]></title><description><![CDATA[L’export en Markdown ne fonctionne pas :
# hermes sessions export --format md --older-than 90 --dry-run
usage: hermes [-h] [--version] [-z PROMPT] [--usage-file PATH] [-m MODEL] [--provider PROVIDER]
              [-t TOOLSETS] [--resume SESSION] [--continue [SESSION_NAME]] [--worktree] [--accept-hooks]
              [--skills SKILLS] [--yolo] [--pass-session-id] [--ignore-user-config] [--ignore-rules]
              [--safe-mode] [--tui] [--cli] [--dev]
              {chat,model,moa,fallback,secrets,migrate,gateway,proxy,lsp,setup,postinstall,whatsapp,whatsapp-cloud,slack,send,login,logout,auth,status,cron,webhook,portal,kanban,project,hooks,doctor,security,dump,debug,backup,checkpoints,import,config,console,pairing,skills,bundles,plugins,curator,pets,journey,learning,memory-graph,memory,tools,computer-use,mcp,sessions,insights,claw,version,update,uninstall,acp,profile,completion,dashboard,serve,desktop,gui,logs,prompt-size}
              ...
hermes: error: unrecognized arguments: --format --older-than 90 --dry-run

]]></description><link>https://lemmy.cyber-neurones.org/topic/374/test-hermes-ia-en-local</link><guid isPermaLink="true">https://lemmy.cyber-neurones.org/topic/374/test-hermes-ia-en-local</guid><dc:creator><![CDATA[Tuxedo17]]></dc:creator><pubDate>Tue, 07 Jul 2026 10:04:40 GMT</pubDate></item><item><title><![CDATA[Installation llama.cpp sous Windows 11 avec Ubuntu 22]]></title><description><![CDATA[Je fait un make install avant puis un nouveau test :
root@pcremi:/workspace/llama.cpp/build/bin# ./llama-bench -m  /models/Qwen3.6-27B-Q4_K_M.gguf
ggml_cuda_init: found 1 ROCm devices (Total VRAM: 16304 MiB):
  Device 0: AMD Radeon RX 9060 XT, gfx1200 (0x1200), VMM: no, Wave Size: 32, VRAM: 16304 MiB
| model                          |       size |     params | backend    | ngl |            test |                  t/s |
| ------------------------------ | ---------: | ---------: | ---------- | --: | --------------: | -------------------: |
pid:3402 tid:0x757b61d04280 [CreateContext] fail 11
llama-bench: /home/remia/librocdxg/src/wddm/queue.cpp:267: wsl::thunk::ComputeQueue::ComputeQueue(wsl::thunk::WDDMDevice*, void*, uint64_t, std::atomic&lt;long unsigned int&gt;*, std::atomic&lt;long unsigned int&gt;*, volatile int64_t*, uint32_t, uint32_t, bool): Assertion `ret' failed.
Aborted (core dumped)

Pas mieux.
]]></description><link>https://lemmy.cyber-neurones.org/topic/371/installation-llama.cpp-sous-windows-11-avec-ubuntu-22</link><guid isPermaLink="true">https://lemmy.cyber-neurones.org/topic/371/installation-llama.cpp-sous-windows-11-avec-ubuntu-22</guid><dc:creator><![CDATA[Remi]]></dc:creator><pubDate>Mon, 22 Jun 2026 17:30:55 GMT</pubDate></item><item><title><![CDATA[llama.cpp : exemple de fichier service]]></title><link>https://lemmy.cyber-neurones.org/topic/370/llama.cpp-exemple-de-fichier-service</link><guid isPermaLink="true">https://lemmy.cyber-neurones.org/topic/370/llama.cpp-exemple-de-fichier-service</guid><pubDate>Mon, 22 Jun 2026 07:54:19 GMT</pubDate></item><item><title><![CDATA[llama.cpp avec Vulkan]]></title><description><![CDATA[Petit test : 40.81 t/s
]]></description><link>https://lemmy.cyber-neurones.org/topic/369/llama.cpp-avec-vulkan</link><guid isPermaLink="true">https://lemmy.cyber-neurones.org/topic/369/llama.cpp-avec-vulkan</guid><dc:creator><![CDATA[farias]]></dc:creator><pubDate>Fri, 19 Jun 2026 14:16:21 GMT</pubDate></item><item><title><![CDATA[llama.cpp : Installation]]></title><description><![CDATA[Mise à jours de Ollama :
# curl -fsSL https://ollama.com/install.sh | sh
&gt;&gt;&gt; Cleaning up old version at /usr/local/lib/ollama
&gt;&gt;&gt; Installing ollama to /usr/local
&gt;&gt;&gt; Downloading ollama-linux-amd64.tar.zst
######################################################################## 100.0%
&gt;&gt;&gt; Adding ollama user to render group...
&gt;&gt;&gt; Adding ollama user to video group...
&gt;&gt;&gt; Adding current user to ollama group...
&gt;&gt;&gt; Creating ollama systemd service...
&gt;&gt;&gt; Enabling and starting ollama service...
&gt;&gt;&gt; NVIDIA GPU installed.


Visiblement même problème :
...
juin 19 10:51:37  ollama[2964]: time=2026-06-19T10:51:37.153Z level=INFO source=model_list_cache.go:111 msg="model list cache hydration complete" models=16 failures=0 elapsed=654.370427ms
juin 19 10:51:42  ollama[2964]: time=2026-06-19T10:51:42.591Z level=WARN source=cuda_compat.go:38 msg="NVIDIA driver too old" device="Quadro M5000" compute=5.2 driver=535 required_driver="570 or newer"
juin 19 10:51:42  ollama[2964]: time=2026-06-19T10:51:42.591Z level=WARN source=cuda_compat.go:38 msg="NVIDIA driver too old" device="Quadro M4000" compute=5.2 driver=535 required_driver="570 or newer"
juin 19 10:51:43  ollama[2964]: time=2026-06-19T10:51:43.181Z level=INFO source=types.go:32 msg="inference compute" id=1 filter_id=1 library=Vulkan compute=0.0 name=Vulkan1 description="Quadro M4000" libd&gt;
juin 19 10:51:43  ollama[2964]: time=2026-06-19T10:51:43.181Z level=INFO source=types.go:32 msg="inference compute" id=0 filter_id=0 library=Vulkan compute=0.0 name=Vulkan0 description="Quadro M5000" libd&gt;
...

]]></description><link>https://lemmy.cyber-neurones.org/topic/368/llama.cpp-installation</link><guid isPermaLink="true">https://lemmy.cyber-neurones.org/topic/368/llama.cpp-installation</guid><dc:creator><![CDATA[farias]]></dc:creator><pubDate>Fri, 19 Jun 2026 09:00:17 GMT</pubDate></item><item><title><![CDATA[llama.cpp :  llama-bench : CPU ( no CUDA )]]></title><link>https://lemmy.cyber-neurones.org/topic/367/llama.cpp-llama-bench-cpu-no-cuda</link><guid isPermaLink="true">https://lemmy.cyber-neurones.org/topic/367/llama.cpp-llama-bench-cpu-no-cuda</guid><pubDate>Thu, 18 Jun 2026 14:40:55 GMT</pubDate></item><item><title><![CDATA[llama.cpp : llama-server : Erreur 404 file not found]]></title><description><![CDATA[Les logs avant le build :
# export LD_LIBRARY_PATH=/usr/local/cuda/lib
# export PATH=$PATH:/usr/local/cuda/bin
# cmake .. -DGGML_CUDA=ON
CMAKE_BUILD_TYPE=Release
-- Warning: ccache not found - consider installing it for faster compilation or disable this warning with GGML_CCACHE=OFF
-- CMAKE_SYSTEM_PROCESSOR: x86_64
-- GGML_SYSTEM_ARCH: x86
-- Including CPU backend
-- x86 detected
-- Adding CPU backend variant ggml-cpu: -march=native 
-- CUDA Toolkit found
-- Using CMAKE_CUDA_ARCHITECTURES=75-virtual;80-virtual;86-real;89-real;90-virtual;120a-real;121a-real CMAKE_CUDA_ARCHITECTURES_NATIVE=
-- CUDA host compiler is GNU 11.4.0
-- Including CUDA backend
-- ggml version: 0.15.1
-- ggml commit:  b4024af6c
-- OpenSSL found: 3.0.2
-- Generating embedded license file for target: llama-app
-- Configuring done
-- Generating done
-- Build files have been written to: /home/XXXX/GIT/llama.cpp/build


]]></description><link>https://lemmy.cyber-neurones.org/topic/366/llama.cpp-llama-server-erreur-404-file-not-found</link><guid isPermaLink="true">https://lemmy.cyber-neurones.org/topic/366/llama.cpp-llama-server-erreur-404-file-not-found</guid><dc:creator><![CDATA[Tuxedo17]]></dc:creator><pubDate>Thu, 18 Jun 2026 13:45:28 GMT</pubDate></item><item><title><![CDATA[Benchmark llama.cpp sur Tuxedo 17]]></title><description><![CDATA[Pendant l’utilisation :
# nvidia-smi 
Tue Jun 30 12:35:32 2026       
+-----------------------------------------------------------------------------------------+
| NVIDIA-SMI 595.71.05              Driver Version: 595.71.05      CUDA Version: 13.2     |
+-----------------------------------------+------------------------+----------------------+
| GPU  Name                 Persistence-M | Bus-Id          Disp.A | Volatile Uncorr. ECC |
| Fan  Temp   Perf          Pwr:Usage/Cap |           Memory-Usage | GPU-Util  Compute M. |
|                                         |                        |               MIG M. |
|=========================================+========================+======================|
|   0  NVIDIA GeForce RTX 3060 ...    Off |   00000000:01:00.0 Off |                  N/A |
| N/A   55C    P0             28W /  115W |    4621MiB /   6144MiB |      0%      Default |
|                                         |                        |                  N/A |
+-----------------------------------------+------------------------+----------------------+

+-----------------------------------------------------------------------------------------+
| Processes:                                                                              |
|  GPU   GI   CI              PID   Type   Process name                        GPU Memory |
|        ID   ID                                                               Usage      |
|=========================================================================================|
|    0   N/A  N/A         1820494      G   /usr/lib/xorg/Xorg                        4MiB |
|    0   N/A  N/A         1845119      C   /usr/local/bin/llama-server            4598MiB |
+-----------------------------------------------------------------------------------------+


]]></description><link>https://lemmy.cyber-neurones.org/topic/365/benchmark-llama.cpp-sur-tuxedo-17</link><guid isPermaLink="true">https://lemmy.cyber-neurones.org/topic/365/benchmark-llama.cpp-sur-tuxedo-17</guid><dc:creator><![CDATA[Tuxedo17]]></dc:creator><pubDate>Wed, 17 Jun 2026 15:19:24 GMT</pubDate></item></channel></rss>