ggml_cuda_init: found 1 CUDA devices (Total VRAM: 32606 MiB):
  Device 0: NVIDIA GeForce RTX 5090, compute capability 12.0, VMM: yes, VRAM: 32606 MiB
llama_model_loader: loaded meta data with 35 key-value pairs and 771 tensors from C:\Users\eatur\.lmstudio\models\mradermacher\Seed-OSS-36B-Base-woSyn-GGUF\Seed-OSS-36B-Base-woSyn.Q4_K_M.gguf (version GGUF V3 (latest))
llama_model_loader: Dumping metadata keys/values. Note: KV overrides do not apply in this output.
llama_model_loader: - kv   0:                       general.architecture str              = seed_oss
llama_model_loader: - kv   1:                               general.type str              = model
llama_model_loader: - kv   2:                               general.name str              = Seed OSS 36B Base woSyn
llama_model_loader: - kv   3:                           general.finetune str              = Base-woSyn
llama_model_loader: - kv   4:                           general.basename str              = Seed-OSS
llama_model_loader: - kv   5:                         general.size_label str              = 36B
llama_model_loader: - kv   6:                            general.license str              = apache-2.0
llama_model_loader: - kv   7:                               general.tags arr[str,2]       = ["vllm", "text-generation"]
llama_model_loader: - kv   8:                       seed_oss.block_count u32              = 64
llama_model_loader: - kv   9:                    seed_oss.context_length u32              = 524288
llama_model_loader: - kv  10:                  seed_oss.embedding_length u32              = 5120
llama_model_loader: - kv  11:               seed_oss.feed_forward_length u32              = 27648
llama_model_loader: - kv  12:              seed_oss.attention.head_count u32              = 80
llama_model_loader: - kv  13:           seed_oss.attention.head_count_kv u32              = 8
llama_model_loader: - kv  14:                    seed_oss.rope.freq_base f32              = 10000000.000000
llama_model_loader: - kv  15:  seed_oss.attention.layer_norm_rms_epsilon f32              = 0.000001
llama_model_loader: - kv  16:              seed_oss.attention.key_length u32              = 128
llama_model_loader: - kv  17:            seed_oss.attention.value_length u32              = 128
llama_model_loader: - kv  18:                       tokenizer.ggml.model str              = gpt2
llama_model_loader: - kv  19:                         tokenizer.ggml.pre str              = seed-coder
llama_model_loader: - kv  20:                      tokenizer.ggml.tokens arr[str,155136]  = ["<seed:bos>", "<seed:pad>", "<seed:e...
llama_model_loader: - kv  21:                  tokenizer.ggml.token_type arr[i32,155136]  = [3, 3, 3, 4, 4, 4, 4, 4, 4, 3, 3, 3, ...
llama_model_loader: - kv  22:                      tokenizer.ggml.merges arr[str,154737]  = ["Ġ Ġ", "Ġ t", "i n", "Ġ a", "e r...
llama_model_loader: - kv  23:                tokenizer.ggml.bos_token_id u32              = 0
llama_model_loader: - kv  24:                tokenizer.ggml.eos_token_id u32              = 2
llama_model_loader: - kv  25:            tokenizer.ggml.padding_token_id u32              = 1
llama_model_loader: - kv  26:               general.quantization_version u32              = 2
llama_model_loader: - kv  27:                          general.file_type u32              = 15
llama_model_loader: - kv  28:                                general.url str              = https://huggingface.co/mradermacher/S...
llama_model_loader: - kv  29:              mradermacher.quantize_version str              = 2
llama_model_loader: - kv  30:                  mradermacher.quantized_by str              = mradermacher
llama_model_loader: - kv  31:                  mradermacher.quantized_at str              = 2025-08-24T20:33:17+02:00
llama_model_loader: - kv  32:                  mradermacher.quantized_on str              = marco
llama_model_loader: - kv  33:                         general.source.url str              = https://huggingface.co/ByteDance-Seed...
llama_model_loader: - kv  34:                  mradermacher.convert_type str              = hf
llama_model_loader: - type  f32:  321 tensors
llama_model_loader: - type q4_K:  385 tensors
llama_model_loader: - type q6_K:   65 tensors
print_info: file format = GGUF V3 (latest)
print_info: file type   = Q4_K - Medium
print_info: file size   = 20.26 GiB (4.81 BPW) 
llama_prepare_model_devices: using device CUDA0 (NVIDIA GeForce RTX 5090) (0000:02:00.0) - 30927 MiB free
init_tokenizer: initializing tokenizer for type 2
load: 0 unused tokens
load: control token:      0 '<seed:bos>' is not marked as EOG
load: control token:     42 '<[PLHD42_never_used]>' is not marked as EOG
load: control token:      1 '<seed:pad>' is not marked as EOG
load: control token:     30 '<[PLHD30_never_used]>' is not marked as EOG
load: control token:      2 '<seed:eos>' is not marked as EOG
load: control token:     33 '<[PLHD33_never_used]>' is not marked as EOG
load: control token:    103 '<[PLHD103_never_used]>' is not marked as EOG
load: control token:     54 '<[PLHD54_never_used]>' is not marked as EOG
load: control token:     18 '<[PLHD18_never_used]>' is not marked as EOG
load: control token:     28 '<[PLHD28_never_used]>' is not marked as EOG
load: control token:     56 '<[PLHD56_never_used]>' is not marked as EOG
load: control token:      9 '<[PLHD9_never_used]>' is not marked as EOG
load: control token:     61 '<[PLHD61_never_used]>' is not marked as EOG
load: control token:     10 '<[PLHD10_never_used]>' is not marked as EOG
load: control token:     60 '<[PLHD60_never_used]>' is not marked as EOG
load: control token:     11 '<[PLHD11_never_used]>' is not marked as EOG
load: control token:     63 '<[PLHD63_never_used]>' is not marked as EOG
load: control token:     12 '<[PLHD12_never_used]>' is not marked as EOG
load: control token:     62 '<[PLHD62_never_used]>' is not marked as EOG
load: control token:     13 '<[PLHD13_never_used]>' is not marked as EOG
load: control token:     58 '<[PLHD58_never_used]>' is not marked as EOG
load: control token:     14 '<[PLHD14_never_used]>' is not marked as EOG
load: control token:     59 '<[PLHD59_never_used]>' is not marked as EOG
load: control token:     15 '<[PLHD15_never_used]>' is not marked as EOG
load: control token:     16 '<[PLHD16_never_used]>' is not marked as EOG
load: control token:     17 '<[PLHD17_never_used]>' is not marked as EOG
load: control token:     55 '<[PLHD55_never_used]>' is not marked as EOG
load: control token:     19 '<[PLHD19_never_used]>' is not marked as EOG
load: control token:     20 '<[PLHD20_never_used]>' is not marked as EOG
load: control token:     21 '<[PLHD21_never_used]>' is not marked as EOG
load: control token:     22 '<[PLHD22_never_used]>' is not marked as EOG
load: control token:     23 '<[PLHD23_never_used]>' is not marked as EOG
load: control token:     24 '<[PLHD24_never_used]>' is not marked as EOG
load: control token:     25 '<[PLHD25_never_used]>' is not marked as EOG
load: control token:     26 '<[PLHD26_never_used]>' is not marked as EOG
load: control token:     27 '<[PLHD27_never_used]>' is not marked as EOG
load: control token:     29 '<[PLHD29_never_used]>' is not marked as EOG
load: control token:     31 '<[PLHD31_never_used]>' is not marked as EOG
load: control token:     32 '<[PLHD32_never_used]>' is not marked as EOG
load: control token:     34 '<[PLHD34_never_used]>' is not marked as EOG
load: control token:     35 '<[PLHD35_never_used]>' is not marked as EOG
load: control token:     36 '<[PLHD36_never_used]>' is not marked as EOG
load: control token:     37 '<[PLHD37_never_used]>' is not marked as EOG
load: control token:     38 '<[PLHD38_never_used]>' is not marked as EOG
load: control token:     39 '<[PLHD39_never_used]>' is not marked as EOG
load: control token:     40 '<[PLHD40_never_used]>' is not marked as EOG
load: control token:     41 '<[PLHD41_never_used]>' is not marked as EOG
load: control token:     43 '<[PLHD43_never_used]>' is not marked as EOG
load: control token:     44 '<[PLHD44_never_used]>' is not marked as EOG
load: control token:     45 '<[PLHD45_never_used]>' is not marked as EOG
load: control token:     46 '<[PLHD46_never_used]>' is not marked as EOG
load: control token:     47 '<[PLHD47_never_used]>' is not marked as EOG
load: control token:     48 '<[PLHD48_never_used]>' is not marked as EOG
load: control token:     49 '<[PLHD49_never_used]>' is not marked as EOG
load: control token:     50 '<[PLHD50_never_used]>' is not marked as EOG
load: control token:     51 '<[PLHD51_never_used]>' is not marked as EOG
load: control token:     52 '<[PLHD52_never_used]>' is not marked as EOG
load: control token:     53 '<[PLHD53_never_used]>' is not marked as EOG
load: control token:     57 '<[PLHD57_never_used]>' is not marked as EOG
load: control token:     64 '<[PLHD64_never_used]>' is not marked as EOG
load: control token:     65 '<[PLHD65_never_used]>' is not marked as EOG
load: control token:     66 '<[PLHD66_never_used]>' is not marked as EOG
load: control token:     67 '<[PLHD67_never_used]>' is not marked as EOG
load: control token:     68 '<[PLHD68_never_used]>' is not marked as EOG
load: control token:     69 '<[PLHD69_never_used]>' is not marked as EOG
load: control token:     70 '<[PLHD70_never_used]>' is not marked as EOG
load: control token:     71 '<[PLHD71_never_used]>' is not marked as EOG
load: control token:     72 '<[PLHD72_never_used]>' is not marked as EOG
load: control token:     73 '<[PLHD73_never_used]>' is not marked as EOG
load: control token:     74 '<[PLHD74_never_used]>' is not marked as EOG
load: control token:     75 '<[PLHD75_never_used]>' is not marked as EOG
load: control token:     76 '<[PLHD76_never_used]>' is not marked as EOG
load: control token:     77 '<[PLHD77_never_used]>' is not marked as EOG
load: control token:     78 '<[PLHD78_never_used]>' is not marked as EOG
load: control token:     79 '<[PLHD79_never_used]>' is not marked as EOG
load: control token:     80 '<[PLHD80_never_used]>' is not marked as EOG
load: control token:     81 '<[PLHD81_never_used]>' is not marked as EOG
load: control token:     82 '<[PLHD82_never_used]>' is not marked as EOG
load: control token:     83 '<[PLHD83_never_used]>' is not marked as EOG
load: control token:     84 '<[PLHD84_never_used]>' is not marked as EOG
load: control token:     85 '<[PLHD85_never_used]>' is not marked as EOG
load: control token:     86 '<[PLHD86_never_used]>' is not marked as EOG
load: control token:     87 '<[PLHD87_never_used]>' is not marked as EOG
load: control token:     88 '<[PLHD88_never_used]>' is not marked as EOG
load: control token:     89 '<[PLHD89_never_used]>' is not marked as EOG
load: control token:     90 '<[PLHD90_never_used]>' is not marked as EOG
load: control token:     91 '<[PLHD91_never_used]>' is not marked as EOG
load: control token:     92 '<[PLHD92_never_used]>' is not marked as EOG
load: control token:     93 '<[PLHD93_never_used]>' is not marked as EOG
load: control token:     94 '<[PLHD94_never_used]>' is not marked as EOG
load: control token:     95 '<[PLHD95_never_used]>' is not marked as EOG
load: control token:     96 '<[PLHD96_never_used]>' is not marked as EOG
load: control token:     97 '<[PLHD97_never_used]>' is not marked as EOG
load: control token:     98 '<[PLHD98_never_used]>' is not marked as EOG
load: control token:     99 '<[PLHD99_never_used]>' is not marked as EOG
load: control token:    100 '<[PLHD100_never_used]>' is not marked as EOG
load: control token:    101 '<[PLHD101_never_used]>' is not marked as EOG
load: control token:    102 '<[PLHD102_never_used]>' is not marked as EOG
load: control token:    104 '<[PLHD104_never_used]>' is not marked as EOG
load: control token:    105 '<[PLHD105_never_used]>' is not marked as EOG
load: control token:    106 '<[PLHD106_never_used]>' is not marked as EOG
load: control token:    107 '<[PLHD107_never_used]>' is not marked as EOG
load: control token:    108 '<[PLHD108_never_used]>' is not marked as EOG
load: control token:    109 '<[PLHD109_never_used]>' is not marked as EOG
load: control token:    110 '<[PLHD110_never_used]>' is not marked as EOG
load: control token:    111 '<[PLHD111_never_used]>' is not marked as EOG
load: control token:    112 '<[PLHD112_never_used]>' is not marked as EOG
load: control token:    113 '<[PLHD113_never_used]>' is not marked as EOG
load: control token:    114 '<[PLHD114_never_used]>' is not marked as EOG
load: control token:    115 '<[PLHD115_never_used]>' is not marked as EOG
load: control token:    116 '<[PLHD116_never_used]>' is not marked as EOG
load: control token:    117 '<[PLHD117_never_used]>' is not marked as EOG
load: control token:    118 '<[PLHD118_never_used]>' is not marked as EOG
load: control token:    119 '<[PLHD119_never_used]>' is not marked as EOG
load: control token:    120 '<[PLHD120_never_used]>' is not marked as EOG
load: control token:    121 '<[PLHD121_never_used]>' is not marked as EOG
load: control token:    122 '<[PLHD122_never_used]>' is not marked as EOG
load: control token:    123 '<[PLHD123_never_used]>' is not marked as EOG
load: control token:    124 '<[PLHD124_never_used]>' is not marked as EOG
load: control token:    125 '<[PLHD125_never_used]>' is not marked as EOG
load: control token:    126 '<[PLHD126_never_used]>' is not marked as EOG
load: control token:    127 '<[PLHD127_never_used]>' is not marked as EOG
load: special_eos_id is not in special_eog_ids - the tokenizer config may be incorrect
load: printing all EOG tokens:
load:   - 2 ('<seed:eos>')
load: special tokens cache size = 128
load: token to piece cache size = 0.9296 MB
print_info: arch                  = seed_oss
print_info: vocab_only            = 0
print_info: no_alloc              = 0
print_info: n_ctx_train           = 524288
print_info: n_embd_inp            = 5120
print_info: n_embd                = 5120
print_info: n_embd_out            = 5120
print_info: n_layer               = 64
print_info: n_layer_all           = 64
print_info: n_head                = 80
print_info: n_head_kv             = 8
print_info: n_rot                 = 128
print_info: n_swa                 = 0
print_info: is_swa_any            = 0
print_info: non_causal_type       = 0
print_info: n_embd_head_k         = 128
print_info: n_embd_head_v         = 128
print_info: n_gqa                 = 10
print_info: n_embd_k_gqa          = 1024
print_info: n_embd_v_gqa          = 1024
print_info: f_norm_eps            = 0.0e+00
print_info: f_norm_rms_eps        = 1.0e-06
print_info: f_clamp_kqv           = 0.0e+00
print_info: f_max_alibi_bias      = 0.0e+00
print_info: f_logit_scale         = 0.0e+00
print_info: f_attn_scale          = 0.0e+00
print_info: f_attn_value_scale    = 0.0000
print_info: n_ff                  = 27648
print_info: n_expert              = 0
print_info: n_expert_used         = 0
print_info: n_expert_groups       = 0
print_info: n_group_used          = 0
print_info: causal attn           = 1
print_info: pooling type          = -1
print_info: rope type             = 2
print_info: rope scaling          = linear
print_info: freq_base_train       = 10000000.0
print_info: freq_scale_train      = 1
print_info: n_ctx_orig_yarn       = 524288
print_info: rope_yarn_log_mul     = 0.0000
print_info: rope_finetuned        = unknown
print_info: model type            = 36B
print_info: model params          = 36.15 B
print_info: general.name          = Seed OSS 36B Base woSyn
print_info: vocab type            = BPE
print_info: n_vocab               = 155136
print_info: n_merges              = 154737
print_info: BOS token             = 0 '<seed:bos>'
print_info: EOS token             = 2 '<seed:eos>'
print_info: PAD token             = 1 '<seed:pad>'
print_info: LF token              = 326 'Ċ'
print_info: EOG token             = 2 '<seed:eos>'
print_info: max token length      = 1024
load_tensors: loading model tensors, this can take a while... (load_mode = mmap)
load_tensors: layer   0 assigned to device CUDA0, is_swa = 0
load_tensors: layer   1 assigned to device CUDA0, is_swa = 0
load_tensors: layer   2 assigned to device CUDA0, is_swa = 0
load_tensors: layer   3 assigned to device CUDA0, is_swa = 0
load_tensors: layer   4 assigned to device CUDA0, is_swa = 0
load_tensors: layer   5 assigned to device CUDA0, is_swa = 0
load_tensors: layer   6 assigned to device CUDA0, is_swa = 0
load_tensors: layer   7 assigned to device CUDA0, is_swa = 0
load_tensors: layer   8 assigned to device CUDA0, is_swa = 0
load_tensors: layer   9 assigned to device CUDA0, is_swa = 0
load_tensors: layer  10 assigned to device CUDA0, is_swa = 0
load_tensors: layer  11 assigned to device CUDA0, is_swa = 0
load_tensors: layer  12 assigned to device CUDA0, is_swa = 0
load_tensors: layer  13 assigned to device CUDA0, is_swa = 0
load_tensors: layer  14 assigned to device CUDA0, is_swa = 0
load_tensors: layer  15 assigned to device CUDA0, is_swa = 0
load_tensors: layer  16 assigned to device CUDA0, is_swa = 0
load_tensors: layer  17 assigned to device CUDA0, is_swa = 0
load_tensors: layer  18 assigned to device CUDA0, is_swa = 0
load_tensors: layer  19 assigned to device CUDA0, is_swa = 0
load_tensors: layer  20 assigned to device CUDA0, is_swa = 0
load_tensors: layer  21 assigned to device CUDA0, is_swa = 0
load_tensors: layer  22 assigned to device CUDA0, is_swa = 0
load_tensors: layer  23 assigned to device CUDA0, is_swa = 0
load_tensors: layer  24 assigned to device CUDA0, is_swa = 0
load_tensors: layer  25 assigned to device CUDA0, is_swa = 0
load_tensors: layer  26 assigned to device CUDA0, is_swa = 0
load_tensors: layer  27 assigned to device CUDA0, is_swa = 0
load_tensors: layer  28 assigned to device CUDA0, is_swa = 0
load_tensors: layer  29 assigned to device CUDA0, is_swa = 0
load_tensors: layer  30 assigned to device CUDA0, is_swa = 0
load_tensors: layer  31 assigned to device CUDA0, is_swa = 0
load_tensors: layer  32 assigned to device CUDA0, is_swa = 0
load_tensors: layer  33 assigned to device CUDA0, is_swa = 0
load_tensors: layer  34 assigned to device CUDA0, is_swa = 0
load_tensors: layer  35 assigned to device CUDA0, is_swa = 0
load_tensors: layer  36 assigned to device CUDA0, is_swa = 0
load_tensors: layer  37 assigned to device CUDA0, is_swa = 0
load_tensors: layer  38 assigned to device CUDA0, is_swa = 0
load_tensors: layer  39 assigned to device CUDA0, is_swa = 0
load_tensors: layer  40 assigned to device CUDA0, is_swa = 0
load_tensors: layer  41 assigned to device CUDA0, is_swa = 0
load_tensors: layer  42 assigned to device CUDA0, is_swa = 0
load_tensors: layer  43 assigned to device CUDA0, is_swa = 0
load_tensors: layer  44 assigned to device CUDA0, is_swa = 0
load_tensors: layer  45 assigned to device CUDA0, is_swa = 0
load_tensors: layer  46 assigned to device CUDA0, is_swa = 0
load_tensors: layer  47 assigned to device CUDA0, is_swa = 0
load_tensors: layer  48 assigned to device CUDA0, is_swa = 0
load_tensors: layer  49 assigned to device CUDA0, is_swa = 0
load_tensors: layer  50 assigned to device CUDA0, is_swa = 0
load_tensors: layer  51 assigned to device CUDA0, is_swa = 0
load_tensors: layer  52 assigned to device CUDA0, is_swa = 0
load_tensors: layer  53 assigned to device CUDA0, is_swa = 0
load_tensors: layer  54 assigned to device CUDA0, is_swa = 0
load_tensors: layer  55 assigned to device CUDA0, is_swa = 0
load_tensors: layer  56 assigned to device CUDA0, is_swa = 0
load_tensors: layer  57 assigned to device CUDA0, is_swa = 0
load_tensors: layer  58 assigned to device CUDA0, is_swa = 0
load_tensors: layer  59 assigned to device CUDA0, is_swa = 0
load_tensors: layer  60 assigned to device CUDA0, is_swa = 0
load_tensors: layer  61 assigned to device CUDA0, is_swa = 0
load_tensors: layer  62 assigned to device CUDA0, is_swa = 0
load_tensors: layer  63 assigned to device CUDA0, is_swa = 0
load_tensors: layer  64 assigned to device CUDA0, is_swa = 0
create_tensor: loading tensor token_embd.weight
tensor token_embd.weight (426 MiB q4_K) buffer type overridden to CUDA0
create_tensor: loading tensor output_norm.weight
create_tensor: loading tensor output.weight
create_tensor: loading tensor blk.0.attn_qkv.weight
create_tensor: loading tensor blk.0.attn_q.weight
create_tensor: loading tensor blk.0.attn_k.weight
create_tensor: loading tensor blk.0.attn_v.weight
create_tensor: loading tensor blk.0.attn_q.bias
create_tensor: loading tensor blk.0.attn_k.bias
create_tensor: loading tensor blk.0.attn_v.bias
create_tensor: loading tensor blk.0.attn_output.weight
create_tensor: loading tensor blk.0.attn_norm.weight
create_tensor: loading tensor blk.0.post_attention_norm.weight
create_tensor: loading tensor blk.0.ffn_gate.weight
create_tensor: loading tensor blk.0.ffn_up.weight
create_tensor: loading tensor blk.0.ffn_down.weight
create_tensor: loading tensor blk.1.attn_qkv.weight
create_tensor: loading tensor blk.1.attn_q.weight
create_tensor: loading tensor blk.1.attn_k.weight
create_tensor: loading tensor blk.1.attn_v.weight
create_tensor: loading tensor blk.1.attn_q.bias
create_tensor: loading tensor blk.1.attn_k.bias
create_tensor: loading tensor blk.1.attn_v.bias
create_tensor: loading tensor blk.1.attn_output.weight
create_tensor: loading tensor blk.1.attn_norm.weight
create_tensor: loading tensor blk.1.post_attention_norm.weight
create_tensor: loading tensor blk.1.ffn_gate.weight
create_tensor: loading tensor blk.1.ffn_up.weight
create_tensor: loading tensor blk.1.ffn_down.weight
create_tensor: loading tensor blk.2.attn_qkv.weight
create_tensor: loading tensor blk.2.attn_q.weight
create_tensor: loading tensor blk.2.attn_k.weight
create_tensor: loading tensor blk.2.attn_v.weight
create_tensor: loading tensor blk.2.attn_q.bias
create_tensor: loading tensor blk.2.attn_k.bias
create_tensor: loading tensor blk.2.attn_v.bias
create_tensor: loading tensor blk.2.attn_output.weight
create_tensor: loading tensor blk.2.attn_norm.weight
create_tensor: loading tensor blk.2.post_attention_norm.weight
create_tensor: loading tensor blk.2.ffn_gate.weight
create_tensor: loading tensor blk.2.ffn_up.weight
create_tensor: loading tensor blk.2.ffn_down.weight
create_tensor: loading tensor blk.3.attn_qkv.weight
create_tensor: loading tensor blk.3.attn_q.weight
create_tensor: loading tensor blk.3.attn_k.weight
create_tensor: loading tensor blk.3.attn_v.weight
create_tensor: loading tensor blk.3.attn_q.bias
create_tensor: loading tensor blk.3.attn_k.bias
create_tensor: loading tensor blk.3.attn_v.bias
create_tensor: loading tensor blk.3.attn_output.weight
create_tensor: loading tensor blk.3.attn_norm.weight
create_tensor: loading tensor blk.3.post_attention_norm.weight
create_tensor: loading tensor blk.3.ffn_gate.weight
create_tensor: loading tensor blk.3.ffn_up.weight
create_tensor: loading tensor blk.3.ffn_down.weight
create_tensor: loading tensor blk.4.attn_qkv.weight
create_tensor: loading tensor blk.4.attn_q.weight
create_tensor: loading tensor blk.4.attn_k.weight
create_tensor: loading tensor blk.4.attn_v.weight
create_tensor: loading tensor blk.4.attn_q.bias
create_tensor: loading tensor blk.4.attn_k.bias
create_tensor: loading tensor blk.4.attn_v.bias
create_tensor: loading tensor blk.4.attn_output.weight
create_tensor: loading tensor blk.4.attn_norm.weight
create_tensor: loading tensor blk.4.post_attention_norm.weight
create_tensor: loading tensor blk.4.ffn_gate.weight
create_tensor: loading tensor blk.4.ffn_up.weight
create_tensor: loading tensor blk.4.ffn_down.weight
create_tensor: loading tensor blk.5.attn_qkv.weight
create_tensor: loading tensor blk.5.attn_q.weight
create_tensor: loading tensor blk.5.attn_k.weight
create_tensor: loading tensor blk.5.attn_v.weight
create_tensor: loading tensor blk.5.attn_q.bias
create_tensor: loading tensor blk.5.attn_k.bias
create_tensor: loading tensor blk.5.attn_v.bias
create_tensor: loading tensor blk.5.attn_output.weight
create_tensor: loading tensor blk.5.attn_norm.weight
create_tensor: loading tensor blk.5.post_attention_norm.weight
create_tensor: loading tensor blk.5.ffn_gate.weight
create_tensor: loading tensor blk.5.ffn_up.weight
create_tensor: loading tensor blk.5.ffn_down.weight
create_tensor: loading tensor blk.6.attn_qkv.weight
create_tensor: loading tensor blk.6.attn_q.weight
create_tensor: loading tensor blk.6.attn_k.weight
create_tensor: loading tensor blk.6.attn_v.weight
create_tensor: loading tensor blk.6.attn_q.bias
create_tensor: loading tensor blk.6.attn_k.bias
create_tensor: loading tensor blk.6.attn_v.bias
create_tensor: loading tensor blk.6.attn_output.weight
create_tensor: loading tensor blk.6.attn_norm.weight
create_tensor: loading tensor blk.6.post_attention_norm.weight
create_tensor: loading tensor blk.6.ffn_gate.weight
create_tensor: loading tensor blk.6.ffn_up.weight
create_tensor: loading tensor blk.6.ffn_down.weight
create_tensor: loading tensor blk.7.attn_qkv.weight
create_tensor: loading tensor blk.7.attn_q.weight
create_tensor: loading tensor blk.7.attn_k.weight
create_tensor: loading tensor blk.7.attn_v.weight
create_tensor: loading tensor blk.7.attn_q.bias
create_tensor: loading tensor blk.7.attn_k.bias
create_tensor: loading tensor blk.7.attn_v.bias
create_tensor: loading tensor blk.7.attn_output.weight
create_tensor: loading tensor blk.7.attn_norm.weight
create_tensor: loading tensor blk.7.post_attention_norm.weight
create_tensor: loading tensor blk.7.ffn_gate.weight
create_tensor: loading tensor blk.7.ffn_up.weight
create_tensor: loading tensor blk.7.ffn_down.weight
create_tensor: loading tensor blk.8.attn_qkv.weight
create_tensor: loading tensor blk.8.attn_q.weight
create_tensor: loading tensor blk.8.attn_k.weight
create_tensor: loading tensor blk.8.attn_v.weight
create_tensor: loading tensor blk.8.attn_q.bias
create_tensor: loading tensor blk.8.attn_k.bias
create_tensor: loading tensor blk.8.attn_v.bias
create_tensor: loading tensor blk.8.attn_output.weight
create_tensor: loading tensor blk.8.attn_norm.weight
create_tensor: loading tensor blk.8.post_attention_norm.weight
create_tensor: loading tensor blk.8.ffn_gate.weight
create_tensor: loading tensor blk.8.ffn_up.weight
create_tensor: loading tensor blk.8.ffn_down.weight
create_tensor: loading tensor blk.9.attn_qkv.weight
create_tensor: loading tensor blk.9.attn_q.weight
create_tensor: loading tensor blk.9.attn_k.weight
create_tensor: loading tensor blk.9.attn_v.weight
create_tensor: loading tensor blk.9.attn_q.bias
create_tensor: loading tensor blk.9.attn_k.bias
create_tensor: loading tensor blk.9.attn_v.bias
create_tensor: loading tensor blk.9.attn_output.weight
create_tensor: loading tensor blk.9.attn_norm.weight
create_tensor: loading tensor blk.9.post_attention_norm.weight
create_tensor: loading tensor blk.9.ffn_gate.weight
create_tensor: loading tensor blk.9.ffn_up.weight
create_tensor: loading tensor blk.9.ffn_down.weight
create_tensor: loading tensor blk.10.attn_qkv.weight
create_tensor: loading tensor blk.10.attn_q.weight
create_tensor: loading tensor blk.10.attn_k.weight
create_tensor: loading tensor blk.10.attn_v.weight
create_tensor: loading tensor blk.10.attn_q.bias
create_tensor: loading tensor blk.10.attn_k.bias
create_tensor: loading tensor blk.10.attn_v.bias
create_tensor: loading tensor blk.10.attn_output.weight
create_tensor: loading tensor blk.10.attn_norm.weight
create_tensor: loading tensor blk.10.post_attention_norm.weight
create_tensor: loading tensor blk.10.ffn_gate.weight
create_tensor: loading tensor blk.10.ffn_up.weight
create_tensor: loading tensor blk.10.ffn_down.weight
create_tensor: loading tensor blk.11.attn_qkv.weight
create_tensor: loading tensor blk.11.attn_q.weight
create_tensor: loading tensor blk.11.attn_k.weight
create_tensor: loading tensor blk.11.attn_v.weight
create_tensor: loading tensor blk.11.attn_q.bias
create_tensor: loading tensor blk.11.attn_k.bias
create_tensor: loading tensor blk.11.attn_v.bias
create_tensor: loading tensor blk.11.attn_output.weight
create_tensor: loading tensor blk.11.attn_norm.weight
create_tensor: loading tensor blk.11.post_attention_norm.weight
create_tensor: loading tensor blk.11.ffn_gate.weight
create_tensor: loading tensor blk.11.ffn_up.weight
create_tensor: loading tensor blk.11.ffn_down.weight
create_tensor: loading tensor blk.12.attn_qkv.weight
create_tensor: loading tensor blk.12.attn_q.weight
create_tensor: loading tensor blk.12.attn_k.weight
create_tensor: loading tensor blk.12.attn_v.weight
create_tensor: loading tensor blk.12.attn_q.bias
create_tensor: loading tensor blk.12.attn_k.bias
create_tensor: loading tensor blk.12.attn_v.bias
create_tensor: loading tensor blk.12.attn_output.weight
create_tensor: loading tensor blk.12.attn_norm.weight
create_tensor: loading tensor blk.12.post_attention_norm.weight
create_tensor: loading tensor blk.12.ffn_gate.weight
create_tensor: loading tensor blk.12.ffn_up.weight
create_tensor: loading tensor blk.12.ffn_down.weight
create_tensor: loading tensor blk.13.attn_qkv.weight
create_tensor: loading tensor blk.13.attn_q.weight
create_tensor: loading tensor blk.13.attn_k.weight
create_tensor: loading tensor blk.13.attn_v.weight
create_tensor: loading tensor blk.13.attn_q.bias
create_tensor: loading tensor blk.13.attn_k.bias
create_tensor: loading tensor blk.13.attn_v.bias
create_tensor: loading tensor blk.13.attn_output.weight
create_tensor: loading tensor blk.13.attn_norm.weight
create_tensor: loading tensor blk.13.post_attention_norm.weight
create_tensor: loading tensor blk.13.ffn_gate.weight
create_tensor: loading tensor blk.13.ffn_up.weight
create_tensor: loading tensor blk.13.ffn_down.weight
create_tensor: loading tensor blk.14.attn_qkv.weight
create_tensor: loading tensor blk.14.attn_q.weight
create_tensor: loading tensor blk.14.attn_k.weight
create_tensor: loading tensor blk.14.attn_v.weight
create_tensor: loading tensor blk.14.attn_q.bias
create_tensor: loading tensor blk.14.attn_k.bias
create_tensor: loading tensor blk.14.attn_v.bias
create_tensor: loading tensor blk.14.attn_output.weight
create_tensor: loading tensor blk.14.attn_norm.weight
create_tensor: loading tensor blk.14.post_attention_norm.weight
create_tensor: loading tensor blk.14.ffn_gate.weight
create_tensor: loading tensor blk.14.ffn_up.weight
create_tensor: loading tensor blk.14.ffn_down.weight
create_tensor: loading tensor blk.15.attn_qkv.weight
create_tensor: loading tensor blk.15.attn_q.weight
create_tensor: loading tensor blk.15.attn_k.weight
create_tensor: loading tensor blk.15.attn_v.weight
create_tensor: loading tensor blk.15.attn_q.bias
create_tensor: loading tensor blk.15.attn_k.bias
create_tensor: loading tensor blk.15.attn_v.bias
create_tensor: loading tensor blk.15.attn_output.weight
create_tensor: loading tensor blk.15.attn_norm.weight
create_tensor: loading tensor blk.15.post_attention_norm.weight
create_tensor: loading tensor blk.15.ffn_gate.weight
create_tensor: loading tensor blk.15.ffn_up.weight
create_tensor: loading tensor blk.15.ffn_down.weight
create_tensor: loading tensor blk.16.attn_qkv.weight
create_tensor: loading tensor blk.16.attn_q.weight
create_tensor: loading tensor blk.16.attn_k.weight
create_tensor: loading tensor blk.16.attn_v.weight
create_tensor: loading tensor blk.16.attn_q.bias
create_tensor: loading tensor blk.16.attn_k.bias
create_tensor: loading tensor blk.16.attn_v.bias
create_tensor: loading tensor blk.16.attn_output.weight
create_tensor: loading tensor blk.16.attn_norm.weight
create_tensor: loading tensor blk.16.post_attention_norm.weight
create_tensor: loading tensor blk.16.ffn_gate.weight
create_tensor: loading tensor blk.16.ffn_up.weight
create_tensor: loading tensor blk.16.ffn_down.weight
create_tensor: loading tensor blk.17.attn_qkv.weight
create_tensor: loading tensor blk.17.attn_q.weight
create_tensor: loading tensor blk.17.attn_k.weight
create_tensor: loading tensor blk.17.attn_v.weight
create_tensor: loading tensor blk.17.attn_q.bias
create_tensor: loading tensor blk.17.attn_k.bias
create_tensor: loading tensor blk.17.attn_v.bias
create_tensor: loading tensor blk.17.attn_output.weight
create_tensor: loading tensor blk.17.attn_norm.weight
create_tensor: loading tensor blk.17.post_attention_norm.weight
create_tensor: loading tensor blk.17.ffn_gate.weight
create_tensor: loading tensor blk.17.ffn_up.weight
create_tensor: loading tensor blk.17.ffn_down.weight
create_tensor: loading tensor blk.18.attn_qkv.weight
create_tensor: loading tensor blk.18.attn_q.weight
create_tensor: loading tensor blk.18.attn_k.weight
create_tensor: loading tensor blk.18.attn_v.weight
create_tensor: loading tensor blk.18.attn_q.bias
create_tensor: loading tensor blk.18.attn_k.bias
create_tensor: loading tensor blk.18.attn_v.bias
create_tensor: loading tensor blk.18.attn_output.weight
create_tensor: loading tensor blk.18.attn_norm.weight
create_tensor: loading tensor blk.18.post_attention_norm.weight
create_tensor: loading tensor blk.18.ffn_gate.weight
create_tensor: loading tensor blk.18.ffn_up.weight
create_tensor: loading tensor blk.18.ffn_down.weight
create_tensor: loading tensor blk.19.attn_qkv.weight
create_tensor: loading tensor blk.19.attn_q.weight
create_tensor: loading tensor blk.19.attn_k.weight
create_tensor: loading tensor blk.19.attn_v.weight
create_tensor: loading tensor blk.19.attn_q.bias
create_tensor: loading tensor blk.19.attn_k.bias
create_tensor: loading tensor blk.19.attn_v.bias
create_tensor: loading tensor blk.19.attn_output.weight
create_tensor: loading tensor blk.19.attn_norm.weight
create_tensor: loading tensor blk.19.post_attention_norm.weight
create_tensor: loading tensor blk.19.ffn_gate.weight
create_tensor: loading tensor blk.19.ffn_up.weight
create_tensor: loading tensor blk.19.ffn_down.weight
create_tensor: loading tensor blk.20.attn_qkv.weight
create_tensor: loading tensor blk.20.attn_q.weight
create_tensor: loading tensor blk.20.attn_k.weight
create_tensor: loading tensor blk.20.attn_v.weight
create_tensor: loading tensor blk.20.attn_q.bias
create_tensor: loading tensor blk.20.attn_k.bias
create_tensor: loading tensor blk.20.attn_v.bias
create_tensor: loading tensor blk.20.attn_output.weight
create_tensor: loading tensor blk.20.attn_norm.weight
create_tensor: loading tensor blk.20.post_attention_norm.weight
create_tensor: loading tensor blk.20.ffn_gate.weight
create_tensor: loading tensor blk.20.ffn_up.weight
create_tensor: loading tensor blk.20.ffn_down.weight
create_tensor: loading tensor blk.21.attn_qkv.weight
create_tensor: loading tensor blk.21.attn_q.weight
create_tensor: loading tensor blk.21.attn_k.weight
create_tensor: loading tensor blk.21.attn_v.weight
create_tensor: loading tensor blk.21.attn_q.bias
create_tensor: loading tensor blk.21.attn_k.bias
create_tensor: loading tensor blk.21.attn_v.bias
create_tensor: loading tensor blk.21.attn_output.weight
create_tensor: loading tensor blk.21.attn_norm.weight
create_tensor: loading tensor blk.21.post_attention_norm.weight
create_tensor: loading tensor blk.21.ffn_gate.weight
create_tensor: loading tensor blk.21.ffn_up.weight
create_tensor: loading tensor blk.21.ffn_down.weight
create_tensor: loading tensor blk.22.attn_qkv.weight
create_tensor: loading tensor blk.22.attn_q.weight
create_tensor: loading tensor blk.22.attn_k.weight
create_tensor: loading tensor blk.22.attn_v.weight
create_tensor: loading tensor blk.22.attn_q.bias
create_tensor: loading tensor blk.22.attn_k.bias
create_tensor: loading tensor blk.22.attn_v.bias
create_tensor: loading tensor blk.22.attn_output.weight
create_tensor: loading tensor blk.22.attn_norm.weight
create_tensor: loading tensor blk.22.post_attention_norm.weight
create_tensor: loading tensor blk.22.ffn_gate.weight
create_tensor: loading tensor blk.22.ffn_up.weight
create_tensor: loading tensor blk.22.ffn_down.weight
create_tensor: loading tensor blk.23.attn_qkv.weight
create_tensor: loading tensor blk.23.attn_q.weight
create_tensor: loading tensor blk.23.attn_k.weight
create_tensor: loading tensor blk.23.attn_v.weight
create_tensor: loading tensor blk.23.attn_q.bias
create_tensor: loading tensor blk.23.attn_k.bias
create_tensor: loading tensor blk.23.attn_v.bias
create_tensor: loading tensor blk.23.attn_output.weight
create_tensor: loading tensor blk.23.attn_norm.weight
create_tensor: loading tensor blk.23.post_attention_norm.weight
create_tensor: loading tensor blk.23.ffn_gate.weight
create_tensor: loading tensor blk.23.ffn_up.weight
create_tensor: loading tensor blk.23.ffn_down.weight
create_tensor: loading tensor blk.24.attn_qkv.weight
create_tensor: loading tensor blk.24.attn_q.weight
create_tensor: loading tensor blk.24.attn_k.weight
create_tensor: loading tensor blk.24.attn_v.weight
create_tensor: loading tensor blk.24.attn_q.bias
create_tensor: loading tensor blk.24.attn_k.bias
create_tensor: loading tensor blk.24.attn_v.bias
create_tensor: loading tensor blk.24.attn_output.weight
create_tensor: loading tensor blk.24.attn_norm.weight
create_tensor: loading tensor blk.24.post_attention_norm.weight
create_tensor: loading tensor blk.24.ffn_gate.weight
create_tensor: loading tensor blk.24.ffn_up.weight
create_tensor: loading tensor blk.24.ffn_down.weight
create_tensor: loading tensor blk.25.attn_qkv.weight
create_tensor: loading tensor blk.25.attn_q.weight
create_tensor: loading tensor blk.25.attn_k.weight
create_tensor: loading tensor blk.25.attn_v.weight
create_tensor: loading tensor blk.25.attn_q.bias
create_tensor: loading tensor blk.25.attn_k.bias
create_tensor: loading tensor blk.25.attn_v.bias
create_tensor: loading tensor blk.25.attn_output.weight
create_tensor: loading tensor blk.25.attn_norm.weight
create_tensor: loading tensor blk.25.post_attention_norm.weight
create_tensor: loading tensor blk.25.ffn_gate.weight
create_tensor: loading tensor blk.25.ffn_up.weight
create_tensor: loading tensor blk.25.ffn_down.weight
create_tensor: loading tensor blk.26.attn_qkv.weight
create_tensor: loading tensor blk.26.attn_q.weight
create_tensor: loading tensor blk.26.attn_k.weight
create_tensor: loading tensor blk.26.attn_v.weight
create_tensor: loading tensor blk.26.attn_q.bias
create_tensor: loading tensor blk.26.attn_k.bias
create_tensor: loading tensor blk.26.attn_v.bias
create_tensor: loading tensor blk.26.attn_output.weight
create_tensor: loading tensor blk.26.attn_norm.weight
create_tensor: loading tensor blk.26.post_attention_norm.weight
create_tensor: loading tensor blk.26.ffn_gate.weight
create_tensor: loading tensor blk.26.ffn_up.weight
create_tensor: loading tensor blk.26.ffn_down.weight
create_tensor: loading tensor blk.27.attn_qkv.weight
create_tensor: loading tensor blk.27.attn_q.weight
create_tensor: loading tensor blk.27.attn_k.weight
create_tensor: loading tensor blk.27.attn_v.weight
create_tensor: loading tensor blk.27.attn_q.bias
create_tensor: loading tensor blk.27.attn_k.bias
create_tensor: loading tensor blk.27.attn_v.bias
create_tensor: loading tensor blk.27.attn_output.weight
create_tensor: loading tensor blk.27.attn_norm.weight
create_tensor: loading tensor blk.27.post_attention_norm.weight
create_tensor: loading tensor blk.27.ffn_gate.weight
create_tensor: loading tensor blk.27.ffn_up.weight
create_tensor: loading tensor blk.27.ffn_down.weight
create_tensor: loading tensor blk.28.attn_qkv.weight
create_tensor: loading tensor blk.28.attn_q.weight
create_tensor: loading tensor blk.28.attn_k.weight
create_tensor: loading tensor blk.28.attn_v.weight
create_tensor: loading tensor blk.28.attn_q.bias
create_tensor: loading tensor blk.28.attn_k.bias
create_tensor: loading tensor blk.28.attn_v.bias
create_tensor: loading tensor blk.28.attn_output.weight
create_tensor: loading tensor blk.28.attn_norm.weight
create_tensor: loading tensor blk.28.post_attention_norm.weight
create_tensor: loading tensor blk.28.ffn_gate.weight
create_tensor: loading tensor blk.28.ffn_up.weight
create_tensor: loading tensor blk.28.ffn_down.weight
create_tensor: loading tensor blk.29.attn_qkv.weight
create_tensor: loading tensor blk.29.attn_q.weight
create_tensor: loading tensor blk.29.attn_k.weight
create_tensor: loading tensor blk.29.attn_v.weight
create_tensor: loading tensor blk.29.attn_q.bias
create_tensor: loading tensor blk.29.attn_k.bias
create_tensor: loading tensor blk.29.attn_v.bias
create_tensor: loading tensor blk.29.attn_output.weight
create_tensor: loading tensor blk.29.attn_norm.weight
create_tensor: loading tensor blk.29.post_attention_norm.weight
create_tensor: loading tensor blk.29.ffn_gate.weight
create_tensor: loading tensor blk.29.ffn_up.weight
create_tensor: loading tensor blk.29.ffn_down.weight
create_tensor: loading tensor blk.30.attn_qkv.weight
create_tensor: loading tensor blk.30.attn_q.weight
create_tensor: loading tensor blk.30.attn_k.weight
create_tensor: loading tensor blk.30.attn_v.weight
create_tensor: loading tensor blk.30.attn_q.bias
create_tensor: loading tensor blk.30.attn_k.bias
create_tensor: loading tensor blk.30.attn_v.bias
create_tensor: loading tensor blk.30.attn_output.weight
create_tensor: loading tensor blk.30.attn_norm.weight
create_tensor: loading tensor blk.30.post_attention_norm.weight
create_tensor: loading tensor blk.30.ffn_gate.weight
create_tensor: loading tensor blk.30.ffn_up.weight
create_tensor: loading tensor blk.30.ffn_down.weight
create_tensor: loading tensor blk.31.attn_qkv.weight
create_tensor: loading tensor blk.31.attn_q.weight
create_tensor: loading tensor blk.31.attn_k.weight
create_tensor: loading tensor blk.31.attn_v.weight
create_tensor: loading tensor blk.31.attn_q.bias
create_tensor: loading tensor blk.31.attn_k.bias
create_tensor: loading tensor blk.31.attn_v.bias
create_tensor: loading tensor blk.31.attn_output.weight
create_tensor: loading tensor blk.31.attn_norm.weight
create_tensor: loading tensor blk.31.post_attention_norm.weight
create_tensor: loading tensor blk.31.ffn_gate.weight
create_tensor: loading tensor blk.31.ffn_up.weight
create_tensor: loading tensor blk.31.ffn_down.weight
create_tensor: loading tensor blk.32.attn_qkv.weight
create_tensor: loading tensor blk.32.attn_q.weight
create_tensor: loading tensor blk.32.attn_k.weight
create_tensor: loading tensor blk.32.attn_v.weight
create_tensor: loading tensor blk.32.attn_q.bias
create_tensor: loading tensor blk.32.attn_k.bias
create_tensor: loading tensor blk.32.attn_v.bias
create_tensor: loading tensor blk.32.attn_output.weight
create_tensor: loading tensor blk.32.attn_norm.weight
create_tensor: loading tensor blk.32.post_attention_norm.weight
create_tensor: loading tensor blk.32.ffn_gate.weight
create_tensor: loading tensor blk.32.ffn_up.weight
create_tensor: loading tensor blk.32.ffn_down.weight
create_tensor: loading tensor blk.33.attn_qkv.weight
create_tensor: loading tensor blk.33.attn_q.weight
create_tensor: loading tensor blk.33.attn_k.weight
create_tensor: loading tensor blk.33.attn_v.weight
create_tensor: loading tensor blk.33.attn_q.bias
create_tensor: loading tensor blk.33.attn_k.bias
create_tensor: loading tensor blk.33.attn_v.bias
create_tensor: loading tensor blk.33.attn_output.weight
create_tensor: loading tensor blk.33.attn_norm.weight
create_tensor: loading tensor blk.33.post_attention_norm.weight
create_tensor: loading tensor blk.33.ffn_gate.weight
create_tensor: loading tensor blk.33.ffn_up.weight
create_tensor: loading tensor blk.33.ffn_down.weight
create_tensor: loading tensor blk.34.attn_qkv.weight
create_tensor: loading tensor blk.34.attn_q.weight
create_tensor: loading tensor blk.34.attn_k.weight
create_tensor: loading tensor blk.34.attn_v.weight
create_tensor: loading tensor blk.34.attn_q.bias
create_tensor: loading tensor blk.34.attn_k.bias
create_tensor: loading tensor blk.34.attn_v.bias
create_tensor: loading tensor blk.34.attn_output.weight
create_tensor: loading tensor blk.34.attn_norm.weight
create_tensor: loading tensor blk.34.post_attention_norm.weight
create_tensor: loading tensor blk.34.ffn_gate.weight
create_tensor: loading tensor blk.34.ffn_up.weight
create_tensor: loading tensor blk.34.ffn_down.weight
create_tensor: loading tensor blk.35.attn_qkv.weight
create_tensor: loading tensor blk.35.attn_q.weight
create_tensor: loading tensor blk.35.attn_k.weight
create_tensor: loading tensor blk.35.attn_v.weight
create_tensor: loading tensor blk.35.attn_q.bias
create_tensor: loading tensor blk.35.attn_k.bias
create_tensor: loading tensor blk.35.attn_v.bias
create_tensor: loading tensor blk.35.attn_output.weight
create_tensor: loading tensor blk.35.attn_norm.weight
create_tensor: loading tensor blk.35.post_attention_norm.weight
create_tensor: loading tensor blk.35.ffn_gate.weight
create_tensor: loading tensor blk.35.ffn_up.weight
create_tensor: loading tensor blk.35.ffn_down.weight
create_tensor: loading tensor blk.36.attn_qkv.weight
create_tensor: loading tensor blk.36.attn_q.weight
create_tensor: loading tensor blk.36.attn_k.weight
create_tensor: loading tensor blk.36.attn_v.weight
create_tensor: loading tensor blk.36.attn_q.bias
create_tensor: loading tensor blk.36.attn_k.bias
create_tensor: loading tensor blk.36.attn_v.bias
create_tensor: loading tensor blk.36.attn_output.weight
create_tensor: loading tensor blk.36.attn_norm.weight
create_tensor: loading tensor blk.36.post_attention_norm.weight
create_tensor: loading tensor blk.36.ffn_gate.weight
create_tensor: loading tensor blk.36.ffn_up.weight
create_tensor: loading tensor blk.36.ffn_down.weight
create_tensor: loading tensor blk.37.attn_qkv.weight
create_tensor: loading tensor blk.37.attn_q.weight
create_tensor: loading tensor blk.37.attn_k.weight
create_tensor: loading tensor blk.37.attn_v.weight
create_tensor: loading tensor blk.37.attn_q.bias
create_tensor: loading tensor blk.37.attn_k.bias
create_tensor: loading tensor blk.37.attn_v.bias
create_tensor: loading tensor blk.37.attn_output.weight
create_tensor: loading tensor blk.37.attn_norm.weight
create_tensor: loading tensor blk.37.post_attention_norm.weight
create_tensor: loading tensor blk.37.ffn_gate.weight
create_tensor: loading tensor blk.37.ffn_up.weight
create_tensor: loading tensor blk.37.ffn_down.weight
create_tensor: loading tensor blk.38.attn_qkv.weight
create_tensor: loading tensor blk.38.attn_q.weight
create_tensor: loading tensor blk.38.attn_k.weight
create_tensor: loading tensor blk.38.attn_v.weight
create_tensor: loading tensor blk.38.attn_q.bias
create_tensor: loading tensor blk.38.attn_k.bias
create_tensor: loading tensor blk.38.attn_v.bias
create_tensor: loading tensor blk.38.attn_output.weight
create_tensor: loading tensor blk.38.attn_norm.weight
create_tensor: loading tensor blk.38.post_attention_norm.weight
create_tensor: loading tensor blk.38.ffn_gate.weight
create_tensor: loading tensor blk.38.ffn_up.weight
create_tensor: loading tensor blk.38.ffn_down.weight
create_tensor: loading tensor blk.39.attn_qkv.weight
create_tensor: loading tensor blk.39.attn_q.weight
create_tensor: loading tensor blk.39.attn_k.weight
create_tensor: loading tensor blk.39.attn_v.weight
create_tensor: loading tensor blk.39.attn_q.bias
create_tensor: loading tensor blk.39.attn_k.bias
create_tensor: loading tensor blk.39.attn_v.bias
create_tensor: loading tensor blk.39.attn_output.weight
create_tensor: loading tensor blk.39.attn_norm.weight
create_tensor: loading tensor blk.39.post_attention_norm.weight
create_tensor: loading tensor blk.39.ffn_gate.weight
create_tensor: loading tensor blk.39.ffn_up.weight
create_tensor: loading tensor blk.39.ffn_down.weight
create_tensor: loading tensor blk.40.attn_qkv.weight
create_tensor: loading tensor blk.40.attn_q.weight
create_tensor: loading tensor blk.40.attn_k.weight
create_tensor: loading tensor blk.40.attn_v.weight
create_tensor: loading tensor blk.40.attn_q.bias
create_tensor: loading tensor blk.40.attn_k.bias
create_tensor: loading tensor blk.40.attn_v.bias
create_tensor: loading tensor blk.40.attn_output.weight
create_tensor: loading tensor blk.40.attn_norm.weight
create_tensor: loading tensor blk.40.post_attention_norm.weight
create_tensor: loading tensor blk.40.ffn_gate.weight
create_tensor: loading tensor blk.40.ffn_up.weight
create_tensor: loading tensor blk.40.ffn_down.weight
create_tensor: loading tensor blk.41.attn_qkv.weight
create_tensor: loading tensor blk.41.attn_q.weight
create_tensor: loading tensor blk.41.attn_k.weight
create_tensor: loading tensor blk.41.attn_v.weight
create_tensor: loading tensor blk.41.attn_q.bias
create_tensor: loading tensor blk.41.attn_k.bias
create_tensor: loading tensor blk.41.attn_v.bias
create_tensor: loading tensor blk.41.attn_output.weight
create_tensor: loading tensor blk.41.attn_norm.weight
create_tensor: loading tensor blk.41.post_attention_norm.weight
create_tensor: loading tensor blk.41.ffn_gate.weight
create_tensor: loading tensor blk.41.ffn_up.weight
create_tensor: loading tensor blk.41.ffn_down.weight
create_tensor: loading tensor blk.42.attn_qkv.weight
create_tensor: loading tensor blk.42.attn_q.weight
create_tensor: loading tensor blk.42.attn_k.weight
create_tensor: loading tensor blk.42.attn_v.weight
create_tensor: loading tensor blk.42.attn_q.bias
create_tensor: loading tensor blk.42.attn_k.bias
create_tensor: loading tensor blk.42.attn_v.bias
create_tensor: loading tensor blk.42.attn_output.weight
create_tensor: loading tensor blk.42.attn_norm.weight
create_tensor: loading tensor blk.42.post_attention_norm.weight
create_tensor: loading tensor blk.42.ffn_gate.weight
create_tensor: loading tensor blk.42.ffn_up.weight
create_tensor: loading tensor blk.42.ffn_down.weight
create_tensor: loading tensor blk.43.attn_qkv.weight
create_tensor: loading tensor blk.43.attn_q.weight
create_tensor: loading tensor blk.43.attn_k.weight
create_tensor: loading tensor blk.43.attn_v.weight
create_tensor: loading tensor blk.43.attn_q.bias
create_tensor: loading tensor blk.43.attn_k.bias
create_tensor: loading tensor blk.43.attn_v.bias
create_tensor: loading tensor blk.43.attn_output.weight
create_tensor: loading tensor blk.43.attn_norm.weight
create_tensor: loading tensor blk.43.post_attention_norm.weight
create_tensor: loading tensor blk.43.ffn_gate.weight
create_tensor: loading tensor blk.43.ffn_up.weight
create_tensor: loading tensor blk.43.ffn_down.weight
create_tensor: loading tensor blk.44.attn_qkv.weight
create_tensor: loading tensor blk.44.attn_q.weight
create_tensor: loading tensor blk.44.attn_k.weight
create_tensor: loading tensor blk.44.attn_v.weight
create_tensor: loading tensor blk.44.attn_q.bias
create_tensor: loading tensor blk.44.attn_k.bias
create_tensor: loading tensor blk.44.attn_v.bias
create_tensor: loading tensor blk.44.attn_output.weight
create_tensor: loading tensor blk.44.attn_norm.weight
create_tensor: loading tensor blk.44.post_attention_norm.weight
create_tensor: loading tensor blk.44.ffn_gate.weight
create_tensor: loading tensor blk.44.ffn_up.weight
create_tensor: loading tensor blk.44.ffn_down.weight
create_tensor: loading tensor blk.45.attn_qkv.weight
create_tensor: loading tensor blk.45.attn_q.weight
create_tensor: loading tensor blk.45.attn_k.weight
create_tensor: loading tensor blk.45.attn_v.weight
create_tensor: loading tensor blk.45.attn_q.bias
create_tensor: loading tensor blk.45.attn_k.bias
create_tensor: loading tensor blk.45.attn_v.bias
create_tensor: loading tensor blk.45.attn_output.weight
create_tensor: loading tensor blk.45.attn_norm.weight
create_tensor: loading tensor blk.45.post_attention_norm.weight
create_tensor: loading tensor blk.45.ffn_gate.weight
create_tensor: loading tensor blk.45.ffn_up.weight
create_tensor: loading tensor blk.45.ffn_down.weight
create_tensor: loading tensor blk.46.attn_qkv.weight
create_tensor: loading tensor blk.46.attn_q.weight
create_tensor: loading tensor blk.46.attn_k.weight
create_tensor: loading tensor blk.46.attn_v.weight
create_tensor: loading tensor blk.46.attn_q.bias
create_tensor: loading tensor blk.46.attn_k.bias
create_tensor: loading tensor blk.46.attn_v.bias
create_tensor: loading tensor blk.46.attn_output.weight
create_tensor: loading tensor blk.46.attn_norm.weight
create_tensor: loading tensor blk.46.post_attention_norm.weight
create_tensor: loading tensor blk.46.ffn_gate.weight
create_tensor: loading tensor blk.46.ffn_up.weight
create_tensor: loading tensor blk.46.ffn_down.weight
create_tensor: loading tensor blk.47.attn_qkv.weight
create_tensor: loading tensor blk.47.attn_q.weight
create_tensor: loading tensor blk.47.attn_k.weight
create_tensor: loading tensor blk.47.attn_v.weight
create_tensor: loading tensor blk.47.attn_q.bias
create_tensor: loading tensor blk.47.attn_k.bias
create_tensor: loading tensor blk.47.attn_v.bias
create_tensor: loading tensor blk.47.attn_output.weight
create_tensor: loading tensor blk.47.attn_norm.weight
create_tensor: loading tensor blk.47.post_attention_norm.weight
create_tensor: loading tensor blk.47.ffn_gate.weight
create_tensor: loading tensor blk.47.ffn_up.weight
create_tensor: loading tensor blk.47.ffn_down.weight
create_tensor: loading tensor blk.48.attn_qkv.weight
create_tensor: loading tensor blk.48.attn_q.weight
create_tensor: loading tensor blk.48.attn_k.weight
create_tensor: loading tensor blk.48.attn_v.weight
create_tensor: loading tensor blk.48.attn_q.bias
create_tensor: loading tensor blk.48.attn_k.bias
create_tensor: loading tensor blk.48.attn_v.bias
create_tensor: loading tensor blk.48.attn_output.weight
create_tensor: loading tensor blk.48.attn_norm.weight
create_tensor: loading tensor blk.48.post_attention_norm.weight
create_tensor: loading tensor blk.48.ffn_gate.weight
create_tensor: loading tensor blk.48.ffn_up.weight
create_tensor: loading tensor blk.48.ffn_down.weight
create_tensor: loading tensor blk.49.attn_qkv.weight
create_tensor: loading tensor blk.49.attn_q.weight
create_tensor: loading tensor blk.49.attn_k.weight
create_tensor: loading tensor blk.49.attn_v.weight
create_tensor: loading tensor blk.49.attn_q.bias
create_tensor: loading tensor blk.49.attn_k.bias
create_tensor: loading tensor blk.49.attn_v.bias
create_tensor: loading tensor blk.49.attn_output.weight
create_tensor: loading tensor blk.49.attn_norm.weight
create_tensor: loading tensor blk.49.post_attention_norm.weight
create_tensor: loading tensor blk.49.ffn_gate.weight
create_tensor: loading tensor blk.49.ffn_up.weight
create_tensor: loading tensor blk.49.ffn_down.weight
create_tensor: loading tensor blk.50.attn_qkv.weight
create_tensor: loading tensor blk.50.attn_q.weight
create_tensor: loading tensor blk.50.attn_k.weight
create_tensor: loading tensor blk.50.attn_v.weight
create_tensor: loading tensor blk.50.attn_q.bias
create_tensor: loading tensor blk.50.attn_k.bias
create_tensor: loading tensor blk.50.attn_v.bias
create_tensor: loading tensor blk.50.attn_output.weight
create_tensor: loading tensor blk.50.attn_norm.weight
create_tensor: loading tensor blk.50.post_attention_norm.weight
create_tensor: loading tensor blk.50.ffn_gate.weight
create_tensor: loading tensor blk.50.ffn_up.weight
create_tensor: loading tensor blk.50.ffn_down.weight
create_tensor: loading tensor blk.51.attn_qkv.weight
create_tensor: loading tensor blk.51.attn_q.weight
create_tensor: loading tensor blk.51.attn_k.weight
create_tensor: loading tensor blk.51.attn_v.weight
create_tensor: loading tensor blk.51.attn_q.bias
create_tensor: loading tensor blk.51.attn_k.bias
create_tensor: loading tensor blk.51.attn_v.bias
create_tensor: loading tensor blk.51.attn_output.weight
create_tensor: loading tensor blk.51.attn_norm.weight
create_tensor: loading tensor blk.51.post_attention_norm.weight
create_tensor: loading tensor blk.51.ffn_gate.weight
create_tensor: loading tensor blk.51.ffn_up.weight
create_tensor: loading tensor blk.51.ffn_down.weight
create_tensor: loading tensor blk.52.attn_qkv.weight
create_tensor: loading tensor blk.52.attn_q.weight
create_tensor: loading tensor blk.52.attn_k.weight
create_tensor: loading tensor blk.52.attn_v.weight
create_tensor: loading tensor blk.52.attn_q.bias
create_tensor: loading tensor blk.52.attn_k.bias
create_tensor: loading tensor blk.52.attn_v.bias
create_tensor: loading tensor blk.52.attn_output.weight
create_tensor: loading tensor blk.52.attn_norm.weight
create_tensor: loading tensor blk.52.post_attention_norm.weight
create_tensor: loading tensor blk.52.ffn_gate.weight
create_tensor: loading tensor blk.52.ffn_up.weight
create_tensor: loading tensor blk.52.ffn_down.weight
create_tensor: loading tensor blk.53.attn_qkv.weight
create_tensor: loading tensor blk.53.attn_q.weight
create_tensor: loading tensor blk.53.attn_k.weight
create_tensor: loading tensor blk.53.attn_v.weight
create_tensor: loading tensor blk.53.attn_q.bias
create_tensor: loading tensor blk.53.attn_k.bias
create_tensor: loading tensor blk.53.attn_v.bias
create_tensor: loading tensor blk.53.attn_output.weight
create_tensor: loading tensor blk.53.attn_norm.weight
create_tensor: loading tensor blk.53.post_attention_norm.weight
create_tensor: loading tensor blk.53.ffn_gate.weight
create_tensor: loading tensor blk.53.ffn_up.weight
create_tensor: loading tensor blk.53.ffn_down.weight
create_tensor: loading tensor blk.54.attn_qkv.weight
create_tensor: loading tensor blk.54.attn_q.weight
create_tensor: loading tensor blk.54.attn_k.weight
create_tensor: loading tensor blk.54.attn_v.weight
create_tensor: loading tensor blk.54.attn_q.bias
create_tensor: loading tensor blk.54.attn_k.bias
create_tensor: loading tensor blk.54.attn_v.bias
create_tensor: loading tensor blk.54.attn_output.weight
create_tensor: loading tensor blk.54.attn_norm.weight
create_tensor: loading tensor blk.54.post_attention_norm.weight
create_tensor: loading tensor blk.54.ffn_gate.weight
create_tensor: loading tensor blk.54.ffn_up.weight
create_tensor: loading tensor blk.54.ffn_down.weight
create_tensor: loading tensor blk.55.attn_qkv.weight
create_tensor: loading tensor blk.55.attn_q.weight
create_tensor: loading tensor blk.55.attn_k.weight
create_tensor: loading tensor blk.55.attn_v.weight
create_tensor: loading tensor blk.55.attn_q.bias
create_tensor: loading tensor blk.55.attn_k.bias
create_tensor: loading tensor blk.55.attn_v.bias
create_tensor: loading tensor blk.55.attn_output.weight
create_tensor: loading tensor blk.55.attn_norm.weight
create_tensor: loading tensor blk.55.post_attention_norm.weight
create_tensor: loading tensor blk.55.ffn_gate.weight
create_tensor: loading tensor blk.55.ffn_up.weight
create_tensor: loading tensor blk.55.ffn_down.weight
create_tensor: loading tensor blk.56.attn_qkv.weight
create_tensor: loading tensor blk.56.attn_q.weight
create_tensor: loading tensor blk.56.attn_k.weight
create_tensor: loading tensor blk.56.attn_v.weight
create_tensor: loading tensor blk.56.attn_q.bias
create_tensor: loading tensor blk.56.attn_k.bias
create_tensor: loading tensor blk.56.attn_v.bias
create_tensor: loading tensor blk.56.attn_output.weight
create_tensor: loading tensor blk.56.attn_norm.weight
create_tensor: loading tensor blk.56.post_attention_norm.weight
create_tensor: loading tensor blk.56.ffn_gate.weight
create_tensor: loading tensor blk.56.ffn_up.weight
create_tensor: loading tensor blk.56.ffn_down.weight
create_tensor: loading tensor blk.57.attn_qkv.weight
create_tensor: loading tensor blk.57.attn_q.weight
create_tensor: loading tensor blk.57.attn_k.weight
create_tensor: loading tensor blk.57.attn_v.weight
create_tensor: loading tensor blk.57.attn_q.bias
create_tensor: loading tensor blk.57.attn_k.bias
create_tensor: loading tensor blk.57.attn_v.bias
create_tensor: loading tensor blk.57.attn_output.weight
create_tensor: loading tensor blk.57.attn_norm.weight
create_tensor: loading tensor blk.57.post_attention_norm.weight
create_tensor: loading tensor blk.57.ffn_gate.weight
create_tensor: loading tensor blk.57.ffn_up.weight
create_tensor: loading tensor blk.57.ffn_down.weight
create_tensor: loading tensor blk.58.attn_qkv.weight
create_tensor: loading tensor blk.58.attn_q.weight
create_tensor: loading tensor blk.58.attn_k.weight
create_tensor: loading tensor blk.58.attn_v.weight
create_tensor: loading tensor blk.58.attn_q.bias
create_tensor: loading tensor blk.58.attn_k.bias
create_tensor: loading tensor blk.58.attn_v.bias
create_tensor: loading tensor blk.58.attn_output.weight
create_tensor: loading tensor blk.58.attn_norm.weight
create_tensor: loading tensor blk.58.post_attention_norm.weight
create_tensor: loading tensor blk.58.ffn_gate.weight
create_tensor: loading tensor blk.58.ffn_up.weight
create_tensor: loading tensor blk.58.ffn_down.weight
create_tensor: loading tensor blk.59.attn_qkv.weight
create_tensor: loading tensor blk.59.attn_q.weight
create_tensor: loading tensor blk.59.attn_k.weight
create_tensor: loading tensor blk.59.attn_v.weight
create_tensor: loading tensor blk.59.attn_q.bias
create_tensor: loading tensor blk.59.attn_k.bias
create_tensor: loading tensor blk.59.attn_v.bias
create_tensor: loading tensor blk.59.attn_output.weight
create_tensor: loading tensor blk.59.attn_norm.weight
create_tensor: loading tensor blk.59.post_attention_norm.weight
create_tensor: loading tensor blk.59.ffn_gate.weight
create_tensor: loading tensor blk.59.ffn_up.weight
create_tensor: loading tensor blk.59.ffn_down.weight
create_tensor: loading tensor blk.60.attn_qkv.weight
create_tensor: loading tensor blk.60.attn_q.weight
create_tensor: loading tensor blk.60.attn_k.weight
create_tensor: loading tensor blk.60.attn_v.weight
create_tensor: loading tensor blk.60.attn_q.bias
create_tensor: loading tensor blk.60.attn_k.bias
create_tensor: loading tensor blk.60.attn_v.bias
create_tensor: loading tensor blk.60.attn_output.weight
create_tensor: loading tensor blk.60.attn_norm.weight
create_tensor: loading tensor blk.60.post_attention_norm.weight
create_tensor: loading tensor blk.60.ffn_gate.weight
create_tensor: loading tensor blk.60.ffn_up.weight
create_tensor: loading tensor blk.60.ffn_down.weight
create_tensor: loading tensor blk.61.attn_qkv.weight
create_tensor: loading tensor blk.61.attn_q.weight
create_tensor: loading tensor blk.61.attn_k.weight
create_tensor: loading tensor blk.61.attn_v.weight
create_tensor: loading tensor blk.61.attn_q.bias
create_tensor: loading tensor blk.61.attn_k.bias
create_tensor: loading tensor blk.61.attn_v.bias
create_tensor: loading tensor blk.61.attn_output.weight
create_tensor: loading tensor blk.61.attn_norm.weight
create_tensor: loading tensor blk.61.post_attention_norm.weight
create_tensor: loading tensor blk.61.ffn_gate.weight
create_tensor: loading tensor blk.61.ffn_up.weight
create_tensor: loading tensor blk.61.ffn_down.weight
create_tensor: loading tensor blk.62.attn_qkv.weight
create_tensor: loading tensor blk.62.attn_q.weight
create_tensor: loading tensor blk.62.attn_k.weight
create_tensor: loading tensor blk.62.attn_v.weight
create_tensor: loading tensor blk.62.attn_q.bias
create_tensor: loading tensor blk.62.attn_k.bias
create_tensor: loading tensor blk.62.attn_v.bias
create_tensor: loading tensor blk.62.attn_output.weight
create_tensor: loading tensor blk.62.attn_norm.weight
create_tensor: loading tensor blk.62.post_attention_norm.weight
create_tensor: loading tensor blk.62.ffn_gate.weight
create_tensor: loading tensor blk.62.ffn_up.weight
create_tensor: loading tensor blk.62.ffn_down.weight
create_tensor: loading tensor blk.63.attn_qkv.weight
create_tensor: loading tensor blk.63.attn_q.weight
create_tensor: loading tensor blk.63.attn_k.weight
create_tensor: loading tensor blk.63.attn_v.weight
create_tensor: loading tensor blk.63.attn_q.bias
create_tensor: loading tensor blk.63.attn_k.bias
create_tensor: loading tensor blk.63.attn_v.bias
create_tensor: loading tensor blk.63.attn_output.weight
create_tensor: loading tensor blk.63.attn_norm.weight
create_tensor: loading tensor blk.63.post_attention_norm.weight
create_tensor: loading tensor blk.63.ffn_gate.weight
create_tensor: loading tensor blk.63.ffn_up.weight
create_tensor: loading tensor blk.63.ffn_down.weight
create_tensor: loading tensor blk.0.attn_q.scale
create_tensor: loading tensor blk.0.attn_k.scale
create_tensor: loading tensor blk.0.attn_v.scale
create_tensor: loading tensor blk.0.attn_output.scale
create_tensor: loading tensor blk.0.ffn_gate.scale
create_tensor: loading tensor blk.0.ffn_down.scale
create_tensor: loading tensor blk.0.ffn_up.scale
create_tensor: loading tensor blk.0.attn_q.input_scale
create_tensor: loading tensor blk.0.attn_k.input_scale
create_tensor: loading tensor blk.0.attn_v.input_scale
create_tensor: loading tensor blk.0.attn_output.input_scale
create_tensor: loading tensor blk.0.ffn_gate.input_scale
create_tensor: loading tensor blk.0.ffn_down.input_scale
create_tensor: loading tensor blk.0.ffn_up.input_scale
create_tensor: loading tensor blk.1.attn_q.scale
create_tensor: loading tensor blk.1.attn_k.scale
create_tensor: loading tensor blk.1.attn_v.scale
create_tensor: loading tensor blk.1.attn_output.scale
create_tensor: loading tensor blk.1.ffn_gate.scale
create_tensor: loading tensor blk.1.ffn_down.scale
create_tensor: loading tensor blk.1.ffn_up.scale
create_tensor: loading tensor blk.1.attn_q.input_scale
create_tensor: loading tensor blk.1.attn_k.input_scale
create_tensor: loading tensor blk.1.attn_v.input_scale
create_tensor: loading tensor blk.1.attn_output.input_scale
create_tensor: loading tensor blk.1.ffn_gate.input_scale
create_tensor: loading tensor blk.1.ffn_down.input_scale
create_tensor: loading tensor blk.1.ffn_up.input_scale
create_tensor: loading tensor blk.2.attn_q.scale
create_tensor: loading tensor blk.2.attn_k.scale
create_tensor: loading tensor blk.2.attn_v.scale
create_tensor: loading tensor blk.2.attn_output.scale
create_tensor: loading tensor blk.2.ffn_gate.scale
create_tensor: loading tensor blk.2.ffn_down.scale
create_tensor: loading tensor blk.2.ffn_up.scale
create_tensor: loading tensor blk.2.attn_q.input_scale
create_tensor: loading tensor blk.2.attn_k.input_scale
create_tensor: loading tensor blk.2.attn_v.input_scale
create_tensor: loading tensor blk.2.attn_output.input_scale
create_tensor: loading tensor blk.2.ffn_gate.input_scale
create_tensor: loading tensor blk.2.ffn_down.input_scale
create_tensor: loading tensor blk.2.ffn_up.input_scale
create_tensor: loading tensor blk.3.attn_q.scale
create_tensor: loading tensor blk.3.attn_k.scale
create_tensor: loading tensor blk.3.attn_v.scale
create_tensor: loading tensor blk.3.attn_output.scale
create_tensor: loading tensor blk.3.ffn_gate.scale
create_tensor: loading tensor blk.3.ffn_down.scale
create_tensor: loading tensor blk.3.ffn_up.scale
create_tensor: loading tensor blk.3.attn_q.input_scale
create_tensor: loading tensor blk.3.attn_k.input_scale
create_tensor: loading tensor blk.3.attn_v.input_scale
create_tensor: loading tensor blk.3.attn_output.input_scale
create_tensor: loading tensor blk.3.ffn_gate.input_scale
create_tensor: loading tensor blk.3.ffn_down.input_scale
create_tensor: loading tensor blk.3.ffn_up.input_scale
create_tensor: loading tensor blk.4.attn_q.scale
create_tensor: loading tensor blk.4.attn_k.scale
create_tensor: loading tensor blk.4.attn_v.scale
create_tensor: loading tensor blk.4.attn_output.scale
create_tensor: loading tensor blk.4.ffn_gate.scale
create_tensor: loading tensor blk.4.ffn_down.scale
create_tensor: loading tensor blk.4.ffn_up.scale
create_tensor: loading tensor blk.4.attn_q.input_scale
create_tensor: loading tensor blk.4.attn_k.input_scale
create_tensor: loading tensor blk.4.attn_v.input_scale
create_tensor: loading tensor blk.4.attn_output.input_scale
create_tensor: loading tensor blk.4.ffn_gate.input_scale
create_tensor: loading tensor blk.4.ffn_down.input_scale
create_tensor: loading tensor blk.4.ffn_up.input_scale
create_tensor: loading tensor blk.5.attn_q.scale
create_tensor: loading tensor blk.5.attn_k.scale
create_tensor: loading tensor blk.5.attn_v.scale
create_tensor: loading tensor blk.5.attn_output.scale
create_tensor: loading tensor blk.5.ffn_gate.scale
create_tensor: loading tensor blk.5.ffn_down.scale
create_tensor: loading tensor blk.5.ffn_up.scale
create_tensor: loading tensor blk.5.attn_q.input_scale
create_tensor: loading tensor blk.5.attn_k.input_scale
create_tensor: loading tensor blk.5.attn_v.input_scale
create_tensor: loading tensor blk.5.attn_output.input_scale
create_tensor: loading tensor blk.5.ffn_gate.input_scale
create_tensor: loading tensor blk.5.ffn_down.input_scale
create_tensor: loading tensor blk.5.ffn_up.input_scale
create_tensor: loading tensor blk.6.attn_q.scale
create_tensor: loading tensor blk.6.attn_k.scale
create_tensor: loading tensor blk.6.attn_v.scale
create_tensor: loading tensor blk.6.attn_output.scale
create_tensor: loading tensor blk.6.ffn_gate.scale
create_tensor: loading tensor blk.6.ffn_down.scale
create_tensor: loading tensor blk.6.ffn_up.scale
create_tensor: loading tensor blk.6.attn_q.input_scale
create_tensor: loading tensor blk.6.attn_k.input_scale
create_tensor: loading tensor blk.6.attn_v.input_scale
create_tensor: loading tensor blk.6.attn_output.input_scale
create_tensor: loading tensor blk.6.ffn_gate.input_scale
create_tensor: loading tensor blk.6.ffn_down.input_scale
create_tensor: loading tensor blk.6.ffn_up.input_scale
create_tensor: loading tensor blk.7.attn_q.scale
create_tensor: loading tensor blk.7.attn_k.scale
create_tensor: loading tensor blk.7.attn_v.scale
create_tensor: loading tensor blk.7.attn_output.scale
create_tensor: loading tensor blk.7.ffn_gate.scale
create_tensor: loading tensor blk.7.ffn_down.scale
create_tensor: loading tensor blk.7.ffn_up.scale
create_tensor: loading tensor blk.7.attn_q.input_scale
create_tensor: loading tensor blk.7.attn_k.input_scale
create_tensor: loading tensor blk.7.attn_v.input_scale
create_tensor: loading tensor blk.7.attn_output.input_scale
create_tensor: loading tensor blk.7.ffn_gate.input_scale
create_tensor: loading tensor blk.7.ffn_down.input_scale
create_tensor: loading tensor blk.7.ffn_up.input_scale
create_tensor: loading tensor blk.8.attn_q.scale
create_tensor: loading tensor blk.8.attn_k.scale
create_tensor: loading tensor blk.8.attn_v.scale
create_tensor: loading tensor blk.8.attn_output.scale
create_tensor: loading tensor blk.8.ffn_gate.scale
create_tensor: loading tensor blk.8.ffn_down.scale
create_tensor: loading tensor blk.8.ffn_up.scale
create_tensor: loading tensor blk.8.attn_q.input_scale
create_tensor: loading tensor blk.8.attn_k.input_scale
create_tensor: loading tensor blk.8.attn_v.input_scale
create_tensor: loading tensor blk.8.attn_output.input_scale
create_tensor: loading tensor blk.8.ffn_gate.input_scale
create_tensor: loading tensor blk.8.ffn_down.input_scale
create_tensor: loading tensor blk.8.ffn_up.input_scale
create_tensor: loading tensor blk.9.attn_q.scale
create_tensor: loading tensor blk.9.attn_k.scale
create_tensor: loading tensor blk.9.attn_v.scale
create_tensor: loading tensor blk.9.attn_output.scale
create_tensor: loading tensor blk.9.ffn_gate.scale
create_tensor: loading tensor blk.9.ffn_down.scale
create_tensor: loading tensor blk.9.ffn_up.scale
create_tensor: loading tensor blk.9.attn_q.input_scale
create_tensor: loading tensor blk.9.attn_k.input_scale
create_tensor: loading tensor blk.9.attn_v.input_scale
create_tensor: loading tensor blk.9.attn_output.input_scale
create_tensor: loading tensor blk.9.ffn_gate.input_scale
create_tensor: loading tensor blk.9.ffn_down.input_scale
create_tensor: loading tensor blk.9.ffn_up.input_scale
create_tensor: loading tensor blk.10.attn_q.scale
create_tensor: loading tensor blk.10.attn_k.scale
create_tensor: loading tensor blk.10.attn_v.scale
create_tensor: loading tensor blk.10.attn_output.scale
create_tensor: loading tensor blk.10.ffn_gate.scale
create_tensor: loading tensor blk.10.ffn_down.scale
create_tensor: loading tensor blk.10.ffn_up.scale
create_tensor: loading tensor blk.10.attn_q.input_scale
create_tensor: loading tensor blk.10.attn_k.input_scale
create_tensor: loading tensor blk.10.attn_v.input_scale
create_tensor: loading tensor blk.10.attn_output.input_scale
create_tensor: loading tensor blk.10.ffn_gate.input_scale
create_tensor: loading tensor blk.10.ffn_down.input_scale
create_tensor: loading tensor blk.10.ffn_up.input_scale
create_tensor: loading tensor blk.11.attn_q.scale
create_tensor: loading tensor blk.11.attn_k.scale
create_tensor: loading tensor blk.11.attn_v.scale
create_tensor: loading tensor blk.11.attn_output.scale
create_tensor: loading tensor blk.11.ffn_gate.scale
create_tensor: loading tensor blk.11.ffn_down.scale
create_tensor: loading tensor blk.11.ffn_up.scale
create_tensor: loading tensor blk.11.attn_q.input_scale
create_tensor: loading tensor blk.11.attn_k.input_scale
create_tensor: loading tensor blk.11.attn_v.input_scale
create_tensor: loading tensor blk.11.attn_output.input_scale
create_tensor: loading tensor blk.11.ffn_gate.input_scale
create_tensor: loading tensor blk.11.ffn_down.input_scale
create_tensor: loading tensor blk.11.ffn_up.input_scale
create_tensor: loading tensor blk.12.attn_q.scale
create_tensor: loading tensor blk.12.attn_k.scale
create_tensor: loading tensor blk.12.attn_v.scale
create_tensor: loading tensor blk.12.attn_output.scale
create_tensor: loading tensor blk.12.ffn_gate.scale
create_tensor: loading tensor blk.12.ffn_down.scale
create_tensor: loading tensor blk.12.ffn_up.scale
create_tensor: loading tensor blk.12.attn_q.input_scale
create_tensor: loading tensor blk.12.attn_k.input_scale
create_tensor: loading tensor blk.12.attn_v.input_scale
create_tensor: loading tensor blk.12.attn_output.input_scale
create_tensor: loading tensor blk.12.ffn_gate.input_scale
create_tensor: loading tensor blk.12.ffn_down.input_scale
create_tensor: loading tensor blk.12.ffn_up.input_scale
create_tensor: loading tensor blk.13.attn_q.scale
create_tensor: loading tensor blk.13.attn_k.scale
create_tensor: loading tensor blk.13.attn_v.scale
create_tensor: loading tensor blk.13.attn_output.scale
create_tensor: loading tensor blk.13.ffn_gate.scale
create_tensor: loading tensor blk.13.ffn_down.scale
create_tensor: loading tensor blk.13.ffn_up.scale
create_tensor: loading tensor blk.13.attn_q.input_scale
create_tensor: loading tensor blk.13.attn_k.input_scale
create_tensor: loading tensor blk.13.attn_v.input_scale
create_tensor: loading tensor blk.13.attn_output.input_scale
create_tensor: loading tensor blk.13.ffn_gate.input_scale
create_tensor: loading tensor blk.13.ffn_down.input_scale
create_tensor: loading tensor blk.13.ffn_up.input_scale
create_tensor: loading tensor blk.14.attn_q.scale
create_tensor: loading tensor blk.14.attn_k.scale
create_tensor: loading tensor blk.14.attn_v.scale
create_tensor: loading tensor blk.14.attn_output.scale
create_tensor: loading tensor blk.14.ffn_gate.scale
create_tensor: loading tensor blk.14.ffn_down.scale
create_tensor: loading tensor blk.14.ffn_up.scale
create_tensor: loading tensor blk.14.attn_q.input_scale
create_tensor: loading tensor blk.14.attn_k.input_scale
create_tensor: loading tensor blk.14.attn_v.input_scale
create_tensor: loading tensor blk.14.attn_output.input_scale
create_tensor: loading tensor blk.14.ffn_gate.input_scale
create_tensor: loading tensor blk.14.ffn_down.input_scale
create_tensor: loading tensor blk.14.ffn_up.input_scale
create_tensor: loading tensor blk.15.attn_q.scale
create_tensor: loading tensor blk.15.attn_k.scale
create_tensor: loading tensor blk.15.attn_v.scale
create_tensor: loading tensor blk.15.attn_output.scale
create_tensor: loading tensor blk.15.ffn_gate.scale
create_tensor: loading tensor blk.15.ffn_down.scale
create_tensor: loading tensor blk.15.ffn_up.scale
create_tensor: loading tensor blk.15.attn_q.input_scale
create_tensor: loading tensor blk.15.attn_k.input_scale
create_tensor: loading tensor blk.15.attn_v.input_scale
create_tensor: loading tensor blk.15.attn_output.input_scale
create_tensor: loading tensor blk.15.ffn_gate.input_scale
create_tensor: loading tensor blk.15.ffn_down.input_scale
create_tensor: loading tensor blk.15.ffn_up.input_scale
create_tensor: loading tensor blk.16.attn_q.scale
create_tensor: loading tensor blk.16.attn_k.scale
create_tensor: loading tensor blk.16.attn_v.scale
create_tensor: loading tensor blk.16.attn_output.scale
create_tensor: loading tensor blk.16.ffn_gate.scale
create_tensor: loading tensor blk.16.ffn_down.scale
create_tensor: loading tensor blk.16.ffn_up.scale
create_tensor: loading tensor blk.16.attn_q.input_scale
create_tensor: loading tensor blk.16.attn_k.input_scale
create_tensor: loading tensor blk.16.attn_v.input_scale
create_tensor: loading tensor blk.16.attn_output.input_scale
create_tensor: loading tensor blk.16.ffn_gate.input_scale
create_tensor: loading tensor blk.16.ffn_down.input_scale
create_tensor: loading tensor blk.16.ffn_up.input_scale
create_tensor: loading tensor blk.17.attn_q.scale
create_tensor: loading tensor blk.17.attn_k.scale
create_tensor: loading tensor blk.17.attn_v.scale
create_tensor: loading tensor blk.17.attn_output.scale
create_tensor: loading tensor blk.17.ffn_gate.scale
create_tensor: loading tensor blk.17.ffn_down.scale
create_tensor: loading tensor blk.17.ffn_up.scale
create_tensor: loading tensor blk.17.attn_q.input_scale
create_tensor: loading tensor blk.17.attn_k.input_scale
create_tensor: loading tensor blk.17.attn_v.input_scale
create_tensor: loading tensor blk.17.attn_output.input_scale
create_tensor: loading tensor blk.17.ffn_gate.input_scale
create_tensor: loading tensor blk.17.ffn_down.input_scale
create_tensor: loading tensor blk.17.ffn_up.input_scale
create_tensor: loading tensor blk.18.attn_q.scale
create_tensor: loading tensor blk.18.attn_k.scale
create_tensor: loading tensor blk.18.attn_v.scale
create_tensor: loading tensor blk.18.attn_output.scale
create_tensor: loading tensor blk.18.ffn_gate.scale
create_tensor: loading tensor blk.18.ffn_down.scale
create_tensor: loading tensor blk.18.ffn_up.scale
create_tensor: loading tensor blk.18.attn_q.input_scale
create_tensor: loading tensor blk.18.attn_k.input_scale
create_tensor: loading tensor blk.18.attn_v.input_scale
create_tensor: loading tensor blk.18.attn_output.input_scale
create_tensor: loading tensor blk.18.ffn_gate.input_scale
create_tensor: loading tensor blk.18.ffn_down.input_scale
create_tensor: loading tensor blk.18.ffn_up.input_scale
create_tensor: loading tensor blk.19.attn_q.scale
create_tensor: loading tensor blk.19.attn_k.scale
create_tensor: loading tensor blk.19.attn_v.scale
create_tensor: loading tensor blk.19.attn_output.scale
create_tensor: loading tensor blk.19.ffn_gate.scale
create_tensor: loading tensor blk.19.ffn_down.scale
create_tensor: loading tensor blk.19.ffn_up.scale
create_tensor: loading tensor blk.19.attn_q.input_scale
create_tensor: loading tensor blk.19.attn_k.input_scale
create_tensor: loading tensor blk.19.attn_v.input_scale
create_tensor: loading tensor blk.19.attn_output.input_scale
create_tensor: loading tensor blk.19.ffn_gate.input_scale
create_tensor: loading tensor blk.19.ffn_down.input_scale
create_tensor: loading tensor blk.19.ffn_up.input_scale
create_tensor: loading tensor blk.20.attn_q.scale
create_tensor: loading tensor blk.20.attn_k.scale
create_tensor: loading tensor blk.20.attn_v.scale
create_tensor: loading tensor blk.20.attn_output.scale
create_tensor: loading tensor blk.20.ffn_gate.scale
create_tensor: loading tensor blk.20.ffn_down.scale
create_tensor: loading tensor blk.20.ffn_up.scale
create_tensor: loading tensor blk.20.attn_q.input_scale
create_tensor: loading tensor blk.20.attn_k.input_scale
create_tensor: loading tensor blk.20.attn_v.input_scale
create_tensor: loading tensor blk.20.attn_output.input_scale
create_tensor: loading tensor blk.20.ffn_gate.input_scale
create_tensor: loading tensor blk.20.ffn_down.input_scale
create_tensor: loading tensor blk.20.ffn_up.input_scale
create_tensor: loading tensor blk.21.attn_q.scale
create_tensor: loading tensor blk.21.attn_k.scale
create_tensor: loading tensor blk.21.attn_v.scale
create_tensor: loading tensor blk.21.attn_output.scale
create_tensor: loading tensor blk.21.ffn_gate.scale
create_tensor: loading tensor blk.21.ffn_down.scale
create_tensor: loading tensor blk.21.ffn_up.scale
create_tensor: loading tensor blk.21.attn_q.input_scale
create_tensor: loading tensor blk.21.attn_k.input_scale
create_tensor: loading tensor blk.21.attn_v.input_scale
create_tensor: loading tensor blk.21.attn_output.input_scale
create_tensor: loading tensor blk.21.ffn_gate.input_scale
create_tensor: loading tensor blk.21.ffn_down.input_scale
create_tensor: loading tensor blk.21.ffn_up.input_scale
create_tensor: loading tensor blk.22.attn_q.scale
create_tensor: loading tensor blk.22.attn_k.scale
create_tensor: loading tensor blk.22.attn_v.scale
create_tensor: loading tensor blk.22.attn_output.scale
create_tensor: loading tensor blk.22.ffn_gate.scale
create_tensor: loading tensor blk.22.ffn_down.scale
create_tensor: loading tensor blk.22.ffn_up.scale
create_tensor: loading tensor blk.22.attn_q.input_scale
create_tensor: loading tensor blk.22.attn_k.input_scale
create_tensor: loading tensor blk.22.attn_v.input_scale
create_tensor: loading tensor blk.22.attn_output.input_scale
create_tensor: loading tensor blk.22.ffn_gate.input_scale
create_tensor: loading tensor blk.22.ffn_down.input_scale
create_tensor: loading tensor blk.22.ffn_up.input_scale
create_tensor: loading tensor blk.23.attn_q.scale
create_tensor: loading tensor blk.23.attn_k.scale
create_tensor: loading tensor blk.23.attn_v.scale
create_tensor: loading tensor blk.23.attn_output.scale
create_tensor: loading tensor blk.23.ffn_gate.scale
create_tensor: loading tensor blk.23.ffn_down.scale
create_tensor: loading tensor blk.23.ffn_up.scale
create_tensor: loading tensor blk.23.attn_q.input_scale
create_tensor: loading tensor blk.23.attn_k.input_scale
create_tensor: loading tensor blk.23.attn_v.input_scale
create_tensor: loading tensor blk.23.attn_output.input_scale
create_tensor: loading tensor blk.23.ffn_gate.input_scale
create_tensor: loading tensor blk.23.ffn_down.input_scale
create_tensor: loading tensor blk.23.ffn_up.input_scale
create_tensor: loading tensor blk.24.attn_q.scale
create_tensor: loading tensor blk.24.attn_k.scale
create_tensor: loading tensor blk.24.attn_v.scale
create_tensor: loading tensor blk.24.attn_output.scale
create_tensor: loading tensor blk.24.ffn_gate.scale
create_tensor: loading tensor blk.24.ffn_down.scale
create_tensor: loading tensor blk.24.ffn_up.scale
create_tensor: loading tensor blk.24.attn_q.input_scale
create_tensor: loading tensor blk.24.attn_k.input_scale
create_tensor: loading tensor blk.24.attn_v.input_scale
create_tensor: loading tensor blk.24.attn_output.input_scale
create_tensor: loading tensor blk.24.ffn_gate.input_scale
create_tensor: loading tensor blk.24.ffn_down.input_scale
create_tensor: loading tensor blk.24.ffn_up.input_scale
create_tensor: loading tensor blk.25.attn_q.scale
create_tensor: loading tensor blk.25.attn_k.scale
create_tensor: loading tensor blk.25.attn_v.scale
create_tensor: loading tensor blk.25.attn_output.scale
create_tensor: loading tensor blk.25.ffn_gate.scale
create_tensor: loading tensor blk.25.ffn_down.scale
create_tensor: loading tensor blk.25.ffn_up.scale
create_tensor: loading tensor blk.25.attn_q.input_scale
create_tensor: loading tensor blk.25.attn_k.input_scale
create_tensor: loading tensor blk.25.attn_v.input_scale
create_tensor: loading tensor blk.25.attn_output.input_scale
create_tensor: loading tensor blk.25.ffn_gate.input_scale
create_tensor: loading tensor blk.25.ffn_down.input_scale
create_tensor: loading tensor blk.25.ffn_up.input_scale
create_tensor: loading tensor blk.26.attn_q.scale
create_tensor: loading tensor blk.26.attn_k.scale
create_tensor: loading tensor blk.26.attn_v.scale
create_tensor: loading tensor blk.26.attn_output.scale
create_tensor: loading tensor blk.26.ffn_gate.scale
create_tensor: loading tensor blk.26.ffn_down.scale
create_tensor: loading tensor blk.26.ffn_up.scale
create_tensor: loading tensor blk.26.attn_q.input_scale
create_tensor: loading tensor blk.26.attn_k.input_scale
create_tensor: loading tensor blk.26.attn_v.input_scale
create_tensor: loading tensor blk.26.attn_output.input_scale
create_tensor: loading tensor blk.26.ffn_gate.input_scale
create_tensor: loading tensor blk.26.ffn_down.input_scale
create_tensor: loading tensor blk.26.ffn_up.input_scale
create_tensor: loading tensor blk.27.attn_q.scale
create_tensor: loading tensor blk.27.attn_k.scale
create_tensor: loading tensor blk.27.attn_v.scale
create_tensor: loading tensor blk.27.attn_output.scale
create_tensor: loading tensor blk.27.ffn_gate.scale
create_tensor: loading tensor blk.27.ffn_down.scale
create_tensor: loading tensor blk.27.ffn_up.scale
create_tensor: loading tensor blk.27.attn_q.input_scale
create_tensor: loading tensor blk.27.attn_k.input_scale
create_tensor: loading tensor blk.27.attn_v.input_scale
create_tensor: loading tensor blk.27.attn_output.input_scale
create_tensor: loading tensor blk.27.ffn_gate.input_scale
create_tensor: loading tensor blk.27.ffn_down.input_scale
create_tensor: loading tensor blk.27.ffn_up.input_scale
create_tensor: loading tensor blk.28.attn_q.scale
create_tensor: loading tensor blk.28.attn_k.scale
create_tensor: loading tensor blk.28.attn_v.scale
create_tensor: loading tensor blk.28.attn_output.scale
create_tensor: loading tensor blk.28.ffn_gate.scale
create_tensor: loading tensor blk.28.ffn_down.scale
create_tensor: loading tensor blk.28.ffn_up.scale
create_tensor: loading tensor blk.28.attn_q.input_scale
create_tensor: loading tensor blk.28.attn_k.input_scale
create_tensor: loading tensor blk.28.attn_v.input_scale
create_tensor: loading tensor blk.28.attn_output.input_scale
create_tensor: loading tensor blk.28.ffn_gate.input_scale
create_tensor: loading tensor blk.28.ffn_down.input_scale
create_tensor: loading tensor blk.28.ffn_up.input_scale
create_tensor: loading tensor blk.29.attn_q.scale
create_tensor: loading tensor blk.29.attn_k.scale
create_tensor: loading tensor blk.29.attn_v.scale
create_tensor: loading tensor blk.29.attn_output.scale
create_tensor: loading tensor blk.29.ffn_gate.scale
create_tensor: loading tensor blk.29.ffn_down.scale
create_tensor: loading tensor blk.29.ffn_up.scale
create_tensor: loading tensor blk.29.attn_q.input_scale
create_tensor: loading tensor blk.29.attn_k.input_scale
create_tensor: loading tensor blk.29.attn_v.input_scale
create_tensor: loading tensor blk.29.attn_output.input_scale
create_tensor: loading tensor blk.29.ffn_gate.input_scale
create_tensor: loading tensor blk.29.ffn_down.input_scale
create_tensor: loading tensor blk.29.ffn_up.input_scale
create_tensor: loading tensor blk.30.attn_q.scale
create_tensor: loading tensor blk.30.attn_k.scale
create_tensor: loading tensor blk.30.attn_v.scale
create_tensor: loading tensor blk.30.attn_output.scale
create_tensor: loading tensor blk.30.ffn_gate.scale
create_tensor: loading tensor blk.30.ffn_down.scale
create_tensor: loading tensor blk.30.ffn_up.scale
create_tensor: loading tensor blk.30.attn_q.input_scale
create_tensor: loading tensor blk.30.attn_k.input_scale
create_tensor: loading tensor blk.30.attn_v.input_scale
create_tensor: loading tensor blk.30.attn_output.input_scale
create_tensor: loading tensor blk.30.ffn_gate.input_scale
create_tensor: loading tensor blk.30.ffn_down.input_scale
create_tensor: loading tensor blk.30.ffn_up.input_scale
create_tensor: loading tensor blk.31.attn_q.scale
create_tensor: loading tensor blk.31.attn_k.scale
create_tensor: loading tensor blk.31.attn_v.scale
create_tensor: loading tensor blk.31.attn_output.scale
create_tensor: loading tensor blk.31.ffn_gate.scale
create_tensor: loading tensor blk.31.ffn_down.scale
create_tensor: loading tensor blk.31.ffn_up.scale
create_tensor: loading tensor blk.31.attn_q.input_scale
create_tensor: loading tensor blk.31.attn_k.input_scale
create_tensor: loading tensor blk.31.attn_v.input_scale
create_tensor: loading tensor blk.31.attn_output.input_scale
create_tensor: loading tensor blk.31.ffn_gate.input_scale
create_tensor: loading tensor blk.31.ffn_down.input_scale
create_tensor: loading tensor blk.31.ffn_up.input_scale
create_tensor: loading tensor blk.32.attn_q.scale
create_tensor: loading tensor blk.32.attn_k.scale
create_tensor: loading tensor blk.32.attn_v.scale
create_tensor: loading tensor blk.32.attn_output.scale
create_tensor: loading tensor blk.32.ffn_gate.scale
create_tensor: loading tensor blk.32.ffn_down.scale
create_tensor: loading tensor blk.32.ffn_up.scale
create_tensor: loading tensor blk.32.attn_q.input_scale
create_tensor: loading tensor blk.32.attn_k.input_scale
create_tensor: loading tensor blk.32.attn_v.input_scale
create_tensor: loading tensor blk.32.attn_output.input_scale
create_tensor: loading tensor blk.32.ffn_gate.input_scale
create_tensor: loading tensor blk.32.ffn_down.input_scale
create_tensor: loading tensor blk.32.ffn_up.input_scale
create_tensor: loading tensor blk.33.attn_q.scale
create_tensor: loading tensor blk.33.attn_k.scale
create_tensor: loading tensor blk.33.attn_v.scale
create_tensor: loading tensor blk.33.attn_output.scale
create_tensor: loading tensor blk.33.ffn_gate.scale
create_tensor: loading tensor blk.33.ffn_down.scale
create_tensor: loading tensor blk.33.ffn_up.scale
create_tensor: loading tensor blk.33.attn_q.input_scale
create_tensor: loading tensor blk.33.attn_k.input_scale
create_tensor: loading tensor blk.33.attn_v.input_scale
create_tensor: loading tensor blk.33.attn_output.input_scale
create_tensor: loading tensor blk.33.ffn_gate.input_scale
create_tensor: loading tensor blk.33.ffn_down.input_scale
create_tensor: loading tensor blk.33.ffn_up.input_scale
create_tensor: loading tensor blk.34.attn_q.scale
create_tensor: loading tensor blk.34.attn_k.scale
create_tensor: loading tensor blk.34.attn_v.scale
create_tensor: loading tensor blk.34.attn_output.scale
create_tensor: loading tensor blk.34.ffn_gate.scale
create_tensor: loading tensor blk.34.ffn_down.scale
create_tensor: loading tensor blk.34.ffn_up.scale
create_tensor: loading tensor blk.34.attn_q.input_scale
create_tensor: loading tensor blk.34.attn_k.input_scale
create_tensor: loading tensor blk.34.attn_v.input_scale
create_tensor: loading tensor blk.34.attn_output.input_scale
create_tensor: loading tensor blk.34.ffn_gate.input_scale
create_tensor: loading tensor blk.34.ffn_down.input_scale
create_tensor: loading tensor blk.34.ffn_up.input_scale
create_tensor: loading tensor blk.35.attn_q.scale
create_tensor: loading tensor blk.35.attn_k.scale
create_tensor: loading tensor blk.35.attn_v.scale
create_tensor: loading tensor blk.35.attn_output.scale
create_tensor: loading tensor blk.35.ffn_gate.scale
create_tensor: loading tensor blk.35.ffn_down.scale
create_tensor: loading tensor blk.35.ffn_up.scale
create_tensor: loading tensor blk.35.attn_q.input_scale
create_tensor: loading tensor blk.35.attn_k.input_scale
create_tensor: loading tensor blk.35.attn_v.input_scale
create_tensor: loading tensor blk.35.attn_output.input_scale
create_tensor: loading tensor blk.35.ffn_gate.input_scale
create_tensor: loading tensor blk.35.ffn_down.input_scale
create_tensor: loading tensor blk.35.ffn_up.input_scale
create_tensor: loading tensor blk.36.attn_q.scale
create_tensor: loading tensor blk.36.attn_k.scale
create_tensor: loading tensor blk.36.attn_v.scale
create_tensor: loading tensor blk.36.attn_output.scale
create_tensor: loading tensor blk.36.ffn_gate.scale
create_tensor: loading tensor blk.36.ffn_down.scale
create_tensor: loading tensor blk.36.ffn_up.scale
create_tensor: loading tensor blk.36.attn_q.input_scale
create_tensor: loading tensor blk.36.attn_k.input_scale
create_tensor: loading tensor blk.36.attn_v.input_scale
create_tensor: loading tensor blk.36.attn_output.input_scale
create_tensor: loading tensor blk.36.ffn_gate.input_scale
create_tensor: loading tensor blk.36.ffn_down.input_scale
create_tensor: loading tensor blk.36.ffn_up.input_scale
create_tensor: loading tensor blk.37.attn_q.scale
create_tensor: loading tensor blk.37.attn_k.scale
create_tensor: loading tensor blk.37.attn_v.scale
create_tensor: loading tensor blk.37.attn_output.scale
create_tensor: loading tensor blk.37.ffn_gate.scale
create_tensor: loading tensor blk.37.ffn_down.scale
create_tensor: loading tensor blk.37.ffn_up.scale
create_tensor: loading tensor blk.37.attn_q.input_scale
create_tensor: loading tensor blk.37.attn_k.input_scale
create_tensor: loading tensor blk.37.attn_v.input_scale
create_tensor: loading tensor blk.37.attn_output.input_scale
create_tensor: loading tensor blk.37.ffn_gate.input_scale
create_tensor: loading tensor blk.37.ffn_down.input_scale
create_tensor: loading tensor blk.37.ffn_up.input_scale
create_tensor: loading tensor blk.38.attn_q.scale
create_tensor: loading tensor blk.38.attn_k.scale
create_tensor: loading tensor blk.38.attn_v.scale
create_tensor: loading tensor blk.38.attn_output.scale
create_tensor: loading tensor blk.38.ffn_gate.scale
create_tensor: loading tensor blk.38.ffn_down.scale
create_tensor: loading tensor blk.38.ffn_up.scale
create_tensor: loading tensor blk.38.attn_q.input_scale
create_tensor: loading tensor blk.38.attn_k.input_scale
create_tensor: loading tensor blk.38.attn_v.input_scale
create_tensor: loading tensor blk.38.attn_output.input_scale
create_tensor: loading tensor blk.38.ffn_gate.input_scale
create_tensor: loading tensor blk.38.ffn_down.input_scale
create_tensor: loading tensor blk.38.ffn_up.input_scale
create_tensor: loading tensor blk.39.attn_q.scale
create_tensor: loading tensor blk.39.attn_k.scale
create_tensor: loading tensor blk.39.attn_v.scale
create_tensor: loading tensor blk.39.attn_output.scale
create_tensor: loading tensor blk.39.ffn_gate.scale
create_tensor: loading tensor blk.39.ffn_down.scale
create_tensor: loading tensor blk.39.ffn_up.scale
create_tensor: loading tensor blk.39.attn_q.input_scale
create_tensor: loading tensor blk.39.attn_k.input_scale
create_tensor: loading tensor blk.39.attn_v.input_scale
create_tensor: loading tensor blk.39.attn_output.input_scale
create_tensor: loading tensor blk.39.ffn_gate.input_scale
create_tensor: loading tensor blk.39.ffn_down.input_scale
create_tensor: loading tensor blk.39.ffn_up.input_scale
create_tensor: loading tensor blk.40.attn_q.scale
create_tensor: loading tensor blk.40.attn_k.scale
create_tensor: loading tensor blk.40.attn_v.scale
create_tensor: loading tensor blk.40.attn_output.scale
create_tensor: loading tensor blk.40.ffn_gate.scale
create_tensor: loading tensor blk.40.ffn_down.scale
create_tensor: loading tensor blk.40.ffn_up.scale
create_tensor: loading tensor blk.40.attn_q.input_scale
create_tensor: loading tensor blk.40.attn_k.input_scale
create_tensor: loading tensor blk.40.attn_v.input_scale
create_tensor: loading tensor blk.40.attn_output.input_scale
create_tensor: loading tensor blk.40.ffn_gate.input_scale
create_tensor: loading tensor blk.40.ffn_down.input_scale
create_tensor: loading tensor blk.40.ffn_up.input_scale
create_tensor: loading tensor blk.41.attn_q.scale
create_tensor: loading tensor blk.41.attn_k.scale
create_tensor: loading tensor blk.41.attn_v.scale
create_tensor: loading tensor blk.41.attn_output.scale
create_tensor: loading tensor blk.41.ffn_gate.scale
create_tensor: loading tensor blk.41.ffn_down.scale
create_tensor: loading tensor blk.41.ffn_up.scale
create_tensor: loading tensor blk.41.attn_q.input_scale
create_tensor: loading tensor blk.41.attn_k.input_scale
create_tensor: loading tensor blk.41.attn_v.input_scale
create_tensor: loading tensor blk.41.attn_output.input_scale
create_tensor: loading tensor blk.41.ffn_gate.input_scale
create_tensor: loading tensor blk.41.ffn_down.input_scale
create_tensor: loading tensor blk.41.ffn_up.input_scale
create_tensor: loading tensor blk.42.attn_q.scale
create_tensor: loading tensor blk.42.attn_k.scale
create_tensor: loading tensor blk.42.attn_v.scale
create_tensor: loading tensor blk.42.attn_output.scale
create_tensor: loading tensor blk.42.ffn_gate.scale
create_tensor: loading tensor blk.42.ffn_down.scale
create_tensor: loading tensor blk.42.ffn_up.scale
create_tensor: loading tensor blk.42.attn_q.input_scale
create_tensor: loading tensor blk.42.attn_k.input_scale
create_tensor: loading tensor blk.42.attn_v.input_scale
create_tensor: loading tensor blk.42.attn_output.input_scale
create_tensor: loading tensor blk.42.ffn_gate.input_scale
create_tensor: loading tensor blk.42.ffn_down.input_scale
create_tensor: loading tensor blk.42.ffn_up.input_scale
create_tensor: loading tensor blk.43.attn_q.scale
create_tensor: loading tensor blk.43.attn_k.scale
create_tensor: loading tensor blk.43.attn_v.scale
create_tensor: loading tensor blk.43.attn_output.scale
create_tensor: loading tensor blk.43.ffn_gate.scale
create_tensor: loading tensor blk.43.ffn_down.scale
create_tensor: loading tensor blk.43.ffn_up.scale
create_tensor: loading tensor blk.43.attn_q.input_scale
create_tensor: loading tensor blk.43.attn_k.input_scale
create_tensor: loading tensor blk.43.attn_v.input_scale
create_tensor: loading tensor blk.43.attn_output.input_scale
create_tensor: loading tensor blk.43.ffn_gate.input_scale
create_tensor: loading tensor blk.43.ffn_down.input_scale
create_tensor: loading tensor blk.43.ffn_up.input_scale
create_tensor: loading tensor blk.44.attn_q.scale
create_tensor: loading tensor blk.44.attn_k.scale
create_tensor: loading tensor blk.44.attn_v.scale
create_tensor: loading tensor blk.44.attn_output.scale
create_tensor: loading tensor blk.44.ffn_gate.scale
create_tensor: loading tensor blk.44.ffn_down.scale
create_tensor: loading tensor blk.44.ffn_up.scale
create_tensor: loading tensor blk.44.attn_q.input_scale
create_tensor: loading tensor blk.44.attn_k.input_scale
create_tensor: loading tensor blk.44.attn_v.input_scale
create_tensor: loading tensor blk.44.attn_output.input_scale
create_tensor: loading tensor blk.44.ffn_gate.input_scale
create_tensor: loading tensor blk.44.ffn_down.input_scale
create_tensor: loading tensor blk.44.ffn_up.input_scale
create_tensor: loading tensor blk.45.attn_q.scale
create_tensor: loading tensor blk.45.attn_k.scale
create_tensor: loading tensor blk.45.attn_v.scale
create_tensor: loading tensor blk.45.attn_output.scale
create_tensor: loading tensor blk.45.ffn_gate.scale
create_tensor: loading tensor blk.45.ffn_down.scale
create_tensor: loading tensor blk.45.ffn_up.scale
create_tensor: loading tensor blk.45.attn_q.input_scale
create_tensor: loading tensor blk.45.attn_k.input_scale
create_tensor: loading tensor blk.45.attn_v.input_scale
create_tensor: loading tensor blk.45.attn_output.input_scale
create_tensor: loading tensor blk.45.ffn_gate.input_scale
create_tensor: loading tensor blk.45.ffn_down.input_scale
create_tensor: loading tensor blk.45.ffn_up.input_scale
create_tensor: loading tensor blk.46.attn_q.scale
create_tensor: loading tensor blk.46.attn_k.scale
create_tensor: loading tensor blk.46.attn_v.scale
create_tensor: loading tensor blk.46.attn_output.scale
create_tensor: loading tensor blk.46.ffn_gate.scale
create_tensor: loading tensor blk.46.ffn_down.scale
create_tensor: loading tensor blk.46.ffn_up.scale
create_tensor: loading tensor blk.46.attn_q.input_scale
create_tensor: loading tensor blk.46.attn_k.input_scale
create_tensor: loading tensor blk.46.attn_v.input_scale
create_tensor: loading tensor blk.46.attn_output.input_scale
create_tensor: loading tensor blk.46.ffn_gate.input_scale
create_tensor: loading tensor blk.46.ffn_down.input_scale
create_tensor: loading tensor blk.46.ffn_up.input_scale
create_tensor: loading tensor blk.47.attn_q.scale
create_tensor: loading tensor blk.47.attn_k.scale
create_tensor: loading tensor blk.47.attn_v.scale
create_tensor: loading tensor blk.47.attn_output.scale
create_tensor: loading tensor blk.47.ffn_gate.scale
create_tensor: loading tensor blk.47.ffn_down.scale
create_tensor: loading tensor blk.47.ffn_up.scale
create_tensor: loading tensor blk.47.attn_q.input_scale
create_tensor: loading tensor blk.47.attn_k.input_scale
create_tensor: loading tensor blk.47.attn_v.input_scale
create_tensor: loading tensor blk.47.attn_output.input_scale
create_tensor: loading tensor blk.47.ffn_gate.input_scale
create_tensor: loading tensor blk.47.ffn_down.input_scale
create_tensor: loading tensor blk.47.ffn_up.input_scale
create_tensor: loading tensor blk.48.attn_q.scale
create_tensor: loading tensor blk.48.attn_k.scale
create_tensor: loading tensor blk.48.attn_v.scale
create_tensor: loading tensor blk.48.attn_output.scale
create_tensor: loading tensor blk.48.ffn_gate.scale
create_tensor: loading tensor blk.48.ffn_down.scale
create_tensor: loading tensor blk.48.ffn_up.scale
create_tensor: loading tensor blk.48.attn_q.input_scale
create_tensor: loading tensor blk.48.attn_k.input_scale
create_tensor: loading tensor blk.48.attn_v.input_scale
create_tensor: loading tensor blk.48.attn_output.input_scale
create_tensor: loading tensor blk.48.ffn_gate.input_scale
create_tensor: loading tensor blk.48.ffn_down.input_scale
create_tensor: loading tensor blk.48.ffn_up.input_scale
create_tensor: loading tensor blk.49.attn_q.scale
create_tensor: loading tensor blk.49.attn_k.scale
create_tensor: loading tensor blk.49.attn_v.scale
create_tensor: loading tensor blk.49.attn_output.scale
create_tensor: loading tensor blk.49.ffn_gate.scale
create_tensor: loading tensor blk.49.ffn_down.scale
create_tensor: loading tensor blk.49.ffn_up.scale
create_tensor: loading tensor blk.49.attn_q.input_scale
create_tensor: loading tensor blk.49.attn_k.input_scale
create_tensor: loading tensor blk.49.attn_v.input_scale
create_tensor: loading tensor blk.49.attn_output.input_scale
create_tensor: loading tensor blk.49.ffn_gate.input_scale
create_tensor: loading tensor blk.49.ffn_down.input_scale
create_tensor: loading tensor blk.49.ffn_up.input_scale
create_tensor: loading tensor blk.50.attn_q.scale
create_tensor: loading tensor blk.50.attn_k.scale
create_tensor: loading tensor blk.50.attn_v.scale
create_tensor: loading tensor blk.50.attn_output.scale
create_tensor: loading tensor blk.50.ffn_gate.scale
create_tensor: loading tensor blk.50.ffn_down.scale
create_tensor: loading tensor blk.50.ffn_up.scale
create_tensor: loading tensor blk.50.attn_q.input_scale
create_tensor: loading tensor blk.50.attn_k.input_scale
create_tensor: loading tensor blk.50.attn_v.input_scale
create_tensor: loading tensor blk.50.attn_output.input_scale
create_tensor: loading tensor blk.50.ffn_gate.input_scale
create_tensor: loading tensor blk.50.ffn_down.input_scale
create_tensor: loading tensor blk.50.ffn_up.input_scale
create_tensor: loading tensor blk.51.attn_q.scale
create_tensor: loading tensor blk.51.attn_k.scale
create_tensor: loading tensor blk.51.attn_v.scale
create_tensor: loading tensor blk.51.attn_output.scale
create_tensor: loading tensor blk.51.ffn_gate.scale
create_tensor: loading tensor blk.51.ffn_down.scale
create_tensor: loading tensor blk.51.ffn_up.scale
create_tensor: loading tensor blk.51.attn_q.input_scale
create_tensor: loading tensor blk.51.attn_k.input_scale
create_tensor: loading tensor blk.51.attn_v.input_scale
create_tensor: loading tensor blk.51.attn_output.input_scale
create_tensor: loading tensor blk.51.ffn_gate.input_scale
create_tensor: loading tensor blk.51.ffn_down.input_scale
create_tensor: loading tensor blk.51.ffn_up.input_scale
create_tensor: loading tensor blk.52.attn_q.scale
create_tensor: loading tensor blk.52.attn_k.scale
create_tensor: loading tensor blk.52.attn_v.scale
create_tensor: loading tensor blk.52.attn_output.scale
create_tensor: loading tensor blk.52.ffn_gate.scale
create_tensor: loading tensor blk.52.ffn_down.scale
create_tensor: loading tensor blk.52.ffn_up.scale
create_tensor: loading tensor blk.52.attn_q.input_scale
create_tensor: loading tensor blk.52.attn_k.input_scale
create_tensor: loading tensor blk.52.attn_v.input_scale
create_tensor: loading tensor blk.52.attn_output.input_scale
create_tensor: loading tensor blk.52.ffn_gate.input_scale
create_tensor: loading tensor blk.52.ffn_down.input_scale
create_tensor: loading tensor blk.52.ffn_up.input_scale
create_tensor: loading tensor blk.53.attn_q.scale
create_tensor: loading tensor blk.53.attn_k.scale
create_tensor: loading tensor blk.53.attn_v.scale
create_tensor: loading tensor blk.53.attn_output.scale
create_tensor: loading tensor blk.53.ffn_gate.scale
create_tensor: loading tensor blk.53.ffn_down.scale
create_tensor: loading tensor blk.53.ffn_up.scale
create_tensor: loading tensor blk.53.attn_q.input_scale
create_tensor: loading tensor blk.53.attn_k.input_scale
create_tensor: loading tensor blk.53.attn_v.input_scale
create_tensor: loading tensor blk.53.attn_output.input_scale
create_tensor: loading tensor blk.53.ffn_gate.input_scale
create_tensor: loading tensor blk.53.ffn_down.input_scale
create_tensor: loading tensor blk.53.ffn_up.input_scale
create_tensor: loading tensor blk.54.attn_q.scale
create_tensor: loading tensor blk.54.attn_k.scale
create_tensor: loading tensor blk.54.attn_v.scale
create_tensor: loading tensor blk.54.attn_output.scale
create_tensor: loading tensor blk.54.ffn_gate.scale
create_tensor: loading tensor blk.54.ffn_down.scale
create_tensor: loading tensor blk.54.ffn_up.scale
create_tensor: loading tensor blk.54.attn_q.input_scale
create_tensor: loading tensor blk.54.attn_k.input_scale
create_tensor: loading tensor blk.54.attn_v.input_scale
create_tensor: loading tensor blk.54.attn_output.input_scale
create_tensor: loading tensor blk.54.ffn_gate.input_scale
create_tensor: loading tensor blk.54.ffn_down.input_scale
create_tensor: loading tensor blk.54.ffn_up.input_scale
create_tensor: loading tensor blk.55.attn_q.scale
create_tensor: loading tensor blk.55.attn_k.scale
create_tensor: loading tensor blk.55.attn_v.scale
create_tensor: loading tensor blk.55.attn_output.scale
create_tensor: loading tensor blk.55.ffn_gate.scale
create_tensor: loading tensor blk.55.ffn_down.scale
create_tensor: loading tensor blk.55.ffn_up.scale
create_tensor: loading tensor blk.55.attn_q.input_scale
create_tensor: loading tensor blk.55.attn_k.input_scale
create_tensor: loading tensor blk.55.attn_v.input_scale
create_tensor: loading tensor blk.55.attn_output.input_scale
create_tensor: loading tensor blk.55.ffn_gate.input_scale
create_tensor: loading tensor blk.55.ffn_down.input_scale
create_tensor: loading tensor blk.55.ffn_up.input_scale
create_tensor: loading tensor blk.56.attn_q.scale
create_tensor: loading tensor blk.56.attn_k.scale
create_tensor: loading tensor blk.56.attn_v.scale
create_tensor: loading tensor blk.56.attn_output.scale
create_tensor: loading tensor blk.56.ffn_gate.scale
create_tensor: loading tensor blk.56.ffn_down.scale
create_tensor: loading tensor blk.56.ffn_up.scale
create_tensor: loading tensor blk.56.attn_q.input_scale
create_tensor: loading tensor blk.56.attn_k.input_scale
create_tensor: loading tensor blk.56.attn_v.input_scale
create_tensor: loading tensor blk.56.attn_output.input_scale
create_tensor: loading tensor blk.56.ffn_gate.input_scale
create_tensor: loading tensor blk.56.ffn_down.input_scale
create_tensor: loading tensor blk.56.ffn_up.input_scale
create_tensor: loading tensor blk.57.attn_q.scale
create_tensor: loading tensor blk.57.attn_k.scale
create_tensor: loading tensor blk.57.attn_v.scale
create_tensor: loading tensor blk.57.attn_output.scale
create_tensor: loading tensor blk.57.ffn_gate.scale
create_tensor: loading tensor blk.57.ffn_down.scale
create_tensor: loading tensor blk.57.ffn_up.scale
create_tensor: loading tensor blk.57.attn_q.input_scale
create_tensor: loading tensor blk.57.attn_k.input_scale
create_tensor: loading tensor blk.57.attn_v.input_scale
create_tensor: loading tensor blk.57.attn_output.input_scale
create_tensor: loading tensor blk.57.ffn_gate.input_scale
create_tensor: loading tensor blk.57.ffn_down.input_scale
create_tensor: loading tensor blk.57.ffn_up.input_scale
create_tensor: loading tensor blk.58.attn_q.scale
create_tensor: loading tensor blk.58.attn_k.scale
create_tensor: loading tensor blk.58.attn_v.scale
create_tensor: loading tensor blk.58.attn_output.scale
create_tensor: loading tensor blk.58.ffn_gate.scale
create_tensor: loading tensor blk.58.ffn_down.scale
create_tensor: loading tensor blk.58.ffn_up.scale
create_tensor: loading tensor blk.58.attn_q.input_scale
create_tensor: loading tensor blk.58.attn_k.input_scale
create_tensor: loading tensor blk.58.attn_v.input_scale
create_tensor: loading tensor blk.58.attn_output.input_scale
create_tensor: loading tensor blk.58.ffn_gate.input_scale
create_tensor: loading tensor blk.58.ffn_down.input_scale
create_tensor: loading tensor blk.58.ffn_up.input_scale
create_tensor: loading tensor blk.59.attn_q.scale
create_tensor: loading tensor blk.59.attn_k.scale
create_tensor: loading tensor blk.59.attn_v.scale
create_tensor: loading tensor blk.59.attn_output.scale
create_tensor: loading tensor blk.59.ffn_gate.scale
create_tensor: loading tensor blk.59.ffn_down.scale
create_tensor: loading tensor blk.59.ffn_up.scale
create_tensor: loading tensor blk.59.attn_q.input_scale
create_tensor: loading tensor blk.59.attn_k.input_scale
create_tensor: loading tensor blk.59.attn_v.input_scale
create_tensor: loading tensor blk.59.attn_output.input_scale
create_tensor: loading tensor blk.59.ffn_gate.input_scale
create_tensor: loading tensor blk.59.ffn_down.input_scale
create_tensor: loading tensor blk.59.ffn_up.input_scale
create_tensor: loading tensor blk.60.attn_q.scale
create_tensor: loading tensor blk.60.attn_k.scale
create_tensor: loading tensor blk.60.attn_v.scale
create_tensor: loading tensor blk.60.attn_output.scale
create_tensor: loading tensor blk.60.ffn_gate.scale
create_tensor: loading tensor blk.60.ffn_down.scale
create_tensor: loading tensor blk.60.ffn_up.scale
create_tensor: loading tensor blk.60.attn_q.input_scale
create_tensor: loading tensor blk.60.attn_k.input_scale
create_tensor: loading tensor blk.60.attn_v.input_scale
create_tensor: loading tensor blk.60.attn_output.input_scale
create_tensor: loading tensor blk.60.ffn_gate.input_scale
create_tensor: loading tensor blk.60.ffn_down.input_scale
create_tensor: loading tensor blk.60.ffn_up.input_scale
create_tensor: loading tensor blk.61.attn_q.scale
create_tensor: loading tensor blk.61.attn_k.scale
create_tensor: loading tensor blk.61.attn_v.scale
create_tensor: loading tensor blk.61.attn_output.scale
create_tensor: loading tensor blk.61.ffn_gate.scale
create_tensor: loading tensor blk.61.ffn_down.scale
create_tensor: loading tensor blk.61.ffn_up.scale
create_tensor: loading tensor blk.61.attn_q.input_scale
create_tensor: loading tensor blk.61.attn_k.input_scale
create_tensor: loading tensor blk.61.attn_v.input_scale
create_tensor: loading tensor blk.61.attn_output.input_scale
create_tensor: loading tensor blk.61.ffn_gate.input_scale
create_tensor: loading tensor blk.61.ffn_down.input_scale
create_tensor: loading tensor blk.61.ffn_up.input_scale
create_tensor: loading tensor blk.62.attn_q.scale
create_tensor: loading tensor blk.62.attn_k.scale
create_tensor: loading tensor blk.62.attn_v.scale
create_tensor: loading tensor blk.62.attn_output.scale
create_tensor: loading tensor blk.62.ffn_gate.scale
create_tensor: loading tensor blk.62.ffn_down.scale
create_tensor: loading tensor blk.62.ffn_up.scale
create_tensor: loading tensor blk.62.attn_q.input_scale
create_tensor: loading tensor blk.62.attn_k.input_scale
create_tensor: loading tensor blk.62.attn_v.input_scale
create_tensor: loading tensor blk.62.attn_output.input_scale
create_tensor: loading tensor blk.62.ffn_gate.input_scale
create_tensor: loading tensor blk.62.ffn_down.input_scale
create_tensor: loading tensor blk.62.ffn_up.input_scale
create_tensor: loading tensor blk.63.attn_q.scale
create_tensor: loading tensor blk.63.attn_k.scale
create_tensor: loading tensor blk.63.attn_v.scale
create_tensor: loading tensor blk.63.attn_output.scale
create_tensor: loading tensor blk.63.ffn_gate.scale
create_tensor: loading tensor blk.63.ffn_down.scale
create_tensor: loading tensor blk.63.ffn_up.scale
create_tensor: loading tensor blk.63.attn_q.input_scale
create_tensor: loading tensor blk.63.attn_k.input_scale
create_tensor: loading tensor blk.63.attn_v.input_scale
create_tensor: loading tensor blk.63.attn_output.input_scale
create_tensor: loading tensor blk.63.ffn_gate.input_scale
create_tensor: loading tensor blk.63.ffn_down.input_scale
create_tensor: loading tensor blk.63.ffn_up.input_scale
done_getting_tensors: tensor 'token_embd.weight' (q4_K) (and 0 others) cannot be used with preferred buffer type CUDA_Host, using CUDA0 instead
load_tensors: offloading output layer to GPU
load_tensors: offloading 63 repeating layers to GPU
load_tensors: offloaded 65/65 layers to GPU
load_tensors:        CUDA0 model buffer size = 20748.00 MiB
.................................................................................................
llama_context: constructing llama_context
llama_context: n_seq_max             = 1
llama_context: n_ctx                 = 16384
llama_context: n_ctx_seq             = 16384
llama_context: n_batch               = 512
llama_context: n_ubatch              = 512
llama_context: causal_attn           = 1
llama_context: flash_attn            = enabled
llama_context: kv_unified            = false
llama_context: freq_base             = 10000000.0
llama_context: freq_scale            = 1
llama_context: n_rs_seq              = 0
llama_context: n_outputs_max         = 512
llama_context: n_outputs_max_per_seq = 1
llama_context: n_ctx_seq (16384) < n_ctx_train (524288) -- the full capacity of the model will not be utilized
set_abort_callback: call
llama_context:  CUDA_Host  output buffer size =     0.59 MiB
llama_kv_cache: layer   0: dev = CUDA0
llama_kv_cache: layer   1: dev = CUDA0
llama_kv_cache: layer   2: dev = CUDA0
llama_kv_cache: layer   3: dev = CUDA0
llama_kv_cache: layer   4: dev = CUDA0
llama_kv_cache: layer   5: dev = CUDA0
llama_kv_cache: layer   6: dev = CUDA0
llama_kv_cache: layer   7: dev = CUDA0
llama_kv_cache: layer   8: dev = CUDA0
llama_kv_cache: layer   9: dev = CUDA0
llama_kv_cache: layer  10: dev = CUDA0
llama_kv_cache: layer  11: dev = CUDA0
llama_kv_cache: layer  12: dev = CUDA0
llama_kv_cache: layer  13: dev = CUDA0
llama_kv_cache: layer  14: dev = CUDA0
llama_kv_cache: layer  15: dev = CUDA0
llama_kv_cache: layer  16: dev = CUDA0
llama_kv_cache: layer  17: dev = CUDA0
llama_kv_cache: layer  18: dev = CUDA0
llama_kv_cache: layer  19: dev = CUDA0
llama_kv_cache: layer  20: dev = CUDA0
llama_kv_cache: layer  21: dev = CUDA0
llama_kv_cache: layer  22: dev = CUDA0
llama_kv_cache: layer  23: dev = CUDA0
llama_kv_cache: layer  24: dev = CUDA0
llama_kv_cache: layer  25: dev = CUDA0
llama_kv_cache: layer  26: dev = CUDA0
llama_kv_cache: layer  27: dev = CUDA0
llama_kv_cache: layer  28: dev = CUDA0
llama_kv_cache: layer  29: dev = CUDA0
llama_kv_cache: layer  30: dev = CUDA0
llama_kv_cache: layer  31: dev = CUDA0
llama_kv_cache: layer  32: dev = CUDA0
llama_kv_cache: layer  33: dev = CUDA0
llama_kv_cache: layer  34: dev = CUDA0
llama_kv_cache: layer  35: dev = CUDA0
llama_kv_cache: layer  36: dev = CUDA0
llama_kv_cache: layer  37: dev = CUDA0
llama_kv_cache: layer  38: dev = CUDA0
llama_kv_cache: layer  39: dev = CUDA0
llama_kv_cache: layer  40: dev = CUDA0
llama_kv_cache: layer  41: dev = CUDA0
llama_kv_cache: layer  42: dev = CUDA0
llama_kv_cache: layer  43: dev = CUDA0
llama_kv_cache: layer  44: dev = CUDA0
llama_kv_cache: layer  45: dev = CUDA0
llama_kv_cache: layer  46: dev = CUDA0
llama_kv_cache: layer  47: dev = CUDA0
llama_kv_cache: layer  48: dev = CUDA0
llama_kv_cache: layer  49: dev = CUDA0
llama_kv_cache: layer  50: dev = CUDA0
llama_kv_cache: layer  51: dev = CUDA0
llama_kv_cache: layer  52: dev = CUDA0
llama_kv_cache: layer  53: dev = CUDA0
llama_kv_cache: layer  54: dev = CUDA0
llama_kv_cache: layer  55: dev = CUDA0
llama_kv_cache: layer  56: dev = CUDA0
llama_kv_cache: layer  57: dev = CUDA0
llama_kv_cache: layer  58: dev = CUDA0
llama_kv_cache: layer  59: dev = CUDA0
llama_kv_cache: layer  60: dev = CUDA0
llama_kv_cache: layer  61: dev = CUDA0
llama_kv_cache: layer  62: dev = CUDA0
llama_kv_cache: layer  63: dev = CUDA0
llama_kv_cache:      CUDA0 KV buffer size =  4096.00 MiB
llama_kv_cache: size = 4096.00 MiB ( 16384 cells,  64 layers,  1/1 seqs), K (f16): 2048.00 MiB, V (f16): 2048.00 MiB
llama_kv_cache: attn_rot_k = 0, n_embd_head_k_all = 128
llama_kv_cache: attn_rot_v = 0, n_embd_head_k_all = 128
llama_context: enumerating backends
llama_context: backend_ptrs.size() = 2
sched_reserve: reserving ...
sched_reserve: max_nodes = 6168
sched_reserve: reserving full memory module
sched_reserve: worst-case: n_tokens = 512, n_seqs = 1, n_outputs = 1
resolve_fused_ops: resolving fused DeepSeek V4 HC support:
graph_reserve: reserving a graph for ubatch with n_tokens =    1, n_seqs =  1, n_outputs =    1
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l0 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l1 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l2 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l3 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l4 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l5 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l6 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l7 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l8 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l9 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l10 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l11 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l12 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l13 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l14 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l15 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l16 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l17 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l18 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l19 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l20 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l21 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l22 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l23 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l24 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l25 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l26 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l27 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l28 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l29 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l30 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l31 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l32 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l33 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l34 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l35 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l36 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l37 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l38 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l39 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l40 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l41 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l42 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l43 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l44 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l45 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l46 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l47 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l48 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l49 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l50 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l51 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l52 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l53 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l54 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l55 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l56 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l57 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l58 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l59 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l60 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l61 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l62 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l63 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                      inp_tokens' [i32, ne = {     1,     1,     1,     1 }] is used by node 'embd' (GET_ROWS)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-0' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-0' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-1' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-1' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-2' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-2' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-3' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-3' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-4' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-4' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-5' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-5' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-6' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-6' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-7' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-7' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-8' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-8' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-9' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-9' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-10' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-10' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-11' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-11' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-12' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-12' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-13' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-13' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-14' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-14' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-15' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-15' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-16' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-16' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-17' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-17' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-18' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-18' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-19' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-19' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-20' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-20' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-21' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-21' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-22' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-22' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-23' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-23' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-24' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-24' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-25' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-25' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-26' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-26' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-27' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-27' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-28' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-28' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-29' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-29' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-30' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-30' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-31' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-31' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-32' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-32' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-33' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-33' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-34' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-34' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-35' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-35' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-36' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-36' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-37' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-37' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-38' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-38' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-39' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-39' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-40' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-40' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-41' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-41' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-42' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-42' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-43' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-43' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-44' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-44' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-45' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-45' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-46' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-46' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-47' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-47' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-48' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-48' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-49' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-49' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-50' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-50' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-51' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-51' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-52' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-52' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-53' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-53' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-54' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-54' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-55' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-55' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-56' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-56' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-57' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-57' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-58' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-58' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-59' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-59' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-60' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-60' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-61' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-61' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-62' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-62' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-63' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-63' (ROPE)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l0 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l1 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l2 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l3 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l4 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l5 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l6 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l7 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l8 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l9 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l10 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l11 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l12 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l13 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l14 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l15 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l16 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l17 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l18 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l19 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l20 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l21 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l22 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l23 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l24 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l25 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l26 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l27 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l28 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l29 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l30 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l31 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l32 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l33 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l34 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l35 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l36 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l37 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l38 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l39 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l40 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l41 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l42 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l43 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l44 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l45 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l46 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l47 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l48 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l49 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l50 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l51 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l52 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l53 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l54 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l55 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l56 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l57 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l58 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l59 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l60 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l61 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l62 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l63 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_24' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_58' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_92' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_126' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_160' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_194' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_228' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_262' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_296' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_330' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_364' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_398' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_432' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_466' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_500' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_534' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_568' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_602' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_636' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_670' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_704' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_738' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_772' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_806' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_840' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_874' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_908' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_942' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_976' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_1010' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_1044' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_1078' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_1112' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_1146' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_1180' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_1214' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_1248' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_1282' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_1316' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_1350' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_1384' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_1418' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_1452' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_1486' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_1520' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_1554' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_1588' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_1622' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_1656' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_1690' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_1724' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_1758' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_1792' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_1826' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_1860' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_1894' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_1928' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_1962' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_1996' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_2030' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_2064' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_2098' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_2132' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_2166' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                         out_ids' [i32, ne = {     1,     1,     1,     1 }] is used by node 'node_2169' (GET_ROWS)
llama_graph_n_input_tensors: input tensor '                         out_ids' [i32, ne = {     1,     1,     1,     1 }] is used by node 'node_2170' (GET_ROWS)
resolve_fused_ops: fused DeepSeek V4 HC pre enabled
graph_reserve: reserving a graph for ubatch with n_tokens =    1, n_seqs =  1, n_outputs =    1
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l0 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l1 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l2 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l3 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l4 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l5 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l6 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l7 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l8 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l9 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l10 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l11 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l12 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l13 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l14 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l15 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l16 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l17 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l18 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l19 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l20 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l21 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l22 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l23 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l24 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l25 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l26 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l27 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l28 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l29 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l30 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l31 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l32 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l33 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l34 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l35 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l36 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l37 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l38 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l39 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l40 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l41 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l42 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l43 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l44 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l45 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l46 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l47 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l48 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l49 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l50 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l51 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l52 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l53 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l54 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l55 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l56 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l57 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l58 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l59 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l60 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l61 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l62 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l63 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                      inp_tokens' [i32, ne = {     1,     1,     1,     1 }] is used by node 'embd' (GET_ROWS)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-0' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-0' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-1' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-1' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-2' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-2' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-3' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-3' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-4' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-4' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-5' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-5' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-6' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-6' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-7' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-7' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-8' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-8' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-9' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-9' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-10' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-10' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-11' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-11' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-12' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-12' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-13' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-13' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-14' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-14' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-15' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-15' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-16' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-16' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-17' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-17' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-18' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-18' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-19' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-19' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-20' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-20' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-21' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-21' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-22' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-22' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-23' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-23' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-24' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-24' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-25' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-25' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-26' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-26' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-27' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-27' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-28' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-28' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-29' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-29' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-30' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-30' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-31' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-31' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-32' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-32' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-33' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-33' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-34' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-34' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-35' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-35' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-36' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-36' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-37' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-37' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-38' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-38' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-39' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-39' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-40' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-40' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-41' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-41' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-42' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-42' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-43' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-43' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-44' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-44' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-45' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-45' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-46' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-46' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-47' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-47' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-48' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-48' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-49' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-49' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-50' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-50' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-51' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-51' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-52' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-52' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-53' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-53' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-54' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-54' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-55' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-55' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-56' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-56' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-57' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-57' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-58' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-58' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-59' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-59' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-60' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-60' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-61' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-61' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-62' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-62' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-63' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-63' (ROPE)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l0 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l1 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l2 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l3 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l4 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l5 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l6 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l7 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l8 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l9 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l10 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l11 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l12 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l13 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l14 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l15 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l16 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l17 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l18 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l19 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l20 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l21 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l22 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l23 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l24 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l25 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l26 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l27 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l28 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l29 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l30 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l31 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l32 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l33 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l34 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l35 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l36 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l37 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l38 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l39 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l40 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l41 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l42 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l43 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l44 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l45 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l46 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l47 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l48 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l49 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l50 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l51 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l52 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l53 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l54 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l55 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l56 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l57 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l58 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l59 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l60 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l61 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l62 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l63 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_24' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_58' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_92' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_126' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_160' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_194' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_228' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_262' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_296' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_330' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_364' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_398' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_432' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_466' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_500' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_534' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_568' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_602' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_636' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_670' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_704' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_738' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_772' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_806' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_840' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_874' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_908' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_942' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_976' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_1010' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_1044' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_1078' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_1112' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_1146' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_1180' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_1214' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_1248' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_1282' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_1316' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_1350' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_1384' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_1418' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_1452' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_1486' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_1520' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_1554' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_1588' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_1622' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_1656' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_1690' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_1724' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_1758' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_1792' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_1826' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_1860' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_1894' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_1928' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_1962' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_1996' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_2030' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_2064' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_2098' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_2132' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_2166' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                         out_ids' [i32, ne = {     1,     1,     1,     1 }] is used by node 'node_2169' (GET_ROWS)
llama_graph_n_input_tensors: input tensor '                         out_ids' [i32, ne = {     1,     1,     1,     1 }] is used by node 'node_2170' (GET_ROWS)
resolve_fused_ops: fused DeepSeek V4 HC comb enabled
graph_reserve: reserving a graph for ubatch with n_tokens =    1, n_seqs =  1, n_outputs =    1
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l0 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l1 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l2 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l3 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l4 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l5 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l6 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l7 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l8 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l9 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l10 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l11 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l12 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l13 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l14 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l15 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l16 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l17 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l18 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l19 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l20 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l21 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l22 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l23 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l24 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l25 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l26 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l27 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l28 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l29 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l30 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l31 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l32 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l33 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l34 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l35 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l36 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l37 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l38 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l39 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l40 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l41 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l42 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l43 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l44 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l45 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l46 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l47 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l48 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l49 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l50 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l51 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l52 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l53 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l54 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l55 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l56 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l57 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l58 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l59 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l60 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l61 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l62 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l63 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                      inp_tokens' [i32, ne = {     1,     1,     1,     1 }] is used by node 'embd' (GET_ROWS)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-0' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-0' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-1' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-1' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-2' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-2' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-3' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-3' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-4' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-4' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-5' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-5' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-6' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-6' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-7' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-7' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-8' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-8' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-9' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-9' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-10' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-10' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-11' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-11' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-12' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-12' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-13' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-13' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-14' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-14' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-15' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-15' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-16' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-16' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-17' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-17' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-18' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-18' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-19' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-19' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-20' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-20' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-21' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-21' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-22' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-22' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-23' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-23' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-24' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-24' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-25' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-25' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-26' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-26' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-27' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-27' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-28' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-28' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-29' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-29' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-30' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-30' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-31' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-31' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-32' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-32' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-33' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-33' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-34' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-34' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-35' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-35' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-36' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-36' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-37' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-37' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-38' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-38' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-39' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-39' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-40' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-40' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-41' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-41' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-42' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-42' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-43' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-43' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-44' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-44' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-45' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-45' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-46' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-46' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-47' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-47' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-48' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-48' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-49' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-49' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-50' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-50' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-51' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-51' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-52' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-52' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-53' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-53' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-54' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-54' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-55' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-55' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-56' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-56' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-57' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-57' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-58' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-58' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-59' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-59' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-60' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-60' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-61' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-61' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-62' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-62' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-63' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-63' (ROPE)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l0 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l1 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l2 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l3 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l4 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l5 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l6 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l7 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l8 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l9 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l10 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l11 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l12 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l13 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l14 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l15 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l16 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l17 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l18 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l19 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l20 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l21 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l22 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l23 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l24 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l25 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l26 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l27 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l28 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l29 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l30 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l31 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l32 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l33 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l34 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l35 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l36 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l37 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l38 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l39 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l40 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l41 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l42 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l43 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l44 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l45 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l46 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l47 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l48 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l49 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l50 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l51 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l52 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l53 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l54 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l55 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l56 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l57 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l58 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l59 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l60 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l61 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l62 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l63 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_24' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_58' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_92' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_126' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_160' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_194' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_228' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_262' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_296' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_330' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_364' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_398' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_432' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_466' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_500' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_534' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_568' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_602' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_636' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_670' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_704' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_738' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_772' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_806' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_840' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_874' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_908' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_942' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_976' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_1010' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_1044' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_1078' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_1112' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_1146' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_1180' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_1214' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_1248' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_1282' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_1316' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_1350' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_1384' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_1418' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_1452' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_1486' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_1520' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_1554' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_1588' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_1622' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_1656' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_1690' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_1724' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_1758' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_1792' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_1826' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_1860' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_1894' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_1928' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_1962' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_1996' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_2030' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_2064' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_2098' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_2132' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_2166' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                         out_ids' [i32, ne = {     1,     1,     1,     1 }] is used by node 'node_2169' (GET_ROWS)
llama_graph_n_input_tensors: input tensor '                         out_ids' [i32, ne = {     1,     1,     1,     1 }] is used by node 'node_2170' (GET_ROWS)
resolve_fused_ops: fused DeepSeek V4 HC post enabled
graph_reserve: reserving a graph for ubatch with n_tokens =  512, n_seqs =  1, n_outputs =  512
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l0 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l1 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l2 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l3 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l4 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l5 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l6 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l7 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l8 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l9 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l10 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l11 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l12 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l13 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l14 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l15 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l16 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l17 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l18 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l19 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l20 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l21 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l22 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l23 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l24 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l25 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l26 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l27 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l28 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l29 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l30 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l31 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l32 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l33 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l34 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l35 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l36 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l37 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l38 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l39 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l40 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l41 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l42 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l43 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l44 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l45 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l46 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l47 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l48 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l49 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l50 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l51 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l52 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l53 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l54 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l55 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l56 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l57 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l58 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l59 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l60 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l61 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l62 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l63 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                      inp_tokens' [i32, ne = {   512,     1,     1,     1 }] is used by node 'embd' (GET_ROWS)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-0' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-0' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-1' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-1' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-2' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-2' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-3' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-3' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-4' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-4' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-5' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-5' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-6' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-6' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-7' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-7' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-8' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-8' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-9' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-9' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-10' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-10' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-11' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-11' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-12' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-12' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-13' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-13' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-14' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-14' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-15' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-15' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-16' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-16' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-17' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-17' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-18' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-18' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-19' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-19' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-20' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-20' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-21' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-21' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-22' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-22' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-23' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-23' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-24' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-24' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-25' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-25' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-26' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-26' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-27' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-27' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-28' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-28' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-29' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-29' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-30' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-30' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-31' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-31' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-32' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-32' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-33' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-33' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-34' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-34' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-35' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-35' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-36' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-36' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-37' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-37' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-38' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-38' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-39' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-39' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-40' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-40' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-41' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-41' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-42' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-42' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-43' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-43' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-44' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-44' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-45' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-45' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-46' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-46' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-47' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-47' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-48' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-48' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-49' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-49' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-50' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-50' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-51' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-51' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-52' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-52' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-53' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-53' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-54' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-54' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-55' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-55' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-56' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-56' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-57' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-57' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-58' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-58' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-59' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-59' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-60' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-60' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-61' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-61' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-62' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-62' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-63' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-63' (ROPE)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l0 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l1 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l2 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l3 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l4 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l5 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l6 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l7 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l8 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l9 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l10 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l11 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l12 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l13 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l14 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l15 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l16 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l17 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l18 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l19 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l20 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l21 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l22 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l23 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l24 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l25 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l26 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l27 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l28 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l29 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l30 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l31 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l32 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l33 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l34 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l35 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l36 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l37 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l38 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l39 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l40 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l41 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l42 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l43 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l44 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l45 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l46 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l47 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l48 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l49 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l50 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l51 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l52 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l53 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l54 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l55 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l56 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l57 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l58 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l59 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l60 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l61 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l62 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l63 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_24' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_58' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_92' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_126' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_160' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_194' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_228' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_262' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_296' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_330' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_364' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_398' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_432' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_466' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_500' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_534' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_568' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_602' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_636' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_670' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_704' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_738' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_772' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_806' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_840' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_874' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_908' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_942' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_976' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_1010' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_1044' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_1078' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_1112' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_1146' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_1180' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_1214' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_1248' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_1282' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_1316' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_1350' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_1384' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_1418' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_1452' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_1486' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_1520' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_1554' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_1588' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_1622' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_1656' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_1690' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_1724' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_1758' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_1792' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_1826' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_1860' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_1894' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_1928' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_1962' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_1996' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_2030' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_2064' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_2098' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_2132' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_2166' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                         out_ids' [i32, ne = {   512,     1,     1,     1 }] is used by node 'node_2169' (GET_ROWS)
llama_graph_n_input_tensors: input tensor '                         out_ids' [i32, ne = {   512,     1,     1,     1 }] is used by node 'node_2170' (GET_ROWS)
graph_reserve: reserving a graph for ubatch with n_tokens =    1, n_seqs =  1, n_outputs =    1
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l0 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l1 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l2 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l3 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l4 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l5 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l6 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l7 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l8 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l9 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l10 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l11 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l12 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l13 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l14 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l15 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l16 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l17 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l18 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l19 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l20 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l21 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l22 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l23 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l24 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l25 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l26 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l27 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l28 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l29 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l30 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l31 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l32 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l33 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l34 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l35 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l36 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l37 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l38 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l39 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l40 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l41 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l42 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l43 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l44 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l45 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l46 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l47 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l48 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l49 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l50 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l51 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l52 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l53 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l54 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l55 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l56 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l57 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l58 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l59 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l60 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l61 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l62 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_v_l63 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                      inp_tokens' [i32, ne = {     1,     1,     1,     1 }] is used by node 'embd' (GET_ROWS)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-0' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-0' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-1' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-1' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-2' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-2' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-3' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-3' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-4' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-4' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-5' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-5' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-6' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-6' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-7' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-7' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-8' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-8' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-9' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-9' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-10' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-10' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-11' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-11' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-12' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-12' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-13' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-13' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-14' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-14' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-15' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-15' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-16' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-16' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-17' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-17' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-18' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-18' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-19' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-19' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-20' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-20' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-21' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-21' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-22' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-22' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-23' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-23' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-24' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-24' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-25' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-25' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-26' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-26' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-27' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-27' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-28' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-28' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-29' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-29' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-30' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-30' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-31' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-31' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-32' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-32' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-33' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-33' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-34' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-34' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-35' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-35' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-36' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-36' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-37' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-37' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-38' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-38' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-39' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-39' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-40' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-40' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-41' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-41' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-42' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-42' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-43' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-43' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-44' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-44' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-45' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-45' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-46' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-46' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-47' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-47' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-48' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-48' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-49' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-49' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-50' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-50' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-51' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-51' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-52' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-52' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-53' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-53' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-54' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-54' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-55' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-55' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-56' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-56' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-57' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-57' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-58' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-58' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-59' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-59' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-60' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-60' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-61' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-61' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-62' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-62' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Qcur-63' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {     1,     1,     1,     1 }] is used by node 'Kcur-63' (ROPE)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l0 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l1 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l2 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l3 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l4 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l5 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l6 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l7 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l8 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l9 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l10 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l11 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l12 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l13 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l14 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l15 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l16 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l17 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l18 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l19 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l20 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l21 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l22 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l23 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l24 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l25 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l26 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l27 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l28 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l29 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l30 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l31 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l32 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l33 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l34 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l35 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l36 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l37 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l38 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l39 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l40 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l41 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l42 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l43 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l44 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l45 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l46 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l47 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l48 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l49 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l50 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l51 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l52 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l53 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l54 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l55 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l56 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l57 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l58 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l59 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l60 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l61 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l62 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {     1,     1,     1,     1 }] is used by node 'cache_k_l63 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_24' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_58' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_92' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_126' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_160' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_194' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_228' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_262' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_296' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_330' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_364' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_398' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_432' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_466' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_500' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_534' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_568' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_602' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_636' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_670' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_704' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_738' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_772' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_806' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_840' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_874' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_908' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_942' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_976' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_1010' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_1044' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_1078' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_1112' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_1146' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_1180' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_1214' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_1248' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_1282' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_1316' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_1350' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_1384' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_1418' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_1452' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_1486' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_1520' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_1554' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_1588' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_1622' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_1656' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_1690' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_1724' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_1758' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_1792' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_1826' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_1860' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_1894' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_1928' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_1962' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_1996' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_2030' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_2064' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_2098' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_2132' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,     1,     1,     1 }] is used by node 'node_2166' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                         out_ids' [i32, ne = {     1,     1,     1,     1 }] is used by node 'node_2169' (GET_ROWS)
llama_graph_n_input_tensors: input tensor '                         out_ids' [i32, ne = {     1,     1,     1,     1 }] is used by node 'node_2170' (GET_ROWS)
graph_reserve: reserving a graph for ubatch with n_tokens =  512, n_seqs =  1, n_outputs =  512
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l0 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l1 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l2 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l3 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l4 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l5 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l6 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l7 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l8 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l9 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l10 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l11 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l12 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l13 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l14 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l15 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l16 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l17 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l18 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l19 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l20 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l21 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l22 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l23 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l24 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l25 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l26 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l27 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l28 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l29 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l30 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l31 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l32 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l33 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l34 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l35 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l36 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l37 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l38 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l39 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l40 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l41 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l42 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l43 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l44 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l45 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l46 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l47 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l48 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l49 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l50 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l51 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l52 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l53 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l54 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l55 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l56 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l57 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l58 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l59 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l60 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l61 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l62 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_v_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_v_l63 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                      inp_tokens' [i32, ne = {   512,     1,     1,     1 }] is used by node 'embd' (GET_ROWS)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-0' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-0' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-1' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-1' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-2' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-2' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-3' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-3' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-4' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-4' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-5' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-5' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-6' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-6' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-7' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-7' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-8' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-8' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-9' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-9' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-10' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-10' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-11' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-11' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-12' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-12' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-13' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-13' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-14' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-14' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-15' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-15' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-16' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-16' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-17' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-17' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-18' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-18' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-19' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-19' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-20' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-20' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-21' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-21' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-22' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-22' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-23' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-23' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-24' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-24' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-25' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-25' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-26' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-26' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-27' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-27' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-28' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-28' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-29' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-29' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-30' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-30' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-31' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-31' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-32' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-32' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-33' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-33' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-34' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-34' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-35' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-35' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-36' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-36' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-37' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-37' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-38' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-38' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-39' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-39' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-40' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-40' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-41' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-41' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-42' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-42' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-43' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-43' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-44' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-44' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-45' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-45' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-46' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-46' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-47' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-47' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-48' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-48' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-49' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-49' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-50' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-50' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-51' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-51' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-52' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-52' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-53' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-53' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-54' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-54' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-55' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-55' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-56' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-56' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-57' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-57' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-58' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-58' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-59' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-59' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-60' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-60' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-61' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-61' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-62' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-62' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Qcur-63' (ROPE)
llama_graph_n_input_tensors: input tensor '                         inp_pos' [i32, ne = {   512,     1,     1,     1 }] is used by node 'Kcur-63' (ROPE)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l0 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l1 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l2 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l3 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l4 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l5 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l6 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l7 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l8 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l9 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l10 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l11 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l12 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l13 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l14 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l15 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l16 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l17 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l18 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l19 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l20 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l21 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l22 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l23 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l24 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l25 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l26 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l27 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l28 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l29 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l30 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l31 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l32 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l33 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l34 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l35 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l36 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l37 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l38 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l39 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l40 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l41 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l42 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l43 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l44 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l45 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l46 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l47 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l48 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l49 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l50 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l51 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l52 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l53 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l54 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l55 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l56 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l57 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l58 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l59 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l60 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l61 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l62 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                 attn_inp_k_idxs' [i64, ne = {   512,     1,     1,     1 }] is used by node 'cache_k_l63 (view)' (SET_ROWS)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_24' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_58' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_92' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_126' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_160' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_194' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_228' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_262' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_296' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_330' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_364' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_398' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_432' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_466' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_500' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_534' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_568' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_602' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_636' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_670' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_704' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_738' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_772' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_806' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_840' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_874' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_908' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_942' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_976' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_1010' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_1044' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_1078' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_1112' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_1146' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_1180' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_1214' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_1248' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_1282' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_1316' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_1350' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_1384' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_1418' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_1452' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_1486' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_1520' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_1554' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_1588' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_1622' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_1656' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_1690' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_1724' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_1758' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_1792' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_1826' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_1860' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_1894' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_1928' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_1962' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_1996' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_2030' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_2064' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_2098' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_2132' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                attn_inp_kq_mask' [f16, ne = { 16384,   512,     1,     1 }] is used by node 'node_2166' (FLASH_ATTN_EXT)
llama_graph_n_input_tensors: input tensor '                         out_ids' [i32, ne = {   512,     1,     1,     1 }] is used by node 'node_2169' (GET_ROWS)
llama_graph_n_input_tensors: input tensor '                         out_ids' [i32, ne = {   512,     1,     1,     1 }] is used by node 'node_2170' (GET_ROWS)
sched_reserve:      CUDA0 compute buffer size =   313.00 MiB
sched_reserve:  CUDA_Host compute buffer size =    26.01 MiB
sched_reserve: graph: nodes = 2182, splits = 1, input objects = 4, input tensors = 6
sched_reserve: reserve took 15.47 ms, sched copies = 1
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
ggml_backend_cuda_graph_compute: CUDA graph warmup reset
ggml_backend_cuda_graph_compute: CUDA graph warmup complete
~llama_context:      CUDA0 compute buffer size is 313.0000 MiB, matches expectation of 313.0000 MiB
~llama_context:  CUDA_Host compute buffer size is  26.0137 MiB, matches expectation of  26.0137 MiB
