Go to file
tlg cef4041500 fix(projekt-matching): disable thinking for structured vLLM calls; retry non-JSON responses
With the vLLM qwen3 reasoning parser active, json_schema guided decoding
plus thinking degenerates: runs burn the whole 65k context
(finish_reason "length", ~63k completion tokens) and return empty or
truncated content. chat_template_kwargs {"enable_thinking": false} fixes
it (extract answers in ~35 s). Also: retry chat_json up to 3x on
non-JSON content, and two prompt calibrations verified against both
gate examples — Nice signal words bind only to their own line (following
unmarked lines stay Must), and compound requirements with clear evidence
for one part rate "unknown" instead of "no".

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-09 15:48:55 +02:00
Description
No description provided
207 KiB
Languages
Python 85.1%
Shell 14.9%