"""Local LLM integration — the voice timeline meets a local model. The voice timeline is *designed* to be handed to a language model: it is the source of truth between speech analysis and editing, layered so a model can reason about the narrative without parsing FCPXML. This module is the client side of that contract. It formats the timeline into the editar-por-voz brief, calls a local model server (Ollama, running Gemma 3 / Llama locally), and parses the model's decisions back into a validated list of VoiceActions — all inside the engine, so there is no wizard, no copy-paste, no manual step. Transport: Ollama's HTTP chat API at ``http://localhost:11434/api/chat``. Any model Ollama serves works; the default is Gemma 3 because that is what runs locally here ("Lama com Gema 3"), but pass ``model=`` to switch. The model is untrusted input: its JSON is validated row-by-row by :func:`fcpxml.voice_actions.parse_actions`, so one malformed decision never discards the edit. The brief is written so the model only ever emits the four action kinds the applier understands. """ from __future__ import annotations import json import logging import re from typing import Any, Dict, Optional, Sequence, Tuple import httpx from .voice_actions import parse_actions logger = logging.getLogger(__name__) DEFAULT_BASE_URL = "http://localhost:11434" # Gemma 3 12B reliably follows the editar-por-voz brief (keep the script, cut # only backstage chatter; the 4B variant skips the "keep the main content" # rule and deletes the script) but doesn't fit an 8GB machine. Qwen2.5 7B # instruct (q4_K_M) is the fallback for constrained hardware — strong at # strict JSON-schema following, the property this brief leans on hardest. # Pass ``model=`` to switch to whatever Ollama serves. DEFAULT_MODEL = "qwen2.5:7b-instruct-q4_K_M" REQUEST_TIMEOUT = 600.0 # The brief. Ported from the editar-por-voz skill criteria (criterios/01..08), # condensed into the instructions a model needs to emit valid actions. Kept in # Portuguese because the decisions and their reasons are read by a human editor. _SYSTEM_PROMPT = """Você é o editor de vídeo por voz deste sistema. Recebe um JSON de "linha do tempo de voz" — a medição de COMO foi falado (ênfase, energia, pausa, falante) de uma gravação — e devolve as DECISÕES de edição em JSON, nada mais. Você nunca escreve XML. Regras (siga rigorosamente): 1. LEIA EM CAMADAS. "summary" dá o formato da peça; "segments" é onde você trabalha (cada fala com seu texto e agregados); "segments[].words" dá o instante exato de cada destaque. Não recalcule energia, tom ou ênfase — use os números do JSON. 2. SEPARAR ROTEIRO DE BASTIDOR. - ROTEIRO = o conteúdo principal que a pessoa quer entregar: explicação, depoimento, roteiro decorado, a mensagem. É isso que VAI FICAR. - BASTIDOR = papo casual de gravação, cumprimentos, conversa com a equipe ("cara, beleza?", "tá gravando?", "deixa eu ver o celular"), piadas fora do assunto, tomadas interrompidas ou repetidas. É isso que VIRA "cut". Exemplo: num vídeo sobre mastopexia, a explicação da cirurgia É o roteiro (mantém); o "tá gravando? pois é" antes dela É bastidor (corta). Use "gap_before" e "take_boundary" (silêncio > ~3s = a câmera parou/recomeçou) para agrupar tomadas — eles marcam ONDE a tomada recomeça, não o que cortar. Nunca corte o conteúdo principal só porque tem ênfase; corte o casual/off-topic. REGRAS DE OURO: - MANTENHA o conteúdo principal (explicação, depoimento, roteiro decorado). Ele É o vídeo. - CORTE SÓ o casual/off-topic: cumprimentos, "tá gravando?", papo com a equipe, olhar o celular, repetições de tomada. - Em dúvida, MANTENHA a fala. É melhor sobrar conteúdo do que cortar o que era pra ficar. 3. ESCOLHER A MELHOR TOMADA de cada frase quando há repetições: mantenha a mais limpa e corte as outras (cut cobrindo a frase inteira). 4. CORTE (kind "cut"): para REMOVER uma frase, cubra ela inteira (start..end = início..fim da frase). Para APARAR só uma hesitação no começo ou fim, corte só da borda até a palavra (corte de meia frase é ambíguo — passe de 60% e apaga a linha toda). Nunca corte o silêncio entre falas. 5. ZOOM (kind "zoom"): só em palavra de CONTEÚDO bem enfatizada (emphasis alto, não artigo). params.scale entre 1.0 e 3.0 (padrão 1.3 se omitido). Posicione em torno da palavra, segurando até o fim da frase. 6. TEXTO (kind "text"): params.content obrigatório (≤120 chars), fixa um termo central ou callout. MARKER (kind "marker"): opcional params.content vira o nome do marcador. Use para emendas/junções que o editor deve conferir. 7. TEMPOS em segundos da MÍDIA ORIGINAL (exatamente como no JSON). Nunca compense para "depois do corte" — o programa desloca sozinho. end sempre > start, ambos ≥ 0. 8. reason OBRIGATÓRIO em cada ação, em português, embasando a decisão (ex.: 'abertura: "Aquela mama" (ênfase 0.42)'). reason vazio é decisão sem critério. Responda APENAS com um objeto JSON válido, sem markdown, sem comentário: {"source": "", "actions": [{"kind": "cut|zoom|text|marker", "start": , "end": , "params": {}, "reason": "", "speaker": ""}]} """ _OUTPUT_REMINDER = """Gere as decisões de edição conforme o brief. Responda SOMENTE o JSON: {"source": "", "actions": [{"kind": "cut|zoom|text|marker", "start": , "end": , "params": {}, "reason": "", "speaker": ""}]} Não inclua explicações nem blocos markdown.""" def ollama_chat( model: str = DEFAULT_MODEL, messages: Optional[Sequence[Dict[str, str]]] = None, base_url: str = DEFAULT_BASE_URL, temperature: float = 0.2, timeout: float = REQUEST_TIMEOUT, num_ctx: int = 32768, ) -> str: """One chat completion from a local Ollama server. Returns the assistant message content. Raises on transport/HTTP errors so the caller can decide whether to retry or report — a model call is the one I/O in this pipeline that can legitimately fail mid-run. """ payload = { "model": model, "messages": list(messages or []), "stream": False, "options": {"temperature": temperature, "num_ctx": num_ctx}, } try: response = httpx.post( f"{base_url.rstrip('/')}/api/chat", json=payload, timeout=timeout ) response.raise_for_status() data = response.json() except Exception as exc: # Covers transport errors AND a dropped connection that yields an empty # body (httpx/JSONDecodeError) — both must become a RuntimeError so the # caller reports the failure instead of crashing the whole pipeline. raise RuntimeError(f"Falha ao falar com o modelo local em {base_url}: {exc}") from exc return (data.get("message") or {}).get("content", "") or "" def list_ollama_models(base_url: str = DEFAULT_BASE_URL) -> list[str]: """Names of the models Ollama currently serves, for a model picker. Returns an empty list when Ollama is unreachable so the UI can fall back to a free-text field instead of erroring. """ try: resp = httpx.get(f"{base_url.rstrip('/')}/api/tags", timeout=10.0) resp.raise_for_status() models = resp.json().get("models", []) names = [m.get("name") for m in models if m.get("name")] return sorted(names) except Exception: return [] def _extract_json(text: str) -> Any: """Pull a JSON value out of a model response, tolerating fences/wrappers.""" if not text: return None candidate = text.strip() # Strip a ```json ... ``` (or bare ```) fence if the model added one. fence = re.search(r"```(?:json)?\s*(.*?)\s*```", candidate, re.DOTALL) if fence: candidate = fence.group(1).strip() # Otherwise take the outermost {...} / [...]. if not candidate.startswith(("{" if True else "", "[")): start = min( (i for i, c in enumerate(candidate) if c in "{["), default=None, ) end = max( (i for i, c in enumerate(candidate) if c in "}"), default=None, ) if start is not None and end is not None and end > start: candidate = candidate[start : end + 1] try: data = json.loads(candidate) except json.JSONDecodeError: return None # Models sometimes wrap the expected `{"source", "actions"}` object inside a # single-element list (`[{...}]`). Unwrap that so the actions aren't treated # as one malformed row. if ( isinstance(data, list) and len(data) == 1 and isinstance(data[0], dict) and "actions" in data[0] # the wrapper carries the actions key ): data = data[0] return data # Only these fields reach the model — the raw timeline also carries heavy # per-word audio features (energy, pitch, arousal...) and speaker `samples` # that blow past the model's context window on any real recording. Dropping # them is what keeps a 3-minute timeline inside `num_ctx`. _SEGMENT_KEEP = ( "start", "end", "speaker", "text", "gap_before", "take_boundary", "avg_energy", "peak_emphasis", "emotion", "emotion_confidence", "arousal", "valence", ) _WORD_KEEP = ("text", "start", "end", "speaker", "emphasis", "pause_before") _SPEAKER_KEEP = ("id", "name") _SKIP_ROOT = ("layers", "scales") def _project_timeline(timeline: dict) -> dict: """Strip the timeline down to what the edit decision actually needs.""" out = {k: v for k, v in timeline.items() if k not in _SKIP_ROOT} speakers = [ {k: sp[k] for k in _SPEAKER_KEEP if k in sp} for sp in timeline.get("speakers", []) ] if speakers: out["speakers"] = speakers segs = [] for seg in timeline.get("segments", []): s = {k: seg[k] for k in _SEGMENT_KEEP if k in seg} s["words"] = [ {k: w[k] for k in _WORD_KEEP if k in w} for w in seg.get("words", []) ] segs.append(s) out["segments"] = segs return out def _shrink_to_fit(compact: dict, max_chars: int) -> dict: """Drop word detail from the lowest-emphasis segments until it fits.""" segs = [dict(s) for s in compact.get("segments", [])] while True: payload = json.dumps( {**compact, "segments": segs}, ensure_ascii=False, indent=1 ) if len(payload) <= max_chars or not any(s.get("words") for s in segs): break idx = min( (i for i, s in enumerate(segs) if s.get("words")), key=lambda i: float(segs[i].get("peak_emphasis", 0.0)), ) segs[idx] = {**segs[idx], "words": []} compact = dict(compact) compact["segments"] = segs return compact def build_edit_messages( timeline: dict, max_words_per_segment: int = 200, max_chars: int = 110000 ) -> Tuple[str, str]: """The (system, user) pair that sends a timeline to the model. The user turn carries a *projected* timeline (see :func:`_project_timeline`) — text, timing, speaker and emphasis only — so a real recording fits in the model's context window. Very long segments still have their word detail capped to ``max_words_per_segment`` (most emphatic + boundaries), and if the whole payload would still exceed ``max_chars`` the lowest-emphasis segments lose their words until it fits, so we never blow ``num_ctx``. """ compact = _project_timeline(timeline) if max_words_per_segment: segs = [] for seg in compact["segments"]: words = seg.get("words", []) if len(words) > max_words_per_segment: ranked = sorted( enumerate(words), key=lambda kv: float(kv[1].get("emphasis", 0.0)), reverse=True, )[: max_words_per_segment - 2] keep = sorted({0, len(words) - 1} | {i for i, _ in ranked}) seg = {**seg, "words": [words[i] for i in keep]} segs.append(seg) compact["segments"] = segs payload = json.dumps(compact, ensure_ascii=False, indent=1) if len(payload) > max_chars: compact = _shrink_to_fit(compact, max_chars) payload = json.dumps(compact, ensure_ascii=False, indent=1) user = ( "Linha do tempo de voz (JSON):\n\n" + payload + "\n\n" + _OUTPUT_REMINDER ) return _SYSTEM_PROMPT, user def generate_voice_actions( timeline: dict, model: str = DEFAULT_MODEL, base_url: str = DEFAULT_BASE_URL, temperature: float = 0.2, timeout: float = REQUEST_TIMEOUT, num_ctx: int = 32768, max_words_per_segment: int = 200, ) -> Dict[str, Any]: """Ask the local model to direct the edit, returning validated actions. Returns ``{"actions": [VoiceAction], "raw": str, "errors": [str]}``. ``actions`` is empty when the model returned nothing usable; ``errors`` carries the per-row rejections from :func:`parse_actions` plus any extraction failure, so the caller can report what went wrong instead of only the wins. """ system, user = build_edit_messages(timeline, max_words_per_segment) try: raw = ollama_chat( model=model, messages=[ {"role": "system", "content": system}, {"role": "user", "content": user}, ], base_url=base_url, temperature=temperature, timeout=timeout, num_ctx=num_ctx, ) except RuntimeError as exc: return {"actions": [], "raw": "", "errors": [str(exc)]} data = _extract_json(raw) if data is None: return { "actions": [], "raw": raw, "errors": ["O modelo não devolveu um JSON de decisões legível."], } actions, errors = parse_actions(data) return {"actions": actions, "raw": raw, "errors": errors}