Files
gart/code/fcpxml/llm_local.py
T
João HenriqueandClaude Sonnet 5 fd791e116a docs(skill): corte deve deixar folga na borda que encosta em fala mantida
Cortes escritos rente ao timestamp da palavra soavam secos (relatado no
projeto Mastopexia) — o critério e o prompt do modelo local mandavam cobrir
a frase inteira sem orientar a borda que toca fala mantida. Adiciona a regra
de recuar ~0,15-0,25s nas duas pontas quando o corte encosta em conteúdo
que fica, tanto no skill (06-texto-corte-marcador.md) quanto no prompt
embutido do Ollama (llm_local.py) — pra não precisar ajustar na mão de novo.

Também registra em 05_EXPERIENCIAS.md/09_MANUTENCAO.md a dívida de
resolve_actions não tolerar margem quando zoom/marker encosta na borda
de um corte (contornado manualmente, não corrigido em código ainda), e
atualiza a lista de dívidas abertas (etapa 6/offset de whisper já resolvidos
nesta sessão).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-21 18:31:27 -04:00

313 lines
14 KiB
Python

"""Local LLM integration — the voice timeline meets a local model.
The voice timeline is *designed* to be handed to a language model: it is the
source of truth between speech analysis and editing, layered so a model can
reason about the narrative without parsing FCPXML. This module is the client
side of that contract. It formats the timeline into the editar-por-voz brief,
calls a local model server (Ollama, running Gemma 3 / Llama locally), and
parses the model's decisions back into a validated list of VoiceActions —
all inside the engine, so there is no wizard, no copy-paste, no manual step.
Transport: Ollama's HTTP chat API at ``http://localhost:11434/api/chat``.
Any model Ollama serves works; the default is Gemma 3 because that is what
runs locally here ("Lama com Gema 3"), but pass ``model=`` to switch.
The model is untrusted input: its JSON is validated row-by-row by
:func:`fcpxml.voice_actions.parse_actions`, so one malformed decision never
discards the edit. The brief is written so the model only ever emits the four
action kinds the applier understands.
"""
from __future__ import annotations
import json
import logging
import re
from typing import Any, Dict, Optional, Sequence, Tuple
import httpx
from .voice_actions import parse_actions
logger = logging.getLogger(__name__)
DEFAULT_BASE_URL = "http://localhost:11434"
# Gemma 3 12B reliably follows the editar-por-voz brief (keep the script, cut
# only backstage chatter; the 4B variant skips the "keep the main content"
# rule and deletes the script) but doesn't fit an 8GB machine. Qwen2.5 7B
# instruct (q4_K_M) is the fallback for constrained hardware — strong at
# strict JSON-schema following, the property this brief leans on hardest.
# Pass ``model=`` to switch to whatever Ollama serves.
DEFAULT_MODEL = "qwen2.5:7b-instruct-q4_K_M"
REQUEST_TIMEOUT = 600.0
# The brief. Ported from the editar-por-voz skill criteria (criterios/01..08),
# condensed into the instructions a model needs to emit valid actions. Kept in
# Portuguese because the decisions and their reasons are read by a human editor.
_SYSTEM_PROMPT = """Você é o editor de vídeo por voz deste sistema. Recebe um JSON de "linha do tempo de voz" — a medição de COMO foi falado (ênfase, energia, pausa, falante) de uma gravação — e devolve as DECISÕES de edição em JSON, nada mais. Você nunca escreve XML.
Regras (siga rigorosamente):
1. LEIA EM CAMADAS. "summary" dá o formato da peça; "segments" é onde você trabalha (cada fala com seu texto e agregados); "segments[].words" dá o instante exato de cada destaque. Não recalcule energia, tom ou ênfase — use os números do JSON.
2. SEPARAR ROTEIRO DE BASTIDOR.
- ROTEIRO = o conteúdo principal que a pessoa quer entregar: explicação, depoimento, roteiro decorado, a mensagem. É isso que VAI FICAR.
- BASTIDOR = papo casual de gravação, cumprimentos, conversa com a equipe ("cara, beleza?", "tá gravando?", "deixa eu ver o celular"), piadas fora do assunto, tomadas interrompidas ou repetidas. É isso que VIRA "cut".
Exemplo: num vídeo sobre mastopexia, a explicação da cirurgia É o roteiro (mantém); o "tá gravando? pois é" antes dela É bastidor (corta).
Use "gap_before" e "take_boundary" (silêncio > ~3s = a câmera parou/recomeçou) para agrupar tomadas — eles marcam ONDE a tomada recomeça, não o que cortar. Nunca corte o conteúdo principal só porque tem ênfase; corte o casual/off-topic.
REGRAS DE OURO:
- MANTENHA o conteúdo principal (explicação, depoimento, roteiro decorado). Ele É o vídeo.
- CORTE SÓ o casual/off-topic: cumprimentos, "tá gravando?", papo com a equipe, olhar o celular, repetições de tomada.
- Em dúvida, MANTENHA a fala. É melhor sobrar conteúdo do que cortar o que era pra ficar.
3. ESCOLHER A MELHOR TOMADA de cada frase quando há repetições: mantenha a mais limpa e corte as outras (cut cobrindo a frase inteira).
4. CORTE (kind "cut"): para REMOVER uma frase, cubra ela inteira (start..end = início..fim da frase). Para APARAR só uma hesitação no começo ou fim, corte só da borda até a palavra (corte de meia frase é ambíguo — passe de 60% e apaga a linha toda). Nunca corte o silêncio entre falas. Quando a borda do corte encosta em fala mantida (não em silêncio puro), recue ~0,15-0,25s para dentro do corte nos dois lados — start ~0,2s DEPOIS do fim real da última palavra mantida, end ~0,2s ANTES do início real da próxima palavra mantida — senão o corte soa seco, engolindo a palavra antes de terminar de soar. Isso vale também pro início/fim do vídeo (ar morto antes da primeira palavra e depois da última).
5. ZOOM (kind "zoom"): só em palavra de CONTEÚDO bem enfatizada (emphasis alto, não artigo). params.scale entre 1.0 e 3.0 (padrão 1.3 se omitido). Posicione em torno da palavra, segurando até o fim da frase.
6. TEXTO (kind "text"): params.content obrigatório (≤120 chars), fixa um termo central ou callout. MARKER (kind "marker"): opcional params.content vira o nome do marcador. Use para emendas/junções que o editor deve conferir.
7. TEMPOS em segundos da MÍDIA ORIGINAL (exatamente como no JSON). Nunca compense para "depois do corte" — o programa desloca sozinho. end sempre > start, ambos ≥ 0.
8. reason OBRIGATÓRIO em cada ação, em português, embasando a decisão (ex.: 'abertura: "Aquela mama" (ênfase 0.42)'). reason vazio é decisão sem critério.
Responda APENAS com um objeto JSON válido, sem markdown, sem comentário:
{"source": "<nome do arquivo>", "actions": [{"kind": "cut|zoom|text|marker", "start": <float>, "end": <float>, "params": {}, "reason": "<pt>", "speaker": "<id>"}]}
"""
_OUTPUT_REMINDER = """Gere as decisões de edição conforme o brief. Responda SOMENTE o JSON:
{"source": "<nome do arquivo>", "actions": [{"kind": "cut|zoom|text|marker", "start": <float>, "end": <float>, "params": {}, "reason": "<pt>", "speaker": "<id>"}]}
Não inclua explicações nem blocos markdown."""
def ollama_chat(
model: str = DEFAULT_MODEL,
messages: Optional[Sequence[Dict[str, str]]] = None,
base_url: str = DEFAULT_BASE_URL,
temperature: float = 0.2,
timeout: float = REQUEST_TIMEOUT,
num_ctx: int = 32768,
) -> str:
"""One chat completion from a local Ollama server.
Returns the assistant message content. Raises on transport/HTTP errors so
the caller can decide whether to retry or report — a model call is the
one I/O in this pipeline that can legitimately fail mid-run.
"""
payload = {
"model": model,
"messages": list(messages or []),
"stream": False,
"options": {"temperature": temperature, "num_ctx": num_ctx},
}
try:
response = httpx.post(
f"{base_url.rstrip('/')}/api/chat", json=payload, timeout=timeout
)
response.raise_for_status()
data = response.json()
except Exception as exc:
# Covers transport errors AND a dropped connection that yields an empty
# body (httpx/JSONDecodeError) — both must become a RuntimeError so the
# caller reports the failure instead of crashing the whole pipeline.
raise RuntimeError(f"Falha ao falar com o modelo local em {base_url}: {exc}") from exc
return (data.get("message") or {}).get("content", "") or ""
def list_ollama_models(base_url: str = DEFAULT_BASE_URL) -> list[str]:
"""Names of the models Ollama currently serves, for a model picker.
Returns an empty list when Ollama is unreachable so the UI can fall back to
a free-text field instead of erroring.
"""
try:
resp = httpx.get(f"{base_url.rstrip('/')}/api/tags", timeout=10.0)
resp.raise_for_status()
models = resp.json().get("models", [])
names = [m.get("name") for m in models if m.get("name")]
return sorted(names)
except Exception:
return []
def _extract_json(text: str) -> Any:
"""Pull a JSON value out of a model response, tolerating fences/wrappers."""
if not text:
return None
candidate = text.strip()
# Strip a ```json ... ``` (or bare ```) fence if the model added one.
fence = re.search(r"```(?:json)?\s*(.*?)\s*```", candidate, re.DOTALL)
if fence:
candidate = fence.group(1).strip()
# Otherwise take the outermost {...} / [...].
if not candidate.startswith(("{" if True else "", "[")):
start = min(
(i for i, c in enumerate(candidate) if c in "{["),
default=None,
)
end = max(
(i for i, c in enumerate(candidate) if c in "}"),
default=None,
)
if start is not None and end is not None and end > start:
candidate = candidate[start : end + 1]
try:
data = json.loads(candidate)
except json.JSONDecodeError:
return None
# Models sometimes wrap the expected `{"source", "actions"}` object inside a
# single-element list (`[{...}]`). Unwrap that so the actions aren't treated
# as one malformed row.
if (
isinstance(data, list)
and len(data) == 1
and isinstance(data[0], dict)
and "actions" in data[0] # the wrapper carries the actions key
):
data = data[0]
return data
# Only these fields reach the model — the raw timeline also carries heavy
# per-word audio features (energy, pitch, arousal...) and speaker `samples`
# that blow past the model's context window on any real recording. Dropping
# them is what keeps a 3-minute timeline inside `num_ctx`.
_SEGMENT_KEEP = (
"start", "end", "speaker", "text", "gap_before", "take_boundary",
"avg_energy", "peak_emphasis", "emotion", "emotion_confidence",
"arousal", "valence",
)
_WORD_KEEP = ("text", "start", "end", "speaker", "emphasis", "pause_before")
_SPEAKER_KEEP = ("id", "name")
_SKIP_ROOT = ("layers", "scales")
def _project_timeline(timeline: dict) -> dict:
"""Strip the timeline down to what the edit decision actually needs."""
out = {k: v for k, v in timeline.items() if k not in _SKIP_ROOT}
speakers = [
{k: sp[k] for k in _SPEAKER_KEEP if k in sp}
for sp in timeline.get("speakers", [])
]
if speakers:
out["speakers"] = speakers
segs = []
for seg in timeline.get("segments", []):
s = {k: seg[k] for k in _SEGMENT_KEEP if k in seg}
s["words"] = [
{k: w[k] for k in _WORD_KEEP if k in w}
for w in seg.get("words", [])
]
segs.append(s)
out["segments"] = segs
return out
def _shrink_to_fit(compact: dict, max_chars: int) -> dict:
"""Drop word detail from the lowest-emphasis segments until it fits."""
segs = [dict(s) for s in compact.get("segments", [])]
while True:
payload = json.dumps(
{**compact, "segments": segs}, ensure_ascii=False, indent=1
)
if len(payload) <= max_chars or not any(s.get("words") for s in segs):
break
idx = min(
(i for i, s in enumerate(segs) if s.get("words")),
key=lambda i: float(segs[i].get("peak_emphasis", 0.0)),
)
segs[idx] = {**segs[idx], "words": []}
compact = dict(compact)
compact["segments"] = segs
return compact
def build_edit_messages(
timeline: dict, max_words_per_segment: int = 200, max_chars: int = 110000
) -> Tuple[str, str]:
"""The (system, user) pair that sends a timeline to the model.
The user turn carries a *projected* timeline (see :func:`_project_timeline`)
— text, timing, speaker and emphasis only — so a real recording fits in the
model's context window. Very long segments still have their word detail
capped to ``max_words_per_segment`` (most emphatic + boundaries), and if the
whole payload would still exceed ``max_chars`` the lowest-emphasis segments
lose their words until it fits, so we never blow ``num_ctx``.
"""
compact = _project_timeline(timeline)
if max_words_per_segment:
segs = []
for seg in compact["segments"]:
words = seg.get("words", [])
if len(words) > max_words_per_segment:
ranked = sorted(
enumerate(words),
key=lambda kv: float(kv[1].get("emphasis", 0.0)),
reverse=True,
)[: max_words_per_segment - 2]
keep = sorted({0, len(words) - 1} | {i for i, _ in ranked})
seg = {**seg, "words": [words[i] for i in keep]}
segs.append(seg)
compact["segments"] = segs
payload = json.dumps(compact, ensure_ascii=False, indent=1)
if len(payload) > max_chars:
compact = _shrink_to_fit(compact, max_chars)
payload = json.dumps(compact, ensure_ascii=False, indent=1)
user = (
"Linha do tempo de voz (JSON):\n\n"
+ payload
+ "\n\n"
+ _OUTPUT_REMINDER
)
return _SYSTEM_PROMPT, user
def generate_voice_actions(
timeline: dict,
model: str = DEFAULT_MODEL,
base_url: str = DEFAULT_BASE_URL,
temperature: float = 0.2,
timeout: float = REQUEST_TIMEOUT,
num_ctx: int = 32768,
max_words_per_segment: int = 200,
) -> Dict[str, Any]:
"""Ask the local model to direct the edit, returning validated actions.
Returns ``{"actions": [VoiceAction], "raw": str, "errors": [str]}``.
``actions`` is empty when the model returned nothing usable; ``errors``
carries the per-row rejections from :func:`parse_actions` plus any
extraction failure, so the caller can report what went wrong instead of
only the wins.
"""
system, user = build_edit_messages(timeline, max_words_per_segment)
try:
raw = ollama_chat(
model=model,
messages=[
{"role": "system", "content": system},
{"role": "user", "content": user},
],
base_url=base_url,
temperature=temperature,
timeout=timeout,
num_ctx=num_ctx,
)
except RuntimeError as exc:
return {"actions": [], "raw": "", "errors": [str(exc)]}
data = _extract_json(raw)
if data is None:
return {
"actions": [],
"raw": raw,
"errors": ["O modelo não devolveu um JSON de decisões legível."],
}
actions, errors = parse_actions(data)
return {"actions": actions, "raw": raw, "errors": errors}