Cortes escritos rente ao timestamp da palavra soavam secos (relatado no projeto Mastopexia) — o critério e o prompt do modelo local mandavam cobrir a frase inteira sem orientar a borda que toca fala mantida. Adiciona a regra de recuar ~0,15-0,25s nas duas pontas quando o corte encosta em conteúdo que fica, tanto no skill (06-texto-corte-marcador.md) quanto no prompt embutido do Ollama (llm_local.py) — pra não precisar ajustar na mão de novo. Também registra em 05_EXPERIENCIAS.md/09_MANUTENCAO.md a dívida de resolve_actions não tolerar margem quando zoom/marker encosta na borda de um corte (contornado manualmente, não corrigido em código ainda), e atualiza a lista de dívidas abertas (etapa 6/offset de whisper já resolvidos nesta sessão). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
313 lines
14 KiB
Python
313 lines
14 KiB
Python
"""Local LLM integration — the voice timeline meets a local model.
|
|
|
|
The voice timeline is *designed* to be handed to a language model: it is the
|
|
source of truth between speech analysis and editing, layered so a model can
|
|
reason about the narrative without parsing FCPXML. This module is the client
|
|
side of that contract. It formats the timeline into the editar-por-voz brief,
|
|
calls a local model server (Ollama, running Gemma 3 / Llama locally), and
|
|
parses the model's decisions back into a validated list of VoiceActions —
|
|
all inside the engine, so there is no wizard, no copy-paste, no manual step.
|
|
|
|
Transport: Ollama's HTTP chat API at ``http://localhost:11434/api/chat``.
|
|
Any model Ollama serves works; the default is Gemma 3 because that is what
|
|
runs locally here ("Lama com Gema 3"), but pass ``model=`` to switch.
|
|
|
|
The model is untrusted input: its JSON is validated row-by-row by
|
|
:func:`fcpxml.voice_actions.parse_actions`, so one malformed decision never
|
|
discards the edit. The brief is written so the model only ever emits the four
|
|
action kinds the applier understands.
|
|
"""
|
|
|
|
from __future__ import annotations
|
|
|
|
import json
|
|
import logging
|
|
import re
|
|
from typing import Any, Dict, Optional, Sequence, Tuple
|
|
|
|
import httpx
|
|
|
|
from .voice_actions import parse_actions
|
|
|
|
logger = logging.getLogger(__name__)
|
|
|
|
DEFAULT_BASE_URL = "http://localhost:11434"
|
|
# Gemma 3 12B reliably follows the editar-por-voz brief (keep the script, cut
|
|
# only backstage chatter; the 4B variant skips the "keep the main content"
|
|
# rule and deletes the script) but doesn't fit an 8GB machine. Qwen2.5 7B
|
|
# instruct (q4_K_M) is the fallback for constrained hardware — strong at
|
|
# strict JSON-schema following, the property this brief leans on hardest.
|
|
# Pass ``model=`` to switch to whatever Ollama serves.
|
|
DEFAULT_MODEL = "qwen2.5:7b-instruct-q4_K_M"
|
|
REQUEST_TIMEOUT = 600.0
|
|
|
|
# The brief. Ported from the editar-por-voz skill criteria (criterios/01..08),
|
|
# condensed into the instructions a model needs to emit valid actions. Kept in
|
|
# Portuguese because the decisions and their reasons are read by a human editor.
|
|
_SYSTEM_PROMPT = """Você é o editor de vídeo por voz deste sistema. Recebe um JSON de "linha do tempo de voz" — a medição de COMO foi falado (ênfase, energia, pausa, falante) de uma gravação — e devolve as DECISÕES de edição em JSON, nada mais. Você nunca escreve XML.
|
|
|
|
Regras (siga rigorosamente):
|
|
|
|
1. LEIA EM CAMADAS. "summary" dá o formato da peça; "segments" é onde você trabalha (cada fala com seu texto e agregados); "segments[].words" dá o instante exato de cada destaque. Não recalcule energia, tom ou ênfase — use os números do JSON.
|
|
|
|
2. SEPARAR ROTEIRO DE BASTIDOR.
|
|
- ROTEIRO = o conteúdo principal que a pessoa quer entregar: explicação, depoimento, roteiro decorado, a mensagem. É isso que VAI FICAR.
|
|
- BASTIDOR = papo casual de gravação, cumprimentos, conversa com a equipe ("cara, beleza?", "tá gravando?", "deixa eu ver o celular"), piadas fora do assunto, tomadas interrompidas ou repetidas. É isso que VIRA "cut".
|
|
Exemplo: num vídeo sobre mastopexia, a explicação da cirurgia É o roteiro (mantém); o "tá gravando? pois é" antes dela É bastidor (corta).
|
|
Use "gap_before" e "take_boundary" (silêncio > ~3s = a câmera parou/recomeçou) para agrupar tomadas — eles marcam ONDE a tomada recomeça, não o que cortar. Nunca corte o conteúdo principal só porque tem ênfase; corte o casual/off-topic.
|
|
|
|
REGRAS DE OURO:
|
|
- MANTENHA o conteúdo principal (explicação, depoimento, roteiro decorado). Ele É o vídeo.
|
|
- CORTE SÓ o casual/off-topic: cumprimentos, "tá gravando?", papo com a equipe, olhar o celular, repetições de tomada.
|
|
- Em dúvida, MANTENHA a fala. É melhor sobrar conteúdo do que cortar o que era pra ficar.
|
|
|
|
3. ESCOLHER A MELHOR TOMADA de cada frase quando há repetições: mantenha a mais limpa e corte as outras (cut cobrindo a frase inteira).
|
|
|
|
4. CORTE (kind "cut"): para REMOVER uma frase, cubra ela inteira (start..end = início..fim da frase). Para APARAR só uma hesitação no começo ou fim, corte só da borda até a palavra (corte de meia frase é ambíguo — passe de 60% e apaga a linha toda). Nunca corte o silêncio entre falas. Quando a borda do corte encosta em fala mantida (não em silêncio puro), recue ~0,15-0,25s para dentro do corte nos dois lados — start ~0,2s DEPOIS do fim real da última palavra mantida, end ~0,2s ANTES do início real da próxima palavra mantida — senão o corte soa seco, engolindo a palavra antes de terminar de soar. Isso vale também pro início/fim do vídeo (ar morto antes da primeira palavra e depois da última).
|
|
|
|
5. ZOOM (kind "zoom"): só em palavra de CONTEÚDO bem enfatizada (emphasis alto, não artigo). params.scale entre 1.0 e 3.0 (padrão 1.3 se omitido). Posicione em torno da palavra, segurando até o fim da frase.
|
|
|
|
6. TEXTO (kind "text"): params.content obrigatório (≤120 chars), fixa um termo central ou callout. MARKER (kind "marker"): opcional params.content vira o nome do marcador. Use para emendas/junções que o editor deve conferir.
|
|
|
|
7. TEMPOS em segundos da MÍDIA ORIGINAL (exatamente como no JSON). Nunca compense para "depois do corte" — o programa desloca sozinho. end sempre > start, ambos ≥ 0.
|
|
|
|
8. reason OBRIGATÓRIO em cada ação, em português, embasando a decisão (ex.: 'abertura: "Aquela mama" (ênfase 0.42)'). reason vazio é decisão sem critério.
|
|
|
|
Responda APENAS com um objeto JSON válido, sem markdown, sem comentário:
|
|
{"source": "<nome do arquivo>", "actions": [{"kind": "cut|zoom|text|marker", "start": <float>, "end": <float>, "params": {}, "reason": "<pt>", "speaker": "<id>"}]}
|
|
"""
|
|
|
|
_OUTPUT_REMINDER = """Gere as decisões de edição conforme o brief. Responda SOMENTE o JSON:
|
|
{"source": "<nome do arquivo>", "actions": [{"kind": "cut|zoom|text|marker", "start": <float>, "end": <float>, "params": {}, "reason": "<pt>", "speaker": "<id>"}]}
|
|
Não inclua explicações nem blocos markdown."""
|
|
|
|
|
|
def ollama_chat(
|
|
model: str = DEFAULT_MODEL,
|
|
messages: Optional[Sequence[Dict[str, str]]] = None,
|
|
base_url: str = DEFAULT_BASE_URL,
|
|
temperature: float = 0.2,
|
|
timeout: float = REQUEST_TIMEOUT,
|
|
num_ctx: int = 32768,
|
|
) -> str:
|
|
"""One chat completion from a local Ollama server.
|
|
|
|
Returns the assistant message content. Raises on transport/HTTP errors so
|
|
the caller can decide whether to retry or report — a model call is the
|
|
one I/O in this pipeline that can legitimately fail mid-run.
|
|
"""
|
|
payload = {
|
|
"model": model,
|
|
"messages": list(messages or []),
|
|
"stream": False,
|
|
"options": {"temperature": temperature, "num_ctx": num_ctx},
|
|
}
|
|
try:
|
|
response = httpx.post(
|
|
f"{base_url.rstrip('/')}/api/chat", json=payload, timeout=timeout
|
|
)
|
|
response.raise_for_status()
|
|
data = response.json()
|
|
except Exception as exc:
|
|
# Covers transport errors AND a dropped connection that yields an empty
|
|
# body (httpx/JSONDecodeError) — both must become a RuntimeError so the
|
|
# caller reports the failure instead of crashing the whole pipeline.
|
|
raise RuntimeError(f"Falha ao falar com o modelo local em {base_url}: {exc}") from exc
|
|
|
|
return (data.get("message") or {}).get("content", "") or ""
|
|
|
|
|
|
def list_ollama_models(base_url: str = DEFAULT_BASE_URL) -> list[str]:
|
|
"""Names of the models Ollama currently serves, for a model picker.
|
|
|
|
Returns an empty list when Ollama is unreachable so the UI can fall back to
|
|
a free-text field instead of erroring.
|
|
"""
|
|
try:
|
|
resp = httpx.get(f"{base_url.rstrip('/')}/api/tags", timeout=10.0)
|
|
resp.raise_for_status()
|
|
models = resp.json().get("models", [])
|
|
names = [m.get("name") for m in models if m.get("name")]
|
|
return sorted(names)
|
|
except Exception:
|
|
return []
|
|
|
|
|
|
def _extract_json(text: str) -> Any:
|
|
"""Pull a JSON value out of a model response, tolerating fences/wrappers."""
|
|
if not text:
|
|
return None
|
|
candidate = text.strip()
|
|
# Strip a ```json ... ``` (or bare ```) fence if the model added one.
|
|
fence = re.search(r"```(?:json)?\s*(.*?)\s*```", candidate, re.DOTALL)
|
|
if fence:
|
|
candidate = fence.group(1).strip()
|
|
# Otherwise take the outermost {...} / [...].
|
|
if not candidate.startswith(("{" if True else "", "[")):
|
|
start = min(
|
|
(i for i, c in enumerate(candidate) if c in "{["),
|
|
default=None,
|
|
)
|
|
end = max(
|
|
(i for i, c in enumerate(candidate) if c in "}"),
|
|
default=None,
|
|
)
|
|
if start is not None and end is not None and end > start:
|
|
candidate = candidate[start : end + 1]
|
|
try:
|
|
data = json.loads(candidate)
|
|
except json.JSONDecodeError:
|
|
return None
|
|
|
|
# Models sometimes wrap the expected `{"source", "actions"}` object inside a
|
|
# single-element list (`[{...}]`). Unwrap that so the actions aren't treated
|
|
# as one malformed row.
|
|
if (
|
|
isinstance(data, list)
|
|
and len(data) == 1
|
|
and isinstance(data[0], dict)
|
|
and "actions" in data[0] # the wrapper carries the actions key
|
|
):
|
|
data = data[0]
|
|
return data
|
|
|
|
|
|
# Only these fields reach the model — the raw timeline also carries heavy
|
|
# per-word audio features (energy, pitch, arousal...) and speaker `samples`
|
|
# that blow past the model's context window on any real recording. Dropping
|
|
# them is what keeps a 3-minute timeline inside `num_ctx`.
|
|
_SEGMENT_KEEP = (
|
|
"start", "end", "speaker", "text", "gap_before", "take_boundary",
|
|
"avg_energy", "peak_emphasis", "emotion", "emotion_confidence",
|
|
"arousal", "valence",
|
|
)
|
|
_WORD_KEEP = ("text", "start", "end", "speaker", "emphasis", "pause_before")
|
|
_SPEAKER_KEEP = ("id", "name")
|
|
_SKIP_ROOT = ("layers", "scales")
|
|
|
|
|
|
def _project_timeline(timeline: dict) -> dict:
|
|
"""Strip the timeline down to what the edit decision actually needs."""
|
|
out = {k: v for k, v in timeline.items() if k not in _SKIP_ROOT}
|
|
speakers = [
|
|
{k: sp[k] for k in _SPEAKER_KEEP if k in sp}
|
|
for sp in timeline.get("speakers", [])
|
|
]
|
|
if speakers:
|
|
out["speakers"] = speakers
|
|
segs = []
|
|
for seg in timeline.get("segments", []):
|
|
s = {k: seg[k] for k in _SEGMENT_KEEP if k in seg}
|
|
s["words"] = [
|
|
{k: w[k] for k in _WORD_KEEP if k in w}
|
|
for w in seg.get("words", [])
|
|
]
|
|
segs.append(s)
|
|
out["segments"] = segs
|
|
return out
|
|
|
|
|
|
def _shrink_to_fit(compact: dict, max_chars: int) -> dict:
|
|
"""Drop word detail from the lowest-emphasis segments until it fits."""
|
|
segs = [dict(s) for s in compact.get("segments", [])]
|
|
while True:
|
|
payload = json.dumps(
|
|
{**compact, "segments": segs}, ensure_ascii=False, indent=1
|
|
)
|
|
if len(payload) <= max_chars or not any(s.get("words") for s in segs):
|
|
break
|
|
idx = min(
|
|
(i for i, s in enumerate(segs) if s.get("words")),
|
|
key=lambda i: float(segs[i].get("peak_emphasis", 0.0)),
|
|
)
|
|
segs[idx] = {**segs[idx], "words": []}
|
|
compact = dict(compact)
|
|
compact["segments"] = segs
|
|
return compact
|
|
|
|
|
|
def build_edit_messages(
|
|
timeline: dict, max_words_per_segment: int = 200, max_chars: int = 110000
|
|
) -> Tuple[str, str]:
|
|
"""The (system, user) pair that sends a timeline to the model.
|
|
|
|
The user turn carries a *projected* timeline (see :func:`_project_timeline`)
|
|
— text, timing, speaker and emphasis only — so a real recording fits in the
|
|
model's context window. Very long segments still have their word detail
|
|
capped to ``max_words_per_segment`` (most emphatic + boundaries), and if the
|
|
whole payload would still exceed ``max_chars`` the lowest-emphasis segments
|
|
lose their words until it fits, so we never blow ``num_ctx``.
|
|
"""
|
|
compact = _project_timeline(timeline)
|
|
if max_words_per_segment:
|
|
segs = []
|
|
for seg in compact["segments"]:
|
|
words = seg.get("words", [])
|
|
if len(words) > max_words_per_segment:
|
|
ranked = sorted(
|
|
enumerate(words),
|
|
key=lambda kv: float(kv[1].get("emphasis", 0.0)),
|
|
reverse=True,
|
|
)[: max_words_per_segment - 2]
|
|
keep = sorted({0, len(words) - 1} | {i for i, _ in ranked})
|
|
seg = {**seg, "words": [words[i] for i in keep]}
|
|
segs.append(seg)
|
|
compact["segments"] = segs
|
|
|
|
payload = json.dumps(compact, ensure_ascii=False, indent=1)
|
|
if len(payload) > max_chars:
|
|
compact = _shrink_to_fit(compact, max_chars)
|
|
payload = json.dumps(compact, ensure_ascii=False, indent=1)
|
|
|
|
user = (
|
|
"Linha do tempo de voz (JSON):\n\n"
|
|
+ payload
|
|
+ "\n\n"
|
|
+ _OUTPUT_REMINDER
|
|
)
|
|
return _SYSTEM_PROMPT, user
|
|
|
|
|
|
def generate_voice_actions(
|
|
timeline: dict,
|
|
model: str = DEFAULT_MODEL,
|
|
base_url: str = DEFAULT_BASE_URL,
|
|
temperature: float = 0.2,
|
|
timeout: float = REQUEST_TIMEOUT,
|
|
num_ctx: int = 32768,
|
|
max_words_per_segment: int = 200,
|
|
) -> Dict[str, Any]:
|
|
"""Ask the local model to direct the edit, returning validated actions.
|
|
|
|
Returns ``{"actions": [VoiceAction], "raw": str, "errors": [str]}``.
|
|
``actions`` is empty when the model returned nothing usable; ``errors``
|
|
carries the per-row rejections from :func:`parse_actions` plus any
|
|
extraction failure, so the caller can report what went wrong instead of
|
|
only the wins.
|
|
"""
|
|
system, user = build_edit_messages(timeline, max_words_per_segment)
|
|
try:
|
|
raw = ollama_chat(
|
|
model=model,
|
|
messages=[
|
|
{"role": "system", "content": system},
|
|
{"role": "user", "content": user},
|
|
],
|
|
base_url=base_url,
|
|
temperature=temperature,
|
|
timeout=timeout,
|
|
num_ctx=num_ctx,
|
|
)
|
|
except RuntimeError as exc:
|
|
return {"actions": [], "raw": "", "errors": [str(exc)]}
|
|
|
|
data = _extract_json(raw)
|
|
if data is None:
|
|
return {
|
|
"actions": [],
|
|
"raw": raw,
|
|
"errors": ["O modelo não devolveu um JSON de decisões legível."],
|
|
}
|
|
actions, errors = parse_actions(data)
|
|
return {"actions": actions, "raw": raw, "errors": errors}
|