feat: etapa 5 do assistente — revisão de ênfases com timeline

Transforma a etapa "colar decisões" numa tela de lapidação: a sugestão da
IA chega carregada e o editor afina frase a frase o que é ênfase e o que
fica fora. Essa marcação é o norte da etapa 6 — só as frases com ênfase
recebem zoom e legenda dinâmica; as demais ficam com legenda comum.

O campo de colar o JSON sobe para a etapa 4, então a numeração das etapas
não muda e a etapa 6 segue intacta.

Backend (fcpxml/phrase_review.py):
- build_phrase_review funde o _voice_timeline.json com as actions da IA
- trim por frase que anda em fronteira de palavra; corte parcial da IA
  chega como trim em vez de ser arredondado fora
- phrase_review_to_actions volta a cuts/zooms + emphasis_spans
- merge_saved_decisions reaplica só as decisões salvas sobre uma revisão
  remontada da análise atual, para reprocessar a voz não ficar mascarado
- resolve_source acha a mídia: o voice timeline guarda só o nome do arquivo

App (SwiftUI):
- layout de sala de edição: preview em cima, inspector à direita, timeline
  atravessando embaixo com seis trilhas rotuladas
- preview enquadra no formato de entrega lido do .fcpxml (fonte horizontal,
  projeto vertical), com alternância para a mídia original
- reprodução pula os trechos removidos e para no fim do trecho
- zoom manual por trecho marcado, sem guardar escala: a forma vem das
  configurações de Análise de Voz no render
- emoção da fala exposta por frase

Correções encontradas no caminho:
- VideoPlayer (AVKit) aborta em runtime no app compilado por swiftc;
  trocado por AVPlayerLayer (ver Engine/docs/05_EXPERIENCIAS.md #22)
- teste que ainda afirmava o default zoom scale=1.3 removido do parser (#21)

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
João Henrique
2026-08-19 21:29:27 -04:00
co-authored by Claude Opus 5
parent e7748c2c58
commit 1bebee4359
31 changed files with 4622 additions and 83 deletions
+60 -7
View File
@@ -15,6 +15,7 @@ from typing import Any, Sequence
from mcp.types import TextContent
from fcpxml.media_intel import media_src_to_path
from fcpxml.model_manager import load_dynamic_subtitle_config, load_voice_analysis_config
from fcpxml.models import (
DuplicateGroup,
FlashFrame,
@@ -25,6 +26,7 @@ from fcpxml.models import (
)
from fcpxml.parser import FCPXMLParser
from fcpxml.rough_cut import RoughCutGenerator
from fcpxml.text_layout import TEXT_TEMPLATE_FONT_SCALE, measure_text
from fcpxml.transcribe import invert_ranges, merge_ranges, transcribe
from fcpxml.writer import FCPXMLModifier
@@ -639,28 +641,79 @@ def _apply_placed_action(modifier, clip_el, action, clip_start: float) -> str:
rel_end = action.end - clip_start
if action.kind == "zoom":
config = load_voice_analysis_config()
# Only forward an explicit ease — otherwise add_zoom's own default
# (a fast ramp in, instant snap back out) is what should apply.
zoom_args = {}
if action.params.get("ease") is not None:
zoom_args["ease"] = float(action.params["ease"])
if action.params.get("ease_out") is not None:
zoom_args["ease_out"] = float(action.params["ease_out"])
zoom_args = {
"ease": float(action.params.get("ease", config["zoom_ease_in"])),
"ease_out": float(action.params.get("ease_out", config["zoom_ease_out"])),
}
mode = str(action.params.get("mode", config["zoom_mode"]))
if mode == "in":
zoom_args["hold_at_end"] = True
zoom_args["start_at_peak"] = False
elif mode == "out":
zoom_args["hold_at_end"] = False
zoom_args["start_at_peak"] = True
elif mode == "in_out":
zoom_args["hold_at_end"] = False
zoom_args["start_at_peak"] = False
modifier.add_zoom(
clip_id=clip_el,
start=rel_start,
end=rel_end,
scale=float(action.params.get("scale", 1.3)),
scale=float(action.params.get("scale", config["zoom_scale"])),
**zoom_args,
)
return f"zoom {action.params.get('scale', 1.3):.2f}x"
return f"zoom {float(action.params.get('scale', config['zoom_scale'])):.2f}x"
if action.kind == "text":
# Default to the "Legendas Dinâmicas" emphasis style (the font used
# to highlight a word in the captions) rather than a hardcoded
# Helvetica Neue, so a callout like "MASTOPEXIA" matches the rest of
# the video's on-screen text instead of looking like a stray default
# title. Any of these the action itself specifies still wins.
subtitle_cfg = load_dynamic_subtitle_config()
font = action.params.get("font", subtitle_cfg["emphasis_font"])
face = action.params.get("face", subtitle_cfg["emphasis_face"])
font_scale = float(subtitle_cfg.get("text_scale", TEXT_TEMPLATE_FONT_SCALE) or 1.0)
requested_size = int(action.params.get("font_size", subtitle_cfg["emphasis_size"]))
requested_kerning = float(action.params.get("kerning", 0.0) or 0.0)
# Voice-action callouts are not part of the dynamic subtitle block.
# When omitted, put them above the subtitle band and shrink wide
# phrases to the title-safe width. The previous default (Position 0 0,
# full emphasis size) made long callouts like "PRÓTESES DE SILICONE"
# collide with captions and run off both sides of a vertical frame.
emitted_size = requested_size * font_scale
emitted_kerning = requested_kerning * font_scale
safe_width = modifier.frame_width() * 0.90
width = measure_text(
action.params["content"],
emitted_size,
bold=bool(action.params.get("bold", False)),
kerning=emitted_kerning,
font=font,
face=face,
)
font_size = requested_size
if width > safe_width and width > 0:
font_size = max(32, int(requested_size * safe_width / width))
position = action.params.get("position")
if not position:
position = f"0 {modifier.frame_height() * 0.23:g}"
modifier.add_text_title(
clip_el,
action.params["content"],
offset=modifier.snap_seconds_to_frame(rel_start).to_fcpxml(),
duration=modifier.snap_seconds_to_frame(action.duration).to_fcpxml(),
position=position,
font=font,
font_size=font_size,
font_color=action.params.get("font_color", subtitle_cfg["emphasis_color"]),
face=face,
bold=action.params.get("bold", False),
)
return f"text \"{action.params['content'][:24]}\""