feat(voz): legenda por ênfase, forced align, IA local e correções de zoom/revisão
Trabalho da branch feat/revisao-enfases: pipeline de edição por voz ganha alinhamento forçado (whisperx), roteirização por LLM local (Ollama), e a etapa 5 (revisão de frases) passa a refletir de verdade o que é aplicado. - generate_subtitles_by_emphasis: legenda comum cobre o clipe inteiro, legenda dinâmica só nas frases de ênfase, e a comum é desativada (enabled="0") onde a dinâmica cobre, em vez de nunca ser gerada ali. - validate_subtitle_layout ignora títulos com enabled="0" — corrige falso positivo de colisão contra o que está desativado no lugar dele. - Corrige zoom/marcador sendo descartado quando a borda encosta exatamente no início de um corte. - Etapa 5 do Assistente: recarrega quando as decisões da IA mudam (com fresh=true, ignorando a revisão salva antiga) — resolve a dessincronia entre "ativa" na tela e o que já foi cortado no FCPXML. - Etapa "Processar" reaplica as decisões da revisão (_phrase_actions.json) antes da cadeia de remoção de silêncio/legendas — antes, desativar uma frase na etapa 5 não tinha efeito nenhum no vídeo final. - Etapa "Concluído" fundida em "Processar" — abrir no Final Cut/Finder aparece assim que termina, sem slide extra. - Palavra clicável na etapa 5 agora funciona como toggle (clique de novo desfaz) e mostra a própria ênfase (sublinhado colorido + peso da fonte). - fcpxml/forced_align.py, fcpxml/llm_local.py, ai_edit.py: alinhamento fonético via whisperx e roteirização local via Ollama/Gemma. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
This commit is contained in:
co-authored by
Claude Sonnet 5
parent
711c397dfe
commit
7b5aed79ee
@@ -93,6 +93,25 @@ TOOLS = [
|
||||
"required": ["filepath"]
|
||||
}
|
||||
),
|
||||
Tool(
|
||||
name="generate_subtitles_by_emphasis",
|
||||
description="Generate BOTH subtitle styles over the FULL clip and let them coexist by visibility, not by splitting words: plain static titles (see generate_plain_subtitles) cover every word from start to end; dynamic progressive-composition titles (see generate_dynamic_subtitles) are additionally generated for whichever whole phrases were marked as emphasis in the phrase-review step (etapa 5, zoom applied, level >= 1). Wherever a dynamic phrase is on screen, the plain titles underneath it are set enabled=\"0\" (still present in the FCPXML, editable/re-enable-able in Final Cut, just not rendered) instead of never being generated there — so disabling emphasis later never leaves a silent gap in the plain track. Reads emphasis spans from the media's cached '<media>_phrase_actions.json' (written by save_phrase_review after the app's etapa 5 review) — run the voice-editing wizard through that step first, or nothing is treated as emphasis and every title stays plain and enabled. Style knobs are the saved 'Legendas Dinâmicas'/plain-subtitle configs (~/.fcp-mcp-server/config.json); this tool does not expose per-call style overrides, only the split logic — use generate_dynamic_subtitles/generate_plain_subtitles directly if you need one-off styling.",
|
||||
inputSchema={
|
||||
"type": "object",
|
||||
"properties": {
|
||||
"filepath": {"type": "string", "description": "Path to FCPXML file"},
|
||||
"clip_name": {"type": "string", "description": "Only caption the clip with this name (default: all spine clips with matched source media)"},
|
||||
"model": {"type": "string", "default": "base", "description": "Whisper model size if transcription is needed"},
|
||||
"language": {"type": "string", "description": "ISO language code hint (e.g. 'pt'); auto-detected if omitted"},
|
||||
"granularity": {"type": "string", "enum": ["phrase", "word"], "default": "phrase", "description": "Passed through to the dynamic half, same meaning as in generate_dynamic_subtitles"},
|
||||
"max_words": {"type": "integer", "description": "Max words per block for the plain half. Falls back to saved plain-subtitle config."},
|
||||
"uppercase": {"type": "boolean", "description": "Uppercase the plain half. Falls back to saved plain-subtitle config."},
|
||||
"keep_punctuation": {"type": "boolean", "description": "Keep punctuation in the plain half. Falls back to saved plain-subtitle config."},
|
||||
"output_path": {"type": "string", "description": "Output path (default: adds _emphasis_subtitles suffix)"},
|
||||
},
|
||||
"required": ["filepath"]
|
||||
}
|
||||
),
|
||||
]
|
||||
|
||||
|
||||
@@ -140,6 +159,85 @@ def _plain_subtitle_blocks(words: Sequence[dict], max_words: int) -> list[list[d
|
||||
return blocks
|
||||
|
||||
|
||||
def _phrase_actions_path(media_path: str) -> Path:
|
||||
"""Where `save_phrase_review` writes emphasis decisions for this media.
|
||||
|
||||
Mirrors `phrase_review.review_paths()`'s naming (stem + "_phrase_actions.json"),
|
||||
without importing that module just for a path — the voice_timeline this would
|
||||
normally derive from is itself named `<media stem>_voice_timeline.json`, so
|
||||
stripping straight from the media stem lands on the same file.
|
||||
"""
|
||||
stem = Path(media_path).stem
|
||||
return Path(media_path).with_name(f"{stem}_phrase_actions.json")
|
||||
|
||||
|
||||
def _load_emphasis_spans(media_path: str) -> list[dict]:
|
||||
"""Load emphasis spans (source-media time) saved by the etapa-5 phrase review.
|
||||
|
||||
Returns [] if the review was never run for this media — callers should treat
|
||||
that as "nothing is emphasis yet", not as an error, since the wizard's later
|
||||
steps are optional.
|
||||
"""
|
||||
path = _phrase_actions_path(media_path)
|
||||
if not path.is_file():
|
||||
return []
|
||||
try:
|
||||
data = json.loads(path.read_text(encoding="utf-8"))
|
||||
except (OSError, json.JSONDecodeError):
|
||||
return []
|
||||
spans = data.get("emphasis_spans", [])
|
||||
return [s for s in spans if isinstance(s, dict) and "start" in s and "end" in s]
|
||||
|
||||
|
||||
def _word_in_spans(word_start: float, word_end: float, spans: Sequence[dict]) -> bool:
|
||||
"""A word belongs to an emphasis span if its midpoint falls inside it.
|
||||
|
||||
Midpoint, not start, so a word straddling a span boundary (which can happen
|
||||
since spans come from phrase trims, not word timestamps) lands on whichever
|
||||
side it mostly belongs to instead of always defaulting to one edge.
|
||||
"""
|
||||
mid = (word_start + word_end) / 2.0
|
||||
return any(float(s["start"]) <= mid < float(s["end"]) for s in spans)
|
||||
|
||||
|
||||
def _words_in_spans(words: Sequence[dict], spans: Sequence[dict]) -> list[dict]:
|
||||
"""The subset of source-time transcript words that fall inside a span.
|
||||
|
||||
Feeds only the DYNAMIC half — the plain half always gets every word, full
|
||||
clip, unfiltered; this is not a partition of the word list into two
|
||||
disjoint sets, it is "which words also get the dynamic treatment on top".
|
||||
"""
|
||||
if not spans:
|
||||
return []
|
||||
return [
|
||||
w for w in words
|
||||
if _word_in_spans(float(w.get("start", 0.0)), float(w.get("end", w.get("start", 0.0))), spans)
|
||||
]
|
||||
|
||||
|
||||
def _segments_in_spans(segments: Sequence[dict], spans: Sequence[dict]) -> list[dict]:
|
||||
"""Keep only the sentences that fall inside an emphasis span (by midpoint).
|
||||
|
||||
Feeds the dynamic half's sentence-block builder; segments outside every span
|
||||
would only produce blocks with no words left in them after the word filter.
|
||||
"""
|
||||
if not spans:
|
||||
return []
|
||||
kept = []
|
||||
for seg in segments:
|
||||
start = float(seg.get("start", 0.0))
|
||||
end = float(seg.get("end", start))
|
||||
mid = (start + end) / 2.0
|
||||
if any(float(s["start"]) <= mid < float(s["end"]) for s in spans):
|
||||
kept.append(seg)
|
||||
return kept
|
||||
|
||||
|
||||
def _overlaps_any_span(start: float, end: float, spans: Sequence[tuple[float, float]]) -> bool:
|
||||
"""Half-open interval overlap: a plain title under this window must hide."""
|
||||
return any(start < span_end and end > span_start for span_start, span_end in spans)
|
||||
|
||||
|
||||
async def handle_validate_subtitle_layout(arguments: dict) -> Sequence[TextContent]:
|
||||
"""Validate title/subtitle layout for spatial collisions and safe-area
|
||||
containment (collision.validate_titles over every <title> in the file)."""
|
||||
@@ -442,8 +540,214 @@ async def handle_generate_plain_subtitles(arguments: dict) -> Sequence[TextConte
|
||||
return _text_result(result)
|
||||
|
||||
|
||||
async def handle_generate_subtitles_by_emphasis(arguments: dict) -> Sequence[TextContent]:
|
||||
"""Generate plain titles for the whole clip and dynamic titles for the
|
||||
emphasis phrases on top, then hide (enabled="0") the plain titles that
|
||||
fall under a dynamic phrase — never split the word list between the two.
|
||||
|
||||
Plain always covers every word, so turning emphasis off later (editing
|
||||
the phrase review and re-running) never leaves a silent gap: the plain
|
||||
title was there all along, just disabled.
|
||||
"""
|
||||
model = arguments.get("model", "base")
|
||||
language = arguments.get("language")
|
||||
output_dir = arguments.get("output_dir")
|
||||
clip_filter = arguments.get("clip_name")
|
||||
granularity = arguments.get("granularity", "phrase")
|
||||
|
||||
saved_dynamic = load_dynamic_subtitle_config()
|
||||
body_color = saved_dynamic["active_color"]
|
||||
dynamic_config = DynamicSubtitleConfig(
|
||||
style=WordStyle(
|
||||
font=saved_dynamic["font"],
|
||||
font_size=int(saved_dynamic["font_size"]),
|
||||
active_color=body_color,
|
||||
inactive_color="0.7 0.7 0.7 1",
|
||||
emphasis_look=WordLook(
|
||||
int(saved_dynamic["emphasis_size"]),
|
||||
saved_dynamic["emphasis_color"] or body_color,
|
||||
font=saved_dynamic["emphasis_font"],
|
||||
face=saved_dynamic["emphasis_face"],
|
||||
kerning=0.0,
|
||||
),
|
||||
body_look=WordLook(
|
||||
int(saved_dynamic["font_size"]),
|
||||
body_color,
|
||||
font=saved_dynamic["font"],
|
||||
face="Bold",
|
||||
kerning=1.2,
|
||||
),
|
||||
),
|
||||
band_height=float(saved_dynamic["band_height"]),
|
||||
block_center_y=float(saved_dynamic["block_center_y"]),
|
||||
granularity=granularity,
|
||||
text_scale=float(saved_dynamic["text_scale"]),
|
||||
line_gap=float(saved_dynamic["line_gap"]),
|
||||
)
|
||||
|
||||
saved_plain = load_plain_subtitle_config()
|
||||
plain_font = saved_plain["font"]
|
||||
plain_font_size = int(saved_plain["font_size"])
|
||||
plain_font_color = saved_plain["font_color"]
|
||||
max_words = max(1, int(arguments.get("max_words", saved_plain["max_words"])))
|
||||
position_y = float(saved_plain["position_y"])
|
||||
uppercase = bool(arguments.get("uppercase", saved_plain["uppercase"]))
|
||||
keep_punctuation = bool(arguments.get("keep_punctuation", saved_plain["keep_punctuation"]))
|
||||
|
||||
filepath, output_path, modifier = _setup_modifier(arguments, "_emphasis_subtitles")
|
||||
|
||||
added: list[tuple[str, int, int, int, int]] = []
|
||||
skipped: list[tuple[str, str]] = []
|
||||
no_review: list[str] = []
|
||||
spine_clips = [el for _, el in modifier._iter_spine_clips()]
|
||||
for el in spine_clips:
|
||||
name = el.get("name", "")
|
||||
if clip_filter and name != clip_filter:
|
||||
continue
|
||||
src = modifier.resources.get(el.get("ref", ""), {}).get("src", "")
|
||||
media_path = media_src_to_path(src)
|
||||
if not media_path or not Path(media_path).is_file():
|
||||
skipped.append((name, "media file missing"))
|
||||
continue
|
||||
data, reason = _load_or_transcribe(media_path, model, language, output_dir)
|
||||
if data is None:
|
||||
skipped.append((name, reason))
|
||||
continue
|
||||
|
||||
spans = _load_emphasis_spans(media_path)
|
||||
if not spans:
|
||||
no_review.append(name)
|
||||
|
||||
clip_source_start = modifier.source_file_start(el).to_seconds()
|
||||
clip_duration = modifier._parse_time(el.get("duration", "0s")).to_seconds()
|
||||
window_end = clip_source_start + clip_duration
|
||||
|
||||
# Clip-relative windows, for deciding which plain titles to hide —
|
||||
# same coordinate space add_text_title's offsets end up in.
|
||||
clip_spans = [
|
||||
(max(0.0, float(s["start"]) - clip_source_start), min(clip_duration, float(s["end"]) - clip_source_start))
|
||||
for s in spans
|
||||
if float(s["end"]) > clip_source_start and float(s["start"]) < window_end
|
||||
]
|
||||
|
||||
all_words = data.get("words", [])
|
||||
|
||||
dynamic_lines = 0
|
||||
dynamic_word_count = 0
|
||||
emphasis_words = _words_in_spans(all_words, spans)
|
||||
clip_emphasis_words = _words_overlapping_clip(emphasis_words, clip_source_start, window_end)
|
||||
if clip_emphasis_words:
|
||||
all_segments = data.get("segments", [])
|
||||
emphasis_segments = _segments_in_spans(all_segments, spans)
|
||||
clip_segments = [
|
||||
{
|
||||
"start": float(s.get("start", 0.0)) - clip_source_start,
|
||||
"end": float(s.get("end", 0.0)) - clip_source_start,
|
||||
}
|
||||
for s in emphasis_segments
|
||||
if float(s.get("end", 0.0)) > clip_source_start
|
||||
and float(s.get("start", 0.0)) < window_end
|
||||
]
|
||||
# Pass the element itself, not `name` — see the same note in
|
||||
# handle_generate_dynamic_subtitles (Engine/docs/05_EXPERIENCIAS.md,
|
||||
# entry 2026-08-17).
|
||||
dynamic_lines = len(
|
||||
modifier.generate_dynamic_subtitles(
|
||||
el, clip_emphasis_words, dynamic_config, segments=clip_segments
|
||||
)
|
||||
)
|
||||
dynamic_word_count = len(clip_emphasis_words)
|
||||
|
||||
# Plain covers EVERY word in the clip — never filtered by emphasis.
|
||||
# Titles landing under a dynamic phrase are disabled below instead of
|
||||
# never being created, so turning emphasis off later never leaves a
|
||||
# silent gap where neither style is on screen.
|
||||
plain_created = 0
|
||||
plain_hidden = 0
|
||||
clip_all_words = _words_overlapping_clip(all_words, clip_source_start, window_end)
|
||||
blocks = _plain_subtitle_blocks(clip_all_words, max_words)
|
||||
for block in blocks:
|
||||
parts = [
|
||||
_plain_word_text(w.get("word", ""), uppercase=uppercase, keep_punctuation=keep_punctuation)
|
||||
for w in block
|
||||
]
|
||||
text = " ".join(p for p in parts if p).strip()
|
||||
if not text:
|
||||
continue
|
||||
start = max(0.0, min(float(w.get("start", 0.0)) for w in block))
|
||||
end = max(float(w.get("end", start)) for w in block)
|
||||
duration = max(end - start, modifier.frame_duration_fraction())
|
||||
title = modifier.add_text_title(
|
||||
el,
|
||||
text,
|
||||
offset=f"{start:.6f}s",
|
||||
duration=f"{duration:.6f}s",
|
||||
lane=20,
|
||||
position=f"0 {position_y:g}",
|
||||
font=plain_font,
|
||||
font_size=plain_font_size,
|
||||
font_color=plain_font_color,
|
||||
bold=True,
|
||||
face=None,
|
||||
font_scale=1.0,
|
||||
size_param=plain_font_size,
|
||||
)
|
||||
plain_created += 1
|
||||
if _overlaps_any_span(start, end, clip_spans):
|
||||
title.set("enabled", "0")
|
||||
plain_hidden += 1
|
||||
|
||||
if dynamic_lines or plain_created:
|
||||
added.append(
|
||||
(name, dynamic_lines, plain_created, plain_hidden, dynamic_word_count + len(clip_all_words))
|
||||
)
|
||||
else:
|
||||
skipped.append((name, "no words in clip's source range"))
|
||||
|
||||
if not added:
|
||||
text = "# Subtitles by Emphasis\n\nNo captions generated — file unchanged (nothing saved)."
|
||||
if skipped:
|
||||
text += "\n\n## Skipped Clips\n" + _markdown_table(
|
||||
["Clip", "Reason"], [[n, r] for n, r in skipped]
|
||||
)
|
||||
return _text_result(text)
|
||||
|
||||
modifier.save(output_path)
|
||||
total_dynamic = sum(d for _, d, _, _, _ in added)
|
||||
total_plain = sum(p for _, _, p, _, _ in added)
|
||||
total_hidden = sum(h for _, _, _, h, _ in added)
|
||||
total_words = sum(w for _, _, _, _, w in added)
|
||||
result = "# Subtitles by Emphasis Generated\n\n## Summary\n"
|
||||
result += (
|
||||
f"- **Clips Captioned**: {len(added)}\n"
|
||||
f"- **Dynamic Title Lines (emphasis)**: {total_dynamic}\n"
|
||||
f"- **Plain Title Blocks (full clip)**: {total_plain}\n"
|
||||
f"- **Plain Blocks Hidden Under Emphasis (enabled=\"0\")**: {total_hidden}\n"
|
||||
f"- **Total Words**: {total_words}\n\n"
|
||||
)
|
||||
result += _markdown_table(
|
||||
["Clip", "Dynamic Lines", "Plain Blocks", "Hidden", "Words"],
|
||||
[[n, str(d), str(p), str(h), str(w)] for n, d, p, h, w in added],
|
||||
)
|
||||
if no_review:
|
||||
result += (
|
||||
"\n## Sem revisão de ênfase\n"
|
||||
"Nenhum `_phrase_actions.json` encontrado para: "
|
||||
+ ", ".join(no_review)
|
||||
+ " — todas as frases desses clipes saíram como legenda comum. "
|
||||
"Rode a etapa 5 do Assistente (revisão de frases) antes, se quiser destaque dinâmico.\n"
|
||||
)
|
||||
if skipped:
|
||||
result += "\n## Skipped Clips\n" + _markdown_table(
|
||||
["Clip", "Reason"], [[n, r] for n, r in skipped]
|
||||
)
|
||||
result += f"\n\nSaved to: `{output_path}`\n\n*Transcripts are cached as _transcript.json; emphasis spans from _phrase_actions.json.*"
|
||||
return _text_result(result)
|
||||
|
||||
|
||||
HANDLERS = {
|
||||
"validate_subtitle_layout": handle_validate_subtitle_layout,
|
||||
"generate_dynamic_subtitles": handle_generate_dynamic_subtitles,
|
||||
"generate_plain_subtitles": handle_generate_plain_subtitles,
|
||||
"generate_subtitles_by_emphasis": handle_generate_subtitles_by_emphasis,
|
||||
}
|
||||
|
||||
Reference in New Issue
Block a user