fix(voz): sobra de fatia interior no corte, e legenda comum poluindo com clipe desativado

Dois problemas reais vistos no projeto Mastopexia:

1. cut_clip_ranges só absorvia um keep-segment curto no INÍCIO/FIM do
   clipe (a lógica já existente do #6). Um keep curto no MEIO (entre dois
   cuts, sem nenhum vizinho mantido pra herdar) nunca era absorvido —
   sobrava como clipe de vídeo de 0,07-0,23s na timeline. Generalizado
   pra qualquer posição, com limiar maior (6 frames / 0,3s, medido no
   material real) — no meio, o pedacinho é descartado (vira parte do
   corte ao redor), nas bordas continua sendo herdado pelo vizinho.

2. generate_subtitles_by_emphasis gerava a legenda comum inteira e
   desativava (enabled="0") onde a dinâmica cobre. Título desativado
   continua aparecendo como clipe riscado na timeline do Final Cut mesmo
   sem renderizar — um corte com bastante ênfase virava dezenas de clipes
   mortos poluindo a trilha (visto ao vivo pelo usuário: "ficou uma
   bosta"). Trocado por não gerar o bloco comum ali, em vez de gerar e
   desativar. Custo: reativar ênfase manualmente depois exige regenerar a
   legenda comum daquele trecho, não só reabilitar.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
This commit is contained in:
João Henrique
2026-08-21 18:59:21 -04:00
co-authored by Claude Sonnet 5
parent 8257155fd3
commit 2ad5854570
4 changed files with 97 additions and 45 deletions
+26 -19
View File
@@ -95,7 +95,7 @@ TOOLS = [
),
Tool(
name="generate_subtitles_by_emphasis",
description="Generate BOTH subtitle styles over the FULL clip and let them coexist by visibility, not by splitting words: plain static titles (see generate_plain_subtitles) cover every word from start to end; dynamic progressive-composition titles (see generate_dynamic_subtitles) are additionally generated for whichever whole phrases were marked as emphasis in the phrase-review step (etapa 5, zoom applied, level >= 1). Wherever a dynamic phrase is on screen, the plain titles underneath it are set enabled=\"0\" (still present in the FCPXML, editable/re-enable-able in Final Cut, just not rendered) instead of never being generated there — so disabling emphasis later never leaves a silent gap in the plain track. Reads emphasis spans from the media's cached '<media>_phrase_actions.json' (written by save_phrase_review after the app's etapa 5 review) — run the voice-editing wizard through that step first, or nothing is treated as emphasis and every title stays plain and enabled. Style knobs are the saved 'Legendas Dinâmicas'/plain-subtitle configs (~/.fcp-mcp-server/config.json); this tool does not expose per-call style overrides, only the split logic — use generate_dynamic_subtitles/generate_plain_subtitles directly if you need one-off styling.",
description="Generate BOTH subtitle styles in one pass, split by word so they never coexist on the same range: dynamic progressive-composition titles (see generate_dynamic_subtitles) cover whichever whole phrases were marked as emphasis in the phrase-review step (etapa 5, zoom applied, level >= 1); plain static titles (see generate_plain_subtitles) cover every OTHER word in the clip. A plain block is simply not created where a dynamic phrase already covers — not created-then-disabled — because a disabled title still shows as its own struck-through clip in Final Cut's timeline even though it never renders, and a heavily emphasized edit ended up with dozens of dead clips cluttering the track. Trade-off: if emphasis is turned off by hand later, the plain line under it has to be regenerated, not just re-enabled. Reads emphasis spans from the media's cached '<media>_phrase_actions.json' (written by save_phrase_review after the app's etapa 5 review) — run the voice-editing wizard through that step first, or nothing is treated as emphasis and every word gets a plain title. Style knobs are the saved 'Legendas Dinâmicas'/plain-subtitle configs (~/.fcp-mcp-server/config.json); this tool does not expose per-call style overrides, only the split logic — use generate_dynamic_subtitles/generate_plain_subtitles directly if you need one-off styling.",
inputSchema={
"type": "object",
"properties": {
@@ -234,7 +234,7 @@ def _segments_in_spans(segments: Sequence[dict], spans: Sequence[dict]) -> list[
def _overlaps_any_span(start: float, end: float, spans: Sequence[tuple[float, float]]) -> bool:
"""Half-open interval overlap: a plain title under this window must hide."""
"""Half-open interval overlap: a plain title under this window is skipped."""
return any(start < span_end and end > span_start for span_start, span_end in spans)
@@ -541,13 +541,16 @@ async def handle_generate_plain_subtitles(arguments: dict) -> Sequence[TextConte
async def handle_generate_subtitles_by_emphasis(arguments: dict) -> Sequence[TextContent]:
"""Generate plain titles for the whole clip and dynamic titles for the
emphasis phrases on top, then hide (enabled="0") the plain titles that
fall under a dynamic phrase — never split the word list between the two.
"""Generate dynamic titles for the emphasis phrases, and plain titles for
every OTHER word — a plain block is simply not created where a dynamic
phrase already covers, rather than created and disabled.
Plain always covers every word, so turning emphasis off later (editing
the phrase review and re-running) never leaves a silent gap: the plain
title was there all along, just disabled.
A disabled ("enabled=0") title still shows as its own struck-through clip
in Final Cut's timeline even though it never renders — a heavily
emphasized edit ended up with dozens of dead clips cluttering the track.
Not generating them there trades that clutter for a smaller gap: if the
emphasis is turned off by hand later, the plain line has to be
regenerated rather than just re-enabled.
"""
model = arguments.get("model", "base")
language = arguments.get("language")
@@ -658,10 +661,14 @@ async def handle_generate_subtitles_by_emphasis(arguments: dict) -> Sequence[Tex
)
dynamic_word_count = len(clip_emphasis_words)
# Plain covers EVERY word in the clip — never filtered by emphasis.
# Titles landing under a dynamic phrase are disabled below instead of
# never being created, so turning emphasis off later never leaves a
# silent gap where neither style is on screen.
# Plain covers every word OUTSIDE an emphasis span. A block landing
# under a dynamic phrase is simply not created there — generating it
# disabled was tried first, but every disabled title still shows up
# as its own clip in Final Cut's timeline (just struck through), so
# a heavily-emphasized edit ended up with dozens of dead clips
# cluttering the track for no visible benefit. The trade-off: if the
# emphasis is later turned off by hand, the plain line under it has
# to be regenerated rather than just re-enabled.
plain_created = 0
plain_hidden = 0
clip_all_words = _words_overlapping_clip(all_words, clip_source_start, window_end)
@@ -676,8 +683,11 @@ async def handle_generate_subtitles_by_emphasis(arguments: dict) -> Sequence[Tex
continue
start = max(0.0, min(float(w.get("start", 0.0)) for w in block))
end = max(float(w.get("end", start)) for w in block)
if _overlaps_any_span(start, end, clip_spans):
plain_hidden += 1
continue
duration = max(end - start, modifier.frame_duration_fraction())
title = modifier.add_text_title(
modifier.add_text_title(
el,
text,
offset=f"{start:.6f}s",
@@ -693,9 +703,6 @@ async def handle_generate_subtitles_by_emphasis(arguments: dict) -> Sequence[Tex
size_param=plain_font_size,
)
plain_created += 1
if _overlaps_any_span(start, end, clip_spans):
title.set("enabled", "0")
plain_hidden += 1
if dynamic_lines or plain_created:
added.append(
@@ -721,12 +728,12 @@ async def handle_generate_subtitles_by_emphasis(arguments: dict) -> Sequence[Tex
result += (
f"- **Clips Captioned**: {len(added)}\n"
f"- **Dynamic Title Lines (emphasis)**: {total_dynamic}\n"
f"- **Plain Title Blocks (full clip)**: {total_plain}\n"
f"- **Plain Blocks Hidden Under Emphasis (enabled=\"0\")**: {total_hidden}\n"
f"- **Plain Title Blocks**: {total_plain}\n"
f"- **Plain Blocks Skipped Under Emphasis (not created there)**: {total_hidden}\n"
f"- **Total Words**: {total_words}\n\n"
)
result += _markdown_table(
["Clip", "Dynamic Lines", "Plain Blocks", "Hidden", "Words"],
["Clip", "Dynamic Lines", "Plain Blocks", "Skipped", "Words"],
[[n, str(d), str(p), str(h), str(w)] for n, d, p, h, w in added],
)
if no_review: