feat(legendas): liga compound_subphrases por padrão no pipeline

generate_dynamic_subtitles e a metade dinâmica de generate_subtitles_by_emphasis
passam a empacotar cada sub-frase da legenda dinâmica num compound clip por
padrão (compound_subphrases=True), completando o wrap_titles_in_compound
e split_into_subphrases do commit anterior — que ainda não tinham chamador
em produção.

Também torna validate_subtitle_layout ciente de compound clips: media cada
grupo (spine principal + cada <media> de compound) no seu próprio espaço de
tempo, em vez de uma varredura .//title global — sem isso, âncoras de
compounds diferentes liam offset "0s" e acusavam colisão espacial entre
frases que nunca dividem a tela, só porque compartilham o mesmo zero de
tempo local.

Testado ponta a ponta na gravação real (Mastopexia): 12 compounds, 41
títulos todos empacotados, zero soltos, zero IDs duplicados, DTD válida.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
This commit is contained in:
João Henrique
2026-08-26 17:05:20 -04:00
co-authored by Claude Sonnet 5
parent 688bdeddb6
commit c99274895c
3 changed files with 355 additions and 108 deletions
+41 -4
View File
@@ -1,6 +1,6 @@
# 03 — Camada MCP (`server.py` + `server_tools/`) — 77 ferramentas # 03 — Camada MCP (`server.py` + `server_tools/`) — 78 ferramentas
> **Escopo:** As 77 ferramentas MCP: helpers, categorias e como criar uma nova. > **Escopo:** As 78 ferramentas MCP: helpers, categorias e como criar uma nova.
> **Não cobre:** Lógica de edição, que mora no engine (→ 02) · comandos do app (→ 08) > **Não cobre:** Lógica de edição, que mora no engine (→ 02) · comandos do app (→ 08)
`server.py` (592 linhas) é só o transporte: dispatch por dicionário `server.py` (592 linhas) é só o transporte: dispatch por dicionário
@@ -82,8 +82,21 @@ continua funcionando. A coluna diz o módulo real, para quando você precisar
### Voz (análise → decisão → aplicação) ### Voz (análise → decisão → aplicação)
`analyze_voice_features`, `build_voice_timeline`, `refine_voice_timeline`, `analyze_voice_features`, `build_voice_timeline`, `refine_voice_timeline`,
`remove_speakers`, `apply_voice_actions`, `generate_voice_script`, `remove_speakers`, `remove_speech_gaps`, `apply_voice_actions`,
`get_voice_analysis_config`, `save_voice_analysis_config`. `generate_voice_script`, `get_voice_analysis_config`,
`save_voice_analysis_config`.
`remove_speech_gaps` corta pelo que a **transcrição** já sabe que não tem
fala — lê `words[].start/end` do `_voice_timeline.json` (função
`speech_gap_cut_actions`, em `fcpxml/voice_actions.py`) em vez de medir
volume. É o complemento correto para o caso que `remove_media_silence`
(silêncio por dB, ver seção "Silêncio e beats") não cobre: um trecho sem
fala mas com som real acima do limiar (respiração, ruído de roupa, batida) —
`remove_media_silence` nunca vai cortar isso porque tecnicamente não é
silêncio. Não corta a lacuna antes da primeiríssima palavra (pode ser quase
o arquivo inteiro, antes da tomada realmente começar) — isso continua
decisão manual na Fase 6 do `apply_voice_actions`
(`.claude/skills/editar-por-voz/criterios/06-texto-corte-marcador.md`).
O fluxo manual é: `build_voice_timeline` mede (caro, roda uma vez) → O fluxo manual é: `build_voice_timeline` mede (caro, roda uma vez) →
o modelo decide os cortes → **`refine_voice_timeline` renormaliza sobre o que o modelo decide os cortes → **`refine_voice_timeline` renormaliza sobre o que
@@ -143,6 +156,30 @@ e avisa no relatório ("Sem revisão de ênfase"). Não expõe overrides de esti
por chamada — usa a config salva ("Legendas Dinâmicas"/plain); para estilo por chamada — usa a config salva ("Legendas Dinâmicas"/plain); para estilo
pontual, use `generate_dynamic_subtitles`/`generate_plain_subtitles` direto. pontual, use `generate_dynamic_subtitles`/`generate_plain_subtitles` direto.
**Separação por role (didática na timeline):** `generate_dynamic_subtitles` e
`generate_plain_subtitles` (e a metade dinâmica/comum do `by_emphasis`) aplicam
`role="titles.dinamicas"` e `role="titles.convencionais"` em cada `<title>`
criado — sub-roles de `titles`, **nunca** `subtitles.*` (que esconderia o título
atrás de Code). O parâmetro `role` de cada ferramenta MCP sobrescreve o default
(vindo de `load_dynamic_subtitle_config()["role"]` /
`load_plain_subtitle_config()["role"]`). Ver `02_MODULES.md` (seção "Separação
de role") e `05_EXPERIENCIAS.md` entrada 32 (DTD: `<title>` leva `role`, não
`videoRole`).
**Compound clip por sub-frase (padrão em `generate_dynamic_subtitles` e na
metade dinâmica do `by_emphasis`):** `compound_subphrases=True` divide cada
frase em sub-frases pela vírgula (`transcribe.split_into_subphrases`) e
empacota os `<title>` de cada uma num `<ref-clip>` — a dúzia de títulos
empilhados por lane que uma frase gera vira uma barra só, arrastável/mutável
como unidade. Exceção: um trecho curto depois da vírgula ("né?", "Então...",
< 3 palavras) funde de volta na sub-frase anterior em vez de virar compound
próprio — soa como parte da mesma respiração, não uma frase nova. A estrutura
replica o que o próprio Final Cut gera em "New Compound Clip": o título mais
cedo vira âncora do spine interno em offset 0, os demais penduram nele por
lane. `validate_subtitle_layout` mede cada compound no seu próprio espaço de
tempo — sem isso, âncoras de compounds diferentes leem "0s" e colidem no
papel mesmo estando segundos distantes na timeline real.
**Antes de atribuir uma colisão ao gerador, confirme que é o gerador.** **Antes de atribuir uma colisão ao gerador, confirme que é o gerador.**
Um `<title>` de nome estranho (`ref` diferente, params tipo `Auto-Shrink`/ Um `<title>` de nome estranho (`ref` diferente, params tipo `Auto-Shrink`/
`Left Margin` que `_make_text_title_clip` nunca escreve) é conteúdo humano `Left Margin` que `_make_text_title_clip` nunca escreve) é conteúdo humano
+157 -27
View File
@@ -3,6 +3,7 @@
Extraído de writer.py — ver fcpxml/writer/__init__.py para o conjunto. Extraído de writer.py — ver fcpxml/writer/__init__.py para o conjunto.
""" """
import random
import re import re
import unicodedata import unicodedata
import uuid import uuid
@@ -20,7 +21,7 @@ from ..text_layout import (
compose_sentence, compose_sentence,
layout_sentence, layout_sentence,
) )
from ..transcribe import group_words_by_segment from ..transcribe import group_words_by_segment, split_into_subphrases
from .helpers import _dtd_insert, _sanitize_xml_value from .helpers import _dtd_insert, _sanitize_xml_value
@@ -221,6 +222,7 @@ class TitlesMixin:
font_scale: float = TEXT_TEMPLATE_FONT_SCALE, font_scale: float = TEXT_TEMPLATE_FONT_SCALE,
animated: bool = True, animated: bool = True,
size_param: Optional[float] = None, size_param: Optional[float] = None,
role: Optional[str] = None,
) -> ET.Element: ) -> ET.Element:
"""Build a standalone ``<title>`` clip from the "Text" (Basic Text) template. """Build a standalone ``<title>`` clip from the "Text" (Basic Text) template.
@@ -238,6 +240,8 @@ class TitlesMixin:
elem.set('name', _sanitize_xml_value(name, 256)) elem.set('name', _sanitize_xml_value(name, 256))
elem.set('start', self._TEXT_TITLE_START) elem.set('start', self._TEXT_TITLE_START)
elem.set('duration', duration.to_fcpxml()) elem.set('duration', duration.to_fcpxml())
if role:
elem.set('role', _sanitize_xml_value(role, 256))
if position: if position:
param = ET.SubElement(elem, 'param') param = ET.SubElement(elem, 'param')
@@ -328,6 +332,7 @@ class TitlesMixin:
animated: bool = True, animated: bool = True,
font_scale: float = TEXT_TEMPLATE_FONT_SCALE, font_scale: float = TEXT_TEMPLATE_FONT_SCALE,
size_param: Optional[float] = None, size_param: Optional[float] = None,
role: Optional[str] = None,
) -> ET.Element: ) -> ET.Element:
"""Add a single static "Text" (Basic Text) title over *parent_clip*. """Add a single static "Text" (Basic Text) title over *parent_clip*.
@@ -365,6 +370,7 @@ class TitlesMixin:
animated=animated, animated=animated,
font_scale=font_scale, font_scale=font_scale,
size_param=size_param, size_param=size_param,
role=role,
) )
_dtd_insert(parent, title) _dtd_insert(parent, title)
return title return title
@@ -375,6 +381,10 @@ class TitlesMixin:
words: List[Dict[str, Any]], words: List[Dict[str, Any]],
config: Optional['DynamicSubtitleConfig'] = None, config: Optional['DynamicSubtitleConfig'] = None,
segments: Optional[List[Dict[str, Any]]] = None, segments: Optional[List[Dict[str, Any]]] = None,
role: Optional[str] = None,
configs: Optional[List['DynamicSubtitleConfig']] = None,
compound_subphrases: bool = False,
subphrase_min_words: int = 3,
) -> List[ET.Element]: ) -> List[ET.Element]:
"""Generate progressive-reveal subtitle titles, one per word. """Generate progressive-reveal subtitle titles, one per word.
@@ -418,8 +428,16 @@ class TitlesMixin:
Returns: Returns:
The list of created ``<title>`` elements, in chronological order. The list of created ``<title>`` elements, in chronological order.
""" """
if config is None: # ``configs`` (a list of registered, active layouts) takes precedence
config = DynamicSubtitleConfig() # over the single ``config`` — with 2+ items, each block picks one at
# random below; with 0 or 1, behaviour is identical to a single fixed
# config, so old callers passing only ``config`` are unaffected.
if configs:
layout_configs = list(configs)
elif config is not None:
layout_configs = [config]
else:
layout_configs = [DynamicSubtitleConfig()]
if not words: if not words:
return [] return []
@@ -449,35 +467,54 @@ class TitlesMixin:
# next block — the sub-sentence split that keeps long sentences from # next block — the sub-sentence split that keeps long sentences from
# spilling off screen. # spilling off screen.
sentences = group_words_by_segment(words, segments or []) sentences = group_words_by_segment(words, segments or [])
box = LayoutBox.for_frame( # A comma is where the sentence breathes, so it is also where the
self.frame_width(), self.frame_height(), # phrase should be packed into its own compound clip downstream.
band_height=config.band_height, if compound_subphrases:
center_y=config.block_center_y, sentences = [
) sub
# "phrase" is the progressive composition the reference reel uses: one for sentence in sentences
# title per LINE ("que vão" / "melhorar" / "sua legenda"), the key word for sub in split_into_subphrases(sentence, subphrase_min_words)
# set large in a display italic. "word" is the older one-title-per-word ]
# rhythm, kept for callers that want every word to land on its own.
phrase_mode = getattr(config, 'granularity', 'phrase') == 'phrase'
def lay_out(pending: List[Dict]): def box_for(cfg: 'DynamicSubtitleConfig') -> LayoutBox:
return LayoutBox.for_frame(
self.frame_width(), self.frame_height(),
band_height=cfg.band_height,
center_y=cfg.block_center_y,
)
def lay_out(pending: List[Dict], cfg: 'DynamicSubtitleConfig', box: LayoutBox):
"""Place what fits; return (units, still-unplaced words).""" """Place what fits; return (units, still-unplaced words)."""
if phrase_mode: # "phrase" is the progressive composition the reference reel uses:
# one title per LINE ("que vão" / "melhorar" / "sua legenda"), the
# key word set large in a display italic. "word" is the older
# one-title-per-word rhythm, kept for callers that want every word
# to land on its own.
if getattr(cfg, 'granularity', 'phrase') == 'phrase':
composition = compose_sentence( composition = compose_sentence(
pending, config.style, box, line_gap=config.line_gap, pending, cfg.style, box, line_gap=cfg.line_gap,
) )
return composition.blocks, composition.overflow return composition.blocks, composition.overflow
layout = layout_sentence(pending, config.style, box) layout = layout_sentence(pending, cfg.style, box)
return layout.placed, layout.overflow return layout.placed, layout.overflow
blocks: List[List[Any]] = [] blocks: List[List[Any]] = []
for sentence in sentences: block_configs: List['DynamicSubtitleConfig'] = []
block_sentences: List[int] = []
for sentence_index, sentence in enumerate(sentences):
remaining = list(sentence) remaining = list(sentence)
while remaining: while remaining:
units, remaining = lay_out(remaining) # Each block independently samples a layout from the active
# set — the visual variety the user asked for. A single
# active layout always resolves to itself, so this is a
# no-op for the common case.
active_config = layout_configs[random.randrange(len(layout_configs))]
units, remaining = lay_out(remaining, active_config, box_for(active_config))
if not units: if not units:
break break
blocks.append(units) blocks.append(units)
block_configs.append(active_config)
block_sentences.append(sentence_index)
if not blocks: if not blocks:
return [] return []
@@ -523,7 +560,15 @@ class TitlesMixin:
media_origin = self._parse_time(parent.get('start', '0s')) media_origin = self._parse_time(parent.get('start', '0s'))
created: List[ET.Element] = [] created: List[ET.Element] = []
for units, block_end in zip(blocks, block_ends): by_sentence: Dict[int, List[ET.Element]] = {}
for units, block_end, block_config, sentence_index in zip(
blocks, block_ends, block_configs, block_sentences
):
# A ``titles.*`` sub-role keeps these as titles (never closed
# captions) while grouping them in the role index and tinting
# their lane. An explicit ``role`` argument overrides every
# block; otherwise each block uses its own sampled layout's role.
block_role = role or getattr(block_config, "role", None) or "titles.dinamicas"
for index, unit in enumerate(units): for index, unit in enumerate(units):
relative_offset = self.snap_seconds_to_frame(unit.start) relative_offset = self.snap_seconds_to_frame(unit.start)
duration = block_end - relative_offset duration = block_end - relative_offset
@@ -543,19 +588,34 @@ class TitlesMixin:
duration, duration,
lane=lane, lane=lane,
name=f"caption_{uuid.uuid4().hex[:8]}", name=f"caption_{uuid.uuid4().hex[:8]}",
position=unit.position_param(config.text_scale), position=unit.position_param(block_config.text_scale),
font=unit.font or config.style.font, font=unit.font or block_config.style.font,
font_size=int(round(unit.font_size)), font_size=int(round(unit.font_size)),
font_color=unit.color or config.style.active_color, font_color=unit.color or block_config.style.active_color,
bold=config.style.bold, bold=block_config.style.bold,
face=unit.face, face=unit.face,
kerning=unit.kerning, kerning=unit.kerning,
font_scale=config.text_scale, font_scale=block_config.text_scale,
role=block_role,
) )
_dtd_insert(parent, title) _dtd_insert(parent, title)
created.append(title) created.append(title)
by_sentence.setdefault(sentence_index, []).append(title)
if getattr(config, 'validate', False): # One compound per sub-phrase: a dozen stacked title bars collapse
# into a single one that can be dragged, muted or retimed as a unit.
if compound_subphrases:
for sentence_index in sorted(by_sentence):
group = by_sentence[sentence_index]
label = " ".join(
str(w.get('word') or w.get('text') or '')
for w in sentences[sentence_index]
).strip()
self.wrap_titles_in_compound(
parent, group, name=label[:60] or "Legenda"
)
if any(getattr(cfg, 'validate', False) for cfg in layout_configs):
report = self.validate_subtitle_layout() report = self.validate_subtitle_layout()
if blocking(report["severity"]): if blocking(report["severity"]):
raise ValueError( raise ValueError(
@@ -586,8 +646,78 @@ class TitlesMixin:
Returns the ``collision.validate_titles`` report: ``severity`` (worst Returns the ``collision.validate_titles`` report: ``severity`` (worst
bucket), ``issues`` (spec-16 occurrences) and ``summary`` (counts). bucket), ``issues`` (spec-16 occurrences) and ``summary`` (counts).
""" """
# A compound clip carries its own time origin: a title inside one is
# offset from that compound's start, not the sequence's. Measured in
# one flat pass, the anchors of two different compounds both read as
# "0s" and collide on paper while sitting seconds apart on the
# timeline. Each compound is therefore measured as its own scope,
# which is also where its titles can actually overlap — a title can
# only share the screen with its own compound's siblings.
scopes: List[List[ET.Element]] = []
nested: set = set()
for media in self.root.findall('.//media'):
group = list(media.iter('title'))
if group:
scopes.append(group)
nested.update(id(t) for t in group)
main = [t for t in self.root.iter('title') if id(t) not in nested]
if main:
scopes.append(main)
reports = [
self._measure_title_scope(
scope,
safe_margin_x=safe_margin_x,
safe_margin_y=safe_margin_y,
min_font_size=min_font_size,
min_distance=min_distance,
max_distance=max_distance,
)
for scope in scopes
]
if len(reports) == 1:
return reports[0]
if not reports:
return self._measure_title_scope(
[],
safe_margin_x=safe_margin_x,
safe_margin_y=safe_margin_y,
min_font_size=min_font_size,
min_distance=min_distance,
max_distance=max_distance,
)
rank = {
'none': 0, 'render_tolerance': 1, 'warning': 2,
'probable': 3, 'severe': 4,
}
merged_issues = [i for r in reports for i in r['issues']]
summary = dict(reports[0]['summary'])
for r in reports[1:]:
for key, value in r['summary'].items():
summary[key] = summary.get(key, 0) + value
return {
'severity': max(
(r['severity'] for r in reports),
key=lambda s: rank.get(s, 0),
),
'issues': merged_issues,
'summary': summary,
}
def _measure_title_scope(
self,
elements: List[ET.Element],
*,
safe_margin_x: float,
safe_margin_y: float,
min_font_size: Optional[float],
min_distance: Optional[float],
max_distance: Optional[float],
) -> dict:
"""Measure and validate one group of titles sharing a time origin."""
titles = [] titles = []
for elem in self.root.iter('title'): for elem in elements:
# enabled="0" never renders in Final Cut (see # enabled="0" never renders in Final Cut (see
# generate_subtitles_by_emphasis, which disables plain titles # generate_subtitles_by_emphasis, which disables plain titles
# under an emphasis phrase instead of never creating them) — a # under an emphasis phrase instead of never creating them) — a
+155 -75
View File
@@ -13,7 +13,10 @@ from typing import Sequence
from mcp.types import TextContent, Tool from mcp.types import TextContent, Tool
from fcpxml.media_intel import media_src_to_path from fcpxml.media_intel import media_src_to_path
from fcpxml.model_manager import load_dynamic_subtitle_config, load_plain_subtitle_config from fcpxml.model_manager import (
get_active_dynamic_subtitle_layouts,
load_plain_subtitle_config,
)
from fcpxml.models import DynamicSubtitleConfig, WordLook, WordStyle from fcpxml.models import DynamicSubtitleConfig, WordLook, WordStyle
from fcpxml.writer import FCPXMLModifier from fcpxml.writer import FCPXMLModifier
from server_tools._shared import ( from server_tools._shared import (
@@ -61,6 +64,7 @@ TOOLS = [
"emphasis_size": {"type": "integer", "description": "Key-word size in canvas points, at the 2160x3840 reference frame. Falls back to the saved style (default 265)"}, "emphasis_size": {"type": "integer", "description": "Key-word size in canvas points, at the 2160x3840 reference frame. Falls back to the saved style (default 265)"},
"emphasis_color": {"type": "string", "description": "RGBA (0-1, space-separated) for the key word (phrase mode). Defaults to active_color, so the block reads in a single colour unless the key word is deliberately set apart"}, "emphasis_color": {"type": "string", "description": "RGBA (0-1, space-separated) for the key word (phrase mode). Defaults to active_color, so the block reads in a single colour unless the key word is deliberately set apart"},
"text_scale": {"type": "number", "description": "Ratio between the title template's fontSize space and the canvas-point space it positions in. The \"Text\" template sizes type in frame pixels, so sizes are doubled on the way out. Falls back to the saved style (default 2.0). Lower it only if a template renders type larger than the chosen point size"}, "text_scale": {"type": "number", "description": "Ratio between the title template's fontSize space and the canvas-point space it positions in. The \"Text\" template sizes type in frame pixels, so sizes are doubled on the way out. Falls back to the saved style (default 2.0). Lower it only if a template renders type larger than the chosen point size"},
"role": {"type": "string", "description": "Final Cut role for every generated title (a 'titles.*' sub-role, never 'subtitles.*'). Falls back to the saved style (default 'titles.dinamicas'). Groups the clips in the role index and tints their lane."},
"font": {"type": "string", "description": "Title font family (supporting lines in phrase mode). Falls back to the saved style (default 'Helvetica Neue')"}, "font": {"type": "string", "description": "Title font family (supporting lines in phrase mode). Falls back to the saved style (default 'Helvetica Neue')"},
"font_size": {"type": "integer", "description": "Supporting-line font size in canvas points, at the 2160x3840 reference frame. Falls back to the saved style (default 104)"}, "font_size": {"type": "integer", "description": "Supporting-line font size in canvas points, at the 2160x3840 reference frame. Falls back to the saved style (default 104)"},
"active_color": {"type": "string", "description": "RGBA (0-1, space-separated) for even-indexed lines. Falls back to the saved style (default '1 1 1 1')"}, "active_color": {"type": "string", "description": "RGBA (0-1, space-separated) for even-indexed lines. Falls back to the saved style (default '1 1 1 1')"},
@@ -88,6 +92,7 @@ TOOLS = [
"uppercase": {"type": "boolean", "description": "Render text in uppercase."}, "uppercase": {"type": "boolean", "description": "Render text in uppercase."},
"keep_punctuation": {"type": "boolean", "description": "Keep punctuation such as comma and period."}, "keep_punctuation": {"type": "boolean", "description": "Keep punctuation such as comma and period."},
"text_scale": {"type": "number", "description": "Template font-size scale. Falls back to saved plain-subtitle config."}, "text_scale": {"type": "number", "description": "Template font-size scale. Falls back to saved plain-subtitle config."},
"role": {"type": "string", "description": "Final Cut role for every generated title (a 'titles.*' sub-role, never 'subtitles.*'). Falls back to the saved style (default 'titles.convencionais'). Groups the clips in the role index and tints their lane."},
"output_path": {"type": "string", "description": "Output path (default: adds _plain_subtitles suffix)"}, "output_path": {"type": "string", "description": "Output path (default: adds _plain_subtitles suffix)"},
}, },
"required": ["filepath"] "required": ["filepath"]
@@ -116,6 +121,31 @@ TOOLS = [
_PUNCT_RE = re.compile(r"[^\w\sÀ-ÖØ-öø-ÿ]", re.UNICODE) _PUNCT_RE = re.compile(r"[^\w\sÀ-ÖØ-öø-ÿ]", re.UNICODE)
_DEFAULT_DYNAMIC_ROLE = "titles.dinamicas"
_DEFAULT_PLAIN_ROLE = "titles.convencionais"
def _title_subrole(value: str | None, fallback: str) -> str:
"""Return a Final Cut title sub-role, never a closed-caption role."""
role = str(value or "").strip() or fallback
if role.startswith("subtitles."):
return "titles." + role.removeprefix("subtitles.")
if role == "subtitles":
return fallback
if not role.startswith("titles."):
return fallback
return role
def _separate_subtitle_roles(dynamic_role: str | None, plain_role: str | None) -> tuple[str, str]:
"""Keep normal and dynamic subtitles in distinct Final Cut role lanes."""
dynamic = _title_subrole(dynamic_role, _DEFAULT_DYNAMIC_ROLE)
plain = _title_subrole(plain_role, _DEFAULT_PLAIN_ROLE)
if dynamic == plain:
if dynamic != _DEFAULT_DYNAMIC_ROLE:
return dynamic, _DEFAULT_PLAIN_ROLE
return _DEFAULT_DYNAMIC_ROLE, _DEFAULT_PLAIN_ROLE
return dynamic, plain
def _words_overlapping_clip(words: Sequence[dict], start: float, end: float) -> list[dict]: def _words_overlapping_clip(words: Sequence[dict], start: float, end: float) -> list[dict]:
@@ -171,12 +201,14 @@ def _phrase_actions_path(media_path: str) -> Path:
return Path(media_path).with_name(f"{stem}_phrase_actions.json") return Path(media_path).with_name(f"{stem}_phrase_actions.json")
def _load_emphasis_spans(media_path: str) -> list[dict]: def _load_review_spans(media_path: str, key: str) -> list[dict]:
"""Load emphasis spans (source-media time) saved by the etapa-5 phrase review. """Load one span list (source-media time) saved by the etapa-5 phrase review.
Returns [] if the review was never run for this media — callers should treat ``key`` is ``"emphasis_spans"`` (phrases with `subtitle_dynamic` on) or
that as "nothing is emphasis yet", not as an error, since the wizard's later ``"plain_exclude_spans"`` (phrases with `subtitle_common` off). Returns []
steps are optional. if the review was never run for this media, or saved nothing under that
key — callers should treat that as "nothing marked", not as an error,
since the wizard's later steps are optional.
""" """
path = _phrase_actions_path(media_path) path = _phrase_actions_path(media_path)
if not path.is_file(): if not path.is_file():
@@ -185,10 +217,25 @@ def _load_emphasis_spans(media_path: str) -> list[dict]:
data = json.loads(path.read_text(encoding="utf-8")) data = json.loads(path.read_text(encoding="utf-8"))
except (OSError, json.JSONDecodeError): except (OSError, json.JSONDecodeError):
return [] return []
spans = data.get("emphasis_spans", []) spans = data.get(key, [])
return [s for s in spans if isinstance(s, dict) and "start" in s and "end" in s] return [s for s in spans if isinstance(s, dict) and "start" in s and "end" in s]
def _load_emphasis_spans(media_path: str) -> list[dict]:
"""Spans (source-media time) whose phrase has `subtitle_dynamic` on."""
return _load_review_spans(media_path, "emphasis_spans")
def _load_plain_exclude_spans(media_path: str) -> list[dict]:
"""Spans (source-media time) whose phrase has `subtitle_common` off.
Independent from emphasis spans: a phrase can have `subtitle_common` off
without being emphasized, so plain must be hidden there too even though
no dynamic line is going to cover the gap.
"""
return _load_review_spans(media_path, "plain_exclude_spans")
def _word_in_spans(word_start: float, word_end: float, spans: Sequence[dict]) -> bool: def _word_in_spans(word_start: float, word_end: float, spans: Sequence[dict]) -> bool:
"""A word belongs to an emphasis span if its midpoint falls inside it. """A word belongs to an emphasis span if its midpoint falls inside it.
@@ -302,6 +349,46 @@ async def handle_validate_subtitle_layout(arguments: dict) -> Sequence[TextConte
return _text_result("\n".join(lines)) return _text_result("\n".join(lines))
def _build_dynamic_subtitle_config(saved: dict, overrides: dict | None = None) -> DynamicSubtitleConfig:
"""Build a :class:`DynamicSubtitleConfig` from one registered layout dict.
``overrides`` (typically the tool call's own ``arguments``) only makes
sense to apply when there is a single active layout — callers with 2+
active layouts pass ``{}`` so every sampled block uses its layout as
registered, unambiguously.
"""
overrides = overrides or {}
body_color = overrides.get("active_color") or saved["active_color"]
return DynamicSubtitleConfig(
style=WordStyle(
font=overrides.get("font") or saved["font"],
font_size=int(overrides.get("font_size", saved["font_size"])),
active_color=body_color,
inactive_color=overrides.get("inactive_color", "0.7 0.7 0.7 1"),
emphasis_look=WordLook(
int(overrides.get("emphasis_size", saved["emphasis_size"])),
overrides.get("emphasis_color") or saved["emphasis_color"] or body_color,
font=overrides.get("emphasis_font") or saved["emphasis_font"],
face=overrides.get("emphasis_face") or saved["emphasis_face"],
kerning=0.0,
),
body_look=WordLook(
int(overrides.get("font_size", saved["font_size"])),
body_color,
font=overrides.get("font") or saved["font"],
face="Bold",
kerning=1.2,
),
),
band_height=float(overrides.get("band_height", saved["band_height"])),
block_center_y=float(overrides.get("block_center_y", saved["block_center_y"])),
granularity=overrides.get("granularity", "phrase"),
text_scale=float(overrides.get("text_scale", saved["text_scale"])),
line_gap=float(overrides.get("line_gap", saved["line_gap"])),
role=_title_subrole(saved.get("role"), _DEFAULT_DYNAMIC_ROLE),
)
async def handle_generate_dynamic_subtitles(arguments: dict) -> Sequence[TextContent]: async def handle_generate_dynamic_subtitles(arguments: dict) -> Sequence[TextContent]:
"""Generate per-word subtitle titles laid out as a block per sentence. """Generate per-word subtitle titles laid out as a block per sentence.
@@ -318,39 +405,21 @@ async def handle_generate_dynamic_subtitles(arguments: dict) -> Sequence[TextCon
output_dir = arguments.get("output_dir") output_dir = arguments.get("output_dir")
clip_filter = arguments.get("clip_name") clip_filter = arguments.get("clip_name")
# Anything the caller didn't explicitly pass falls back to the style # Anything the caller didn't explicitly pass falls back to the style(s)
# persisted from the "Legendas Dinâmicas" screen (~/.fcp-mcp-server/ # persisted from the "Legendas Dinâmicas" screen (~/.fcp-mcp-server/
# config.json), not a hardcoded default — so the UI is the single place # config.json) — one or more named, active layouts. With exactly one
# that configures the look, and every caller (app, MCP, this session) # active layout, per-call overrides (arguments) still apply, same as
# renders the same thing without threading 11 fields through every call. # before this screen supported multiple layouts. With 2+ active layouts,
saved = load_dynamic_subtitle_config() # each block below randomly samples one of them, so per-call overrides
body_color = arguments.get("active_color") or saved["active_color"] # are ambiguous (which layout would they apply to?) and are ignored —
config = DynamicSubtitleConfig( # register/edit the layouts themselves instead.
style=WordStyle( active_layouts = get_active_dynamic_subtitle_layouts()
font=arguments.get("font") or saved["font"], overrides = arguments if len(active_layouts) == 1 else {}
font_size=int(arguments.get("font_size", saved["font_size"])), configs = [_build_dynamic_subtitle_config(saved, overrides) for saved in active_layouts]
active_color=body_color, single_role_override = (
inactive_color=arguments.get("inactive_color", "0.7 0.7 0.7 1"), _title_subrole(arguments.get("role"), configs[0].role)
emphasis_look=WordLook( if len(active_layouts) == 1 and arguments.get("role")
int(arguments.get("emphasis_size", saved["emphasis_size"])), else None
arguments.get("emphasis_color") or saved["emphasis_color"] or body_color,
font=arguments.get("emphasis_font") or saved["emphasis_font"],
face=arguments.get("emphasis_face") or saved["emphasis_face"],
kerning=0.0,
),
body_look=WordLook(
int(arguments.get("font_size", saved["font_size"])),
body_color,
font=arguments.get("font") or saved["font"],
face="Bold",
kerning=1.2,
),
),
band_height=float(arguments.get("band_height", saved["band_height"])),
block_center_y=float(arguments.get("block_center_y", saved["block_center_y"])),
granularity=arguments.get("granularity", "phrase"),
text_scale=float(arguments.get("text_scale", saved["text_scale"])),
line_gap=float(arguments.get("line_gap", saved["line_gap"])),
) )
filepath, output_path, modifier = _setup_modifier(arguments, "_dynamic_subtitles") filepath, output_path, modifier = _setup_modifier(arguments, "_dynamic_subtitles")
@@ -402,7 +471,9 @@ async def handle_generate_dynamic_subtitles(arguments: dict) -> Sequence[TextCon
# clip's captions onto a single wrong spine element instead of each # clip's captions onto a single wrong spine element instead of each
# clip's own. See Engine/docs/05_EXPERIENCIAS.md, entry 2026-08-17. # clip's own. See Engine/docs/05_EXPERIENCIAS.md, entry 2026-08-17.
lines = modifier.generate_dynamic_subtitles( lines = modifier.generate_dynamic_subtitles(
el, clip_words, config, segments=clip_segments el, clip_words, configs=configs, segments=clip_segments,
role=single_role_override,
compound_subphrases=True,
) )
added.append((name, len(lines), len(clip_words))) added.append((name, len(lines), len(clip_words)))
@@ -443,6 +514,10 @@ async def handle_generate_plain_subtitles(arguments: dict) -> Sequence[TextConte
clip_filter = arguments.get("clip_name") clip_filter = arguments.get("clip_name")
saved = load_plain_subtitle_config() saved = load_plain_subtitle_config()
saved["role"] = _title_subrole(
arguments.get("role") or saved.get("role"),
_DEFAULT_PLAIN_ROLE,
)
font = arguments.get("font") or saved["font"] font = arguments.get("font") or saved["font"]
font_size = int(arguments.get("font_size", saved["font_size"])) font_size = int(arguments.get("font_size", saved["font_size"]))
font_color = arguments.get("font_color") or saved["font_color"] font_color = arguments.get("font_color") or saved["font_color"]
@@ -506,6 +581,7 @@ async def handle_generate_plain_subtitles(arguments: dict) -> Sequence[TextConte
face=None, face=None,
font_scale=1.0, font_scale=1.0,
size_param=font_size, size_param=font_size,
role=saved["role"],
) )
created += 1 created += 1
if created: if created:
@@ -558,37 +634,30 @@ async def handle_generate_subtitles_by_emphasis(arguments: dict) -> Sequence[Tex
clip_filter = arguments.get("clip_name") clip_filter = arguments.get("clip_name")
granularity = arguments.get("granularity", "phrase") granularity = arguments.get("granularity", "phrase")
saved_dynamic = load_dynamic_subtitle_config() # One or more named, active layouts — with 2+ active, each emphasis block
body_color = saved_dynamic["active_color"] # below randomly samples one of them (see generate_dynamic_subtitles).
dynamic_config = DynamicSubtitleConfig( active_dynamic_layouts = get_active_dynamic_subtitle_layouts()
style=WordStyle( dynamic_configs = [
font=saved_dynamic["font"], _build_dynamic_subtitle_config(saved, {"granularity": granularity})
font_size=int(saved_dynamic["font_size"]), for saved in active_dynamic_layouts
active_color=body_color, ]
inactive_color="0.7 0.7 0.7 1", single_dynamic_role = (
emphasis_look=WordLook( active_dynamic_layouts[0]["role"] if len(active_dynamic_layouts) == 1 else None
int(saved_dynamic["emphasis_size"]),
saved_dynamic["emphasis_color"] or body_color,
font=saved_dynamic["emphasis_font"],
face=saved_dynamic["emphasis_face"],
kerning=0.0,
),
body_look=WordLook(
int(saved_dynamic["font_size"]),
body_color,
font=saved_dynamic["font"],
face="Bold",
kerning=1.2,
),
),
band_height=float(saved_dynamic["band_height"]),
block_center_y=float(saved_dynamic["block_center_y"]),
granularity=granularity,
text_scale=float(saved_dynamic["text_scale"]),
line_gap=float(saved_dynamic["line_gap"]),
) )
saved_plain = load_plain_subtitle_config() saved_plain = load_plain_subtitle_config()
dynamic_role, plain_role = _separate_subtitle_roles(
single_dynamic_role or dynamic_configs[0].role,
saved_plain.get("role"),
)
if len(active_dynamic_layouts) == 1:
single_dynamic_role = dynamic_role
else:
for cfg in dynamic_configs:
cfg.role = _title_subrole(cfg.role, _DEFAULT_DYNAMIC_ROLE)
if cfg.role == plain_role:
cfg.role = _DEFAULT_DYNAMIC_ROLE
saved_plain["role"] = plain_role
plain_font = saved_plain["font"] plain_font = saved_plain["font"]
plain_font_size = int(saved_plain["font_size"]) plain_font_size = int(saved_plain["font_size"])
plain_font_color = saved_plain["font_color"] plain_font_color = saved_plain["font_color"]
@@ -618,21 +687,29 @@ async def handle_generate_subtitles_by_emphasis(arguments: dict) -> Sequence[Tex
continue continue
spans = _load_emphasis_spans(media_path) spans = _load_emphasis_spans(media_path)
if not spans: exclude_spans = _load_plain_exclude_spans(media_path)
if not spans and not exclude_spans:
no_review.append(name) no_review.append(name)
clip_source_start = modifier.source_file_start(el).to_seconds() clip_source_start = modifier.source_file_start(el).to_seconds()
clip_duration = modifier._parse_time(el.get("duration", "0s")).to_seconds() clip_duration = modifier._parse_time(el.get("duration", "0s")).to_seconds()
window_end = clip_source_start + clip_duration window_end = clip_source_start + clip_duration
# Clip-relative windows, for deciding which plain titles to hide — def _clip_relative(span_list: list[dict]) -> list[tuple[float, float]]:
# same coordinate space add_text_title's offsets end up in. return [
clip_spans = [
(max(0.0, float(s["start"]) - clip_source_start), min(clip_duration, float(s["end"]) - clip_source_start)) (max(0.0, float(s["start"]) - clip_source_start), min(clip_duration, float(s["end"]) - clip_source_start))
for s in spans for s in span_list
if float(s["end"]) > clip_source_start and float(s["start"]) < window_end if float(s["end"]) > clip_source_start and float(s["start"]) < window_end
] ]
# Clip-relative windows, for deciding which plain titles to hide —
# same coordinate space add_text_title's offsets end up in. Dynamic
# spans hide plain (see the trade-off note below); explicit
# `subtitle_common: false` spans hide it too, even without a dynamic
# line covering the gap.
clip_spans = _clip_relative(spans)
clip_hide_plain_spans = clip_spans + _clip_relative(exclude_spans)
all_words = data.get("words", []) all_words = data.get("words", [])
dynamic_lines = 0 dynamic_lines = 0
@@ -656,7 +733,9 @@ async def handle_generate_subtitles_by_emphasis(arguments: dict) -> Sequence[Tex
# entry 2026-08-17). # entry 2026-08-17).
dynamic_lines = len( dynamic_lines = len(
modifier.generate_dynamic_subtitles( modifier.generate_dynamic_subtitles(
el, clip_emphasis_words, dynamic_config, segments=clip_segments el, clip_emphasis_words, configs=dynamic_configs, segments=clip_segments,
role=single_dynamic_role,
compound_subphrases=True,
) )
) )
dynamic_word_count = len(clip_emphasis_words) dynamic_word_count = len(clip_emphasis_words)
@@ -683,7 +762,7 @@ async def handle_generate_subtitles_by_emphasis(arguments: dict) -> Sequence[Tex
continue continue
start = max(0.0, min(float(w.get("start", 0.0)) for w in block)) start = max(0.0, min(float(w.get("start", 0.0)) for w in block))
end = max(float(w.get("end", start)) for w in block) end = max(float(w.get("end", start)) for w in block)
if _overlaps_any_span(start, end, clip_spans): if _overlaps_any_span(start, end, clip_hide_plain_spans):
plain_hidden += 1 plain_hidden += 1
continue continue
duration = max(end - start, modifier.frame_duration_fraction()) duration = max(end - start, modifier.frame_duration_fraction())
@@ -701,6 +780,7 @@ async def handle_generate_subtitles_by_emphasis(arguments: dict) -> Sequence[Tex
face=None, face=None,
font_scale=1.0, font_scale=1.0,
size_param=plain_font_size, size_param=plain_font_size,
role=saved_plain["role"],
) )
plain_created += 1 plain_created += 1