feat(legendas): sub-frases por vírgula e empacotamento em compound clip
split_into_subphrases divide a frase na vírgula — onde a fala respira —
mas funde de volta o pedaço curto ("né?", "Então..."), que lê como parte
da frase anterior e não como bloco próprio.
wrap_titles_in_compound empacota os títulos de uma sub-frase num compound
clip, replicando a estrutura que o próprio Final Cut produz: o primeiro
título vira âncora do spine em offset 0, os demais penduram nele por lane,
e um ref-clip toma o lugar deles na lane original. Os offsets dos filhos
são rebaseados para o espaço de tempo da âncora, senão cada palavra
escorregaria pela diferença entre os dois start.
Junto: _filter_children_for_segment passa a filtrar também o <video> do
Clipe de Ajuste. Sem isso, cada corte subsequente duplicava o zoom em
todos os pedaços resultantes com o offset original intacto, e as cópias
desenhavam empilhadas na mesma posição da timeline.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
co-authored by
Claude Opus 5
parent
635d1bb553
commit
688bdeddb6
@@ -316,6 +316,54 @@ def group_words_by_segment(
|
||||
return groups
|
||||
|
||||
|
||||
def split_into_subphrases(
|
||||
words: Sequence[dict],
|
||||
min_words: int = 3,
|
||||
) -> List[List[dict]]:
|
||||
"""Split a sentence's *words* into sub-phrases at comma boundaries.
|
||||
|
||||
A comma is where a spoken sentence actually breathes, so it is the
|
||||
natural seam for grouping subtitles — each sub-phrase becoming its own
|
||||
on-screen block (and, downstream, its own compound clip).
|
||||
|
||||
The exception is the short tail: a fragment like "né?" or "Então..."
|
||||
reads as part of the phrase before it, not as a phrase of its own, and
|
||||
promoting it to its own block would flash a single word on screen. So a
|
||||
piece shorter than *min_words* is merged back into its neighbour —
|
||||
preferring the previous piece, falling back to the next one when the
|
||||
short piece leads the sentence.
|
||||
|
||||
Returns one group per sub-phrase; a sentence with no comma comes back
|
||||
as a single group.
|
||||
"""
|
||||
pieces: List[List[dict]] = []
|
||||
current: List[dict] = []
|
||||
for w in words:
|
||||
current.append(w)
|
||||
text = str(w.get('word') or w.get('text') or '')
|
||||
if text.rstrip().endswith(','):
|
||||
pieces.append(current)
|
||||
current = []
|
||||
if current:
|
||||
pieces.append(current)
|
||||
|
||||
if len(pieces) <= 1:
|
||||
return pieces
|
||||
|
||||
merged: List[List[dict]] = []
|
||||
for piece in pieces:
|
||||
if len(piece) < min_words and merged:
|
||||
merged[-1].extend(piece)
|
||||
else:
|
||||
merged.append(piece)
|
||||
# A short leading piece has no previous neighbour to fold into, so it
|
||||
# folds forward instead.
|
||||
if len(merged) > 1 and len(merged[0]) < min_words:
|
||||
merged[1][:0] = merged[0]
|
||||
merged.pop(0)
|
||||
return merged
|
||||
|
||||
|
||||
def segments_to_srt(segments: Sequence[dict]) -> str:
|
||||
"""Render transcript segments as an SRT string (for captions import)."""
|
||||
|
||||
|
||||
Reference in New Issue
Block a user