chore: atualização geral

This commit is contained in:
João Henrique
2026-08-19 16:35:29 -04:00
parent 8fca456ceb
commit e7748c2c58
66 changed files with 13037 additions and 4237 deletions
+93
View File
@@ -0,0 +1,93 @@
---
name: editar-por-voz
description: Edita um vídeo a partir da análise de voz — lê a timeline de voz (JSON) de uma gravação, separa o roteiro da conversa de bastidor, escolhe a melhor tomada de cada frase, decide cortes/zooms/textos e devolve a lista de decisões em JSON. Use quando o usuário pedir para editar por voz, montar um corte automático, limpar tomadas repetidas, escolher as melhores tomadas, ou marcar os momentos de ênfase de uma fala.
---
# Editar por voz
Você recebe um JSON com o que foi dito, por quem e **como**; devolve um JSON
com **o que fazer**. Quem aplica é o programa — você nunca escreve XML.
## Princípio
O sistema já mede *como* a pessoa falou, de forma reprodutível. Não
recalcule nada disso nem reestime tempos "no olho".
Seu trabalho é o que nenhum limiar resolve: **o que aquilo significa.** O
índice diz que uma palavra foi dita com força; só você sabe se ela é o
argumento central ou uma piada com o câmera.
## Entrada e saída
**Entrada:** `<mídia>_voice_timeline.json` (de `build_voice_timeline`). Se
não existir, rode a ferramenta; se existir, leia direto — a análise leva
minutos.
**Saída:** JSON com a lista de ações → `criterios/08-formato-de-saida.md`
`apply_voice_actions` aplica direto, mas é para **teste**. O produto do seu
trabalho é a lista de decisões.
## Ordem de trabalho
Siga nesta ordem. Pular a Fase 2 ou a 3 leva a decisões erradas.
| Fase | O que fazer | Critérios |
|---|---|---|
| **0** | Ler `layers` — saber o que rodou | `criterios/09-analise-incompleta.md` |
| **1** | Ler o JSON em camadas | `criterios/01-leitura-do-json.md` |
| **2** | Separar roteiro de conversa de bastidor | `criterios/02-triagem-roteiro-vs-conversa.md` |
| **3** | Escolher a melhor tomada de cada frase | `criterios/03-escolha-da-melhor-tomada.md` |
| **4** | **Reanalisar** o material que sobrou | `criterios/04-reanalise-do-material-restante.md` |
| **5** | Decidir zooms | `criterios/05-zoom.md` |
| **6** | Decidir textos, cortes e marcadores | `criterios/06-texto-corte-marcador.md` |
| **7** | Cortar a lista pelo ritmo | `criterios/07-ritmo.md` |
| **8** | Montar o JSON de saída | `criterios/08-formato-de-saida.md` |
## As três armadilhas
Cada uma já causou erro silencioso em material real:
1. **A conversa de bastidor tem a maior ênfase do vídeo.** O índice acústico
favorece a fala solta sobre o texto decorado. Separar é tarefa de texto,
nunca de limiar — e nem de diarização, já que costuma ser a mesma pessoa.
2. **Ênfase é relativa ao conjunto analisado.** Depois de cortar, os números
da análise bruta apontam para as palavras erradas. Sempre reanalise
(Fase 4) antes de escolher zooms.
3. **Tempos sempre na mídia original.** Nunca compense para "depois do
corte" — o programa faz isso sozinho, e compensar por conta própria joga
todo destaque no frame errado, sem erro visível.
## Ordem no sistema
Sua edição roda **primeiro, na timeline intacta**. Os posicionamentos (zoom,
texto, marcador) são deslocados a partir da *sua* lista de cortes; se outra
ferramenta já tiver rippado a timeline antes, esse deslocamento não sabe
disso e o efeito cai no frame errado — sem erro visível.
```
build_voice_timeline → [você decide] → refine_voice_timeline → [você corta
pelo ritmo] → apply_voice_actions → remove_media_silence →
generate_dynamic_subtitles
```
Vícios de linguagem e lacunas longas entram na **sua** lista, num ripple só
(`06-texto-corte-marcador.md`). Silêncio fino e legendas vêm depois.
`generate_dynamic_subtitles` nunca é o passo final por si só — sempre
seguido de `validate_subtitle_layout` antes de dar a legenda como pronta.
Detalhe do porquê e das regras específicas dessa etapa:
`code/Engine/docs/03_SERVER_TOOLS.md`, seção "Legendas dinâmicas".
## Ao relatar
Sempre em **português**, com o raciocínio e não só o resultado:
- quantas tomadas encontrou de cada frase e **qual escolheu, com o motivo**;
- o que descartou como bastidor;
- por que cada zoom caiu onde caiu (palavra + ênfase);
- **o que foi descartado ou rejeitado** — nunca relate só os acertos;
- o que você **não** conseguiu decidir. Em dúvida entre duas tomadas,
marque as duas e deixe para o editor.
@@ -0,0 +1,71 @@
# 01 — Leitura do JSON
O arquivo `<mídia>_voice_timeline.json` é a entrada de todo o trabalho.
Leia em camadas, de cima para baixo, e só desça quando precisar.
## Camadas
| Camada | O que traz | Para quê |
|---|---|---|
| `layers` | o que de fato rodou na análise | **leia primeiro** — ver `07-analise-incompleta.md` |
| `summary` | forma da peça, `peak_moments`, contagens | visão geral em poucos números |
| `speakers` | quem fala, % do tempo, frases de exemplo | identificar papéis |
| `segments` | cada fala com seus agregados | **onde você mais trabalha** |
| `segments[].words` | detalhe por palavra | achar o instante exato de um destaque |
| `scales` | o que cada número significa | documentação dentro do próprio arquivo |
## Campos que decidem quase tudo
**`gap_before`** — silêncio antes da fala, em segundos. É o mapa estrutural
da gravação: acima de ~3s (`take_boundary: true`) a câmera parou ou a
tomada recomeçou. Num material real de 3min17s isso identificou 6
fronteiras, todas exatamente onde a pessoa recomeçava o roteiro.
**`take_boundary`** — booleano derivado do `gap_before`. Use para agrupar
tomadas.
**`emphasis`** (0–1) — índice combinado de energia, variação de tom,
variação de ritmo, pausa anterior e duração. **É relativo ao material
analisado**, nunca uma medida absoluta. Ver `02-enfase-e-reanalise.md`.
**`energy`** (0–1) — intensidade relativa ao trecho mais alto da gravação.
**`pitch_delta`** (0–1) — quanto o tom se afasta da média do falante.
**`peak_emphasis`** e **`avg_energy`** (por segmento) — permitem julgar uma
frase inteira sem ler palavra por palavra. É por aqui que você avalia o
arco narrativo.
**`energy_raw`** e **`pitch_hz`** — valores brutos, sem normalização. Não
use para decidir; existem para permitir a reanálise da Fase 2.
## O que NÃO fazer
- **Não recalcule** energia, tom ou ênfase. O sistema mede melhor e de
forma reprodutível.
- **Não reestime tempos "no olho".** Use os timestamps do JSON.
- **Não trate `emphasis` como valor absoluto.** Um 0,35 pode ser o pico de
uma gravação e ruído em outra.
## O timestamp por palavra tem um viés conhecido
O início de cada palavra vem sistematicamente **adiantado em ~0,3-0,5s** em
relação ao ataque real da fala — medido em material real com ffmpeg
(`astats`), consistente em 6 pontos do mesmo vídeo. O fim da palavra não
tem esse problema (erro de poucos centésimos). Causa: `word_timestamps` do
faster-whisper deriva por atenção cruzada, sem alinhamento forçado — ver
`05_EXPERIENCIAS.md`, entrada de 2026-08-19.
Isso não é "reestimar no olho" — é um bug de medição na fonte, não um
julgamento seu. Na prática:
- Ao posicionar um `zoom` cujo `start` precisa cair exatamente na palavra
(não uma frase inteira), some **+0,3 a +0,4s** ao timestamp do JSON antes
de decidir, ou confira com `ffmpeg -af astats` se a precisão importar
para o frame.
- **Não aplique essa correção a `gap_before` para decidir corte** — a régua
de silêncio (`06-texto-corte-marcador.md`) já é conservadora o bastante
para absorver esse erro; corrigir os dois ao mesmo tempo é redundante.
- Se um dia o pipeline ganhar alinhamento forçado (WhisperX), este aviso
perde a razão de existir — confira se `layers` ou a versão do documento
já indicam isso antes de aplicar o offset manualmente.
@@ -0,0 +1,51 @@
# 02 — Triagem: roteiro vs. conversa de bastidor
**Primeira coisa a fazer, antes de qualquer decisão de efeito.**
Material bruto de gravação quase nunca é uma tomada só. A pessoa lê o
roteiro, erra, conversa com a equipe e recomeça.
## Por que isso é tarefa sua, e não do sistema
No áudio essa separação é **invisível** — e pior: o índice de ênfase
*favorece* a conversa, que é mais solta e mais alta que o texto decorado.
Caso real: a fala mais enfática de um vídeo inteiro (energia **1,00**, o
topo absoluto da gravação) era *"Amor, eu tô intacto!"*, dita para o marido
fora de quadro. Três das sete palavras de maior ênfase do vídeo vinham
dessa única frase de bastidor.
Nenhum limiar acústico separa isso. O **texto** separa sem erro.
> Atenção: isso também **não é diarização**. Num caso real, a pessoa da
> equipe estava fora do microfone — a diarização a ouvia, mas o Whisper não
> a transcrevia. As falas a descartar eram da **própria protagonista**:
> mesma voz, contexto diferente. "Quem fala" e "isso é tomada válida" são
> perguntas diferentes.
## Descartar — conversa com a equipe
Reconhece-se pelo **conteúdo**:
- **vocativo para alguém da sala** — *"Amor, eu tô intacto!"*
- **pergunta operacional** — *"Posso começar da mastopexia?"*,
*"E aí, continua?"*, *"Mas eu vou ter que falar tudo de novo?"*
- **instrução técnica** — *"Só clica aí agora na tela."*, *"Aumenta."*
- **comentário sobre a própria gravação** — *"Vou falar só a última frase,
só um pouquinho, não pegou?"*
## Descartar — frases interrompidas
Texto que morre no meio, tipicamente em reticências ou emendando numa
pergunta:
- *"E tudo isso associado à medida..."*
- *"Aquela mama com um formato mais estruturado, com o colo que..."*
- *"de pele..."*
## Sinais estruturais que ajudam
Use `take_boundary` para achar onde cada tomada recomeça. Num material
real, as fronteiras (gaps de 3,6s a 19,8s) caíam exatamente nos pontos onde
a médica reiniciava o roteiro — inclusive nas duas retomadas da frase de
abertura.
@@ -0,0 +1,60 @@
# 03 — Escolha da melhor tomada
A mesma frase costuma aparecer 2, 3, 4 vezes. Seu trabalho é ficar com
**uma**.
## Como agrupar
1. Use `take_boundary` para localizar onde cada tomada recomeça.
2. Agrupe as repetições **pelo texto**, não pelo tempo — a mesma frase
reaparece em pontos distantes da gravação. Num caso real, a abertura
*"Aquela mama com um formato mais estruturado"* apareceu aos 2,0s, 64,9s
e 86,4s.
## Critérios, nesta ordem
### 1. Completa
Não morre no meio, não emenda numa pergunta. Uma tomada incompleta está
descartada por definição, mesmo que a dicção seja ótima.
### 2. Dicção limpa
Sem tropeço, sem repetição de palavra, sem vício de linguagem. **Compare os
textos lado a lado:**
| Tomada 1 | Tomada 3 | Escolha |
|---|---|---|
| *"isso é desejo de muitas mulheres"* | *"**aí** isso é desejo de muitas mulheres"* | Tomada 1 |
### 3. Formulação melhor
Quando as duas estão limpas, prefira a mais direta — normalmente a última,
porque é onde a pessoa já se ajustou:
| Antes | Depois | Escolha |
|---|---|---|
| *"a gente **faz a inserção de** próteses"* | *"a gente **insere** próteses"* | a segunda |
| *"reestrutura a mama"* | *"reestrutura a **sua** mama"* | a segunda |
### 4. Entrega
**Só então** desempate por `avg_energy` / `peak_emphasis`.
Quando duas tomadas têm texto **idêntico palavra por palavra**, aí a
energia decide sozinha — é o único sinal disponível. Caso real: o fecho
tinha duas tomadas iguais, energia **0,38** e **0,19**. A de 0,38 é a boa,
e o texto sozinho jamais diria isso.
## Regra de ouro
A **última** tomada costuma ser a melhor — é onde a pessoa acertou. Mas
**confirme lendo o texto**; nunca assuma.
## Quando estiver em dúvida
Não decida no escuro. Coloque um `marker` nas duas candidatas, explique a
dúvida no `reason`, e deixe a escolha para o editor humano.
## Continuidade
Ao montar o corte final você pode misturar blocos de tomadas diferentes —
abertura da tomada 1, corpo da tomada 3. Isso é normal. Mas **avise nas
emendas**: coloque um `marker` em cada junção para o editor conferir se o
enquadramento e a posição da pessoa combinam.
@@ -0,0 +1,70 @@
# 04 — Reanálise do material que sobrou
**Não escolha zooms com os números da análise bruta.**
## O problema
Ênfase e energia são **relativas ao conjunto analisado**. Energia é
normalizada contra o momento mais alto da gravação; ênfase deriva dela.
Se esse momento mais alto foi cortado — uma piada, um grito, uma conversa
de bastidor — tudo que sobrou continua pontuado contra uma referência que o
espectador **nunca verá**. As notas do corte final ficam artificialmente
comprimidas, e o ranking aponta para as palavras erradas.
Caso real: o pico do vídeo era *"Amor, eu tô intacto!"* (energia 1,00),
descartado na triagem. Todo o material restante estava sendo medido contra
ele.
## A solução
Depois de definir os cortes, renormalize sobre os sobreviventes. Chame a
ferramenta **`refine_voice_timeline`**, passando a mídia e a lista de
cortes que você já decidiu:
```
refine_voice_timeline(media_path, cuts=[{start, end}, ...], min_gap=8.0)
```
Ela devolve, numa chamada só, a comparação bruto × sobreviventes, os picos
re-ranqueados e os candidatos a zoom. É barata: renormaliza os números já
medidos, sem reabrir o áudio.
Efeito medido no mesmo material:
| | Bruto | Só o que sobrou |
|---|---|---|
| Ênfase média | 0,179 | **0,197** |
| *"Aquela"* | 0,39 | **0,42** |
| *"mastopexia"* | 0,26 | **0,34** |
| *"devolver"* | — | **0,35** |
*"mastopexia"* só virou candidata legítima depois da reanálise.
## As janelas que ela propõe
A seção **Zoom Candidates** da resposta já vem com três coisas resolvidas:
1. **Pega a palavra de conteúdo mais enfática de cada frase.** Artigos e
conectivos são filtrados — um *"a"* falado alto continua sendo um artigo.
Sem esse filtro, o ranking bruto apontava para "o", "a", "eu": picos de
*entrega*, não de *sentido*.
2. **Estende a janela até o fim da frase**, não do segmento (ver
`05-zoom.md`).
3. **Mantém distância mínima** entre zooms.
São **candidatos, não obrigações.** Corte a lista pelo ritmo
(`07-ritmo.md`). Os tempos continuam na mídia original — vão direto para
`apply_voice_actions`.
`max_zooms` limita a lista, mas prefira cortá-la você mesmo: o corte por
ritmo é decisão editorial, não um teto numérico.
## Princípio geral
> "Qual o momento mais forte da **gravação**?" e "qual o momento mais forte
> do **vídeo final**?" são perguntas diferentes sempre que a métrica for
> relativa.
Toda métrica normalizada precisa ser recalculada quando o conjunto muda —
senão ela responde a pergunta errada, silenciosamente.
@@ -0,0 +1,71 @@
# 05 — Zoom (punch-in)
## Quando usar
No momento em que o argumento vira. Um pico acústico só merece zoom se for
também um pico **de sentido**.
Palavra gritada sem peso narrativo não ganha nada — e isso inclui os picos
que caem em artigos e conectivos, que são picos de entrega, não de conteúdo.
## A janela
**`start`** — na palavra de ênfase.
**`end`** — no **fim da frase**. A frase inteira, não o fim do segmento da
transcrição.
O Whisper corta frases no meio, por respiração e não por gramática:
> *"Aquela mama com um formato mais estruturado, que valoriza o seu colo,
> que dá aquele ar"* **|** *"de elegância, isso é desejo de muitas mulheres,
> né?"*
Soltar o zoom no fim do primeiro segmento libera **no meio do pensamento** —
é o que faz um punch-in parecer arbitrário. `suggest_zoom_windows` já
estende até a pontuação final (`.` `!` `?` `…`), e nunca atravessa uma
fronteira de tomada.
## A forma — o programa decide sozinho
Você escolhe `start` e `end`; a forma sai da posição da janela dentro do
trecho:
| Situação | Comportamento | Por quê |
|---|---|---|
| Começa a **>0,5s** do início do trecho | entrada rápida (~0,25s) | o movimento chega junto com a palavra |
| Começa a **≤0,5s** do início | **entra já ampliado, sem transição** | o corte já foi a transição; uma rampa ali lê como a imagem se acomodando |
| Frase termina no meio do trecho | **saída seca**, 1 frame | volta ao enquadramento sem chamar atenção |
| Frase termina a **≤1s** do corte | **não volta** — segura até o corte | o próximo trecho já abre no enquadramento dele; voltar antes é movimento desperdiçado |
Os limiares são diferentes de propósito: no fim o corte esconde um retorno
inacabado, mas no início a rampa é visível desde o primeiro frame.
Para forçar manualmente, existem `start_at_peak` e `hold_at_end` — mas o
automático acerta na quase totalidade dos casos.
## Escala
| Valor | Uso |
|---|---|
| 1,15 | sutil |
| 1,18 – 1,3 | padrão |
| 1,5 | forte |
Em vídeo institucional, fique na faixa baixa. Acima de 3,0 é rejeitado.
O zoom é **relativo ao enquadramento existente**: se o clipe já tem escala
1,77 (material gravado de lado e reenquadrado), um zoom 1,18 anima de 1,77
para 2,09 e preserva rotação e posição.
## Dois zooms no mesmo clipe
Depois do corte, dois picos que você escolheu podem cair no **mesmo**
trecho sobrevivente (nenhum corte os separou em clipes distintos) — é
comum quando a corrida limpa de uma tomada é longa. O sistema resolve isso
sozinho, e a regra é a mesma que rege o resto: janelas **distantes**
empilham (os dois zooms convivem, cada um voltando ao enquadramento real
entre um e outro); janelas que **se sobrepõem** substituem (é o mesmo
evento sendo reajustado, não dois). Você não precisa calcular isso na
hora de decidir — só respeitar o `min_gap` de `07-ritmo.md`, que já
garante que dois zooms escolhidos por você nunca se sobrepõem.
@@ -0,0 +1,88 @@
# 06 — Texto, corte e marcador
## Texto
Para fixar um **conceito, número ou nome** que o espectador precisa reter.
- Use a palavra **dita**, não uma paráfrase.
- Curta, em caixa alta. Até 120 caracteres (é truncado além disso).
- Uma por frase, no máximo.
**Não legende a frase inteira.** Para isso existe
`generate_dynamic_subtitles`, que é outra ferramenta e outro propósito.
Boas candidatas são as palavras-chave que sobram depois de filtrar as
funcionais — num caso real: *mastopexia*, *flacidez*, *próteses*,
*devolver*, *desejo*.
## Corte
Digressão, repetição, frase abandonada, conversa de bastidor, tomada pior —
e mais duas coisas que **são** seu trabalho, ao contrário do que parece.
### Vícios de linguagem entram na sua lista
Não delegue para `remove_filler_words`. Você já está percorrendo palavra por
palavra na triagem; marcar as muletas é uma linha a mais, sem custo. E você
tem o que a lista fixa não tem: **contexto**.
Um *"tipo"* em *"tipo assim, sabe"* é muleta. Em *"esse tipo de cirurgia"*
é a palavra principal. Um *"não não não"* pode ser gagueira ou ênfase. A
lista fixa não distingue; você distingue.
### Lacunas longas entram na sua lista — curtas, nunca
**O tamanho da lacuna muda o que ela é.** A régua está medida em material
real (`pause_weight()` em `emphasis.py`, `TAKE_BOUNDARY_GAP` em
`voice_timeline.py`):
| `pause_before` | O que é | O que fazer |
|---|---|---|
| até ~1,5s | o falante montando a frase — **isso É a ênfase** | **nunca cortar** |
| 1,5–3s | zona cinza | julgue pela frase |
| acima de 3s | troca de tomada, ar morto, outra pessoa falando | **cortar** |
Cortar a pausa curta é o erro grave: ela é uma das cinco entradas do índice
de ênfase, então você estaria apagando justamente a batida que faz a palavra
seguinte pontuar alto. Uma frase fluida não se aperta.
Acima de 3s a pausa deixa de contar como ênfase por construção — medido em
material real, lacunas de 6–9s rankeavam como os momentos mais enfáticos da
gravação só porque a escala saturava.
### O que continua NÃO sendo seu trabalho
| Tarefa | Ferramenta | Por quê |
|---|---|---|
| Apertar o ar **dentro** da fala | `remove_media_silence` | Lê o áudio real com ffmpeg; você só tem os intervalos entre palavras transcritas |
E cuidado: **ausência de fala não é ausência de som.** Respiração, riso,
suspiro, a reação depois da frase — nada disso vira palavra, então aparece
como lacuna, e às vezes é o melhor frame do vídeo. Lacuna longa é candidata
a corte, não corte automático.
`remove_media_silence` roda **depois** de `apply_voice_actions`, como
acabamento opcional, sobre o material que sobrou.
**Antes de rodar `remove_media_silence` sobre o corte final, sempre rode a
detecção primeiro** (sem aplicar) e leia os spans um a um contra a régua
acima. O detector corta por limiar de dB — ele não sabe distinguir "batida
de 0,8s entre duas frases", que a régua protege, de "ar morto de emenda",
que deveria ser apertado. Aplicar direto, sem essa checagem, é o mesmo erro
de cortar pausa curta, só que por outra ferramenta.
## Marcador
Quando você quer **sinalizar para o editor humano decidir**, em vez de
decidir por ele.
Use em:
- **Emendas entre tomadas** — sempre. O editor precisa conferir se o
enquadramento e a posição da pessoa combinam na junção.
- **Dúvida entre duas tomadas** — marque as duas, explique no `reason`.
- **Momentos que talvez mereçam efeito** mas que você não tem confiança
para decidir.
Marcador é um **ponto**, não um trecho: sobrevive mesmo encostado na borda
de um corte, o que é justamente o caso das emendas.
@@ -0,0 +1,37 @@
# 07 — Ritmo
**O erro mais comum é efeito demais.** Cansa mais que efeito de menos, e
denuncia edição automática.
## Limites
| Regra | Valor |
|---|---|
| Distância mínima entre dois zooms | **8–10 segundos** |
| Zooms por minuto de vídeo | **2 a 4** (teto) |
| Zoom + texto no mesmo instante | só com motivo claro |
Se dois picos estiverem colados, **escolha o mais forte e abra mão do
outro**. Não tente encaixar os dois.
## Candidatos ≠ obrigações
`suggest_zoom_windows` devolve uma lista de candidatos. Normalmente você usa
uma **fração** dela.
Caso real: num corte de 47,6s a ferramenta sugeriu **5** janelas. O certo
foram **3** — 5 violaria o teto de 2–4 por minuto. Ficaram a abertura, o
termo central e o fecho; as duas descartadas eram frases de apoio.
O mesmo vale para `peak_moments` no `summary`: é lista de candidatos.
## Como escolher quais manter
Quando precisar cortar a lista, priorize por **função narrativa**, não por
nota:
1. **A abertura** — prende o espectador.
2. **O conceito central** — o termo que o vídeo existe para explicar.
3. **O fecho** — a frase que fica.
Só depois disso, as frases de apoio, por ordem de ênfase.
@@ -0,0 +1,65 @@
# 08 — Formato de saída
O produto do seu trabalho é **este JSON**. É ele que vai para o programa
gerar o FCPXML. Você nunca escreve XML.
## Estrutura
```json
{
"source": "0E6A8290.mp4",
"actions": [
{"kind": "cut", "start": 21.9, "end": 127.6,
"reason": "tomadas descartadas, frases interrompidas e conversa com a equipe"},
{"kind": "zoom", "start": 2.0, "end": 10.7,
"params": {"scale": 1.15}, "reason": "abertura: \"Aquela mama\" (ênfase 0.42)"},
{"kind": "text", "start": 127.7, "end": 129.0,
"params": {"content": "MASTOPEXIA"}, "reason": "fixa o termo central"},
{"kind": "marker", "start": 21.85, "end": 22.0,
"reason": "EMENDA 1 — conferir junção entre tomadas"}
]
}
```
## Regras
### 1. Tempos em segundos da mídia ORIGINAL
Exatamente como aparecem no `voice_timeline.json`.
**Nunca compense para "depois do corte".** O programa faz esse deslocamento
sozinho: ele resolve os cortes primeiro e reposiciona todo o resto. Se você
compensar por conta própria, **todo destaque cai no frame errado** — e o
erro é silencioso.
### 2. `end` sempre maior que `start`
Ambos ≥ 0. Um `end <= start` é rejeitado.
### 3. Tipos
`cut` · `zoom` · `text` · `marker`
### 4. Parâmetros por tipo
| Tipo | `params` |
|---|---|
| `cut` | nenhum |
| `zoom` | `scale` entre 1.0 e 3.0 (padrão 1.3 se omitido) |
| `text` | `content` **obrigatório**, até 120 caracteres |
| `marker` | opcional: `content` vira o nome do marcador |
### 5. `reason` — sempre preencha
É o que o usuário lê para revisar sua decisão, e o que te obriga a **ter**
uma. Um `reason` vazio é sinal de decisão sem critério.
Inclua o dado que embasou: *"abertura: 'Aquela mama' (ênfase 0.42)"* é útil;
*"zoom"* não é.
## Como o programa trata erros
- **Ação inválida** → rejeitada e reportada **individualmente**. Uma linha
malformada nunca derruba as outras.
- **Ação apontando para material cortado** → descartada e reportada, nunca
deslizada para o conteúdo vizinho.
- **Ação fora da mídia** → reportada como não colocada.
Você recebe o relatório dos três casos. **Repasse ao usuário** — nunca
relate só os acertos.
@@ -0,0 +1,48 @@
# 09 — Quando a análise veio incompleta
O bloco `layers` no topo do JSON diz **o que de fato rodou**. Leia antes de
qualquer outra coisa.
```json
"layers": {"transcript": true, "acoustics": false, "speakers": false}
```
## Por que esse bloco existe
Fala monótona e acústica que não carregou deixam **os mesmos zeros** nos
dados. Sem o `layers`, é impossível distinguir "esta pessoa fala de forma
uniforme" de "a análise acústica falhou".
## Os casos
### `acoustics: false`
Todos os valores acústicos são 0. **Você não tem ênfase real.**
- Decida só pelo texto.
- **Avise o usuário** explicitamente.
- Prefira `marker` a `zoom` — sinalize em vez de decidir.
Causa comum: o componente librosa não está instalado, ou o ffmpeg não
conseguiu extrair o áudio do container.
### `speakers: false` num vídeo com várias pessoas
A diarização não rodou — falta o token do HuggingFace (aba Modelos do app).
- Avise antes de tratar tudo como uma voz só.
- Lembre que isso **não impede** a triagem roteiro/conversa, que é feita
pelo texto (ver `02-triagem-roteiro-vs-conversa.md`).
### `peak_count: 0`
Nada cruzou o piso de ênfase. Duas causas possíveis:
1. A fala é uniforme mesmo — material sem picos.
2. O limiar está alto para esse material.
Sugira ajustar em **Análise de Voz** no app. **Não force destaques
inexistentes** só para entregar alguma coisa.
## Regra geral
Não finja precisão que você não tem. Uma edição entregue com a ressalva
certa é útil; uma entregue como se estivesse completa, quando metade dos
dados faltou, custa a confiança do usuário no sistema inteiro.
+2 -2
View File
@@ -9,7 +9,7 @@ normalmente; a regra é sobre a comunicação com o usuário.
## What This Is
MCP server that reads/writes Final Cut Pro XML (FCPXML) files. 62 tools for timeline analysis, batch editing, QC, generation, multi-track support, media relink, NLE export, transcript-based editing (local Whisper), and LIVE FCP control (push_to_fcp / list_fcp_libraries via Apple events). Reads FCPXML 1.8–1.14 (incl. `.fcpxmld` bundles with sidecar preservation), writes 1.13 by default. Dual-mode (XML + Live) direction: `code/docs/CAPABILITY-AUDIT-2026-06.md`.
MCP server that reads/writes Final Cut Pro XML (FCPXML) files. 73 tools for timeline analysis, batch editing, QC, generation, multi-track support, media relink, NLE export, transcript-based editing (local Whisper), and LIVE FCP control (push_to_fcp / list_fcp_libraries via Apple events). Reads FCPXML 1.8–1.14 (incl. `.fcpxmld` bundles with sidecar preservation), writes 1.13 by default. Dual-mode (XML + Live) direction: `code/docs/CAPABILITY-AUDIT-2026-06.md`.
## Architecture
@@ -88,7 +88,7 @@ CI runs both on every push to main. If either fails, the commit gets an X on Git
## Testing
1032 tests across 24 files. `test_models.py` covers TimeValue arithmetic, Timecode parsing/formatting, Clip properties, validation models, and Timeline helpers. `test_writer.py` covers insert_clip, add_marker (all types), trim_clip, delete_clip, split_clip, and change_speed operations. `test_server.py` covers MCP tool handlers, parsers, and dispatch. `test_rough_cut.py` covers RoughCutGenerator. `test_features_v05.py` covers connected clips, roles, timeline diff, reformat, silence detection, export, and backward compatibility. `test_marker_pipeline.py` covers build_marker_element shared builder, batch auto-modes, clip index duplicate-name behavior, and write_fcpxml output format. `test_refactored_helpers.py` covers _index_elements, _iter_spine_clips, _find_spine_clip_at_seconds, _resolve_clip_duration, _make_asset_clip, _format_batch_result, and serialize_xml edge cases. `test_transcribe.py` covers phrase/filler span matching, range merge/invert algebra, whisper graceful degradation, and transcript-driven handler cuts against cached transcripts. `test_media_intel.py` covers silencedetect stderr parsing, source-to-timeline mapping, parameter bounds, and real-WAV ffmpeg integration (skips without ffmpeg; CI installs it). Tests use `examples/sample.fcpxml` as fixture data and inline XML fixtures. Tests create temp files and clean up after.
1342 tests across 34 files. `test_models.py` covers TimeValue arithmetic, Timecode parsing/formatting, Clip properties, validation models, and Timeline helpers. `test_writer.py` covers insert_clip, add_marker (all types), trim_clip, delete_clip, split_clip, and change_speed operations. `test_server.py` covers MCP tool handlers, parsers, and dispatch. `test_rough_cut.py` covers RoughCutGenerator. `test_features_v05.py` covers connected clips, roles, timeline diff, reformat, silence detection, export, and backward compatibility. `test_marker_pipeline.py` covers build_marker_element shared builder, batch auto-modes, clip index duplicate-name behavior, and write_fcpxml output format. `test_refactored_helpers.py` covers _index_elements, _iter_spine_clips, _find_spine_clip_at_seconds, _resolve_clip_duration, _make_asset_clip, _format_batch_result, and serialize_xml edge cases. `test_transcribe.py` covers phrase/filler span matching, range merge/invert algebra, whisper graceful degradation, and transcript-driven handler cuts against cached transcripts. `test_media_intel.py` covers silencedetect stderr parsing, source-to-timeline mapping, parameter bounds, and real-WAV ffmpeg integration (skips without ffmpeg; CI installs it). Tests use `examples/sample.fcpxml` as fixture data and inline XML fixtures. Tests create temp files and clean up after.
## FCPXML Gotchas
+258 -3
View File
@@ -41,6 +41,34 @@ Commands:
clips, cuts, connected, markers}]}
or {"ok": false, "error": "..."}
analyze_voice {"path": "...", "output_dir": "...", "model": "...",
"language": "pt"|"auto"|null, "hf_token": "..."|null,
"num_speakers": ""|null}
Build the voice timeline (transcript+diarization+acoustics) for
every unique source media — analysis only, writes _voice_timeline.json
next to each media, `path` passes through unchanged. Meant as one
entry in the batch operations list (see processBatchStep), so
`refine_voice_timeline` never has to reopen the audio later.
-> {"ok": true, "path": "...", "message": "..."} or {"ok": false, "error": "..."}
dynamic_subtitle_config {}
-> {"ok": true, "band_height", "block_center_y", "line_gap", "font",
"font_size", "emphasis_font", "emphasis_face", "emphasis_size",
"active_color", "emphasis_color", "text_scale"}
set_dynamic_subtitle_config {<any of the fields above>}
Persists only the given fields to ~/.fcp-mcp-server/config.json.
generate_dynamic_subtitles reads this as its own fallback default.
-> {"ok": true, <same shape as dynamic_subtitle_config>}
silence_config {}
-> {"ok": true, "noise_db": -30.0, "min_silence": 0.5, "padding": 0.05}
set_silence_config {"noise_db": -30.0, "min_silence": 0.5, "padding": 0.05}
Persists only the given fields. detect_media_silence and
remove_media_silence read this as their own fallback default.
-> {"ok": true, <same shape as silence_config>}
transcribe {"path": "...", "model": "small", "language": "pt"|null,
"hf_token": "..."|null, "num_speakers": ""|null}
-> JSON-lines:
@@ -76,7 +104,9 @@ Commands:
generate_dynamic_subtitles {"path": "...", "clip_name": "..."|null,
"band_height": 0.22, "block_center_y": -167,
"font": "Helvetica Neue", "font_size": 128,
"active_color": "1 1 1 1", "inactive_color": "0.7 0.7 0.7 1",
"emphasis_font": "Playfair Display",
"emphasis_face": "Medium Italic", "emphasis_size": 265,
"active_color": "1 1 1 1", "emphasis_color": "1 1 1 1",
"model": "small", "language": "pt"|null}
-> {"ok": true, "path": "..._dynamic_subtitles.fcpxml", "message": "..."}
or {"ok": false, "error": "..."}
@@ -89,6 +119,16 @@ Commands:
-> {"ok": true, "diarization": bool, "diarization_message": "...",
"num_speakers": "..."}
voice_analysis
-> {"ok": true, "energy_threshold": 0.5, "emphasis_threshold": 0.85,
"emphasis_weights": {...}, "emotion_enabled": false,
"emotion_sensitivity": 0.5}
set_voice_analysis {"energy_threshold": 0.6, "emphasis_threshold": 0.9,
"emphasis_weights": {"energy": 0.4}|null,
"emotion_enabled": true, "emotion_sensitivity": 0.5}
-> same shape as voice_analysis (only given fields change)
Exit code 0 on success, 1 on error.
"""
@@ -122,16 +162,24 @@ from fcpxml.model_manager import ( # noqa: E402
is_model_downloaded,
list_installed_models,
load_catalog,
load_dynamic_subtitle_config,
load_hf_token,
load_num_speakers,
load_project_config,
load_selected_model,
load_silence_config,
load_transcript_language,
load_voice_analysis_config,
model_cache_dir,
save_dynamic_subtitle_config,
save_hf_token,
save_models_dir,
save_num_speakers,
save_project_config,
save_selected_model,
save_silence_config,
save_transcript_language,
save_voice_analysis_config,
)
from fcpxml.parser import parse_fcpxml # noqa: E402
from fcpxml.transcribe import transcribe # noqa: E402
@@ -600,14 +648,23 @@ def cmd_transcribe(args: dict) -> int:
total = len(media_paths)
results: list[dict] = []
for i, mp in enumerate(media_paths, 1):
_emit({"type": "progress", "fraction": i / total, "stage": f"Transcrevendo {Path(mp).name} ({i}/{total})…"})
stage = f"Transcrevendo {Path(mp).name} ({i}/{total})…"
_emit({"type": "progress", "fraction": (i - 1) / total, "stage": stage})
json_path = _transcript_json_path(mp, output_dir)
cached = _load_cached_transcript(json_path)
if cached is not None:
_emit({"type": "progress", "fraction": i / total, "stage": stage})
results.append(_result_row(mp, cached))
continue
data = transcribe(mp, model_size=model, language=language)
def _on_progress(file_fraction: float, _i: int = i, _stage: str = stage) -> None:
# Blend this file's own progress into the overall fraction so a
# single-media project doesn't jump straight to 100% before the
# actual (slow) decoding work has even started.
overall = (_i - 1 + file_fraction) / total
_emit({"type": "progress", "fraction": overall, "stage": _stage})
data = transcribe(mp, model_size=model, language=language, progress_cb=_on_progress)
if data is None:
_emit({"type": "error", "message": f"Não foi possível transcrever: {Path(mp).name}"})
return 1
@@ -638,6 +695,65 @@ def cmd_transcribe(args: dict) -> int:
return 0
def cmd_analyze_voice(args: dict) -> int:
"""Build the voice timeline (transcript+diarization+acoustics -> emphasis)
for every unique source media in the project, so `refine_voice_timeline`
and friends have something to read without ever reopening the audio.
Analysis only — writes _voice_timeline.json next to each media, doesn't
touch the project XML. `path` passes through unchanged so it composes
with the other batch steps (silence removal, captions) regardless of
where in the list it runs.
"""
path = str(args.get("path", ""))
if not path or not Path(path).exists():
_emit({"ok": False, "error": "Arquivo de projeto não encontrado."})
return 1
model = str(args.get("model", "") or load_selected_model() or "")
language = args.get("language")
if language is None:
language = load_transcript_language()
if language == "auto":
language = None
token = str(args.get("hf_token") or load_hf_token() or "")
num_speakers = str(args.get("num_speakers") or load_num_speakers() or "")
try:
proj = parse_fcpxml(path)
except Exception as exc:
_emit({"ok": False, "error": f"Erro ao ler o projeto: {exc}"})
return 1
tl = proj.primary_timeline or (proj.timelines[0] if proj.timelines else None)
media_paths: list[str] = []
if tl is not None:
for clip in getattr(tl, "clips", []):
mp = media_src_to_path(clip.media_path or "")
if mp and Path(mp).is_file() and mp not in media_paths:
media_paths.append(mp)
if not media_paths:
_emit({"ok": False, "error": "Nenhum arquivo de mídia acessível encontrado."})
return 1
from server import handle_build_voice_timeline
messages: list[str] = []
for mp in media_paths:
try:
contents = asyncio.run(handle_build_voice_timeline({
"media_path": mp, "model": model, "language": language,
"hf_token": token, "num_speakers": num_speakers,
"output_dir": args.get("output_dir"),
}))
except Exception as exc:
_emit({"ok": False, "error": f"Falha analisando {Path(mp).name}: {exc}"})
return 1
messages.append("\n".join(getattr(c, "text", str(c)) for c in contents))
_emit({"ok": True, "path": path, "message": "\n\n---\n\n".join(messages)})
return 0
def cmd_export_srt(args: dict) -> int:
"""Write a captions .srt synced to the edited timeline.
@@ -822,6 +938,135 @@ def cmd_set_diarization(args: dict) -> int:
return 0
def cmd_voice_analysis(args: dict) -> int:
"""Read the persisted voice-analysis settings (energy/emphasis/emotion)."""
_emit({"ok": True, **load_voice_analysis_config()})
return 0
def cmd_set_voice_analysis(args: dict) -> int:
"""Persist voice-analysis settings. Only the given fields change."""
weights = args.get("emphasis_weights")
config = save_voice_analysis_config(
energy_threshold=args.get("energy_threshold"),
emphasis_weights=weights if isinstance(weights, dict) else None,
emphasis_threshold=args.get("emphasis_threshold"),
emotion_enabled=args.get("emotion_enabled"),
emotion_sensitivity=args.get("emotion_sensitivity"),
)
_emit({"ok": True, **config})
return 0
def cmd_dynamic_subtitle_config(args: dict) -> int:
"""Read the persisted dynamic-subtitle style (font, size, color, layout)."""
_emit({"ok": True, **load_dynamic_subtitle_config()})
return 0
def cmd_set_dynamic_subtitle_config(args: dict) -> int:
"""Persist dynamic-subtitle style fields. Only the given fields change."""
config = save_dynamic_subtitle_config(**{
k: args.get(k) for k in (
"band_height", "block_center_y", "line_gap", "font", "font_size",
"emphasis_font", "emphasis_face", "emphasis_size",
"active_color", "emphasis_color", "text_scale",
)
})
_emit({"ok": True, **config})
return 0
def cmd_silence_config(args: dict) -> int:
"""Read the persisted silence thresholds (noise floor, duration, padding)."""
_emit({"ok": True, **load_silence_config()})
return 0
def cmd_set_silence_config(args: dict) -> int:
"""Persist silence thresholds. Only the given fields change."""
config = save_silence_config(
noise_db=args.get("noise_db"),
min_silence=args.get("min_silence"),
padding=args.get("padding"),
)
_emit({"ok": True, **config})
return 0
def cmd_apply_voice_actions(args: dict) -> int:
"""Apply a decision list (cuts/zooms/texts/markers) to the project XML.
The list is produced by a model reading the _voice_timeline.json — this
is the step that turns those decisions into an edit, and the one the
batch chain was missing: without it the app could measure the voice and
caption the result, but never cut by it.
`actions_path` points at the JSON; either a bare list or the
``{"actions": [...]}`` wrapper the skill emits is accepted. Times stay in
ORIGINAL source seconds — the handler resolves cuts first and shifts
everything else itself.
"""
path = str(args.get("path", ""))
if not path or not Path(path).exists():
_emit({"ok": False, "error": "Arquivo de projeto não encontrado."})
return 1
actions = args.get("actions")
if actions is None:
actions_path = str(args.get("actions_path", ""))
if not actions_path or not Path(actions_path).exists():
_emit({"ok": False, "error": "Arquivo de decisões (JSON) não encontrado."})
return 1
try:
with open(actions_path, encoding="utf-8") as fh:
loaded = json.load(fh)
except (OSError, ValueError) as exc:
_emit({"ok": False, "error": f"Erro ao ler as decisões: {exc}"})
return 1
actions = loaded.get("actions") if isinstance(loaded, dict) else loaded
if not isinstance(actions, list) or not actions:
_emit({"ok": False, "error": "A lista de decisões está vazia ou malformada."})
return 1
from server import handle_apply_voice_actions
try:
contents = asyncio.run(handle_apply_voice_actions({
"filepath": path,
"actions": actions,
"output_dir": args.get("output_dir"),
}))
except Exception as exc:
_emit({"ok": False, "error": f"Falha ao aplicar as decisões: {exc}"})
return 1
message = "\n".join(getattr(c, "text", str(c)) for c in contents)
# The handler reports dropped/rejected actions individually; hand the
# whole report back so the app can surface them instead of only the count.
out_path = path
for line in message.splitlines():
if line.startswith("- **Saved to**:"):
out_path = line.split("`")[1] if "`" in line else path
break
_emit({"ok": True, "path": out_path, "message": message})
return 0
def cmd_project_config(args: dict) -> int:
"""Read the last project folder/file the app was working on."""
_emit({"ok": True, **load_project_config()})
return 0
def cmd_set_project_config(args: dict) -> int:
"""Persist the last project folder/file. Only the given fields change."""
config = save_project_config(folder=args.get("folder"), file=args.get("file"))
_emit({"ok": True, **config})
return 0
def _result_row(mp: str, data: dict) -> dict:
words = data.get("words", [])
preview = (data.get("text", "") or "")[:160]
@@ -875,6 +1120,16 @@ def main() -> int:
"zoom_segments": cmd_zoom_segments,
"rename_speakers": cmd_rename_speakers,
"set_diarization": cmd_set_diarization,
"voice_analysis": cmd_voice_analysis,
"set_voice_analysis": cmd_set_voice_analysis,
"analyze_voice": cmd_analyze_voice,
"dynamic_subtitle_config": cmd_dynamic_subtitle_config,
"set_dynamic_subtitle_config": cmd_set_dynamic_subtitle_config,
"apply_voice_actions": cmd_apply_voice_actions,
"project_config": cmd_project_config,
"set_project_config": cmd_set_project_config,
"silence_config": cmd_silence_config,
"set_silence_config": cmd_set_silence_config,
}
handler = handlers.get(command)
if handler is None:
+2 -2
View File
@@ -16,7 +16,7 @@ O sistema é um **servidor MCP em Python** que lê/analisa/reescreve arquivos
│ graphify.sh/.md Pipeline de graphify do código │
├─────────────────────────────────────────────────────────────┤
│ server.py — CAMADA MCP / TRANSPORTE (NÃO tem lógica) │
│ 62 tools, handlers, prompts, resources, dispatch │
│ 73 tools, handlers, prompts, resources, dispatch │
│ Só valida entrada/saída e traduz JSON-RPC → chamadas │
├─────────────────────────────────────────────────────────────┤
│ fcpxml/ — "ENGINE" = NÚCLEO PURO Python (desacoplado) │
@@ -90,7 +90,7 @@ programático — round-trips voltam pelas ferramentas XML.
| Controle Live do FCP | `fcpxml/live.py` |
| Segurança XML (`defusedxml`, `serialize_xml`) | `fcpxml/safe_xml.py` |
| Validação contra DTDs da Apple | `fcpxml/dtd.py` |
| Transporte MCP (62 tools) | `server.py` |
| Transporte MCP (73 tools) | `server.py` |
## 6. Mapa de dependências (você está aqui se for mexer no X → quem tocar)
+68 -2
View File
@@ -1,4 +1,4 @@
# 03 — Camada MCP (`server.py`) — 62 ferramentas
# 03 — Camada MCP (`server.py`) — 73 ferramentas
`server.py` (3824 linhas) é a camada de transporte. Não tem lógica de timeline —
mapeia nome → handler e delega ao Engine. O dispatch é um dicionário
@@ -20,7 +20,7 @@ mapeia nome → handler e delega ao Engine. O dispatch é um dicionário
| `_parse_timestamp_parts()` | 433 | Parse de timestamps (min:seg, H:MM:SS, SMPTE) |
| `_detect_flash_frames/gaps/duplicate_groups()` | 1667+ | Detectores de QC |
## As 62 ferramentas por categoria
## As 73 ferramentas por categoria
### Timeline & análise (Projeto)
`list_projects`, `analyze_timeline`, `list_clips`, `list_markers`, `list_connected_clips`,
@@ -57,6 +57,72 @@ mapeia nome → handler e delega ao Engine. O dispatch é um dicionário
### Reformat
`reformat_timeline`.
### Voz (análise → decisão → aplicação)
`analyze_voice_features`, `build_voice_timeline`, `refine_voice_timeline`,
`remove_speakers`, `apply_voice_actions`, `get_voice_analysis_config`,
`save_voice_analysis_config`.
O fluxo é sempre o mesmo: `build_voice_timeline` mede (caro, roda uma vez) →
o modelo decide os cortes → **`refine_voice_timeline` renormaliza sobre o que
sobrou** (barato, sem reabrir áudio) e propõe as janelas de zoom → o modelo
corta a lista pelo ritmo → `apply_voice_actions` aplica. Pular a renormalização
faz o ranking de ênfase apontar para as palavras erradas (ver
`05_EXPERIENCIAS.md`).
### Legendas dinâmicas (geração → validação → aplicação)
`generate_dynamic_subtitles`, `validate_subtitle_layout`, `transcript_markers`.
**Sempre gere e depois valide — nunca dê a geração como pronta sem
`validate_subtitle_layout`.** A composição garante "sem sobreposição" só
*por construção* dentro do que ela mesma sabe medir; um título editado à
mão, uma palavra fora do alcance do que foi calibrado, ou conteúdo antigo
no mesmo arquivo escapam dessa garantia. Fluxo:
```
generate_dynamic_subtitles(filepath)
↓
validate_subtitle_layout(output_path) ← sempre, mesmo quando "parece certo"
↓
severidade none/warning → entregar
severidade probable/severe → investigar CADA colisão pela fração exata do
XML antes de mudar código (ver checklist abaixo)
```
**Antes de atribuir uma colisão ao gerador, confirme que é o gerador.**
Um `<title>` de nome estranho (`ref` diferente, params tipo `Auto-Shrink`/
`Left Margin` que `_make_text_title_clip` nunca escreve) é conteúdo humano
ou de outra ferramenta, não um bug — comparar contra o arquivo original
(`grep` pelo texto) resolve em segundos. Caso real: uma colisão "severa"
era um título manual feito no FCP que sobrou no arquivo reaproveitado como
base de teste (`05_EXPERIENCIAS.md`, 2026-08-19).
**Antes de atribuir uma colisão a uma sobreposição real, confirme pela
fração exata do FCPXML, não pelo float arredondado.** Dois títulos que só
se tocam na borda (um bloco some exatamente quando o próximo começa, por
design) podem imprimir tempos "iguais" e ainda assim colidir no relatório
por ruído de ponto flutuante — `float(a+b) != float(c)` mesmo quando as
frações `a+b` e `c` são idênticas. `temporal_overlap()` já tem uma
tolerância (`_BOUNDARY_EPSILON = 1e-6`, muitas ordens abaixo de um frame)
para absorver isso; se uma colisão nova parecer nascer do nada, comparar
`m._parse_time(...)` dos dois títulos por igualdade exata antes de
suspeitar de sobreposição de verdade.
**Constantes que resolvem os três bugs já encontrados nesta área** (todas
em `fcpxml/text_layout.py`, exceto a última):
| Constante | O que resolve | Por quê |
|---|---|---|
| `TEXT_TEMPLATE_FONT_SCALE = 2.0` | Posição e tamanho de fonte dessincronizados | O template "Text" do FCP posiciona no espaço do **frame** (2160×3840), mas o layout mede em pontos de meia-escala (1080×1920). Escalar só o tamanho da fonte e não a posição espalha o texto errado — os dois têm que ser convertidos pelo mesmo fator na saída. |
| `_EMPHASIS_ITALIC_CUSHION_RATIO = 0.06` | Linha de corpo lendo apertada sob a linha de ênfase | O itálico da Playfair inclina as hastes além da caixa de tinta que a métrica mede; ~14pt de respiro extra só nessa fronteira corrige sem tocar no `line_gap` do resto. |
| `fit_emphasis()` (função, não constante) | Palavra de ênfase estourando o frame inteiro | Só linhas de corpo faziam wrap contra `box.width`; a linha de ênfase (sempre uma palavra só) nunca foi checada. Uma palavra longa ou toda maiúscula podia medir mais que o frame inteiro sozinha. Encolhe `font_size`+`kerning` pelo mesmo fator até caber — nunca abaixo do tamanho do corpo, senão ênfase deixa de ser ênfase. |
| `_BOUNDARY_EPSILON = 1e-6` (`fcpxml/collision.py`) | Falso positivo de colisão em títulos que só se tocam | Ver parágrafo acima. |
Detalhe de implementação e efeito medido de cada um: `05_EXPERIENCIAS.md`,
entradas de 2026-08-19 (#15 zoom, #16 colisão por float, #17 auto-fit da
ênfase — a #15 é do módulo de voz, não de legendas, mas mesma causa-raiz
de fundo: um mecanismo que lê a própria saída anterior precisa continuar
sendo fonte de verdade legível, não só efeito colateral write-only).
### Live (macOS)
`push_to_fcp`, `list_fcp_libraries`.
+1 -1
View File
@@ -63,7 +63,7 @@ uv run --extra dev pytest tests/ -v # Testes com extra de dev
## 4. Estado atual do sistema (resumo "até agora")
- **v0.6.35** — núcleo FCPXML completo em Python (`fcpxml/`).
- **62 ferramentas MCP** em `server.py`, organizadas por dispatch `TOOL_HANDLERS`.
- **73 ferramentas MCP** em `server.py`, organizadas por dispatch `TOOL_HANDLERS`.
- **Suporte FCPXML 1.8–1.14** (`.fcpxml` e bundles `.fcpxmld` com sidecars),
escrita padrão 1.13.
- **Dual-mode:** XML (principal) + Live (push_to_fcp / list_fcp_libraries via Apple events).
+501
View File
@@ -33,6 +33,407 @@ Use o bloco abaixo como modelo. Uma entrada = um problema resolvido/reconhecido.
## Registro de Experiências
### 2026-08-19 — Linha de ênfase das legendas dinâmicas sem limite de largura: auto-fit implementado
- **Contexto:** validando `generate_dynamic_subtitles` sobre o corte real do Mastopexia (ver entrada anterior sobre `validate_subtitle_layout`), sobraram 15 títulos `outside_frame` mesmo depois de eliminados os falsos positivos de colisão.
- **Causa raiz:** `compose_sentence()` (`fcpxml/text_layout.py`) faz wrap das linhas de corpo contra `box.width`, mas a linha de ênfase — sempre uma palavra só, a "key word" em itálico grande — nunca era checada contra largura nenhuma, porque uma palavra sozinha não tem como quebrar em duas linhas. Uma palavra longa (`estruturado,`, `sustentação`, `proporcional`) ou toda maiúscula media mais que o frame inteiro sozinha: `estruturado,` a 460pt (ponto emitido, após `TEXT_TEMPLATE_FONT_SCALE`) mediu **2538px de largura contra 2160px de frame**, estourando os dois lados centrada.
- **Onde:** `fcpxml/text_layout.py::compose_sentence`.
- **Solução adotada:** `fit_emphasis()` mede a linha de ênfase e, se ultrapassar `box.width`, encolhe `font_size` e `kerning` pelo mesmo fator — a largura é linear nesses dois parâmetros juntos, então o fator exato é `box.width / width_medida`, sem iteração. Nunca encolhe abaixo do tamanho do corpo (`body_look.font_size`): ênfase do mesmo tamanho que o texto normal deixa de ser ênfase. Como a extensão vertical de tinta (`ink_extent`) também é linear em `font_size`, encolher a largura encolhe a altura usada no empilhamento junto — restaurando de brinde a garantia "sem sobreposição por construção" que o resto da função já tinha, sem precisar de lógica extra para isso.
- **Efeito medido:** revalidando o mesmo corte, `outside_frame` caiu de 15 para 0, severidade de `severe` para `warning` (restam avisos de área segura, não erros de frame).
- **Achado colateral:** a única colisão que sobrou depois da correção (`MASTOPEXIA` × `é a cirurgia que`) não era bug nenhum — era um título manual (`Auto-Shrink`, template "Text" nativo do FCP, sem relação com `_make_text_title_clip`) presente no arquivo base reutilizado para o teste, provavelmente de uma edição manual no FCP, não gerado por nenhuma chamada da sessão. Revalidando sobre uma base limpa, sem esse título estranho: **0 colisões**. Vale sempre revalidar sobre uma base conhecida antes de atribuir um achado ao código.
- **Estado:** `resolvido` — corrigido em `text_layout.py`, suíte completa e lint passando, revalidado sobre o corte real.
---
### 2026-08-19 — `validate_subtitle_layout` acusava colisão em 7 de 8 casos por ruído de ponto flutuante, não por sobreposição real
- **Contexto:** primeiro uso real de `validate_subtitle_layout` (ferramenta nova) sobre o corte do Mastopexia com legendas dinâmicas geradas. Relatou severidade `severe`: 8 colisões, 15 títulos fora do frame, 10 fora da área segura.
- **Sintoma:** ao rastrear cada colisão reportada pelas frações exatas do FCPXML, 7 dos 8 pares eram títulos **consecutivos que terminam exatamente quando o próximo começa** — o design pretendido ("cada bloco some quando o próximo aparece", `writer.py`'s `block_ends[i] = block_starts[i+1]`) funcionando corretamente. O oitavo (`MASTOPEXIA` × `é a cirurgia que`) era uma sobreposição real de ~0,46s.
- **Causa raiz:** `temporal_overlap()` em `fcpxml/collision.py` compara `end = start.to_seconds() + duration.to_seconds()` (soma de dois floats já arredondados) contra `start.to_seconds()` de outro título (uma única divisão) — mesmo quando a fração exata subjacente é bit-idêntica nos dois casos, a soma de dois floats arredondados não bate com uma única divisão da soma exata dos numeradores (não-associatividade de ponto flutuante). Medido: diferença de ~4,5×10⁻¹³s — treze ordens de grandeza menor que um frame (~0,04s) — suficiente para inverter `start_b < end_a` de `False` para `True` e disparar uma colisão `severe` fantasma. Contradizia o próprio comentário do código ("a title ending exactly as the next begins is never flagged").
- **Onde:** `fcpxml/collision.py::temporal_overlap`.
- **Solução adotada:** tolerância `_BOUNDARY_EPSILON = 1e-6` subtraída de ambos os lados da comparação — muitas ordens de grandeza abaixo de qualquer fronteira de frame real, então não mascara nenhuma sobreposição genuína, só absorve o ruído de arredondamento entre dois caminhos de cálculo do mesmo instante.
- **Aprendizado:** comparar dois floats derivados do MESMO valor exato por caminhos aritméticos diferentes (soma vs. divisão direta) nunca deve usar igualdade/desigualdade estrita — vale para qualquer checagem "toca a borda mas não deveria contar", não só tempo de título. O sintoma (severidade `severe` sem nenhuma sobreposição visível no material) é o sinal de alerta: sempre rastrear a colisão até as frações exatas do XML antes de aceitar o relatório da ferramenta de validação como verdade.
- **Estado:** `resolvido` — corrigido em `collision.py`, teste de regressão com os números reais do caso (`test_boundary_survives_float_noise_from_the_writer`), suíte completa (1369 testes) e lint passando, revalidado sobre o corte real: 8 colisões → 1.
---
### 2026-08-19 — `add_zoom` perdia o enquadramento real quando dois zooms caíam no mesmo clipe pós-corte
- **Sintoma:** no mesmo teste real (Mastopexia), o clipe de abertura do corte apareceu "achatado" no FCP — Scale 100% em vez do enquadramento real (~177%) que o projeto original já tinha, enquanto os clipes seguintes apareciam corretos. Rotação e posição estavam certas; só a escala quebrava, e só no primeiro trecho.
- **Causa raiz:** `add_zoom()` (`fcpxml/writer.py`) lê o enquadramento-base de um clipe só de um jeito: o atributo estático `scale="X Y"` em `<adjust-transform>`. Isso funciona na primeira chamada. Mas quando dois zooms editoriais caem dentro do **mesmo** clipe sobrevivente (dois picos de ênfase que o corte não separou em clipes distintos), a segunda chamada de `add_zoom` encontra não mais um atributo estático, e sim um `<param name="scale">` já **animado** pela primeira — e o código só sabia ler atributo. `stale.get('scale')` voltava `None`, o base virava `1.0` por padrão, e a linha seguinte (`clip.remove(stale)`) **apagava a animação da primeira chamada inteira**, substituindo por uma segunda com base errada.
- **Onde:** `fcpxml/writer.py::add_zoom` (base scale + merge de keyframes); reproduzido isolando `cut_clip_ranges` + duas chamadas de `add_zoom` no mesmo elemento.
- **Por que passou despercebido:** o teste existente (`test_only_one_transform_remains`) já chamava `add_zoom` duas vezes no mesmo clipe, mas só checava que sobrava **um** `<adjust-transform>` na árvore — nunca verificou se a base da segunda chamada estava certa. A suíte cobria a estrutura, não o valor.
- **Solução adotada (duas partes):**
1. Quando não há atributo `scale` estático, `add_zoom` agora lê o `<param name="scale">` existente e recupera a base como o **menor** valor entre as keyframes — válido porque `MIN_ZOOM_SCALE == 1.0` garante que todo pico é `>= base`, então o menor valor keyframeado é sempre o resting scale, seja ele o de abertura, o de fecho ou qualquer um no meio.
2. Duas janelas de zoom no mesmo clipe agora só se **substituem** quando as janelas de tempo se sobrepõem (é o mesmo evento sendo reajustado); quando são **disjuntas** (dois picos editoriais distintos que um corte não separou), as keyframes são **empilhadas** no mesmo `keyframeAnimation` em vez de uma apagar a outra — FCPXML aceita quantas keyframes forem necessárias num único `<param>`.
- **Aprendizado:** "preservar o enquadramento existente" precisa valer em **toda** leitura subsequente do mesmo clipe, não só na primeira. Um mecanismo que lê corretamente da fonte original mas degrada ao reler sua própria saída anterior é o mesmo bug de fundo da entrada #13 (renormalização) por outro ângulo: qualquer estado que o sistema regrava precisa continuar sendo uma fonte de verdade legível, não só um efeito colateral write-only. Vale desconfiar de qualquer `findall()`/leitura de atributo que tenha um "senão assume 1.0/padrão" — é aí que a segunda chamada perde o que a primeira escreveu.
- **Estado:** `resolvido` — corrigido em `writer.py`, dois testes de regressão adicionados (`test_second_zoom_on_same_clip_keeps_the_real_base_scale`, `test_overlapping_zoom_on_same_clip_replaces_instead_of_stacking`), suíte completa (1344 testes) e lint passando, corte real do Mastopexia regravado e conferido.
---
### 2026-08-19 — Teste real fechou o ciclo, e revelou um offset sistemático de ~0,4s no timing por palavra
- **Contexto:** primeiro teste ponta a ponta de `editar-por-voz` num projeto real (Mastopexia, 196,6s de gravação de roteiro com 6 tomadas). Fluxo completo: `build_voice_timeline` → triagem manual (tomada/bastidor/frase abandonada) → `refine_voice_timeline` sobre os sobreviventes → escolha de zoom por função narrativa → `apply_voice_actions`. Corte final: 196,6s → ~50s, 3 clipes.
- **Sintoma:** antes de rodar `remove_media_silence`, medi manualmente o RMS do áudio real nas emendas propostas pelo corte e achei folgas de ~0,4-0,6s onde o JSON dizia que a fala começava/terminava. Comparando timestamp da transcrição contra o ataque real medido em 6 pontos do vídeo (ffmpeg `astats`), o erro era **sistemático, sempre no início da palavra**, entre +0,35s e +0,51s — os finais de palavra batiam certo (+0,01 a +0,15s).
- **Causa raiz:** `transcribe.py` usa `word_timestamps=True` do faster-whisper, que deriva os tempos por atenção cruzada — aproximado por natureza, sem alinhamento forçado. O submódulo WHISPERX existe no repositório mas **não é usado** em nenhum ponto do código; não há etapa de alinhamento fonético.
- **Por que isso importa mais do que parece:** o erro contamina toda decisão temporal a jusante — zoom disparava ~0,4s antes da palavra-alvo, `gap_before` subestimava pausas reais na mesma medida (o que afeta diretamente a régua de silêncio recém-adotada), e as folgas de corte saíam erradas nas emendas.
- **Decisão tomada:** não rodei `remove_media_silence` bruto sobre o corte. A detecção (ffmpeg, limiar -30dB/0,5s) não distingue "batida entre frases dentro da régua de 1,5s" de "ar morto de emenda" — cortar ambos teria apertado frases fluidas. Corrigi os tempos manualmente medindo o ataque real nos pontos críticos (cabeça, 2 emendas, cauda, 3 zooms) e refiz o corte numa passada só.
- **Solução adotada (paliativa, aplicada manualmente neste teste):** medir o RMS real com `ffmpeg -af astats=metadata=1:reset=1:length=0.05,ametadata=print` em janelas curtas ao redor de cada ponto crítico antes de fixar um corte ou zoom que dependa de precisão de frame. Não é o padrão do sistema — é o que cobre a lacuna até o alinhamento forçado existir.
- **Solução estrutural ainda pendente:** ligar o WhisperX (ou alinhamento forçado equivalente) em `transcribe.py`, o que levaria o erro de ~400ms para ~30ms e corrigiria zoom, corte e `gap_before` de uma vez, sem paliativo por projeto. Não implementado ainda — é mudança de pipeline, exige regerar todos os `_transcript.json`/`_voice_timeline.json` existentes.
- **Aprendizado:** "não reestime tempos no olho" (critério 01) continua certo para decisão *editorial* — mas não cobre erro sistemático de *medição* na fonte dos tempos. Um offset constante e na mesma direção, em vários pontos do material, é sinal de bug no pipeline de transcrição, não de julgamento errado sobre o material. Vale conferir com uma amostra de áudio real antes de confiar cegamente em timestamp de word-level de qualquer fonte nova.
- **Estado:** `parcialmente resolvido` — paliativo documentado e aplicado neste teste; correção estrutural (WhisperX) pendente de implementação.
---
### 2026-08-19 — Capacidade existente sem porta de entrada: a Fase 4 da skill era letra morta
- **Sintoma:** a skill `editar-por-voz` manda, como fase obrigatória, reanalisar o material sobrevivente antes de escolher zooms — e o modelo não tinha como cumprir isso. `restrict_to_kept()` e `suggest_zoom_windows()` existiam, estavam testadas e documentadas, mas **nenhuma ferramenta MCP as expunha**. Na prática, todo zoom continuava sendo escolhido com o ranking bruto, exatamente o erro que a entrada anterior descreve.
- **Causa raiz:** a implementação parou na camada Engine. O critério foi escrito descrevendo chamadas Python, que só os testes conseguiam fazer — a distância entre "existe no `fcpxml/`" e "o modelo consegue chamar" passou despercebida porque a suíte cobria a função, não o caminho.
- **Onde:** `server.py` (nova tool `refine_voice_timeline` + handler + dispatch), `tests/test_refine_voice_timeline_tool.py`, `.claude/skills/editar-por-voz/criterios/04-reanalise-do-material-restante.md`.
- **Solução adotada:** ferramenta `refine_voice_timeline(media_path, cuts, min_gap, max_zooms, save)`, que numa chamada devolve a comparação bruto × sobreviventes, os picos re-ranqueados e os candidatos a zoom. Handler fino: nada de lógica nova, só o caminho até o que já existia. O critério 04 passou a citar a ferramenta em vez das funções Python.
- **Aprendizado:** função coberta por teste unitário **não** é capacidade entregue. Toda vez que um critério de skill mandar "rode X", verifique que X é chamável pelo modelo — senão o critério vira instrução impossível, e o modelo segue em frente sem erro visível. Vale um teste de registro (`a tool está em list_tools` + `está no TOOL_HANDLERS`) para cada ferramenta nova.
- **Estado:** `resolvido`
---
### 2026-08-19 — Ênfase é relativa: analisar o bruto e editar o corte final são perguntas diferentes
- **Problema:** os zooms estavam sendo escolhidos a partir da análise do material **bruto**. Mas energia é normalizada contra o momento mais alto da gravação — que era `"Amor, eu tô intacto!"` (energia 1,00), justamente uma das falas **cortadas**. Todo o material que sobrou estava pontuado contra uma referência que o espectador nunca veria, comprimindo artificialmente as notas do corte final.
- **Solução:** `restrict_to_kept()` filtra a timeline pelos ranges removidos e **re-normaliza sobre os sobreviventes**. Efeito medido: ênfase média subiu de 0,179 (bruto) para 0,197 (só o que ficou) — o material restante passou a usar a escala inteira. E o ranking mudou de figura: `"Aquela"` 0,39→0,42, `"mastopexia"` 0,26→0,34, `"devolver"` entrando com 0,35.
- **Pré-requisito que virou bug na hora:** re-normalizar exige os valores **brutos**, e o JSON só guardava os normalizados. Adicionados `energy_raw` e `pitch_hz` em cada palavra. A primeira execução saiu com todas as notas caindo — sintoma de estar lendo campo ausente num JSON gerado antes da mudança. Regerar o JSON resolveu; vale lembrar que mudança de esquema exige regerar os caches antes de interpretar qualquer resultado.
- **Armadilha de API evitada:** `enrich_words` chamava `word_pitch_energy`, que **sobrescreve** `energy`/`pitch_hz` com `None` quando não há frame tracks — então re-analisar um subconjunto zerava tudo. Em vez de remendar restaurando os valores depois (que foi a primeira tentativa, e ficou ilegível), entrou o parâmetro `already_measured`.
- **`suggest_zoom_windows()`:** propõe uma janela por frase, a partir da palavra de **conteúdo** mais enfática (artigos e conectivos filtrados por `_FUNCTION_WORDS` — um "a" falado alto continua sendo um artigo), indo até o fim da frase; `min_gap` mantém os zooms afastados.
- **Aprendizado:** "qual o momento mais forte da gravação?" e "qual o momento mais forte do vídeo final?" são perguntas distintas sempre que a métrica for relativa. Toda métrica normalizada precisa ser recalculada quando o conjunto muda — caso contrário ela responde a pergunta errada, silenciosamente.
- **Estado:** `resolvido` (1325 testes verdes; 3 zooms escolhidos pela re-análise aplicados em clipes distintos, DTD 1.14 válido).
---
### 2026-08-19 — Forma do punch-in é assimétrica: entrada de 0,5s, saída de 1 frame
- **Regra editorial (do usuário):** na palavra de ênfase, zoom in **rápido** (~meio segundo); segura durante a frase de impacto; e no fim **volta de um quadro para o outro, sem transição nenhuma** — o vídeo simplesmente retoma o enquadramento e segue o fluxo.
- **O que havia:** `add_zoom` tinha um único `ease` (padrão 0,3s) aplicado **simetricamente** na entrada e na saída, produzindo um retorno lento que chama atenção para si.
- **Solução:** `ease` passou a valer só para a entrada (padrão **0,5s**) e a saída virou **um frame**, calculado do `frameDuration` real da sequência (`ease_out` opcional para quem quiser retorno gradual). Medido no material: entrada 0,501s, hold 3,378s, saída 0,042s = 1 frame a 23,976fps.
- **Detalhe que quase passou:** o handler em `server.py` forçava `ease=float(action.params.get("ease", 0.3))`, então o padrão novo do writer nunca chegava a valer — o zoom saía com 0,33s de entrada. Um default duplicado em duas camadas é sempre o errado das duas; o handler passou a repassar `ease` **só quando explicitamente informado**, deixando o writer ser o dono do padrão.
- **Aprendizado:** ao mudar um default, procurar quem já o repassa. Um `params.get("x", <default>)` numa camada acima anula silenciosamente o default da camada que de fato conhece o assunto.
- **Estado:** `resolvido` (1311 testes verdes, `TestZoomShapeIsAsymmetric` fixa a forma; DTD 1.14 válido) — pendente de conferência visual no FCP.
---
### 2026-08-19 — Export de calibração do FCP fecha o zoom: keyframe só com `time` e `value`
- **Como veio:** depois de o zoom continuar não aparecendo, o usuário fez o zoom **à mão no FCP** sobre o mesmo material e exportou o FCPXML (`Mastopexia - exemplo de zoom.fcpxmld`) — o padrão de calibração já registrado em 2026-08-15 e 2026-08-18, agora aplicado a keyframes.
- **O que o export real mostrou:**
```xml
<adjust-transform position="0.160319 0.663249" rotation="90.1008">
<param name="scale">
<keyframeAnimation>
<keyframe time="2329601280/720000s" value="1.77311 1.77311"/>
```
1. `position`/`rotation` mantidos como atributos e o atributo `scale` **removido** quando a escala é animada — confirmou a correção de preservação de enquadramento;
2. o primeiro keyframe cai exatamente no `start` do clipe (3235,557s) — **confirmou** a correção de timebase de origem;
3. **`<keyframe>` carrega apenas `time` e `value`** — sem `interp` e **sem `curve`**.
- **Correção final:** removido o `curve="smooth"` que eu havia adicionado ao trocar o `interp`. O DTD permite `curve` (default `smooth`), mas como o importador já havia rejeitado `interp` neste mesmo param vetorial, não há razão para apostar que `curve` sobrevive — o export real do FCP não escreve nenhum dos dois, então passamos a escrever nenhum dos dois. Estrutura agora idêntica à do FCP.
- **Aprendizado:** ao corrigir um atributo rejeitado pelo importador, **não basta trocar por outro plausível do DTD** — foi o que fiz (`interp` → `curve`) e ficou uma segunda aposta não verificada em cima da primeira. A resposta certa era pedir um export de calibração e copiar. Vale a regra: diante de qualquer incerteza sobre o que o FCP aceita, o caminho mais curto é um export real, não uma segunda leitura do DTD.
- **Estado:** `resolvido` (1306 testes verdes, estrutura conferida atributo a atributo contra o export do usuário, DTD 1.14 válido) — pendente de confirmação de importação.
---
### 2026-08-19 — Zoom importava como nada: keyframe em tempo relativo, não no timebase de origem
- **Sintoma:** usuário importou o FCPXML no FCP e relatou "não tem zoom, não tem corte, não tem nada".
- **Causa 1 (real) — o zoom:** os `<keyframe>` do `adjust-transform` eram escritos em segundos **relativos ao clipe** (0,08s a 4,80s), mas o clipe tem `start="74637363/24000s"` = **3109,9s** (timecode de origem). O FCP procura a animação no timebase do próprio clipe, não encontra keyframe nenhum na janela dele, e importa o zoom como **nada** — sem erro, sem aviso. Corrigido somando o `start` do clipe (`media_origin + tempo relativo`), exatamente o que `add_text_title` já fazia e **documentava**: *"Anchored in SOURCE media coordinates… so the title lands on screen instead of at ~0s of the media (which FCP silently drops)"*. O `add_zoom` nunca recebeu o mesmo tratamento.
- **Por que escapou:** todo fixture sintético e o projeto da Erika têm clipe com `start="0s"`, onde relativo e absoluto coincidem. Pior: o teste `test_add_zoom_creates_keyframed_transform` usava uma fixture com `start="10s"` — tinha tudo para pegar o bug — mas afirmava `times == [1.0, 1.5, 2.5, 3.0]`, ou seja, **fixava o comportamento errado**. Reescrito para ancorar em `origin + relativo`, com `assert origin > 0` garantindo que a fixture continue exercitando o caso.
- **Causa 2 (percepção) — o corte:** os cortes **estavam** no arquivo (3 clipes, 47,6s contra 196,7s do bruto). Mas o XML mantinha `<event name="17-08-2026">` e `<project name="Mastopexia">` idênticos ao original, então a importação criava um projeto homônimo no mesmo evento e o usuário abriu o antigo. Passou-se a renomear o projeto para `<nome> — corte por voz`.
- **Aprendizado 1:** "não fez nada" pode ser duas coisas muito diferentes — não gerou, ou gerou e o usuário não achou. Vale sempre inspecionar o arquivo antes de concluir, e nunca deixar a saída indistinguível da entrada dentro do app de destino.
- **Aprendizado 2 (repetição do padrão de 2026-08-19/`interp`):** quando um módulo já resolve um problema de coordenadas e **documenta** a solução no docstring, procurar os irmãos que fazem operação equivalente. `add_text_title` sabia ancorar em coordenadas de origem; `add_zoom` e qualquer outro futuro escritor de keyframes precisam da mesma regra.
- **Estado:** `resolvido` (1306 testes verdes, keyframes conferidos dentro da janela do clipe: 3297,74–3302,46s para um clipe de 3297,7–3302,6s; DTD 1.14 válido) — pendente de nova importação no FCP.
---
### 2026-08-19 — Primeira edição real ponta a ponta: três bugs de ordem/identidade que a suíte não pegava
Triagem editorial completa de um vídeo institucional (196,7s → 47,6s, 76% de redução),
feita a partir do `_voice_timeline.json`. A geração do FCPXML expôs três bugs, todos
invisíveis em fixture sintética porque dependem de um projeto **já editado**.
**1. `add_zoom` destruía o enquadramento do editor.** O clipe original trazia
`<adjust-transform position="0.160319 0.663249" rotation="90.1008" scale="1.77311 1.77311"/>`
— material gravado de lado e reenquadrado à mão. `add_zoom` removia qualquer
`adjust-transform` existente ("Replace rather than stack a prior zoom") e criava o seu do
zero, então o trecho com zoom voltava **girado 90°**. Corrigido: os atributos estáticos
(position/rotation/anchor) são preservados e a escala passa a ser animada **relativa** à
base (1,77311 → 1,77311 × 1,18 → 1,77311). Comentário "replace rather than stack" estava
certo na intenção e errado no alcance — nem todo `adjust-transform` é um zoom anterior.
**2. Posicionar antes de cortar espalhava o zoom e apagava marcadores.**
`cut_clip_ranges` divide o clipe e reescreve a spine; o `adjust-transform` era **copiado
para os 3 pedaços** e os `<marker>` sumiam. Invertida a ordem: **cortes primeiro**,
posicionamentos depois — `resolve_actions` já converte os tempos para a timeline
pós-corte, então continuam apontando para o mesmo instante.
**3. Depois de cortar, todos os pedaços têm o MESMO nome.** `add_zoom`/`add_marker`
resolviam o clipe por nome (`_require_clip`), então toda edição caía no **primeiro**
pedaço. `_require_clip` passou a aceitar um `Element` direto (como `add_text_title` já
fazia), e o handler passa o elemento exato — mapeando pelo **offset na timeline**, não
mais pela janela de origem.
**Bônus — marcador é ponto, não trecho.** `resolve_actions` exigia que início *e* fim
sobrevivessem ao corte, descartando justamente os marcadores úteis: os que sinalizam uma
emenda e por definição encostam na borda do corte. Marcadores passaram a resolver só pelo
início.
- **Aprendizado:** os três bugs são a mesma família — **ordem de operações e identidade de
elemento** depois de uma operação que reestrutura a árvore. Fixture sintética tem um
clipe limpo, sem transformações prévias e sem nomes duplicados, então nada disso
aparece. Testar contra um projeto **real já editado** é categoricamente diferente de
testar contra XML gerado por nós.
- **Regressões adicionadas:** `TestZoomPreservesExistingFraming` (5),
`TestPlacementsLandOnTheRightPieceAfterCuts` (3), `TestMarkersSurviveCutEdges` (4).
- **Estado:** `resolvido` (1301 testes verdes, DTD 1.14 válido, enquadramento conferido no
XML) — pendente de importação real no FCP pelo usuário.
---
### 2026-08-19 — Diarização quebrada pelo torchcodec + descoberta: separar tomada de conversa NÃO é diarização
**Parte 1 — a falha técnica.** Com token e termos válidos, `diarize()` retornava `None` e o log dizia só "diarization failed" (o `except Exception:` amplo engolia a causa). Rodando o pyannote direto, a causa apareceu: `pyannote.audio` 4.x decodifica áudio via **torchcodec**, que linka contra uma versão específica do FFmpeg — `dlopen(libtorchcodec_core4.dylib): Library not loaded: @rpath/libavutil.56.dylib`. O FFmpeg instalado é outro major, e a diarização caía inteira numa máquina em que tudo o mais funcionava.
- **Solução:** `_load_waveform()` em `diarize.py` decodifica o áudio por conta própria (reusando `decodable_audio()` do `voice_features.py`, que já extrai WAV mono 16 kHz via ffmpeg) e passa a `{"waveform": tensor, "sample_rate": sr}` que o pyannote aceita — pulando o torchcodec por completo. Bônus: containers de vídeo passam a funcionar direto. Resultado: 33 turnos, 2 participantes, ~1min40s para 3min17s de áudio.
- **Aprendizado:** `except Exception` sem registrar a exceção transforma falha diagnosticável em mistério. O log deveria carregar a causa; sem isso, foi preciso reexecutar a biblioteca à mão para ver o erro real.
**Parte 2 — a descoberta de produto, mais importante.** Com a diarização funcionando, o `remove_speakers` encontrou só **uma** fala do SPEAKER_01 — e ainda por cima uma atribuição errada. Cruzando os turnos com a transcrição, a explicação apareceu: o pyannote detecta a segunda voz em 44,1–46,0s e 51,4–54,8s, mas a transcrição **não tem nada** nesses intervalos (vãos de 43,8→47,5 e 48,8→55,4). A pessoa da equipe está **fora do microfone**: a diarização a ouve, o Whisper não a transcreve.
- **Consequência:** as falas que o usuário quer descartar (*"Amor, eu tô intacto!"*, *"Só clica aí agora na tela."*, *"Posso começar da mastopexia?"*) são **da própria protagonista** — mesma voz, contexto diferente. Cortar por participante não resolve esse caso.
- **Aprendizado:** "quem fala" e "isso é tomada válida?" são perguntas **diferentes**, e é tentador confundi-las porque ambas soam como "separar as partes do vídeo". Diarização resolve a primeira; só a linguagem resolve a segunda. `remove_speakers` continua válido para o caso em que o entrevistador está microfonado (ex. depoimento da Erika), mas não é a ferramenta para triar tomada de conversa.
- **Estado:** `resolvido` (diarização funcional, 1294 testes verdes); triagem tomada/conversa fica na camada de linguagem, critérios em `.claude/skills/editar-por-voz/SKILL.md`.
---
### 2026-08-19 — Pausa longa: ruído para ênfase, sinal para estrutura (o mesmo dado, dois usos opostos)
- **Sintoma:** no material real, palavras de energia baixíssima lideravam o ranking de ênfase. No Mastopexia, 4 dos 7 picos eram assim: `"mastopexia"` (energia 0,25, pausa 6,2s), `"Aquela"` (0,32, 8,7s), `"Aumenta."` (0,25, 5,7s). O mesmo padrão aparecia no depoimento da Erika.
- **Causa raiz:** `compute_emphasis` normalizava a pausa contra `max_pause=1,5s` **saturando** — ou seja, uma pausa de 8,7s e uma de 1,5s recebiam nota idêntica (1,0). Mas gap de 6–9s não é ênfase dramática: é troca de tomada, inserção de B-roll ou a outra pessoa falando. O índice estava premiando corte de cena como se fosse entrega enfática.
- **Solução adotada:** `pause_weight()` — a contribuição sobe até `max_pause` e **cai a zero** acima de `pause_ignore_above` (3s), em vez de saturar. Resultado imediato no mesmo vídeo: o topo passou a ser `"eu"` (energia 1,00), `"o"` (0,87), `"intacto!"` (0,70), `"tô"` (0,74) — todas da mesma frase, que é de fato a fala de impacto.
- **A virada:** o mesmo dado que era ruído virou o sinal mais útil do documento. Os gaps descartados marcam **onde a tomada recomeçou**. Adicionados `gap_before` e `take_boundary` (>= 3s) em cada segmento; num material de 3min17s isso detectou 6 fronteiras, exatamente onde a médica recomeçava o roteiro.
- **Descoberta de produto:** o bruto de consultório não é uma tomada — é o mesmo roteiro gravado 3–4 vezes, entremeado de conversa com a equipe (*"Amor, eu tô intacto!"*, *"Posso começar da mastopexia?"*, *"Só clica aí agora na tela."*). Separar tomada válida de conversa é **tarefa de linguagem, não de acústica**: no áudio a conversa é mais solta e mais alta que o texto decorado, então qualquer limiar acústico erra o alvo por construção. Critérios registrados em `.claude/skills/editar-por-voz/SKILL.md`.
- **Aprendizado:** antes de descartar um sinal por estar poluindo uma métrica, perguntar **para que outra pergunta ele é a resposta**. Aqui a mesma pausa respondia mal "isso foi enfático?" e otimamente "a tomada recomeçou aqui?".
- **Bug pego pelo próprio teste:** ao adicionar `gap_before`, o teste `test_scales_document_every_word_metric` quebrou — `scales` documentava tudo como métrica de palavra, e `gap_before` é de segmento. `VALUE_SCALES` passou a ser aninhado (`word`/`segment`), com um teste por nível. Um contrato auto-descritivo só vale se um teste garantir que ele não mente.
- **Estado:** `resolvido` (1294 testes verdes, validado nos dois vídeos reais).
---
### 2026-08-19 — FCP descartava TODO zoom gerado: `interp` em param vetorial (DTD-válido ≠ FCP-aceito)
- **Sintoma:** ao importar o FCPXML no Final Cut, o aviso `This param element was ignored because it does not support the interpolation attribute on its keyframes (.../adjust-transform[1]/param[1])`. O `<param name="scale">` inteiro era **descartado** — ou seja, o zoom simplesmente não existia no projeto importado, sem erro nem falha visível.
- **Causa raiz:** `add_zoom` escrevia `interp="ease"` em cada `<keyframe>` do parâmetro `scale`. O importador do FCP só aceita `interp` em parâmetros **escalares** (opacidade, volume); `scale` é vetorial (`value="1.25 1.25"`) e admite apenas `curve`.
- **Por que passou por tudo:** o DTD oficial da Apple declara `<!ATTLIST keyframe interp (linear|ease|easeIn|easeOut) "linear">` — ou seja, `interp` é **DTD-válido em qualquer keyframe**. A validação contra o DTD passava com 100% de sucesso, e o teste `test_add_zoom_creates_keyframed_transform` **afirmava** `interp == 'ease'`, travando o comportamento errado. Só a importação real no FCP revelou.
- **Solução adotada:** `curve="smooth"` no lugar de `interp` (o `curve` já é `smooth` por padrão no DTD, mas explícito documenta a intenção e protege contra mudança de default). Teste invertido: agora exige `curve == 'smooth'` **e** ausência de `interp`.
- **Atenção — não confundir com `timept`:** o `<timept>` do `timeMap` (usado em `change_speed`) aceita `interp` normalmente e **não** foi alterado. A restrição é do `<keyframe>` em param vetorial.
- **Aprendizado (o mais importante desta série):** **DTD-válido ≠ aceito pelo Final Cut.** O DTD descreve a gramática, não as regras semânticas do importador. Para qualquer construção nova de XML, validar contra o DTD é o piso, não o teto — só a importação real fecha a verificação. E um teste escrito a partir do próprio código gerado (em vez de um export real do FCP) apenas congela o erro: reforça o padrão já registrado em 2026-08-15 e 2026-08-18 — **extrair a verdade de um export real do FCP, nunca do que nós mesmos geramos.**
- **Estado:** `resolvido` no XML (1286 testes verdes, DTD 1.14 válido) — **pendente de nova confirmação de importação no FCP pelo usuário.**
---
### 2026-08-19 — Primeiro teste em material real: quatro bugs que só apareceram fora dos testes
Rodar o pipeline completo num depoimento real (17 min, 4K, 2320 palavras) expôs quatro
problemas que a suíte inteira, verde, não pegava. Todos vinham de premissas
que só material sintético sustentava.
**1. Teto de 100 MB rejeitava a mídia (14,7 GB).** `MAX_FILE_SIZE` existe para
documentos que lemos **inteiros na memória** (FCPXML, JSON) — onde um arquivo
gigante é o próprio ataque. Mídia nunca é carregada assim: ffmpeg e librosa
leem em fluxo, com timeout e limite de duração próprios. Criado
`MAX_MEDIA_FILE_SIZE` (32 GB) e `_validate_filepath(..., max_size=)`. Afetava
também o `detect_beats`, que já rejeitava qualquer WAV acima de ~10 minutos.
**2. librosa não lia `.mp4` — faltava a extração de áudio.** O PDF previa
"FFmpeg para extração do áudio" e eu pulei essa etapa, analisando o container
direto. Criado `decodable_audio()` em `voice_features.py`: passa adiante
arquivos de áudio nativos e extrai um WAV mono 16 kHz temporário via ffmpeg
para containers de vídeo (16 kHz basta — o teto de pitch é 1 kHz).
**3. O relatório MENTIA sobre o que rodou.** A tabela dizia "Acoustics: yes"
enquanto as duas extrações falhavam, porque reportava `features_capability()`
— se a *biblioteca está instalada* — e não se a *análise funcionou*. Agora o
JSON carrega um bloco `layers` com o que de fato executou. Fundamental porque
"fala monótona" e "acústica não carregou" deixam **os mesmos zeros** nos dados:
sem esse bloco, nem o usuário nem a IA que lê o arquivo conseguem distinguir.
**4. Limiar absoluto de ênfase não generaliza — trocado por percentil.** O
0,85 do PDF eu já havia recalibrado para 0,60 usando dados sintéticos; no
material real o índice **nunca passou de 0,544** (mediana 0,127), então 0,60
ainda selecionava nada. Corrigido de vez trocando o mecanismo: `select_peaks`
pega o **top N%** (padrão 2%), com um piso mínimo apenas como guarda para
áudio genuinamente plano. Qualquer corte fixo ou inunda um material ou zera
o outro; percentil entrega um punhado útil nos dois casos.
- **Aprendizado central:** limiar calibrado em dado sintético é chute. Duas
recalibrações erradas seguidas (0,85 → 0,60, ambas inúteis) só pararam
quando a régua virou **relativa à distribuição do próprio material**.
Sempre que um número governar seleção, prefira percentil a valor absoluto.
- **Observação de qualidade ainda aberta:** no top de ênfase real aparecem
palavras com energia baixíssima (`"No"`, energia 0,07) pontuando alto só
por virem depois de pausa longa. Em entrevista, pausa longa costuma ser o
entrevistador falando — não ênfase. O peso `pause_before` (0,15) com
saturação em 1,5s recompensa o sinal errado; avaliar reduzir o peso ou
ignorar pausas acima de ~3s.
- **Estado:** `resolvido` (1286 testes verdes; saída **validada contra o DTD
oficial FCPXML 1.14 da Apple**) — pendente de importação real no FCP.
---
### 2026-08-19 — Validador acusava desalinhamento de frame em todo projeto NTSC (falso positivo)
- **Sintoma:** o FCPXML gerado a partir de um projeto real 23,976fps acusava `Duration ... is not frame-aligned at 24fps` em clipes que estavam perfeitamente alinhados. Conferido na mão: `1200199/12000s ÷ 1001/24000s = 2398` frames exatos — inteiro, sem resto. O aviso do *arquivo original*, intocado, também era falso.
- **Causa raiz:** `_check_frame_alignment` fazia `fps_int = int(fps)` e multiplicava os segundos por esse inteiro. A 23,976 (`1001/24000s`), uma duração exatamente alinhada **não** é múltiplo inteiro de "24fps" — então todo projeto NTSC (23,976 / 29,97 / 59,94, ou seja, a maioria) era reportado como quebrado. Detalhe irônico: o docstring de `serialize_xml` já alertava para passar a taxa real "so NTSC projects don't get spurious warnings", mas o `int()` logo adiante destruía a correção.
- **Solução adotada:** o validador passou a ler o `frameDuration` exato do formato que a `<sequence>` referencia (`_document_frame_duration`) e a comparar com aritmética de `Fraction` — alinhado é quando `duração / frameDuration` tem denominador 1. A mensagem também passou a nomear a taxa real (`23.976fps`), não uma arredondada.
- **Aprendizado:** nunca converter timebase para inteiro/float para checar alinhamento — o projeto inteiro é construído sobre tempo racional justamente por isso (`TimeValue`), e a validação precisa seguir a mesma regra que a escrita. Falso positivo em validador é pior que ausência de validação: ensina o usuário a ignorar avisos, e aí o aviso verdadeiro passa batido.
- **Estado:** `resolvido` (3 testes de regressão em `TestNTSCFrameAlignment`, incluindo um que garante que desalinhamento **real** continua sendo detectado).
---
### 2026-08-18 — Divisão IA × sistema: a IA devolve DECISÕES, nunca XML; e tudo em tempo de origem
- **Contexto:** definido como a inteligência entra no editor automático. A ideia inicial era mandar o JSON para um serviço externo que devolveria o material já editado; evoluiu para fazer a decisão aqui dentro, com um skill versionado no repo (`.claude/skills/editar-por-voz/SKILL.md`).
- **Decisão 1 — o que a IA NÃO faz:** silêncio, vícios de linguagem, extração acústica, índice de ênfase, diarização e **geração de FCPXML** continuam determinísticos. Tudo que tem resposta objetiva (um limiar decide) não ganha nada indo para um modelo — só custo, latência e perda de reprodutibilidade. Geração de XML em particular é matemática de tempo racional frame a frame: modelo gerando XML produz arquivo sutilmente quebrado.
- **Decisão 2 — a IA devolve uma lista de ações, não mídia editada.** Contrato em `fcpxml/voice_actions.py` (`kind`/`start`/`end`/`params`/`reason`). Motivos: dá para **validar** antes de aplicar; é **reprodutível** (mesma lista → mesmo FCPXML); e o usuário **revisa** antes de qualquer coisa tocar a timeline. `parse_actions` trata a lista como entrada não confiável — uma linha malformada é reportada e pulada, nunca derruba a edição inteira.
- **Decisão 3 (a que evita a pior classe de bug) — todos os tempos em segundos da mídia ORIGINAL.** Cortes deslocam tudo que vem depois: se as decisões viessem em tempo pós-corte, cada destaque cairia silenciosamente no frame errado assim que um corte fosse adicionado. `resolve_actions`/`shift_after_cuts` resolvem o deslocamento na hora de aplicar, e ação que aponta para material removido é **descartada e reportada**, nunca deslizada para o conteúdo vizinho.
- **Bug pego pelo próprio relatório:** título e marcador não apareciam no XML. A causa era `str(TimeValue)` devolvendo o `__repr__` (`TimeValue(3/1s = 3.000s)`) em vez da string racional — o método certo é `to_fcpxml()`. Só foi visível na hora porque o handler reporta o que **não** conseguiu colocar, com a exceção real, em vez de aplicar em silêncio.
- **Aprendizado:** todo handler que aplica uma lista de operações deve relatar as três categorias — aplicadas, descartadas e rejeitadas. Um handler que só conta sucessos transforma bug em "não aconteceu nada" e some do radar. Nunca converter `TimeValue` para string com `str()`: sempre `to_fcpxml()`.
- **Estado:** `resolvido` (1276 testes verdes; aplicação validada contra FCPXML real, incluindo o `examples/sample.fcpxml`) — **pendente de confirmação de importação real no FCP pelo usuário.**
---
### 2026-08-18 — Limiar de ênfase de 0,85 do PDF era inalcançável: média ponderada não chega lá
- **Sintoma:** com a timeline de voz montada, o campo `peak_count` vinha **sempre 0**. Nem a palavra mais alta e mais aguda de um trecho de demonstração ("segurança", energia normalizada 1,0 e pico de tom) era marcada como candidata a punch-in.
- **Investigação:** medido o teto real da fórmula. `compute_emphasis` é uma **média ponderada** de cinco fatores normalizados (energia 0,30 / tom 0,25 / ritmo 0,20 / pausa 0,15 / duração 0,10). Para o resultado passar de 0,85 seria preciso que quase todos os cinco estivessem no máximo **simultaneamente** — o que a fala real não produz: uma palavra com pausa dramática antes dela quase por definição não tem desvio de ritmo alto. Valores medidos: todos os fatores no máximo = 1,00; pico realista (energia e tom máximos, pausa longa, palavra longa, ritmo normal) = **0,80**; pico comum = **0,66**.
- **Causa raiz:** o 0,85 veio literalmente da especificação do PDF (`SE emphasis > 0.85 ENTÃO aplicar punch-in`), que pressupunha outra normalização — provavelmente um índice de máximo, não de média. Copiar a constante sem conferir a distribuição da nossa fórmula tornou o recurso inerte.
- **Solução adotada:** padrão recalibrado para **0,60** (`DEFAULT_VOICE_ANALYSIS_CONFIG` em `fcpxml/model_manager.py`), com o porquê comentado no próprio código. O texto da tela e a descrição da tool passaram a dizer que picos reais ficam na faixa 0,55–0,80 — para que ninguém volte a subir o valor achando que "quanto maior, mais seletivo" sem saber onde fica o teto.
- **Aprendizado:** constante numérica herdada de especificação externa precisa ser **validada contra a distribuição real da fórmula implementada** antes de virar padrão. O sintoma aqui foi silencioso (nenhum erro, nenhum teste vermelho — só um recurso que nunca disparava), e só apareceu porque a saída de demonstração foi inspecionada com dados realistas. Vale gerar uma amostra de verdade e olhar os números sempre que um limiar governar um comportamento.
- **Estado:** `resolvido` (padrão 0,60 verificado: o mesmo trecho passou a marcar corretamente 1 pico).
---
### 2026-08-18 — Configurações de análise de voz: uma fonte de verdade só (config.json do backend), não UserDefaults
- **Contexto:** Fases 2–3 da arquitetura de análise de voz (features acústicas + índice de ênfase) e a tela de configurações pedida pelo usuário para regular limiares de energia, ênfase e emoção.
- **Decisão:** os parâmetros de análise ficam **só** em `~/.fcp-mcp-server/config.json` (via `model_manager.load/save_voice_analysis_config`), diferente do padrão `@AppStorage`/UserDefaults usado por `CaptionsView.swift` para estilo de legenda. Motivo: estilo de legenda é preferência de UI que só é lida na hora de montar os argumentos de uma chamada; já os limiares de análise são lidos **pelo próprio motor** (`handle_analyze_voice_features`) mesmo quando a análise é disparada fora do app (tool MCP direta, script). Duplicar em UserDefaults criaria duas verdades divergentes — a tela mostraria um valor e a análise usaria outro.
- **Onde:** `fcpxml/model_manager.py` (`DEFAULT_VOICE_ANALYSIS_CONFIG`, `load/save_voice_analysis_config`), comandos `voice_analysis`/`set_voice_analysis` em `admin/models_api.py`, tools MCP `get/save_voice_analysis_config`, tela `MacApp/Sources/VoiceAnalysisView.swift`.
- **Cuidado que rendeu teste:** `load_voice_analysis_config` precisa devolver uma **cópia** dos defaults — a primeira versão devolvia o dict aninhado `emphasis_weights` por referência, e quem mutasse o resultado corrompia o default do módulo para o resto do processo. Coberto por `test_defaults_are_not_shared_mutable_state`.
- **Aprendizado:** ao adicionar configuração nova, perguntar "quem lê esse valor?" — se for o motor Python, ele mora no config.json do backend; se for só a montagem de argumentos na UI, UserDefaults serve. E todo default composto (dict/lista) devolvido de um `load_*` precisa ser cópia, nunca a constante do módulo.
- **Estado:** `resolvido` (1201 testes verdes, lint zero erros, ciclo salvar→reler validado pelo bridge) — **a renderização visual da tela no app não pôde ser confirmada por captura de tela** (janela do app em outro Space); a aba foi confirmada via árvore de acessibilidade.
---
### 2026-08-18 — Diarização de locutor já existia pronta e testada, mas órfã (nenhuma tool MCP a expunha)
- **Contexto:** início da implementação da "Arquitetura de Análise de Voz para Editor Automático" (locutor, energia, pitch, ênfase, motor de regras → FCPXML), especificada num PDF trazido pelo usuário. Plano salvo em `~/.claude/plans/volumes-merongo-downloads-arquitetura-a-mossy-micali.md`.
- **Descoberta:** `fcpxml/diarize.py` (diarização via `pyannote/speaker-diarization-3.1`, com `diarization_capability`, `diarize`, `assign_speakers`, `build_speakers`) e `tests/test_diarize.py` já existiam completos e passando, mas nenhuma tool em `server.py` chamava esse módulo — código morto do ponto de vista de uso real. `model_manager.py` também já tinha `load_hf_token`/`save_hf_token` prontos para o token do HuggingFace exigido pelo pyannote.
- **Decisão de arquitetura:** usar pyannote (já é dependência declarada em `pyproject.toml` como extra `diarization`) para diarização bruta por turno, em vez de treinar/rodar SpeechBrain ECAPA-TDNN do zero como o PDF sugeria em primeiro lugar. ECAPA-TDNN fica reservado para uma fase futura (reconhecimento de pessoa cadastrada por cima dos turnos já diarizados), evitando duas libs pesadas resolvendo o mesmo problema.
- **Onde:** nova tool `diarize_media` em `server.py` (handler `handle_diarize_media`), reaproveitando `diarize.py` sem alterá-lo; cache em `_diarization.json` ao lado da mídia, seguindo exatamente o padrão de `_transcript.json`/`_beats.json` já usados por `transcribe_media`/`detect_beats`.
- **Aprendizado:** antes de implementar uma fase "do zero" a partir de uma spec externa, vale sempre grepar o `fcpxml/` por nomes prováveis (`diarize`, `speaker`, etc.) — pode já existir motor pronto e testado, só faltando a camada de exposição via MCP tool.
- **Estado:** `resolvido` (tool nova + testes, 1154 testes verdes, lint zero erros).
---
### 2026-08-18 — Terceira linha "colando" na linha de ênfase: o gap simétrico não bastava para o itálico
- **Sintoma:** no bloco de composição "phrase" (uma palavra de ênfase em itálico grande, cercada por linhas de corpo), a linha logo abaixo da ênfase aparecia quase tocando o texto — mesmo com o slider "Espaçamento entre linhas" da tela de legendas dinâmicas configurado.
- **Investigação:** reproduzido o cálculo de `compose_sentence` fora do FCP com a frase exata do usuário ("de" / "encontrar" / "roupa,") — o gap entre as caixas de tinta dava **exatamente 8pt nos dois lados** (acima e abaixo da ênfase), confirmando que o valor do slider chega corretamente até o layout (`CaptionsView.swift` → `admin/models_api.py` → `server.py` → `compose_sentence`). Não era bug de configuração não aplicada.
- **Causa raiz:** a caixa de tinta medida (`ink_extent`, `fcpxml/text_layout.py`) é vertical e simétrica, mas a inclinação itálica do Playfair Display faz os traços "vazarem" visualmente para baixo além do que a métrica vertical mede — então o mesmo gap numérico lê como mais apertado abaixo da linha de ênfase do que acima dela.
- **Onde:** `fcpxml/text_layout.py::compose_sentence` (função `pair_gap` nova) e constante `_EMPHASIS_ITALIC_CUSHION_RATIO`.
- **Solução adotada:** gap por par de linhas em vez de um valor único para todo o bloco — quando a linha anterior é a de ênfase, soma-se uma folga extra proporcional ao seu `font_size` (`ratio = 0.06`, ~14pt a 230pt) só naquele par; todos os outros pares continuam usando exatamente o `line_gap` do usuário. A folga entra tanto no teste de "cabe na banda" quanto na centralização da pilha, senão o bloco vazaria do box.height ou ficaria descentrado.
- **Aprendizado:** medir a caixa de tinta (ascendente/descendente reais) resolve colisão entre glifos retos, mas não captura o "peso visual" da inclinação itálica — para faces itálicas grandes ao lado de corpo reto, a folga simétrica por ink-box ainda pode ler como assimétrica no render final. Um cushion proporcional ao tamanho da fonte, aplicado só no lado que precisa, corrige sem inflar o espaçamento nos pares que já estavam certos.
- **Estado:** `resolvido` no cálculo (1150 testes verdes, valores conferidos numericamente) — **pendente de confirmação visual real no FCP pelo usuário**.
---
### 2026-08-18 — Build Out desligado libera toda a janela de "Per Object" para o Build In terminar de revelar
- **Sintoma:** em blocos de palavras curtos, a animação de entrada do título ("Text"/Basic Text template) às vezes cortava antes de terminar de revelar a palavra — o corte pro próximo bloco acontecia no meio do reveal.
- **Causa raiz:** `Apply Speed = "2 (Per Object)"` (já presente em `_TEXT_TITLE_PARAMS`) faz o FCP comprimir/esticar a animação **inteira** do template (build in + build out) para caber exatamente na duração real do `<title>`. Com as duas fases ativas, build in e build out disputam a mesma janela comprimida — em clipes curtos, build in não tinha tempo suficiente.
- **Onde:** `fcpxml/writer.py::_TEXT_TITLE_PARAMS` (`FCPXMLModifier._make_text_title_clip`).
- **Descoberta do `key`:** não havia como adivinhar — o usuário desmarcou manualmente "Build Out" no Inspector de um título "Text" isolado no FCP e exportou o FCPXML. O override só aparece no XML quando o valor difere do default do template (por isso um export sem a alteração real não mostra o `param` nenhum). Valor capturado: `<param name="Build Out" key="9999/10000/2/102" value="0"/>`.
- **Solução adotada:** `Build Out` adicionado como primeiro item de `_TEXT_TITLE_PARAMS`, sempre `"0"` (desligado) em todo título gerado. Não foi necessário nenhum parâmetro extra de velocidade — desligar o build out já entrega toda a janela "Per Object" comprimida ao build in, que é o efeito de "sempre acelerado" pedido pelo usuário.
- **Aprendizado:** pra descobrir o `key` de um checkbox/param publicado num template Motion, o export de calibração **precisa** ter o valor realmente alterado no Inspector antes de exportar — reexportar o projeto sem mexer em nada não revela nada (o FCP só escreve params que divergem do default). Reforça o padrão já registrado em 2026-08-15: nunca adivinhar `key`, sempre extrair de um export real.
- **Estado:** `resolvido` no XML (1 novo param verificado no writer) — **pendente de confirmação de importação real no FCP pelo usuário**.
---
### 2026-08-18 — Espaço de coordenadas do modelo de título: tamanho E posição
- **Sintoma (1ª metade):** o bloco caía exatamente onde o preview mostrava, mas
o texto renderizava cerca de **metade** do tamanho configurado — com 213pt a
ênfase deveria ocupar ~87% da largura do quadro e ocupava ~35%.
- **Sintoma (2ª metade, causado pela primeira correção):** ao dobrar só o
`fontSize`, o tamanho ficou certo e as **linhas passaram a se sobrepor** — o
bloco mantinha o espalhamento antigo com o dobro de letra dentro.
- **Causa raiz:** a calibração de tamanhos veio do export manual feito com o
modelo **"Essencial - Título"**, cujo espaço de coordenadas é o canvas de
pontos (metade do quadro). Esse modelo nunca renderizou quando gerado por nós
(entrada de 2026-08-17), então o writer passou a emitir o **"Basic Text >
Text" (Text.moti)** — cujo espaço é o **quadro inteiro** (2160×3840). Tudo o
que esse modelo lê está nesse espaço: `fontSize`, `kerning` **e** `Position`.
- **Onde:** `fcpxml/text_layout.py` (`TEXT_TEMPLATE_FONT_SCALE`,
`position_param`), `fcpxml/writer.py`, `fcpxml/models.py`
(`DynamicSubtitleConfig.text_scale`), `server.py`,
`MacApp/Sources/CaptionsView.swift`.
- **Tentativas que falharam:** (a) procurar a diferença nos params do título
(`Auto-Shrink`, margens, `Layout Method`) — todos idênticos ao export manual;
(b) **converter só o `fontSize`** — corrigiu o tamanho e quebrou o
espaçamento, que é o erro registrado aqui como aprendizado principal.
- **Solução adotada:** um único fator, `TEXT_TEMPLATE_FONT_SCALE = 2.0`,
aplicado ao `fontSize`, ao `kerning` **e** à `Position` na saída. O layout
continua medindo em pontos de canvas — toda constante calibrada depende
disso — e a conversão acontece só na emissão, que é a única forma de os dois
andarem juntos. Exposto como `text_scale`.
- **Aprendizado:** um espaço de coordenadas é indivisível. Converter metade das
grandezas que vivem nele é **pior** do que não converter nenhuma: sem
conversão o erro é uniforme e parece "só um ajuste de tamanho"; pela metade,
tipo e espaçamento se descolam e o defeito muda de cara. Ao trocar o modelo
de título, toda constante calibrada contra o modelo antigo vira suspeita — a
posição foi re-verificada em 2026-08-17 e o tamanho não, e o bug ficou
invisível porque "está no lugar certo" parece "está certo".
- **Estado:** `resolvido`
---
### 2026-08-18 — Preview das legendas dinâmicas desproporcional ao render do FCP
- **Sintoma:** o painel "Legendas Dinâmicas" mostrava um preview que não batia
com o resultado no Final Cut: linhas de apoio coladas nas bordas do quadro,
espaçamento entre linhas errado, a banda do bloco invisível e as cores
aplicadas de forma trocada. O formulário de controles também estava confuso,
com blocos de texto explicativo a cada slider.
- **Causa raiz:** `SubtitlePreviewView` era um desenho aproximado feito à mão
(VStack + Spacer, gap fixo de 14pt, canvas mapeado só na altura) e não
reproduzia `compose_sentence` de `fcpxml/text_layout.py`. Além disso, o app
mandava `inactive_color`, que em `granularity="phrase"` o backend **nunca
usa** — o preview pintava a ênfase com uma cor que o FCP ignoraria.
- **Onde:** `MacApp/Sources/SubtitlePreviewView.swift`,
`MacApp/Sources/CaptionsView.swift`, `server.py`
(`handle_generate_dynamic_subtitles`), `admin/models_api.py` (docstring).
- **Tentativas que falharam:** apenas re-escalar as fontes do preview — a
posição continuava errada, porque o desalinhamento vinha do *stagger* e do
gap, não do tamanho.
- **Solução adotada:** o preview passou a espelhar a geometria do backend —
canvas 1080×1920 pt (largura inclusa), margem lateral de 4%,
`REFERENCE_BLOCK_LINE_GAP` (8 pt), `REFERENCE_STAGGER_RATIO` (0.8) com lados
alternados a partir da esquerda, corpo em Bold, empilhamento sobre a tinta
(cap-height + descida) e compensação do centro do frame do `Text`. O
backend ganhou `emphasis_color` (padrão = `active_color`), e a UI foi
reagrupada em "Linhas de apoio" / "Palavra de ênfase" com sliders em
`LabeledContent` e explicação em tooltip.
- **Aprendizado:** um preview só é útil se for derivado das MESMAS constantes
do gerador. Quando o preview é redesenhado "de olho", ele vira uma segunda
fonte de verdade que diverge silenciosamente. E todo controle exposto na UI
precisa existir de fato no caminho de código que ele diz configurar.
- **Estado:** `resolvido`
---
### 2026-08-17 — Garantir que dois blocos nunca se sobreponham: empilhar pela TINTA real, não pela cap-height
- **Sintoma:** na composição progressiva, a cedilha de "começar" (Playfair
@@ -676,6 +1077,42 @@ Use o bloco abaixo como modelo. Uma entrada = um problema resolvido/reconhecido.
<!-- NOVAS ENTRADAS DEVEM SER ADICIONADAS ACIMA DESTA LINHA, SEMPRE NO TOPO
DA LISTA, PARA QUE A MAIS RECENTE FIQUE EM PRIMEIRO LUGAR. -->
### 2026-08-19 — Legendas dinâmicas geradas com `bold="0" fontFace="Bold"` não renderizam no FCP
- **Sintoma:** no corte real da Mastopexia, as legendas dinâmicas (composição
progressiva) não apareciam no Final Cut — só as primeiras linhas de cada
bloco surgiam e o restante sumia. O arquivo que o usuário re-exportou do FCP
("legendas dinamicas.fcpxmld") renderizava normalmente.
- **Causa raiz:** o corpo das legendas era definido como
`EDITORIAL_BODY_LOOK = WordLook(88, ..., face="Bold")` e o gravador emitia
`bold="0"` + `fontFace="Bold"` (pois `WordStyle.bold` é `False` por padrão).
Essa combinação é contraditória: no FCPXML negrito é o **atributo** `bold="1"`
(nunca um `fontFace="Bold"`), e itálico é `fontFace="... Italic"` **mais**
`italic="1"`. O FCP re-exporta `bold="1"` (sem `fontFace`) e
`fontFace="Medium Italic"` + `italic="1"`, provando o formato correto.
- **Onde:** `fcpxml/writer.py::_make_text_title_clip` (emissão do `text-style`);
o estilo em si em `fcpxml/models.py::EDITORIAL_BODY_LOOK`.
- **Tentativas que falharam:** corrigir manualmente o XML gerado trocando
`bold="0" fontFace="Bold"` por `bold="1"` — resolvia só aquele arquivo e o
bug reaparecia a cada geração. Também tentei "corrigir" os offsets dos
títulos (achando que estavam fora da realidade por estarem em coordenadas de
source) e quebrei o arquivo com timebases errados (24000, 30000) — os offsets
em source coords estavam corretos o tempo todo (ver entrada de 2026-08-17
sobre "anchored in SOURCE media coordinates").
- **Solução adotada:** em `_make_text_title_clip`, traduzir a face "bold" para
`bold="1"` sem `fontFace`; emitir `italic="1"` quando a face contém "italic";
e não mais emitir `bold="0"` junto de uma face. Agora a saída bate com a
re-exportação do FCP (corpo `bold="1"`, palavra-chave `fontFace` + `italic="1"`).
- **Aprendizado:** o FCPXML do template "Text" usa `bold` (atributo) para peso e
`fontFace`+`italic` para a face itálica; "Bold" não é um valor válido de
`fontFace`. Ao duvidar de um formato, confiar na re-exportação do FCP (saída
canônica) e nunca "corrigir" offsets/times que já seguem a convenção do
gerador. Também: comparar a saída gerada contra o FCP byte a byte por campo
(bold/fontFace/italic) antes de assumir o problema em outro lugar.
- **Estado:** `resolvido`
---
### 2026-08-14 — Início do registro de experiências
- **Sintoma:** não havia um local centralizado para registrar erros/estruturas
@@ -692,6 +1129,60 @@ Use o bloco abaixo como modelo. Uma entrada = um problema resolvido/reconhecido.
---
## 19 — `output_dir` aplicado só como cerca, nunca como destino
- **Data:** 2026-08-19
- **Sintoma:** toda chamada com `output_dir` diferente da pasta do arquivo de
entrada morria com `Output path escapes allowed directory`, apontando para um
caminho que a própria função tinha acabado de montar. Na prática o ajuste
"Pasta do projeto" do app só funcionava quando apontava para a pasta onde o
arquivo já ia cair sozinho — ou seja, nunca fazia nada.
- **Causa raiz:** em `_resolve_io_paths` (`server_tools/_shared.py`) o
`output_dir` virava apenas `anchor_dir` da validação, enquanto o nome do
arquivo continuava saindo de `generate_output_path(filepath, suffix)`, que
preserva o diretório da ENTRADA. Cerca em um lugar, destino em outro: o
caminho gerado ficava fora da própria cerca. Afetava os 18+ handlers de
escrita, não só as legendas onde o erro apareceu.
- **Solução adotada:** quando `output_dir` é passado, o destino padrão passa a
ser `<output_dir>/<nome derivado>`; sem ele, mantém-se o comportamento antigo
(ao lado da entrada). Um `output_path` explícito continua vencendo e continua
obrigado a ficar dentro da âncora. `build_voice_timeline` e
`refine_voice_timeline` passaram a aceitar e repassar `output_dir`; os
leitores procuram na pasta do projeto primeiro e caem para o lado da mídia,
para não perder timelines geradas antes da mudança.
- **Aprendizado:** validação e destino não podem ser derivados de fontes
diferentes. Quando um parâmetro tem dois papéis (permissão e endereço),
aplicar só um dos dois produz um erro que acusa o próprio código — e some da
vista porque o caso que funciona é justamente o caso trivial.
- **Estado:** `resolvido`
---
## 20 — Cadeia de processamento sem o passo que corta
- **Data:** 2026-08-19
- **Sintoma:** o encadeamento do app ia de `analyze_voice` direto para
`remove_silences`/legendas. Dava para medir a voz e legendar o resultado, mas
não para aplicar as decisões de edição — o corte por voz tinha que ser rodado
à mão, fora do app, e era fácil parar no primeiro passo achando que o arquivo
estava pronto.
- **Causa raiz:** `apply_voice_actions` existia como handler MCP mas nunca foi
exposto na ponte `admin/models_api.py`, então o batch não tinha como chamá-lo.
- **Solução adotada:** comando `apply_voice_actions` na ponte (aceita
`actions_path` apontando para o JSON de decisões, com ou sem o embrulho
`{"actions": [...]}`), e a etapa correspondente no batch do app, posicionada
logo após a análise e **antes** de qualquer passo que faça ripple — os
zooms/textos/marcadores são posicionados deslocando a partir da própria lista
de cortes, então rodar depois de outro corte os joga no frame errado sem erro
visível. O botão fica bloqueado se a etapa estiver ligada sem arquivo
escolhido, para a cadeia não quebrar no meio.
- **Aprendizado:** um passo que só existe como ferramenta MCP não existe para
quem usa o app. Vale conferir se toda etapa documentada no fluxo tem
representação na cadeia que o usuário de fato executa.
- **Estado:** `resolvido`
---
## Resumo rápido (índice)
| # | Data | Problema | Estado |
@@ -704,5 +1195,15 @@ Use o bloco abaixo como modelo. Uma entrada = um problema resolvido/reconhecido.
| 8 | 2026-08-17 | Importação recusada: `id` de `<text-style-def>` derivado do texto (acentos/espaços/dígito inicial) não é XML Name válido | `resolvido` |
| 9 | 2026-08-17 | Legendas palavra a palavra centradas em vez da composição progressiva diagramada (bloco por trecho, palavra-chave em display italic) | `resolvido` |
| 10 | 2026-08-17 | Cedilha/acentos da display italic invadindo a linha vizinha: empilhamento passou a usar a tinta real por classe de glifo | `resolvido` |
| 11 | 2026-08-18 | Preview das legendas dinâmicas desproporcional ao render do FCP (stagger/gap/canvas divergentes) e `inactive_color` exposto sem efeito | `resolvido` |
| 12 | 2026-08-18 | Espaço de coordenadas do modelo "Text": `fontSize`, `kerning` e `Position` no espaço do quadro — converter só o tamanho descolou o espaçamento | `resolvido` |
| 13 | 2026-08-19 | Reanálise de ênfase implementada no Engine mas sem ferramenta MCP — Fase 4 da skill era inexecutável | `resolvido` |
| 14 | 2026-08-19 | Offset sistemático de ~0,4s no timing por palavra (faster-whisper sem alinhamento forçado) — corrigido manualmente no teste, WhisperX pendente | `parcialmente resolvido` |
| 15 | 2026-08-19 | `add_zoom` perdia o enquadramento real (voltava a 100%) quando dois zooms caiam no mesmo clipe pós-corte; agora empilha ou substitui conforme as janelas se sobrepõem | `resolvido` |
| 16 | 2026-08-19 | `validate_subtitle_layout` acusava colisão severa em títulos que só se tocam na borda, por não-associatividade de float; 7 de 8 colisões reportadas no teste real eram falso positivo | `resolvido` |
| 17 | 2026-08-19 | Linha de ênfase das legendas dinâmicas sem limite de largura — palavra longa/maiúscula estourava o frame inteiro; auto-fit encolhe até caber, nunca abaixo do corpo | `resolvido` |
| 18 | 2026-08-19 | Legendas dinâmicas geradas com `bold="0" fontFace="Bold"` não renderizam no FCP — negrito deve ser `bold="1"` (atributo) e itálico `fontFace`+`italic="1"` | `resolvido` |
| 19 | 2026-08-19 | `output_dir` usado só como cerca de validação e nunca como destino — toda chamada entre pastas falhava acusando o caminho que ela mesma gerou | `resolvido` |
| 20 | 2026-08-19 | `apply_voice_actions` ausente da ponte e do encadeamento do app — dava para analisar e legendar, não para cortar | `resolvido` |
> Mantenha o índice acima sempre sincronizado com as entradas mais recentes.
+8 -2
View File
@@ -14,6 +14,7 @@ struct GArtApp: App {
enum ActiveTab: Hashable {
case project
case captions
case voiceAnalysis
case models
case about
}
@@ -26,8 +27,10 @@ struct ContentView: View {
List(selection: $activeTab) {
Label("Projeto", systemImage: "film")
.tag(ActiveTab.project)
Label("Legendas", systemImage: "captions.bubble")
Label("Legendas Dinâmicas", systemImage: "captions.bubble")
.tag(ActiveTab.captions)
Label("Análise de Voz", systemImage: "waveform")
.tag(ActiveTab.voiceAnalysis)
Label("Modelos", systemImage: "tray.and.arrow.down")
.tag(ActiveTab.models)
Label("Sobre", systemImage: "info.circle")
@@ -42,7 +45,10 @@ struct ContentView: View {
.navigationTitle("Projeto")
case .captions:
CaptionsView().id(UUID())
.navigationTitle("Legendas")
.navigationTitle("Legendas Dinâmicas")
case .voiceAnalysis:
VoiceAnalysisView().id(UUID())
.navigationTitle("Análise de Voz")
case .models:
ModelDownloadView().id(UUID())
.navigationTitle("Modelos")
+374 -153
View File
@@ -1,195 +1,416 @@
import SwiftUI
import UniformTypeIdentifiers
/// Guia "Legendas" — configura e gera legendas dinâmicas como clipes de título
/// editáveis no Final Cut Pro (template "Essencial - Título"). Cada palavra
/// vira um clipe de título posicionado: as palavras da frase vão surgindo
/// conforme são faladas, se acumulam num bloco centralizado, e somem todas
/// juntas no fim da frase. Sem Compound Clip.
/// Guia "Legendas Dinâmicas" — apenas configuração de estilo. Nenhum
/// processamento acontece aqui: a geração das legendas roda na aba
/// "Projeto".
///
/// Os valores ficam em `~/.fcp-mcp-server/config.json` (via
/// `model_manager.save_dynamic_subtitle_config`), os mesmos lidos por
/// `generate_dynamic_subtitles` como padrão — não em UserDefaults/
/// `@AppStorage`, para que a geração renderize exatamente o que esta tela
/// mostra, e para que qualquer chamador (app, MCP, uma sessão de IA) veja o
/// mesmo estilo sem precisar repassar os 11 campos a cada chamada. Mesmo
/// desenho de `VoiceAnalysisView`.
///
/// Layout em duas colunas: à esquerda, um **preview vertical 9:16** fixo que
/// reproduz a composição real (posição, escala, quebra de linhas e o
/// escalonamento das linhas de apoio); à direita, os controles agrupados por
/// assunto.
struct CaptionsView: View {
@State private var projectPath: String?
@State private var clipName = ""
// Posicionamento do bloco
@State private var bandHeight: Double = 0.22
@State private var blockCenterY: Double = -167
// Estilo
@State private var font = "Helvetica Neue"
@State private var fontSize: Double = 90
@State private var activeColor = Color.white
@State private var inactiveColor = Color(white: 0.7)
@State private var isGenerating = false
@State private var resultPath: String?
@State private var config = CaptionStyleConfig.defaults
@State private var isLoading = true
@State private var errorMessage: String?
// Frase de amostra do preview — só conveniência local, não afeta a
// geração real (que usa as palavras de verdade da transcrição), então
// continua em @AppStorage em vez do config compartilhado.
@AppStorage("capSampleBefore") private var sampleBefore = "que vão"
@AppStorage("capSampleEmphasis") private var sampleEmphasis = "melhorar"
@AppStorage("capSampleAfter") private var sampleAfter = "sua legenda"
@AppStorage("capShowGuides") private var showsGuides = true
private let fontChoices = [
"Helvetica Neue", "Helvetica", "Arial", "Avenir Next",
"Futura", "SF Pro Display", "Georgia", "Impact",
]
private let emphasisFontChoices = [
"Playfair Display", "Georgia", "Didot", "Futura",
"Avenir Next", "Times New Roman", "Helvetica Neue", "Impact",
]
private let emphasisFaceChoices = [
"Medium Italic", "Italic", "Bold Italic", "Bold", "Regular", "Light Italic",
]
/// Salva no arquivo a cada mudança e devolve um Binding, para os controles
/// continuarem simples (`$config.x` viraria só memória local).
private func bound<T>(_ keyPath: WritableKeyPath<CaptionStyleConfig, T>) -> Binding<T> {
Binding(
get: { config[keyPath: keyPath] },
set: { config[keyPath: keyPath] = $0; save() }
)
}
private func colorBound(_ keyPath: WritableKeyPath<CaptionStyleConfig, String>) -> Binding<Color> {
Binding(
get: { Color(rgbaString: config[keyPath: keyPath]) },
set: { config[keyPath: keyPath] = $0.fcpxmlColorString; save() }
)
}
var body: some View {
HSplitView {
previewColumn
.frame(minWidth: 250, idealWidth: 300, maxWidth: 380)
controlsColumn
.frame(minWidth: 380, idealWidth: 460)
}
.task { await load() }
}
// MARK: - Coluna da pré-visualização
private var previewColumn: some View {
VStack(alignment: .leading, spacing: 14) {
HStack {
Label("Pré-visualização", systemImage: "rectangle.on.rectangle.angled")
.font(.headline)
Spacer()
Toggle("Guias", isOn: $showsGuides)
.toggleStyle(.switch)
.controlSize(.mini)
.labelsHidden()
.help("Mostra a faixa do bloco e a linha de centro do quadro.")
}
SubtitlePreviewView(
bodyFont: config.font,
bodySize: config.fontSize,
bodyColor: Color(rgbaString: config.activeColor),
emphasisFont: config.emphasisFont,
emphasisFace: config.emphasisFace,
emphasisSize: config.emphasisSize,
emphasisColor: Color(rgbaString: config.emphasisColor),
bandHeight: config.bandHeight,
blockCenterY: config.blockCenterY,
lineGap: config.lineGap,
beforeText: $sampleBefore,
emphasisText: $sampleEmphasis,
afterText: $sampleAfter,
showsGuides: showsGuides
)
.frame(maxWidth: .infinity, maxHeight: 420)
GroupBox("Frase de amostra") {
VStack(spacing: 6) {
LabeledContent("Antes") {
TextField("", text: $sampleBefore).textFieldStyle(.roundedBorder)
}
LabeledContent("Ênfase") {
TextField("", text: $sampleEmphasis).textFieldStyle(.roundedBorder)
}
LabeledContent("Depois") {
TextField("", text: $sampleAfter).textFieldStyle(.roundedBorder)
}
}
.padding(.vertical, 4)
}
if emphasisOverflows {
Label(
"A palavra de ênfase é mais larga que o quadro nesse tamanho — reduza o tamanho da ênfase para não sair cortada.",
systemImage: "exclamationmark.triangle.fill"
)
.font(.caption)
.foregroundStyle(.orange)
.fixedSize(horizontal: false, vertical: true)
}
Text("Quadro vertical 9:16 na mesma geometria do Final Cut: a ênfase fica centrada e as linhas de apoio se deslocam para os lados alternados. Larguras são estimadas — a posição e a escala são reais.")
.font(.caption2)
.foregroundStyle(.secondary)
.fixedSize(horizontal: false, vertical: true)
Spacer(minLength: 0)
}
.padding(20)
.frame(maxHeight: .infinity, alignment: .top)
}
// MARK: - Coluna de controles
private var controlsColumn: some View {
Form {
Section("Projeto do Final Cut Pro") {
HStack {
Text(projectName)
.foregroundStyle(projectPath == nil ? .secondary : .primary)
.lineLimit(1)
Spacer()
Button("Escolher…") { pickProjectFile() }
if isLoading {
Section {
ProgressView().controlSize(.small)
.frame(maxWidth: .infinity, alignment: .center)
}
if let projectPath {
HStack {
Image(systemName: "doc.text").foregroundStyle(.secondary)
Text(projectPath).font(.caption).foregroundStyle(.secondary).lineLimit(1)
}
}
TextField("Nome do clipe (opcional — vazio = todos os clipes)", text: $clipName)
.textFieldStyle(.roundedBorder)
} else {
positionSection
bodySection
emphasisSection
calibrationSection
}
Section {
VStack(alignment: .leading, spacing: 6) {
HStack {
Text("Altura do bloco").font(.callout.weight(.medium))
Spacer()
Text("\(Int(bandHeight * 100))%").font(.caption).foregroundStyle(.secondary).monospacedDigit()
}
Slider(value: $bandHeight, in: 0.10...0.50, step: 0.01)
Text("Quanto da altura do quadro a frase pode ocupar antes de quebrar em outro bloco. Maior = mais palavras juntas na tela.")
.font(.caption2).foregroundStyle(.secondary)
}
.padding(.vertical, 4)
VStack(alignment: .leading, spacing: 6) {
HStack {
Text("Altura na tela").font(.callout.weight(.medium))
Spacer()
Text("\(Int(blockCenterY))").font(.caption).foregroundStyle(.secondary).monospacedDigit()
}
Slider(value: $blockCenterY, in: -700...300, step: 1)
Text("Posição vertical do bloco. 0 é o centro do quadro; valores negativos descem.")
.font(.caption2).foregroundStyle(.secondary)
}
.padding(.vertical, 4)
} header: {
Text("Posicionamento do Bloco")
} footer: {
Text("Cada palavra vira um clipe de título solto na timeline (sem Compound Clip), posicionado para não sobrepor as outras palavras da frase.")
.font(.caption2).foregroundStyle(.secondary)
}
Section("Estilo do Texto") {
Picker("Fonte", selection: $font) {
ForEach(fontChoices, id: \.self) { Text($0).tag($0) }
}
VStack(alignment: .leading, spacing: 6) {
HStack {
Text("Tamanho da fonte").font(.callout.weight(.medium))
Spacer()
Text("\(Int(fontSize))pt").font(.caption).foregroundStyle(.secondary).monospacedDigit()
}
Slider(value: $fontSize, in: 20...200, step: 1)
}
.padding(.vertical, 4)
ColorPicker("Cor A", selection: $activeColor, supportsOpacity: true)
ColorPicker("Cor B", selection: $inactiveColor, supportsOpacity: true)
Text("Tamanho, cor e estilo variam por palavra seguindo um ritmo fixo, calibrado a partir de um projeto real do Final Cut. As cores acima entram nesse ritmo.")
.font(.caption2).foregroundStyle(.secondary)
}
Section {
Button {
generate()
} label: {
if isGenerating { ProgressView().controlSize(.small) }
Label("Gerar Legendas Dinâmicas", systemImage: "captions.bubble.fill")
}
.disabled(isGenerating || projectPath == nil)
.frame(maxWidth: .infinity)
.buttonStyle(.borderedProminent)
if let errorMessage {
if let errorMessage {
Section {
Label(errorMessage, systemImage: "exclamationmark.triangle.fill")
.foregroundStyle(.red)
}
if let resultPath {
Text(resultPath).font(.caption).lineLimit(1)
HStack {
Button("Abrir no Final Cut Pro") { NSWorkspace.shared.open(URL(fileURLWithPath: resultPath)) }
Button("Mostrar no Finder") {
NSWorkspace.shared.activateFileViewerSelecting([URL(fileURLWithPath: resultPath)])
}
}
}
} footer: {
Text("Usa a transcrição local (Whisper) já feita na aba Projeto/Transcrição. Se ainda não houver transcrição salva, ela é gerada automaticamente.")
.font(.caption2).foregroundStyle(.secondary)
}
}
.formStyle(.grouped)
}
private var projectName: String {
guard let p = projectPath else { return "Nenhum projeto selecionado" }
return URL(fileURLWithPath: p).lastPathComponent
private var positionSection: some View {
Section {
slider(
"Altura do bloco",
value: bound(\.bandHeight), in: 0.10...0.50, step: 0.01,
readout: "\(Int(config.bandHeight * 100))%",
help: "Quanto da altura do quadro a frase pode ocupar antes de quebrar em outro bloco."
)
slider(
"Posição vertical",
value: bound(\.blockCenterY), in: -700...300, step: 1,
readout: "\(Int(config.blockCenterY))",
help: "0 é o centro do quadro; negativos descem, positivos sobem."
)
slider(
"Espaço entre linhas",
value: bound(\.lineGap), in: -80...120, step: 1,
readout: "\(Int(config.lineGap))pt",
help: "Distância entre uma linha e a outra, além das próprias letras. Negativo sobrepõe."
)
} header: {
Text("Posicionamento no Quadro")
} footer: {
Text("Medidas em pontos do canvas de referência 2160×3840. O espaço entre linhas é contado a partir da tinta real de cada linha — em 0 elas se encostam, e em negativo uma entra na outra.")
.font(.caption2).foregroundStyle(.secondary)
}
}
private func pickProjectFile() {
let panel = NSOpenPanel()
panel.canChooseFiles = true
panel.canChooseDirectories = false
panel.allowsMultipleSelection = false
panel.prompt = "Selecionar"
panel.message = "Selecione o arquivo de projeto (.fcpxml) exportado pelo Final Cut Pro."
if panel.runModal() == .OK, let url = panel.url {
let ext = url.pathExtension.lowercased()
if ext == "fcpxml" || ext == "xml" || ext == "fcpxmld" {
projectPath = url.path
} else {
errorMessage = "Selecione um arquivo .fcpxml ou .xml do Final Cut Pro."
private var bodySection: some View {
Section("Linhas de apoio") {
Picker("Fonte", selection: bound(\.font)) {
ForEach(fontChoices, id: \.self) { Text($0).tag($0) }
}
slider(
"Tamanho",
value: bound(\.fontSize), in: 20...200, step: 1,
readout: "\(Int(config.fontSize))pt",
help: "Tamanho das linhas que acompanham a palavra-chave."
)
ColorPicker("Cor", selection: colorBound(\.activeColor), supportsOpacity: true)
}
}
private var emphasisSection: some View {
Section("Palavra de ênfase") {
Picker("Fonte", selection: bound(\.emphasisFont)) {
ForEach(emphasisFontChoices, id: \.self) { Text($0).tag($0) }
}
Picker("Estilo", selection: bound(\.emphasisFace)) {
ForEach(emphasisFaceChoices, id: \.self) { Text($0).tag($0) }
}
slider(
"Tamanho",
value: bound(\.emphasisSize), in: 60...400, step: 1,
readout: "\(Int(config.emphasisSize))pt",
help: "A palavra-chave da frase, sozinha em sua linha e maior."
)
ColorPicker("Cor", selection: colorBound(\.emphasisColor), supportsOpacity: true)
}
}
private var calibrationSection: some View {
Section {
slider(
"Escala no Final Cut",
value: bound(\.textScale), in: 0.5...3.0, step: 0.05,
readout: String(format: "%.2f×", config.textScale),
help: "Converte o tamanho escolhido para o espaço em que o modelo de título do Final Cut desenha o texto."
)
HStack {
Spacer()
Button("Restaurar padrão") {
config = CaptionStyleConfig.defaults
save()
}
.buttonStyle(.link)
}
} header: {
Text("Calibração")
} footer: {
VStack(alignment: .leading, spacing: 6) {
Text("O modelo \"Text\" do Final Cut trabalha no espaço do quadro inteiro (2160×3840), enquanto o layout é calculado em pontos — metade disso. Por isso o padrão é 2,00×, aplicado ao tamanho E à posição juntos. Ajuste só se o seu modelo usar outra proporção.")
Text("Este estilo fica salvo em ~/.fcp-mcp-server/config.json e é usado automaticamente ao gerar legendas dinâmicas — nada é processado aqui.")
}
.font(.caption2).foregroundStyle(.secondary)
}
}
/// Slider com rótulo, leitura numérica alinhada e explicação em tooltip —
/// mantém as linhas do formulário com a mesma altura.
private func slider(
_ title: String,
value: Binding<Double>,
in range: ClosedRange<Double>,
step: Double,
readout: String,
help: String
) -> some View {
LabeledContent(title) {
HStack(spacing: 10) {
Slider(value: value, in: range, step: step)
Text(readout)
.font(.caption).monospacedDigit()
.foregroundStyle(.secondary)
.frame(width: 46, alignment: .trailing)
}
}
.help(help)
}
private func generate() {
guard let projectPath else { return }
isGenerating = true
errorMessage = nil
resultPath = nil
/// Largura útil do canvas de referência (1080 pt menos 4% de cada lado).
private static let usableCanvasWidth: CGFloat = 1080 * 0.92
var args: [String: Any] = [
"path": projectPath,
"band_height": bandHeight,
"block_center_y": blockCenterY,
"font": font,
"font_size": Int(fontSize),
"active_color": activeColor.fcpxmlColorString,
"inactive_color": inactiveColor.fcpxmlColorString,
]
let trimmedClip = clipName.trimmingCharacters(in: .whitespaces)
if !trimmedClip.isEmpty {
args["clip_name"] = trimmedClip
/// A palavra de ênfase da amostra passa da largura do quadro no tamanho
/// escolhido? Medida no mesmo canvas que o backend usa.
private var emphasisOverflows: Bool {
let text = sampleEmphasis.trimmingCharacters(in: .whitespaces)
guard !text.isEmpty else { return false }
var descriptor = NSFontDescriptor(fontAttributes: [.family: config.emphasisFont])
if !config.emphasisFace.isEmpty {
descriptor = descriptor.addingAttributes([.face: config.emphasisFace])
}
let nsFont = NSFont(descriptor: descriptor, size: CGFloat(config.emphasisSize))
?? NSFont.systemFont(ofSize: CGFloat(config.emphasisSize))
let width = (text as NSString).size(withAttributes: [.font: nsFont]).width
return width > Self.usableCanvasWidth
}
PythonBridge.call(command: "generate_dynamic_subtitles", arguments: args) { result, err in
DispatchQueue.main.async {
isGenerating = false
if result?["ok"] as? Bool == true {
resultPath = result?["path"] as? String
} else {
errorMessage = result?["error"] as? String ?? err ?? "Falha ao gerar legendas dinâmicas."
// MARK: - Backend
@MainActor
private func load() async {
await withCheckedContinuation { continuation in
PythonBridge.call(command: "dynamic_subtitle_config") { result, error in
DispatchQueue.main.async {
if let result {
config = CaptionStyleConfig(from: result)
} else if let error {
errorMessage = error
}
isLoading = false
continuation.resume()
}
}
}
}
private func save() {
PythonBridge.call(command: "set_dynamic_subtitle_config", arguments: config.arguments()) { _, error in
DispatchQueue.main.async { errorMessage = error }
}
}
}
/// O estilo das legendas dinâmicas, no formato que a tela edita e o bridge
/// (`admin/models_api.py` → `set_dynamic_subtitle_config`) persiste.
struct CaptionStyleConfig {
var bandHeight: Double
var blockCenterY: Double
var lineGap: Double
var font: String
var fontSize: Double
var emphasisFont: String
var emphasisFace: String
var emphasisSize: Double
var activeColor: String
var emphasisColor: String
var textScale: Double
static let defaults = CaptionStyleConfig(
bandHeight: 0.22,
blockCenterY: -167,
lineGap: 8,
font: "Helvetica Neue",
fontSize: 104,
emphasisFont: "Playfair Display",
emphasisFace: "Medium Italic",
emphasisSize: 265,
activeColor: "1 1 1 1",
emphasisColor: "1 1 1 1",
textScale: 2.0
)
/// Lê a resposta do bridge, caindo no padrão para qualquer campo ausente.
init(from json: [String: Any]) {
let d = CaptionStyleConfig.defaults
self.init(
bandHeight: json["band_height"] as? Double ?? d.bandHeight,
blockCenterY: json["block_center_y"] as? Double ?? d.blockCenterY,
lineGap: json["line_gap"] as? Double ?? d.lineGap,
font: json["font"] as? String ?? d.font,
fontSize: (json["font_size"] as? NSNumber)?.doubleValue ?? d.fontSize,
emphasisFont: json["emphasis_font"] as? String ?? d.emphasisFont,
emphasisFace: json["emphasis_face"] as? String ?? d.emphasisFace,
emphasisSize: (json["emphasis_size"] as? NSNumber)?.doubleValue ?? d.emphasisSize,
activeColor: json["active_color"] as? String ?? d.activeColor,
emphasisColor: json["emphasis_color"] as? String ?? d.emphasisColor,
textScale: json["text_scale"] as? Double ?? d.textScale
)
}
init(
bandHeight: Double, blockCenterY: Double, lineGap: Double,
font: String, fontSize: Double,
emphasisFont: String, emphasisFace: String, emphasisSize: Double,
activeColor: String, emphasisColor: String, textScale: Double
) {
self.bandHeight = bandHeight
self.blockCenterY = blockCenterY
self.lineGap = lineGap
self.font = font
self.fontSize = fontSize
self.emphasisFont = emphasisFont
self.emphasisFace = emphasisFace
self.emphasisSize = emphasisSize
self.activeColor = activeColor
self.emphasisColor = emphasisColor
self.textScale = textScale
}
func arguments() -> [String: Any] {
[
"band_height": bandHeight,
"block_center_y": blockCenterY,
"line_gap": lineGap,
"font": font,
"font_size": Int(fontSize),
"emphasis_font": emphasisFont,
"emphasis_face": emphasisFace,
"emphasis_size": Int(emphasisSize),
"active_color": activeColor,
"emphasis_color": emphasisColor,
"text_scale": textScale,
]
}
}
extension Color {
/// Parses an FCPXML "R G B A" space-separated 0-1 string into a Color.
init(rgbaString: String) {
let parts = rgbaString.split(separator: " ").compactMap { Double($0) }
guard parts.count >= 3 else { self = .white; return }
let r = parts[0], g = parts[1], b = parts[2], a = parts.count >= 4 ? parts[3] : 1.0
self.init(.sRGB, red: r, green: g, blue: b, opacity: a)
}
/// Converts to FCPXML's "R G B A" space-separated 0-1 string (sRGB).
var fcpxmlColorString: String {
let ns = NSColor(self).usingColorSpace(.sRGB) ?? NSColor(self)
+52 -4
View File
@@ -105,25 +105,64 @@ struct ModelDownloadView: View {
.font(.caption)
.foregroundStyle(catalog?.diarization == true ? Color.secondary : Color.orange)
}
SecureField("Token HuggingFace (diarização)", text: $hfTokenText)
.textFieldStyle(.roundedBorder)
.help("Token com acesso aos modelos gated pyannote (segmentation + speaker-diarization)")
tokenField
HStack {
TextField("Nº de participantes (vazio = automático)", text: $numSpeakersText)
.textFieldStyle(.roundedBorder)
Button("Salvar") { saveDiarization() }
.disabled(isLoading)
}
setupSteps
}
} header: {
Text("Diarização (participantes)")
} footer: {
Text("Opcional. Sem token, cada fala é atribuída a um único participante padrão.")
Text("Opcional. Sem token, cada fala é atribuída a um único participante padrão. O token só é usado para baixar o modelo uma vez — depois disso a análise roda offline, nesta máquina.")
.font(.caption)
.foregroundStyle(.secondary)
}
}
/// Campo do token. Quando já existe um salvo, mostra o estado em vez de um
/// campo vazio ambíguo — e oferece a remoção, já que salvar vazio não apaga.
@ViewBuilder
private var tokenField: some View {
if catalog?.hfTokenSet == true {
HStack {
Label("Token salvo nesta máquina", systemImage: "key.fill")
.font(.caption)
.foregroundStyle(.secondary)
Spacer()
Button("Remover") { removeToken() }
.buttonStyle(.link)
.disabled(isLoading)
}
}
SecureField(
catalog?.hfTokenSet == true ? "Substituir token…" : "Token HuggingFace (diarização)",
text: $hfTokenText
)
.textFieldStyle(.roundedBorder)
.help("Token com acesso aos modelos gated pyannote (segmentation + speaker-diarization)")
}
/// O token sozinho não basta: os dois modelos pyannote são "gated" e exigem
/// aceitar os termos na conta antes do download funcionar.
private var setupSteps: some View {
VStack(alignment: .leading, spacing: 4) {
Text("Para ativar, uma vez só:")
.font(.caption).bold()
.foregroundStyle(.secondary)
Link("1. Aceitar os termos do modelo de segmentação",
destination: URL(string: "https://huggingface.co/pyannote/segmentation-3.0")!)
Link("2. Aceitar os termos do modelo de diarização",
destination: URL(string: "https://huggingface.co/pyannote/speaker-diarization-3.1")!)
Link("3. Criar um token de acesso e colar acima",
destination: URL(string: "https://huggingface.co/settings/tokens")!)
}
.font(.caption)
}
private func saveDiarization() {
var args: [String: Any] = ["num_speakers": numSpeakersText]
if !hfTokenText.isEmpty {
@@ -138,6 +177,15 @@ struct ModelDownloadView: View {
}
}
private func removeToken() {
PythonBridge.call(command: "set_diarization", arguments: ["token": ""]) { _, _ in
DispatchQueue.main.async {
hfTokenText = ""
Task { await refresh() }
}
}
}
// MARK: - Storage
private var storageSection: some View {
+49 -7
View File
@@ -101,6 +101,25 @@ struct ProjectView: View {
}
}
.formStyle(.grouped)
.task { restoreLastProject() }
}
/// Reabre o último projeto salvo em ~/.fcp-mcp-server/config.json, para o
/// app voltar onde parou em vez de pedir o arquivo de novo a cada abertura.
/// É aqui que a restauração precisa morar: a TranscriptionView embutida só
/// existe depois que há um projeto carregado, então ela não consegue se
/// restaurar sozinha. Caminhos que sumiram do disco voltam vazios do
/// Python, e nesse caso a tela abre limpa como antes.
private func restoreLastProject() {
guard project == nil else { return }
PythonBridge.call(command: "project_config") { result, _ in
DispatchQueue.main.async {
guard project == nil,
let result, result["ok"] as? Bool == true,
let file = result["file"] as? String, !file.isEmpty else { return }
inspect(file)
}
}
}
/// Presents an NSOpenPanel configured to select a single file (not a
@@ -156,6 +175,23 @@ struct ProjectView: View {
private func handleDrop(_ providers: [NSItemProvider]) -> Bool {
guard let provider = providers.first else { return false }
// `loadObject(ofClass: URL.self)` is Foundation's own bridge for a
// dropped file URL and handles every representation Finder/FCP may
// hand back (NSURL via secure coding, a bookmark, a plain path).
// The item can ALSO be manually pulled as raw `Data` — but that only
// works if the bytes are exactly `URL.dataRepresentation`'s format
// (a UTF-8 absolute-string encoding), which a `.fcpxmld` *bundle*
// (a package macOS treats as a directory, not a plain file) does not
// always arrive as: the decode silently returns nil, so the drop
// looks like it does nothing. Try the robust path first.
if provider.canLoadObject(ofClass: URL.self) {
_ = provider.loadObject(ofClass: URL.self) { url, _ in
self.handleDroppedURL(url)
}
return true
}
provider.loadItem(forTypeIdentifier: UTType.fileURL.identifier, options: nil) { item, _ in
var url: URL?
if let data = item as? Data {
@@ -163,17 +199,21 @@ struct ProjectView: View {
} else if let u = item as? URL {
url = u
}
if let url, isAccepted(url) {
DispatchQueue.main.async { inspect(url.path) }
} else {
DispatchQueue.main.async {
errorMessage = "Este arquivo não parece ser um projeto do Final Cut Pro."
}
}
self.handleDroppedURL(url)
}
return true
}
private func handleDroppedURL(_ url: URL?) {
DispatchQueue.main.async {
if let url, isAccepted(url) {
inspect(url.path)
} else {
errorMessage = "Este arquivo não parece ser um projeto do Final Cut Pro."
}
}
}
private func isAccepted(_ url: URL) -> Bool {
let ext = url.pathExtension.lowercased()
return ext == "fcpxml" || ext == "fcpxmld" || ext == "xml"
@@ -242,6 +282,8 @@ struct ProjectView: View {
if let result, result["ok"] as? Bool == true {
project = ProjectInfo(json: result)
showTranscription = true
PythonBridge.call(command: "set_project_config",
arguments: ["file": path]) { _, _ in }
} else {
errorMessage = result?["error"] as? String ?? err ?? "Falha ao ler o projeto."
}
@@ -0,0 +1,218 @@
import SwiftUI
/// Simulação ao vivo da composição `phrase` das legendas dinâmicas como um
/// **quadro vertical 9:16**, na mesma geometria que `fcpxml/text_layout.py`
/// usa para posicionar os títulos no Final Cut.
///
/// Mapeamento fiel ao backend (`compose_sentence`):
/// * o canvas de referência é **1080×1920 pt** (2160×3840 a `POINT_SCALE = 0.5`);
/// * `bodySize`/`emphasisSize` e `blockCenterY` são pontos nesse canvas — o
/// preview multiplica tudo por ``altura do preview / 1920``;
/// * a linha de ênfase fica centrada e as linhas de apoio "penduram" nas
/// bordas dela, alternando lados a partir da esquerda, deslocadas por
/// `REFERENCE_STAGGER_RATIO` (0.8) da folga em relação à linha mais larga;
/// * as linhas empilham com `lineGap` pontos de tinta a tinta entre elas — 0
/// as encosta e negativo sobrepõe — e o bloco inteiro é centrado em
/// `blockCenterY`;
/// * o corpo é sempre **Bold**, a ênfase usa a face escolhida.
///
/// As larguras são estimadas pelo próprio layout de texto do macOS, então o
/// resultado é uma **aproximação** do render do FCP — mas a posição relativa,
/// as proporções e o escalonamento são reais.
struct SubtitlePreviewView: View {
var bodyFont: String
var bodySize: Double
var bodyColor: Color
var emphasisFont: String
var emphasisFace: String
var emphasisSize: Double
var emphasisColor: Color
var bandHeight: Double
var blockCenterY: Double
var lineGap: Double
@Binding var beforeText: String
@Binding var emphasisText: String
@Binding var afterText: String
/// Mostra a faixa (banda) e a linha de centro por cima do quadro.
var showsGuides: Bool = true
/// Canvas de referência do FCP: 2160×3840 px a POINT_SCALE 0.5.
private let canvasHeight: CGFloat = 1920
private let canvasWidth: CGFloat = 1080
/// REFERENCE_STAGGER_RATIO.
private let staggerRatio: CGFloat = 0.8
/// `side_margin` de LayoutBox.for_frame.
private let sideMargin: CGFloat = 0.04
var body: some View {
GeometryReader { geo in
let scale = geo.size.height / canvasHeight
let lines = composedLines(scale: scale, usableWidth: geo.size.width * (1 - 2 * sideMargin))
let bandPixels = geo.size.height * CGFloat(bandHeight)
// y cresce para cima no FCP; na tela cresce para baixo.
let centerY = geo.size.height * 0.5 - CGFloat(blockCenterY) * scale
ZStack {
LinearGradient(
colors: [Color(white: 0.14), Color(white: 0.03)],
startPoint: .top, endPoint: .bottom
)
if showsGuides {
Rectangle()
.fill(Color.accentColor.opacity(0.10))
.frame(height: bandPixels)
.overlay(alignment: .top) { guideRule }
.overlay(alignment: .bottom) { guideRule }
.position(x: geo.size.width / 2, y: centerY)
Rectangle()
.fill(Color.white.opacity(0.16))
.frame(height: 1)
.position(x: geo.size.width / 2, y: geo.size.height * 0.5)
}
ForEach(Array(lines.enumerated()), id: \.offset) { _, line in
Text(line.text)
.font(line.font)
.foregroundColor(line.color)
.lineLimit(1)
.fixedSize()
.position(
x: geo.size.width / 2 + line.x,
y: centerY + line.y
)
}
}
.clipShape(RoundedRectangle(cornerRadius: 10))
.overlay(
RoundedRectangle(cornerRadius: 10)
.strokeBorder(Color.white.opacity(0.15), lineWidth: 1)
)
}
.aspectRatio(canvasWidth / canvasHeight, contentMode: .fit)
}
private var guideRule: some View {
Rectangle().fill(Color.accentColor.opacity(0.45)).frame(height: 1)
}
// MARK: - Composição
private struct Line {
var text: String
var font: Font
var nsFont: NSFont
var color: Color
var isEmphasis: Bool
var width: CGFloat
var height: CGFloat
var x: CGFloat = 0
var y: CGFloat = 0
}
/// Reproduz `compose_sentence`: quebra as linhas de apoio na largura útil,
/// empilha os blocos com `lineGap` e desloca as linhas de apoio para os
/// lados alternados da linha mais larga.
private func composedLines(scale: CGFloat, usableWidth: CGFloat) -> [Line] {
let bodyPt = max(3, CGFloat(bodySize) * scale)
let emphasisPt = max(3, CGFloat(emphasisSize) * scale)
var lines: [Line] = []
lines += bodyLines(before(), size: bodyPt, maxWidth: usableWidth)
let emphasis = emphasisText.trimmingCharacters(in: .whitespaces)
lines.append(makeLine(
emphasis.isEmpty ? " " : emphasis,
nsFont: resolvedFont(emphasisFont, size: emphasisPt, face: emphasisFace),
color: emphasisColor,
isEmphasis: true
))
lines += bodyLines(after(), size: bodyPt, maxWidth: usableWidth)
guard !lines.isEmpty else { return [] }
let gap = CGFloat(lineGap) * scale
let stackHeight = lines.reduce(0) { $0 + $1.height } + gap * CGFloat(lines.count - 1)
let anchor = lines.map(\.width).max() ?? 0
// Topo do bloco relativo ao seu próprio centro.
var edge = -stackHeight / 2
var side: CGFloat = -1
for index in lines.indices {
if index > 0 { edge += gap }
// `.position` centra o frame do Text, cujo centro fica acima do
// centro da tinta — corrige para a tinta cair onde o FCP a põe.
let f = lines[index].nsFont
lines[index].y = edge + lines[index].height / 2 + (f.capHeight - f.ascender) / 2
edge += lines[index].height
if !lines[index].isEmphasis {
lines[index].x = side * (anchor - lines[index].width) / 2 * staggerRatio
side = -side
}
}
return lines
}
private func before() -> String { beforeText.trimmingCharacters(in: .whitespaces) }
private func after() -> String { afterText.trimmingCharacters(in: .whitespaces) }
/// Quebra um trecho de apoio em linhas que cabem em `maxWidth`, palavra a
/// palavra — o mesmo critério de `body_lines` no backend.
private func bodyLines(_ run: String, size: CGFloat, maxWidth: CGFloat) -> [Line] {
guard !run.isEmpty else { return [] }
let nsFont = resolvedFont(bodyFont, size: size, face: "Bold")
var out: [Line] = []
var current = ""
for word in run.split(separator: " ").map(String.init) {
let trial = current.isEmpty ? word : current + " " + word
if !current.isEmpty, textWidth(trial, nsFont) > maxWidth {
out.append(makeLine(current, nsFont: nsFont, color: bodyColor, isEmphasis: false))
current = word
} else {
current = trial
}
}
if !current.isEmpty {
out.append(makeLine(current, nsFont: nsFont, color: bodyColor, isEmphasis: false))
}
return out
}
private func makeLine(_ text: String, nsFont: NSFont, color: Color, isEmphasis: Bool) -> Line {
Line(
text: text,
font: Font(nsFont as CTFont),
nsFont: nsFont,
color: color,
isEmphasis: isEmphasis,
// Empilha sobre a tinta real (cap-height + descida), como ink_extent.
width: textWidth(text, nsFont),
height: nsFont.capHeight + abs(nsFont.descender)
)
}
private func textWidth(_ text: String, _ font: NSFont) -> CGFloat {
(text as NSString).size(withAttributes: [.font: font]).width
}
/// Resolve família + face para uma `NSFont`, caindo para a fonte do sistema
/// quando a família não está instalada — assim o preview nunca some.
private func resolvedFont(_ family: String, size: CGFloat, face: String?) -> NSFont {
var descriptor = NSFontDescriptor(fontAttributes: [.family: family])
if let face, !face.isEmpty {
descriptor = descriptor.addingAttributes([.face: face])
}
if let font = NSFont(descriptor: descriptor, size: size) {
return font
}
var traits: NSFontDescriptor.SymbolicTraits = []
if face?.localizedCaseInsensitiveContains("italic") == true { traits.insert(.italic) }
if face?.localizedCaseInsensitiveContains("bold") == true { traits.insert(.bold) }
let fallback = NSFontDescriptor(fontAttributes: [.family: family])
.withSymbolicTraits(traits)
return NSFont(descriptor: fallback, size: size)
?? NSFont.systemFont(ofSize: size)
}
}
+290 -117
View File
@@ -16,6 +16,8 @@ struct TranscriptionView: View {
@State private var processedPath: String?
@State private var isRemovingSilences = false
@State private var silencePadding: Double = 0.05
@State private var silenceNoiseDb: Double = -30
@State private var silenceMinDuration: Double = 0.5
@State private var isRemovingFillers = false
@State private var fillerRemovedPath: String?
@State private var phraseInput = ""
@@ -39,27 +41,19 @@ struct TranscriptionView: View {
@State private var selectedZoomEndID: Int?
@State private var zoomMessage = ""
@State private var outputFolder: String?
@State private var batchVoiceAnalysis = false
@State private var batchVoiceEdit = false
@State private var voiceActionsPath: String?
@State private var batchSilences = true
@State private var batchFillers = false
@State private var batchPhrases = false
@State private var batchMarkers = false
@State private var batchSubtitles = true
@State private var batchDynamicSubtitles = false
@State private var dynamicSubtitlesBandHeight: Double = 0.22
@State private var dynamicSubtitlesBlockCenterY: Double = -167
@State private var dynamicSubtitlesFont = "Helvetica Neue"
@State private var dynamicSubtitlesFontSize: Double = 90
@State private var dynamicSubtitlesActiveColor = Color.white
@State private var dynamicSubtitlesInactiveColor = Color(white: 0.7)
@State private var dynamicSubtitlesPath: String?
@State private var isBatchProcessing = false
@State private var batchStatus = ""
private let dynamicSubtitlesFontChoices = [
"Helvetica Neue", "Helvetica", "Arial", "Avenir Next",
"Futura", "SF Pro Display", "Georgia", "Impact",
]
init(projectPath: String? = nil, embedded: Bool = false) {
self.embedded = embedded
_projectPath = State(initialValue: projectPath)
@@ -133,6 +127,19 @@ struct TranscriptionView: View {
.disabled(isRunning || projectPath == nil || outputFolder == nil || hasNoInstalledModel)
.frame(maxWidth: .infinity)
.buttonStyle(.borderedProminent)
if isRunning {
VStack(alignment: .leading, spacing: 6) {
ProgressView(value: progress)
HStack {
Text(stage.isEmpty ? "Processando o áudio…" : stage)
Spacer()
Text("\(Int(progress * 100))%").monospacedDigit()
}
.font(.caption).foregroundStyle(.secondary)
}
.padding(.top, 4)
}
}
.disabled(outputFolder == nil)
@@ -143,15 +150,6 @@ struct TranscriptionView: View {
}
}
if isRunning {
Section {
VStack(alignment: .leading, spacing: 8) {
ProgressView(value: progress)
Text("\(Int(progress * 100))%").font(.caption).foregroundStyle(.secondary)
}
}
}
if !results.isEmpty {
Section("Transcrição Concluída") {
ForEach(results, id: \.media) { r in
@@ -176,105 +174,174 @@ struct TranscriptionView: View {
if outputFolder == nil {
Text("Selecione a pasta do projeto para habilitar o processamento.")
.font(.caption2).foregroundStyle(.secondary)
} else if results.isEmpty {
Label(
isRunning ? "Aguardando a transcrição terminar…" : "Transcreva o projeto acima para liberar o processamento.",
systemImage: "lock.fill"
)
.font(.caption).foregroundStyle(.orange)
}
GroupBox("Processar em lote") {
VStack(alignment: .leading, spacing: 6) {
Toggle("Remover silêncios do áudio", isOn: $batchSilences)
VStack(alignment: .leading, spacing: 4) {
HStack {
Text("Tolerância do corte")
.foregroundStyle(batchSilences ? .primary : .secondary)
Spacer()
Text(String(format: "%.2fs", silencePadding))
.font(.caption).foregroundStyle(.secondary).monospacedDigit()
}
Slider(value: $silencePadding, in: 0...2, step: 0.01)
.disabled(!batchSilences)
Text("Quanto de silêncio sobra em volta de cada corte.")
.font(.caption2).foregroundStyle(.secondary)
GroupBox {
VStack(alignment: .leading, spacing: 18) {
batchOptionRow(
toggle: Toggle("Analisar voz (transcrição, locutor, ênfase)", isOn: $batchVoiceAnalysis),
expanded: batchVoiceAnalysis
) {
Text("Gera o JSON com transcrição, diarização e intensidade (pitch/energia/ritmo) por palavra — a base que o corte por voz usa para decidir tomadas e zooms sem reabrir o áudio depois. Só análise: não corta nada.")
.font(.caption).foregroundStyle(.secondary)
}
.padding(.leading, 20)
Divider()
batchOptionRow(
toggle: Toggle("Aplicar edição por voz (lista de decisões)", isOn: $batchVoiceEdit),
expanded: batchVoiceEdit
) {
VStack(alignment: .leading, spacing: 6) {
Text("Aplica um JSON de decisões — cortes, zooms, textos e marcadores — gerado a partir da análise de voz. É o passo que descarta bastidor e tomadas repetidas.")
.font(.caption).foregroundStyle(.secondary)
HStack {
Image(systemName: "doc.text")
Text(voiceActionsPath.map { URL(fileURLWithPath: $0).lastPathComponent }
?? "Nenhum arquivo de decisões")
.foregroundStyle(voiceActionsPath == nil ? .secondary : .primary)
.lineLimit(1).truncationMode(.middle)
Spacer()
Button("Escolher…") { pickVoiceActions() }
}
}
}
Divider()
batchOptionRow(
toggle: Toggle("Remover silêncios do áudio", isOn: $batchSilences),
expanded: true
) {
VStack(alignment: .leading, spacing: 6) {
HStack {
Text("Tolerância do corte")
.font(.callout)
.foregroundStyle(.secondary)
Spacer()
Text(String(format: "%.2fs", silencePadding))
.font(.callout).foregroundStyle(.secondary).monospacedDigit()
}
Slider(value: $silencePadding, in: 0...2, step: 0.01)
.disabled(!batchSilences)
.onChange(of: silencePadding) { _, _ in saveSilenceConfig() }
HStack {
Text("Limiar de silêncio")
.font(.callout)
.foregroundStyle(.secondary)
Spacer()
Text(String(format: "%.0f dB", silenceNoiseDb))
.font(.callout).foregroundStyle(.secondary).monospacedDigit()
}
Slider(value: $silenceNoiseDb, in: -60...(-10), step: 1)
.disabled(!batchSilences)
.onChange(of: silenceNoiseDb) { _, _ in saveSilenceConfig() }
HStack {
Text("Duração mínima")
.font(.callout)
.foregroundStyle(.secondary)
Spacer()
Text(String(format: "%.2fs", silenceMinDuration))
.font(.callout).foregroundStyle(.secondary).monospacedDigit()
}
Slider(value: $silenceMinDuration, in: 0.1...5, step: 0.05)
.disabled(!batchSilences)
.onChange(of: silenceMinDuration) { _, _ in saveSilenceConfig() }
Text("Quanto de silêncio sobra em volta de cada corte, abaixo de que volume conta como silêncio, e quanto tempo ele precisa durar. Fica salvo e vale também fora do app.")
.font(.caption).foregroundStyle(.secondary)
}
}
Divider()
Toggle("Remover palavras de preenchimento", isOn: $batchFillers)
Toggle("Cortar frases ditas", isOn: $batchPhrases)
if batchPhrases {
Divider()
batchOptionRow(
toggle: Toggle("Cortar frases ditas", isOn: $batchPhrases),
expanded: batchPhrases
) {
TextField("Frases separadas por vírgula", text: $phraseInput)
.textFieldStyle(.roundedBorder)
.padding(.leading, 20)
}
Divider()
Toggle("Marcar o que foi dito na timeline", isOn: $batchMarkers)
Divider()
Toggle("Exportar legendas SRT", isOn: $batchSubtitles)
Toggle("Gerar legendas dinâmicas (cascata, editáveis no FCP)", isOn: $batchDynamicSubtitles)
if batchDynamicSubtitles {
VStack(alignment: .leading, spacing: 6) {
Picker("Fonte", selection: $dynamicSubtitlesFont) {
ForEach(dynamicSubtitlesFontChoices, id: \.self) { Text($0).tag($0) }
}
HStack {
Text("Tamanho da fonte")
Spacer()
Text("\(Int(dynamicSubtitlesFontSize))pt")
.font(.caption).foregroundStyle(.secondary).monospacedDigit()
}
Slider(value: $dynamicSubtitlesFontSize, in: 20...200, step: 1)
ColorPicker("Cor A", selection: $dynamicSubtitlesActiveColor, supportsOpacity: true)
ColorPicker("Cor B", selection: $dynamicSubtitlesInactiveColor, supportsOpacity: true)
HStack {
Text("Altura do bloco")
Spacer()
Text("\(Int(dynamicSubtitlesBandHeight * 100))%")
.font(.caption).foregroundStyle(.secondary).monospacedDigit()
}
Slider(value: $dynamicSubtitlesBandHeight, in: 0.10...0.50, step: 0.01)
HStack {
Text("Altura na tela")
Spacer()
Text("\(Int(dynamicSubtitlesBlockCenterY))")
.font(.caption).foregroundStyle(.secondary).monospacedDigit()
}
Slider(value: $dynamicSubtitlesBlockCenterY, in: -700...300, step: 1)
Text("Roda por último, depois dos outros passos marcados acima, usando o timing já cortado da timeline.")
.font(.caption2).foregroundStyle(.secondary)
}
.padding(.leading, 20)
}
Button {
processBatch()
} label: {
if isBatchProcessing { ProgressView().controlSize(.small) }
Label("Processar selecionados", systemImage: "play.fill")
}
.disabled(isBatchProcessing || isRunning || projectPath == nil || outputFolder == nil
|| !batchSilences && !batchFillers && !batchPhrases && !batchMarkers
&& !batchSubtitles && !batchDynamicSubtitles)
if !batchStatus.isEmpty {
Text(batchStatus).font(.caption).foregroundStyle(.secondary)
}
HStack {
Button("Abrir pasta selecionada") {
if let outputFolder {
NSWorkspace.shared.open(URL(fileURLWithPath: outputFolder))
}
}
.disabled(outputFolder == nil)
Button("Abrir no Final Cut Pro") {
if let path = dynamicSubtitlesPath ?? processedPath ?? transcriptEditPath ?? transcriptMarkersPath {
NSWorkspace.shared.open(URL(fileURLWithPath: path))
}
}
.disabled(dynamicSubtitlesPath == nil && processedPath == nil
&& transcriptEditPath == nil && transcriptMarkersPath == nil)
Divider()
batchOptionRow(
toggle: Toggle("Gerar legendas dinâmicas (cascata, editáveis no FCP)", isOn: $batchDynamicSubtitles),
expanded: batchDynamicSubtitles
) {
Text("Usa o estilo configurado na aba \"Legendas Dinâmicas\". Roda por último, depois dos outros passos marcados acima, usando o timing já cortado da timeline.")
.font(.caption).foregroundStyle(.secondary)
}
}
.padding(.vertical, 8)
}
if results.isEmpty {
Text("Transcreva primeiro para habilitar os cortes por texto.")
.font(.caption2).foregroundStyle(.secondary)
VStack(alignment: .leading, spacing: 12) {
Button {
processBatch()
} label: {
if isBatchProcessing {
HStack {
ProgressView().controlSize(.small)
Text("Processando…")
}
.frame(maxWidth: .infinity)
} else {
Label("Processar selecionados", systemImage: "play.fill")
.frame(maxWidth: .infinity)
}
}
.buttonStyle(.borderedProminent)
.controlSize(.large)
.disabled(isBatchProcessing || isRunning || projectPath == nil || outputFolder == nil
|| !batchVoiceAnalysis && !batchVoiceEdit && !batchSilences && !batchFillers
&& !batchPhrases && !batchMarkers && !batchSubtitles && !batchDynamicSubtitles
// Ligada sem arquivo, a etapa falharia no meio da
// cadeia e interromperia tudo o que vem depois.
|| batchVoiceEdit && voiceActionsPath == nil)
if !batchStatus.isEmpty {
Text(batchStatus).font(.caption).foregroundStyle(.secondary)
}
HStack(spacing: 12) {
Button("Abrir pasta selecionada") {
if let outputFolder {
NSWorkspace.shared.open(URL(fileURLWithPath: outputFolder))
}
}
.disabled(outputFolder == nil)
Button("Abrir no Final Cut Pro") {
if let path = dynamicSubtitlesPath ?? processedPath ?? transcriptEditPath ?? transcriptMarkersPath {
NSWorkspace.shared.open(URL(fileURLWithPath: path))
}
}
.disabled(dynamicSubtitlesPath == nil && processedPath == nil
&& transcriptEditPath == nil && transcriptMarkersPath == nil)
Spacer()
}
}
.padding(.top, 4)
}
.disabled(outputFolder == nil)
.disabled(outputFolder == nil || results.isEmpty)
Section("Zoom (aproximar em um trecho)") {
if outputFolder != nil && results.isEmpty {
Label(
isRunning ? "Aguardando a transcrição terminar…" : "Transcreva o projeto acima para liberar o zoom.",
systemImage: "lock.fill"
)
.font(.caption).foregroundStyle(.orange)
}
if zoomClips.isEmpty {
Text("Carregando clipes do projeto…")
.font(.caption).foregroundStyle(.secondary)
@@ -347,12 +414,47 @@ struct TranscriptionView: View {
resultRow(zoomResultPath)
}
}
.disabled(outputFolder == nil)
.disabled(outputFolder == nil || results.isEmpty)
}
.formStyle(.grouped)
.task { await loadCatalog(); loadZoomClips() }
.onChange(of: projectPath) { _, _ in loadZoomClips() }
.onChange(of: outputFolder) { _, _ in loadZoomClips() }
.task { await loadCatalog(); loadProjectConfig(); loadZoomClips(); loadSilenceConfig() }
.onChange(of: projectPath) { _, newValue in
loadZoomClips()
saveProjectConfig(file: newValue ?? "")
}
.onChange(of: outputFolder) { _, newValue in
loadZoomClips()
saveProjectConfig(folder: newValue ?? "")
}
}
/// Reabre o último projeto salvo em ~/.fcp-mcp-server/config.json — pasta
/// de saída e arquivo — para a tela não começar vazia a cada abertura.
/// Só preenche o que ainda está vazio, então o projeto aberto pela aba
/// Projeto (embedded) nunca é sobrescrito pelo que ficou salvo. Caminhos
/// que sumiram do disco voltam vazios do Python e são ignorados aqui.
private func loadProjectConfig() {
PythonBridge.call(command: "project_config") { result, _ in
DispatchQueue.main.async {
guard let result, result["ok"] as? Bool == true else { return }
if outputFolder == nil, let folder = result["folder"] as? String, !folder.isEmpty {
outputFolder = folder
}
if projectPath == nil, let file = result["file"] as? String, !file.isEmpty {
projectPath = file
}
}
}
}
/// Grava a pasta e/ou o arquivo do projeto no config compartilhado. Campos
/// omitidos ficam como estão; string vazia limpa o campo.
private func saveProjectConfig(folder: String? = nil, file: String? = nil) {
var arguments: [String: Any] = [:]
if let folder { arguments["folder"] = folder }
if let file { arguments["file"] = file }
guard !arguments.isEmpty else { return }
PythonBridge.call(command: "set_project_config", arguments: arguments) { _, _ in }
}
/// Presents an NSOpenPanel configured to select a single file (not a
@@ -374,6 +476,45 @@ struct TranscriptionView: View {
}
}
/// Lê os limiares de silêncio salvos em ~/.fcp-mcp-server/config.json —
/// os mesmos que remove_media_silence usa por padrão — para os controles
/// abrirem já mostrando o que de fato vai rodar.
private func loadSilenceConfig() {
PythonBridge.call(command: "silence_config") { result, _ in
DispatchQueue.main.async {
guard let result, result["ok"] as? Bool == true else { return }
silencePadding = result["padding"] as? Double ?? silencePadding
silenceNoiseDb = result["noise_db"] as? Double ?? silenceNoiseDb
silenceMinDuration = result["min_silence"] as? Double ?? silenceMinDuration
}
}
}
private func saveSilenceConfig() {
PythonBridge.call(command: "set_silence_config", arguments: [
"padding": silencePadding,
"noise_db": silenceNoiseDb,
"min_silence": silenceMinDuration,
]) { _, _ in }
}
/// Escolhe o JSON de decisões (cortes/zooms/textos/marcadores) que
/// `apply_voice_actions` vai aplicar. Pré-seleciona a pasta do projeto,
/// que é onde o arquivo costuma ser gravado.
private func pickVoiceActions() {
let panel = NSOpenPanel()
panel.canChooseFiles = true
panel.canChooseDirectories = false
panel.allowsMultipleSelection = false
panel.allowedContentTypes = [.json]
panel.prompt = "Usar este arquivo"
panel.message = "Selecione o JSON com a lista de decisões da edição por voz."
if let outputFolder { panel.directoryURL = URL(fileURLWithPath: outputFolder) }
if panel.runModal() == .OK, let url = panel.url {
voiceActionsPath = url.path
}
}
private func pickOutputFolder() {
let panel = NSOpenPanel()
panel.canChooseFiles = false
@@ -438,6 +579,28 @@ struct TranscriptionView: View {
}
}
/// Renders a toggle plus its optional expanded detail block, indented and
/// visually tied together so batch options don't collapse into one wall
/// of controls with no breathing room.
@ViewBuilder
private func batchOptionRow<Toggle: View, Detail: View>(
toggle: Toggle, expanded: Bool, @ViewBuilder detail: () -> Detail
) -> some View {
VStack(alignment: .leading, spacing: 12) {
toggle
if expanded {
detail()
.padding(.leading, 20)
.padding(.vertical, 10)
.padding(.trailing, 8)
.background(
RoundedRectangle(cornerRadius: 8)
.fill(Color.secondary.opacity(0.06))
)
}
}
}
@ViewBuilder
private func resultRow(_ path: String) -> some View {
Text(path).font(.caption).lineLimit(1)
@@ -596,9 +759,11 @@ struct TranscriptionView: View {
guard let projectPath else { return }
isRemovingSilences = true
errorMessage = nil
// Sem repassar limiares: remove_silences lê os valores salvos em
// ~/.fcp-mcp-server/config.json (os mesmos que os controles acima
// gravam), então há uma só fonte da verdade.
PythonBridge.call(command: "remove_silences", arguments: [
"path": projectPath,
"padding": silencePadding,
]) { result, err in
DispatchQueue.main.async {
isRemovingSilences = false
@@ -614,6 +779,16 @@ struct TranscriptionView: View {
private func processBatch() {
guard let projectPath, let outputFolder else { return }
var operations: [String] = []
// Analysis first: it only writes a JSON sidecar next to the source
// media and never touches the project XML, so its position relative
// to the other steps doesn't change what they do — but running it
// before any cut keeps the mental model simple (measure, then edit).
if batchVoiceAnalysis { operations.append("analyze_voice") }
// The edit itself, and it has to run FIRST on the intact timeline:
// every zoom/text/marker it places is positioned by shifting from its
// own cut list, so a step that already rippled the timeline would put
// them on the wrong frame — with no visible error.
if batchVoiceEdit { operations.append("apply_voice_actions") }
if batchSilences { operations.append("remove_silences") }
if batchFillers { operations.append("remove_filler_words") }
if batchPhrases { operations.append("edit_by_transcript") }
@@ -638,19 +813,17 @@ struct TranscriptionView: View {
let operation = operations[index]
batchStatus = "Processando: \(operation)…"
var arguments: [String: Any] = ["path": currentPath, "output_dir": outputFolder]
if operation == "remove_silences" { arguments["padding"] = silencePadding }
if operation == "apply_voice_actions", let voiceActionsPath {
arguments["actions_path"] = voiceActionsPath
}
if operation == "edit_by_transcript" {
let phrases = phraseInput.split(separator: ",").map { $0.trimmingCharacters(in: .whitespaces) }.filter { !$0.isEmpty }
arguments["phrases"] = phrases
}
if operation == "generate_dynamic_subtitles" {
arguments["band_height"] = dynamicSubtitlesBandHeight
arguments["block_center_y"] = dynamicSubtitlesBlockCenterY
arguments["font"] = dynamicSubtitlesFont
arguments["font_size"] = Int(dynamicSubtitlesFontSize)
arguments["active_color"] = dynamicSubtitlesActiveColor.fcpxmlColorString
arguments["inactive_color"] = dynamicSubtitlesInactiveColor.fcpxmlColorString
}
// O estilo das legendas não é repassado aqui: generate_dynamic_subtitles
// lê ~/.fcp-mcp-server/config.json (o mesmo arquivo que a aba "Legendas
// Dinâmicas" grava) como seu próprio padrão, então há uma só fonte da
// verdade em vez de duas cópias podendo divergir.
PythonBridge.call(command: operation, arguments: arguments) { result, err in
DispatchQueue.main.async {
guard result?["ok"] as? Bool == true else {
+258
View File
@@ -0,0 +1,258 @@
import SwiftUI
/// Guia "Análise de Voz" — parâmetros do motor de análise acústica da fala.
///
/// Controla os limiares que decidem quais palavras são candidatas a
/// punch-in/destaque: energia (intensidade da fala), o índice de ênfase e
/// seus pesos, e a camada opcional de emoção. Os valores ficam em
/// `~/.fcp-mcp-server/config.json` (via `model_manager.save_voice_analysis_config`),
/// os mesmos lidos pelas ferramentas de análise — não em UserDefaults, para
/// que a análise rode com exatamente o que a tela mostra.
struct VoiceAnalysisView: View {
@State private var config = VoiceAnalysisConfig.defaults
@State private var isLoading = true
@State private var errorMessage: String?
var body: some View {
Form {
if isLoading {
Section {
ProgressView().controlSize(.small)
.frame(maxWidth: .infinity, alignment: .center)
}
} else {
energySection
emphasisSection
weightsSection
emotionSection
resetSection
}
if let errorMessage {
Section {
Label(errorMessage, systemImage: "exclamationmark.triangle.fill")
.foregroundStyle(.red)
}
}
}
.formStyle(.grouped)
.task { await load() }
}
// MARK: - Energia
private var energySection: some View {
Section {
sliderRow(
title: "Limiar de energia",
value: $config.energyThreshold,
help: "Acima deste valor a fala conta como \"alta energia\"."
)
} header: {
Text("Energia da Fala")
} footer: {
Text("A energia de cada palavra é normalizada (0–1) pelo trecho mais alto do áudio. Valores mais baixos marcam mais palavras como intensas.")
.font(.caption)
.foregroundStyle(.secondary)
}
}
// MARK: - Ênfase
private var emphasisSection: some View {
Section {
sliderRow(
title: "Limiar de ênfase",
value: $config.emphasisThreshold,
help: "Acima deste valor a palavra vira candidata a punch-in/destaque."
)
} header: {
Text("Índice de Ênfase")
} footer: {
Text("O índice é a média ponderada de energia, variação de tom, variação de ritmo, pausa anterior e duração — com os pesos abaixo. Por ser média, picos reais ficam entre 0,55 e 0,80: acima de 0,85 quase nada é selecionado.")
.font(.caption)
.foregroundStyle(.secondary)
}
}
private var weightsSection: some View {
Section {
sliderRow(title: "Energia", value: $config.weightEnergy, range: 0...1)
sliderRow(title: "Variação de tom", value: $config.weightPitch, range: 0...1)
sliderRow(title: "Variação de ritmo", value: $config.weightRate, range: 0...1)
sliderRow(title: "Pausa anterior", value: $config.weightPause, range: 0...1)
sliderRow(title: "Duração da palavra", value: $config.weightDuration, range: 0...1)
} header: {
Text("Pesos do Índice de Ênfase")
} footer: {
Text("Não precisam somar 1 — são normalizados internamente. O que importa é a proporção entre eles.")
.font(.caption)
.foregroundStyle(.secondary)
}
}
// MARK: - Emoção
private var emotionSection: some View {
Section {
Toggle("Detectar emoção durante a fala", isOn: $config.emotionEnabled)
.onChange(of: config.emotionEnabled) { _, _ in save() }
if config.emotionEnabled {
sliderRow(
title: "Sensibilidade",
value: $config.emotionSensitivity,
help: "Confiança mínima para aceitar uma emoção detectada."
)
}
} header: {
Text("Emoção")
} footer: {
Text("A emoção nunca decide um corte sozinha — entra combinada com energia, tom e ênfase. Exige o componente opcional de emoção instalado; sem ele, a análise segue normalmente sem essa camada.")
.font(.caption)
.foregroundStyle(.secondary)
}
}
private var resetSection: some View {
Section {
Button("Restaurar padrões") {
config = .defaults
save()
}
}
}
// MARK: - Componentes
/// Slider com rótulo à esquerda e valor numérico à direita, salvando no
/// backend só quando o arraste termina (evita uma escrita por quadro).
private func sliderRow(
title: String,
value: Binding<Double>,
range: ClosedRange<Double> = 0...1,
help: String? = nil
) -> some View {
VStack(alignment: .leading, spacing: 2) {
HStack {
Text(title)
Spacer()
Text(String(format: "%.2f", value.wrappedValue))
.monospacedDigit()
.foregroundStyle(.secondary)
}
Slider(value: value, in: range) { editing in
if !editing { save() }
}
if let help {
Text(help)
.font(.caption)
.foregroundStyle(.secondary)
}
}
}
// MARK: - Backend
@MainActor
private func load() async {
await withCheckedContinuation { continuation in
PythonBridge.call(command: "voice_analysis") { result, error in
DispatchQueue.main.async {
if let result {
config = VoiceAnalysisConfig(from: result)
} else if let error {
errorMessage = error
}
isLoading = false
continuation.resume()
}
}
}
}
private func save() {
PythonBridge.call(command: "set_voice_analysis", arguments: config.arguments()) { _, error in
DispatchQueue.main.async { errorMessage = error }
}
}
}
/// Os parâmetros de análise de voz, no formato que a tela edita e o bridge
/// (`admin/models_api.py` → `set_voice_analysis`) persiste.
struct VoiceAnalysisConfig {
var energyThreshold: Double
var emphasisThreshold: Double
var weightEnergy: Double
var weightPitch: Double
var weightRate: Double
var weightPause: Double
var weightDuration: Double
var emotionEnabled: Bool
var emotionSensitivity: Double
static let defaults = VoiceAnalysisConfig(
energyThreshold: 0.5,
emphasisThreshold: 0.60,
weightEnergy: 0.30,
weightPitch: 0.25,
weightRate: 0.20,
weightPause: 0.15,
weightDuration: 0.10,
emotionEnabled: false,
emotionSensitivity: 0.5
)
init(
energyThreshold: Double,
emphasisThreshold: Double,
weightEnergy: Double,
weightPitch: Double,
weightRate: Double,
weightPause: Double,
weightDuration: Double,
emotionEnabled: Bool,
emotionSensitivity: Double
) {
self.energyThreshold = energyThreshold
self.emphasisThreshold = emphasisThreshold
self.weightEnergy = weightEnergy
self.weightPitch = weightPitch
self.weightRate = weightRate
self.weightPause = weightPause
self.weightDuration = weightDuration
self.emotionEnabled = emotionEnabled
self.emotionSensitivity = emotionSensitivity
}
/// Lê a resposta do bridge, caindo no padrão para qualquer campo ausente.
init(from json: [String: Any]) {
let defaults = VoiceAnalysisConfig.defaults
let weights = json["emphasis_weights"] as? [String: Any] ?? [:]
self.init(
energyThreshold: json["energy_threshold"] as? Double ?? defaults.energyThreshold,
emphasisThreshold: json["emphasis_threshold"] as? Double ?? defaults.emphasisThreshold,
weightEnergy: weights["energy"] as? Double ?? defaults.weightEnergy,
weightPitch: weights["pitch_variation"] as? Double ?? defaults.weightPitch,
weightRate: weights["rate_variation"] as? Double ?? defaults.weightRate,
weightPause: weights["pause_before"] as? Double ?? defaults.weightPause,
weightDuration: weights["duration"] as? Double ?? defaults.weightDuration,
emotionEnabled: json["emotion_enabled"] as? Bool ?? defaults.emotionEnabled,
emotionSensitivity: json["emotion_sensitivity"] as? Double ?? defaults.emotionSensitivity
)
}
func arguments() -> [String: Any] {
[
"energy_threshold": energyThreshold,
"emphasis_threshold": emphasisThreshold,
"emphasis_weights": [
"energy": weightEnergy,
"pitch_variation": weightPitch,
"rate_variation": weightRate,
"pause_before": weightPause,
"duration": weightDuration,
],
"emotion_enabled": emotionEnabled,
"emotion_sensitivity": emotionSensitivity,
]
}
}
+4 -4
View File
@@ -1,6 +1,6 @@
# FCPXML MCP
**The bridge between Final Cut Pro and AI. 62 tools that turn timeline XML into structured data Claude can read, edit, and generate.**
**The bridge between Final Cut Pro and AI. 73 tools that turn timeline XML into structured data Claude can read, edit, and generate.**
[![CI](https://github.com/DareDev256/fcp-mcp-server/actions/workflows/test.yml/badge.svg)](https://github.com/DareDev256/fcp-mcp-server/actions)
[![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg)](https://opensource.org/licenses/MIT)
@@ -410,7 +410,7 @@ Select these from Claude's prompt menu (⌘/) — they chain multiple tools auto
```
fcp-mcp-server/ ~9.4k lines Python
├── server.py MCP entry point — 62 tools, 5 prompts, resource discovery
├── server.py MCP entry point — 73 tools, 5 prompts, resource discovery
│ _resolve_io_paths() / _setup_modifier() / _setup_generator()
│ _format_clip_table() / _markdown_table() / _format_batch_result()
│ _raw_markers_to_batch()
@@ -430,7 +430,7 @@ fcp-mcp-server/ ~9.4k lines Python
│ ├── safe_xml.py Centralized defusedxml wrappers (XXE/entity-bomb protection) + serialize_xml()
│ ├── dtd.py Validate output against Apple's official DTDs (located in the FCP app bundle)
│ └── templates.py Template system (intro/outro, lower thirds, music video)
├── tests/ 1032 tests across 24 suites
├── tests/ 1342 tests across 34 suites
│ ├── test_models.py TimeValue math, Timecode formatting, MarkerType contracts
│ ├── test_parser.py FCPXML parsing, connected clips, edge cases
│ ├── test_writer.py Clip editing, marker writing, speed changes
@@ -568,7 +568,7 @@ uv run --extra dev pytest tests/ -v # or: python3 -m pytest tests/ -v
ruff check . --exclude docs/ # lint — must pass before committing
```
1032 tests across 24 suites covering models, parser, writer, FCPXMLWriter generation, server handlers, rough cut generation, speed cutting & pacing curves, marker pipeline, refactored helper functions, regression fixes, security hardening (XXE, entity expansion, path traversal, sandbox boundaries, minidom defense-in-depth, JSON depth limits, input validation, ffmpeg bounds, write-handler sandboxing), connected clips, roles, diff, export, compound clip flattening, audio track generation, templates, effects, `.fcpxmld` bundles with sidecar preservation, bulk media relink, real media silence detection (parser, timeline mapping, real-WAV ffmpeg integration), and DTD validation against Apple's official DTDs (auto-skipped on machines without Final Cut Pro).
1342 tests across 34 suites covering models, parser, writer, FCPXMLWriter generation, server handlers, rough cut generation, speed cutting & pacing curves, marker pipeline, refactored helper functions, regression fixes, security hardening (XXE, entity expansion, path traversal, sandbox boundaries, minidom defense-in-depth, JSON depth limits, input validation, ffmpeg bounds, write-handler sandboxing), connected clips, roles, diff, export, compound clip flattening, audio track generation, templates, effects, `.fcpxmld` bundles with sidecar preservation, bulk media relink, real media silence detection (parser, timeline mapping, real-WAV ffmpeg integration), and DTD validation against Apple's official DTDs (auto-skipped on machines without Final Cut Pro).
---
+23
View File
@@ -168,6 +168,29 @@ Mark mode adds markers instead of deleting — safer for first pass.
---
## Dynamic Subtitles: Generate, Then Always Validate
**Scenario:** Word-by-word progressive-composition subtitles (the diagrammed look — small supporting words, one key word large in a display italic) need to go on a cut before delivery.
```
"Generate dynamic subtitles for /path/to/project.fcpxml"
```
**Tool chain:** `generate_dynamic_subtitles` → `validate_subtitle_layout`
Run `generate_dynamic_subtitles` on the *final* cut, after cuts/zooms are already applied — a connected title anchors to its parent clip's source-media coordinates, so re-cutting the timeline afterward can silently detach captions from the words they were built for.
**Never treat generation as done without the second call.** The layout only guarantees non-overlap *by construction* for what it itself lays out — it cannot see a hand-edited title, stray content left over in a reused base file, or a word long/uppercase enough to have needed shrinking. `validate_subtitle_layout` re-measures every `<title>` independently and reports a severity (`none`/`warning`/`probable`/`severe`) plus per-issue suggested corrections:
```
"Validate the subtitle layout in /path/to/project_dynamic_subtitles.fcpxmld"
```
- `none`/`warning` (only `outside_safe_area`, no `outside_frame` or `collision`) — safe to deliver; the emphasis word sitting close to the 5% margin is expected on the diagrammed look.
- `probable`/`severe` — investigate before touching code. Read the issue's exact FCPXML fraction times (not the rounded float) before deciding whether it's a real overlap; two titles that are only touching at a shared boundary can still round to "equal-looking but not bit-identical" floats and misreport. See `Engine/docs/03_SERVER_TOOLS.md`'s "Legendas dinâmicas" section and `Engine/docs/05_EXPERIENCIAS.md` (2026-08-19 entries) for the concrete bugs already found and fixed this way, and the checklist for the next one.
---
## Composing Tools in AI Agent Workflows
Each tool in this MCP server follows the same pattern: read FCPXML → process → write modified FCPXML. This makes them composable — the output of one tool is valid input for the next.
+472
View File
@@ -0,0 +1,472 @@
"""Collision detection and layout validation for dynamic-subtitle titles.
Pure functions — no I/O, no FCPXML parsing — that answer one question over and
over: given the boxes a set of titles occupy on screen, do any two titles that
are on screen at the same time intersect? And are they inside the frame, inside
the safe area, and using a font the layout actually measured?
This is the post-generation guarantee the layout engine only provides *by
construction* (``text_layout.compose_sentence`` stacks lines so their ink boxes
never touch). Re-running it over already-emitted titles catches the cases the
layout cannot see: a hand-edited position, a template whose type scales
differently than ``text_scale`` assumed, a font that fell back to an estimate,
or a word pushed off frame by a long emphasis line.
Boxes are measured in the *emitted* template space (frame pixels) — the same
numbers the writer wrote to the FCPXML (``fontSize``, ``kerning`` and
``Position`` are all already scaled by ``text_scale``), so validation re-measures
with ``measure_text``/``ink_extent`` against those same numbers and never
re-applies the scale factor. See ``writer.validate_subtitle_layout``.
"""
from dataclasses import dataclass
from math import hypot
from typing import Dict, List, Optional, Sequence
from .text_layout import (
ink_extent,
measure_text,
metrics_for,
vertical_metrics_for,
)
# Severity buckets for a spatial overlap, ordered from harmless to blocking.
# ``render_tolerance`` is the 5px the renderer can round off; ``severe`` is a
# real collision that must be fixed before export.
OVERLAP_NONE = "none"
OVERLAP_RENDER_TOLERANCE = "render_tolerance"
OVERLAP_WARNING = "warning"
OVERLAP_PROBABLE = "probable"
OVERLAP_SEVERE = "severe"
# Max fraction of the smaller box a severe collision may cover (spec 7.2).
SEVERE_OVERLAP_RATIO = 0.15
# issue types (spec 16)
SPATIAL_COLLISION = "spatial_collision"
OUTSIDE_FRAME = "outside_frame"
OUTSIDE_SAFE_AREA = "outside_safe_area"
INSUFFICIENT_SPACING = "insufficient_spacing"
EXCESSIVE_SPACING = "excessive_spacing"
FONT_MISSING = "font_missing"
FONT_TOO_SMALL = "font_too_small"
INVALID_ANCHOR = "invalid_anchor"
UNRESOLVED_TRANSFORM = "unresolved_transform"
@dataclass
class Box:
"""An axis-aligned rectangle in frame coordinates, y growing upward."""
left: float
right: float
bottom: float
top: float
@property
def width(self) -> float:
return self.right - self.left
@property
def height(self) -> float:
return self.top - self.bottom
@property
def area(self) -> float:
return self.width * self.height
def overlaps(self, other: "Box") -> bool:
"""True if the two boxes intersect (strict — touching edges do not)."""
return (
self.left < other.right
and other.left < self.right
and self.bottom < other.top
and other.bottom < self.top
)
# A boundary the writer places deliberately exact — one block's title
# duration set to literally equal the next block's start (see writer.py's
# ``block_ends``) — can still land a few float-ULPs apart by the time it
# gets here: an ``end`` re-derived as ``start + duration`` from two already-
# rounded floats isn't bit-identical to a ``start`` read as one division of
# the same exact fraction, even though both trace back to one FCPXML value.
# Found on real footage: 7 of 8 "severe" collisions in one clip were exactly
# this — same instant, off by ~1e-13s, nowhere near a real frame boundary
# (~0.04s). A tolerance many orders below one frame absorbs the artifact
# without hiding a genuine overlap.
_BOUNDARY_EPSILON = 1e-6
def temporal_overlap(
start_a: float, end_a: float, start_b: float, end_b: float
) -> bool:
"""Whether the half-open intervals ``[start, end)`` intersect (spec 7.1).
Strict on both sides, so a title that ends exactly when the next begins is
never treated as simultaneous — see ``_BOUNDARY_EPSILON`` for why "exactly"
needs a tolerance rather than bare float comparison.
"""
return (
start_a < end_b - _BOUNDARY_EPSILON
and start_b < end_a - _BOUNDARY_EPSILON
)
def overlap_metrics(a: Box, b: Box) -> Dict[str, float]:
"""Width, height, area and ratio of the intersection of ``a`` and ``b``.
``overlap_ratio`` is the shared area over the *smaller* box's area, so a
small box swallowed by a big one reads as the severe case it is.
"""
overlap_width = min(a.right, b.right) - max(a.left, b.left)
overlap_height = min(a.top, b.top) - max(a.bottom, b.bottom)
overlap_area = max(0.0, overlap_width) * max(0.0, overlap_height)
smaller = min(a.area, b.area)
ratio = overlap_area / smaller if smaller > 0 else 0.0
return {
"overlap_width": overlap_width,
"overlap_height": overlap_height,
"overlap_area": overlap_area,
"overlap_ratio": ratio,
}
def classify_overlap(metrics: Dict[str, float]) -> str:
"""Severity bucket for an overlap, following spec 7.2.
Zero area is no conflict at all; a ratio above ``SEVERE_OVERLAP_RATIO`` is
severe regardless of absolute size; otherwise the vertical penetration is
bucketed into tolerance / warning / probable / severe.
"""
height = metrics["overlap_height"]
if metrics["overlap_area"] <= 0:
return OVERLAP_NONE
if metrics["overlap_ratio"] > SEVERE_OVERLAP_RATIO:
return OVERLAP_SEVERE
if height <= 5:
return OVERLAP_RENDER_TOLERANCE
if height <= 20:
return OVERLAP_WARNING
if height <= 50:
return OVERLAP_PROBABLE
return OVERLAP_SEVERE
def distance_between(a: Box, b: Box) -> Dict[str, float]:
"""Gap between two non-overlapping boxes, per axis and euclidean (spec 8)."""
if a.right < b.left:
distance_x = b.left - a.right
elif b.right < a.left:
distance_x = a.left - b.right
else:
distance_x = 0.0
if a.top < b.bottom:
distance_y = b.bottom - a.top
elif b.top < a.bottom:
distance_y = a.bottom - b.top
else:
distance_y = 0.0
return {
"distance_x": distance_x,
"distance_y": distance_y,
"distance": hypot(distance_x, distance_y),
}
def separation_suggestion(
a: Box, b: Box, min_gap: float = 0.0
) -> Dict[str, float]:
"""The minimum translation that separates two overlapping boxes (spec 10).
Picks the smallest of the four penetrations (move left/right/up/down) and
reports that axis plus the required movement (penetration + ``min_gap``).
"""
move_left = a.right - b.left
move_right = b.right - a.left
move_down = a.top - b.bottom
move_up = b.top - a.bottom
candidates = [
("horizontal", move_left),
("horizontal", move_right),
("vertical", move_down),
("vertical", move_up),
]
axis, penetration = min(candidates, key=lambda kv: kv[1])
return {
"axis": axis,
"minimum_movement": max(0.0, penetration + min_gap),
}
def measure_title_box(
text: str,
font_size: float,
*,
x: float,
y: float,
font: Optional[str] = None,
face: Optional[str] = None,
kerning: float = 0.0,
) -> Box:
"""The on-screen box of one title, measured in the emitted template space.
``font_size``/``kerning``/``x``/``y`` are the values the writer put into the
FCPXML, so the box is comparable across every title in the document without
any further scaling. Width comes from the real advance table, vertical
extent from the real ink (accents and descenders included); the anchor is
the title's centre.
"""
width = measure_text(
text, font_size, kerning=kerning, font=font, face=face
)
top, bottom = ink_extent(text, font_size, font=font, face=face)
return Box(
left=x - width / 2,
right=x + width / 2,
bottom=y + bottom,
top=y + top,
)
def _is_font_measured(font: Optional[str], face: Optional[str]) -> bool:
"""Whether both advance and vertical metrics for ``font``/``face`` exist."""
if metrics_for(font, face) is None:
return False
_, measured = vertical_metrics_for(font, face)
return measured
def _box_within(box: Box, limits: Box) -> bool:
return (
box.left >= limits.left
and box.right <= limits.right
and box.bottom >= limits.bottom
and box.top <= limits.top
)
def _issue(severity: str, type_: str, **fields) -> Dict:
return {"severity": severity, "type": type_, **fields}
def validate_titles(
titles: Sequence[Dict],
frame_width: float,
frame_height: float,
*,
safe_margin_x: float = 0.05,
safe_margin_y: float = 0.05,
min_font_size: Optional[float] = None,
min_distance: Optional[float] = None,
max_distance: Optional[float] = None,
) -> Dict:
"""Validate a set of already-positioned titles and return a report.
Each title dict must carry the emitted values:
- ``text`` (str)
- ``font_size`` (float), ``kerning`` (float), ``x``/``y`` (floats)
- ``font`` (str) and ``face`` (str|None)
- ``start``/``end`` (seconds) for temporal overlap
- ``group`` (hashable) for spacing checks: titles sharing a group are one
block, expected to sit near each other (spec 14). Optional.
Returns ``{"severity", "issues", "summary"}`` where ``severity`` is the
worst bucket seen and ``issues`` are the spec-16-shaped occurrences.
"""
frame_left = -frame_width / 2
frame_right = frame_width / 2
frame_bottom = -frame_height / 2
frame_top = frame_height / 2
frame_box = Box(frame_left, frame_right, frame_bottom, frame_top)
safe_box = Box(
left=frame_left + safe_margin_x * frame_width,
right=frame_right - safe_margin_x * frame_width,
bottom=frame_bottom + safe_margin_y * frame_height,
top=frame_top - safe_margin_y * frame_height,
)
issues: List[Dict] = []
boxes: List[Box] = []
measured_flags: List[bool] = []
for title in titles:
text = str(title.get("text", "") or "")
font = title.get("font") or None
face = title.get("face") or None
box = measure_title_box(
text,
float(title.get("font_size", 0.0)),
x=float(title.get("x", 0.0)),
y=float(title.get("y", 0.0)),
font=font,
face=face,
kerning=float(title.get("kerning", 0.0)),
)
boxes.append(box)
measured_flags.append(_is_font_measured(font, face))
if not _is_font_measured(font, face):
issues.append(
_issue(
"warning",
FONT_MISSING,
title=text,
font=font,
face=face,
message=(
f"Font '{font or '?'}"
+ (f" {face}" if face else "")
+ "' has no embedded metrics; widths are estimated"
),
)
)
if min_font_size is not None and float(title.get("font_size", 0.0)) < min_font_size:
issues.append(
_issue(
"warning",
FONT_TOO_SMALL,
title=text,
font_size=float(title.get("font_size", 0.0)),
minimum=min_font_size,
)
)
if not _box_within(box, frame_box):
issues.append(
_issue(
"error",
OUTSIDE_FRAME,
title=text,
left=box.left,
right=box.right,
bottom=box.bottom,
top=box.top,
)
)
elif not _box_within(box, safe_box):
issues.append(
_issue(
"warning",
OUTSIDE_SAFE_AREA,
title=text,
left=box.left,
right=box.right,
bottom=box.bottom,
top=box.top,
)
)
# Spatial collisions between temporally overlapping titles.
for i in range(len(titles)):
for j in range(i + 1, len(titles)):
a, b = titles[i], titles[j]
if not temporal_overlap(
float(a.get("start", 0.0)), float(a.get("end", 0.0)),
float(b.get("start", 0.0)), float(b.get("end", 0.0)),
):
continue
box_a, box_b = boxes[i], boxes[j]
if not box_a.overlaps(box_b):
continue
metrics = overlap_metrics(box_a, box_b)
severity = classify_overlap(metrics)
issues.append(
_issue(
severity,
SPATIAL_COLLISION,
first_title=str(a.get("text", "")),
second_title=str(b.get("text", "")),
time_start=float(a.get("start", 0.0)),
time_end=float(b.get("end", 0.0)),
overlap_width=metrics["overlap_width"],
overlap_height=metrics["overlap_height"],
overlap_area=metrics["overlap_area"],
overlap_ratio=metrics["overlap_ratio"],
suggested_correction=separation_suggestion(box_a, box_b),
)
)
# Spacing within a block (spec 8/14). Only when the caller asked for it —
# a generic minimum can fire on the reference look's own tight stacking.
if min_distance is not None or max_distance is not None:
groups: Dict = {}
for index, title in enumerate(titles):
groups.setdefault(title.get("group", index), []).append(index)
for members in groups.values():
for m in range(len(members)):
for n in range(m + 1, len(members)):
i, j = members[m], members[n]
box_a, box_b = boxes[i], boxes[j]
if box_a.overlaps(box_b):
continue
gap = distance_between(box_a, box_b)["distance"]
if min_distance is not None and gap < min_distance:
issues.append(
_issue(
"warning",
INSUFFICIENT_SPACING,
first_title=str(titles[i].get("text", "")),
second_title=str(titles[j].get("text", "")),
distance=gap,
minimum=min_distance,
)
)
if max_distance is not None and gap > max_distance:
issues.append(
_issue(
"warning",
EXCESSIVE_SPACING,
first_title=str(titles[i].get("text", "")),
second_title=str(titles[j].get("text", "")),
distance=gap,
maximum=max_distance,
)
)
_rank = {
OVERLAP_NONE: 0,
OVERLAP_RENDER_TOLERANCE: 1,
OVERLAP_WARNING: 2,
OVERLAP_PROBABLE: 3,
OVERLAP_SEVERE: 4,
}
severities = [issue["severity"] for issue in issues]
worst = max(severities, key=lambda s: _rank.get(s, 0), default=OVERLAP_NONE)
return {
"severity": worst,
"issues": issues,
"summary": {
"title_count": len(titles),
"issue_count": len(issues),
"spatial_collision": sum(
1 for i in issues if i["type"] == SPATIAL_COLLISION
),
"outside_frame": sum(
1 for i in issues if i["type"] == OUTSIDE_FRAME
),
"outside_safe_area": sum(
1 for i in issues if i["type"] == OUTSIDE_SAFE_AREA
),
"font_missing": sum(
1 for i in issues if i["type"] == FONT_MISSING
),
"font_too_small": sum(
1 for i in issues if i["type"] == FONT_TOO_SMALL
),
"insufficient_spacing": sum(
1 for i in issues if i["type"] == INSUFFICIENT_SPACING
),
"excessive_spacing": sum(
1 for i in issues if i["type"] == EXCESSIVE_SPACING
),
},
}
def blocking(severity: str) -> bool:
"""Whether a validation severity should block export (spec 16)."""
return severity in (OVERLAP_SEVERE, OVERLAP_PROBABLE)
+31 -1
View File
@@ -37,6 +37,36 @@ def diarization_capability(token: Optional[str]) -> Tuple[bool, str]:
return True, "Identificação de participantes disponível."
def _load_waveform(path: str) -> Optional[dict]:
"""Decode ``path`` ourselves into the waveform dict pyannote accepts.
pyannote 4.x decodes audio through torchcodec, which links against a
specific FFmpeg major version and fails outright when the installed one
differs (``libavutil.56.dylib`` not found) — taking diarization down on
an otherwise working machine. Handing it an already-decoded waveform
skips that path entirely and reuses the ffmpeg extraction the acoustic
analysis already relies on, so video containers work too.
Returns ``None`` when decoding is not possible, letting the caller fall
back to passing the path and whatever pyannote can do with it.
"""
try:
import soundfile
import torch
from .voice_features import decodable_audio
with decodable_audio(path) as audio_path:
if audio_path is None:
return None
data, sample_rate = soundfile.read(audio_path, dtype="float32", always_2d=True)
# soundfile gives (samples, channels); pyannote wants (channels, samples)
return {"waveform": torch.from_numpy(data.T), "sample_rate": int(sample_rate)}
except Exception:
logger.info("could not pre-decode %s for diarization", path)
return None
def diarize(
path: str,
token: Optional[str],
@@ -67,7 +97,7 @@ def diarize(
n = str(num_speakers or "").strip()
if n.isdigit() and int(n) > 0:
kwargs["num_speakers"] = int(n)
result = pipe(path, **kwargs)
result = pipe(_load_waveform(path) or path, **kwargs)
# pyannote.audio >= 4.0 wraps the annotation; normalize to the raw one.
if hasattr(result, "exclusive_speaker_diarization"):
result = result.exclusive_speaker_diarization
+133
View File
@@ -0,0 +1,133 @@
"""Emphasis index — how much a spoken word "pops" acoustically.
Pure functions over already-extracted per-word features (energy, pitch
delta, rate delta, pause before, duration); no I/O, no external dependency.
Combines them into a single ``[0, 1]`` score, configurable via
:class:`EmphasisWeights` so the weighting can be tuned (and persisted,
see ``model_manager.load_voice_analysis_config``) without touching code.
"""
from dataclasses import dataclass
from typing import List, Sequence
_FIELDS = ("energy", "pitch_variation", "rate_variation", "pause_before", "duration")
@dataclass
class EmphasisWeights:
energy: float = 0.30
pitch_variation: float = 0.25
rate_variation: float = 0.20
pause_before: float = 0.15
duration: float = 0.10
def as_dict(self) -> dict:
return {field: getattr(self, field) for field in _FIELDS}
@classmethod
def from_dict(cls, data: dict) -> "EmphasisWeights":
defaults = cls()
return cls(**{field: float(data.get(field, getattr(defaults, field))) for field in _FIELDS})
def _clamp01(x: float) -> float:
return max(0.0, min(1.0, x))
def pause_weight(
pause_before: float, max_pause: float = 1.5, ignore_above: float = 3.0
) -> float:
"""How much a preceding silence counts as emphasis, in ``[0, 1]``.
A short beat before a word is real emphasis: the speaker is setting it
up. A *long* gap is not — it is an edit point, a B-roll insert, or the
other person in the room talking. Measured on real footage, gaps of
6-9s were scoring as the most emphatic moments in the recording purely
because the scale saturated, ranking a scene change above a word the
speaker actually hit hard.
So the contribution rises up to ``max_pause`` and then drops to zero
past ``ignore_above``, instead of saturating. Set ``ignore_above`` to
``0`` to disable the cutoff and keep the old saturating behaviour.
"""
if pause_before <= 0 or max_pause <= 0:
return 0.0
if ignore_above > 0 and pause_before > ignore_above:
return 0.0
return _clamp01(pause_before / max_pause)
def compute_emphasis(
energy: float,
pitch_delta: float,
rate_delta: float,
pause_before: float,
word_duration: float,
weights: EmphasisWeights = EmphasisWeights(),
*,
max_pause: float = 1.5,
max_duration: float = 1.0,
pause_ignore_above: float = 3.0,
) -> float:
"""Emphasis score in ``[0, 1]`` for one word.
``energy``/``pitch_delta``/``rate_delta`` are expected already
normalized to roughly ``[0, 1]`` (deltas may be negative — only their
magnitude counts as emphasis). ``pause_before``/``word_duration`` are
raw seconds; duration saturates at ``max_duration``, while the pause
contribution is shaped by :func:`pause_weight`.
"""
energy_n = _clamp01(energy)
pitch_n = _clamp01(abs(pitch_delta))
rate_n = _clamp01(abs(rate_delta))
pause_n = pause_weight(pause_before, max_pause, pause_ignore_above)
duration_n = _clamp01(word_duration / max_duration) if max_duration > 0 else 0.0
total_weight = sum(getattr(weights, field) for field in _FIELDS)
if total_weight <= 0:
return 0.0
score = (
weights.energy * energy_n
+ weights.pitch_variation * pitch_n
+ weights.rate_variation * rate_n
+ weights.pause_before * pause_n
+ weights.duration * duration_n
)
return _clamp01(score / total_weight)
def annotate_emphasis(
words: Sequence[dict],
weights: EmphasisWeights = EmphasisWeights(),
*,
max_pause: float = 1.5,
max_duration: float = 1.0,
pause_ignore_above: float = 3.0,
) -> List[dict]:
"""Return copies of ``words`` with an ``"emphasis"`` key added.
Each word dict is expected to carry ``energy``, ``pitch_delta``,
``rate_delta``, ``pause_before`` (all pre-computed, e.g. by
``voice_features.py``), plus ``start``/``end`` — or an explicit
``duration`` — to derive word length.
"""
out: List[dict] = []
for w in words:
ww = dict(w)
duration = ww.get("duration")
if duration is None:
duration = max(0.0, float(ww.get("end", 0.0)) - float(ww.get("start", 0.0)))
ww["emphasis"] = compute_emphasis(
energy=float(ww.get("energy") or 0.0),
pitch_delta=float(ww.get("pitch_delta") or 0.0),
rate_delta=float(ww.get("rate_delta") or 0.0),
pause_before=float(ww.get("pause_before") or 0.0),
word_duration=float(duration),
weights=weights,
max_pause=max_pause,
max_duration=max_duration,
pause_ignore_above=pause_ignore_above,
)
out.append(ww)
return out
+288
View File
@@ -359,3 +359,291 @@ def save_num_speakers(num: str) -> str:
data["num_speakers"] = val
_write_config(data)
return val
DEFAULT_VOICE_ANALYSIS_CONFIG: dict = {
"energy_threshold": 0.5,
"emphasis_weights": {
"energy": 0.30,
"pitch_variation": 0.25,
"rate_variation": 0.20,
"pause_before": 0.15,
"duration": 0.10,
},
# Peaks are selected RELATIVELY — the top slice of the distribution —
# because the emphasis index is a weighted average whose real range
# depends on the material. Measured on a 17-minute interview the index
# never passed 0.55, so any absolute cutoff near the spec's 0.85 selects
# nothing; on punchier material the same cutoff would flood the edit.
# 2% of words is roughly one highlight every 50 words.
"peak_percentile": 0.02,
# Guard for genuinely flat audio, where even the top of the distribution
# carries no emphasis worth cutting on.
"emphasis_floor": 0.25,
"emotion_enabled": False,
"emotion_sensitivity": 0.5,
}
def load_voice_analysis_config() -> dict:
"""The persisted voice-analysis thresholds/weights, merged over defaults.
Backs the "Análise de Voz" settings screen: energy threshold (how loud
counts as "high energy"), the emphasis-index weights (see
``emphasis.EmphasisWeights``), the punch-in emphasis cutoff, and the
emotion-detection toggle/sensitivity. Unknown/malformed stored values
fall back to the default rather than raising, so a hand-edited or
partially-written config.json never breaks the settings screen.
"""
cfg = {
**DEFAULT_VOICE_ANALYSIS_CONFIG,
"emphasis_weights": dict(DEFAULT_VOICE_ANALYSIS_CONFIG["emphasis_weights"]),
}
stored = _load_config().get("voice_analysis")
if not isinstance(stored, dict):
return cfg
for key in ("energy_threshold", "peak_percentile", "emphasis_floor", "emotion_sensitivity"):
if key in stored:
try:
cfg[key] = max(0.0, min(1.0, float(stored[key])))
except (TypeError, ValueError):
pass
if "emotion_enabled" in stored:
cfg["emotion_enabled"] = bool(stored["emotion_enabled"])
weights = stored.get("emphasis_weights")
if isinstance(weights, dict):
for key in cfg["emphasis_weights"]:
if key in weights:
try:
cfg["emphasis_weights"][key] = max(0.0, float(weights[key]))
except (TypeError, ValueError):
pass
return cfg
def save_voice_analysis_config(
energy_threshold: float | None = None,
emphasis_weights: dict | None = None,
peak_percentile: float | None = None,
emphasis_floor: float | None = None,
emotion_enabled: bool | None = None,
emotion_sensitivity: float | None = None,
) -> dict:
"""Persist voice-analysis thresholds/weights. Only given fields change.
Returns the full merged config (same shape as
:func:`load_voice_analysis_config`) so callers can render it back
immediately without a second round-trip.
"""
cfg = load_voice_analysis_config()
if energy_threshold is not None:
cfg["energy_threshold"] = max(0.0, min(1.0, float(energy_threshold)))
if peak_percentile is not None:
cfg["peak_percentile"] = max(0.0, min(1.0, float(peak_percentile)))
if emphasis_floor is not None:
cfg["emphasis_floor"] = max(0.0, min(1.0, float(emphasis_floor)))
if emotion_enabled is not None:
cfg["emotion_enabled"] = bool(emotion_enabled)
if emotion_sensitivity is not None:
cfg["emotion_sensitivity"] = max(0.0, min(1.0, float(emotion_sensitivity)))
if emphasis_weights is not None:
for key, value in emphasis_weights.items():
if key in cfg["emphasis_weights"] and value is not None:
cfg["emphasis_weights"][key] = max(0.0, float(value))
data = _load_config()
data["voice_analysis"] = cfg
_write_config(data)
return cfg
# Mirrors the "Legendas Dinâmicas" tab's own defaults (MacApp/Sources/
# CaptionsView.swift), so a fresh install shows the same look in the UI and
# in what generate_dynamic_subtitles renders when no override is passed.
DEFAULT_DYNAMIC_SUBTITLE_CONFIG: dict = {
"band_height": 0.22,
"block_center_y": -167.0,
"line_gap": 8.0,
"font": "Helvetica Neue",
"font_size": 104,
"emphasis_font": "Playfair Display",
"emphasis_face": "Medium Italic",
"emphasis_size": 265,
"active_color": "1 1 1 1",
"emphasis_color": "1 1 1 1",
"text_scale": 2.0,
}
def load_dynamic_subtitle_config() -> dict:
"""The persisted dynamic-subtitle style, merged over defaults.
Backs the "Legendas Dinâmicas" settings screen AND is the fallback
``generate_dynamic_subtitles`` reads for any field the caller doesn't
explicitly override — so the style configured in the UI is what actually
renders, without the app having to thread every field through each call.
Unknown/malformed stored values fall back to the default, same as
:func:`load_voice_analysis_config`.
"""
cfg = dict(DEFAULT_DYNAMIC_SUBTITLE_CONFIG)
stored = _load_config().get("dynamic_subtitles")
if not isinstance(stored, dict):
return cfg
for key in ("band_height", "block_center_y", "line_gap", "text_scale"):
if key in stored:
try:
cfg[key] = float(stored[key])
except (TypeError, ValueError):
pass
for key in ("font_size", "emphasis_size"):
if key in stored:
try:
cfg[key] = int(stored[key])
except (TypeError, ValueError):
pass
for key in ("font", "emphasis_font", "emphasis_face", "active_color", "emphasis_color"):
if key in stored and isinstance(stored[key], str) and stored[key]:
cfg[key] = stored[key]
return cfg
def save_dynamic_subtitle_config(**fields) -> dict:
"""Persist dynamic-subtitle style fields. Only given fields change.
Accepts the same keys as :data:`DEFAULT_DYNAMIC_SUBTITLE_CONFIG`; unknown
keys are ignored so a newer app talking to an older config shape degrades
quietly. Returns the full merged config, mirroring
:func:`save_voice_analysis_config`.
"""
cfg = load_dynamic_subtitle_config()
for key, value in fields.items():
if key not in DEFAULT_DYNAMIC_SUBTITLE_CONFIG or value is None:
continue
if isinstance(DEFAULT_DYNAMIC_SUBTITLE_CONFIG[key], float):
try:
cfg[key] = float(value)
except (TypeError, ValueError):
continue
elif isinstance(DEFAULT_DYNAMIC_SUBTITLE_CONFIG[key], int):
try:
cfg[key] = int(value)
except (TypeError, ValueError):
continue
else:
cfg[key] = str(value)
data = _load_config()
data["dynamic_subtitles"] = cfg
_write_config(data)
return cfg
# Mirrors the silence thresholds the detection/removal handlers use when no
# argument is passed (server_tools/qc.py). Persisted so the app's slider and
# any later run agree without threading three fields through every call.
DEFAULT_SILENCE_CONFIG: dict = {
# dBFS below which audio counts as silence.
"noise_db": -30.0,
# Seconds a quiet stretch must last before it's a cut candidate.
"min_silence": 0.5,
# Seconds left inside each cut so speech never gets clipped at the edges.
"padding": 0.05,
}
def load_silence_config() -> dict:
"""The persisted silence-detection thresholds, merged over defaults.
Read by ``detect_media_silence``/``remove_media_silence`` as their
fallback, so the tolerance chosen in the app is what actually runs.
Malformed stored values fall back to the default rather than raising,
matching :func:`load_voice_analysis_config`.
"""
cfg = dict(DEFAULT_SILENCE_CONFIG)
stored = _load_config().get("silence")
if not isinstance(stored, dict):
return cfg
for key in cfg:
if key in stored:
try:
cfg[key] = float(stored[key])
except (TypeError, ValueError):
pass
return cfg
def save_silence_config(
noise_db: float | None = None,
min_silence: float | None = None,
padding: float | None = None,
) -> dict:
"""Persist silence thresholds. Only the given fields change.
Values are clamped to the same ranges the handlers validate against, so
a bad write here can't produce a config the tools would later reject.
"""
cfg = load_silence_config()
if noise_db is not None:
try:
cfg["noise_db"] = max(-120.0, min(0.0, float(noise_db)))
except (TypeError, ValueError):
pass
if min_silence is not None:
try:
cfg["min_silence"] = max(0.01, min(3600.0, float(min_silence)))
except (TypeError, ValueError):
pass
if padding is not None:
try:
cfg["padding"] = max(0.0, min(5.0, float(padding)))
except (TypeError, ValueError):
pass
data = _load_config()
data["silence"] = cfg
_write_config(data)
return cfg
# Last project worked on, so the app reopens where the user left off instead of
# making them pick the folder again every launch. Only paths that still exist
# are handed back — a project on an unmounted volume degrades to "none selected"
# rather than to a dead path the tools would later fail on.
DEFAULT_PROJECT_CONFIG: dict = {
# Folder every generated file (transcript .json, XML, SRT) is written to.
"folder": "",
# The .fcpxml/.fcpxmld that was loaded from it.
"file": "",
}
def load_project_config() -> dict:
"""The persisted last project (folder + file), merged over defaults.
Paths that no longer exist on disk come back empty, matching what the app
shows for "nothing selected". Malformed stored values fall back to the
default rather than raising, same as :func:`load_voice_analysis_config`.
"""
cfg = dict(DEFAULT_PROJECT_CONFIG)
stored = _load_config().get("project")
if not isinstance(stored, dict):
return cfg
for key in cfg:
value = stored.get(key)
if isinstance(value, str) and value and Path(value).exists():
cfg[key] = value
return cfg
def save_project_config(folder: str | None = None, file: str | None = None) -> dict:
"""Persist the last project folder/file. Only the given fields change.
Passing an empty string clears a field (the app does this when the user
deselects), while ``None`` leaves it untouched.
"""
cfg = load_project_config()
for key, value in (("folder", folder), ("file", file)):
if value is None:
continue
cfg[key] = str(Path(value).expanduser()) if str(value).strip() else ""
data = _load_config()
data["project"] = cfg
_write_config(data)
return cfg
+18
View File
@@ -13,6 +13,8 @@ from functools import total_ordering
from math import gcd
from typing import Any, Callable, Dict, List, Optional, Tuple
from .text_layout import REFERENCE_BLOCK_LINE_GAP, TEXT_TEMPLATE_FONT_SCALE
# ============================================================================
# ENUMS
# ============================================================================
@@ -1071,3 +1073,19 @@ class DynamicSubtitleConfig:
# grouped, the key word alone and large (the reference look). "word": one
# title per word, the earlier rhythm.
granularity: str = "phrase"
# Ratio between the template's fontSize space and the canvas-point space
# its Position uses. See text_layout.TEXT_TEMPLATE_FONT_SCALE: the "Text"
# (Text.moti) template sizes type in frame pixels, so a size chosen in
# points renders half as large unless it is converted on the way out.
text_scale: float = TEXT_TEMPLATE_FONT_SCALE
# Vertical air between stacked lines, in canvas points. Negative values
# deliberately overlap the lines — the display italic tucking under the
# line above is a real editorial look, and the stacking arithmetic places
# ink boxes edge to edge, so a negative gap moves them by exactly that
# much rather than colliding unpredictably.
line_gap: float = REFERENCE_BLOCK_LINE_GAP
# Run the post-generation collision validation (collision.validate_titles)
# and refuse to emit when it reports a blocking overlap. Off by default so
# generation stays byte-identical to before this flag existed; flip it on
# for a guaranteed no-collision export.
validate: bool = False
+91 -16
View File
@@ -122,6 +122,33 @@ REFERENCE_BLOCK_LINE_GAP = 8.0
# emphasis line's edge; the reference leaves a little air.
REFERENCE_STAGGER_RATIO = 0.8
# Extra gap, as a fraction of the emphasis line's font size, added only to
# the boundary right below it. The display italic's slant leans its stems
# past the vertical ink box the metrics measure, so a body line directly
# under the emphasis line reads tighter than the same nominal gap anywhere
# else in the stack — this cushion (~14pt at the 230pt reference size)
# closes that optical gap without touching the user's `line_gap` elsewhere.
_EMPHASIS_ITALIC_CUSHION_RATIO = 0.06
# The numbers above were read off a hand export that used the "Essencial -
# Título" template. That template never rendered when we generated it (see
# Engine/docs/05_EXPERIENCIAS.md, 2026-08-17), so the writer switched to FCP's
# own "Basic Text > Text" (Text.moti) — whose coordinate space is the FRAME
# ITSELF (2160x3840), not the half-scale point canvas the numbers above were
# measured in. Everything the template reads is in that space: fontSize,
# kerning AND Position alike.
#
# Getting this half-right is worse than getting it wrong. Scaling only the type
# left the block at the old spread with twice the type in it, so the lines
# collided; scaling only the positions would spread a block of half-size type
# across the frame. The layout keeps measuring in canvas points — every
# constant above depends on that — and this single factor converts the whole
# result on the way out, which is the only way the two stay in step.
#
# Exposed as `text_scale` on DynamicSubtitleConfig for a template authored
# against a different space.
TEXT_TEMPLATE_FONT_SCALE = 2.0
def metrics_for(font: Optional[str], face: Optional[str] = None) -> Optional[Dict]:
"""The embedded advance table for *font*/*face*, or None if uncovered.
@@ -296,9 +323,14 @@ class PlacedWord:
and other.bottom < self.top
)
def position_param(self) -> str:
"""The value for the title's "Posição" param, as FCP writes it."""
return f"{self.x:g} {self.y:g}"
def position_param(self, scale: float = 1.0) -> str:
"""The value for the title's "Posição" param, as FCP writes it.
*scale* converts from canvas points to the template's own space; see
TEXT_TEMPLATE_FONT_SCALE. It must be the same factor the emitted
fontSize uses, or the type and the spacing drift apart.
"""
return f"{self.x * scale:g} {self.y * scale:g}"
# The rest of this block is the interface a placed unit shares with
# PlacedBlock, so the writer emits titles from either without caring
@@ -649,9 +681,14 @@ class PlacedBlock:
"""When this block finishes being spoken, in seconds."""
return max(float(w.get('end', 0.0)) for w in self.words)
def position_param(self) -> str:
"""The value for the title's "Posição" param, as FCP writes it."""
return f"{self.x:g} {self.y:g}"
def position_param(self, scale: float = 1.0) -> str:
"""The value for the title's "Posição" param, as FCP writes it.
*scale* converts from canvas points to the template's own space; see
TEXT_TEMPLATE_FONT_SCALE. It must be the same factor the emitted
fontSize uses, or the type and the spacing drift apart.
"""
return f"{self.x * scale:g} {self.y * scale:g}"
def overlaps(self, other: 'PlacedBlock') -> bool:
return (
@@ -742,10 +779,33 @@ def compose_sentence(
for run in body_lines(entries[emphasis_index + 1:]):
lines.append((run, body_look, False))
# A body run wraps onto a new line when it doesn't fit — but the
# emphasis line is always exactly one word, so it can't wrap, and
# nothing capped its size against the box. A long or all-caps word (an
# emphasis pass sometimes upper-cases its pick) could run past both
# edges of the frame — found on real footage, wide enough to spill off
# BOTH sides while centred. Shrinking it back to the box scales its
# font_size and kerning by the same factor, so the ink height used for
# stacking below shrinks with it too — restoring the vertical
# non-overlap the rest of this function already guarantees by
# construction. Never shrunk below the body size: emphasis smaller
# than body text isn't emphasis anymore, it's just a different font.
def fit_emphasis(text: str, look) -> tuple:
size, kerning, width = measure(text, look)
if width <= box.width:
return size, kerning, width
floor = float(body_look.font_size) * scale
fit = max(box.width / width, floor / size) if size > 0 else 1.0
fit = min(fit, 1.0)
return size * fit, kerning * fit, width * fit
measured = []
for run, look, is_emphasis in lines:
text = ' '.join(t for _, t in run)
size, kerning, width = measure(text, look)
if is_emphasis:
size, kerning, width = fit_emphasis(text, look)
else:
size, kerning, width = measure(text, look)
# Stack on the real ink each line contains, not on a nominal
# cap-height: the display italic's accents and descenders run well
# past it, and a nominal box lets them collide with the neighbour.
@@ -764,14 +824,29 @@ def compose_sentence(
# loses the whole point of the look — so if it does not fit, everything
# from the emphasis on overflows together.
gap = line_gap * scale
# The emphasis line's italic slant carries visual weight below its own
# ink box — Playfair's stems lean past what the vertical metrics measure
# — so a body line sitting right under it reads tighter than the same
# nominal gap elsewhere, even though the ink boxes themselves never
# touch. Add a size-proportional cushion only to the boundary right
# after the emphasis line; every other pair keeps exactly the caller's
# ``line_gap``.
def pair_gap(prev_line: dict) -> float:
if prev_line['emphasis']:
return gap + _EMPHASIS_ITALIC_CUSHION_RATIO * prev_line['font_size']
return gap
kept = 0
total = 0.0
prev = None
for line in measured:
advance = line['height'] if not kept else line['height'] + gap
if kept and total + advance > box.height:
advance = line['height'] if prev is None else line['height'] + pair_gap(prev)
if prev is not None and total + advance > box.height:
break
total += advance
kept += 1
prev = line
kept = max(kept, 1)
if not any(line['emphasis'] for line in measured[:kept]):
kept = min(kept, next(
@@ -783,12 +858,12 @@ def compose_sentence(
result.overflow.extend(w for w, _ in line['run'])
visible = measured[:kept]
# Stack the ink boxes edge to edge with exactly *gap* between them, then
# centre the whole stack on the band. Because the boxes are the real ink,
# "no overlap" is a property of the arithmetic, not of a safety factor.
stack_height = (
sum(line['height'] for line in visible) + gap * (len(visible) - 1)
)
# Stack the ink boxes edge to edge with exactly *gap* between them (plus
# the emphasis cushion where it applies), then centre the whole stack on
# the band. Because the boxes are the real ink, "no overlap" is a
# property of the arithmetic, not of a safety factor.
gaps = [pair_gap(visible[i - 1]) for i in range(1, len(visible))]
stack_height = sum(line['height'] for line in visible) + sum(gaps)
edge = box.center_y + stack_height / 2
# Body lines hang off the emphasis line's edges, alternating sides in
@@ -797,7 +872,7 @@ def compose_sentence(
side = -1
for index, line in enumerate(visible):
if index:
edge -= gap
edge -= gaps[index - 1]
cursor_y = edge - line['ink_top']
edge = cursor_y + line['ink_bottom']
if line['emphasis']:
+11 -2
View File
@@ -16,7 +16,7 @@ import logging
import os
import re
from pathlib import Path
from typing import List, Optional, Sequence, Tuple
from typing import Callable, List, Optional, Sequence, Tuple
logger = logging.getLogger(__name__)
@@ -118,7 +118,10 @@ def invert_ranges(
def transcribe(
path: str, model_size: str = "base", language: Optional[str] = None
path: str,
model_size: str = "base",
language: Optional[str] = None,
progress_cb: Optional[Callable[[float], None]] = None,
) -> Optional[dict]:
"""Transcribe an audio/video file locally with word-level timestamps.
@@ -170,6 +173,10 @@ def transcribe(
)
segments: List[dict] = []
words: List[dict] = []
# `info.duration` is known upfront (from the container), so each
# segment's end time — yielded lazily as faster-whisper decodes —
# gives real, granular progress instead of a single before/after step.
total_duration = float(info.duration) if info.duration else 0.0
for seg in segments_iter:
start = float(seg.start)
end = float(seg.end)
@@ -182,6 +189,8 @@ def transcribe(
"end_fmt": format_timestamp(end),
}
)
if progress_cb is not None and total_duration > 0:
progress_cb(min(end / total_duration, 1.0))
for w in seg.words or []:
ws = float(w.start)
we = float(w.end)
+248
View File
@@ -0,0 +1,248 @@
"""Voice actions — the editing decisions produced from a voice timeline.
This is the contract between *deciding* and *applying*. Whoever makes the
editorial call — the deterministic rules engine, or a model reading the
voice timeline JSON — emits the same list of actions, and one applier turns
it into FCPXML. Nothing that produces actions ever touches XML.
Every action's ``start``/``end`` is in **original source seconds**, matching
the voice timeline. That matters: cuts shift everything after them, so if
decisions were expressed in post-cut time they would silently land in the
wrong place the moment a cut was added. Keeping one origin and resolving the
shift at apply time (:func:`shift_after_cuts`) removes that whole class of bug.
Actions arriving from a model are untrusted input: :func:`parse_actions`
validates and reports what it rejected rather than raising, so one malformed
row never discards a whole edit.
"""
from dataclasses import dataclass, field
from typing import Any, List, Optional, Sequence, Tuple
# What an action can ask for. Deliberately small — each maps onto one
# existing writer capability, so no new XML knowledge lives here.
ACTION_KINDS = ("cut", "zoom", "text", "marker")
# Bounds for a zoom's scale factor. Below 1.0 is a pull-back, not a punch-in;
# above 3x the image falls apart on any normal footage.
MIN_ZOOM_SCALE = 1.0
MAX_ZOOM_SCALE = 3.0
MAX_TEXT_LENGTH = 120
@dataclass
class VoiceAction:
"""One editing decision, in original source time."""
kind: str
start: float
end: float
params: dict = field(default_factory=dict)
reason: str = ""
speaker: str = ""
@property
def duration(self) -> float:
return max(0.0, self.end - self.start)
def as_dict(self) -> dict:
return {
"kind": self.kind,
"start": round(self.start, 3),
"end": round(self.end, 3),
"params": self.params,
"reason": self.reason,
"speaker": self.speaker,
}
def _validate_one(raw: Any, index: int) -> Tuple[Optional[VoiceAction], str]:
"""Turn one raw row into a VoiceAction, or explain why it can't be."""
where = f"action[{index}]"
if not isinstance(raw, dict):
return None, f"{where}: expected an object, got {type(raw).__name__}"
kind = str(raw.get("kind", "")).strip().lower()
if kind not in ACTION_KINDS:
return None, f"{where}: unknown kind {raw.get('kind')!r} (expected one of {', '.join(ACTION_KINDS)})"
try:
start = float(raw.get("start"))
end = float(raw.get("end"))
except (TypeError, ValueError):
return None, f"{where}: start/end must be numbers (seconds)"
if start < 0:
return None, f"{where}: start is negative ({start})"
if end <= start:
return None, f"{where}: end ({end}) must be after start ({start})"
params = raw.get("params")
params = dict(params) if isinstance(params, dict) else {}
if kind == "zoom":
try:
scale = float(params.get("scale", 1.3))
except (TypeError, ValueError):
return None, f"{where}: zoom scale must be a number"
if not (MIN_ZOOM_SCALE <= scale <= MAX_ZOOM_SCALE):
return None, (
f"{where}: zoom scale {scale} outside {MIN_ZOOM_SCALE}-{MAX_ZOOM_SCALE}"
)
params["scale"] = scale
if kind == "text":
content = str(params.get("content", "")).strip()
if not content:
return None, f"{where}: text action needs params.content"
params["content"] = content[:MAX_TEXT_LENGTH]
return (
VoiceAction(
kind=kind,
start=start,
end=end,
params=params,
reason=str(raw.get("reason", "")),
speaker=str(raw.get("speaker", "")),
),
"",
)
def parse_actions(data: Any) -> Tuple[List[VoiceAction], List[str]]:
"""Validate a decision list into actions, collecting rejections.
Accepts either a bare list of actions or ``{"actions": [...]}`` — the
shape a model is most likely to return. Returns ``(actions, errors)``;
a row that fails validation is reported and skipped, never fatal.
"""
if isinstance(data, dict):
data = data.get("actions", [])
if not isinstance(data, Sequence) or isinstance(data, (str, bytes)):
return [], ["expected a list of actions, or an object with an 'actions' list"]
actions: List[VoiceAction] = []
errors: List[str] = []
for i, raw in enumerate(data):
action, error = _validate_one(raw, i)
if action is not None:
actions.append(action)
else:
errors.append(error)
return actions, errors
def speaker_cut_actions(
timeline: dict,
speaker_ids: Sequence[str],
padding: float = 0.15,
) -> List[VoiceAction]:
"""Cut actions removing everything the given speakers say.
The everyday case on a testimonial shoot: an interviewer or a crew
member talks over the take, and only the subject should survive the
edit. ``padding`` trims slightly *inside* each segment rather than
around it — speech boundaries from a transcript are approximate, and
eating into the neighbouring silence is far safer than clipping the
first syllable of the person being kept.
"""
wanted = {str(s) for s in speaker_ids}
actions: List[VoiceAction] = []
for segment in timeline.get("segments", []):
if str(segment.get("speaker", "")) not in wanted:
continue
start = float(segment.get("start", 0.0)) + padding
end = float(segment.get("end", 0.0)) - padding
if end <= start:
continue
actions.append(
VoiceAction(
kind="cut",
start=start,
end=end,
reason=f"fala de {segment.get('speaker')}",
speaker=str(segment.get("speaker", "")),
)
)
return actions
def merge_cut_ranges(actions: Sequence[VoiceAction]) -> List[Tuple[float, float]]:
"""The cut actions as merged, sorted, non-overlapping source ranges."""
cuts = sorted((a.start, a.end) for a in actions if a.kind == "cut")
merged: List[Tuple[float, float]] = []
for start, end in cuts:
if merged and start <= merged[-1][1]:
merged[-1] = (merged[-1][0], max(merged[-1][1], end))
else:
merged.append((start, end))
return merged
def shift_after_cuts(
time: float, cuts: Sequence[Tuple[float, float]]
) -> Optional[float]:
"""Where source ``time`` lands once ``cuts`` are removed.
Returns ``None`` when the time falls *inside* a cut — the material it
referred to no longer exists, so the action that pointed at it must be
dropped rather than silently slid onto neighbouring content.
``cuts`` must be merged and sorted (see :func:`merge_cut_ranges`).
"""
shift = 0.0
for start, end in cuts:
if time < start:
break
if time < end:
return None
shift += end - start
return time - shift
def resolve_actions(
actions: Sequence[VoiceAction],
) -> Tuple[List[Tuple[float, float]], List[VoiceAction], List[VoiceAction]]:
"""Split a decision list into what to cut and what to place afterwards.
Returns ``(cut_ranges, placed, dropped)``. Non-cut actions are moved onto
their post-cut times; any that pointed into removed material land in
``dropped`` so the caller can report them instead of losing them quietly.
"""
cut_ranges = merge_cut_ranges(actions)
placed: List[VoiceAction] = []
dropped: List[VoiceAction] = []
for action in actions:
if action.kind == "cut":
continue
new_start = shift_after_cuts(action.start, cut_ranges)
if new_start is None:
dropped.append(action)
continue
if action.kind == "marker":
# A marker is a point, not a span: it survives as long as its own
# instant does. Requiring its nominal end to survive too would
# drop exactly the markers worth keeping — the ones flagging a
# join, which sit right against a cut edge by definition.
new_end = new_start + action.duration
else:
new_end = shift_after_cuts(action.end, cut_ranges)
if new_end is None:
dropped.append(action)
continue
if new_end <= new_start:
dropped.append(action)
continue
placed.append(
VoiceAction(
kind=action.kind,
start=new_start,
end=new_end,
params=action.params,
reason=action.reason,
speaker=action.speaker,
)
)
return cut_ranges, placed, dropped
+220
View File
@@ -0,0 +1,220 @@
"""Acoustic features for voice analysis — pitch, energy, rate, pauses.
Mirrors the ``media_intel.py`` contract: librosa is an optional dependency
(``pip install 'fcp-mcp-server[intelligence]'``, already required by beat
detection), imported lazily, and every extractor degrades to ``None`` when
the library is missing or the file cannot be analyzed — never crashes.
``compute_speech_rate``/``compute_pauses`` are pure functions over
word-timestamp dicts (the shape ``transcribe.py`` already produces) and need
no audio file at all.
"""
import contextlib
import logging
import shutil
import subprocess
import tempfile
from pathlib import Path
from typing import Iterator, List, Optional, Sequence, Tuple
logger = logging.getLogger(__name__)
# Human voice fundamental frequency range (covers low male to high female/child).
PITCH_FMIN_HZ = 65.0
PITCH_FMAX_HZ = 1000.0
# Formats librosa reads directly through soundfile. Anything else — notably
# the .mov/.mp4 that source footage actually arrives in — must be decoded by
# ffmpeg first, or analysis fails outright.
NATIVE_AUDIO_SUFFIXES = {".wav", ".aif", ".aiff", ".flac"}
# Voice analysis only needs the speech band: 16 kHz mono is well above the
# Nyquist limit for our 1 kHz pitch ceiling, and keeps the extracted file
# small and fast to decode even for hour-long footage.
EXTRACT_SAMPLE_RATE = 16000
EXTRACT_TIMEOUT_SECONDS = 600
@contextlib.contextmanager
def decodable_audio(path: str) -> Iterator[Optional[str]]:
"""Yield a path librosa can read, extracting the audio track if needed.
Audio files pass straight through. Video containers are decoded to a
temporary mono WAV with ffmpeg and cleaned up on exit. Yields ``None``
when the audio cannot be obtained (no ffmpeg, no audio track, failure),
keeping the graceful-degradation contract of this module.
"""
file_path = Path(path)
if file_path.suffix.lower() in NATIVE_AUDIO_SUFFIXES:
yield str(file_path)
return
if shutil.which("ffmpeg") is None:
logger.info("ffmpeg not found on PATH; cannot extract audio from %s", file_path)
yield None
return
tmp_dir = tempfile.mkdtemp(prefix="fcp_voice_")
wav_path = Path(tmp_dir) / "audio.wav"
try:
result = subprocess.run(
[
"ffmpeg", "-hide_banner", "-nostdin", "-y",
"-i", str(file_path),
"-vn", # audio only: decoding video would dominate the runtime
"-ac", "1",
"-ar", str(EXTRACT_SAMPLE_RATE),
str(wav_path),
],
capture_output=True,
text=True,
timeout=EXTRACT_TIMEOUT_SECONDS,
)
if result.returncode != 0 or not wav_path.is_file():
logger.warning("ffmpeg could not extract audio from %s", file_path)
yield None
else:
yield str(wav_path)
except (OSError, subprocess.TimeoutExpired):
logger.warning("audio extraction failed for %s", file_path)
yield None
finally:
shutil.rmtree(tmp_dir, ignore_errors=True)
def features_capability() -> Tuple[bool, str]:
"""Whether pitch/energy extraction is available (librosa installed)."""
try:
import librosa # noqa: F401
except Exception:
return False, "Análise acústica indisponível: componente librosa ausente."
return True, "Análise acústica disponível."
def extract_pitch(
path: str, hop_length: int = 512, max_analysis_seconds: float = 1200.0
) -> Optional[List[Tuple[float, float]]]:
"""Frame-level pitch (F0) track via librosa's ``pyin``.
Returns ``[(time_seconds, hz), ...]`` for voiced frames only (unvoiced
frames, where ``pyin`` reports no pitch, are dropped), or ``None`` when
librosa is unavailable or the file cannot be analyzed.
"""
file_path = Path(path)
if not file_path.is_file():
return None
try:
import librosa
except ImportError:
logger.info("librosa not installed; pitch extraction unavailable")
return None
try:
with decodable_audio(str(file_path)) as audio_path:
if audio_path is None:
return None
y, sr = librosa.load(audio_path, sr=None, mono=True, duration=max_analysis_seconds)
f0, voiced_flag, _voiced_prob = librosa.pyin(
y, fmin=PITCH_FMIN_HZ, fmax=PITCH_FMAX_HZ, sr=sr, hop_length=hop_length
)
times = librosa.times_like(f0, sr=sr, hop_length=hop_length)
except Exception:
logger.warning("librosa pitch analysis failed for %s", file_path)
return None
return [
(float(t), float(hz))
for t, hz, voiced in zip(times, f0, voiced_flag)
if voiced and hz == hz # ``hz == hz`` filters NaN without importing math/numpy here
]
def extract_energy(
path: str, hop_length: int = 512, max_analysis_seconds: float = 1200.0
) -> Optional[List[Tuple[float, float]]]:
"""Frame-level RMS energy track via librosa.
Returns ``[(time_seconds, rms), ...]``, or ``None`` when librosa is
unavailable or the file cannot be analyzed.
"""
file_path = Path(path)
if not file_path.is_file():
return None
try:
import librosa
except ImportError:
logger.info("librosa not installed; energy extraction unavailable")
return None
try:
with decodable_audio(str(file_path)) as audio_path:
if audio_path is None:
return None
y, sr = librosa.load(audio_path, sr=None, mono=True, duration=max_analysis_seconds)
rms = librosa.feature.rms(y=y, hop_length=hop_length)[0]
times = librosa.times_like(rms, sr=sr, hop_length=hop_length)
except Exception:
logger.warning("librosa energy analysis failed for %s", file_path)
return None
return [(float(t), float(r)) for t, r in zip(times, rms)]
def _window_average(track: Sequence[Tuple[float, float]], start: float, end: float) -> Optional[float]:
"""Average of ``track`` values whose timestamp falls in ``[start, end]``."""
values = [v for t, v in track if start <= t <= end]
if not values:
return None
return sum(values) / len(values)
def word_pitch_energy(
words: Sequence[dict],
pitch_track: Optional[Sequence[Tuple[float, float]]],
energy_track: Optional[Sequence[Tuple[float, float]]],
) -> List[dict]:
"""Attach average pitch/energy over each word's ``[start, end]`` span.
Words carry ``pitch_hz``/``energy`` (``None`` when the span has no
voiced frames or a track is unavailable). Both tracks are the output of
:func:`extract_pitch`/:func:`extract_energy`.
"""
out: List[dict] = []
for w in words:
ww = dict(w)
start = float(w.get("start", 0.0))
end = float(w.get("end", start))
ww["pitch_hz"] = _window_average(pitch_track, start, end) if pitch_track else None
ww["energy"] = _window_average(energy_track, start, end) if energy_track else None
out.append(ww)
return out
def compute_speech_rate(words: Sequence[dict], window_seconds: float = 3.0) -> List[float]:
"""Local speech rate (words/second) around each word.
For word *i*, counts every word whose start falls within
``[start_i - window_seconds, start_i]`` and divides by
``window_seconds`` — a trailing local rate, cheap to compute and stable
against a single long/short word skewing the whole utterance's average.
"""
starts = [float(w.get("start", 0.0)) for w in words]
rates: List[float] = []
for i, s in enumerate(starts):
lo = s - window_seconds
count = sum(1 for t in starts[: i + 1] if t >= lo)
rates.append(count / window_seconds if window_seconds > 0 else 0.0)
return rates
def compute_pauses(words: Sequence[dict]) -> List[float]:
"""Silence (seconds) immediately before each word.
The first word's "pause before" is the time from the start of the audio
to its own start; every other word measures the gap since the previous
word's end (clamped to ``0`` for overlapping/adjacent words).
"""
pauses: List[float] = []
prev_end = 0.0
for w in words:
start = float(w.get("start", 0.0))
pauses.append(max(0.0, start - prev_end))
prev_end = float(w.get("end", start))
return pauses
+505
View File
@@ -0,0 +1,505 @@
"""Voice timeline — the consolidated, AI-readable view of how a video is spoken.
This is the *source of truth* between analysis and editing: it merges what
was said (transcript), who said it (diarization), and how it was said
(pitch/energy/rate/pauses → emphasis) into one JSON document, decoupling the
audio analysis from FCPXML generation entirely.
The shape is designed to be handed to a language model so it can reason about
the narrative — which beats carry weight, where a speaker changes, where the
delivery peaks — and decide how to direct the edit. Two design choices serve
that goal:
* **Layered, not flat.** A ``summary`` gives the whole picture in a few
numbers, ``segments`` group words into utterances with their own
aggregates, and ``words`` hold the fine detail. A model can reason from
the top layer and only descend where it matters, instead of parsing
thousands of word rows to find the shape of the piece.
* **Normalized, self-describing values.** Every acoustic value is 0–1 and
relative to *this* recording (a quiet podcast and a shouted ad both use
the full range), and ``scales`` documents that contract inline, so the
numbers are interpretable without external context.
"""
import json
import logging
from pathlib import Path
from typing import Callable, List, Optional, Sequence, Tuple
from .diarize import DEFAULT_SPEAKER, assign_speakers, build_speakers, diarize
from .emphasis import EmphasisWeights, annotate_emphasis
from .voice_features import (
compute_pauses,
compute_speech_rate,
extract_energy,
extract_pitch,
word_pitch_energy,
)
logger = logging.getLogger(__name__)
VOICE_TIMELINE_VERSION = "1.0"
# Silence long enough to mean the take stopped rather than the speaker paused.
# On real footage, boundaries between retakes showed gaps of 3.6-19.8s while
# dramatic beats inside a delivered line stayed under ~2s.
TAKE_BOUNDARY_GAP = 3.0
# How to read the values in this document, split by the level they live on.
# Embedded in the output so a model consuming the JSON needs no external
# documentation — and kept honest: a metric listed under "word" must exist on
# every word row, and one under "segment" on every segment row.
VALUE_SCALES = {
"word": {
"energy": "0-1, loudness relative to the loudest moment of this recording",
"pitch_delta": "0-1, how far this word's pitch sits from the speaker's average",
"rate_delta": "0-1, how much the local speaking rate departs from the average",
"pause_before": "seconds of silence immediately before the word",
"emphasis": "0-1 combined index; high values are punch-in/highlight candidates",
},
"segment": {
"gap_before": "seconds of silence before this line",
"take_boundary": "true when the gap is long enough that the take likely restarted here",
"avg_energy": "0-1 mean loudness across the line",
"peak_emphasis": "0-1 highest emphasis of any word in the line",
},
}
def _normalize(value: Optional[float], maximum: float) -> float:
"""Scale ``value`` into 0-1 against ``maximum`` (0.0 when unavailable)."""
if value is None or maximum <= 0:
return 0.0
return max(0.0, min(1.0, value / maximum))
def _round_word(word: dict) -> dict:
"""One word row, rounded to a size a model can read without noise.
The raw ``energy_raw``/``pitch_hz`` ride along beside the normalized
values so the document can be re-analyzed over a subset later. That
matters after cutting: every normalized value is relative to the
loudest moment of the *whole* recording, and if that moment gets cut
the survivors are scored against something that no longer exists.
"""
return {
"text": word.get("word", ""),
"start": round(float(word.get("start", 0.0)), 3),
"end": round(float(word.get("end", 0.0)), 3),
"speaker": word.get("speaker_id", DEFAULT_SPEAKER),
"energy": round(word.get("energy_norm", 0.0), 3),
"pitch_delta": round(word.get("pitch_delta", 0.0), 3),
"rate_delta": round(word.get("rate_delta", 0.0), 3),
"pause_before": round(word.get("pause_before", 0.0), 3),
"emphasis": round(word.get("emphasis", 0.0), 3),
"energy_raw": word.get("energy"),
"pitch_hz": word.get("pitch_hz"),
}
def enrich_words(
words: Sequence[dict],
pitch_track: Optional[Sequence] = None,
energy_track: Optional[Sequence] = None,
weights: EmphasisWeights = EmphasisWeights(),
already_measured: bool = False,
) -> List[dict]:
"""Attach normalized acoustic features + the emphasis index to each word.
Normalization is per-recording: energy against the loudest word, pitch
against the spread around this recording's average, rate against the
largest local departure. That makes the numbers comparable within a
piece regardless of how it was recorded.
Set ``already_measured`` when the words already carry ``energy`` and
``pitch_hz`` from a previous pass — re-analyzing a subset, say. The
frame tracks are then unnecessary, and sampling them again would
overwrite good values with ``None``.
"""
if not words:
return []
if not already_measured:
words = word_pitch_energy(words, pitch_track, energy_track)
rates = compute_speech_rate(words)
pauses = compute_pauses(words)
energies = [w["energy"] for w in words if w.get("energy") is not None]
max_energy = max(energies) if energies else 0.0
pitches = [w["pitch_hz"] for w in words if w.get("pitch_hz") is not None]
avg_pitch = sum(pitches) / len(pitches) if pitches else 0.0
pitch_span = (max(pitches) - min(pitches)) if len(pitches) > 1 else 0.0
avg_rate = sum(rates) / len(rates) if rates else 0.0
max_rate = max(rates) if rates else 0.0
enriched: List[dict] = []
for i, w in enumerate(words):
ww = dict(w)
ww["energy_norm"] = _normalize(w.get("energy"), max_energy)
pitch = w.get("pitch_hz")
ww["pitch_delta"] = (
_normalize(abs(pitch - avg_pitch), pitch_span) if pitch is not None else 0.0
)
ww["rate_delta"] = _normalize(abs(rates[i] - avg_rate), max_rate)
ww["pause_before"] = pauses[i]
enriched.append(ww)
annotated = annotate_emphasis(
[{**w, "energy": w["energy_norm"]} for w in enriched], weights=weights
)
for word, scored in zip(enriched, annotated):
word["emphasis"] = scored["emphasis"]
return enriched
def _segment_rows(segments: Sequence[dict], words: Sequence[dict]) -> List[dict]:
"""Group enriched words under their segment, with per-segment aggregates.
The aggregates are what let a model judge a whole utterance ("this line
is delivered hot, that one trails off") without reading every word.
"""
rows: List[dict] = []
previous_end = 0.0
for seg in segments:
start = float(seg.get("start", 0.0))
end = float(seg.get("end", 0.0))
in_seg = [w for w in words if start <= float(w.get("start", 0.0)) < end]
energies = [w["energy_norm"] for w in in_seg]
emphases = [w["emphasis"] for w in in_seg]
gap = max(0.0, start - previous_end)
rows.append(
{
"start": round(start, 3),
"end": round(end, 3),
"speaker": seg.get("speaker_id", DEFAULT_SPEAKER),
"text": (seg.get("text") or "").strip(),
# Silence before this line. Long gaps are where the camera
# stopped or the take restarted, so this is the structural
# hint for splitting a recording into takes — the same signal
# that is *noise* for emphasis (see emphasis.pause_weight).
"gap_before": round(gap, 3),
"take_boundary": gap >= TAKE_BOUNDARY_GAP,
"avg_energy": round(sum(energies) / len(energies), 3) if energies else 0.0,
"peak_emphasis": round(max(emphases), 3) if emphases else 0.0,
"words": [_round_word(w) for w in in_seg],
}
)
previous_end = end
return rows
# Words too common to ever be the point of a punch-in. A zoom lands on what a
# sentence is *about*, and an article spoken loudly is still an article.
_FUNCTION_WORDS = {
"a", "o", "e", "de", "da", "do", "que", "é", "em", "um", "uma", "as", "os",
"no", "na", "com", "pra", "para", "por", "se", "mais", "isso", "aí", "tudo",
"ao", "à", "dos", "das", "nos", "nas", "ou", "mas", "já", "ele", "ela",
"eu", "você", "seu", "sua", "meu", "minha", "esse", "essa", "aquele",
}
def _survives(start: float, end: float, cuts: Sequence[Tuple[float, float]]) -> bool:
"""Whether a span lies entirely outside every removed range."""
return all(end <= cut_start or start >= cut_end for cut_start, cut_end in cuts)
def restrict_to_kept(
timeline: dict,
cut_ranges: Sequence[Tuple[float, float]],
weights: EmphasisWeights = EmphasisWeights(),
peak_percentile: float = 0.02,
emphasis_floor: float = 0.25,
) -> dict:
"""Re-analyze a timeline over only the material that survives ``cut_ranges``.
Emphasis is *relative*: energy is scored against the loudest word,
pitch against the spread of the recording. Cut the loudest moment out —
a laugh, an aside to the crew — and every remaining score is measured
against something the viewer will never see. Re-running the
normalization over just the survivors is what makes "the most emphatic
line of the final video" a meaningful question.
Returns a timeline of the same shape, with times still in original
source seconds so the result can be fed straight back as actions.
"""
kept_words = [
w
for segment in timeline.get("segments", [])
for w in segment.get("words", [])
if _survives(w["start"], w["end"], cut_ranges)
]
# enrich_words expects the raw analysis keys, not the normalized ones.
raw = [
{
"word": w["text"],
"start": w["start"],
"end": w["end"],
"speaker_id": w.get("speaker", DEFAULT_SPEAKER),
"energy": w.get("energy_raw"),
"pitch_hz": w.get("pitch_hz"),
}
for w in kept_words
]
enriched = enrich_words(raw, weights=weights, already_measured=True)
kept_segments = [
{**s, "words": [w for w in s.get("words", []) if _survives(w["start"], w["end"], cut_ranges)]}
for s in timeline.get("segments", [])
]
kept_segments = [s for s in kept_segments if s["words"]]
rows = _segment_rows(
[{"text": s["text"], "start": s["start"], "end": s["end"],
"speaker_id": s.get("speaker", DEFAULT_SPEAKER)} for s in kept_segments],
enriched,
)
duration = sum(s["end"] - s["start"] for s in rows)
return {
**timeline,
"summary": _summary(enriched, rows, timeline.get("speakers", []),
duration, peak_percentile, emphasis_floor),
"segments": rows,
}
def sentence_end(segments: Sequence[dict], index: int) -> float:
"""Where the sentence starting at ``segments[index]`` actually finishes.
Transcription segments break on breath and timing, not on grammar — a
sentence routinely spans two or three of them ("…que dá aquele ar" /
"de elegância, isso é desejo de muitas mulheres, né?"). A zoom that
ends on a segment boundary would therefore release mid-thought, so the
window is extended until a segment closes with terminal punctuation.
"""
last = float(segments[index]["end"])
for offset, segment in enumerate(segments[index:]):
# A long gap means the take stopped; never run a zoom across that.
# Checked before adopting the end, or the boundary segment's own
# end would already have been taken.
if offset > 0 and segment.get("take_boundary"):
break
last = float(segment["end"])
if (segment.get("text") or "").strip().endswith((".", "!", "?", "…")):
break
return last
def suggest_zoom_windows(
timeline: dict,
min_gap: float = 8.0,
max_zooms: Optional[int] = None,
) -> List[dict]:
"""Propose punch-in windows over a timeline's strongest lines.
One zoom per line at most, taken from the line's most emphatic
*content* word — a loudly spoken "a" is still an article, so function
words are skipped. The window runs from that word to the end of its
line, which is the shape the edit wants: the move lands with the word
and holds through the rest of the phrase.
``min_gap`` keeps successive zooms apart; effects stacked close
together read as nervous editing rather than emphasis.
"""
segments = timeline.get("segments", [])
candidates: List[dict] = []
for i, segment in enumerate(segments):
content = [
w for w in segment.get("words", [])
if w["text"].strip(",.!?;:").lower() not in _FUNCTION_WORDS
]
if not content:
continue
best = max(content, key=lambda w: w["emphasis"])
candidates.append({
"start": best["start"],
# Hold through to the end of the sentence, not of the segment —
# releasing mid-thought is what makes a punch-in feel arbitrary.
"end": sentence_end(segments, i),
"word": best["text"],
"emphasis": best["emphasis"],
"line": segment["text"],
})
chosen: List[dict] = []
for candidate in sorted(candidates, key=lambda c: c["emphasis"], reverse=True):
if max_zooms is not None and len(chosen) >= max_zooms:
break
if any(abs(candidate["start"] - c["start"]) < min_gap for c in chosen):
continue
chosen.append(candidate)
return sorted(chosen, key=lambda c: c["start"])
def speaker_profiles(segments: Sequence[dict], duration: float) -> List[dict]:
"""Per-speaker statistics and sample lines, so a person can tell who is who.
A bare ``SPEAKER_00`` label is useless for deciding whose audio to cut.
What identifies a role is *how* someone participates: an interviewer or
a crew member asks short questions and holds little of the runtime,
while the subject speaks in long stretches. ``avg_segment`` and
``share`` capture exactly that contrast, and the sample lines confirm
it in the person's own words.
"""
by_speaker: dict = {}
for seg in segments:
sid = seg.get("speaker", seg.get("speaker_id", DEFAULT_SPEAKER))
length = max(0.0, float(seg.get("end", 0.0)) - float(seg.get("start", 0.0)))
entry = by_speaker.setdefault(sid, {"seconds": 0.0, "segments": [], "words": 0})
entry["seconds"] += length
entry["words"] += len(seg.get("words", []))
entry["segments"].append(seg)
profiles: List[dict] = []
for i, (sid, entry) in enumerate(
sorted(by_speaker.items(), key=lambda kv: kv[1]["seconds"], reverse=True)
):
count = len(entry["segments"])
# Longest lines identify a role far better than the first ones: a
# question and an answer look alike at the start of a recording.
longest = sorted(
entry["segments"],
key=lambda s: float(s.get("end", 0)) - float(s.get("start", 0)),
reverse=True,
)[:3]
profiles.append({
"id": sid,
"name": f"Speaker {i + 1}",
"speaking_seconds": round(entry["seconds"], 2),
"share": round(entry["seconds"] / duration, 3) if duration > 0 else 0.0,
"segment_count": count,
"avg_segment": round(entry["seconds"] / count, 2) if count else 0.0,
"word_count": entry["words"],
"samples": [(s.get("text") or "").strip()[:160] for s in longest],
})
return profiles
def select_peaks(
words: Sequence[dict], percentile: float, floor: float
) -> List[dict]:
"""The most emphatic words: the top ``percentile`` fraction, above ``floor``.
Selection is relative on purpose. The emphasis index is a weighted
average whose real range depends entirely on the material — a measured
interview peaks around 0.5 while an energetic ad reaches much higher —
so any fixed cutoff either floods one and selects nothing in the other.
Asking for "the top 2%" instead yields a usable handful either way.
``floor`` is only a sanity guard for genuinely flat audio, where even
the top of the distribution carries no emphasis worth cutting on.
"""
ranked = sorted(words, key=lambda w: w["emphasis"], reverse=True)
keep = max(1, round(len(ranked) * percentile)) if ranked else 0
return [w for w in ranked[:keep] if w["emphasis"] >= floor]
def _summary(words: Sequence[dict], segments: Sequence[dict], speakers: Sequence[dict],
duration: float, peak_percentile: float, emphasis_floor: float) -> dict:
"""The top layer: the shape of the piece in a handful of numbers."""
emphases = [w["emphasis"] for w in words]
peaks = select_peaks(words, peak_percentile, emphasis_floor)
return {
"duration": round(duration, 3),
"speaker_count": len(speakers),
"segment_count": len(segments),
"word_count": len(words),
"avg_emphasis": round(sum(emphases) / len(emphases), 3) if emphases else 0.0,
"peak_selection": f"top {peak_percentile:.0%} of words, minimum emphasis {emphasis_floor:.2f}",
"peak_count": len(peaks),
"peak_moments": [
{
"time": round(float(w.get("start", 0.0)), 3),
"text": w.get("word", ""),
"speaker": w.get("speaker_id", DEFAULT_SPEAKER),
"emphasis": round(w["emphasis"], 3),
}
for w in sorted(peaks, key=lambda w: w["emphasis"], reverse=True)[:20]
],
}
def build_voice_timeline(
media_path: str,
transcript: dict,
hf_token: Optional[str] = None,
num_speakers: str = "",
weights: EmphasisWeights = EmphasisWeights(),
peak_percentile: float = 0.02,
emphasis_floor: float = 0.25,
progress_cb: Optional[Callable[[float, str], None]] = None,
) -> dict:
"""Build the consolidated voice timeline for one media file.
Every analysis layer is optional and degrades independently: without
librosa the acoustic values are ``0.0``; without a diarization token
every word belongs to ``SPEAKER_00``. The document's shape never
changes, so downstream consumers (the rules engine, or a model reading
the JSON) can rely on it.
"""
def report(fraction: float, stage: str) -> None:
if progress_cb:
progress_cb(fraction, stage)
report(0.1, "Analisando tom e energia...")
pitch_track = extract_pitch(media_path)
energy_track = extract_energy(media_path)
report(0.5, "Calculando ênfase...")
words = enrich_words(transcript.get("words", []), pitch_track, energy_track, weights)
report(0.7, "Identificando participantes...")
tracks = diarize(media_path, hf_token, num_speakers) if hf_token else None
segments, words = assign_speakers(transcript.get("segments", []), words, tracks)
speakers = build_speakers(segments)
report(0.9, "Montando linha do tempo...")
duration = float(transcript.get("duration", 0.0))
segment_rows = _segment_rows(segments, words)
return {
"version": VOICE_TIMELINE_VERSION,
"source": Path(media_path).name,
"language": transcript.get("language", ""),
# What actually ran, not what was installed — a consumer must be able
# to tell "this speech is flat" from "the acoustics never loaded",
# since both leave the same zeros in the data.
"layers": {
"transcript": bool(transcript.get("words")),
"acoustics": pitch_track is not None or energy_track is not None,
"speakers": tracks is not None,
},
"scales": VALUE_SCALES,
"summary": _summary(
words, segments, speakers, duration, peak_percentile, emphasis_floor
),
"speakers": speaker_profiles(segment_rows, duration),
"segments": segment_rows,
}
def voice_timeline_path(media_path: str, output_dir: Optional[str] = None) -> Path:
"""Where the ``_voice_timeline.json`` for ``media_path`` lives.
Mirrors ``_transcript.json``: next to the media, or in the chosen
project folder when one is set.
"""
p = Path(media_path)
if output_dir:
directory = Path(output_dir).expanduser()
directory.mkdir(parents=True, exist_ok=True)
return directory / f"{p.stem}_voice_timeline.json"
return p.with_name(p.stem + "_voice_timeline.json")
def save_voice_timeline(timeline: dict, path: Path) -> None:
"""Write the timeline as UTF-8 JSON (accented transcripts stay readable)."""
with open(path, "w", encoding="utf-8") as f:
json.dump(timeline, f, ensure_ascii=False, indent=2)
def load_voice_timeline(path: Path) -> Optional[dict]:
"""Read a cached voice timeline, or ``None`` when absent/unreadable."""
try:
with open(path, encoding="utf-8") as f:
data = json.load(f)
except (OSError, json.JSONDecodeError, UnicodeDecodeError):
return None
return data if isinstance(data, dict) and "segments" in data else None
+355 -39
View File
@@ -35,9 +35,11 @@ import unicodedata
import uuid
import xml.etree.ElementTree as ET
from datetime import datetime
from fractions import Fraction
from pathlib import Path
from typing import Any, Dict, List, Optional, Tuple
from .collision import blocking, validate_titles
from .models import (
_FCPXML_STANDARD_TIMEBASES,
DynamicSubtitleConfig,
@@ -50,7 +52,12 @@ from .models import (
ValidationIssue,
ValidationIssueType,
)
from .text_layout import LayoutBox, compose_sentence, layout_sentence
from .text_layout import (
TEXT_TEMPLATE_FONT_SCALE,
LayoutBox,
compose_sentence,
layout_sentence,
)
from .transcribe import group_words_by_segment
# Maximum lengths for XML attribute values to prevent memory abuse
@@ -154,6 +161,23 @@ _ASSET_CLIP_CHILD_ORDER = [
_CHILD_ORDER_INDEX = {tag: i for i, tag in enumerate(_ASSET_CLIP_CHILD_ORDER)}
# How close to the end of a clip a zoom must finish for the return to be
# skipped. Within this margin the cut arrives before the eye registers the
# move back, so the return reads as a twitch rather than a resolution.
HOLD_AT_CUT_THRESHOLD = 1.0
# How close to the start of a clip a zoom must begin for the ramp-in to be
# skipped and the shot to simply open already zoomed. Tighter than the end
# margin on purpose: at the end the cut hides an unfinished return, but at
# the start a ramp is visible from frame one and reads as the shot settling.
START_AT_CUT_THRESHOLD = 0.5
def _fmt_scale(value: float) -> str:
"""Format a scale factor without trailing float noise (1.0 -> "1")."""
return f"{value:.6f}".rstrip("0").rstrip(".") or "0"
def _dtd_insert(parent: ET.Element, child: ET.Element) -> ET.Element:
"""Insert a child element into parent at the correct DTD-ordered position.
@@ -541,10 +565,43 @@ def _check_timebases(root: ET.Element) -> List[ValidationIssue]:
return issues
def _document_frame_duration(root: ET.Element) -> Optional[Fraction]:
"""The sequence's exact ``frameDuration`` as a fraction, if declared.
Read from the format the ``<sequence>`` references (falling back to the
first declared format), so the value is the document's own timebase
rather than an assumed rate.
"""
formats = {f.get('id'): f for f in root.findall('.//format') if f.get('id')}
sequence = root.find('.//sequence')
fmt = formats.get(sequence.get('format')) if sequence is not None else None
if fmt is None:
fmt = next(iter(formats.values()), None)
if fmt is None:
return None
raw = fmt.get('frameDuration', '')
if not (raw.endswith('s') and '/' in raw):
return None
numerator, denominator = raw[:-1].split('/', 1)
try:
value = Fraction(int(numerator), int(denominator))
except (ValueError, ZeroDivisionError):
return None
return value if value > 0 else None
def _check_frame_alignment(root: ET.Element, fps: float = 24.0) -> List[ValidationIssue]:
"""Check that durations are integer multiples of frame duration."""
"""Check that durations are integer multiples of the frame duration.
Uses the document's exact ``frameDuration`` fraction and rational
arithmetic. Comparing against an integer fps instead would flag every
NTSC project as broken: at 1001/24000s (23.976fps) a perfectly aligned
duration is not an integer number of "24fps" frames, so whole timelines
would be reported misaligned when nothing is wrong.
"""
issues = []
fps_int = int(fps)
frame_duration = _document_frame_duration(root)
label = f"{1 / float(frame_duration):.3f}".rstrip('0').rstrip('.') if frame_duration else str(fps)
for elem in root.iter():
dur_str = elem.get('duration')
if not dur_str or not dur_str.endswith('s'):
@@ -553,14 +610,19 @@ def _check_frame_alignment(root: ET.Element, fps: float = 24.0) -> List[Validati
continue
try:
tv = TimeValue.from_timecode(dur_str)
frames = tv.to_seconds() * fps_int
if abs(frames - round(frames)) > 0.01:
if frame_duration is not None:
frames = Fraction(tv.numerator, tv.denominator) / frame_duration
aligned = frames.denominator == 1
else:
approx = tv.to_seconds() * fps
aligned = abs(approx - round(approx)) <= 0.01
if not aligned:
issues.append(ValidationIssue(
issue_type=ValidationIssueType.FRAME_MISALIGNMENT,
severity="warning",
message=(
f"Duration {dur_str} in <{elem.tag}> "
f"'{elem.get('name', '')}' is not frame-aligned at {fps_int}fps."
f"'{elem.get('name', '')}' is not frame-aligned at {label}fps."
),
clip_name=elem.get('name'),
))
@@ -1030,13 +1092,21 @@ class FCPXMLModifier:
elem.set(attr, val)
return elem
def _require_clip(self, clip_id: str) -> ET.Element:
def _require_clip(self, clip_id: 'str | ET.Element') -> ET.Element:
"""Look up a clip by ID/name, raising if not found.
Centralises the get-or-raise pattern used by every clip-mutating
method so the error message stays consistent and future
enhancements (fuzzy matching, suggestions) only need one site.
An Element is returned as-is. That matters after ``split_clip`` or
``cut_clip_ranges``: the resulting pieces all carry the *same* name,
so a name lookup would always resolve to the first one and silently
put the edit on the wrong piece. Callers holding the exact element
pass it directly.
"""
if isinstance(clip_id, ET.Element):
return clip_id
clip = self.clips.get(clip_id)
if clip is None:
raise ValueError(f"Clip not found: {clip_id}")
@@ -1421,7 +1491,7 @@ class FCPXMLModifier:
def add_marker(
self,
clip_id: str,
clip_id: 'str | ET.Element',
timecode: str,
name: str,
marker_type: "MarkerType | str" = MarkerType.STANDARD,
@@ -1937,35 +2007,50 @@ class FCPXMLModifier:
def add_zoom(
self,
clip_id: str,
clip_id: 'str | ET.Element',
start: float,
end: float,
scale: float = 1.3,
ease: float = 0.3,
ease: float = 0.25,
position: str = "0 0",
ease_out: Optional[float] = None,
hold_at_end: Optional[bool] = None,
start_at_peak: Optional[bool] = None,
) -> ET.Element:
"""Add a smooth ease-in/ease-out punch-in zoom to a clip.
"""Add a punch-in zoom to a clip, snapping back to its framing at the end.
Animates ``<adjust-transform>``'s ``scale`` param (per the FCPXML
DTD: ``<param>`` + ``<keyframeAnimation>`` of ``<keyframe>``
elements, ``interp="ease"``) from 100% up to *scale* and back down
to 100%, entirely within ``[start, end]`` — clip-relative seconds
(seconds from the clip's own head, same convention as
``cut_clip_ranges``). The ease portions each last *ease* seconds;
the zoom holds at *scale* in between.
Animates ``<adjust-transform>``'s ``scale`` param (``<param>`` +
``<keyframeAnimation>`` of ``<keyframe>``) from the clip's current
scale up to *scale* times it, holds, then returns — all within
``[start, end]`` — clip-relative seconds (same convention as
``cut_clip_ranges``).
The two ends are deliberately asymmetric. *ease* ramps the zoom
**in** over half a second by default, fast enough to land with the
emphasised word. The way **out** is instant — a single frame — so
the moment the impact phrase ends the shot is simply back to its
normal framing and the video resumes its flow, with no drift
drawing attention to itself. Pass *ease_out* to ramp the return
gradually instead.
*hold_at_end* keeps the peak instead of returning, and
*start_at_peak* opens already zoomed with no ramp. Left as ``None``
both decide on their own from how close the window sits to the
clip's edges: a cut is itself the transition, so ramping away from
one — or back toward one — is motion the viewer reads as a wobble
rather than as emphasis.
"""
if end <= start:
raise ValueError(f"end ({end}) must be greater than start ({start})")
if ease <= 0:
raise ValueError(f"ease must be positive, got {ease}")
if ease * 2 > (end - start):
raise ValueError(
f"ease ({ease}s x2 = {ease * 2}s) doesn't fit in the zoom "
f"window ({end - start}s) — shorten ease or widen start/end"
)
if scale <= 0:
raise ValueError(f"scale must be positive, got {scale}")
frame = float(self.frame_duration_fraction())
ramp_out = frame if ease_out is None else ease_out
if ramp_out <= 0:
raise ValueError(f"ease_out must be positive, got {ease_out}")
clip = self._require_clip(clip_id)
clip_duration = self._parse_time(clip.get('duration', '0s')).to_seconds()
if start < 0 or end > clip_duration:
@@ -1974,26 +2059,146 @@ class FCPXMLModifier:
f"duration (0 to {clip_duration:.3f}s)"
)
# Replace rather than stack a prior zoom on the same clip.
# Replace a prior zoom, but never the clip's framing. A clip can
# already carry an <adjust-transform> holding the editor's own
# reframe — rotation for footage shot sideways, position, a scale
# that makes the shot work at all. Dropping it outright (the old
# behaviour) silently destroyed that framing; on real footage the
# zoomed section came back rotated. So: keep the static attributes,
# and animate *relative to* the existing scale.
base_x, base_y = 1.0, 1.0
carried: dict = {}
old_keyframes: list = []
for stale in clip.findall('adjust-transform'):
carried = {k: v for k, v in stale.attrib.items() if k != 'scale'}
parts = (stale.get('scale') or '').split()
if len(parts) == 2:
try:
base_x, base_y = float(parts[0]), float(parts[1])
except ValueError:
base_x, base_y = 1.0, 1.0
else:
# No static attribute — a PRIOR zoom on this same clip left
# an animated <param name="scale"> instead, and the true
# resting framing lives in its keyframes, not in 1.0.
# Reading it as 1.0 here doesn't just miss the framing: it
# replaces the earlier zoom's whole animation with a wrong
# one, since this loop unconditionally removes `stale`
# right after. The rest value is recoverable without
# knowing which keyframe it is: MIN_ZOOM_SCALE == 1.0 means
# every keyframed value is >= the rest scale, so the
# smallest one keyframed is the rest value, peak or not.
for old_param in stale.findall("param[@name='scale']"):
xs, ys = [], []
for kf in old_param.findall('.//keyframe'):
kv = (kf.get('value') or '').split()
if len(kv) == 2:
try:
xs.append(float(kv[0]))
ys.append(float(kv[1]))
except ValueError:
pass
# Kept for merging: a second zoom on the same clip
# (two emphatic beats a cut didn't separate) should
# stack alongside the first, not erase it — the
# earlier peak is still a real editorial decision.
old_keyframes.append((kf.get('time', '0s'), kf.get('value', '')))
if xs and ys:
base_x, base_y = min(xs), min(ys)
clip.remove(stale)
transform = ET.Element('adjust-transform')
for key, value in carried.items():
transform.set(key, value)
scale_param = ET.SubElement(transform, 'param')
scale_param.set('name', 'scale')
anim = ET.SubElement(scale_param, 'keyframeAnimation')
scale_value = f"{scale} {scale}"
for seconds, value in (
(start, "1 1"),
(start + ease, scale_value),
(end - ease, scale_value),
(end, "1 1"),
):
# Keyframe times live in the clip's SOURCE timebase — the same origin
# as its own ``start`` — not in clip-relative seconds. A clip whose
# media starts at, say, 3109.9s of timecode looks for the animation
# there; keyframes written at 0-5s land outside the clip entirely and
# Final Cut imports the zoom as nothing at all, silently. Matches what
# add_text_title already does, and only shows up on footage whose
# start isn't 0s — every synthetic fixture starts at 0s and hides it.
media_origin = self._parse_time(clip.get('start', '0s'))
rest_value = f"{_fmt_scale(base_x)} {_fmt_scale(base_y)}"
scale_value = f"{_fmt_scale(base_x * scale)} {_fmt_scale(base_y * scale)}"
# A return that lands right before a cut is wasted motion: the next
# clip begins on its own framing anyway, so all the viewer sees is a
# twitch on the way out. When the zoom runs to the end of the clip,
# hold the peak and let the cut do the resetting.
holds_to_cut = (
hold_at_end
if hold_at_end is not None
else (clip_duration - end) <= HOLD_AT_CUT_THRESHOLD
)
opens_at_peak = (
start_at_peak
if start_at_peak is not None
else start <= START_AT_CUT_THRESHOLD
)
# Only the ramps actually written have to fit in the window: a zoom
# that opens at the peak spends no time ramping in, and one held to
# the cut spends none ramping out.
needed = (0.0 if opens_at_peak else ease) + (0.0 if holds_to_cut else ramp_out)
if needed > (end - start):
raise ValueError(
f"the ramps ({needed}s) don't fit in the zoom window "
f"({end - start}s) — shorten them or widen start/end"
)
if opens_at_peak:
# The cut already delivered the change of framing; ramping up
# from it just looks like the shot settling.
keyframes = [(start, scale_value)]
else:
keyframes = [(start, rest_value), (start + ease, scale_value)]
if holds_to_cut:
keyframes.append((end, scale_value))
else:
# Hold the peak right up to the end, then drop back on the very
# next frame — the snap-back the edit wants, not a slow drift.
keyframes.append((end - ramp_out, scale_value))
keyframes.append((end, rest_value))
new_entries = [
((media_origin + self.snap_seconds_to_frame(seconds)), value)
for seconds, value in keyframes
]
new_start_time = new_entries[0][0]
new_end_time = new_entries[-1][0]
# Two calls on the same clip mean two different things depending on
# whether their windows overlap. Overlapping = redoing the *same*
# zoom with new numbers — the old keyframes are stale and all of
# them go. Disjoint = a second, separate beat that a cut didn't
# separate onto its own clip — that one stacks alongside the first
# instead of erasing it, since both are real editorial decisions.
old_times = [self._parse_time(t) for t, _ in old_keyframes]
old_span_overlaps_new = bool(old_times) and not (
max(old_times) < new_start_time or min(old_times) > new_end_time
)
if old_span_overlaps_new:
surviving_old: list = []
else:
surviving_old = [(self._parse_time(t), v) for t, v in old_keyframes]
all_entries = sorted(surviving_old + new_entries, key=lambda e: e[0])
for time_value, value in all_entries:
kf = ET.SubElement(anim, 'keyframe')
kf.set('time', self.snap_seconds_to_frame(seconds).to_fcpxml())
kf.set('time', time_value.to_fcpxml())
kf.set('value', value)
kf.set('interp', 'ease')
# Only 'time' and 'value' — no 'interp', no 'curve'. The DTD allows
# both, but Final Cut rejected 'interp' on this vector param
# ("does not support the interpolation attribute") and discarded
# the whole <param>. A hand-made zoom exported from FCP itself
# writes bare keyframes and relies on the DTD default
# (curve="smooth"), so we match that export exactly rather than
# guess which attributes survive its importer.
if position != "0 0":
pos_param = ET.SubElement(transform, 'param')
@@ -2707,7 +2912,18 @@ class FCPXMLModifier:
_TEXT_POSITION_KEY = '9999/10003/13260/3296672360/1/100/101'
# Layout params the "Text" template ships with. These keys are the
# template's own defaults and never vary between instances.
#
# "Build Out" is the one deliberate override: with "Apply Speed" set to
# "2 (Per Object)" below, the template's whole built-in animation (build
# in + build out) is always compressed to exactly fill the title's own
# on-screen duration — so on a short word-length clip, build out was
# eating time that build in needed to finish revealing the text before
# the cut. Disabling build out hands that entire compressed window to
# build in alone, which is what "sempre acelerado" turned out to mean:
# no separate speed knob needed. Value captured from a real FCP export
# with "Build Out" unchecked in the Inspector (see chat, 2026-08-18).
_TEXT_TITLE_PARAMS = (
('Build Out', '9999/10000/2/102', '0'),
('Layout Method', '9999/10003/13260/3296672360/2/314', '1 (Paragraph)'),
('Left Margin', '9999/10003/13260/3296672360/2/323', '-1210'),
('Right Margin', '9999/10003/13260/3296672360/2/324', '1210'),
@@ -2795,6 +3011,7 @@ class FCPXMLModifier:
bold: bool = True,
face: Optional[str] = None,
kerning: Optional[float] = None,
font_scale: float = TEXT_TEMPLATE_FONT_SCALE,
) -> ET.Element:
"""Build a standalone ``<title>`` clip from the "Text" (Basic Text) template.
@@ -2849,13 +3066,31 @@ class FCPXMLModifier:
style_def.set('id', ts_id)
text_style = ET.SubElement(style_def, 'text-style')
text_style.set('font', font)
text_style.set('fontSize', str(font_size))
# Text.moti sizes type in frame pixels but positions in canvas points.
# See TEXT_TEMPLATE_FONT_SCALE: layout measures in points, so only the
# emitted size (and its kerning, to keep the same letter spacing) is
# converted here.
scale = float(font_scale) or 1.0
text_style.set('fontSize', f"{float(font_size) * scale:g}")
text_style.set('fontColor', font_color)
text_style.set('bold', '1' if bold else '0')
if face:
# FCP represents bold weight as the bold attribute — never as a
# fontFace. Writing ``bold="0" fontFace="Bold"`` (the previous
# behaviour) is contradictory and FCP refuses to render the text.
# Italic, by contrast, IS a face: FCP writes both ``fontFace`` and
# ``italic="1"``. See Engine/docs/05_EXPERIENCIAS.md, entry 2026-08-19.
face_lower = (face or '').strip().lower()
if face_lower == 'bold':
text_style.set('bold', '1')
elif 'italic' in face_lower:
text_style.set('fontFace', face)
text_style.set('italic', '1')
else:
if bold:
text_style.set('bold', '1')
if face:
text_style.set('fontFace', face)
if kerning:
text_style.set('kerning', f"{float(kerning):g}")
text_style.set('kerning', f"{float(kerning) * scale:g}")
text_style.set('alignment', 'center')
text_style.set('lineSpacing', '-19')
@@ -3005,7 +3240,9 @@ class FCPXMLModifier:
def lay_out(pending: List[Dict]):
"""Place what fits; return (units, still-unplaced words)."""
if phrase_mode:
composition = compose_sentence(pending, config.style, box)
composition = compose_sentence(
pending, config.style, box, line_gap=config.line_gap,
)
return composition.blocks, composition.overflow
layout = layout_sentence(pending, config.style, box)
return layout.placed, layout.overflow
@@ -3083,19 +3320,98 @@ class FCPXMLModifier:
duration,
lane=lane,
name=f"caption_{uuid.uuid4().hex[:8]}",
position=unit.position_param(),
position=unit.position_param(config.text_scale),
font=unit.font or config.style.font,
font_size=int(round(unit.font_size)),
font_color=unit.color or config.style.active_color,
bold=config.style.bold,
face=unit.face,
kerning=unit.kerning,
font_scale=config.text_scale,
)
_dtd_insert(parent, title)
created.append(title)
if getattr(config, 'validate', False):
report = self.validate_subtitle_layout()
if blocking(report["severity"]):
raise ValueError(
"Subtitle layout validation failed: "
+ str(report["summary"])
)
return created
def validate_subtitle_layout(
self,
*,
safe_margin_x: float = 0.05,
safe_margin_y: float = 0.05,
min_font_size: Optional[float] = None,
min_distance: Optional[float] = None,
max_distance: Optional[float] = None,
) -> dict:
"""Re-measure every ``<title>`` in the document and report collisions.
Reconstructs each title's on-screen box from the values the writer
emitted (``fontSize``/``kerning``/``Position`` are already in template
space), then checks for temporal+spatial collisions, frame/safe-area
containment, and font fallbacks. This is the spec-16 validation pass the
layout engine does not do on its own — it only guarantees non-overlap
*by construction* while composing, and cannot see a hand-edited title.
Returns the ``collision.validate_titles`` report: ``severity`` (worst
bucket), ``issues`` (spec-16 occurrences) and ``summary`` (counts).
"""
titles = []
for elem in self.root.iter('title'):
text_el = elem.find('text/text-style')
text = (text_el.text or '').strip() if text_el is not None else ''
style = elem.find('text-style-def/text-style')
font = style.get('font') if style is not None else None
face = style.get('fontFace') if style is not None else None
font_size = (
float(style.get('fontSize', '0')) if style is not None else 0.0
)
kerning = (
float(style.get('kerning', '0') or 0)
if style is not None else 0.0
)
x = y = 0.0
for param in elem.findall('param'):
if param.get('name') == 'Position' and param.get('value'):
parts = param.get('value').split()
if len(parts) >= 2:
x, y = float(parts[0]), float(parts[1])
start = self._parse_time(elem.get('offset', '0s')).to_seconds()
duration = self._parse_time(elem.get('duration', '0s')).to_seconds()
titles.append({
'text': text,
'font': font,
'face': face,
'font_size': font_size,
'kerning': kerning,
'x': x,
'y': y,
'start': start,
'end': start + duration,
'group': start + duration,
})
return validate_titles(
titles,
self.frame_width(),
self.frame_height(),
safe_margin_x=safe_margin_x,
safe_margin_y=safe_margin_y,
min_font_size=min_font_size,
min_distance=min_distance,
max_distance=max_distance,
)
# ========================================================================
# AUDIO CLIP OPERATIONS (v0.6.0)
# ========================================================================
Executable → Regular
+318 -3870
View File
File diff suppressed because it is too large Load Diff
+8
View File
@@ -0,0 +1,8 @@
"""Tool handlers and schemas for the FCPXML MCP server, split by category.
server.py is the composition root: it imports each module's TOOLS/HANDLERS
and concatenates them for list_tools()/TOOL_HANDLERS. Each module here owns
one category (see Engine/docs/03_SERVER_TOOLS.md) — its Tool() schemas and
handle_<name> functions live together, so a tool's contract and its
implementation are never in different files.
"""
+829
View File
@@ -0,0 +1,829 @@
"""Shared internal helpers used by tool handlers across categories.
Extracted from server.py — validation, formatting, and small parsing utilities
that more than one server_tools/*.py module needs.
"""
from __future__ import annotations
import json
import os
import re
from pathlib import Path
from typing import Any, Sequence
from mcp.types import TextContent
from fcpxml.media_intel import media_src_to_path
from fcpxml.models import (
DuplicateGroup,
FlashFrame,
FlashFrameSeverity,
GapInfo,
Timecode,
TimeValue,
)
from fcpxml.parser import FCPXMLParser
from fcpxml.rough_cut import RoughCutGenerator
from fcpxml.transcribe import invert_ranges, merge_ranges, transcribe
from fcpxml.writer import FCPXMLModifier
PROJECTS_DIR = os.environ.get("FCP_PROJECTS_DIR", os.path.expanduser("~/Movies"))
_SANDBOX_ENABLED = "FCP_PROJECTS_DIR" in os.environ
MAX_FILE_SIZE = 100 * 1024 * 1024
MAX_MEDIA_FILE_SIZE = 32 * 1024 * 1024 * 1024
_MAX_JSON_DEPTH = 50
def _check_json_depth(obj: object, _depth: int = 0) -> None:
"""Reject JSON structures nested beyond _MAX_JSON_DEPTH.
Prevents denial-of-service via deeply nested objects that exhaust the
call stack or memory during downstream processing. Called after
json.load() since Python's json module has no built-in depth limit.
"""
if _depth > _MAX_JSON_DEPTH:
raise ValueError(
f"JSON nesting depth exceeds {_MAX_JSON_DEPTH} — "
"file may be malformed or adversarial"
)
if isinstance(obj, dict):
for v in obj.values():
_check_json_depth(v, _depth + 1)
elif isinstance(obj, list):
for item in obj:
_check_json_depth(item, _depth + 1)
def _validate_filepath(
filepath: str,
allowed_extensions: tuple[str, ...] | None = None,
max_size: int = MAX_FILE_SIZE,
) -> str:
"""Validate a user-provided file path against traversal and size attacks.
Resolves symlinks, blocks null bytes, enforces extension whitelist, and
checks file size before any parsing takes place.
``max_size`` defaults to the document limit; callers handling source
media pass ``MAX_MEDIA_FILE_SIZE``, since media is streamed rather than
parsed into memory (see the constant for why).
Raises:
ValueError: For invalid paths (null bytes, bad extensions, oversized).
FileNotFoundError: When the resolved path does not exist.
"""
if '\x00' in filepath:
raise ValueError("Invalid file path: null byte detected")
resolved = Path(filepath).resolve()
if not resolved.exists():
raise FileNotFoundError(f"File not found: {filepath}")
# .fcpxmld bundles are directories (a package wrapping Info.fcpxml plus
# sidecar data files for object tracking / Cinematic mode). The size
# check applies to the inner Info.fcpxml, which is what gets parsed.
if resolved.is_dir():
if resolved.suffix.lower() != '.fcpxmld':
raise ValueError(f"Not a regular file: {filepath}")
inner = resolved / 'Info.fcpxml'
if not inner.is_file():
raise ValueError(f"Invalid bundle (no Info.fcpxml): {filepath}")
size_target = inner
elif not resolved.is_file():
raise ValueError(f"Not a regular file: {filepath}")
else:
size_target = resolved
if allowed_extensions and resolved.suffix.lower() not in allowed_extensions:
raise ValueError(
f"Invalid file type '{resolved.suffix}'. "
f"Allowed: {', '.join(allowed_extensions)}"
)
if size_target.stat().st_size > max_size:
size_mb = size_target.stat().st_size / (1024 * 1024)
raise ValueError(f"File too large ({size_mb:.1f} MB). Maximum: {max_size // (1024 * 1024)} MB")
return str(resolved)
def _validate_output_path(output_path: str, *, anchor_dir: str | None = None) -> str:
"""Validate an output path with optional sandbox enforcement.
Resolves traversal, blocks null bytes, ensures parent exists, and — when
*anchor_dir* is provided — verifies the resolved output lives under that
directory. This prevents LLM-generated tool calls from writing to
arbitrary filesystem locations (e.g. ``/etc/cron.d/backdoor``).
Args:
output_path: The raw output path to validate.
anchor_dir: If set, the resolved output must be a child of this
directory. Typically the parent directory of the input file so
outputs stay co-located with their sources.
Raises:
ValueError: For null bytes, missing parent, or sandbox escape.
"""
if '\x00' in output_path:
raise ValueError("Invalid output path: null byte detected")
resolved = Path(output_path).resolve()
if not resolved.parent.exists():
raise ValueError(f"Output directory does not exist: {resolved.parent}")
if anchor_dir is not None:
anchor = Path(anchor_dir).resolve()
try:
resolved.relative_to(anchor)
except ValueError:
raise ValueError(
f"Output path escapes allowed directory: "
f"{resolved} is not under {anchor}"
)
return str(resolved)
def _validate_directory(directory: str, *, allowed_root: str | None = None) -> str:
"""Validate a user-provided directory path against traversal and injection.
Resolves symlinks, blocks null bytes, and verifies the path is a real
directory. When *allowed_root* is given, the resolved path must be a
descendant of (or equal to) that root — preventing filesystem enumeration
beyond the project workspace.
Raises:
ValueError: For invalid paths (null bytes, not a directory, sandbox escape).
"""
if '\x00' in directory:
raise ValueError("Invalid directory path: null byte detected")
resolved = Path(directory).resolve()
if not resolved.is_dir():
raise ValueError(f"Not a valid directory: {directory}")
if allowed_root is not None:
root = Path(allowed_root).resolve()
try:
resolved.relative_to(root)
except ValueError:
raise ValueError(
f"Directory escapes allowed root: "
f"{resolved} is not under {root}"
)
return str(resolved)
def find_fcpxml_files(directory: str) -> list[str]:
"""Find all FCPXML files in a directory."""
path = Path(directory)
files = list(str(f) for f in path.rglob("*.fcpxml"))
files.extend(str(f) for f in path.rglob("*.fcpxmld"))
return sorted(files)
def format_timecode(tc) -> str:
"""Format a Timecode object to SMPTE string."""
return tc.to_smpte() if tc else "00:00:00:00"
def format_duration(seconds: float) -> str:
"""Format seconds into human-readable duration."""
if seconds < 1:
return f"{seconds*1000:.0f}ms"
elif seconds < 60:
return f"{seconds:.2f}s"
return f"{int(seconds // 60)}m {seconds % 60:.1f}s"
def _format_clip_table(clips: list, header: str) -> str:
"""Render a list of clips as a markdown table with timecodes and durations.
Shared by handlers that filter clips by duration threshold
(find_short_cuts, find_long_clips).
"""
result = f"{header}\n\n| Name | TC | Duration |\n|------|----|---------|\n"
result += "\n".join(
f"| {c.name} | {format_timecode(c.start)} | {format_duration(c.duration_seconds)} |"
for c in clips
)
return result
def _markdown_table(headers: list[str], rows: list[list[str]]) -> str:
"""Build a markdown table from headers and rows.
Returns header row, separator row, and data rows as a single string.
Callers avoid repeating the ``| H1 | H2 |\\n|---|---|`` boilerplate
that appears in 15+ handlers.
"""
header_line = "| " + " | ".join(headers) + " |"
sep_line = "|" + "|".join("------" for _ in headers) + "|"
data_lines = "\n".join(
"| " + " | ".join(str(c) for c in row) + " |" for row in rows
)
return f"{header_line}\n{sep_line}\n{data_lines}"
def _format_batch_result(
title: str,
summary: dict[str, str],
headers: list[str],
rows: list[list[str]],
output_path: str,
) -> str:
"""Build a standard batch-operation result with summary, table, and save footer.
Used by batch fix handlers (flash frames, rapid trim, fill gaps) that all
share the same markdown structure: ``# Title → ## Summary → ## Details table
→ Saved to`` footer.
"""
summary_lines = "\n".join(f"- **{k}**: {v}" for k, v in summary.items())
table = _markdown_table(headers, rows)
return (
f"# {title}\n\n"
f"## Summary\n{summary_lines}\n\n"
f"## Details\n{table}\n\n"
f"Saved to: `{output_path}`"
)
def _fmt_suggestions(suggestions: list[str]) -> str:
"""Format pacing suggestions as markdown list (Python 3.10 compatible)."""
if not suggestions:
return "- Pacing looks good!"
nl = "\n"
return nl.join(f"- {s}" for s in suggestions)
def generate_output_path(input_path: str, suffix: str = "_modified") -> str:
"""Generate output path from input path.
The suffix is sanitized to prevent path-component injection — only
alphanumeric, hyphen, underscore, and dot characters survive.
"""
# Strip anything that could inject path separators or traversal sequences
clean_suffix = re.sub(r'[^a-zA-Z0-9._-]', '', suffix)
if not clean_suffix:
clean_suffix = "_modified"
p = Path(input_path)
return str(p.parent / f"{p.stem}{clean_suffix}{p.suffix}")
def _parse_project(filepath: str):
"""Parse an FCPXML file and return the project with its primary timeline."""
filepath = _validate_filepath(filepath, ('.fcpxml', '.fcpxmld'))
project = FCPXMLParser().parse_file(filepath)
if not project.timelines:
return None, None
return project, project.primary_timeline
def _text_result(text: str) -> list[TextContent]:
"""Wrap a string in the MCP TextContent list that every tool handler returns."""
return [TextContent(type="text", text=text)]
def _no_timeline():
"""Standard response when no timelines are found."""
return _text_result("No timelines found")
def _require_timeline(filepath: str):
"""Parse FCPXML and return (project, timeline), raising if no timeline exists.
Centralises the repeated _parse_project + _no_timeline guard that
appears in every read-only timeline handler. Returns a tuple so
callers can destructure directly::
project, tl = _require_timeline(arguments["filepath"])
"""
project, tl = _parse_project(filepath)
if not tl:
raise _NoTimelineError()
return project, tl
class _NoTimelineError(Exception):
"""Sentinel raised by _require_timeline when no timelines exist."""
def _resolve_io_paths(
arguments: dict,
suffix: str = "_modified",
) -> tuple[str, str]:
"""Validate input filepath and resolve the output path.
Shared foundation for every handler that reads an FCPXML and writes
a derived file. Validates the input, falls back to a suffixed
output name when ``output_path`` is not supplied, and sandbox-checks
the result.
Args:
arguments: Tool arguments dict (must contain ``filepath``; may
contain ``output_path``).
suffix: Default output filename suffix when ``output_path`` is
not provided (e.g. ``"_modified"``, ``"_beats"``).
Returns:
``(filepath, output_path)`` tuple with both paths validated.
"""
filepath = _validate_filepath(arguments["filepath"], ('.fcpxml', '.fcpxmld'))
# Anchor write operations to the input file's directory so LLM-generated
# tool calls cannot write to arbitrary filesystem locations (e.g.
# /etc/cron.d/backdoor). When the explicit sandbox is off, the anchor
# still prevents writes outside the source directory tree.
# `output_dir` is where the caller wants the file written, not merely a
# sandbox boundary: the app's "Pasta do projeto" promises that everything
# generated lands there. Deriving the name from the input but keeping the
# input's directory made every cross-directory call fail its own anchor
# check ("output path escapes allowed directory"), so the setting silently
# only worked when it pointed at the directory the file was already going
# to. An explicit `output_path` still wins, and still has to sit inside
# the anchor.
output_dir = arguments.get("output_dir")
if output_dir:
anchor = _validate_directory(str(output_dir))
default_output = str(Path(anchor) / Path(generate_output_path(filepath, suffix)).name)
else:
anchor = str(Path(filepath).resolve().parent)
default_output = generate_output_path(filepath, suffix)
output_path = _validate_output_path(
arguments.get("output_path") or default_output,
anchor_dir=anchor,
)
return filepath, output_path
def _setup_modifier(
arguments: dict,
suffix: str = "_modified",
) -> tuple[str, str, "FCPXMLModifier"]:
"""Common setup for write handlers: validate paths and create modifier.
Consolidates the repeated validate-filepath → resolve-output-path →
create-modifier boilerplate shared by 18+ write handlers.
Args:
arguments: Tool arguments dict (must contain ``filepath``; may
contain ``output_path``).
suffix: Default output filename suffix when ``output_path`` is
not provided (e.g. ``"_modified"``, ``"_flash_fixed"``).
Returns:
``(filepath, output_path, modifier)`` tuple ready for the
handler's domain-specific operation.
"""
filepath, output_path = _resolve_io_paths(arguments, suffix)
modifier = FCPXMLModifier(filepath)
return filepath, output_path, modifier
def _setup_generator(
arguments: dict,
suffix: str = "_roughcut",
) -> tuple[str, str, "RoughCutGenerator"]:
"""Common setup for generation handlers: validate paths and create generator.
Args:
arguments: Tool arguments dict (must contain ``filepath`` and
``output_path``).
suffix: Default output filename suffix.
Returns:
``(filepath, output_path, generator)`` tuple.
"""
filepath, output_path = _resolve_io_paths(arguments, suffix)
generator = RoughCutGenerator(filepath)
return filepath, output_path, generator
def _parse_timestamp_parts(
parts: list[str], *, frame_rate: float = 24.0
) -> float | None:
"""Convert colon-separated timestamp parts to total seconds.
Handles 2-part (M:SS), 3-part (H:MM:SS / HH:MM:SS.ms), and
4-part (HH:MM:SS:FF SMPTE) formats. Returns ``None`` when the
part count is unrecognised so callers can skip.
Args:
parts: Colon-split timestamp components.
frame_rate: FPS used to convert the frame component of SMPTE
timecodes into fractional seconds (default 24.0).
"""
if len(parts) == 2:
return int(parts[0]) * 60 + float(parts[1])
elif len(parts) == 3:
return int(parts[0]) * 3600 + int(parts[1]) * 60 + float(parts[2])
elif len(parts) == 4:
# SMPTE: HH:MM:SS:FF — convert frames to fractional seconds
base = int(parts[0]) * 3600 + int(parts[1]) * 60 + float(parts[2])
frames = int(parts[3])
return base + (frames / frame_rate) if frame_rate > 0 else base
return None
def _raw_markers_to_batch(
raw_markers: list[dict],
marker_type: str = "chapter",
max_label: int | None = None,
) -> list[dict]:
"""Convert raw {seconds, text} marker dicts to batch_add_markers format.
Shared by import_srt_markers and import_transcript_markers.
"""
batch = []
for m in raw_markers:
label = m["text"]
if max_label and len(label) > max_label:
label = label[:max_label]
batch.append({
"timecode": f"{m['seconds']}s",
"name": label,
"marker_type": marker_type.upper(),
})
return batch
def _extract_subtitle_blocks(text: str, *, strip_vtt_tags: bool = False) -> list[dict]:
"""Extract timestamp/text pairs from subtitle cue blocks (SRT or VTT).
Both SRT and VTT use the same ``start --> end`` cue syntax with
text lines underneath; only header stripping and tag cleaning differ.
"""
markers = []
blocks = re.split(r'\n\s*\n', text.strip())
for block in blocks:
lines = block.strip().split('\n')
if len(lines) < 2:
continue
ts_line = None
text_lines = []
for line in lines:
if '-->' in line:
ts_line = line
elif ts_line is not None:
if strip_vtt_tags:
line = re.sub(r'<[^>]+>', '', line)
cleaned = line.strip()
if cleaned:
text_lines.append(cleaned)
if not ts_line or not text_lines:
continue
start_str = ts_line.split('-->')[0].strip().replace(',', '.')
seconds = _parse_timestamp_parts(start_str.split(':'))
if seconds is not None:
markers.append({'seconds': seconds, 'text': ' '.join(text_lines)})
return markers
def parse_srt(text: str) -> list[dict]:
"""Parse SRT subtitle format into timestamp/text pairs."""
return _extract_subtitle_blocks(text)
def parse_vtt(text: str) -> list[dict]:
"""Parse WebVTT subtitle format into timestamp/text pairs."""
text = re.sub(r'^WEBVTT.*?\n', '', text, flags=re.MULTILINE)
text = re.sub(r'NOTE\n.*?\n\n', '', text, flags=re.DOTALL)
return _extract_subtitle_blocks(text, strip_vtt_tags=True)
def parse_transcript_timestamps(text: str) -> list[dict]:
"""Parse timestamped text (YouTube description format) into markers.
Supports formats like:
0:00 Introduction
00:01:30 Main Topic
1:05:30 Conclusion
00:00:00:00 SMPTE timecode
"""
markers = []
for line in text.strip().split('\n'):
line = line.strip()
if not line:
continue
match = re.match(r'^(\d{1,2}:\d{2}(?::\d{2}){0,2})\s+(.+)$', line)
if match:
seconds = _parse_timestamp_parts(match.group(1).split(':'))
if seconds is not None:
markers.append({'seconds': seconds, 'text': match.group(2).strip()})
return markers
def _detect_flash_frames(
tl: Any, *, critical_threshold: int = 2, warning_threshold: int = 6,
) -> list:
"""Find clips shorter than *warning_threshold* frames.
Returns a list of ``FlashFrame`` objects sorted by severity. Shared by
``handle_detect_flash_frames`` and ``handle_validate_timeline`` so the
detection logic lives in exactly one place.
"""
fps = tl.frame_rate
flash_frames: list[FlashFrame] = []
for clip in tl.clips:
duration_frames = int(clip.duration_seconds * fps)
if duration_frames < warning_threshold:
severity = (
FlashFrameSeverity.CRITICAL
if duration_frames < critical_threshold
else FlashFrameSeverity.WARNING
)
flash_frames.append(FlashFrame(
clip_name=clip.name, clip_id=clip.name,
start=clip.start, duration_frames=duration_frames,
duration_seconds=clip.duration_seconds, severity=severity,
))
return flash_frames
def _detect_gaps(tl: Any, *, min_gap_frames: int = 1) -> list:
"""Find inter-clip gaps of at least *min_gap_frames* length.
Returns a list of ``GapInfo`` objects. Shared by ``handle_detect_gaps``
and ``handle_validate_timeline``.
"""
fps = tl.frame_rate
min_gap_seconds = min_gap_frames / fps
gaps: list[GapInfo] = []
sorted_clips = sorted(tl.clips, key=lambda c: c.start.seconds)
for i in range(len(sorted_clips) - 1):
current_end = sorted_clips[i].end.seconds
next_start = sorted_clips[i + 1].start.seconds
gap_duration = next_start - current_end
if gap_duration >= min_gap_seconds:
gaps.append(GapInfo(
start=Timecode(frames=int(current_end * fps), frame_rate=fps),
duration_frames=int(gap_duration * fps),
duration_seconds=gap_duration,
previous_clip=sorted_clips[i].name,
next_clip=sorted_clips[i + 1].name,
))
return gaps
def _detect_duplicate_groups(tl: Any, *, mode: str = "same_source") -> list:
"""Group clips that share a source media reference.
Returns a list of ``DuplicateGroup`` objects. Shared by
``handle_detect_duplicates`` and ``handle_validate_timeline``.
"""
source_groups: dict[str, list[dict]] = {}
for clip in tl.clips:
source_key = clip.media_path or clip.name
if source_key not in source_groups:
source_groups[source_key] = []
source_groups[source_key].append({
'name': clip.name,
'start': clip.start.seconds,
'duration': clip.duration_seconds,
'source_start': clip.source_start.seconds if clip.source_start else 0,
'source_duration': clip.duration_seconds,
'timecode': format_timecode(clip.start),
})
duplicates: list[DuplicateGroup] = []
for source_key, clips in source_groups.items():
if len(clips) <= 1:
continue
group = DuplicateGroup(
source_ref=source_key,
source_name=source_key.split('/')[-1] if '/' in source_key else source_key,
clips=clips,
)
if mode == "same_source":
duplicates.append(group)
elif mode == "overlapping_ranges" and group.has_overlapping_ranges:
duplicates.append(group)
elif mode == "identical":
seen_ranges: set[tuple] = set()
identical_clips = []
for c in clips:
range_key = (c['source_start'], c['source_duration'])
if range_key in seen_ranges:
identical_clips.append(c)
seen_ranges.add(range_key)
if identical_clips:
group.clips = identical_clips
duplicates.append(group)
return duplicates
AUDIO_MEDIA_EXTENSIONS = (
'.wav', '.aif', '.aiff', '.mp3', '.m4a', '.aac', '.flac', '.mov', '.mp4',
)
_DIARIZATION_INSTALL_HINT = (
"\n\nInstall the optional diarization extra:\n\n"
" pip install 'fcp-mcp-server[diarization]'\n\n"
"and set a HuggingFace token with access to "
"pyannote/speaker-diarization-3.1 (pass hf_token= or persist one via "
"save_hf_token)."
)
_FEATURES_INSTALL_HINT = (
"\n\nInstall the optional media-intelligence extra:\n\n"
" pip install 'fcp-mcp-server[intelligence]'"
)
def _voice_analysis_config_text(config: dict) -> str:
w = config["emphasis_weights"]
text = "# Voice Analysis Settings\n\n"
text += _markdown_table(
["Setting", "Value"],
[
["Energy threshold", f"{config['energy_threshold']:.2f}"],
["Peak selection", f"top {config['peak_percentile']:.1%} of words"],
["Emphasis floor", f"{config['emphasis_floor']:.2f}"],
["Emotion detection", "on" if config["emotion_enabled"] else "off"],
["Emotion sensitivity", f"{config['emotion_sensitivity']:.2f}"],
],
) + "\n\n## Emphasis Weights\n"
text += _markdown_table(
["Factor", "Weight"],
[[k.replace("_", " ").title(), f"{v:.2f}"] for k, v in w.items()],
)
return text
def _apply_placed_action(modifier, clip_el, action, clip_start: float) -> str:
"""Apply one non-cut action to the clip that hosts it.
``clip_start`` is where that clip begins on the timeline; the writer
wants times relative to the clip's own head, so the rebase happens here
— the single place that knows about the conversion. The clip *element*
is passed through rather than its name: after a cut the pieces share a
name, and a name lookup would land every edit on the first piece.
"""
rel_start = action.start - clip_start
rel_end = action.end - clip_start
if action.kind == "zoom":
# Only forward an explicit ease — otherwise add_zoom's own default
# (a fast ramp in, instant snap back out) is what should apply.
zoom_args = {}
if action.params.get("ease") is not None:
zoom_args["ease"] = float(action.params["ease"])
if action.params.get("ease_out") is not None:
zoom_args["ease_out"] = float(action.params["ease_out"])
modifier.add_zoom(
clip_id=clip_el,
start=rel_start,
end=rel_end,
scale=float(action.params.get("scale", 1.3)),
**zoom_args,
)
return f"zoom {action.params.get('scale', 1.3):.2f}x"
if action.kind == "text":
modifier.add_text_title(
clip_el,
action.params["content"],
offset=modifier.snap_seconds_to_frame(rel_start).to_fcpxml(),
duration=modifier.snap_seconds_to_frame(action.duration).to_fcpxml(),
)
return f"text \"{action.params['content'][:24]}\""
# marker
modifier.add_marker(
clip_id=clip_el,
timecode=modifier.snap_seconds_to_frame(rel_start).to_fcpxml(),
name=action.params.get("content") or action.reason or "Voice action",
note=action.reason or None,
)
return "marker"
def _speaker_table(profiles: Sequence[dict]) -> str:
"""Who was detected, ordered by how much of the runtime each holds."""
return _markdown_table(
["ID", "Name", "Share", "Speaking", "Lines", "Avg line"],
[
[
p["id"],
p.get("name", ""),
f"{p['share']:.0%}",
format_duration(p["speaking_seconds"]),
str(p["segment_count"]),
f"{p['avg_segment']:.1f}s",
]
for p in profiles
],
)
TRANSCRIBE_MAX_MEDIA = 10
_TRANSCRIBE_INSTALL_HINT = (
"\n\nInstall the optional transcription extra:\n\n"
" pip install 'fcp-mcp-server[transcribe]'\n\n"
"or run via uvx:\n\n"
" uvx --from \"fcp-mcp-server[transcribe]\" fcp-mcp-server"
)
def _transcript_json_path(media_path: str, output_dir: str | None = None) -> Path:
"""Where the ``_transcript.json`` for ``media_path`` lives.
When ``output_dir`` (the user-selected project folder) is set, the
transcript is saved/read there instead of next to the source media.
"""
p = Path(media_path)
if output_dir:
directory = Path(output_dir).expanduser()
directory.mkdir(parents=True, exist_ok=True)
return directory / f"{p.stem}_transcript.json"
return p.with_name(p.stem + "_transcript.json")
def _load_or_transcribe(
media_path: str, model: str, language: str | None, output_dir: str | None = None
) -> tuple[dict | None, str]:
"""Load a cached ``_transcript.json`` for a media file, else transcribe and cache it.
Returns ``(transcript, "")`` or ``(None, reason)``. The cache makes
transcription a one-time cost per media file across all transcript tools.
"""
json_path = _transcript_json_path(media_path, output_dir)
if json_path.is_file():
try:
with open(json_path) as f:
data = json.load(f)
if isinstance(data, dict) and isinstance(data.get("words"), list):
return data, ""
except (OSError, json.JSONDecodeError, UnicodeDecodeError):
pass # unreadable cache falls through to re-transcribe
result = transcribe(media_path, model_size=model, language=language)
if result is None:
return None, "untranscribable (faster-whisper not installed or media unreadable)"
anchor = str(Path(output_dir).expanduser()) if output_dir else str(Path(media_path).parent)
out_path = _validate_output_path(str(json_path), anchor_dir=anchor)
with open(out_path, "w") as f:
json.dump({"source": Path(media_path).name, **result}, f, indent=2)
return result, ""
def _cut_transcript_spans(modifier, clip_filter, model, language, padding, spans_fn, keep_only=False, output_dir=None):
"""Shared cut engine for transcript-driven editing.
``spans_fn(words) -> [(start, end), ...]`` in source seconds. Spans are
padded, clamped to each clip's used source window, optionally inverted
(keep_only), snapped to the frame grid, and cut with ripple.
"""
to_frame = modifier.snap_seconds_to_frame
cache: dict[str, tuple] = {}
cuts_made: list[tuple[str, int, float]] = []
skipped: list[tuple[str, str]] = []
spine_clips = [el for _, el in modifier._iter_spine_clips()]
for el in spine_clips:
name = el.get("name", "")
if clip_filter and name != clip_filter:
continue
src = modifier.resources.get(el.get("ref", ""), {}).get("src", "")
media_path = media_src_to_path(src)
if not media_path or not Path(media_path).is_file():
skipped.append((name, "media file missing"))
continue
if media_path not in cache:
if len(cache) >= TRANSCRIBE_MAX_MEDIA:
skipped.append((name, f"transcription cap reached ({TRANSCRIBE_MAX_MEDIA} media files)"))
continue
cache[media_path] = _load_or_transcribe(media_path, model, language, output_dir)
data, reason = cache[media_path]
if data is None:
skipped.append((name, reason))
continue
clip_source_start = modifier.source_file_start(el).to_seconds()
clip_duration = modifier._parse_time(el.get("duration", "0s")).to_seconds()
window_start = clip_source_start
window_end = clip_source_start + clip_duration
spans = spans_fn(data.get("words", []))
padded = merge_ranges([(s - padding, e + padding) for s, e in spans])
clamped = [
(max(s, window_start), min(e, window_end))
for s, e in padded
if min(e, window_end) > max(s, window_start)
]
if keep_only:
if not clamped:
# Never delete a whole clip just because nothing matched in it.
skipped.append((name, "no phrase matches — left untouched (keep_only)"))
continue
cut_source = invert_ranges(clamped, window_start, window_end)
else:
cut_source = clamped
cut_ranges = [
(to_frame(s - clip_source_start), to_frame(e - clip_source_start))
for s, e in cut_source
]
cut_ranges = [(a, b) for a, b in cut_ranges if b > a]
if not cut_ranges:
continue
removed = modifier.cut_clip_ranges(el, cut_ranges)
if removed > TimeValue.zero():
cuts_made.append((name, len(cut_ranges), removed.to_seconds()))
return cuts_made, skipped
def _transcript_cut_report(title, summary_lines, cuts_made, skipped, output_path, footer):
if not cuts_made:
text = f"# {title}\n\nNo cuts to make — file unchanged (nothing saved)."
if skipped:
text += "\n\n## Skipped Clips\n" + _markdown_table(
["Clip", "Reason"], [[name, reason] for name, reason in skipped]
)
if any("faster-whisper" in reason for _, reason in skipped):
text += _TRANSCRIBE_INSTALL_HINT
return _text_result(text)
total_removed = sum(seconds for _, _, seconds in cuts_made)
result = f"# {title}\n\n## Summary\n"
result += "\n".join(summary_lines) + "\n"
result += f"- **Clips Cut**: {len(cuts_made)}\n- **Total Removed**: {format_duration(total_removed)}\n"
result += "\n## Cuts\n"
result += _markdown_table(
["Clip", "Ranges Cut", "Removed"],
[[name, str(count), f"{seconds:.2f}s"] for name, count, seconds in cuts_made],
) + "\n"
if skipped:
result += "\n## Skipped Clips\n" + _markdown_table(
["Clip", "Reason"], [[name, reason] for name, reason in skipped]
) + "\n"
result += f"\nSaved to: {output_path}\n\n{footer}"
return _text_result(result)
+649
View File
@@ -0,0 +1,649 @@
"""Edição — tool schemas and handlers.
Extracted from server.py; see Engine/docs/03_SERVER_TOOLS.md for the tool catalog.
"""
from __future__ import annotations
from typing import Sequence
from mcp.types import TextContent, Tool
from fcpxml.models import MarkerType
from fcpxml.writer import FCPXMLModifier
from server_tools._shared import (
_format_batch_result,
_resolve_io_paths,
_setup_modifier,
_text_result,
format_duration,
)
TOOLS = [
Tool(
name="add_marker",
description="Add a marker at a specific timecode",
inputSchema={
"type": "object",
"properties": {
"filepath": {"type": "string", "description": "Path to FCPXML file"},
"timecode": {"type": "string", "description": "Position (00:00:10:00 or 10s)"},
"name": {"type": "string", "description": "Marker label"},
"marker_type": {"type": "string", "enum": ["standard", "chapter", "todo", "completed"], "default": "standard"},
"note": {"type": "string", "description": "Optional note"},
"output_path": {"type": "string", "description": "Output path (default: adds _modified suffix)"}
},
"required": ["filepath", "timecode", "name"]
}
),
Tool(
name="batch_add_markers",
description="Add multiple markers at once, or auto-generate at cuts/intervals",
inputSchema={
"type": "object",
"properties": {
"filepath": {"type": "string"},
"markers": {
"type": "array",
"items": {
"type": "object",
"properties": {
"timecode": {"type": "string"},
"name": {"type": "string"},
"marker_type": {"type": "string"},
"note": {"type": "string"}
}
},
"description": "List of markers to add"
},
"auto_at_cuts": {"type": "boolean", "description": "Add marker at every cut"},
"auto_at_intervals": {"type": "string", "description": "Add markers every N seconds (e.g., '30s')"},
"output_path": {"type": "string"}
},
"required": ["filepath"]
}
),
Tool(
name="trim_clip",
description="Trim a clip's in-point and/or out-point",
inputSchema={
"type": "object",
"properties": {
"filepath": {"type": "string"},
"clip_id": {"type": "string", "description": "Clip name or ID"},
"trim_start": {"type": "string", "description": "New in-point or delta (+1s, -10f)"},
"trim_end": {"type": "string", "description": "New out-point or delta"},
"ripple": {"type": "boolean", "default": True, "description": "Shift subsequent clips"},
"output_path": {"type": "string"}
},
"required": ["filepath", "clip_id"]
}
),
Tool(
name="reorder_clips",
description="Move clips to a new position in the timeline",
inputSchema={
"type": "object",
"properties": {
"filepath": {"type": "string"},
"clip_ids": {"type": "array", "items": {"type": "string"}, "description": "Clips to move"},
"target_position": {"type": "string", "description": "'start', 'end', timecode, or 'after:clip_id'"},
"ripple": {"type": "boolean", "default": True},
"output_path": {"type": "string"}
},
"required": ["filepath", "clip_ids", "target_position"]
}
),
Tool(
name="add_transition",
description="Add a transition between clips",
inputSchema={
"type": "object",
"properties": {
"filepath": {"type": "string"},
"clip_id": {"type": "string", "description": "Clip to add transition to"},
"position": {"type": "string", "enum": ["start", "end", "both"], "default": "end"},
"transition_type": {"type": "string", "enum": ["cross-dissolve", "fade-to-black", "fade-from-black", "wipe"], "default": "cross-dissolve"},
"duration": {"type": "string", "default": "00:00:00:15"},
"output_path": {"type": "string"}
},
"required": ["filepath", "clip_id"]
}
),
Tool(
name="change_speed",
description="Change clip playback speed (slow motion or speed up)",
inputSchema={
"type": "object",
"properties": {
"filepath": {"type": "string"},
"clip_id": {"type": "string"},
"speed": {"type": "number", "description": "Speed multiplier (0.5 = half, 2.0 = double)"},
"preserve_pitch": {"type": "boolean", "default": True},
"output_path": {"type": "string"}
},
"required": ["filepath", "clip_id", "speed"]
}
),
Tool(
name="add_zoom",
description="Add a smooth ease-in/ease-out punch-in zoom to a clip, animating <adjust-transform>'s scale param via keyframes (100% -> scale -> 100%) entirely within [start, end] (clip-relative seconds, i.e. seconds from the clip's own head). The ease portions each last `ease` seconds; the zoom holds at `scale` in between. Replaces any existing zoom on the same clip rather than stacking.",
inputSchema={
"type": "object",
"properties": {
"filepath": {"type": "string"},
"clip_id": {"type": "string", "description": "Name/ID of the clip to zoom"},
"start": {"type": "number", "description": "Clip-relative seconds where the ease-in begins"},
"end": {"type": "number", "description": "Clip-relative seconds where the ease-out ends (back to 100%)"},
"scale": {"type": "number", "default": 1.3, "description": "Zoom scale, e.g. 1.3 = 130%"},
"ease": {"type": "number", "default": 0.3, "description": "Seconds for each of the ease-in/ease-out portions (must fit: 2*ease <= end-start)"},
"position": {"type": "string", "default": "0 0", "description": "Optional pan offset \"x y\" applied for the duration of the transform"},
"output_path": {"type": "string"}
},
"required": ["filepath", "clip_id", "start", "end"]
}
),
Tool(
name="delete_clips",
description="Delete clips from timeline",
inputSchema={
"type": "object",
"properties": {
"filepath": {"type": "string"},
"clip_ids": {"type": "array", "items": {"type": "string"}},
"ripple": {"type": "boolean", "default": True, "description": "Close gaps after deletion"},
"output_path": {"type": "string"}
},
"required": ["filepath", "clip_ids"]
}
),
Tool(
name="split_clip",
description="Split a clip at specified timecodes",
inputSchema={
"type": "object",
"properties": {
"filepath": {"type": "string"},
"clip_id": {"type": "string"},
"split_points": {"type": "array", "items": {"type": "string"}, "description": "Timecodes to split at"},
"output_path": {"type": "string"}
},
"required": ["filepath", "clip_id", "split_points"]
}
),
Tool(
name="insert_clip",
description="Insert a library clip onto the timeline at a specific position",
inputSchema={
"type": "object",
"properties": {
"filepath": {"type": "string", "description": "Path to FCPXML file"},
"asset_id": {"type": "string", "description": "Asset reference ID (e.g., 'r3')"},
"asset_name": {"type": "string", "description": "Asset name (alternative to asset_id)"},
"position": {"type": "string", "description": "'start', 'end', timecode, or 'after:clip_name'"},
"duration": {"type": "string", "description": "Clip duration (if not using in/out points)"},
"in_point": {"type": "string", "description": "Source in-point for subclip"},
"out_point": {"type": "string", "description": "Source out-point for subclip"},
"ripple": {"type": "boolean", "default": True, "description": "Shift subsequent clips"},
"output_path": {"type": "string", "description": "Output path (default: adds _modified suffix)"}
},
"required": ["filepath", "position"]
}
),
Tool(
name="fix_flash_frames",
description="Automatically fix detected flash frames by extending neighbors or deleting",
inputSchema={
"type": "object",
"properties": {
"filepath": {"type": "string", "description": "Path to FCPXML file"},
"mode": {"type": "string", "enum": ["extend_previous", "extend_next", "delete", "auto"], "default": "auto", "description": "How to fix: extend previous/next clip, delete, or auto"},
"threshold_frames": {"type": "integer", "default": 6, "description": "Frames below this threshold are flash frames"},
"output_path": {"type": "string", "description": "Output path (default: adds _modified suffix)"}
},
"required": ["filepath"]
}
),
Tool(
name="rapid_trim",
description="Batch trim clips to a maximum duration for fast-paced montages",
inputSchema={
"type": "object",
"properties": {
"filepath": {"type": "string", "description": "Path to FCPXML file"},
"max_duration": {"type": "string", "description": "Maximum clip duration (e.g., '2s', '00:00:02:00')"},
"min_duration": {"type": "string", "description": "Minimum clip duration (optional)"},
"keywords": {"type": "array", "items": {"type": "string"}, "description": "Only trim clips with these keywords"},
"trim_from": {"type": "string", "enum": ["start", "end", "center"], "default": "end", "description": "Where to trim from"},
"output_path": {"type": "string", "description": "Output path (default: adds _modified suffix)"}
},
"required": ["filepath", "max_duration"]
}
),
Tool(
name="fill_gaps",
description="Automatically fill gaps in the timeline by extending adjacent clips",
inputSchema={
"type": "object",
"properties": {
"filepath": {"type": "string", "description": "Path to FCPXML file"},
"mode": {"type": "string", "enum": ["extend_previous", "extend_next", "delete"], "default": "extend_previous", "description": "How to fill gaps"},
"max_gap": {"type": "string", "description": "Only fill gaps smaller than this (e.g., '1s')"},
"output_path": {"type": "string", "description": "Output path (default: adds _modified suffix)"}
},
"required": ["filepath"]
}
),
Tool(
name="add_connected_clip",
description="Connect a library clip to an existing timeline clip (B-roll overlay, audio, title)",
inputSchema={
"type": "object",
"properties": {
"filepath": {"type": "string", "description": "Path to FCPXML file"},
"parent_clip_id": {"type": "string", "description": "Name/ID of the clip to attach to"},
"asset_id": {"type": "string", "description": "Asset reference ID"},
"asset_name": {"type": "string", "description": "Asset name (alternative to asset_id)"},
"offset": {"type": "string", "default": "0s", "description": "Position relative to parent clip start"},
"duration": {"type": "string", "description": "Duration (default: full asset)"},
"lane": {"type": "integer", "default": 1, "description": "Lane number (positive=above, negative=below)"},
"output_path": {"type": "string", "description": "Output path (default: adds _modified suffix)"}
},
"required": ["filepath", "parent_clip_id"]
}
),
Tool(
name="reformat_timeline",
description="Create new FCPXML with different resolution/aspect ratio (9:16 for TikTok, 1:1 for Instagram, etc.)",
inputSchema={
"type": "object",
"properties": {
"filepath": {"type": "string", "description": "Path to FCPXML file"},
"format": {"type": "string", "enum": ["9:16", "1:1", "4:5", "16:9", "4:3", "custom"], "description": "Target format preset"},
"width": {"type": "integer", "description": "Custom width (only with format='custom')"},
"height": {"type": "integer", "description": "Custom height (only with format='custom')"},
"output_path": {"type": "string", "description": "Output path (default: adds _reformatted suffix)"}
},
"required": ["filepath", "format"]
}
),
Tool(
name="add_audio",
description="Add an audio clip or music bed to the timeline",
inputSchema={
"type": "object",
"properties": {
"filepath": {"type": "string", "description": "Path to FCPXML file"},
"parent_clip_id": {"type": "string", "description": "Clip to attach audio to (omit for music bed spanning full timeline)"},
"asset_id": {"type": "string", "description": "Existing asset reference ID"},
"src": {"type": "string", "description": "Path to audio file (creates new asset)"},
"offset": {"type": "string", "description": "Position relative to parent clip start", "default": "0s"},
"duration": {"type": "string", "description": "Duration of audio clip"},
"role": {"type": "string", "description": "Audio role (dialogue, music, effects, etc.)", "default": "dialogue"},
"lane": {"type": "integer", "description": "Lane number (negative = below)", "default": -1},
"output_path": {"type": "string", "description": "Output path"},
},
"required": ["filepath"]
}
),
Tool(
name="create_compound_clip",
description="Group spine clips into a compound clip",
inputSchema={
"type": "object",
"properties": {
"filepath": {"type": "string", "description": "Path to FCPXML file"},
"clip_ids": {"type": "array", "items": {"type": "string"}, "description": "Clip IDs to group"},
"name": {"type": "string", "description": "Name for the compound clip", "default": "Compound Clip"},
"output_path": {"type": "string", "description": "Output path"},
},
"required": ["filepath", "clip_ids"]
}
),
Tool(
name="flatten_compound_clip",
description="Flatten a compound clip back into individual clips in the spine",
inputSchema={
"type": "object",
"properties": {
"filepath": {"type": "string", "description": "Path to FCPXML file"},
"ref_clip_id": {"type": "string", "description": "ID of the ref-clip to flatten"},
"output_path": {"type": "string", "description": "Output path"},
},
"required": ["filepath", "ref_clip_id"]
}
),
]
async def handle_add_marker(arguments: dict) -> Sequence[TextContent]:
filepath, output_path, modifier = _setup_modifier(arguments)
marker_type = MarkerType.from_string(arguments.get("marker_type", "standard"))
modifier.add_marker_at_timeline(
timecode=arguments["timecode"], name=arguments["name"],
marker_type=marker_type, note=arguments.get("note"),
)
modifier.save(output_path)
return _text_result(f"Added marker '{arguments['name']}' at {arguments['timecode']}\n\nSaved to: {output_path}")
async def handle_batch_add_markers(arguments: dict) -> Sequence[TextContent]:
filepath, output_path, modifier = _setup_modifier(arguments)
markers_added = modifier.batch_add_markers(
markers=arguments.get("markers", []),
auto_at_cuts=arguments.get("auto_at_cuts", False),
auto_at_intervals=arguments.get("auto_at_intervals"),
)
modifier.save(output_path)
return _text_result(f"Added {len(markers_added)} markers\n\nSaved to: {output_path}")
async def handle_trim_clip(arguments: dict) -> Sequence[TextContent]:
filepath, output_path, modifier = _setup_modifier(arguments)
modifier.trim_clip(
clip_id=arguments["clip_id"],
trim_start=arguments.get("trim_start"),
trim_end=arguments.get("trim_end"),
ripple=arguments.get("ripple", True),
)
modifier.save(output_path)
return _text_result(f"Trimmed clip '{arguments['clip_id']}'\n\nSaved to: {output_path}")
async def handle_reorder_clips(arguments: dict) -> Sequence[TextContent]:
filepath, output_path, modifier = _setup_modifier(arguments)
modifier.reorder_clips(
clip_ids=arguments["clip_ids"],
target_position=arguments["target_position"],
ripple=arguments.get("ripple", True),
)
modifier.save(output_path)
clips_moved = ", ".join(arguments["clip_ids"])
return _text_result(f"Moved clips [{clips_moved}] to {arguments['target_position']}\n\nSaved to: {output_path}")
async def handle_add_transition(arguments: dict) -> Sequence[TextContent]:
filepath, output_path, modifier = _setup_modifier(arguments)
modifier.add_transition(
clip_id=arguments["clip_id"],
position=arguments.get("position", "end"),
transition_type=arguments.get("transition_type", "cross-dissolve"),
duration=arguments.get("duration", "00:00:00:15"),
)
modifier.save(output_path)
return _text_result(f"Added {arguments.get('transition_type', 'cross-dissolve')} to '{arguments['clip_id']}'\n\nSaved to: {output_path}")
async def handle_change_speed(arguments: dict) -> Sequence[TextContent]:
speed = arguments["speed"]
if not isinstance(speed, (int, float)) or speed <= 0 or speed > 100:
raise ValueError(
f"Speed must be a positive number between 0 (exclusive) and 100, got {speed!r}"
)
filepath, output_path, modifier = _setup_modifier(arguments)
modifier.change_speed(
clip_id=arguments["clip_id"],
speed=speed,
preserve_pitch=arguments.get("preserve_pitch", True),
)
modifier.save(output_path)
speed_desc = f"{speed}x" if speed >= 1 else f"{int(1/speed)}x slow motion"
return _text_result(f"Changed speed of '{arguments['clip_id']}' to {speed_desc}\n\nSaved to: {output_path}")
async def handle_add_zoom(arguments: dict) -> Sequence[TextContent]:
start = float(arguments["start"])
end = float(arguments["end"])
scale = float(arguments.get("scale", 1.3))
ease = float(arguments.get("ease", 0.3))
position = arguments.get("position", "0 0")
filepath, output_path, modifier = _setup_modifier(arguments)
modifier.add_zoom(
clip_id=arguments["clip_id"], start=start, end=end,
scale=scale, ease=ease, position=position,
)
modifier.save(output_path)
return _text_result(
f"# Zoom Added\n\n"
f"- **Clip**: {arguments['clip_id']}\n"
f"- **Window**: {start}s → {end}s (clip-relative)\n"
f"- **Scale**: {int(scale * 100)}%\n"
f"- **Ease**: {ease}s in/out\n\n"
f"Saved to: {output_path}"
)
async def handle_delete_clips(arguments: dict) -> Sequence[TextContent]:
filepath, output_path, modifier = _setup_modifier(arguments)
modifier.delete_clip(
clip_ids=arguments["clip_ids"],
ripple=arguments.get("ripple", True),
)
modifier.save(output_path)
return _text_result(f"Deleted {len(arguments['clip_ids'])} clip(s)\n\nSaved to: {output_path}")
async def handle_split_clip(arguments: dict) -> Sequence[TextContent]:
filepath, output_path, modifier = _setup_modifier(arguments)
new_clips = modifier.split_clip(
clip_id=arguments["clip_id"],
split_points=arguments["split_points"],
)
modifier.save(output_path)
return _text_result(f"Split '{arguments['clip_id']}' into {len(new_clips)} clips\n\nSaved to: {output_path}")
async def handle_insert_clip(arguments: dict) -> Sequence[TextContent]:
filepath, output_path, modifier = _setup_modifier(arguments)
new_clip = modifier.insert_clip(
asset_id=arguments.get("asset_id"),
asset_name=arguments.get("asset_name"),
position=arguments["position"],
duration=arguments.get("duration"),
in_point=arguments.get("in_point"),
out_point=arguments.get("out_point"),
ripple=arguments.get("ripple", True),
)
modifier.save(output_path)
clip_name = new_clip.get('name', 'Unknown')
pos = arguments["position"]
return _text_result(f"Inserted '{clip_name}' at position '{pos}'\n\nSaved to: {output_path}")
async def handle_fix_flash_frames(arguments: dict) -> Sequence[TextContent]:
filepath, output_path, modifier = _setup_modifier(arguments, "_flash_fixed")
fixed = modifier.fix_flash_frames(
mode=arguments.get("mode", "auto"),
threshold_frames=arguments.get("threshold_frames", 6),
)
modifier.save(output_path)
if not fixed:
return _text_result("No flash frames found to fix.")
result = _format_batch_result(
title="Flash Frames Fixed",
summary={"Fixed": f"{len(fixed)} flash frames", "Mode": arguments.get('mode', 'auto')},
headers=["Clip", "Frames", "Action", "Result"],
rows=[
[f['clip_name'], f"{f['duration_frames']}f", f['action'], f"Extended: {f.get('extended_clip', 'N/A')}"]
for f in fixed
],
output_path=output_path,
)
return _text_result(result)
async def handle_rapid_trim(arguments: dict) -> Sequence[TextContent]:
filepath, output_path, modifier = _setup_modifier(arguments, "_rapid_trim")
trimmed = modifier.rapid_trim(
max_duration=arguments["max_duration"],
min_duration=arguments.get("min_duration"),
keywords=arguments.get("keywords"),
trim_from=arguments.get("trim_from", "end"),
)
modifier.save(output_path)
if not trimmed:
return _text_result(f"No clips exceeded {arguments['max_duration']} - nothing trimmed.")
total_before = sum(t['original_duration'] for t in trimmed)
total_after = sum(t['new_duration'] for t in trimmed)
result = _format_batch_result(
title="Rapid Trim Complete",
summary={
"Clips Trimmed": str(len(trimmed)),
"Max Duration": str(arguments['max_duration']),
"Trim From": arguments.get('trim_from', 'end'),
"Time Saved": format_duration(total_before - total_after),
},
headers=["Clip", "Before", "After"],
rows=[
[t['clip_name'], format_duration(t['original_duration']), format_duration(t['new_duration'])]
for t in trimmed
],
output_path=output_path,
)
return _text_result(result)
async def handle_fill_gaps(arguments: dict) -> Sequence[TextContent]:
filepath, output_path, modifier = _setup_modifier(arguments, "_gaps_filled")
filled = modifier.fill_gaps(
mode=arguments.get("mode", "extend_previous"),
max_gap=arguments.get("max_gap"),
)
modifier.save(output_path)
if not filled:
return _text_result("No gaps found to fill.")
result = _format_batch_result(
title="Gaps Filled",
summary={"Gaps Filled": str(len(filled)), "Mode": arguments.get('mode', 'extend_previous')},
headers=["Position", "Duration", "Action"],
rows=[[g['timecode'], f"{g['duration_frames']}f", g['action']] for g in filled],
output_path=output_path,
)
return _text_result(result)
async def handle_add_connected_clip(arguments: dict) -> Sequence[TextContent]:
filepath, output_path, modifier = _setup_modifier(arguments)
modifier.add_connected_clip(
parent_clip_id=arguments["parent_clip_id"],
asset_id=arguments.get("asset_id"),
asset_name=arguments.get("asset_name"),
offset=arguments.get("offset", "0s"),
duration=arguments.get("duration"),
lane=arguments.get("lane", 1),
)
modifier.save(output_path)
return _text_result((
f"Connected clip added to '{arguments['parent_clip_id']}' on lane {arguments.get('lane', 1)}\n\n"
f"Saved to: `{output_path}`"
))
async def handle_reformat_timeline(arguments: dict) -> Sequence[TextContent]:
filepath, output_path = _resolve_io_paths(arguments, "_reformatted")
fmt = arguments["format"]
if fmt == "custom":
width = arguments.get("width")
height = arguments.get("height")
if not width or not height:
return _text_result("Custom format requires both 'width' and 'height' parameters.")
else:
formats = FCPXMLModifier.SOCIAL_FORMATS
if fmt not in formats:
return _text_result(f"Unknown format: {fmt}. Valid: {', '.join(formats.keys())}")
width, height = formats[fmt]
modifier = FCPXMLModifier(filepath)
modifier.reformat_resolution(width, height)
modifier.save(output_path)
return _text_result((
f"# Timeline Reformatted\n\n"
f"- **Format**: {fmt} ({width}x{height})\n"
f"- **Aspect ratio**: {width}:{height}\n\n"
f"Saved to: `{output_path}`\n\n"
f"**Next step**: Import into FCP (File > Import > XML). "
f"FCP will handle spatial conforming automatically."
))
async def handle_add_audio(arguments: dict) -> Sequence[TextContent]:
filepath, output_path, modifier = _setup_modifier(arguments, "_audio")
parent_clip_id = arguments.get("parent_clip_id")
if parent_clip_id:
modifier.add_audio_clip(
parent_clip_id=parent_clip_id,
asset_id=arguments.get("asset_id"),
offset=arguments.get("offset", "0s"),
duration=arguments.get("duration"),
role=arguments.get("role", "dialogue"),
lane=arguments.get("lane", -1),
src=arguments.get("src"),
)
action = f"Added audio clip to '{parent_clip_id}'"
else:
modifier.add_music_bed(
asset_id=arguments.get("asset_id"),
duration=arguments.get("duration"),
role=arguments.get("role", "music"),
src=arguments.get("src"),
)
action = "Added music bed spanning full timeline"
modifier.save(output_path)
return _text_result(f"{action}\nSaved to: `{output_path}`")
async def handle_create_compound_clip(arguments: dict) -> Sequence[TextContent]:
filepath, output_path, modifier = _setup_modifier(arguments, "_compound")
clip_ids = arguments["clip_ids"]
name = arguments.get("name", "Compound Clip")
modifier.create_compound_clip(clip_ids, name)
modifier.save(output_path)
return _text_result((
f"Created compound clip '{name}' from {len(clip_ids)} clips.\n"
f"Saved to: `{output_path}`"
))
async def handle_flatten_compound_clip(arguments: dict) -> Sequence[TextContent]:
filepath, output_path, modifier = _setup_modifier(arguments, "_flattened")
ref_clip_id = arguments["ref_clip_id"]
extracted = modifier.flatten_compound_clip(ref_clip_id)
modifier.save(output_path)
return _text_result((
f"Flattened compound clip '{ref_clip_id}' into {len(extracted)} clips.\n"
f"Saved to: `{output_path}`"
))
HANDLERS = {
"add_marker": handle_add_marker,
"batch_add_markers": handle_batch_add_markers,
"trim_clip": handle_trim_clip,
"reorder_clips": handle_reorder_clips,
"add_transition": handle_add_transition,
"change_speed": handle_change_speed,
"add_zoom": handle_add_zoom,
"delete_clips": handle_delete_clips,
"split_clip": handle_split_clip,
"insert_clip": handle_insert_clip,
"fix_flash_frames": handle_fix_flash_frames,
"rapid_trim": handle_rapid_trim,
"fill_gaps": handle_fill_gaps,
"add_connected_clip": handle_add_connected_clip,
"reformat_timeline": handle_reformat_timeline,
"add_audio": handle_add_audio,
"create_compound_clip": handle_create_compound_clip,
"flatten_compound_clip": handle_flatten_compound_clip,
}
+185
View File
@@ -0,0 +1,185 @@
"""Export / relink — tool schemas and handlers.
Extracted from server.py; see Engine/docs/03_SERVER_TOOLS.md for the tool catalog.
"""
from __future__ import annotations
from typing import Sequence
from mcp.types import TextContent, Tool
from fcpxml.export import DaVinciExporter
from fcpxml.writer import FCPXMLModifier
from server_tools._shared import (
_require_timeline,
_resolve_io_paths,
_setup_modifier,
_text_result,
_validate_filepath,
format_timecode,
)
TOOLS = [
Tool(
name="export_edl",
description="Generate EDL (Edit Decision List) from timeline",
inputSchema={
"type": "object",
"properties": {"filepath": {"type": "string"}},
"required": ["filepath"]
}
),
Tool(
name="export_csv",
description="Export timeline data to CSV format",
inputSchema={
"type": "object",
"properties": {
"filepath": {"type": "string"},
"include": {"type": "array", "items": {"type": "string"}}
},
"required": ["filepath"]
}
),
Tool(
name="export_resolve_xml",
description="Export timeline as DaVinci Resolve compatible FCPXML (simplified v1.9)",
inputSchema={
"type": "object",
"properties": {
"filepath": {"type": "string", "description": "Path to FCPXML file"},
"flatten_compounds": {"type": "boolean", "default": True, "description": "Flatten compound clips for compatibility"},
"output_path": {"type": "string", "description": "Output path (default: adds _resolve suffix)"},
},
"required": ["filepath"]
}
),
Tool(
name="export_fcp7_xml",
description="Export timeline as FCP7 XML (XMEML) for Premiere Pro, DaVinci Resolve, and Avid compatibility",
inputSchema={
"type": "object",
"properties": {
"filepath": {"type": "string", "description": "Path to FCPXML file"},
"output_path": {"type": "string", "description": "Output path (default: adds _fcp7.xml suffix)"},
},
"required": ["filepath"]
}
),
Tool(
name="relink_media",
description="Bulk-rewrite media source paths (asset/media-rep src URLs) to relink moved or renamed media folders without opening FCP. Prefix-based: find='/Volumes/OldDrive/Media' replace='/Volumes/NewDrive/Media'. Use dry_run to preview.",
inputSchema={
"type": "object",
"properties": {
"filepath": {"type": "string", "description": "Path to FCPXML file or .fcpxmld bundle"},
"find": {"type": "string", "description": "Old path prefix to match (plain path or file:// URL)"},
"replace": {"type": "string", "description": "New path prefix to substitute"},
"dry_run": {"type": "boolean", "description": "Preview changes without writing", "default": False},
"output_path": {"type": "string", "description": "Output path (default: adds _relinked suffix)"},
},
"required": ["filepath", "find", "replace"]
}
),
]
async def handle_export_edl(arguments: dict) -> Sequence[TextContent]:
project, tl = _require_timeline(arguments["filepath"])
edl = f"TITLE: {tl.name}\nFCM: NON-DROP FRAME\n\n"
for i, c in enumerate(tl.clips, 1):
edl += f"{i:03d} AX V C {format_timecode(c.source_start)} {format_timecode(c.end)} {format_timecode(c.start)} {format_timecode(c.end)}\n"
edl += f"* FROM CLIP NAME: {c.name}\n\n"
return _text_result(f"```edl\n{edl}```")
async def handle_export_csv(arguments: dict) -> Sequence[TextContent]:
project, tl = _require_timeline(arguments["filepath"])
csv = "Name,Start,End,Duration,Keywords\n"
for c in tl.clips:
kws = "|".join(k.value for k in c.keywords)
csv += f'"{c.name}",{format_timecode(c.start)},{format_timecode(c.end)},{c.duration_seconds:.3f},"{kws}"\n'
return _text_result(f"```csv\n{csv}```")
async def handle_export_resolve_xml(arguments: dict) -> Sequence[TextContent]:
filepath, output_path = _resolve_io_paths(arguments, "_resolve")
exporter = DaVinciExporter(filepath)
exporter.export_simplified_fcpxml(
output_path,
flatten_compounds=arguments.get("flatten_compounds", True),
)
return _text_result((
f"# Exported for DaVinci Resolve\n\n"
f"- **Format**: Simplified FCPXML v1.9\n"
f"- **Compound clips flattened**: {arguments.get('flatten_compounds', True)}\n\n"
f"Saved to: `{output_path}`\n\n"
f"**Next step**: In DaVinci Resolve, go to File > Import > Timeline > Import AAF/EDL/XML"
))
async def handle_export_fcp7_xml(arguments: dict) -> Sequence[TextContent]:
filepath, output_path = _resolve_io_paths(arguments, "_fcp7")
exporter = DaVinciExporter(filepath)
exporter.export_xmeml(output_path)
return _text_result((
f"# Exported as FCP7 XML (XMEML)\n\n"
f"- **Format**: XMEML v5\n"
f"- **Compatible with**: Premiere Pro, DaVinci Resolve, Avid Media Composer\n\n"
f"Saved to: `{output_path}`\n\n"
f"**Next step**: Import via File > Import in your target NLE"
))
async def handle_relink_media(arguments: dict) -> Sequence[TextContent]:
dry_run = arguments.get("dry_run", False)
if dry_run:
filepath = _validate_filepath(arguments["filepath"], ('.fcpxml', '.fcpxmld'))
modifier = FCPXMLModifier(filepath)
result = modifier.relink_media(
arguments["find"], arguments["replace"], dry_run=True
)
footer = "Dry run — no file written."
else:
filepath, output_path, modifier = _setup_modifier(arguments, "_relinked")
result = modifier.relink_media(arguments["find"], arguments["replace"])
saved = modifier.save(output_path)
footer = f"Saved to: {saved}"
if not result["relinked"]:
return _text_result(
f"No media paths matched prefix '{arguments['find']}' "
f"({result['total_assets']} assets scanned). Nothing to relink."
)
lines = [
f"{'Would relink' if dry_run else 'Relinked'} "
f"{result['relinked']} media reference(s) "
f"across {result['total_assets']} asset(s):",
"",
]
missing = 0
for change in result["changes"]:
mark = "✓" if change["target_exists"] else "⚠ target missing"
if not change["target_exists"]:
missing += 1
lines.append(f" {change['asset']}: {change['new']} [{mark}]")
if missing:
lines.append("")
lines.append(
f"⚠ {missing} new path(s) do not exist on this machine — "
f"FCP will show those clips as missing until the media is present."
)
lines.append("")
lines.append(footer)
return _text_result("\n".join(lines))
HANDLERS = {
"export_edl": handle_export_edl,
"export_csv": handle_export_csv,
"export_resolve_xml": handle_export_resolve_xml,
"export_fcp7_xml": handle_export_fcp7_xml,
"relink_media": handle_relink_media,
}
+271
View File
@@ -0,0 +1,271 @@
"""Geração — tool schemas and handlers.
Extracted from server.py; see Engine/docs/03_SERVER_TOOLS.md for the tool catalog.
"""
from __future__ import annotations
from typing import Sequence
from mcp.types import TextContent, Tool
from fcpxml.models import SegmentSpec
from fcpxml.templates import ClipSpec, apply_template, list_templates
from server_tools._shared import (
PROJECTS_DIR,
_setup_generator,
_text_result,
_validate_output_path,
format_duration,
)
TOOLS = [
Tool(
name="auto_rough_cut",
description="Generate a rough cut from source clips based on keywords, duration, and pacing",
inputSchema={
"type": "object",
"properties": {
"filepath": {"type": "string", "description": "Source FCPXML with clips"},
"output_path": {"type": "string", "description": "Where to save rough cut"},
"target_duration": {"type": "string", "description": "Target length (3m, 00:03:00:00)"},
"pacing": {"type": "string", "enum": ["slow", "medium", "fast", "dynamic"], "default": "medium"},
"keywords": {"type": "array", "items": {"type": "string"}, "description": "Filter clips by keywords"},
"segments": {
"type": "array",
"items": {
"type": "object",
"properties": {
"name": {"type": "string"},
"keywords": {"type": "array", "items": {"type": "string"}},
"duration": {"type": "number"}
}
},
"description": "Segment structure [{name, keywords, duration_seconds}]"
},
"priority": {"type": "string", "enum": ["best", "favorites", "longest", "shortest", "random"], "default": "best"},
"favorites_only": {"type": "boolean", "default": False},
"add_transitions": {"type": "boolean", "default": False}
},
"required": ["filepath", "output_path", "target_duration"]
}
),
Tool(
name="generate_montage",
description="Create rapid-fire montages with pacing curves (accelerating, decelerating, pyramid)",
inputSchema={
"type": "object",
"properties": {
"filepath": {"type": "string", "description": "Source FCPXML with clips"},
"output_path": {"type": "string", "description": "Where to save montage"},
"target_duration": {"type": "string", "description": "Total montage length (e.g., '30s', '00:00:30:00')"},
"pacing_curve": {"type": "string", "enum": ["accelerating", "decelerating", "pyramid", "constant"], "default": "accelerating", "description": "How clip duration changes over time"},
"start_duration": {"type": "number", "default": 2.0, "description": "Clip duration at start (seconds)"},
"end_duration": {"type": "number", "default": 0.5, "description": "Clip duration at end (seconds)"},
"keywords": {"type": "array", "items": {"type": "string"}, "description": "Filter clips by keywords"},
"add_transitions": {"type": "boolean", "default": False, "description": "Add quick dissolves"}
},
"required": ["filepath", "output_path", "target_duration"]
}
),
Tool(
name="generate_ab_roll",
description="Create documentary-style A/B roll edits alternating between main content and cutaways",
inputSchema={
"type": "object",
"properties": {
"filepath": {"type": "string", "description": "Source FCPXML with clips"},
"output_path": {"type": "string", "description": "Where to save A/B roll edit"},
"target_duration": {"type": "string", "description": "Total duration (e.g., '3m', '00:03:00:00')"},
"a_keywords": {"type": "array", "items": {"type": "string"}, "description": "Keywords for A-roll (main content, interviews)"},
"b_keywords": {"type": "array", "items": {"type": "string"}, "description": "Keywords for B-roll (cutaways, visuals)"},
"a_duration": {"type": "string", "default": "5s", "description": "Duration of each A-roll segment"},
"b_duration": {"type": "string", "default": "3s", "description": "Duration of each B-roll cutaway"},
"start_with": {"type": "string", "enum": ["a", "b"], "default": "a", "description": "Which roll to start with"},
"add_transitions": {"type": "boolean", "default": True, "description": "Add cross-dissolves"}
},
"required": ["filepath", "output_path", "target_duration", "a_keywords", "b_keywords"]
}
),
Tool(
name="list_templates",
description="List available timeline templates with slot definitions",
inputSchema={
"type": "object",
"properties": {},
}
),
Tool(
name="apply_template",
description="Fill a timeline template with clips and generate FCPXML",
inputSchema={
"type": "object",
"properties": {
"template_name": {"type": "string", "description": "Template name (intro_outro, lower_thirds, music_video)"},
"clips": {"type": "object", "description": "Map of slot_name -> {src, name, duration} or {asset_id, name, duration}"},
"output_path": {"type": "string", "description": "Output FCPXML path"},
"fps": {"type": "number", "description": "Frame rate", "default": 24},
},
"required": ["template_name", "clips", "output_path"]
}
),
]
async def handle_auto_rough_cut(arguments: dict) -> Sequence[TextContent]:
filepath, output_path, generator = _setup_generator(arguments, "_roughcut")
segments = None
if arguments.get("segments"):
segments = [
SegmentSpec(
name=s.get("name", "Segment"),
keywords=s.get("keywords", []),
duration_seconds=s.get("duration", 0),
priority=s.get("priority", "best"),
)
for s in arguments["segments"]
]
result = generator.generate(
output_path=output_path,
target_duration=arguments["target_duration"],
pacing=arguments.get("pacing", "medium"),
keywords=arguments.get("keywords"),
segments=segments,
priority=arguments.get("priority", "best"),
favorites_only=arguments.get("favorites_only", False),
add_transitions=arguments.get("add_transitions", False),
)
return _text_result(f"""# Rough Cut Generated
## Summary
- **Clips Used**: {result.clips_used} of {result.clips_available} available
- **Target Duration**: {format_duration(result.target_duration)}
- **Actual Duration**: {format_duration(result.actual_duration)}
- **Average Clip**: {format_duration(result.average_clip_duration)}
## Output
Saved to: `{result.output_path}`
**Next step**: Import this FCPXML into Final Cut Pro (File > Import > XML)
""")
async def handle_generate_montage(arguments: dict) -> Sequence[TextContent]:
filepath, output_path, generator = _setup_generator(arguments, "_montage")
result = generator.generate_montage(
output_path=output_path,
target_duration=arguments["target_duration"],
pacing_curve=arguments.get("pacing_curve", "accelerating"),
start_duration=arguments.get("start_duration", 2.0),
end_duration=arguments.get("end_duration", 0.5),
keywords=arguments.get("keywords"),
add_transitions=arguments.get("add_transitions", False),
)
curve_desc = {
'accelerating': 'slow to fast (builds energy)',
'decelerating': 'fast to slow (winds down)',
'pyramid': 'slow to fast to slow (dramatic arc)',
'constant': 'same duration throughout',
}
return _text_result(f"""# Montage Generated
## Summary
- **Clips Used**: {result['clips_used']} of {result['clips_available']} available
- **Target Duration**: {format_duration(result['target_duration'])}
- **Actual Duration**: {format_duration(result['actual_duration'])}
- **Pacing Curve**: {result['pacing_curve']} - {curve_desc.get(result['pacing_curve'], '')}
## Pacing
- **Start Clip Duration**: {format_duration(result['start_clip_duration'])}
- **End Clip Duration**: {format_duration(result['end_clip_duration'])}
## Output
Saved to: `{result['output_path']}`
""")
async def handle_generate_ab_roll(arguments: dict) -> Sequence[TextContent]:
filepath, output_path, generator = _setup_generator(arguments, "_ab_roll")
result = generator.generate_ab_roll(
output_path=output_path,
target_duration=arguments["target_duration"],
a_keywords=arguments["a_keywords"],
b_keywords=arguments["b_keywords"],
a_duration=arguments.get("a_duration", "5s"),
b_duration=arguments.get("b_duration", "3s"),
start_with=arguments.get("start_with", "a"),
add_transitions=arguments.get("add_transitions", True),
)
return _text_result(f"""# A/B Roll Edit Generated
## Summary
- **A-Roll Segments**: {result['a_segments']} (from {result['a_clips_available']} available)
- **B-Roll Segments**: {result['b_segments']} (from {result['b_clips_available']} available)
- **Total Clips**: {result['clips_used']}
## Timing
- **Target Duration**: {format_duration(result['target_duration'])}
- **Actual Duration**: {format_duration(result['actual_duration'])}
- **A-Roll Duration**: {result['a_duration_setting']} per segment
- **B-Roll Duration**: {result['b_duration_setting']} per cutaway
## Output
Saved to: `{result['output_path']}`
**Next step**: Import this FCPXML into Final Cut Pro (File > Import > XML)
""")
async def handle_list_templates(arguments: dict) -> Sequence[TextContent]:
templates = list_templates()
lines = ["# Available Timeline Templates\n"]
for tmpl in templates:
lines.append(f"## {tmpl['name']}")
lines.append(f"{tmpl['description']}\n")
lines.append("| Slot | Type | Default Duration | Lane | Required |")
lines.append("|------|------|-----------------|------|----------|")
for s in tmpl['slots']:
lines.append(
f"| {s['name']} | {s['slot_type']} | {s['default_duration']}s "
f"| {s['lane']} | {'Yes' if s['required'] else 'No'} |"
)
lines.append("")
return _text_result("\n".join(lines))
async def handle_apply_template(arguments: dict) -> Sequence[TextContent]:
template_name = arguments["template_name"]
clips_raw = arguments["clips"]
output_path = _validate_output_path(arguments["output_path"], anchor_dir=PROJECTS_DIR)
fps = arguments.get("fps", 24)
# Convert raw clips dict to ClipSpec objects
clips_map = {}
for slot_name, spec_data in clips_raw.items():
if isinstance(spec_data, dict):
clips_map[slot_name] = ClipSpec(
asset_id=spec_data.get("asset_id"),
src=spec_data.get("src"),
name=spec_data.get("name", slot_name),
duration=spec_data.get("duration"),
)
result_path = apply_template(template_name, clips_map, output_path, fps)
return _text_result((
f"Applied template '{template_name}' with {len(clips_map)} clips.\n"
f"Saved to: `{result_path}`"
))
HANDLERS = {
"auto_rough_cut": handle_auto_rough_cut,
"generate_montage": handle_generate_montage,
"generate_ab_roll": handle_generate_ab_roll,
"list_templates": handle_list_templates,
"apply_template": handle_apply_template,
}
+110
View File
@@ -0,0 +1,110 @@
"""Live (macOS) — tool schemas and handlers.
Extracted from server.py; see Engine/docs/03_SERVER_TOOLS.md for the tool catalog.
"""
from __future__ import annotations
from pathlib import Path
from typing import Sequence
from mcp.types import TextContent, Tool
from server_tools._shared import (
_text_result,
_validate_filepath,
_validate_output_path,
generate_output_path,
)
TOOLS = [
Tool(
name="push_to_fcp",
description="LIVE: send an FCPXML file into the running Final Cut Pro with zero clicks (official Open Document Apple event). Creates/targets a library via import-options. Launches FCP if needed. macOS-only; first use triggers an Automation permission prompt. For true zero-click, pass a library_location ending in .fcpbundle (a new path is auto-created); omitting it makes FCP show a modal library picker.",
inputSchema={
"type": "object",
"properties": {
"filepath": {"type": "string", "description": "Path to FCPXML file or .fcpxmld bundle to import"},
"library_location": {"type": "string", "description": "Target .fcpbundle library path (auto-created if it doesn't exist; the extension is normalized to .fcpbundle). Omit to import into the active library, but note FCP then shows a modal 'Open Library' picker that blocks until answered"},
"suppress_warnings": {"type": "boolean", "description": "Suppress non-fatal import warning dialogs", "default": True},
"copy_assets": {"type": "boolean", "description": "Copy media into the library (true) or link in place (false). Omit for FCP default"},
},
"required": ["filepath"]
}
),
Tool(
name="list_fcp_libraries",
description="LIVE: enumerate the running Final Cut Pro's open libraries, events, and projects via Apple's read-only scripting dictionary. Refuses to launch FCP unless allow_launch is true. macOS-only.",
inputSchema={
"type": "object",
"properties": {
"allow_launch": {"type": "boolean", "description": "Launch FCP if it isn't running", "default": False},
},
}
),
]
async def handle_push_to_fcp(arguments: dict) -> Sequence[TextContent]:
from fcpxml.live import push_to_fcp
filepath = _validate_filepath(arguments["filepath"], ('.fcpxml', '.fcpxmld'))
# Flat files get an options-injected sibling copy (never touch the
# original); the copy path goes through the same write sandbox as
# every other derived output.
import_copy = None
if Path(filepath).suffix.lower() == '.fcpxml':
anchor = str(Path(filepath).resolve().parent)
import_copy = _validate_output_path(
generate_output_path(filepath, "_import"), anchor_dir=anchor
)
result = push_to_fcp(
filepath,
library_location=arguments.get("library_location"),
suppress_warnings=arguments.get("suppress_warnings", True),
copy_assets=arguments.get("copy_assets"),
import_copy_path=import_copy,
)
lines = [
f"Sent to Final Cut Pro: {result['sent']}",
f"FCP {'was launched' if result['launched_fcp'] else 'was already running'} — "
f"import happens in-app (libraries/events are created or merged per import-options).",
]
if arguments.get("library_location"):
lines.append(f"Target library: {arguments['library_location']}")
lines.append(
"Note: Apple offers no programmatic export — to round-trip edits "
"back, use File > Export XML in FCP."
)
return _text_result("\n".join(lines))
async def handle_list_fcp_libraries(arguments: dict) -> Sequence[TextContent]:
from fcpxml.live import list_fcp_libraries
try:
libraries = list_fcp_libraries(
allow_launch=arguments.get("allow_launch", False)
)
except RuntimeError as exc:
return _text_result(str(exc))
if not libraries:
return _text_result("Final Cut Pro is running but reports no open libraries.")
lines = [f"Open libraries in Final Cut Pro ({len(libraries)}):", ""]
for lib in libraries:
lines.append(f"📚 {lib['name']}")
for event in lib["events"]:
lines.append(f" └─ {event['name']}")
for proj in event["projects"]:
lines.append(f" • {proj}")
return _text_result("\n".join(lines))
HANDLERS = {
"push_to_fcp": handle_push_to_fcp,
"list_fcp_libraries": handle_list_fcp_libraries,
}
+458
View File
@@ -0,0 +1,458 @@
"""Beats / markers importados — tool schemas and handlers.
Extracted from server.py; see Engine/docs/03_SERVER_TOOLS.md for the tool catalog.
"""
from __future__ import annotations
import json
import re
from pathlib import Path
from typing import Sequence
from mcp.types import TextContent, Tool
from fcpxml.media_intel import media_src_to_path
from fcpxml.parser import FCPXMLParser
from fcpxml.writer import FCPXMLModifier
from server_tools._shared import (
_check_json_depth,
_load_or_transcribe,
_markdown_table,
_no_timeline,
_raw_markers_to_batch,
_resolve_io_paths,
_setup_modifier,
_text_result,
_validate_filepath,
format_duration,
parse_srt,
parse_transcript_timestamps,
parse_vtt,
)
TOOLS = [
Tool(
name="import_beat_markers",
description="Import beat markers from external audio analysis (JSON format)",
inputSchema={
"type": "object",
"properties": {
"filepath": {"type": "string", "description": "Path to FCPXML file"},
"beats_path": {"type": "string", "description": "Path to beats JSON file"},
"marker_type": {"type": "string", "enum": ["standard", "chapter"], "default": "standard"},
"beat_filter": {"type": "string", "enum": ["all", "downbeat", "measure"], "default": "all", "description": "Which beats to import"},
"output_path": {"type": "string", "description": "Output path (default: adds _beats suffix)"}
},
"required": ["filepath", "beats_path"]
}
),
Tool(
name="snap_to_beats",
description="Align cuts to nearest beat markers for music-synced edits",
inputSchema={
"type": "object",
"properties": {
"filepath": {"type": "string", "description": "Path to FCPXML file with beat markers"},
"max_shift_frames": {"type": "integer", "default": 6, "description": "Maximum frames to shift a cut"},
"prefer": {"type": "string", "enum": ["earlier", "later", "nearest"], "default": "nearest", "description": "Which beat to prefer when equidistant"},
"output_path": {"type": "string", "description": "Output path (default: adds _synced suffix)"}
},
"required": ["filepath"]
}
),
Tool(
name="import_srt_markers",
description="Import SRT or VTT subtitles as chapter markers on the timeline",
inputSchema={
"type": "object",
"properties": {
"filepath": {"type": "string", "description": "Path to FCPXML file"},
"srt_path": {"type": "string", "description": "Path to SRT or VTT subtitle file"},
"mode": {"type": "string", "enum": ["all", "first_per_minute", "scene_changes"], "default": "first_per_minute", "description": "How to create markers: every subtitle, first per minute, or on text changes"},
"marker_type": {"type": "string", "enum": ["standard", "chapter"], "default": "chapter"},
"max_label_length": {"type": "integer", "default": 50, "description": "Truncate marker labels to this length"},
"output_path": {"type": "string", "description": "Output path (default: adds _subtitled suffix)"}
},
"required": ["filepath", "srt_path"]
}
),
Tool(
name="import_transcript_markers",
description="Import timestamped transcript (YouTube chapter format) as markers. Supports '0:00 Title' and 'HH:MM:SS Title' formats",
inputSchema={
"type": "object",
"properties": {
"filepath": {"type": "string", "description": "Path to FCPXML file"},
"transcript": {"type": "string", "description": "Timestamped text (one per line: '0:00 Introduction')"},
"transcript_path": {"type": "string", "description": "Path to text file with timestamps (alternative to inline transcript)"},
"marker_type": {"type": "string", "enum": ["standard", "chapter"], "default": "chapter"},
"output_path": {"type": "string", "description": "Output path (default: adds _chapters suffix)"}
},
"required": ["filepath"]
}
),
Tool(
name="transcript_markers",
description="Add a marker at the start of every transcribed segment (sentence-level), using each media file's local Whisper transcript. Maps each segment's source-media timestamp to its correct timeline position per clip, so it stays accurate across multiple clips/trims — unlike import_transcript_markers (plain timestamp text) or import_srt_markers (a caption track already synced to the whole export). Uses each media file's _transcript.json (auto-transcribes if missing). Non-destructive: writes a _transcript_markers copy.",
inputSchema={
"type": "object",
"properties": {
"filepath": {"type": "string", "description": "Path to FCPXML file"},
"clip_name": {"type": "string", "description": "Only mark the clip with this name"},
"marker_type": {"type": "string", "default": "chapter", "description": "Marker type: standard, chapter, todo, completed"},
"max_label_length": {"type": "integer", "default": 50, "description": "Truncate marker labels to this many characters (0 = no truncation)"},
"model": {"type": "string", "default": "base", "description": "Whisper model size if transcription is needed"},
"output_path": {"type": "string", "description": "Output path (default: adds _transcript_markers suffix)"},
},
"required": ["filepath"]
}
),
]
async def handle_import_beat_markers(arguments: dict) -> Sequence[TextContent]:
filepath, output_path = _resolve_io_paths(arguments, "_beats")
beats_path = _validate_filepath(arguments["beats_path"], ('.json',))
with open(beats_path, 'r') as f:
beats_data = json.load(f)
_check_json_depth(beats_data)
beat_times = []
if isinstance(beats_data, list):
beat_times = beats_data
elif isinstance(beats_data, dict):
beat_times = beats_data.get('beats', beats_data.get('times', beats_data.get('markers', [])))
beat_filter = arguments.get("beat_filter", "all")
if beat_filter == "downbeat" and isinstance(beats_data, dict):
beat_times = beats_data.get('downbeats', beat_times[::4])
elif beat_filter == "measure" and isinstance(beats_data, dict):
beat_times = beats_data.get('measures', beat_times[::4])
markers = []
marker_type = arguments.get("marker_type", "standard")
for i, beat_time in enumerate(beat_times):
if isinstance(beat_time, (int, float)):
markers.append({
'timecode': f"{beat_time}s",
'name': f"Beat {i+1}",
'marker_type': marker_type.upper(),
})
elif isinstance(beat_time, dict):
markers.append({
'timecode': f"{beat_time.get('time', beat_time.get('position', 0))}s",
'name': beat_time.get('label', f"Beat {i+1}"),
'marker_type': marker_type.upper(),
})
modifier = FCPXMLModifier(filepath)
# Songs routinely run longer than the edit — beats past the timeline's
# end are skipped (add_marker_at_timeline would raise on them).
timeline_end = modifier._timeline_duration().to_seconds()
in_range = [m for m in markers if float(m['timecode'].rstrip('s')) < timeline_end]
skipped_count = len(markers) - len(in_range)
added = modifier.batch_add_markers(markers=in_range)
modifier.save(output_path)
skipped_note = (
f"- **Skipped**: {skipped_count} beat(s) beyond the timeline end "
f"({format_duration(timeline_end)})\n" if skipped_count else ""
)
return _text_result(f"""# Beat Markers Imported
## Summary
- **Beats Found**: {len(beat_times)}
- **Markers Added**: {len(added)}
{skipped_note}- **Filter**: {beat_filter}
- **Marker Type**: {marker_type}
## Output
Saved to: `{output_path}`
*Use `snap_to_beats` to align your cuts to these markers.*
""")
async def handle_snap_to_beats(arguments: dict) -> Sequence[TextContent]:
filepath, output_path = _resolve_io_paths(arguments, "_synced")
max_shift = arguments.get("max_shift_frames", 6)
prefer = arguments.get("prefer", "nearest")
parser = FCPXMLParser()
project = parser.parse_file(filepath)
if not project.timelines:
return _no_timeline()
tl = project.primary_timeline
fps = tl.frame_rate
markers = list(tl.markers)
for clip in tl.clips:
markers.extend(clip.markers)
if not markers:
return _text_result("No markers found. Use `import_beat_markers` first.")
marker_times = sorted([m.start.seconds for m in markers])
modifier = FCPXMLModifier(filepath)
spine = modifier._get_spine()
adjusted_count = 0
total_shift = 0
clips_list = [c for c in spine if c.tag in ('clip', 'asset-clip', 'video', 'ref-clip')]
for i, clip in enumerate(clips_list[1:], 1):
cut_offset = modifier._parse_time(clip.get('offset', '0s'))
cut_seconds = cut_offset.to_seconds()
best_marker = None
best_distance = float('inf')
for marker_time in marker_times:
distance = abs(marker_time - cut_seconds)
distance_frames = distance * fps
if distance_frames <= max_shift:
if prefer == "earlier" and marker_time <= cut_seconds:
if distance < best_distance:
best_distance = distance
best_marker = marker_time
elif prefer == "later" and marker_time >= cut_seconds:
if distance < best_distance:
best_distance = distance
best_marker = marker_time
elif prefer == "nearest":
if distance < best_distance:
best_distance = distance
best_marker = marker_time
if best_marker is not None and best_distance > 0.001:
shift = best_marker - cut_seconds
shift_frames = int(shift * fps)
prev_clip = clips_list[i - 1]
prev_dur = modifier._parse_time(prev_clip.get('duration', '0s'))
new_prev_dur = prev_dur + modifier._parse_time(f"{shift}s")
prev_clip.set('duration', new_prev_dur.to_fcpxml())
new_offset = modifier._parse_time(f"{best_marker}s")
clip.set('offset', new_offset.to_fcpxml())
adjusted_count += 1
total_shift += abs(shift_frames)
modifier.save(output_path)
avg_shift = total_shift / adjusted_count if adjusted_count > 0 else 0
return _text_result(f"""# Cuts Snapped to Beats
## Summary
- **Cuts Adjusted**: {adjusted_count}
- **Max Shift Allowed**: {max_shift} frames
- **Preference**: {prefer}
- **Average Shift**: {avg_shift:.1f} frames
## Output
Saved to: `{output_path}`
Your edits are now synced to the beat!
""")
async def handle_import_srt_markers(arguments: dict) -> Sequence[TextContent]:
filepath, output_path = _resolve_io_paths(arguments, "_subtitled")
srt_path = _validate_filepath(arguments["srt_path"], ('.srt', '.vtt'))
mode = arguments.get("mode", "first_per_minute")
marker_type = arguments.get("marker_type", "chapter")
max_label = arguments.get("max_label_length", 50)
text = Path(srt_path).read_text(encoding='utf-8')
# Detect format and parse
if srt_path.endswith('.vtt') or text.strip().startswith('WEBVTT'):
raw_markers = parse_vtt(text)
fmt_name = "WebVTT"
else:
raw_markers = parse_srt(text)
fmt_name = "SRT"
if not raw_markers:
return _text_result(f"No subtitles found in {srt_path}")
# Apply mode filtering
filtered = []
if mode == "all":
filtered = raw_markers
elif mode == "first_per_minute":
seen_minutes = set()
for m in raw_markers:
minute = int(m['seconds'] // 60)
if minute not in seen_minutes:
seen_minutes.add(minute)
filtered.append(m)
elif mode == "scene_changes":
# Group by similar text, take first occurrence of each unique line
seen_texts = set()
for m in raw_markers:
# Normalize: lowercase, strip punctuation
normalized = re.sub(r'[^\w\s]', '', m['text'].lower()).strip()
words = normalized.split()[:3] # First 3 words as key
key = ' '.join(words)
if key and key not in seen_texts:
seen_texts.add(key)
filtered.append(m)
markers = _raw_markers_to_batch(filtered, marker_type, max_label=max_label)
modifier = FCPXMLModifier(filepath)
added = modifier.batch_add_markers(markers=markers)
modifier.save(output_path)
return _text_result(f"""# Subtitle Markers Imported
## Summary
- **Format**: {fmt_name}
- **Subtitles Parsed**: {len(raw_markers)}
- **Mode**: {mode}
- **Markers Added**: {len(added)}
- **Marker Type**: {marker_type}
## Output
Saved to: `{output_path}`
""")
async def handle_import_transcript_markers(arguments: dict) -> Sequence[TextContent]:
filepath, output_path = _resolve_io_paths(arguments, "_chapters")
marker_type = arguments.get("marker_type", "chapter")
# Get transcript text from inline or file
transcript = arguments.get("transcript")
transcript_path = arguments.get("transcript_path")
if not transcript and not transcript_path:
return _text_result("Provide either 'transcript' (inline text) or 'transcript_path' (path to file)")
if transcript_path:
# .txt only: the parser below understands "0:00 Title" lines, not real
# SRT/VTT cue syntax — that's import_srt_markers (parse_srt/parse_vtt).
transcript_path = _validate_filepath(transcript_path, ('.txt',))
transcript = Path(transcript_path).read_text(encoding='utf-8')
raw_markers = parse_transcript_timestamps(transcript or "")
if not raw_markers:
return _text_result("No timestamps found. Expected format: '0:00 Title' or 'HH:MM:SS Title', one per line.")
markers = _raw_markers_to_batch(raw_markers, marker_type)
modifier = FCPXMLModifier(filepath)
added = modifier.batch_add_markers(markers=markers)
modifier.save(output_path)
return _text_result(f"""# Transcript Markers Imported
## Summary
- **Timestamps Found**: {len(raw_markers)}
- **Markers Added**: {len(added)}
- **Marker Type**: {marker_type}
## Markers
""" + "\n".join(f"- `{m['timecode']}` {m['name']}" for m in markers) + f"""
## Output
Saved to: `{output_path}`
""")
async def handle_transcript_markers(arguments: dict) -> Sequence[TextContent]:
"""Add a marker at the start of each transcribed segment, using each
media's cached (or freshly transcribed) local Whisper transcript.
Unlike ``import_transcript_markers`` (plain "0:00 Title" text) or
``import_srt_markers`` (a caption track already synced to the whole
exported video), this maps each segment's SOURCE-media timestamp to its
TIMELINE position per spine clip — the same source->timeline mapping
``detect_media_silence`` uses — so it stays correct across multiple
clips built from different (and differently-trimmed) source files.
"""
marker_type = arguments.get("marker_type", "chapter")
max_label = int(arguments.get("max_label_length", 50))
model = arguments.get("model", "base")
language = arguments.get("language")
output_dir = arguments.get("output_dir")
clip_filter = arguments.get("clip_name")
filepath, output_path, modifier = _setup_modifier(arguments, "_transcript_markers")
added: list[tuple[str, float, str]] = []
skipped: list[tuple[str, str]] = []
spine_clips = [el for _, el in modifier._iter_spine_clips()]
for el in spine_clips:
name = el.get("name", "")
if clip_filter and name != clip_filter:
continue
src = modifier.resources.get(el.get("ref", ""), {}).get("src", "")
media_path = media_src_to_path(src)
if not media_path or not Path(media_path).is_file():
skipped.append((name, "media file missing"))
continue
data, reason = _load_or_transcribe(media_path, model, language, output_dir)
if data is None:
skipped.append((name, reason))
continue
clip_source_start = modifier.source_file_start(el).to_seconds()
clip_duration = modifier._parse_time(el.get("duration", "0s")).to_seconds()
clip_offset = modifier._parse_time(el.get("offset", "0s")).to_seconds()
window_end = clip_source_start + clip_duration
for seg in data.get("segments", []):
seg_start = float(seg.get("start", 0.0))
if seg_start < clip_source_start or seg_start >= window_end:
continue
label = seg.get("text", "").strip()
if not label:
continue
if max_label and len(label) > max_label:
label = label[:max_label]
timeline_seconds = clip_offset + (seg_start - clip_source_start)
modifier.add_marker_at_timeline(
timecode=f"{timeline_seconds}s", name=label, marker_type=marker_type,
)
added.append((name, seg_start, label))
if not added:
text = "# Transcript Markers\n\nNo segments to mark — file unchanged (nothing saved)."
if skipped:
text += "\n\n## Skipped Clips\n" + _markdown_table(
["Clip", "Reason"], [[n, r] for n, r in skipped]
)
return _text_result(text)
modifier.save(output_path)
result = "# Transcript Markers Imported (local Whisper)\n\n## Summary\n"
result += f"- **Markers Added**: {len(added)}\n- **Marker Type**: {marker_type}\n\n"
result += _markdown_table(
["Clip", "Start", "Label"], [[n, f"{s:.2f}s", label] for n, s, label in added]
)
if skipped:
result += "\n## Skipped Clips\n" + _markdown_table(
["Clip", "Reason"], [[n, r] for n, r in skipped]
)
result += f"\n\nSaved to: `{output_path}`\n\n*Transcripts are cached as _transcript.json.*"
return _text_result(result)
HANDLERS = {
"import_beat_markers": handle_import_beat_markers,
"snap_to_beats": handle_snap_to_beats,
"import_srt_markers": handle_import_srt_markers,
"import_transcript_markers": handle_import_transcript_markers,
"transcript_markers": handle_transcript_markers,
}
+696
View File
@@ -0,0 +1,696 @@
"""QC e detecção — tool schemas and handlers.
Extracted from server.py; see Engine/docs/03_SERVER_TOOLS.md for the tool catalog.
"""
from __future__ import annotations
import json
from pathlib import Path
from typing import Sequence
from mcp.types import TextContent, Tool
from fcpxml.media_intel import (
detect_beats,
detect_silence,
map_silence_to_timeline,
media_src_to_path,
)
from fcpxml.model_manager import load_silence_config
from fcpxml.models import FlashFrameSeverity, TimeValue
from fcpxml.writer import FCPXMLModifier
from server_tools._shared import (
AUDIO_MEDIA_EXTENSIONS,
MAX_MEDIA_FILE_SIZE,
_detect_duplicate_groups,
_detect_flash_frames,
_detect_gaps,
_fmt_suggestions,
_format_clip_table,
_markdown_table,
_require_timeline,
_setup_modifier,
_text_result,
_validate_filepath,
_validate_output_path,
format_duration,
format_timecode,
)
TOOLS = [
Tool(
name="find_short_cuts",
description="Find clips shorter than threshold (flash frame detection)",
inputSchema={
"type": "object",
"properties": {
"filepath": {"type": "string"},
"threshold_seconds": {"type": "number", "default": 0.5}
},
"required": ["filepath"]
}
),
Tool(
name="find_long_clips",
description="Find clips longer than threshold",
inputSchema={
"type": "object",
"properties": {
"filepath": {"type": "string"},
"threshold_seconds": {"type": "number", "default": 10.0}
},
"required": ["filepath"]
}
),
Tool(
name="analyze_pacing",
description="Analyze edit pacing with suggestions for improvements",
inputSchema={
"type": "object",
"properties": {"filepath": {"type": "string"}},
"required": ["filepath"]
}
),
Tool(
name="detect_flash_frames",
description="Find ultra-short clips (flash frames) that are likely errors, with severity categorization",
inputSchema={
"type": "object",
"properties": {
"filepath": {"type": "string", "description": "Path to FCPXML file"},
"critical_threshold_frames": {"type": "integer", "default": 2, "description": "Frames below this = critical (default: 2)"},
"warning_threshold_frames": {"type": "integer", "default": 6, "description": "Frames below this = warning (default: 6)"}
},
"required": ["filepath"]
}
),
Tool(
name="detect_duplicates",
description="Find clips using the same source media (potential duplicates)",
inputSchema={
"type": "object",
"properties": {
"filepath": {"type": "string", "description": "Path to FCPXML file"},
"mode": {"type": "string", "enum": ["same_source", "overlapping_ranges", "identical"], "default": "same_source", "description": "Detection mode"}
},
"required": ["filepath"]
}
),
Tool(
name="detect_gaps",
description="Find unintentional gaps in the timeline",
inputSchema={
"type": "object",
"properties": {
"filepath": {"type": "string", "description": "Path to FCPXML file"},
"min_gap_frames": {"type": "integer", "default": 1, "description": "Minimum gap size to detect (default: 1 frame)"}
},
"required": ["filepath"]
}
),
Tool(
name="validate_timeline",
description="Comprehensive timeline health check for flash frames, gaps, duplicates, and issues",
inputSchema={
"type": "object",
"properties": {
"filepath": {"type": "string", "description": "Path to FCPXML file"},
"checks": {"type": "array", "items": {"type": "string", "enum": ["all", "flash_frames", "gaps", "duplicates", "offsets"]}, "default": ["all"], "description": "Which checks to run"}
},
"required": ["filepath"]
}
),
Tool(
name="detect_media_silence",
description="Detect REAL silence by analyzing each clip's source audio with ffmpeg silencedetect, mapped into timeline time. Unlike detect_silence_candidates (XML-only heuristics), this reads the actual media files referenced by the timeline. Requires ffmpeg; clips whose media is missing or unreadable are reported, not failed.",
inputSchema={
"type": "object",
"properties": {
"filepath": {"type": "string", "description": "Path to FCPXML file"},
"noise_db": {"type": "number", "description": "Silence threshold in dBFS, -120 to 0. Falls back to the saved silence settings (default -30)"},
"min_silence": {"type": "number", "description": "Minimum silence duration in seconds to report. Falls back to the saved silence settings (default 0.5)"},
"clip_name": {"type": "string", "description": "Only analyze the clip with this name"},
},
"required": ["filepath"]
}
),
Tool(
name="detect_beats",
description="Detect musical beats and tempo in an audio/video file (librosa beat tracker). Writes a beats JSON next to the media file that plugs directly into import_beat_markers + snap_to_beats for beat-synced editing. Requires the optional [intelligence] extra (librosa); degrades to an install hint without it.",
inputSchema={
"type": "object",
"properties": {
"media_path": {"type": "string", "description": "Path to audio/video file (.wav, .mp3, .m4a, .aac, .aif, .flac, .mov, .mp4)"},
},
"required": ["media_path"]
}
),
Tool(
name="remove_media_silence",
description="Detect REAL silence in each clip's source audio (ffmpeg) and CUT it out of the timeline with ripple. Clips are split around silence; the silent middles are removed and everything after shifts earlier. Non-destructive: writes a _silence_removed copy. Preview with detect_media_silence first.",
inputSchema={
"type": "object",
"properties": {
"filepath": {"type": "string", "description": "Path to FCPXML file"},
"noise_db": {"type": "number", "description": "Silence threshold in dBFS, -120 to 0. Falls back to the saved silence settings (default -30)"},
"min_silence": {"type": "number", "description": "Minimum silence duration in seconds to cut. Falls back to the saved silence settings (default 0.5)"},
"padding": {"type": "number", "description": "Seconds of silence to keep on each side of a cut so edits breathe (max 5). Falls back to the saved silence settings (default 0.05)"},
"clip_name": {"type": "string", "description": "Only cut silence in the clip with this name"},
"output_path": {"type": "string", "description": "Output path (default: adds _silence_removed suffix)"},
},
"required": ["filepath"]
}
),
Tool(
name="detect_silence_candidates",
description="Detect potential silence/dead air using timeline heuristics (gaps, ultra-short clips, name patterns, duration anomalies)",
inputSchema={
"type": "object",
"properties": {
"filepath": {"type": "string", "description": "Path to FCPXML file"},
"min_gap_seconds": {"type": "number", "default": 0.5, "description": "Minimum gap duration to flag"},
"patterns": {"type": "array", "items": {"type": "string"}, "description": "Name patterns to match (default: gap, silence, room tone)"},
},
"required": ["filepath"]
}
),
Tool(
name="remove_silence_candidates",
description="Remove or mark detected silence candidates from timeline",
inputSchema={
"type": "object",
"properties": {
"filepath": {"type": "string", "description": "Path to FCPXML file"},
"mode": {"type": "string", "enum": ["delete", "mark"], "default": "mark", "description": "delete=remove clips/gaps, mark=add red markers"},
"min_gap_seconds": {"type": "number", "default": 0.5},
"min_confidence": {"type": "number", "default": 0.7, "description": "Only act on candidates above this confidence"},
"output_path": {"type": "string", "description": "Output path (default: adds _silence_cleaned suffix)"}
},
"required": ["filepath"]
}
),
]
async def handle_find_short_cuts(arguments: dict) -> Sequence[TextContent]:
project, tl = _require_timeline(arguments["filepath"])
threshold = arguments.get("threshold_seconds", 0.5)
short = tl.get_clips_shorter_than(threshold)
if not short:
return _text_result(f"No clips shorter than {threshold}s")
return _text_result(_format_clip_table(
short, f"# Short Clips (< {threshold}s) - {len(short)} found",
))
async def handle_find_long_clips(arguments: dict) -> Sequence[TextContent]:
project, tl = _require_timeline(arguments["filepath"])
threshold = arguments.get("threshold_seconds", 10.0)
long = tl.get_clips_longer_than(threshold)
if not long:
return _text_result(f"No clips longer than {threshold}s")
return _text_result(_format_clip_table(
long, f"# Long Clips (> {threshold}s) - {len(long)} found",
))
async def handle_analyze_pacing(arguments: dict) -> Sequence[TextContent]:
project, tl = _require_timeline(arguments["filepath"])
if not tl.clips:
return _text_result("No clips to analyze")
durs = [c.duration_seconds for c in tl.clips]
avg = sum(durs) / len(durs)
q_len = len(durs) // 4 or 1
segments = [durs[i:i+q_len] for i in range(0, len(durs), q_len)][:4]
seg_avgs = [sum(s)/len(s) if s else 0 for s in segments]
suggestions = []
flash = [c for c in tl.clips if c.duration_seconds < 0.2]
if flash:
suggestions.append(f" {len(flash)} potential flash frames (< 0.2s)")
long = [c for c in tl.clips if c.duration_seconds > 30]
if long:
suggestions.append(f" {len(long)} long takes (> 30s) - consider trimming")
if len(seg_avgs) >= 4 and seg_avgs[3] < seg_avgs[0] * 0.7:
suggestions.append(" Pacing accelerates toward end - good for building energy")
elif len(seg_avgs) >= 4 and seg_avgs[3] > seg_avgs[0] * 1.3:
suggestions.append(" Pacing slows toward end - consider tightening")
return _text_result(f"""# Pacing Analysis: {tl.name}
## Overall
- **Avg Cut**: {format_duration(avg)}
- **Cuts/Min**: {tl.cuts_per_minute:.1f}
## By Section
| Q1 | Q2 | Q3 | Q4 |
|----|----|----|----|
| {format_duration(seg_avgs[0]) if len(seg_avgs) > 0 else 'N/A'} | {format_duration(seg_avgs[1]) if len(seg_avgs) > 1 else 'N/A'} | {format_duration(seg_avgs[2]) if len(seg_avgs) > 2 else 'N/A'} | {format_duration(seg_avgs[3]) if len(seg_avgs) > 3 else 'N/A'} |
## Suggestions
{_fmt_suggestions(suggestions)}
""")
async def handle_detect_flash_frames(arguments: dict) -> Sequence[TextContent]:
project, tl = _require_timeline(arguments["filepath"])
critical_threshold = arguments.get("critical_threshold_frames", 2)
warning_threshold = arguments.get("warning_threshold_frames", 6)
flash_frames = _detect_flash_frames(
tl, critical_threshold=critical_threshold, warning_threshold=warning_threshold,
)
if not flash_frames:
return _text_result(f"No flash frames detected (threshold: {warning_threshold} frames)")
critical = [f for f in flash_frames if f.severity == FlashFrameSeverity.CRITICAL]
warnings = [f for f in flash_frames if f.severity == FlashFrameSeverity.WARNING]
result = f"""# Flash Frame Detection
## Summary
- **Critical** (< {critical_threshold} frames): {len(critical)} found
- **Warning** (< {warning_threshold} frames): {len(warnings)} found
- **Total**: {len(flash_frames)} flash frames
## Critical Flash Frames
"""
flash_headers = ["Clip", "Timecode", "Frames", "Duration"]
if critical:
result += _markdown_table(flash_headers, [
[f.clip_name, format_timecode(f.start), f"{f.duration_frames}f", format_duration(f.duration_seconds)]
for f in critical
]) + "\n"
else:
result += "_None_\n"
result += "\n## Warning Flash Frames\n"
if warnings:
result += _markdown_table(flash_headers, [
[f.clip_name, format_timecode(f.start), f"{f.duration_frames}f", format_duration(f.duration_seconds)]
for f in warnings
]) + "\n"
else:
result += "_None_\n"
result += "\n*Use `fix_flash_frames` to automatically resolve these issues.*"
return _text_result(result)
async def handle_detect_duplicates(arguments: dict) -> Sequence[TextContent]:
project, tl = _require_timeline(arguments["filepath"])
mode = arguments.get("mode", "same_source")
duplicates = _detect_duplicate_groups(tl, mode=mode)
if not duplicates:
return _text_result(f"No duplicate clips found (mode: {mode})")
result = f"""# Duplicate Clip Detection
## Summary
- **Mode**: {mode}
- **Duplicate Groups**: {len(duplicates)}
- **Total Duplicate Clips**: {sum(g.count for g in duplicates)}
## Duplicate Groups
"""
for group in duplicates:
result += f"\n### {group.source_name} ({group.count} uses)\n"
result += "| Clip Name | Timeline Position | Duration |\n|-----------|-------------------|----------|\n"
for c in group.clips:
result += f"| {c['name']} | {c['timecode']} | {format_duration(c['duration'])} |\n"
return _text_result(result)
async def handle_detect_gaps(arguments: dict) -> Sequence[TextContent]:
project, tl = _require_timeline(arguments["filepath"])
min_gap_frames = arguments.get("min_gap_frames", 1)
gaps = _detect_gaps(tl, min_gap_frames=min_gap_frames)
if not gaps:
return _text_result(f"No gaps detected (minimum: {min_gap_frames} frame(s))")
result = f"""# Gap Detection
## Summary
- **Gaps Found**: {len(gaps)}
- **Total Gap Duration**: {format_duration(sum(g.duration_seconds for g in gaps))}
- **Minimum Detection**: {min_gap_frames} frame(s)
## Gaps
"""
result += _markdown_table(
["Position", "Duration", "Between"],
[[gap.timecode, f"{gap.duration_frames}f ({format_duration(gap.duration_seconds)})",
f"{gap.previous_clip} -> {gap.next_clip}"] for gap in gaps],
) + "\n"
result += "\n*Use `fill_gaps` to automatically close these gaps.*"
return _text_result(result)
async def handle_validate_timeline(arguments: dict) -> Sequence[TextContent]:
project, tl = _require_timeline(arguments["filepath"])
checks = arguments.get("checks", ["all"])
run_all = "all" in checks
issues: list[str] = []
flash_count = 0
gap_count = 0
duplicate_count = 0
if run_all or "flash_frames" in checks:
flashes = _detect_flash_frames(tl)
flash_count = len(flashes)
for f in flashes:
severity = "error" if f.severity == FlashFrameSeverity.CRITICAL else "warning"
issues.append(
f"- [{severity.upper()}] Flash frame: {f.clip_name} "
f"({f.duration_frames}f) at {format_timecode(f.start)}"
)
if run_all or "gaps" in checks:
detected_gaps = _detect_gaps(tl)
gap_count = len(detected_gaps)
for g in detected_gaps:
issues.append(f"- [WARNING] Gap: {g.duration_frames}f at {g.timecode}")
if run_all or "duplicates" in checks:
dup_groups = _detect_duplicate_groups(tl)
for group in dup_groups:
duplicate_count += group.count
issues.append(
f"- [INFO] Duplicate source: {group.source_name} ({group.count} uses)"
)
error_weight = 10
warning_weight = 3
info_weight = 1
errors = len([i for i in issues if "[ERROR]" in i])
warnings = len([i for i in issues if "[WARNING]" in i])
infos = len([i for i in issues if "[INFO]" in i])
penalty = (errors * error_weight) + (warnings * warning_weight) + (infos * info_weight)
health_score = max(0, 100 - penalty)
result = f"""# Timeline Validation: {tl.name}
## Health Score: {health_score}%
## Summary
| Check | Count | Status |
|-------|-------|--------|
| Flash Frames | {flash_count} | {'PASS' if flash_count == 0 else 'FAIL'} |
| Gaps | {gap_count} | {'PASS' if gap_count == 0 else 'WARN'} |
| Duplicate Sources | {duplicate_count} | {'PASS' if duplicate_count == 0 else 'INFO'} |
## Issues ({len(issues)})
"""
if issues:
result += "\n".join(issues[:20])
if len(issues) > 20:
result += f"\n... and {len(issues) - 20} more issues"
else:
result += "_No issues found!_"
result += "\n\n*Use `fix_flash_frames` and `fill_gaps` to automatically resolve issues.*"
return _text_result(result)
async def handle_detect_media_silence(arguments: dict) -> Sequence[TextContent]:
# Unpassed thresholds come from the persisted silence settings (the app's
# own slider), not a hardcoded constant, so detection previews exactly
# what removal would cut.
saved = load_silence_config()
noise_db = float(arguments.get("noise_db", saved["noise_db"]))
min_silence = float(arguments.get("min_silence", saved["min_silence"]))
# Same bounds detect_silence() enforces — validated here so a bad request
# fails before any media file is opened.
if not (-120.0 <= noise_db <= 0.0):
raise ValueError(f"noise_db must be between -120 and 0 dB, got {noise_db}")
if not (0 < min_silence <= 3600):
raise ValueError(f"min_silence must be between 0 and 3600 seconds, got {min_silence}")
filepath = _validate_filepath(arguments["filepath"], ('.fcpxml', '.fcpxmld'))
modifier = FCPXMLModifier(filepath)
clip_filter = arguments.get("clip_name")
max_media_probes = 100
findings: list[tuple[str, float, float]] = []
skipped: list[tuple[str, str]] = []
probe_cache: dict[str, list | None] = {}
for el in [el for _, el in modifier._iter_spine_clips()]:
name = el.get("name", "")
if clip_filter and name != clip_filter:
continue
src = modifier.resources.get(el.get("ref", ""), {}).get("src", "")
media_path = media_src_to_path(src)
if not media_path or not Path(media_path).is_file():
skipped.append((name, "media file missing"))
continue
if media_path not in probe_cache:
if len(probe_cache) >= max_media_probes:
skipped.append((name, f"probe cap reached ({max_media_probes} media files)"))
continue
probe_cache[media_path] = detect_silence(
media_path, noise_db=noise_db, min_duration=min_silence
)
silences = probe_cache[media_path]
if silences is None:
skipped.append((name, "unanalyzable (ffmpeg missing or media unreadable)"))
continue
source_start = modifier.source_file_start(el).to_seconds()
clip_duration = modifier._parse_time(el.get("duration", "0s")).to_seconds()
timeline_offset = modifier._parse_time(el.get("offset", "0s")).to_seconds()
mapped = map_silence_to_timeline(
silences, source_start, clip_duration, timeline_offset
)
findings.extend((name, start, end) for start, end in mapped)
total_silence = sum(end - start for _, start, end in findings)
result = f"""# Media Silence Detection (real audio analysis)
## Summary
- **Threshold**: {noise_db} dB for >= {min_silence}s
- **Media Files Probed**: {len(probe_cache)}
- **Silence Spans Found**: {len(findings)} ({format_duration(total_silence)} total)
"""
if findings:
result += "\n## Silence Spans (timeline time)\n"
result += _markdown_table(
["Clip", "Start", "End", "Duration"],
[[name, f"{start:.2f}s", f"{end:.2f}s", f"{end - start:.2f}s"]
for name, start, end in findings],
) + "\n"
result += "\n*To remove: `split_clip` at each boundary, then `delete_clips` with ripple.*"
if skipped:
result += "\n## Skipped Clips\n"
result += _markdown_table(
["Clip", "Reason"], [[name, reason] for name, reason in skipped]
) + "\n"
if not findings and not skipped:
result += "\nNo silence detected in any clip's source audio."
return _text_result(result)
async def handle_detect_beats(arguments: dict) -> Sequence[TextContent]:
media_path = _validate_filepath(
arguments["media_path"], AUDIO_MEDIA_EXTENSIONS, max_size=MAX_MEDIA_FILE_SIZE
)
result = detect_beats(media_path)
if result is None:
return _text_result(
"Beat detection unavailable — librosa is not installed or the file "
"could not be analyzed.\n\nInstall the optional media-intelligence "
"extra:\n\n pip install 'fcp-mcp-server[intelligence]'"
)
bpm, beats = result["bpm"], result["beats"]
beats_data = {
"source": str(Path(media_path).name),
"bpm": round(bpm, 2),
"beats": [round(b, 4) for b in beats],
"downbeats": [round(b, 4) for b in beats[::4]],
}
json_path = _validate_output_path(
str(Path(media_path).with_name(Path(media_path).stem + "_beats.json")),
anchor_dir=str(Path(media_path).parent),
)
with open(json_path, "w") as f:
json.dump(beats_data, f, indent=2)
preview = beats[:16]
result_text = f"""# Beat Detection
## Summary
- **Source**: {Path(media_path).name}
- **Estimated Tempo**: {bpm:.1f} BPM
- **Beats Detected**: {len(beats)} ({format_duration(beats[-1]) if beats else '0s'} span)
- **Beats JSON**: {json_path}
## First Beats
"""
result_text += _markdown_table(
["#", "Time"],
[[str(i + 1), f"{b:.3f}s"] for i, b in enumerate(preview)],
) + "\n"
result_text += (
f"\n*Next: `import_beat_markers` with beats_path=\"{json_path}\" to place "
"markers, then `snap_to_beats` to align your cuts.*"
)
return _text_result(result_text)
async def handle_remove_media_silence(arguments: dict) -> Sequence[TextContent]:
saved = load_silence_config()
noise_db = float(arguments.get("noise_db", saved["noise_db"]))
min_silence = float(arguments.get("min_silence", saved["min_silence"]))
padding = float(arguments.get("padding", saved["padding"]))
if not (-120.0 <= noise_db <= 0.0):
raise ValueError(f"noise_db must be between -120 and 0 dB, got {noise_db}")
if not (0 < min_silence <= 3600):
raise ValueError(f"min_silence must be between 0 and 3600 seconds, got {min_silence}")
if not (0 <= padding <= 5):
raise ValueError(f"padding must be between 0 and 5 seconds, got {padding}")
filepath, output_path, modifier = _setup_modifier(arguments, "_silence_removed")
clip_filter = arguments.get("clip_name")
to_frame_timevalue = modifier.snap_seconds_to_frame
max_media_probes = 100
cuts_made: list[tuple[str, int, float]] = []
skipped: list[tuple[str, str]] = []
probe_cache: dict[str, list | None] = {}
spine_clips = [el for _, el in modifier._iter_spine_clips()]
for el in spine_clips:
name = el.get("name", "")
if clip_filter and name != clip_filter:
continue
src = modifier.resources.get(el.get("ref", ""), {}).get("src", "")
media_path = media_src_to_path(src)
if not media_path or not Path(media_path).is_file():
skipped.append((name, "media file missing"))
continue
if media_path not in probe_cache:
if len(probe_cache) >= max_media_probes:
skipped.append((name, f"probe cap reached ({max_media_probes} media files)"))
continue
probe_cache[media_path] = detect_silence(
media_path, noise_db=noise_db, min_duration=min_silence
)
silences = probe_cache[media_path]
if silences is None:
skipped.append((name, "unanalyzable (ffmpeg missing or media unreadable)"))
continue
clip_source_start = modifier.source_file_start(el).to_seconds()
clip_duration = modifier._parse_time(el.get("duration", "0s")).to_seconds()
cut_ranges = []
for sil_start, sil_end in silences:
# Source time -> clip-relative, padded so cuts breathe.
cut_start = max(sil_start, clip_source_start) - clip_source_start + padding
cut_end = min(sil_end, clip_source_start + clip_duration) - clip_source_start - padding
if cut_end > cut_start:
cut_ranges.append((to_frame_timevalue(cut_start), to_frame_timevalue(cut_end)))
if not cut_ranges:
continue
removed = modifier.cut_clip_ranges(el, cut_ranges)
if removed > TimeValue.zero():
cuts_made.append((name, len(cut_ranges), removed.to_seconds()))
if not cuts_made:
text = "# Media Silence Removal\n\nNo silence found to remove — file unchanged (nothing saved)."
if skipped:
text += "\n\n## Skipped Clips\n" + _markdown_table(
["Clip", "Reason"], [[name, reason] for name, reason in skipped]
)
return _text_result(text)
modifier.remove_trailing_gaps()
modifier.save(output_path)
total_removed = sum(seconds for _, _, seconds in cuts_made)
result = f"""# Media Silence Removal (real audio analysis)
## Summary
- **Threshold**: {noise_db} dB for >= {min_silence}s, padding {padding}s
- **Clips Cut**: {len(cuts_made)}
- **Total Removed**: {format_duration(total_removed)}
## Cuts
"""
result += _markdown_table(
["Clip", "Silence Spans Cut", "Removed"],
[[name, str(count), f"{seconds:.2f}s"] for name, count, seconds in cuts_made],
) + "\n"
if skipped:
result += "\n## Skipped Clips\n" + _markdown_table(
["Clip", "Reason"], [[name, reason] for name, reason in skipped]
) + "\n"
result += f"\nSaved to: {output_path}\n\n*Preview first next time with `detect_media_silence`. Original file untouched.*"
return _text_result(result)
async def handle_detect_silence_candidates(arguments: dict) -> Sequence[TextContent]:
filepath = _validate_filepath(arguments["filepath"], ('.fcpxml', '.fcpxmld'))
modifier = FCPXMLModifier(filepath)
candidates = modifier.detect_silence_candidates(
min_gap_seconds=arguments.get("min_gap_seconds", 0.5),
patterns=arguments.get("patterns"),
)
if not candidates:
return _text_result("No silence candidates detected.")
result = f"# Silence Candidates Detected\n\n**Found**: {len(candidates)}\n\n"
result += "| # | Timecode | Duration | Reason | Confidence | Clip |\n"
result += "|---|----------|----------|--------|------------|------|\n"
for i, c in enumerate(candidates, 1):
result += (
f"| {i} | {c['start_timecode']} | {format_duration(c['duration_seconds'])} | "
f"{c['reason']} | {c['confidence']:.0%} | {c.get('clip_name') or '-'} |\n"
)
result += (
"\n**Note**: Detection uses timeline heuristics (gaps, ultra-short clips, name patterns). "
"Review candidates before removing — some may be intentional."
)
return _text_result(result)
async def handle_remove_silence_candidates(arguments: dict) -> Sequence[TextContent]:
filepath, output_path, modifier = _setup_modifier(arguments, "_silence_cleaned")
actions = modifier.remove_silence_candidates(
mode=arguments.get("mode", "mark"),
min_gap_seconds=arguments.get("min_gap_seconds", 0.5),
min_confidence=arguments.get("min_confidence", 0.7),
)
modifier.save(output_path)
if not actions:
return _text_result("No silence candidates met the confidence threshold.")
mode = arguments.get("mode", "mark")
result = f"# Silence Candidates {'Marked' if mode == 'mark' else 'Removed'}\n\n"
result += f"**Actions taken**: {len(actions)}\n\n"
for a in actions:
result += f"- **{a['action']}** {a.get('clip_name', 'gap')} ({a['reason']})\n"
result += f"\nSaved to: `{output_path}`"
return _text_result(result)
HANDLERS = {
"find_short_cuts": handle_find_short_cuts,
"find_long_clips": handle_find_long_clips,
"analyze_pacing": handle_analyze_pacing,
"detect_flash_frames": handle_detect_flash_frames,
"detect_duplicates": handle_detect_duplicates,
"detect_gaps": handle_detect_gaps,
"validate_timeline": handle_validate_timeline,
"detect_media_silence": handle_detect_media_silence,
"detect_beats": handle_detect_beats,
"remove_media_silence": handle_remove_media_silence,
"detect_silence_candidates": handle_detect_silence_candidates,
"remove_silence_candidates": handle_remove_silence_candidates,
}
+133
View File
@@ -0,0 +1,133 @@
"""Roles — tool schemas and handlers.
Extracted from server.py; see Engine/docs/03_SERVER_TOOLS.md for the tool catalog.
"""
from __future__ import annotations
from typing import Sequence
from mcp.types import TextContent, Tool
from server_tools._shared import (
_require_timeline,
_setup_modifier,
_text_result,
format_duration,
)
TOOLS = [
Tool(
name="assign_role",
description="Set the audio or video role on a clip (dialogue, music, effects, titles, etc.)",
inputSchema={
"type": "object",
"properties": {
"filepath": {"type": "string", "description": "Path to FCPXML file"},
"clip_id": {"type": "string", "description": "Clip name or ID"},
"audio_role": {"type": "string", "description": "Audio role (e.g., dialogue, music, effects)"},
"video_role": {"type": "string", "description": "Video role (e.g., video, titles)"},
"output_path": {"type": "string", "description": "Output path (default: adds _modified suffix)"}
},
"required": ["filepath", "clip_id"]
}
),
Tool(
name="filter_by_role",
description="List all clips matching a specific audio or video role",
inputSchema={
"type": "object",
"properties": {
"filepath": {"type": "string", "description": "Path to FCPXML file"},
"role": {"type": "string", "description": "Role name to filter by"},
"role_type": {"type": "string", "enum": ["audio", "video", "any"], "default": "any", "description": "Which role type to search"},
},
"required": ["filepath", "role"]
}
),
Tool(
name="export_role_stems",
description="Export clip list grouped by role for audio mixing stem planning",
inputSchema={
"type": "object",
"properties": {
"filepath": {"type": "string", "description": "Path to FCPXML file"},
},
"required": ["filepath"]
}
),
]
async def handle_assign_role(arguments: dict) -> Sequence[TextContent]:
filepath, output_path, modifier = _setup_modifier(arguments)
modifier.assign_role(
clip_id=arguments["clip_id"],
audio_role=arguments.get("audio_role"),
video_role=arguments.get("video_role"),
)
modifier.save(output_path)
roles_set = []
if arguments.get("audio_role"):
roles_set.append(f"audioRole={arguments['audio_role']}")
if arguments.get("video_role"):
roles_set.append(f"videoRole={arguments['video_role']}")
return _text_result((
f"Set {', '.join(roles_set)} on '{arguments['clip_id']}'\n\n"
f"Saved to: `{output_path}`"
))
async def handle_filter_by_role(arguments: dict) -> Sequence[TextContent]:
project, tl = _require_timeline(arguments["filepath"])
role = arguments["role"].lower()
role_type = arguments.get("role_type", "any")
matches = []
for clip in tl.clips:
if role_type in ("audio", "any") and clip.audio_role.lower() == role:
matches.append((clip.name, "audio", clip.audio_role, format_duration(clip.duration_seconds)))
if role_type in ("video", "any") and clip.video_role.lower() == role:
matches.append((clip.name, "video", clip.video_role, format_duration(clip.duration_seconds)))
if not matches:
return _text_result(f"No clips found with role '{role}'.")
result = f"# Clips with role '{role}'\n\n"
result += "| Clip | Type | Role | Duration |\n|------|------|------|----------|\n"
for name, rtype, rval, dur in matches:
result += f"| {name} | {rtype} | {rval} | {dur} |\n"
return _text_result(result)
async def handle_export_role_stems(arguments: dict) -> Sequence[TextContent]:
project, tl = _require_timeline(arguments["filepath"])
stems: dict[str, list] = {}
for clip in tl.clips:
role = clip.audio_role or "unassigned"
stems.setdefault(role, []).append(clip)
for cc in tl.connected_clips:
role = cc.role or "unassigned"
stems.setdefault(role, []).append(cc)
result = f"# Audio Stem Plan for {tl.name}\n\n"
for role, clips in sorted(stems.items()):
total_dur = sum(c.duration_seconds for c in clips)
result += f"## {role.title()} ({len(clips)} clips, {format_duration(total_dur)})\n\n"
for c in clips:
result += f"- {c.name} ({format_duration(c.duration_seconds)})\n"
result += "\n"
return _text_result(result)
HANDLERS = {
"assign_role": handle_assign_role,
"filter_by_role": handle_filter_by_role,
"export_role_stems": handle_export_role_stems,
}
+283
View File
@@ -0,0 +1,283 @@
"""Legendas dinâmicas (geração → validação) — tool schemas and handlers.
Extracted from server.py; see Engine/docs/03_SERVER_TOOLS.md for the tool catalog.
"""
from __future__ import annotations
import json
from pathlib import Path
from typing import Sequence
from mcp.types import TextContent, Tool
from fcpxml.media_intel import media_src_to_path
from fcpxml.model_manager import load_dynamic_subtitle_config
from fcpxml.models import DynamicSubtitleConfig, WordLook, WordStyle
from fcpxml.writer import FCPXMLModifier
from server_tools._shared import (
_load_or_transcribe,
_markdown_table,
_setup_modifier,
_text_result,
_validate_filepath,
)
TOOLS = [
Tool(
name="validate_subtitle_layout",
description="Re-measure every title/subtitle in an FCPXML and report spatial collisions, frame and safe-area violations, and font fallbacks. Detects overlapping boxes only for titles on screen at the same time (half-open time intervals, so a title ending exactly as the next begins is never flagged). Returns a severity (none/render_tolerance/warning/probable/severe), the list of issues with suggested corrections, and summary counts.",
inputSchema={
"type": "object",
"properties": {
"filepath": {"type": "string", "description": "Path to FCPXML file"},
"safe_margin_x": {"type": "number", "default": 0.05, "description": "Fraction of frame width to inset from each side (0.05 = 5%)"},
"safe_margin_y": {"type": "number", "default": 0.05, "description": "Fraction of frame height to inset from top/bottom"},
"min_font_size": {"type": "number", "description": "Flag titles whose emitted fontSize is below this readable minimum"},
"min_distance": {"type": "number", "description": "Flag same-block titles closer than this many pixels (insufficient_spacing)"},
"max_distance": {"type": "number", "description": "Flag same-block titles farther than this many pixels (excessive_spacing)"},
"output_format": {"type": "string", "enum": ["markdown", "json"], "default": "markdown", "description": "Report format"}
},
"required": ["filepath"]
}
),
Tool(
name="generate_dynamic_subtitles",
description="Generate progressive-composition subtitles as real, editable FCPXML title clips (the 'Text'/Basic Text template). Whisper's segments become sentences; each sentence is diagrammed as stacked blocks — supporting words grouped small in a grotesque, the sentence's key word alone and large in a display italic, body lines staggered to opposite edges. One <title> per block: each enters as its own words are spoken and stays on screen, so the sentence assembles itself, and every block clears at the same instant. Set granularity='word' for the older one-title-per-word rhythm. A sentence too tall for the band splits into successive compositions. These are TITLES, not captions: no subtitles role, so they render over the video without enabling caption display. Uses each media file's local Whisper word-level transcript (_transcript.json, auto-transcribes if missing). Style fields below (band_height through inactive_color) fall back to the style saved from the app's 'Legendas Dinâmicas' screen (~/.fcp-mcp-server/config.json via save_dynamic_subtitle_config) when omitted — pass a value here only to override that for one call. Non-destructive: writes a _dynamic_subtitles copy.",
inputSchema={
"type": "object",
"properties": {
"filepath": {"type": "string", "description": "Path to FCPXML file"},
"clip_name": {"type": "string", "description": "Only caption the clip with this name (default: all spine clips with matched source media)"},
"model": {"type": "string", "default": "base", "description": "Whisper model size if transcription is needed"},
"language": {"type": "string", "description": "ISO language code hint (e.g. 'en'); auto-detected if omitted"},
"band_height": {"type": "number", "description": "Fraction of frame height the sentence block may fill before splitting into another block. Falls back to the saved style (default 0.22 — about three lines)"},
"block_center_y": {"type": "number", "description": "Vertical centre of the block in canvas points; negative sits below frame centre. Falls back to the saved style (default -167, just under centre)"},
"line_gap": {"type": "number", "description": "Air between stacked lines in canvas points. Lines are stacked on their real ink, so this is the whole distance beyond the glyphs themselves; negative values deliberately tuck each line into the one above. Falls back to the saved style (default 8)"},
"granularity": {"type": "string", "enum": ["phrase", "word"], "default": "phrase", "description": "'phrase': one title per LINE of the composition, key word set large (the reference look). 'word': one title per word."},
"emphasis_font": {"type": "string", "description": "Family for the key word (phrase mode). Must be installed on the editing Mac; unmeasured families fall back to estimated widths. Falls back to the saved style (default 'Playfair Display')"},
"emphasis_face": {"type": "string", "description": "Face for the key word, e.g. 'Medium Italic' or a script/calligraphic face. Falls back to the saved style (default 'Medium Italic')"},
"emphasis_size": {"type": "integer", "description": "Key-word size in canvas points, at the 2160x3840 reference frame. Falls back to the saved style (default 265)"},
"emphasis_color": {"type": "string", "description": "RGBA (0-1, space-separated) for the key word (phrase mode). Defaults to active_color, so the block reads in a single colour unless the key word is deliberately set apart"},
"text_scale": {"type": "number", "description": "Ratio between the title template's fontSize space and the canvas-point space it positions in. The \"Text\" template sizes type in frame pixels, so sizes are doubled on the way out. Falls back to the saved style (default 2.0). Lower it only if a template renders type larger than the chosen point size"},
"font": {"type": "string", "description": "Title font family (supporting lines in phrase mode). Falls back to the saved style (default 'Helvetica Neue')"},
"font_size": {"type": "integer", "description": "Supporting-line font size in canvas points, at the 2160x3840 reference frame. Falls back to the saved style (default 104)"},
"active_color": {"type": "string", "description": "RGBA (0-1, space-separated) for even-indexed lines. Falls back to the saved style (default '1 1 1 1')"},
"inactive_color": {"type": "string", "default": "0.7 0.7 0.7 1", "description": "RGBA (0-1, space-separated) for odd-indexed lines — alternates with active_color for visual variety between stacked lines"},
"output_path": {"type": "string", "description": "Output path (default: adds _dynamic_subtitles suffix)"},
},
"required": ["filepath"]
}
),
]
async def handle_validate_subtitle_layout(arguments: dict) -> Sequence[TextContent]:
"""Validate title/subtitle layout for spatial collisions and safe-area
containment (collision.validate_titles over every <title> in the file)."""
filepath = _validate_filepath(arguments["filepath"], (".fcpxml", ".fcpxmld"))
modifier = FCPXMLModifier(filepath)
report = modifier.validate_subtitle_layout(
safe_margin_x=float(arguments.get("safe_margin_x", 0.05)),
safe_margin_y=float(arguments.get("safe_margin_y", 0.05)),
min_font_size=(
float(arguments["min_font_size"])
if arguments.get("min_font_size") is not None else None
),
min_distance=(
float(arguments["min_distance"])
if arguments.get("min_distance") is not None else None
),
max_distance=(
float(arguments["max_distance"])
if arguments.get("max_distance") is not None else None
),
)
if arguments.get("output_format") == "json":
return _text_result(json.dumps(report, indent=2))
summary = report["summary"]
lines = [
"# Subtitle Layout Validation",
"",
f"## Summary (severity: {report['severity']})",
f"- **Titles**: {summary['title_count']}",
f"- **Issues**: {summary['issue_count']}",
f"- **Collisions**: {summary['spatial_collision']}",
f"- **Outside frame**: {summary['outside_frame']}",
f"- **Outside safe area**: {summary['outside_safe_area']}",
f"- **Font fallback**: {summary['font_missing']}",
f"- **Font too small**: {summary['font_too_small']}",
"",
]
issues = report["issues"]
if issues:
lines.append(f"## Issues ({len(issues)})")
for issue in issues:
sev = issue["severity"].upper()
if issue["type"] == "spatial_collision":
corr = issue["suggested_correction"]
lines.append(
f"- [{sev}] collision: \"{issue['first_title']}\" x "
f"\"{issue['second_title']}\" "
f"(overlap {issue['overlap_width']:.0f}x"
f"{issue['overlap_height']:.0f} = "
f"{issue['overlap_area']:.0f}px, ratio "
f"{issue['overlap_ratio']:.2f}, move "
f"{corr['axis']} {corr['minimum_movement']:.0f}px)"
)
else:
detail = issue.get("title", "") or issue.get("font", "")
lines.append(f"- [{sev}] {issue['type']}: {detail}".rstrip())
else:
lines.append("_No issues found — no simultaneous titles overlap._")
return _text_result("\n".join(lines))
async def handle_generate_dynamic_subtitles(arguments: dict) -> Sequence[TextContent]:
"""Generate per-word subtitle titles laid out as a block per sentence.
Whisper's segments become sentences; each word becomes its own positioned
<title> connected clip, appearing as it is spoken and accumulating on
screen until the whole block clears at once. No compound clip.
Reuses the same SOURCE-media -> TIMELINE mapping as ``transcript_markers``
(``modifier.source_file_start`` per spine clip) so word timestamps land
at the correct position even across trimmed/multiple clips.
"""
model = arguments.get("model", "base")
language = arguments.get("language")
output_dir = arguments.get("output_dir")
clip_filter = arguments.get("clip_name")
# Anything the caller didn't explicitly pass falls back to the style
# persisted from the "Legendas Dinâmicas" screen (~/.fcp-mcp-server/
# config.json), not a hardcoded default — so the UI is the single place
# that configures the look, and every caller (app, MCP, this session)
# renders the same thing without threading 11 fields through every call.
saved = load_dynamic_subtitle_config()
body_color = arguments.get("active_color") or saved["active_color"]
config = DynamicSubtitleConfig(
style=WordStyle(
font=arguments.get("font") or saved["font"],
font_size=int(arguments.get("font_size", saved["font_size"])),
active_color=body_color,
inactive_color=arguments.get("inactive_color", "0.7 0.7 0.7 1"),
emphasis_look=WordLook(
int(arguments.get("emphasis_size", saved["emphasis_size"])),
arguments.get("emphasis_color") or saved["emphasis_color"] or body_color,
font=arguments.get("emphasis_font") or saved["emphasis_font"],
face=arguments.get("emphasis_face") or saved["emphasis_face"],
kerning=0.0,
),
body_look=WordLook(
int(arguments.get("font_size", saved["font_size"])),
body_color,
font=arguments.get("font") or saved["font"],
face="Bold",
kerning=1.2,
),
),
band_height=float(arguments.get("band_height", saved["band_height"])),
block_center_y=float(arguments.get("block_center_y", saved["block_center_y"])),
granularity=arguments.get("granularity", "phrase"),
text_scale=float(arguments.get("text_scale", saved["text_scale"])),
line_gap=float(arguments.get("line_gap", saved["line_gap"])),
)
filepath, output_path, modifier = _setup_modifier(arguments, "_dynamic_subtitles")
added: list[tuple[str, int, int]] = []
skipped: list[tuple[str, str]] = []
spine_clips = [el for _, el in modifier._iter_spine_clips()]
for el in spine_clips:
name = el.get("name", "")
if clip_filter and name != clip_filter:
continue
src = modifier.resources.get(el.get("ref", ""), {}).get("src", "")
media_path = media_src_to_path(src)
if not media_path or not Path(media_path).is_file():
skipped.append((name, "media file missing"))
continue
data, reason = _load_or_transcribe(media_path, model, language, output_dir)
if data is None:
skipped.append((name, reason))
continue
clip_source_start = modifier.source_file_start(el).to_seconds()
clip_duration = modifier._parse_time(el.get("duration", "0s")).to_seconds()
window_end = clip_source_start + clip_duration
clip_words = [
{
"word": w.get("word", ""),
"start": float(w.get("start", 0.0)) - clip_source_start,
"end": float(w.get("end", 0.0)) - clip_source_start,
}
for w in data.get("words", [])
if clip_source_start <= float(w.get("start", 0.0)) < window_end
]
if not clip_words:
skipped.append((name, "no words in clip's source range"))
continue
# Sentence boundaries, rebased the same way, so each sentence becomes
# its own block of titles that builds up and then clears together.
# Overlap rather than containment: a segment straddling the clip's
# in-point still governs the words that made the cut.
clip_segments = [
{
"start": float(s.get("start", 0.0)) - clip_source_start,
"end": float(s.get("end", 0.0)) - clip_source_start,
}
for s in data.get("segments", [])
if float(s.get("end", 0.0)) > clip_source_start
and float(s.get("start", 0.0)) < window_end
]
# Pass the element itself, not `name` — after ripple-cut/silence
# removal every fragment of an originally-named clip keeps the same
# `name`, so a name lookup here would resolve every clip in this
# loop to whichever one `self.clips` last indexed, stacking every
# clip's captions onto a single wrong spine element instead of each
# clip's own. See Engine/docs/05_EXPERIENCIAS.md, entry 2026-08-17.
lines = modifier.generate_dynamic_subtitles(
el, clip_words, config, segments=clip_segments
)
added.append((name, len(lines), len(clip_words)))
if not added:
text = "# Dynamic Subtitles\n\nNo captions generated — file unchanged (nothing saved)."
if skipped:
text += "\n\n## Skipped Clips\n" + _markdown_table(
["Clip", "Reason"], [[n, r] for n, r in skipped]
)
return _text_result(text)
modifier.save(output_path)
total_lines = sum(lines for _, lines, _ in added)
total_words = sum(words for _, _, words in added)
result = "# Dynamic Subtitles Generated (local Whisper)\n\n## Summary\n"
result += (
f"- **Clips Captioned**: {len(added)}\n"
f"- **Caption Lines (Title Clips)**: {total_lines}\n"
f"- **Total Words**: {total_words}\n\n"
)
result += _markdown_table(
["Clip", "Caption Lines", "Words"],
[[n, str(lines), str(words)] for n, lines, words in added],
)
if skipped:
result += "\n## Skipped Clips\n" + _markdown_table(
["Clip", "Reason"], [[n, r] for n, r in skipped]
)
result += f"\n\nSaved to: `{output_path}`\n\n*Transcripts are cached as _transcript.json.*"
return _text_result(result)
HANDLERS = {
"validate_subtitle_layout": handle_validate_subtitle_layout,
"generate_dynamic_subtitles": handle_generate_dynamic_subtitles,
}
+400
View File
@@ -0,0 +1,400 @@
"""Timeline & análise (Projeto) — tool schemas and handlers.
Extracted from server.py; see Engine/docs/03_SERVER_TOOLS.md for the tool catalog.
"""
from __future__ import annotations
from typing import Sequence
from mcp.types import TextContent, Tool
from fcpxml.diff import compare_timelines
from fcpxml.models import MarkerType
from fcpxml.parser import FCPXMLParser
from fcpxml.writer import list_effects
from server_tools._shared import (
_SANDBOX_ENABLED,
PROJECTS_DIR,
_require_timeline,
_text_result,
_validate_directory,
_validate_filepath,
find_fcpxml_files,
format_duration,
format_timecode,
)
TOOLS = [
Tool(
name="list_projects",
description="List all FCPXML projects in directory",
inputSchema={
"type": "object",
"properties": {
"directory": {"type": "string", "description": "Directory to search (default: ~/Movies)"}
}
}
),
Tool(
name="analyze_timeline",
description="Get comprehensive timeline statistics including duration, resolution, clip count, pacing metrics",
inputSchema={
"type": "object",
"properties": {"filepath": {"type": "string", "description": "Path to FCPXML file"}},
"required": ["filepath"]
}
),
Tool(
name="list_clips",
description="List all clips with timecodes, durations, and metadata",
inputSchema={
"type": "object",
"properties": {
"filepath": {"type": "string"},
"limit": {"type": "integer", "description": "Max clips to return"}
},
"required": ["filepath"]
}
),
Tool(
name="list_markers",
description="Extract markers (chapter, todo, standard) with timestamps",
inputSchema={
"type": "object",
"properties": {
"filepath": {"type": "string"},
"marker_type": {"type": "string", "enum": ["all", "chapter", "todo", "standard", "completed"]},
"format": {"type": "string", "enum": ["detailed", "youtube", "simple"]}
},
"required": ["filepath"]
}
),
Tool(
name="list_keywords",
description="Extract all keywords/tags from project",
inputSchema={
"type": "object",
"properties": {"filepath": {"type": "string"}},
"required": ["filepath"]
}
),
Tool(
name="list_library_clips",
description="List all available clips in the library (source media, not yet on timeline)",
inputSchema={
"type": "object",
"properties": {
"filepath": {"type": "string", "description": "Path to FCPXML file"},
"keywords": {"type": "array", "items": {"type": "string"}, "description": "Filter by keywords"},
"limit": {"type": "integer", "description": "Max clips to return"}
},
"required": ["filepath"]
}
),
Tool(
name="list_connected_clips",
description="List all connected clips (B-roll, titles, audio) with their lanes and parent clips",
inputSchema={
"type": "object",
"properties": {
"filepath": {"type": "string", "description": "Path to FCPXML file"},
"lane": {"type": "integer", "description": "Filter by lane number (positive=above, negative=below)"},
},
"required": ["filepath"]
}
),
Tool(
name="list_compound_clips",
description="List compound clips (ref-clips) and their nested content",
inputSchema={
"type": "object",
"properties": {
"filepath": {"type": "string", "description": "Path to FCPXML file"},
},
"required": ["filepath"]
}
),
Tool(
name="list_roles",
description="List all audio/video roles used in the timeline with clip counts",
inputSchema={
"type": "object",
"properties": {
"filepath": {"type": "string", "description": "Path to FCPXML file"},
},
"required": ["filepath"]
}
),
Tool(
name="diff_timelines",
description="Compare two FCPXML files and report differences in clips, markers, transitions, and format",
inputSchema={
"type": "object",
"properties": {
"filepath_a": {"type": "string", "description": "Path to first FCPXML file (baseline)"},
"filepath_b": {"type": "string", "description": "Path to second FCPXML file (comparison)"},
},
"required": ["filepath_a", "filepath_b"]
}
),
Tool(
name="list_effects",
description="List all available FCP transition effects with slugs and UUIDs",
inputSchema={
"type": "object",
"properties": {},
}
),
]
async def handle_list_projects(arguments: dict) -> Sequence[TextContent]:
directory = arguments.get("directory", PROJECTS_DIR)
resolved_dir = _validate_directory(
directory, allowed_root=PROJECTS_DIR if _SANDBOX_ENABLED else None
)
files = find_fcpxml_files(resolved_dir)
if not files:
return _text_result(f"No FCPXML files found in {directory}")
return _text_result(f"Found {len(files)} FCPXML file(s):\n" + "\n".join(f" - {f}" for f in files))
async def handle_analyze_timeline(arguments: dict) -> Sequence[TextContent]:
project, tl = _require_timeline(arguments["filepath"])
durs = [c.duration_seconds for c in tl.clips]
avg, med, mn, mx = (0, 0, 0, 0) if not durs else (
sum(durs)/len(durs), sorted(durs)[len(durs)//2], min(durs), max(durs))
return _text_result(f"""# Timeline Analysis: {tl.name}
## Overview
- **Duration**: {format_duration(tl.duration.seconds)}
- **Resolution**: {tl.width}x{tl.height} @ {tl.frame_rate}fps
## Clip Statistics
- **Total Clips**: {tl.total_clips}
- **Total Cuts**: {tl.total_cuts}
- **Transitions**: {len(tl.transitions)}
## Pacing
- **Average**: {format_duration(avg)}
- **Median**: {format_duration(med)}
- **Shortest**: {format_duration(mn)}
- **Longest**: {format_duration(mx)}
- **Cuts/Minute**: {tl.cuts_per_minute:.1f}
## Markers
- **Total**: {len(tl.markers)}
- **Chapters**: {len([m for m in tl.markers if m.marker_type == MarkerType.CHAPTER])}
""")
async def handle_list_clips(arguments: dict) -> Sequence[TextContent]:
project, tl = _require_timeline(arguments["filepath"])
limit = arguments.get("limit")
clips = tl.clips[:limit] if limit else tl.clips
result = f"# Clips in {tl.name}\n\n| # | Name | Start | Duration | Keywords |\n|---|------|-------|----------|----------|\n"
for i, c in enumerate(clips, 1):
kws = ", ".join(k.value for k in c.keywords) if c.keywords else "-"
result += f"| {i} | {c.name} | {format_timecode(c.start)} | {format_duration(c.duration_seconds)} | {kws} |\n"
return _text_result(result)
async def handle_list_markers(arguments: dict) -> Sequence[TextContent]:
project, tl = _require_timeline(arguments["filepath"])
markers = list(tl.markers)
for clip in tl.clips:
markers.extend(clip.markers)
marker_type = arguments.get("marker_type", "all")
if marker_type != "all":
markers = [m for m in markers if m.marker_type == MarkerType.from_string(marker_type)]
markers.sort(key=lambda m: m.start.frames)
fmt = arguments.get("format", "detailed")
if fmt == "youtube":
result = "# YouTube Chapters\n\n" + "\n".join(f"{m.to_youtube_timestamp()} {m.name}" for m in markers)
elif fmt == "simple":
result = "\n".join(f"{format_timecode(m.start)} - {m.name}" for m in markers)
else:
result = f"# Markers ({len(markers)})\n\n| TC | Name | Type |\n|---|------|------|\n"
result += "\n".join(f"| {format_timecode(m.start)} | {m.name} | {m.marker_type.value} |" for m in markers)
return _text_result(result)
async def handle_list_keywords(arguments: dict) -> Sequence[TextContent]:
project, tl = _require_timeline(arguments["filepath"])
keywords = {}
for clip in tl.clips:
for kw in clip.keywords:
keywords.setdefault(kw.value, []).append(clip.name)
if not keywords:
return _text_result("No keywords found")
result = f"# Keywords ({len(keywords)})\n\n"
for kw, clips in sorted(keywords.items()):
result += f"**{kw}** ({len(clips)} clips)\n"
return _text_result(result)
async def handle_list_library_clips(arguments: dict) -> Sequence[TextContent]:
filepath = _validate_filepath(arguments["filepath"], ('.fcpxml', '.fcpxmld'))
parser = FCPXMLParser()
parser.parse_file(filepath)
keywords = arguments.get("keywords")
library_clips = parser.get_library_clips(keywords=keywords)
limit = arguments.get("limit")
if limit:
library_clips = library_clips[:limit]
if not library_clips:
return _text_result("No library clips found")
result = f"# Library Clips ({len(library_clips)} available)\n\n"
result += "| ID | Name | Duration | Has Video | Has Audio |\n"
result += "|----|------|----------|-----------|----------|\n"
for c in library_clips:
result += f"| {c['asset_id']} | {c['name']} | {format_duration(c['duration_seconds'])} | {'Y' if c['has_video'] else 'N'} | {'Y' if c['has_audio'] else 'N'} |\n"
result += "\n*Use `insert_clip` to add these to your timeline.*"
return _text_result(result)
async def handle_list_connected_clips(arguments: dict) -> Sequence[TextContent]:
project, tl = _require_timeline(arguments["filepath"])
lane_filter = arguments.get("lane")
clips = tl.connected_clips
if lane_filter is not None:
clips = [c for c in clips if c.lane == lane_filter]
if not clips:
return _text_result("No connected clips found in timeline.")
result = f"# Connected Clips in {tl.name}\n\n**Total**: {len(clips)}\n\n"
result += "| # | Name | Lane | Type | Duration | Parent | Role |\n"
result += "|---|------|------|------|----------|--------|------|\n"
for i, c in enumerate(clips, 1):
result += (
f"| {i} | {c.name} | {c.lane} | {c.clip_type} | "
f"{format_duration(c.duration_seconds)} | {c.parent_clip_name} | "
f"{c.role or '-'} |\n"
)
return _text_result(result)
async def handle_list_compound_clips(arguments: dict) -> Sequence[TextContent]:
project, tl = _require_timeline(arguments["filepath"])
if not tl.compound_clips:
return _text_result("No compound clips found in timeline.")
result = f"# Compound Clips in {tl.name}\n\n"
for i, cc in enumerate(tl.compound_clips, 1):
result += f"### {i}. {cc.name}\n"
result += f"- **Ref ID**: {cc.ref_id}\n"
result += f"- **Duration**: {format_duration(cc.duration_seconds)}\n"
result += f"- **Clips inside**: {len(cc.clips)}\n\n"
return _text_result(result)
async def handle_list_roles(arguments: dict) -> Sequence[TextContent]:
project, tl = _require_timeline(arguments["filepath"])
audio_roles: dict[str, int] = {}
video_roles: dict[str, int] = {}
for clip in tl.clips:
if clip.audio_role:
audio_roles[clip.audio_role] = audio_roles.get(clip.audio_role, 0) + 1
if clip.video_role:
video_roles[clip.video_role] = video_roles.get(clip.video_role, 0) + 1
for cc in tl.connected_clips:
if cc.role:
# Determine type from clip_type
if cc.clip_type in ('audio', 'audio-clip'):
audio_roles[cc.role] = audio_roles.get(cc.role, 0) + 1
else:
video_roles[cc.role] = video_roles.get(cc.role, 0) + 1
result = f"# Roles in {tl.name}\n\n"
if audio_roles:
result += "## Audio Roles\n\n| Role | Clips |\n|------|-------|\n"
for role, count in sorted(audio_roles.items()):
result += f"| {role} | {count} |\n"
else:
result += "## Audio Roles\n\nNo audio roles assigned.\n"
result += "\n"
if video_roles:
result += "## Video Roles\n\n| Role | Clips |\n|------|-------|\n"
for role, count in sorted(video_roles.items()):
result += f"| {role} | {count} |\n"
else:
result += "## Video Roles\n\nNo video roles assigned.\n"
return _text_result(result)
async def handle_diff_timelines(arguments: dict) -> Sequence[TextContent]:
filepath_a = _validate_filepath(arguments["filepath_a"], ('.fcpxml', '.fcpxmld'))
filepath_b = _validate_filepath(arguments["filepath_b"], ('.fcpxml', '.fcpxmld'))
diff = compare_timelines(filepath_a, filepath_b)
if not diff.has_changes:
return _text_result((
f"# Timeline Diff: No Changes\n\n"
f"**{diff.timeline_a_name}** vs **{diff.timeline_b_name}** are identical."
))
result = (
f"# Timeline Diff\n\n"
f"**Baseline**: {diff.timeline_a_name}\n"
f"**Comparison**: {diff.timeline_b_name}\n"
f"**Total changes**: {diff.total_changes}\n\n"
)
if diff.format_changes:
result += "## Format Changes\n\n"
for change in diff.format_changes:
result += f"- {change}\n"
result += "\n"
clip_changes = [d for d in diff.clip_diffs if d.action != "unchanged"]
if clip_changes:
result += "## Clip Changes\n\n| Action | Clip | Details |\n|--------|------|--------|\n"
for d in clip_changes:
result += f"| {d.action.upper()} | {d.clip_name} | {d.details} |\n"
result += "\n"
if diff.marker_diffs:
result += "## Marker Changes\n\n| Action | Marker | Details |\n|--------|--------|--------|\n"
for d in diff.marker_diffs:
result += f"| {d.action.upper()} | {d.marker_name} | {d.details} |\n"
result += "\n"
if diff.transition_diffs:
result += "## Transition Changes\n\n"
for change in diff.transition_diffs:
result += f"- {change}\n"
return _text_result(result)
async def handle_list_effects(arguments: dict) -> Sequence[TextContent]:
effects = list_effects()
lines = ["# Available FCP Transition Effects\n"]
for eff in effects:
lines.append(f"- **{eff['slug']}**: {eff['name']} (`{eff['uuid']}`)")
return _text_result("\n".join(lines))
HANDLERS = {
"list_projects": handle_list_projects,
"analyze_timeline": handle_analyze_timeline,
"list_clips": handle_list_clips,
"list_markers": handle_list_markers,
"list_keywords": handle_list_keywords,
"list_library_clips": handle_list_library_clips,
"list_connected_clips": handle_list_connected_clips,
"list_compound_clips": handle_list_compound_clips,
"list_roles": handle_list_roles,
"diff_timelines": handle_diff_timelines,
"list_effects": handle_list_effects,
}
+235
View File
@@ -0,0 +1,235 @@
"""Transcrição & edição por transcrição — tool schemas and handlers.
Extracted from server.py; see Engine/docs/03_SERVER_TOOLS.md for the tool catalog.
"""
from __future__ import annotations
from pathlib import Path
from typing import Sequence
from mcp.types import TextContent, Tool
from fcpxml.media_intel import media_src_to_path
from fcpxml.transcribe import (
DEFAULT_FILLERS,
find_filler_spans,
find_phrase_spans,
merge_ranges,
segments_to_srt,
)
from server_tools._shared import (
_TRANSCRIBE_INSTALL_HINT,
TRANSCRIBE_MAX_MEDIA,
_cut_transcript_spans,
_load_or_transcribe,
_markdown_table,
_require_timeline,
_setup_modifier,
_text_result,
_transcript_cut_report,
_validate_output_path,
format_duration,
)
TOOLS = [
Tool(
name="transcribe_media",
description="Transcribe each clip's source media locally with word-level timestamps (faster-whisper). Writes a _transcript.json next to each media file (reused by edit_by_transcript / remove_filler_words so media is only transcribed once) and optionally an SRT for captions. Requires the optional [transcribe] extra; degrades to an install hint without it.",
inputSchema={
"type": "object",
"properties": {
"filepath": {"type": "string", "description": "Path to FCPXML file"},
"clip_name": {"type": "string", "description": "Only transcribe the clip with this name"},
"model": {"type": "string", "default": "base", "description": "Whisper model size: tiny, base, small, medium, large-v3 (default base; larger = slower + more accurate)"},
"language": {"type": "string", "description": "ISO language code hint (e.g. 'en'); auto-detected if omitted"},
"write_srt": {"type": "boolean", "default": False, "description": "Also write a _transcript.srt next to each media file (plugs into import_srt_markers)"},
},
"required": ["filepath"]
}
),
Tool(
name="edit_by_transcript",
description="Text-based editing: cut timeline content by what was SAID. mode=remove cuts every occurrence of the given phrases (with ripple); mode=keep_only keeps only the matched phrases and cuts everything else in each matched clip (clips with no matches are left untouched). Uses each media file's _transcript.json (auto-transcribes if missing). Non-destructive: writes a _transcript_edit copy.",
inputSchema={
"type": "object",
"properties": {
"filepath": {"type": "string", "description": "Path to FCPXML file"},
"phrases": {"type": "array", "items": {"type": "string"}, "description": "Spoken phrases to match (case/punctuation-insensitive)"},
"mode": {"type": "string", "enum": ["remove", "keep_only"], "default": "remove", "description": "remove=cut matches out; keep_only=keep only matches"},
"clip_name": {"type": "string", "description": "Only edit the clip with this name"},
"model": {"type": "string", "default": "base", "description": "Whisper model size if transcription is needed"},
"padding": {"type": "number", "default": 0.0, "description": "Seconds to widen each cut on both sides (0-2, default 0)"},
"output_path": {"type": "string", "description": "Output path (default: adds _transcript_edit suffix)"},
},
"required": ["filepath", "phrases"]
}
),
Tool(
name="remove_filler_words",
description="Cut filler words (um, uh, erm...) out of the timeline with ripple, using word-level transcripts of the real source audio. Conservative default filler list — words like 'like' and 'so' are only cut if you pass them explicitly. Uses each media file's _transcript.json (auto-transcribes if missing). Non-destructive: writes a _defillered copy.",
inputSchema={
"type": "object",
"properties": {
"filepath": {"type": "string", "description": "Path to FCPXML file"},
"fillers": {"type": "array", "items": {"type": "string"}, "description": "Filler words/phrases to cut (default: um, uh, uhh, umm, erm, ehm, mmm, hmm, mhm)"},
"clip_name": {"type": "string", "description": "Only clean the clip with this name"},
"model": {"type": "string", "default": "base", "description": "Whisper model size if transcription is needed"},
"padding": {"type": "number", "default": 0.02, "description": "Seconds to widen each cut on both sides (0-2, default 0.02)"},
"output_path": {"type": "string", "description": "Output path (default: adds _defillered suffix)"},
},
"required": ["filepath"]
}
),
]
async def handle_transcribe_media(arguments: dict) -> Sequence[TextContent]:
model = arguments.get("model", "base")
language = arguments.get("language")
output_dir = arguments.get("output_dir")
write_srt = bool(arguments.get("write_srt", False))
_, tl = _require_timeline(arguments["filepath"])
clip_filter = arguments.get("clip_name")
done: dict[str, dict | None] = {}
skipped: list[tuple[str, str]] = []
rows: list[list[str]] = []
srt_paths: list[str] = []
for clip in tl.clips:
if clip_filter and clip.name != clip_filter:
continue
media_path = media_src_to_path(clip.media_path or "")
if not media_path or not Path(media_path).is_file():
skipped.append((clip.name, "media file missing"))
continue
if media_path in done:
continue
if len(done) >= TRANSCRIBE_MAX_MEDIA:
skipped.append((clip.name, f"transcription cap reached ({TRANSCRIBE_MAX_MEDIA} media files)"))
continue
data, reason = _load_or_transcribe(media_path, model, language, output_dir)
done[media_path] = data
if data is None:
skipped.append((clip.name, reason))
continue
if write_srt and data.get("segments"):
srt_name = Path(media_path).stem + "_transcript.srt"
srt_anchor = str(Path(output_dir).expanduser()) if output_dir else str(Path(media_path).parent)
srt_path = _validate_output_path(
str(Path(srt_anchor) / srt_name),
anchor_dir=srt_anchor,
)
with open(srt_path, "w") as f:
f.write(segments_to_srt(data["segments"]))
srt_paths.append(srt_path)
preview = data.get("text", "")[:160]
rows.append([
Path(media_path).name,
data.get("language", "?"),
str(len(data.get("words", []))),
format_duration(float(data.get("duration", 0.0))),
preview + ("…" if len(data.get("text", "")) > 160 else ""),
])
result = f"""# Media Transcription (local Whisper)
## Summary
- **Model**: {model}
- **Media Files Transcribed**: {len(rows)}
"""
if rows:
result += "\n## Transcripts (saved as _transcript.json next to each media file)\n"
result += _markdown_table(
["Media", "Language", "Words", "Duration", "Preview"], rows
) + "\n"
result += (
"\n*Next: `edit_by_transcript` to cut by what was said, or "
"`remove_filler_words` to clean ums/uhs. Transcripts are cached — "
"media is only transcribed once.*"
)
if srt_paths:
result += "\n\n## SRT Files\n" + "\n".join(f"- {p}" for p in srt_paths)
if skipped:
result += "\n## Skipped Clips\n" + _markdown_table(
["Clip", "Reason"], [[name, reason] for name, reason in skipped]
) + "\n"
if not rows and any("faster-whisper" in reason for _, reason in skipped):
result += _TRANSCRIBE_INSTALL_HINT
return _text_result(result)
async def handle_edit_by_transcript(arguments: dict) -> Sequence[TextContent]:
phrases = arguments.get("phrases") or []
if not isinstance(phrases, list) or not all(isinstance(p, str) for p in phrases):
raise ValueError("phrases must be a list of strings")
phrases = [p for p in phrases if p.strip()]
if not phrases:
raise ValueError("phrases must contain at least one non-empty string")
mode = arguments.get("mode", "remove")
if mode not in ("remove", "keep_only"):
raise ValueError(f"mode must be 'remove' or 'keep_only', got {mode!r}")
padding = float(arguments.get("padding", 0.0))
if not (0 <= padding <= 2):
raise ValueError(f"padding must be between 0 and 2 seconds, got {padding}")
model = arguments.get("model", "base")
language = arguments.get("language")
output_dir = arguments.get("output_dir")
filepath, output_path, modifier = _setup_modifier(arguments, "_transcript_edit")
def spans_fn(words):
return merge_ranges(
[span for phrase in phrases for span in find_phrase_spans(words, phrase)]
)
cuts_made, skipped = _cut_transcript_spans(
modifier, arguments.get("clip_name"), model, language, padding,
spans_fn, keep_only=(mode == "keep_only"), output_dir=output_dir,
)
if cuts_made:
modifier.save(output_path)
verb = "kept only" if mode == "keep_only" else "removed"
return _transcript_cut_report(
"Transcript Edit",
[f"- **Mode**: {mode} ({verb} the matched phrases)",
f"- **Phrases**: {', '.join(repr(p) for p in phrases)}",
f"- **Padding**: {padding}s"],
cuts_made, skipped, output_path,
"*Transcripts are cached as _transcript.json. Original file untouched.*",
)
async def handle_remove_filler_words(arguments: dict) -> Sequence[TextContent]:
fillers = arguments.get("fillers") or list(DEFAULT_FILLERS)
if not isinstance(fillers, list) or not all(isinstance(f, str) for f in fillers):
raise ValueError("fillers must be a list of strings")
padding = float(arguments.get("padding", 0.02))
if not (0 <= padding <= 2):
raise ValueError(f"padding must be between 0 and 2 seconds, got {padding}")
model = arguments.get("model", "base")
language = arguments.get("language")
output_dir = arguments.get("output_dir")
filepath, output_path, modifier = _setup_modifier(arguments, "_defillered")
cuts_made, skipped = _cut_transcript_spans(
modifier, arguments.get("clip_name"), model, language, padding,
lambda words: merge_ranges(find_filler_spans(words, fillers)),
output_dir=output_dir,
)
if cuts_made:
modifier.save(output_path)
return _transcript_cut_report(
"Filler Word Removal",
[f"- **Fillers**: {', '.join(fillers)}", f"- **Padding**: {padding}s"],
cuts_made, skipped, output_path,
"*Transcripts are cached as _transcript.json. Original file untouched.*",
)
HANDLERS = {
"transcribe_media": handle_transcribe_media,
"edit_by_transcript": handle_edit_by_transcript,
"remove_filler_words": handle_remove_filler_words,
}
+751
View File
@@ -0,0 +1,751 @@
"""Voz (análise → decisão → aplicação) — tool schemas and handlers.
Extracted from server.py; see Engine/docs/03_SERVER_TOOLS.md for the tool catalog.
"""
from __future__ import annotations
import json
from pathlib import Path
from typing import List, Optional, Sequence, Tuple
from mcp.types import TextContent, Tool
from fcpxml.diarize import assign_speakers, build_speakers, diarization_capability, diarize
from fcpxml.emphasis import EmphasisWeights
from fcpxml.media_intel import media_src_to_path
from fcpxml.model_manager import (
load_hf_token,
load_num_speakers,
load_voice_analysis_config,
save_voice_analysis_config,
)
from fcpxml.models import TimeValue
from fcpxml.voice_actions import parse_actions, resolve_actions, speaker_cut_actions
from fcpxml.voice_features import extract_energy, extract_pitch, features_capability
from fcpxml.voice_timeline import (
build_voice_timeline,
enrich_words,
load_voice_timeline,
restrict_to_kept,
save_voice_timeline,
select_peaks,
suggest_zoom_windows,
voice_timeline_path,
)
from fcpxml.writer import FCPXMLModifier
from server_tools._shared import (
_DIARIZATION_INSTALL_HINT,
_FEATURES_INSTALL_HINT,
_TRANSCRIBE_INSTALL_HINT,
AUDIO_MEDIA_EXTENSIONS,
MAX_MEDIA_FILE_SIZE,
_apply_placed_action,
_load_or_transcribe,
_markdown_table,
_setup_modifier,
_speaker_table,
_text_result,
_validate_filepath,
_validate_output_path,
_voice_analysis_config_text,
format_duration,
)
TOOLS = [
Tool(
name="diarize_media",
description="Identify WHO is speaking (speaker diarization) in an audio/video file using pyannote.audio, and assign SPEAKER_NN labels to each word/segment of its cached transcript. Writes a _diarization.json next to the media file. Requires the optional [diarization] extra (pyannote.audio) and a HuggingFace token with access to pyannote/speaker-diarization-3.1 (set once via save_hf_token or the HF_TOKEN argument); degrades to an install/token hint without them. Transcribes first if no _transcript.json is cached yet.",
inputSchema={
"type": "object",
"properties": {
"media_path": {"type": "string", "description": "Path to audio/video file (.wav, .mp3, .m4a, .aac, .aif, .flac, .mov, .mp4)"},
"hf_token": {"type": "string", "description": "HuggingFace token with pyannote/speaker-diarization-3.1 access (default: the persisted token from save_hf_token, if any)"},
"num_speakers": {"type": "string", "description": "Known number of speakers, if you know it (speeds up and improves accuracy). Leave empty to auto-detect."},
"model": {"type": "string", "default": "base", "description": "Whisper model size to use if transcription is needed (default base)"},
"language": {"type": "string", "description": "ISO language code hint for transcription, if needed"},
},
"required": ["media_path"]
}
),
Tool(
name="analyze_voice_features",
description="Analyze HOW a voice is speaking: pitch, energy, local speech rate, pauses, and a combined emphasis index (0-1) per transcribed word, using the persisted Voice Analysis settings (energy threshold, emphasis weights, emphasis cutoff — see save_voice_analysis_config). Writes a _voice_features.json next to the media file. Requires the optional [intelligence] extra (librosa); degrades to an install hint without it. Transcribes first if no _transcript.json is cached yet.",
inputSchema={
"type": "object",
"properties": {
"media_path": {"type": "string", "description": "Path to audio/video file (.wav, .mp3, .m4a, .aac, .aif, .flac, .mov, .mp4)"},
"model": {"type": "string", "default": "base", "description": "Whisper model size to use if transcription is needed (default base)"},
"language": {"type": "string", "description": "ISO language code hint for transcription, if needed"},
},
"required": ["media_path"]
}
),
Tool(
name="build_voice_timeline",
description="Build the consolidated voice timeline: WHAT was said (transcript), WHO said it (diarization), and HOW it was said (pitch/energy/rate/pauses -> emphasis index), merged into one AI-readable JSON written next to the media as _voice_timeline.json. This is the source of truth for automated editing — layered as summary -> segments -> words, with all acoustic values normalized 0-1 and documented inline, so a model can read the narrative shape and decide how to direct the edit. Every layer degrades independently: no librosa means acoustic values are 0, no HuggingFace token means a single default speaker; the document shape never changes.",
inputSchema={
"type": "object",
"properties": {
"media_path": {"type": "string", "description": "Path to audio/video file (.wav, .mp3, .m4a, .aac, .aif, .flac, .mov, .mp4)"},
"model": {"type": "string", "default": "base", "description": "Whisper model size to use if transcription is needed (default base)"},
"language": {"type": "string", "description": "ISO language code hint for transcription, if needed"},
"hf_token": {"type": "string", "description": "HuggingFace token for speaker diarization (default: the persisted token; omit to skip diarization)"},
"num_speakers": {"type": "string", "description": "Known number of speakers, if any (default: the persisted setting, else auto-detect)"},
"output_dir": {"type": "string", "description": "Folder to write _voice_timeline.json into (default: next to the media file)"},
},
"required": ["media_path"]
}
),
Tool(
name="remove_speakers",
description="Cut everything one or more speakers say out of the timeline — the standard cleanup on an interview shoot, where the interviewer or a crew member talks during the take and only the subject should survive. Reads the media's _voice_timeline.json (build it first with build_voice_timeline, which must have run with diarization so speakers are separated). Call without speaker_ids to just LIST who was detected, with speaking share and sample lines, so you can tell who is who before cutting anything. Non-destructive: writes a _voice_edit copy.",
inputSchema={
"type": "object",
"properties": {
"filepath": {"type": "string", "description": "Path to FCPXML file"},
"media_path": {"type": "string", "description": "Media whose _voice_timeline.json holds the speakers (default: the timeline's first clip media)"},
"speaker_ids": {"type": "array", "items": {"type": "string"}, "description": "Speakers to REMOVE (e.g. [\"SPEAKER_01\"]). Omit to only list the detected speakers without editing."},
"padding": {"type": "number", "default": 0.15, "description": "Seconds trimmed inside each cut so the kept speaker's first syllable is never clipped (default 0.15)"},
"output_path": {"type": "string", "description": "Output path (default: adds _voice_edit suffix)"},
},
"required": ["filepath"]
}
),
Tool(
name="refine_voice_timeline",
description="Re-analyze a voice timeline over only the material that survives a set of cuts, then propose punch-in windows over it. Emphasis is RELATIVE — energy is scored against the loudest word of the recording — so once the loudest moment is cut (a laugh, an aside to the crew), every remaining score is measured against something the viewer will never see and the ranking points at the wrong words. Run this after deciding cuts and before deciding zooms. Cheap: it re-normalizes the already-measured numbers, never re-reads the audio. Times stay in ORIGINAL source seconds, so the result feeds straight back into apply_voice_actions.",
inputSchema={
"type": "object",
"properties": {
"media_path": {"type": "string", "description": "Media whose _voice_timeline.json will be refined (build it first with build_voice_timeline)"},
"cuts": {
"type": "array",
"description": "The ranges being REMOVED, in original source seconds. Pass the cut actions you already decided; anything overlapping them is excluded from the re-analysis.",
"items": {
"type": "object",
"properties": {
"start": {"type": "number", "description": "Start in original source seconds"},
"end": {"type": "number", "description": "End in original source seconds"},
},
"required": ["start", "end"],
},
},
"min_gap": {"type": "number", "default": 8.0, "description": "Minimum seconds between two proposed zooms — effects stacked close together read as nervous editing (default 8.0)"},
"max_zooms": {"type": "integer", "description": "Cap on how many zoom candidates to return (default: no cap — cut the list by rhythm yourself)"},
"save": {"type": "boolean", "default": False, "description": "Also write the refined timeline as _voice_timeline_refined.json next to the media"},
"output_dir": {"type": "string", "description": "Folder holding _voice_timeline.json (default: next to the media file)"},
},
"required": ["media_path", "cuts"]
}
),
Tool(
name="apply_voice_actions",
description="Apply a list of editing decisions (from the rules engine, or from a model that read the _voice_timeline.json) to a timeline, producing FCPXML. Actions are validated first and reported per row, so one malformed decision never discards the edit. All action times are in ORIGINAL source seconds: cuts are resolved first and every other action is moved onto its post-cut position automatically, so decisions never land on the wrong frame. Actions pointing into removed material are dropped and reported, not silently slid. Non-destructive: writes a _voice_edit copy.",
inputSchema={
"type": "object",
"properties": {
"filepath": {"type": "string", "description": "Path to FCPXML file"},
"actions": {
"type": "array",
"description": "The decision list. Each item: {kind, start, end, params, reason, speaker}. kind is cut | zoom | text | marker. Times in original source seconds. zoom takes params.scale (1.0-3.0, default 1.3); text requires params.content.",
"items": {
"type": "object",
"properties": {
"kind": {"type": "string", "enum": ["cut", "zoom", "text", "marker"]},
"start": {"type": "number", "description": "Start in original source seconds"},
"end": {"type": "number", "description": "End in original source seconds"},
"params": {"type": "object", "description": "kind-specific: {scale} for zoom, {content} for text"},
"reason": {"type": "string", "description": "Why this decision was made — kept for review"},
"speaker": {"type": "string", "description": "Speaker id this decision relates to, if any"},
},
"required": ["kind", "start", "end"],
},
},
"output_path": {"type": "string", "description": "Output path (default: adds _voice_edit suffix)"},
},
"required": ["filepath", "actions"]
}
),
Tool(
name="get_voice_analysis_config",
description="Read the persisted Voice Analysis settings: energy threshold, emphasis-index weights (energy/pitch_variation/rate_variation/pause_before/duration), emphasis cutoff for punch-in candidates, and emotion detection toggle/sensitivity. Shared with the MacApp settings screen (~/.fcp-mcp-server/config.json).",
inputSchema={"type": "object", "properties": {}}
),
Tool(
name="save_voice_analysis_config",
description="Persist Voice Analysis settings. Only the fields you pass are changed; omitted fields keep their current value. emphasis_weights don't need to sum to 1 (normalized internally). Shared with the MacApp settings screen (~/.fcp-mcp-server/config.json).",
inputSchema={
"type": "object",
"properties": {
"energy_threshold": {"type": "number", "description": "0-1, how loud (normalized RMS) counts as 'high energy' (default 0.5)"},
"emphasis_weights": {
"type": "object",
"description": "Any subset of {energy, pitch_variation, rate_variation, pause_before, duration} weights for the emphasis index",
"properties": {
"energy": {"type": "number"},
"pitch_variation": {"type": "number"},
"rate_variation": {"type": "number"},
"pause_before": {"type": "number"},
"duration": {"type": "number"},
},
},
"peak_percentile": {"type": "number", "description": "Fraction of words selected as peaks, 0-1 (default 0.02 = top 2%). Selection is relative because the emphasis index's real range depends on the material — measured on a real interview it never passed 0.55."},
"emphasis_floor": {"type": "number", "description": "0-1 minimum emphasis for a peak, guarding genuinely flat audio (default 0.25)"},
"emotion_enabled": {"type": "boolean", "description": "Whether emotion detection runs as part of voice analysis (default false)"},
"emotion_sensitivity": {"type": "number", "description": "0-1 confidence threshold to accept an emotion label (default 0.5)"},
},
}
),
]
async def handle_diarize_media(arguments: dict) -> Sequence[TextContent]:
media_path = _validate_filepath(
arguments["media_path"], AUDIO_MEDIA_EXTENSIONS, max_size=MAX_MEDIA_FILE_SIZE
)
token = str(arguments.get("hf_token") or "").strip() or load_hf_token() or None
num_speakers = str(arguments.get("num_speakers") or "").strip()
model = arguments.get("model", "base")
language = arguments.get("language")
ok, message = diarization_capability(token)
if not ok:
return _text_result(f"# Speaker Diarization\n\n{message}{_DIARIZATION_INSTALL_HINT}")
transcript, reason = _load_or_transcribe(media_path, model, language)
if transcript is None:
return _text_result(
f"# Speaker Diarization\n\nCould not obtain a transcript to diarize "
f"({reason}).{_TRANSCRIBE_INSTALL_HINT}"
)
tracks = diarize(media_path, token, num_speakers)
if tracks is None:
return _text_result(
"# Speaker Diarization\n\nDiarization failed — check the HuggingFace "
"token has accepted the pyannote/speaker-diarization-3.1 model terms, "
"and that the media file is readable."
)
segments, words = assign_speakers(
transcript.get("segments", []), transcript.get("words", []), tracks
)
speakers = build_speakers(segments)
diarization_data = {
"source": Path(media_path).name,
"speakers": speakers,
"segments": segments,
"words": words,
}
json_path = _validate_output_path(
str(Path(media_path).with_name(Path(media_path).stem + "_diarization.json")),
anchor_dir=str(Path(media_path).parent),
)
with open(json_path, "w") as f:
json.dump(diarization_data, f, indent=2)
result_text = f"""# Speaker Diarization
## Summary
- **Source**: {Path(media_path).name}
- **Speakers Detected**: {len(speakers)}
- **Segments**: {len(segments)}
- **Diarization JSON**: {json_path}
## Speakers
"""
result_text += _markdown_table(
["ID", "Name"], [[s["id"], s["name"]] for s in speakers]
) + "\n"
result_text += (
"\n*Next: `build_voice_timeline` to cross this with acoustic features, "
"or use the segments/words directly for speaker-aware editing.*"
)
return _text_result(result_text)
async def handle_analyze_voice_features(arguments: dict) -> Sequence[TextContent]:
media_path = _validate_filepath(
arguments["media_path"], AUDIO_MEDIA_EXTENSIONS, max_size=MAX_MEDIA_FILE_SIZE
)
model = arguments.get("model", "base")
language = arguments.get("language")
ok, message = features_capability()
if not ok:
return _text_result(f"# Voice Feature Analysis\n\n{message}{_FEATURES_INSTALL_HINT}")
transcript, reason = _load_or_transcribe(media_path, model, language)
if transcript is None:
return _text_result(
f"# Voice Feature Analysis\n\nCould not obtain a transcript to analyze "
f"({reason}).{_TRANSCRIBE_INSTALL_HINT}"
)
words = transcript.get("words", [])
if not words:
return _text_result("# Voice Feature Analysis\n\nNo words in transcript — nothing to analyze.")
config = load_voice_analysis_config()
weights = EmphasisWeights.from_dict(config["emphasis_weights"])
enriched = enrich_words(
words, extract_pitch(media_path), extract_energy(media_path), weights
)
energy_threshold = config["energy_threshold"]
high_energy_words = [w for w in enriched if w["energy_norm"] >= energy_threshold]
high_emphasis_words = select_peaks(
enriched, config["peak_percentile"], config["emphasis_floor"]
)
features_data = {"source": Path(media_path).name, "config": config, "words": enriched}
json_path = _validate_output_path(
str(Path(media_path).with_name(Path(media_path).stem + "_voice_features.json")),
anchor_dir=str(Path(media_path).parent),
)
with open(json_path, "w") as f:
json.dump(features_data, f, indent=2)
result_text = f"""# Voice Feature Analysis
## Summary
- **Source**: {Path(media_path).name}
- **Words Analyzed**: {len(enriched)}
- **High-Energy Words** (>= {energy_threshold:.2f}): {len(high_energy_words)}
- **Peak Words** (top {config["peak_percentile"]:.1%}): {len(high_emphasis_words)}
- **Features JSON**: {json_path}
## Top Emphasis Words
"""
top = sorted(enriched, key=lambda w: w["emphasis"], reverse=True)[:10]
result_text += _markdown_table(
["Word", "Time", "Emphasis", "Energy", "Pitch Δ"],
[
[
w.get("word", ""),
f"{w.get('start', 0):.2f}s",
f"{w['emphasis']:.2f}",
f"{w['energy_norm']:.2f}",
f"{w['pitch_delta']:.2f}",
]
for w in top
],
) + "\n"
result_text += (
"\n*Thresholds and emphasis weights are configurable in Voice Analysis "
"settings (`save_voice_analysis_config`). Next: `diarize_media` to add "
"speaker labels.*"
)
return _text_result(result_text)
async def handle_build_voice_timeline(arguments: dict) -> Sequence[TextContent]:
media_path = _validate_filepath(
arguments["media_path"], AUDIO_MEDIA_EXTENSIONS, max_size=MAX_MEDIA_FILE_SIZE
)
model = arguments.get("model", "base")
language = arguments.get("language")
token = str(arguments.get("hf_token") or "").strip() or load_hf_token() or None
num_speakers = str(arguments.get("num_speakers") or "").strip() or load_num_speakers()
transcript, reason = _load_or_transcribe(media_path, model, language)
if transcript is None:
return _text_result(
f"# Voice Timeline\n\nCould not obtain a transcript "
f"({reason}).{_TRANSCRIBE_INSTALL_HINT}"
)
config = load_voice_analysis_config()
timeline = build_voice_timeline(
media_path,
transcript,
hf_token=token,
num_speakers=num_speakers,
weights=EmphasisWeights.from_dict(config["emphasis_weights"]),
peak_percentile=config["peak_percentile"],
emphasis_floor=config["emphasis_floor"],
)
output_dir = arguments.get("output_dir")
json_path = Path(_validate_output_path(
str(voice_timeline_path(media_path, output_dir)),
anchor_dir=str(Path(output_dir) if output_dir else Path(media_path).parent),
))
save_voice_timeline(timeline, json_path)
summary = timeline["summary"]
layers = timeline["layers"]
result_text = f"""# Voice Timeline
## Summary
- **Source**: {timeline["source"]}
- **Duration**: {format_duration(summary["duration"])}
- **Speakers**: {summary["speaker_count"]}
- **Segments**: {summary["segment_count"]} ({summary["word_count"]} words)
- **Average Emphasis**: {summary["avg_emphasis"]:.2f}
- **Peak Moments** ({summary["peak_selection"]}): {summary["peak_count"]}
- **Timeline JSON**: {json_path}
## Analysis Layers
"""
result_text += _markdown_table(
["Layer", "Status"],
[
["Transcript", "yes" if layers["transcript"] else "empty"],
[
"Acoustics (pitch/energy)",
"yes" if layers["acoustics"] else "FAILED — every acoustic value is 0",
],
["Speakers", "yes" if layers["speakers"] else "not run — single default speaker"],
],
) + "\n"
if summary["peak_moments"]:
result_text += "\n## Peak Moments\n"
result_text += _markdown_table(
["Time", "Word", "Speaker", "Emphasis"],
[
[
f"{m['time']:.2f}s",
m["text"],
m["speaker"],
f"{m['emphasis']:.2f}",
]
for m in summary["peak_moments"][:10]
],
) + "\n"
result_text += (
"\n*The JSON is layered summary -> segments -> words with normalized "
"0-1 values, ready to hand to a model for edit direction.*"
)
return _text_result(result_text)
async def handle_remove_speakers(arguments: dict) -> Sequence[TextContent]:
filepath = _validate_filepath(arguments["filepath"], ('.fcpxml', '.fcpxmld'))
speaker_ids = arguments.get("speaker_ids") or []
media_path = arguments.get("media_path")
if not media_path:
modifier = FCPXMLModifier(filepath)
for _, clip_el in modifier._iter_spine_clips():
src = modifier.resources.get(clip_el.get("ref", ""), {}).get("src", "")
candidate = media_src_to_path(src)
if candidate and Path(candidate).is_file():
media_path = candidate
break
if not media_path:
return _text_result(
"# Speakers\n\nNo source media found for this timeline — pass `media_path` explicitly."
)
timeline = _read_voice_timeline(media_path, arguments.get("output_dir"))
if timeline is None:
return _text_result(
f"# Speakers\n\nNo voice timeline for `{Path(media_path).name}` yet.\n\n"
"Run `build_voice_timeline` on it first."
)
profiles = timeline.get("speakers", [])
header = f"# Speakers in {timeline.get('source', '')}\n\n" + _speaker_table(profiles) + "\n"
if len(profiles) < 2:
header += (
"\n> Only one speaker is present. Either the recording really has one "
"voice, or diarization did not run — check the Models tab for the "
"HuggingFace token.\n"
)
for p in profiles:
samples = [s for s in p.get("samples", []) if s]
if samples:
header += f"\n**{p['id']}** ({p.get('name', '')}) says things like:\n"
header += "".join(f"> {s}\n" for s in samples[:2])
if not speaker_ids:
return _text_result(
header
+ "\n*Nothing was edited. Re-run with `speaker_ids` naming who to REMOVE — "
"typically the interviewer or crew, keeping the subject.*"
)
known = {p["id"] for p in profiles}
unknown = [s for s in speaker_ids if s not in known]
if unknown:
return _text_result(
header + f"\n**Unknown speaker(s): {', '.join(unknown)}** — nothing was edited."
)
if set(speaker_ids) >= known:
return _text_result(
header + "\n**That would remove every speaker**, leaving nothing — nothing was edited."
)
actions = speaker_cut_actions(
timeline, speaker_ids, padding=float(arguments.get("padding", 0.15))
)
if not actions:
return _text_result(header + "\n No speech found for those speakers — nothing was edited.")
result = await handle_apply_voice_actions({
**arguments,
"actions": [a.as_dict() for a in actions],
})
removed = sum(a.duration for a in actions)
return _text_result(
header
+ f"\n## Removed\n- **Speakers cut**: {', '.join(speaker_ids)}\n"
+ f"- **Speech removed**: {format_duration(removed)} across {len(actions)} segments\n\n"
+ result[0].text
)
def _read_voice_timeline(media_path: str, output_dir: Optional[str] = None) -> Optional[dict]:
"""Load the cached voice timeline, project folder first.
``build_voice_timeline`` writes to the chosen project folder when one is
set and beside the media otherwise, so a reader that only checks one of
the two reports "no voice timeline yet" for a file that exists. Checking
both also keeps timelines built before the project folder existed
readable.
"""
if output_dir:
timeline = load_voice_timeline(voice_timeline_path(media_path, output_dir))
if timeline is not None:
return timeline
return load_voice_timeline(voice_timeline_path(media_path))
async def handle_refine_voice_timeline(arguments: dict) -> Sequence[TextContent]:
media_path = _validate_filepath(
arguments["media_path"], AUDIO_MEDIA_EXTENSIONS, max_size=MAX_MEDIA_FILE_SIZE
)
timeline = _read_voice_timeline(media_path, arguments.get("output_dir"))
if timeline is None:
return _text_result(
f"# Refined Voice Timeline\n\nNo voice timeline for "
f"`{Path(media_path).name}` yet.\n\nRun `build_voice_timeline` on it first."
)
cut_ranges: List[Tuple[float, float]] = []
rejected: List[str] = []
for i, raw in enumerate(arguments.get("cuts") or []):
try:
start, end = float(raw["start"]), float(raw["end"])
except (TypeError, ValueError, KeyError):
rejected.append(f"cut #{i}: start/end must be numbers")
continue
if end <= start:
rejected.append(f"cut #{i}: end ({end}) must be after start ({start})")
continue
cut_ranges.append((start, end))
config = load_voice_analysis_config()
refined = restrict_to_kept(
timeline,
cut_ranges,
weights=EmphasisWeights.from_dict(config["emphasis_weights"]),
peak_percentile=config["peak_percentile"],
emphasis_floor=config["emphasis_floor"],
)
before, after = timeline["summary"], refined["summary"]
zooms = suggest_zoom_windows(
refined,
min_gap=float(arguments.get("min_gap", 8.0)),
max_zooms=arguments.get("max_zooms"),
)
result = f"""# Refined Voice Timeline
Re-normalized over the material that survives {len(cut_ranges)} cut(s).
"""
result += _markdown_table(
["Measure", "Raw recording", "Survivors only"],
[
["Duration", format_duration(before["duration"]), format_duration(after["duration"])],
["Segments", str(before["segment_count"]), str(after["segment_count"])],
["Words", str(before["word_count"]), str(after["word_count"])],
["Average emphasis", f"{before['avg_emphasis']:.3f}", f"{after['avg_emphasis']:.3f}"],
["Peak moments", str(before["peak_count"]), str(after["peak_count"])],
],
) + "\n"
if after["word_count"] == 0:
result += "\n> The cuts removed every word — nothing left to analyze.\n"
if after["peak_moments"]:
result += "\n## Peak Moments (re-ranked)\n"
result += _markdown_table(
["Time", "Word", "Speaker", "Emphasis"],
[
[f"{m['time']:.2f}s", m["text"], m["speaker"], f"{m['emphasis']:.2f}"]
for m in after["peak_moments"][:10]
],
) + "\n"
if zooms:
result += "\n## Zoom Candidates\n"
result += _markdown_table(
["Start", "End", "Word", "Emphasis", "Line"],
[
[f"{z['start']:.2f}s", f"{z['end']:.2f}s", z["word"],
f"{z['emphasis']:.2f}", z["line"][:60]]
for z in zooms
],
) + "\n"
else:
result += "\n## Zoom Candidates\n\nNone — no content words survived the cuts.\n"
if arguments.get("save"):
p = Path(media_path)
json_path = Path(_validate_output_path(
str(p.with_name(p.stem + "_voice_timeline_refined.json")),
anchor_dir=str(p.parent),
))
save_voice_timeline(refined, json_path)
result += f"\n**Refined JSON**: {json_path}\n"
if rejected:
result += "\n## Rejected cuts\n" + "\n".join(f"- {r}" for r in rejected) + "\n"
result += (
"\n*Candidates, not obligations — cut the list by rhythm. Times are in "
"original source seconds, ready for `apply_voice_actions`.*"
)
return _text_result(result)
async def handle_apply_voice_actions(arguments: dict) -> Sequence[TextContent]:
raw_actions = arguments.get("actions")
if raw_actions is None:
return _text_result("# Voice Actions\n\nNo `actions` provided — nothing to apply.")
actions, errors = parse_actions(raw_actions)
if not actions:
text = "# Voice Actions\n\nNo valid actions to apply."
if errors:
text += "\n\n## Rejected\n" + "\n".join(f"- {e}" for e in errors)
return _text_result(text)
cut_ranges, placed, dropped = resolve_actions(actions)
filepath, output_path, modifier = _setup_modifier(arguments, "_voice_edit")
def source_window(clip_el) -> tuple[float, float]:
"""The span of source media a spine clip actually uses."""
start = modifier.source_file_start(clip_el).to_seconds()
duration = modifier._parse_time(clip_el.get("duration", "0s")).to_seconds()
return start, start + duration
applied: list[list[str]] = []
unplaced: list[str] = []
# Placements go on before cuts: they are anchored in source coordinates,
# and cutting afterwards ripples the spine around them.
# Cuts go FIRST. Cutting splits a clip into pieces and rewrites the
# spine around them, which would duplicate a zoom onto every piece and
# lose markers entirely. Cutting first means placements land on final,
# stable clips — and `resolve_actions` already moved their times onto
# the post-cut timeline, so they still point at the same moment.
cuts_made = 0
for _, clip_el in modifier._iter_spine_clips():
clip_start, clip_end = source_window(clip_el)
to_frame = modifier.snap_seconds_to_frame
ranges = [
(to_frame(max(s, clip_start) - clip_start), to_frame(min(e, clip_end) - clip_start))
for s, e in cut_ranges
if min(e, clip_end) > max(s, clip_start)
]
ranges = [(a, b) for a, b in ranges if b > a]
if ranges and modifier.cut_clip_ranges(clip_el, ranges) > TimeValue.zero():
cuts_made += len(ranges)
def timeline_window(clip_el) -> tuple[float, float]:
"""Where a spine clip sits on the timeline, in seconds."""
offset = modifier._parse_time(clip_el.get("offset", "0s")).to_seconds()
duration = modifier._parse_time(clip_el.get("duration", "0s")).to_seconds()
return offset, offset + duration
for action in placed:
host = next(
(
clip_el
for _, clip_el in modifier._iter_spine_clips()
if timeline_window(clip_el)[0] <= action.start < timeline_window(clip_el)[1]
),
None,
)
if host is None:
unplaced.append(
f"{action.kind} @ {action.start:.2f}s — falls outside the edited timeline"
)
continue
try:
what = _apply_placed_action(modifier, host, action, timeline_window(host)[0])
applied.append([f"{action.start:.2f}s", what, host.get("name", ""), action.reason])
except (ValueError, KeyError) as exc:
unplaced.append(f"{action.kind} @ {action.start:.2f}s — {exc}")
# Rename the project so it does not land in the library indistinguishable
# from the original. FCP imports by the name in the XML, so an untouched
# name puts two same-named projects in the same event — and the edit looks
# like it did nothing, because the original is what gets opened.
project_name = ""
project_el = modifier.root.find(".//project")
if project_el is not None:
project_name = f"{project_el.get('name', 'Projeto')} — corte por voz"
project_el.set("name", project_name)
modifier.save(output_path)
result = f"""# Voice Actions Applied
## Summary
- **Actions received**: {len(actions)}
- **Placed** (zoom/text/marker): {len(applied)}
- **Cuts applied**: {cuts_made}
- **Project name**: {project_name or '(unchanged)'}
- **Saved to**: `{output_path}`
"""
if applied:
result += "## Applied\n" + _markdown_table(
["Time", "Action", "Clip", "Reason"], applied[:40]
) + "\n"
if dropped:
result += "\n## Dropped (pointed into removed material)\n" + "\n".join(
f"- {a.kind} @ {a.start:.2f}s — {a.reason or 'no reason given'}" for a in dropped
) + "\n"
if unplaced:
result += "\n## Not placed\n" + "\n".join(f"- {u}" for u in unplaced) + "\n"
if errors:
result += "\n## Rejected\n" + "\n".join(f"- {e}" for e in errors) + "\n"
result += "\n*Non-destructive: the original file is untouched.*"
return _text_result(result)
async def handle_get_voice_analysis_config(arguments: dict) -> Sequence[TextContent]:
return _text_result(_voice_analysis_config_text(load_voice_analysis_config()))
async def handle_save_voice_analysis_config(arguments: dict) -> Sequence[TextContent]:
config = save_voice_analysis_config(
energy_threshold=arguments.get("energy_threshold"),
emphasis_weights=arguments.get("emphasis_weights"),
peak_percentile=arguments.get("peak_percentile"),
emphasis_floor=arguments.get("emphasis_floor"),
emotion_enabled=arguments.get("emotion_enabled"),
emotion_sensitivity=arguments.get("emotion_sensitivity"),
)
return _text_result(_voice_analysis_config_text(config))
HANDLERS = {
"diarize_media": handle_diarize_media,
"analyze_voice_features": handle_analyze_voice_features,
"build_voice_timeline": handle_build_voice_timeline,
"remove_speakers": handle_remove_speakers,
"refine_voice_timeline": handle_refine_voice_timeline,
"apply_voice_actions": handle_apply_voice_actions,
"get_voice_analysis_config": handle_get_voice_analysis_config,
"save_voice_analysis_config": handle_save_voice_analysis_config,
}
+266
View File
@@ -0,0 +1,266 @@
"""Tests for collision detection and subtitle-layout validation (Fase 1).
Covers the pure functions in ``fcpxml.collision`` (spatial/temporal overlap,
area/ratio classification, distance, separation suggestion, box measurement)
and the integration through ``FCPXMLModifier.validate_subtitle_layout`` over a
generated document — the post-generation guarantee the layout engine only
provides by construction.
"""
import shutil
import tempfile
from pathlib import Path
import pytest
from fcpxml.collision import (
FONT_MISSING,
FONT_TOO_SMALL,
OUTSIDE_FRAME,
OUTSIDE_SAFE_AREA,
OVERLAP_PROBABLE,
OVERLAP_RENDER_TOLERANCE,
OVERLAP_SEVERE,
SPATIAL_COLLISION,
Box,
blocking,
classify_overlap,
distance_between,
measure_title_box,
overlap_metrics,
separation_suggestion,
temporal_overlap,
validate_titles,
)
from fcpxml.models import DynamicSubtitleConfig
from fcpxml.writer import FCPXMLModifier
SAMPLE = Path(__file__).parent.parent / "examples" / "sample.fcpxml"
WORD_MODE = DynamicSubtitleConfig(granularity="word")
WORDS = [
{"word": "Hello", "start": 0.0, "end": 0.4},
{"word": "there", "start": 0.4, "end": 0.8},
{"word": "friend", "start": 0.8, "end": 1.3},
]
@pytest.fixture
def temp_fcpxml():
with tempfile.NamedTemporaryFile(suffix=".fcpxml", delete=False) as f:
shutil.copy(SAMPLE, f.name)
yield f.name
Path(f.name).unlink(missing_ok=True)
def _title(text, x=0.0, y=0.0, *, start=0.0, end=1.0, font="Helvetica Neue",
face=None, font_size=100.0, kerning=0.0, group=0):
return {
"text": text,
"x": x,
"y": y,
"start": start,
"end": end,
"font": font,
"face": face,
"font_size": font_size,
"kerning": kerning,
"group": group,
}
class TestBoxOverlap:
def test_touching_boxes_do_not_overlap(self):
a = Box(0, 100, 0, 50)
b = Box(100, 200, 0, 50) # shares the right edge
assert not a.overlaps(b)
assert overlap_metrics(a, b)["overlap_area"] == 0
def test_overlap_by_one_pixel_detected(self):
a = Box(0, 100, 0, 50)
b = Box(99, 200, 0, 50) # one-pixel horizontal overlap
assert a.overlaps(b)
metrics = overlap_metrics(a, b)
assert metrics["overlap_width"] == 1
assert metrics["overlap_area"] == 50
def test_vertical_only_overlap(self):
a = Box(0, 100, 0, 50)
b = Box(0, 100, 49, 100) # one-pixel vertical overlap
assert a.overlaps(b)
assert overlap_metrics(a, b)["overlap_height"] == 1
class TestTemporalOverlap:
def test_adjacent_intervals_are_not_simultaneous(self):
# [0, 1) and [1, 2) share no instant.
assert not temporal_overlap(0.0, 1.0, 1.0, 2.0)
assert not temporal_overlap(1.0, 2.0, 0.0, 1.0)
def test_interleaved_intervals_overlap(self):
assert temporal_overlap(0.0, 2.0, 1.0, 3.0)
def test_contained_interval_overlaps(self):
assert temporal_overlap(0.0, 5.0, 1.0, 2.0)
def test_boundary_survives_float_noise_from_the_writer(self):
"""Found on real footage: the writer sets one block's title duration
to make its end land EXACTLY on the next block's start (same exact
FCPXML fraction), but end here is re-derived as start + duration —
two independently-rounded floats — which isn't bit-identical to the
other title's start read as a single division of that same
fraction. 7 of 8 "severe" collisions from one real clip were this,
off by ~1e-13s, far below any frame boundary."""
start_a, duration_a = 74883609 / 24000, 43043 / 24000
end_a = start_a + duration_a # float addition, like validate_titles does
start_b = 74926652 / 24000 # the exact same instant, read directly
assert end_a != start_b # the float noise is real
assert abs(end_a - start_b) < 1e-6 # ...and far below one frame
assert not temporal_overlap(start_a, end_a, start_b, start_b + 1.0)
class TestClassifyOverlap:
def test_zero_area_is_no_conflict(self):
assert classify_overlap(
{"overlap_area": 0, "overlap_height": 0, "overlap_ratio": 0}
) == "none"
def test_ratio_dominates_small_height(self):
# A sliver of 5px that covers most of a tiny box is still severe.
metrics = {"overlap_area": 50, "overlap_height": 5, "overlap_ratio": 0.9}
assert classify_overlap(metrics) == OVERLAP_SEVERE
def test_height_buckets(self):
def metrics(height):
return {"overlap_area": 1, "overlap_height": height,
"overlap_ratio": 0.0}
assert classify_overlap(metrics(5)) == OVERLAP_RENDER_TOLERANCE
assert classify_overlap(metrics(6)) == "warning"
assert classify_overlap(metrics(21)) == OVERLAP_PROBABLE
assert classify_overlap(metrics(51)) == OVERLAP_SEVERE
class TestDistanceAndSeparation:
def test_horizontal_distance(self):
a = Box(0, 100, 0, 50)
b = Box(150, 200, 0, 50)
d = distance_between(a, b)
assert d["distance_x"] == 50
assert d["distance_y"] == 0
assert d["distance"] == 50
def test_vertical_distance(self):
a = Box(0, 100, 0, 50)
b = Box(0, 100, 80, 130)
d = distance_between(a, b)
assert d["distance_y"] == 30
assert d["distance_x"] == 0
def test_separation_picks_smallest_axis(self):
a = Box(0, 100, 0, 100)
b = Box(90, 190, 95, 195) # 10px horizontal, 5px vertical penetration
suggestion = separation_suggestion(a, b)
assert suggestion["axis"] == "vertical"
assert suggestion["minimum_movement"] == 5
def test_blocking_severities(self):
assert blocking(OVERLAP_SEVERE)
assert blocking(OVERLAP_PROBABLE)
assert not blocking("warning")
assert not blocking("none")
class TestValidateTitles:
def test_clean_layout_has_no_issues(self):
report = validate_titles(
[_title("um", x=-200), _title("dois", x=200)],
2160, 3840,
)
assert report["severity"] == "none"
assert report["issues"] == []
def test_simultaneous_collision_reported(self):
titles = [_title("A", x=0), _title("B", x=0)]
report = validate_titles(titles, 2160, 3840)
collisions = [
i for i in report["issues"] if i["type"] == SPATIAL_COLLISION
]
assert len(collisions) == 1
assert collisions[0]["severity"] == OVERLAP_SEVERE
assert collisions[0]["suggested_correction"]["axis"] in (
"vertical", "horizontal",
)
def test_collision_across_times_ignored(self):
titles = [
_title("A", x=0, start=0.0, end=1.0),
_title("B", x=0, start=1.0, end=2.0),
]
report = validate_titles(titles, 2160, 3840)
assert not any(
i["type"] == SPATIAL_COLLISION for i in report["issues"]
)
def test_outside_frame_is_error(self):
report = validate_titles([_title("fora", x=5000)], 2160, 3840)
assert any(i["type"] == OUTSIDE_FRAME for i in report["issues"])
def test_outside_safe_area_is_warning_not_frame(self):
# Near the right edge: inside the frame, outside the 5% safe area.
report = validate_titles(
[_title("a", x=1000, font_size=40)], 2160, 3840,
)
assert any(i["type"] == OUTSIDE_SAFE_AREA for i in report["issues"])
assert not any(i["type"] == OUTSIDE_FRAME for i in report["issues"])
def test_resolution_changes_safe_area(self):
# Same x is fine on a wide frame but out of the safe area on a narrow one.
narrow = validate_titles([_title("a", x=1000, font_size=40)], 1920, 1080)
assert any(i["type"] == OUTSIDE_FRAME for i in narrow["issues"])
def test_font_missing_reported(self):
report = validate_titles(
[_title("oi", font="Comic Sans MS")], 2160, 3840,
)
assert any(i["type"] == FONT_MISSING for i in report["issues"])
def test_font_too_small_reported(self):
report = validate_titles(
[_title("oi", font_size=10)], 2160, 3840, min_font_size=20,
)
assert any(i["type"] == FONT_TOO_SMALL for i in report["issues"])
def test_measure_title_box_uses_real_width(self):
box = measure_title_box("ii", 100, x=0, y=0, font="Helvetica Neue")
# "ii" is the narrowest glyph; a single "W" is much wider.
wide = measure_title_box("W", 100, x=0, y=0, font="Helvetica Neue")
assert wide.width > box.width
class TestIntegration:
def test_generated_layout_is_clean(self, temp_fcpxml):
modifier = FCPXMLModifier(temp_fcpxml)
modifier.generate_dynamic_subtitles("Interview_A", WORDS, WORD_MODE)
report = modifier.validate_subtitle_layout()
assert report["summary"]["title_count"] >= 3
assert report["summary"]["spatial_collision"] == 0
assert not blocking(report["severity"])
def test_hand_edited_position_is_caught(self, temp_fcpxml):
modifier = FCPXMLModifier(temp_fcpxml)
titles = modifier.generate_dynamic_subtitles("Interview_A", WORDS, WORD_MODE)
def position(el):
for p in el.findall("param"):
if p.get("name") == "Position":
return p
return None
p0 = position(titles[0])
position(titles[1]).set("value", p0.get("value"))
report = modifier.validate_subtitle_layout()
assert report["summary"]["spatial_collision"] >= 1
assert blocking(report["severity"])
+108
View File
@@ -0,0 +1,108 @@
"""Tests for the diarize_media MCP tool (server.handle_diarize_media).
Diarization itself (pyannote.audio) is monkeypatched so these tests run
without the optional [diarization] extra or a HuggingFace token — matching
the existing TestDetectBeatsHandler pattern in test_media_intel.py.
"""
import json
import pytest
def _write_tiny_wav(path: str, seconds: float = 1.0) -> None:
import struct
import wave
n_frames = int(44100 * seconds)
with wave.open(path, "w") as f:
f.setnchannels(1)
f.setsampwidth(2)
f.setframerate(44100)
f.writeframes(struct.pack("<%dh" % n_frames, *([0] * n_frames)))
class TestDiarizeMediaHandler:
async def test_reports_when_pyannote_unavailable(self, tmp_path, monkeypatch):
import server_tools.voice as server_mod
from server import handle_diarize_media
wav = tmp_path / "clip.wav"
_write_tiny_wav(str(wav))
monkeypatch.setattr(
server_mod, "diarization_capability", lambda token: (False, "Diarização indisponível: componente pyannote.audio ausente.")
)
result = await handle_diarize_media({"media_path": str(wav)})
text = result[0].text
assert "indisponível" in text.lower() or "unavailable" in text.lower()
assert "diarization" in text.lower()
async def test_rejects_disallowed_extension(self, tmp_path):
from server import handle_diarize_media
bad = tmp_path / "clip.txt"
bad.write_text("not audio")
with pytest.raises(ValueError):
await handle_diarize_media({"media_path": str(bad)})
async def test_writes_diarization_json_and_reports(self, tmp_path, monkeypatch):
import server_tools._shared as _shared_mod
import server_tools.voice as server_mod
from server import handle_diarize_media
wav = tmp_path / "clip.wav"
_write_tiny_wav(str(wav), seconds=2.0)
fake_transcript = {
"language": "en",
"duration": 2.0,
"text": "hello world",
"segments": [
{"text": "hello", "start": 0.0, "end": 1.0},
{"text": "world", "start": 1.0, "end": 2.0},
],
"words": [
{"word": "hello", "start": 0.0, "end": 0.5, "confidence": 0.9},
{"word": "world", "start": 1.0, "end": 1.5, "confidence": 0.9},
],
}
monkeypatch.setattr(_shared_mod, "transcribe", lambda *a, **k: fake_transcript)
monkeypatch.setattr(server_mod, "diarization_capability", lambda token: (True, "ok"))
monkeypatch.setattr(
server_mod,
"diarize",
lambda path, token, num_speakers="": [(0.0, 1.0, "A"), (1.0, 2.0, "B")],
)
result = await handle_diarize_media({"media_path": str(wav), "hf_token": "fake-token"})
text = result[0].text
assert "Speaker" in text or "speaker" in text.lower()
json_path = tmp_path / "clip_diarization.json"
assert str(json_path) in text
data = json.loads(json_path.read_text())
assert len(data["speakers"]) == 2
assert data["words"][0]["speaker_id"] == "SPEAKER_00"
assert data["words"][1]["speaker_id"] == "SPEAKER_01"
async def test_reports_when_diarization_fails(self, tmp_path, monkeypatch):
import server_tools._shared as _shared_mod
import server_tools.voice as server_mod
from server import handle_diarize_media
wav = tmp_path / "clip.wav"
_write_tiny_wav(str(wav))
fake_transcript = {
"language": "en",
"duration": 1.0,
"text": "hi",
"segments": [{"text": "hi", "start": 0.0, "end": 1.0}],
"words": [{"word": "hi", "start": 0.0, "end": 0.5, "confidence": 0.9}],
}
monkeypatch.setattr(_shared_mod, "transcribe", lambda *a, **k: fake_transcript)
monkeypatch.setattr(server_mod, "diarization_capability", lambda token: (True, "ok"))
monkeypatch.setattr(server_mod, "diarize", lambda *a, **k: None)
result = await handle_diarize_media({"media_path": str(wav), "hf_token": "fake-token"})
assert "failed" in result[0].text.lower()
+168 -7
View File
@@ -17,12 +17,27 @@ from pathlib import Path
import pytest
from fcpxml.models import DynamicSubtitleConfig, TimeValue, WordStyle
from fcpxml.models import DynamicSubtitleConfig, TimeValue, WordLook, WordStyle
from fcpxml.parser import parse_fcpxml
from fcpxml.text_layout import POINT_SCALE, REFERENCE_CANVAS_HEIGHT, ink_extent
from fcpxml.text_layout import (
POINT_SCALE,
REFERENCE_BLOCK_LINE_GAP,
REFERENCE_CANVAS_HEIGHT,
TEXT_TEMPLATE_FONT_SCALE,
LayoutBox,
compose_sentence,
ink_extent,
)
from fcpxml.writer import FCPXMLModifier
SAMPLE = Path(__file__).parent.parent / "examples" / "sample.fcpxml"
def font_points(style) -> float:
"""The style's size back in CANVAS POINTS.
The emitted fontSize lives in the template's own space, which is
TEXT_TEMPLATE_FONT_SCALE times bigger than the space positions use, so any
check that mixes the two has to convert first."""
return float(style.get("fontSize")) / TEXT_TEMPLATE_FONT_SCALE
# The earlier rhythm: one title per WORD. The default is now the
# progressive composition (one title per LINE), covered in
@@ -106,7 +121,7 @@ class TestGenerateDynamicSubtitles:
param_names = [p.get("name") for p in title.findall("param")]
assert param_names == [
"Position", "Layout Method", "Left Margin", "Right Margin",
"Position", "Build Out", "Layout Method", "Left Margin", "Right Margin",
"Top Margin", "Bottom Margin", "Alignment", "Line Spacing",
"Auto-Shrink", "Alignment", "Opacity", "Speed", "Custom Speed",
"Apply Speed",
@@ -132,7 +147,8 @@ class TestGenerateDynamicSubtitles:
assert style_def.get("fontFace") == first.face
scale = (1080 * POINT_SCALE) / REFERENCE_CANVAS_HEIGHT
assert style_def.get("fontSize") == str(round(first.font_size * scale))
expected = round(first.font_size * scale) * TEXT_TEMPLATE_FONT_SCALE
assert float(style_def.get("fontSize")) == expected
def test_font_size_scales_with_the_frame(self, temp_fcpxml):
"""A vertical 2160x3840 timeline must get the reference sizes back
@@ -144,7 +160,9 @@ class TestGenerateDynamicSubtitles:
title = modifier.generate_dynamic_subtitles("Interview_A", WORDS, WORD_MODE)[0]
run = title.find("text/text-style")
style_def = title.find(f"text-style-def[@id='{run.get('ref')}']/text-style")
assert style_def.get("fontSize") == str(WordStyle().rhythm[0].font_size)
assert float(style_def.get("fontSize")) == (
WordStyle().rhythm[0].font_size * TEXT_TEMPLATE_FONT_SCALE
)
def test_position_is_keyframed_constant_hold(self, temp_fcpxml):
"""Regression (2026-08-17): position is a STATIC param in the "Text"
@@ -685,7 +703,7 @@ class TestProgressiveComposition:
style = self._style(t)
top, bottom = ink_extent(
t.find("text/text-style").text,
float(style.get("fontSize")),
font_points(style),
font=style.get("font"),
face=style.get("fontFace"),
)
@@ -718,7 +736,7 @@ class TestProgressiveComposition:
style = self._style(t)
top, bottom = ink_extent(
t.find("text/text-style").text,
float(style.get("fontSize")),
font_points(style),
font=style.get("font"),
face=style.get("fontFace"),
)
@@ -762,3 +780,146 @@ class TestProgressiveComposition:
spoken = " ".join(w["word"] for w in words)
emitted = " ".join(t.find("text/text-style").text for t in titles)
assert emitted == spoken, "no spoken word may be dropped or reordered"
class TestTemplateFontScale:
"""The "Text" template sizes type in frame pixels but positions in canvas
points, so the emitted fontSize must be converted or the block renders in
the right place at half the chosen size."""
STYLE = WordStyle(
emphasis_look=WordLook(200, "1 1 1 1", font="Georgia", kerning=3.0),
body_look=WordLook(100, "1 1 1 1", font="Helvetica Neue", kerning=3.0),
)
def _styles(self, titles):
return [t.find(".//text-style-def/text-style") for t in titles]
def _generate(self, path, **kwargs):
modifier = FCPXMLModifier(path)
titles = modifier.generate_dynamic_subtitles(
"Interview_A", WORDS,
DynamicSubtitleConfig(style=self.STYLE, **kwargs),
)
return titles
def test_emitted_size_is_the_layout_size_times_the_template_scale(self, temp_fcpxml):
unscaled = self._generate(temp_fcpxml, text_scale=1.0)
scaled = self._generate(temp_fcpxml, text_scale=TEXT_TEMPLATE_FONT_SCALE)
for plain, big in zip(self._styles(unscaled), self._styles(scaled)):
assert float(big.get("fontSize")) == (
float(plain.get("fontSize")) * TEXT_TEMPLATE_FONT_SCALE
)
def test_default_config_applies_the_template_scale(self, temp_fcpxml):
assert DynamicSubtitleConfig().text_scale == TEXT_TEMPLATE_FONT_SCALE
default = self._styles(self._generate(temp_fcpxml))
unscaled = self._styles(self._generate(temp_fcpxml, text_scale=1.0))
assert [s.get("fontSize") for s in default] != [
s.get("fontSize") for s in unscaled
]
def test_kerning_scales_with_the_font_size(self, temp_fcpxml):
"""Kerning is in font units too — leaving it behind would tighten the
letter spacing to half as the type doubled."""
unscaled = self._styles(self._generate(temp_fcpxml, text_scale=1.0))
scaled = self._styles(self._generate(temp_fcpxml, text_scale=2.0))
for plain, big in zip(unscaled, scaled):
if plain.get("kerning"):
assert float(big.get("kerning")) == float(plain.get("kerning")) * 2
def _positions(self, titles):
return [
tuple(float(v) for v in t.find(
"param[@name='Position']").get("value").split())
for t in titles
]
def test_position_is_converted_with_the_type(self, temp_fcpxml):
"""The template reads fontSize and Position in the SAME space, so the
conversion has to reach both. Scaling only the type leaves the block at
the old spread with twice the type in it, and the lines collide."""
plain = self._generate(temp_fcpxml, text_scale=1.0)
big = self._generate(temp_fcpxml, text_scale=2.0)
for (x, y), (x2, y2) in zip(self._positions(plain), self._positions(big)):
assert (x2, y2) == pytest.approx((x * 2, y * 2), rel=1e-4, abs=0.01)
def test_type_and_spacing_keep_their_ratio_at_any_scale(self, temp_fcpxml):
"""The invariant that broke in the field: the distance between two
lines, measured in font sizes, must not depend on the scale."""
ratios = []
for scale in (1.0, 2.0, 3.5):
titles = self._generate(temp_fcpxml, text_scale=scale)
ys = [y for _, y in self._positions(titles)]
sizes = [float(s.get("fontSize")) for s in self._styles(titles)]
ratios.append([
(a - b) / size
for a, b, size in zip(ys, ys[1:], sizes)
])
for other in ratios[1:]:
assert other == pytest.approx(ratios[0], rel=1e-4, abs=1e-4)
class TestLineGap:
"""The air between stacked lines is a design choice, negative included."""
STYLE = WordStyle(
emphasis_look=WordLook(200, "1 1 1 1", font="Georgia", kerning=0.0),
body_look=WordLook(100, "1 1 1 1", font="Helvetica Neue", kerning=0.0),
)
PHRASE = "eu tinha muita dificuldade de encontrar roupa"
def _blocks(self, gap):
"""Compose in a band tall enough to hold every line, so the gap is the
only thing that changes — a short band would also change how many
lines fit, which is a different effect."""
words = [
{"word": w, "start": i * 0.3, "end": i * 0.3 + 0.3}
for i, w in enumerate(self.PHRASE.split())
]
box = LayoutBox(width=1080 * 0.92, height=100_000, center_y=0)
return compose_sentence(words, self.STYLE, box, line_gap=gap).blocks
def test_default_matches_the_reference_gap(self):
assert DynamicSubtitleConfig().line_gap == REFERENCE_BLOCK_LINE_GAP
def test_each_step_changes_by_exactly_the_gap(self):
"""The stack places ink boxes edge to edge, so the gap is the whole
distance between two lines beyond their own ink."""
zero = [b.y for b in self._blocks(0.0)]
loose = [b.y for b in self._blocks(50.0)]
assert len(zero) == len(loose) >= 2
for plain, spaced in zip(
[a - b for a, b in zip(zero, zero[1:])],
[a - b for a, b in zip(loose, loose[1:])],
):
assert spaced == pytest.approx(plain + 50.0)
def test_a_negative_gap_overlaps_by_exactly_that_much(self):
"""Negative is a supported look, not a failure: -40 tucks each line 40
points into the one above rather than colliding by some amount the
caller cannot predict."""
zero = [b.y for b in self._blocks(0.0)]
tucked = [b.y for b in self._blocks(-40.0)]
assert len(zero) == len(tucked) >= 2
for plain, tight in zip(
[a - b for a, b in zip(zero, zero[1:])],
[a - b for a, b in zip(tucked, tucked[1:])],
):
assert tight == pytest.approx(plain - 40.0)
def test_the_gap_reaches_the_generated_titles(self, temp_fcpxml):
"""The config field has to survive the trip to the XML."""
def spread(gap):
modifier = FCPXMLModifier(temp_fcpxml)
titles = modifier.generate_dynamic_subtitles(
"Interview_A", WORDS,
DynamicSubtitleConfig(style=self.STYLE, line_gap=gap, text_scale=1.0),
)
ys = [
float(t.find("param[@name='Position']").get("value").split()[1])
for t in titles
]
return max(ys) - min(ys)
assert spread(0.0) < spread(80.0)
+112
View File
@@ -0,0 +1,112 @@
"""Tests for fcpxml/emphasis.py — the emphasis index (pure, no audio needed)."""
from fcpxml.emphasis import (
EmphasisWeights,
annotate_emphasis,
compute_emphasis,
pause_weight,
)
def test_zero_input_gives_zero_score():
assert compute_emphasis(0.0, 0.0, 0.0, 0.0, 0.0) == 0.0
def test_max_input_gives_max_score():
# pause sits inside the "dramatic beat" window, not a scene-change gap
score = compute_emphasis(
energy=1.0, pitch_delta=1.0, rate_delta=1.0, pause_before=1.5, word_duration=10.0
)
assert score == 1.0
def test_negative_pitch_delta_uses_magnitude():
a = compute_emphasis(0.0, pitch_delta=0.5, rate_delta=0.0, pause_before=0.0, word_duration=0.0)
b = compute_emphasis(0.0, pitch_delta=-0.5, rate_delta=0.0, pause_before=0.0, word_duration=0.0)
assert a == b > 0.0
def test_duration_saturates_beyond_cap():
at_cap = compute_emphasis(0.0, 0.0, 0.0, 0.0, word_duration=1.0, max_duration=1.0)
beyond = compute_emphasis(0.0, 0.0, 0.0, 0.0, word_duration=30.0, max_duration=1.0)
assert at_cap == beyond
class TestPauseWeight:
"""A long gap is a scene change, not emphasis — it must not outrank a
word the speaker actually hit hard. Grounded in real footage where 6-9s
gaps were topping the emphasis ranking."""
def test_dramatic_beat_counts_fully(self):
assert pause_weight(1.5, max_pause=1.5) == 1.0
def test_short_beat_counts_proportionally(self):
assert pause_weight(0.75, max_pause=1.5) == 0.5
def test_long_gap_is_ignored(self):
assert pause_weight(8.7, max_pause=1.5, ignore_above=3.0) == 0.0
def test_no_pause_is_zero(self):
assert pause_weight(0.0) == 0.0
def test_cutoff_can_be_disabled(self):
assert pause_weight(30.0, max_pause=1.5, ignore_above=0.0) == 1.0
def test_scene_change_scores_below_a_loud_word(self):
gap = compute_emphasis(0.25, 0.04, 0.0, pause_before=6.2, word_duration=0.5)
loud = compute_emphasis(1.00, 0.28, 0.0, pause_before=1.9, word_duration=0.2)
assert loud > gap
def test_higher_energy_weight_increases_energy_contribution():
low_weight = EmphasisWeights(energy=0.1, pitch_variation=0.0, rate_variation=0.0, pause_before=0.0, duration=0.0)
high_weight = EmphasisWeights(energy=1.0, pitch_variation=0.0, rate_variation=0.0, pause_before=0.0, duration=0.0)
# energy is the only nonzero factor for both weight sets, so normalized
# score should be identical regardless of the absolute weight value.
score_low = compute_emphasis(0.6, 0.0, 0.0, 0.0, 0.0, weights=low_weight)
score_high = compute_emphasis(0.6, 0.0, 0.0, 0.0, 0.0, weights=high_weight)
assert score_low == score_high
def test_all_zero_weights_returns_zero_not_error():
zero_weights = EmphasisWeights(0.0, 0.0, 0.0, 0.0, 0.0)
assert compute_emphasis(1.0, 1.0, 1.0, 1.0, 1.0, weights=zero_weights) == 0.0
def test_weights_round_trip_dict():
w = EmphasisWeights(energy=0.4, pitch_variation=0.3, rate_variation=0.1, pause_before=0.1, duration=0.1)
restored = EmphasisWeights.from_dict(w.as_dict())
assert restored == w
def test_from_dict_fills_missing_with_defaults():
restored = EmphasisWeights.from_dict({"energy": 0.9})
defaults = EmphasisWeights()
assert restored.energy == 0.9
assert restored.pitch_variation == defaults.pitch_variation
def test_annotate_emphasis_adds_score_per_word():
words = [
{"word": "hi", "start": 0.0, "end": 0.3, "energy": 0.2, "pitch_delta": 0.1, "rate_delta": 0.1, "pause_before": 0.0},
{"word": "WOW", "start": 1.0, "end": 1.5, "energy": 0.9, "pitch_delta": 0.8, "rate_delta": 0.7, "pause_before": 1.0},
]
annotated = annotate_emphasis(words)
assert len(annotated) == 2
assert all("emphasis" in w for w in annotated)
assert annotated[1]["emphasis"] > annotated[0]["emphasis"]
def test_annotate_emphasis_does_not_mutate_input():
words = [{"word": "hi", "start": 0.0, "end": 0.3, "energy": 0.5}]
annotate_emphasis(words)
assert "emphasis" not in words[0]
def test_annotate_emphasis_derives_duration_from_start_end():
"""A word with no explicit "duration" key gets it from end - start."""
base = {"energy": 0.0, "pitch_delta": 0.0, "rate_delta": 0.0, "pause_before": 0.0}
derived = annotate_emphasis([{"word": "hi", "start": 1.0, "end": 1.5, **base}], max_duration=0.5)
explicit = annotate_emphasis([{"word": "hi", "start": 0.0, "end": 0.0, "duration": 0.5, **base}], max_duration=0.5)
# duration is the only nonzero factor in both, so scores must match
assert derived[0]["emphasis"] == explicit[0]["emphasis"] > 0.0
+1 -1
View File
@@ -518,7 +518,7 @@ class TestDetectBeatsHandler:
await handle_detect_beats({"media_path": str(bad)})
async def test_reports_when_librosa_unavailable(self, tmp_path, monkeypatch):
import server as server_mod
import server_tools.qc as server_mod
from server import handle_detect_beats
wav = tmp_path / "song.wav"
+63
View File
@@ -0,0 +1,63 @@
"""Tests for `output_dir` routing — the app's "Pasta do projeto" promise.
The setting is documented in the UI as "everything generated is saved in
here". It used to be applied as a sandbox anchor only, while the filename
was still derived in the INPUT's directory — so any call whose output_dir
differed from the input's folder failed its own anchor check.
"""
import pytest
from server_tools._shared import _resolve_io_paths
@pytest.fixture
def project(tmp_path):
"""An .fcpxml in one folder, with a separate chosen output folder."""
source_dir = tmp_path / "media"
source_dir.mkdir()
fcpxml = source_dir / "Projeto.fcpxml"
fcpxml.write_text("<fcpxml version='1.13'/>")
chosen = tmp_path / "pasta do projeto"
chosen.mkdir()
return fcpxml, chosen
class TestOutputDirRouting:
def test_output_lands_in_the_chosen_folder(self, project):
fcpxml, chosen = project
_, output_path = _resolve_io_paths({"filepath": str(fcpxml), "output_dir": str(chosen)}, "_voice_edit")
assert output_path.startswith(str(chosen))
def test_filename_keeps_the_suffix_convention(self, project):
fcpxml, chosen = project
_, output_path = _resolve_io_paths({"filepath": str(fcpxml), "output_dir": str(chosen)}, "_voice_edit")
assert output_path.endswith("Projeto_voice_edit.fcpxml")
def test_cross_directory_call_does_not_raise(self, project):
"""The regression: output_dir different from the input folder used to
raise "output path escapes allowed directory" every single time."""
fcpxml, chosen = project
_resolve_io_paths({"filepath": str(fcpxml), "output_dir": str(chosen)}, "_dynamic_subtitles")
def test_without_output_dir_it_still_writes_beside_the_input(self, project):
fcpxml, _ = project
_, output_path = _resolve_io_paths({"filepath": str(fcpxml)}, "_modified")
assert output_path == str(fcpxml.parent / "Projeto_modified.fcpxml")
def test_explicit_output_path_still_wins(self, project):
fcpxml, chosen = project
target = chosen / "nome escolhido.fcpxml"
_, output_path = _resolve_io_paths(
{"filepath": str(fcpxml), "output_dir": str(chosen), "output_path": str(target)}, "_voice_edit"
)
assert output_path == str(target)
def test_explicit_output_path_outside_the_anchor_is_rejected(self, project, tmp_path):
"""The anchor must keep constraining explicit paths, not just names."""
fcpxml, chosen = project
with pytest.raises(ValueError):
_resolve_io_paths(
{"filepath": str(fcpxml), "output_dir": str(chosen),
"output_path": str(tmp_path / "fora.fcpxml")}, "_voice_edit"
)
+83
View File
@@ -0,0 +1,83 @@
"""Tests for the last-project settings — the folder/file the app reopens with.
The config file (~/.fcp-mcp-server/config.json) is redirected to a tmp_path
so these never touch the developer's real settings.
"""
import json
import pytest
from fcpxml import model_manager
@pytest.fixture(autouse=True)
def isolated_config(tmp_path, monkeypatch):
"""Point model_manager's config file at a throwaway directory."""
monkeypatch.setattr(model_manager, "_CONFIG_DIR", tmp_path)
monkeypatch.setattr(model_manager, "_CONFIG_FILE", tmp_path / "config.json")
return tmp_path / "config.json"
class TestLoadProjectConfig:
def test_defaults_when_nothing_stored(self):
assert model_manager.load_project_config() == model_manager.DEFAULT_PROJECT_CONFIG
def test_non_dict_stored_value_falls_back(self, isolated_config):
isolated_config.write_text(json.dumps({"project": "nonsense"}))
assert model_manager.load_project_config() == model_manager.DEFAULT_PROJECT_CONFIG
def test_reads_back_what_was_stored(self, tmp_path):
folder = tmp_path / "03 - Mastopexia"
folder.mkdir()
model_manager.save_project_config(folder=str(folder))
assert model_manager.load_project_config()["folder"] == str(folder)
def test_path_that_no_longer_exists_comes_back_empty(self, isolated_config, tmp_path):
"""An unmounted volume must degrade to "nothing selected", not a dead path."""
isolated_config.write_text(json.dumps({"project": {"folder": str(tmp_path / "gone")}}))
assert model_manager.load_project_config()["folder"] == ""
def test_non_string_stored_value_is_ignored(self, isolated_config):
isolated_config.write_text(json.dumps({"project": {"folder": 42}}))
assert model_manager.load_project_config()["folder"] == ""
class TestSaveProjectConfig:
def test_omitted_field_keeps_its_current_value(self, tmp_path):
folder = tmp_path / "projeto"
folder.mkdir()
project = tmp_path / "projeto" / "Mastopexia.fcpxml"
project.write_text("<fcpxml/>")
model_manager.save_project_config(folder=str(folder), file=str(project))
model_manager.save_project_config(file=str(project))
assert model_manager.load_project_config()["folder"] == str(folder)
def test_empty_string_clears_a_field(self, tmp_path):
folder = tmp_path / "projeto"
folder.mkdir()
model_manager.save_project_config(folder=str(folder))
model_manager.save_project_config(folder="")
assert model_manager.load_project_config()["folder"] == ""
def test_home_relative_path_is_expanded(self, isolated_config):
model_manager.save_project_config(folder="~/Movies")
stored = json.loads(isolated_config.read_text())["project"]["folder"]
assert not stored.startswith("~")
def test_does_not_disturb_other_config_sections(self, isolated_config, tmp_path):
isolated_config.write_text(json.dumps({"language": "pt", "selected_model": "large-v3"}))
folder = tmp_path / "projeto"
folder.mkdir()
model_manager.save_project_config(folder=str(folder))
data = json.loads(isolated_config.read_text())
assert data["language"] == "pt"
assert data["selected_model"] == "large-v3"
def test_returns_the_merged_config(self, tmp_path):
folder = tmp_path / "projeto"
folder.mkdir()
assert model_manager.save_project_config(folder=str(folder)) == {
"folder": str(folder),
"file": "",
}
@@ -0,0 +1,167 @@
"""Tests for the refine_voice_timeline MCP tool.
The tool exists because emphasis is *relative*: cut the loudest moment of a
recording and every surviving score is still measured against something the
viewer never sees. These tests pin the re-normalization actually happening,
and the times staying in original source seconds so the result can be fed
straight back to apply_voice_actions.
"""
import json
import pytest
from tests.test_voice_features_tool import _write_silent_wav
from tests.test_voice_timeline_tool import _TRANSCRIPT, patched, wav # noqa: F401
_ = _write_silent_wav, _TRANSCRIPT # re-exported fixtures need the imports
async def _build(wav_path):
from server import handle_build_voice_timeline
await handle_build_voice_timeline({"media_path": str(wav_path)})
@pytest.fixture
def two_candidates(monkeypatch):
"""A transcript where BOTH lines carry a content word.
The shared fixture's first line is "isso e" — two function words, which
``suggest_zoom_windows`` skips by design, so it can never produce more
than one candidate to cap.
"""
import fcpxml.voice_timeline as vt
import server_tools._shared as _shared_mod
transcript = {
"language": "pt",
"duration": 4.0,
"text": "cirurgia rapida seguranca total",
"segments": [
{"text": "cirurgia rapida", "start": 0.0, "end": 1.0},
{"text": "seguranca total", "start": 2.0, "end": 4.0},
],
"words": [
{"word": "cirurgia", "start": 0.0, "end": 0.4, "confidence": 0.9},
{"word": "rapida", "start": 0.5, "end": 0.7, "confidence": 0.9},
{"word": "seguranca", "start": 2.0, "end": 2.9, "confidence": 0.9},
{"word": "total", "start": 3.0, "end": 3.5, "confidence": 0.9},
],
}
monkeypatch.setattr(_shared_mod, "transcribe", lambda *a, **k: transcript)
monkeypatch.setattr(vt, "extract_pitch", lambda *a, **k: [(2.4, 260.0), (0.2, 120.0)])
monkeypatch.setattr(vt, "extract_energy", lambda *a, **k: [(2.4, 0.95), (0.2, 0.10)])
class TestRefineVoiceTimelineHandler:
async def test_requires_an_existing_timeline(self, wav): # noqa: F811
from server import handle_refine_voice_timeline
result = await handle_refine_voice_timeline({"media_path": str(wav), "cuts": []})
assert "build_voice_timeline" in result[0].text
async def test_rejects_disallowed_extension(self, tmp_path):
from server import handle_refine_voice_timeline
bad = tmp_path / "clip.txt"
bad.write_text("not audio")
with pytest.raises(ValueError):
await handle_refine_voice_timeline({"media_path": str(bad), "cuts": []})
async def test_compares_raw_against_survivors(self, wav, patched): # noqa: F811
from server import handle_refine_voice_timeline
await _build(wav)
result = await handle_refine_voice_timeline(
{"media_path": str(wav), "cuts": [{"start": 0.0, "end": 1.0}]}
)
text = result[0].text
assert "Survivors only" in text
assert "Average emphasis" in text
async def test_cut_words_are_excluded(self, wav, patched): # noqa: F811
from server import handle_refine_voice_timeline
await _build(wav)
result = await handle_refine_voice_timeline(
{"media_path": str(wav), "cuts": [{"start": 0.0, "end": 1.0}], "save": True}
)
assert "_voice_timeline_refined.json" in result[0].text
data = json.loads(
(wav.parent / "clip_voice_timeline_refined.json").read_text(encoding="utf-8")
)
words = [w["text"] for s in data["segments"] for w in s["words"]]
assert "isso" not in words
assert "seguranca" in words
async def test_times_stay_in_original_source_seconds(self, wav, patched): # noqa: F811
"""A cut at the head must NOT slide the survivors back to zero."""
from server import handle_refine_voice_timeline
await _build(wav)
await handle_refine_voice_timeline(
{"media_path": str(wav), "cuts": [{"start": 0.0, "end": 1.0}], "save": True}
)
data = json.loads(
(wav.parent / "clip_voice_timeline_refined.json").read_text(encoding="utf-8")
)
first = data["segments"][0]["words"][0]
assert first["start"] == pytest.approx(2.0)
async def test_proposes_zoom_candidates(self, wav, patched): # noqa: F811
from server import handle_refine_voice_timeline
await _build(wav)
result = await handle_refine_voice_timeline({"media_path": str(wav), "cuts": []})
assert "Zoom Candidates" in result[0].text
async def test_max_zooms_caps_the_list(self, wav, two_candidates): # noqa: F811
from server import handle_refine_voice_timeline
await _build(wav)
async def zoom_rows(**extra):
result = await handle_refine_voice_timeline(
{"media_path": str(wav), "cuts": [], "min_gap": 0.0, **extra}
)
section = result[0].text.split("## Zoom Candidates", 1)[1]
return [
ln for ln in section.splitlines()
if ln.startswith("| ") and ln.rstrip().endswith("|") and "Start" not in ln
]
assert len(await zoom_rows()) == 2
assert len(await zoom_rows(max_zooms=1)) == 1
async def test_malformed_cut_is_reported_not_raised(self, wav, patched): # noqa: F811
from server import handle_refine_voice_timeline
await _build(wav)
result = await handle_refine_voice_timeline(
{"media_path": str(wav), "cuts": [{"start": 3.0, "end": 1.0}]}
)
text = result[0].text
assert "Rejected cuts" in text
assert "must be after start" in text
async def test_cutting_everything_says_so(self, wav, patched): # noqa: F811
from server import handle_refine_voice_timeline
await _build(wav)
result = await handle_refine_voice_timeline(
{"media_path": str(wav), "cuts": [{"start": 0.0, "end": 60.0}]}
)
assert "removed every word" in result[0].text
class TestRefineVoiceTimelineRegistration:
async def test_tool_is_listed(self):
from server import list_tools
assert "refine_voice_timeline" in {t.name for t in await list_tools()}
async def test_tool_is_dispatched(self):
from server import TOOL_HANDLERS, handle_refine_voice_timeline
assert TOOL_HANDLERS["refine_voice_timeline"] is handle_refine_voice_timeline
+227
View File
@@ -0,0 +1,227 @@
"""Tests for fcpxml/voice_actions.py — the decision contract.
Pure functions over untrusted input (a model's decision list), so these
cover the rejection paths as carefully as the happy path.
"""
import pytest
from fcpxml.voice_actions import (
MAX_TEXT_LENGTH,
VoiceAction,
merge_cut_ranges,
parse_actions,
resolve_actions,
shift_after_cuts,
)
class TestParseActions:
def test_accepts_bare_list(self):
actions, errors = parse_actions([{"kind": "cut", "start": 1.0, "end": 2.0}])
assert len(actions) == 1 and errors == []
def test_accepts_actions_envelope(self):
actions, errors = parse_actions({"actions": [{"kind": "cut", "start": 1.0, "end": 2.0}]})
assert len(actions) == 1 and errors == []
def test_rejects_non_list(self):
actions, errors = parse_actions("cortar tudo")
assert actions == [] and len(errors) == 1
def test_one_bad_row_does_not_discard_the_good_ones(self):
actions, errors = parse_actions([
{"kind": "cut", "start": 1.0, "end": 2.0},
{"kind": "teleport", "start": 3.0, "end": 4.0},
{"kind": "zoom", "start": 5.0, "end": 6.0},
])
assert len(actions) == 2
assert len(errors) == 1 and "teleport" in errors[0]
def test_rejects_unknown_kind(self):
_, errors = parse_actions([{"kind": "explode", "start": 0.0, "end": 1.0}])
assert "explode" in errors[0]
def test_rejects_non_numeric_times(self):
_, errors = parse_actions([{"kind": "cut", "start": "início", "end": 2.0}])
assert "numbers" in errors[0]
def test_rejects_negative_start(self):
_, errors = parse_actions([{"kind": "cut", "start": -1.0, "end": 2.0}])
assert "negative" in errors[0]
def test_rejects_end_before_start(self):
_, errors = parse_actions([{"kind": "cut", "start": 5.0, "end": 2.0}])
assert "must be after" in errors[0]
def test_rejects_zero_length(self):
_, errors = parse_actions([{"kind": "cut", "start": 2.0, "end": 2.0}])
assert errors
def test_rejects_row_that_is_not_an_object(self):
_, errors = parse_actions(["cortar aos 5s"])
assert "expected an object" in errors[0]
def test_kind_is_case_insensitive(self):
actions, _ = parse_actions([{"kind": "ZOOM", "start": 1.0, "end": 2.0}])
assert actions[0].kind == "zoom"
def test_preserves_reason_and_speaker(self):
actions, _ = parse_actions([
{"kind": "zoom", "start": 1.0, "end": 2.0,
"reason": "argumento central", "speaker": "SPEAKER_01"}
])
assert actions[0].reason == "argumento central"
assert actions[0].speaker == "SPEAKER_01"
class TestZoomValidation:
def test_default_scale_when_absent(self):
actions, _ = parse_actions([{"kind": "zoom", "start": 1.0, "end": 2.0}])
assert actions[0].params["scale"] == 1.3
def test_rejects_scale_below_one(self):
_, errors = parse_actions([
{"kind": "zoom", "start": 1.0, "end": 2.0, "params": {"scale": 0.5}}
])
assert "outside" in errors[0]
def test_rejects_absurd_scale(self):
_, errors = parse_actions([
{"kind": "zoom", "start": 1.0, "end": 2.0, "params": {"scale": 50}}
])
assert "outside" in errors[0]
def test_rejects_non_numeric_scale(self):
_, errors = parse_actions([
{"kind": "zoom", "start": 1.0, "end": 2.0, "params": {"scale": "muito"}}
])
assert "must be a number" in errors[0]
class TestTextValidation:
def test_requires_content(self):
_, errors = parse_actions([{"kind": "text", "start": 1.0, "end": 2.0}])
assert "params.content" in errors[0]
def test_rejects_blank_content(self):
_, errors = parse_actions([
{"kind": "text", "start": 1.0, "end": 2.0, "params": {"content": " "}}
])
assert "params.content" in errors[0]
def test_truncates_overlong_content(self):
actions, _ = parse_actions([
{"kind": "text", "start": 1.0, "end": 2.0, "params": {"content": "A" * 500}}
])
assert len(actions[0].params["content"]) == MAX_TEXT_LENGTH
class TestMergeCutRanges:
def test_sorts_and_merges_overlaps(self):
actions = [
VoiceAction("cut", 5.0, 7.0),
VoiceAction("cut", 1.0, 3.0),
VoiceAction("cut", 2.0, 4.0),
]
assert merge_cut_ranges(actions) == [(1.0, 4.0), (5.0, 7.0)]
def test_merges_touching_ranges(self):
actions = [VoiceAction("cut", 1.0, 2.0), VoiceAction("cut", 2.0, 3.0)]
assert merge_cut_ranges(actions) == [(1.0, 3.0)]
def test_ignores_non_cut_actions(self):
assert merge_cut_ranges([VoiceAction("zoom", 1.0, 2.0)]) == []
class TestShiftAfterCuts:
def test_time_before_any_cut_is_unchanged(self):
assert shift_after_cuts(0.5, [(2.0, 4.0)]) == 0.5
def test_time_after_a_cut_moves_earlier(self):
assert shift_after_cuts(6.0, [(2.0, 4.0)]) == pytest.approx(4.0)
def test_time_inside_a_cut_is_dropped(self):
assert shift_after_cuts(3.0, [(2.0, 4.0)]) is None
def test_multiple_cuts_accumulate(self):
cuts = [(1.0, 2.0), (5.0, 7.0)]
assert shift_after_cuts(10.0, cuts) == pytest.approx(7.0)
def test_no_cuts_is_identity(self):
assert shift_after_cuts(3.0, []) == 3.0
def test_boundary_start_of_cut_is_inside(self):
assert shift_after_cuts(2.0, [(2.0, 4.0)]) is None
def test_boundary_end_of_cut_survives(self):
assert shift_after_cuts(4.0, [(2.0, 4.0)]) == pytest.approx(2.0)
class TestResolveActions:
def test_zoom_after_a_cut_is_moved_earlier(self):
actions = [VoiceAction("cut", 2.0, 4.0), VoiceAction("zoom", 6.0, 7.0)]
cuts, placed, dropped = resolve_actions(actions)
assert cuts == [(2.0, 4.0)]
assert dropped == []
assert placed[0].start == pytest.approx(4.0)
assert placed[0].end == pytest.approx(5.0)
def test_zoom_inside_a_cut_is_dropped_not_slid(self):
actions = [VoiceAction("cut", 2.0, 8.0), VoiceAction("zoom", 3.0, 4.0)]
_, placed, dropped = resolve_actions(actions)
assert placed == []
assert len(dropped) == 1
def test_zoom_straddling_a_cut_edge_is_dropped(self):
actions = [VoiceAction("cut", 4.0, 8.0), VoiceAction("zoom", 3.0, 5.0)]
_, placed, dropped = resolve_actions(actions)
assert placed == [] and len(dropped) == 1
def test_cuts_are_not_returned_as_placed(self):
_, placed, _ = resolve_actions([VoiceAction("cut", 1.0, 2.0)])
assert placed == []
def test_without_cuts_everything_keeps_its_time(self):
actions = [VoiceAction("zoom", 3.0, 4.0), VoiceAction("text", 5.0, 6.0)]
cuts, placed, dropped = resolve_actions(actions)
assert cuts == [] and dropped == []
assert [(a.start, a.end) for a in placed] == [(3.0, 4.0), (5.0, 6.0)]
def test_params_survive_the_shift(self):
actions = [
VoiceAction("cut", 1.0, 2.0),
VoiceAction("text", 5.0, 6.0, params={"content": "SEGURANÇA"}),
]
_, placed, _ = resolve_actions(actions)
assert placed[0].params["content"] == "SEGURANÇA"
class TestMarkersSurviveCutEdges:
"""A marker is a point in time, not a span. The markers worth keeping are
precisely the ones flagging a join, which sit against a cut edge — so
requiring their nominal end to survive would drop exactly those."""
def test_marker_at_a_cut_edge_survives(self):
actions = [VoiceAction("cut", 21.9, 127.6), VoiceAction("marker", 21.85, 22.0)]
_, placed, dropped = resolve_actions(actions)
assert dropped == []
assert placed[0].kind == "marker"
assert placed[0].start == pytest.approx(21.85)
def test_marker_inside_removed_material_is_still_dropped(self):
actions = [VoiceAction("cut", 20.0, 100.0), VoiceAction("marker", 50.0, 50.2)]
_, placed, dropped = resolve_actions(actions)
assert placed == [] and len(dropped) == 1
def test_marker_keeps_its_length_after_shifting(self):
actions = [VoiceAction("cut", 0.0, 10.0), VoiceAction("marker", 20.0, 20.5)]
_, placed, _ = resolve_actions(actions)
assert placed[0].start == pytest.approx(10.0)
assert placed[0].duration == pytest.approx(0.5)
def test_zoom_straddling_an_edge_is_still_dropped(self):
"""Only markers get the point-action treatment — a span must fit."""
actions = [VoiceAction("cut", 21.9, 127.6), VoiceAction("zoom", 21.0, 22.5)]
_, placed, dropped = resolve_actions(actions)
assert placed == [] and len(dropped) == 1
+267
View File
@@ -0,0 +1,267 @@
"""Tests for the apply_voice_actions MCP tool — decisions -> real FCPXML.
Uses an inline fixture rather than examples/sample.fcpxml so the source
windows are explicit and the assertions can be exact.
"""
import shutil
import pytest
from fcpxml.safe_xml import safe_parse
_FIXTURE = """<?xml version="1.0" encoding="UTF-8"?>
<fcpxml version="1.13">
<resources>
<format id="r1" name="FFVideoFormat1080p30" frameDuration="100/3000s" width="1920" height="1080"/>
<asset id="a1" name="entrevista" start="0s" duration="600/30s" hasVideo="1" hasAudio="1" format="r1">
<media-rep kind="original-media" src="file:///media/entrevista.mov"/>
</asset>
</resources>
<library>
<event name="Ev">
<project name="Proj">
<sequence format="r1" duration="600/30s" tcStart="0s">
<spine>
<asset-clip name="entrevista" ref="a1" offset="0s" start="0s" duration="600/30s"/>
</spine>
</sequence>
</project>
</event>
</library>
</fcpxml>
"""
@pytest.fixture
def project(tmp_path):
path = tmp_path / "proj.fcpxml"
path.write_text(_FIXTURE)
return path
def _out(project):
return project.with_name("proj_voice_edit.fcpxml")
class TestApplyVoiceActionsHandler:
async def test_no_actions_reports_instead_of_writing(self, project):
from server import handle_apply_voice_actions
result = await handle_apply_voice_actions({"filepath": str(project)})
assert "nothing to apply" in result[0].text.lower()
assert not _out(project).exists()
async def test_all_invalid_actions_writes_nothing(self, project):
from server import handle_apply_voice_actions
result = await handle_apply_voice_actions({
"filepath": str(project),
"actions": [{"kind": "teleport", "start": 1.0, "end": 2.0}],
})
assert "No valid actions" in result[0].text
assert "teleport" in result[0].text
assert not _out(project).exists()
async def test_applies_zoom_into_the_hosting_clip(self, project):
from server import handle_apply_voice_actions
result = await handle_apply_voice_actions({
"filepath": str(project),
"actions": [{
"kind": "zoom", "start": 5.0, "end": 6.0,
"params": {"scale": 1.4}, "reason": "argumento central",
}],
})
assert "argumento central" in result[0].text
tree = safe_parse(str(_out(project)))
transforms = tree.getroot().findall(".//adjust-transform")
assert len(transforms) == 1
async def test_applies_text_title(self, project):
from server import handle_apply_voice_actions
await handle_apply_voice_actions({
"filepath": str(project),
"actions": [{
"kind": "text", "start": 3.0, "end": 4.0,
"params": {"content": "SEGURANÇA"},
}],
})
titles = safe_parse(str(_out(project))).getroot().findall(".//title")
assert len(titles) == 1
texts = [t.text for t in titles[0].iter() if t.text]
assert any("SEGURANÇA" in t for t in texts)
async def test_applies_marker(self, project):
from server import handle_apply_voice_actions
await handle_apply_voice_actions({
"filepath": str(project),
"actions": [{
"kind": "marker", "start": 2.0, "end": 2.5, "reason": "virada",
}],
})
markers = safe_parse(str(_out(project))).getroot().findall(".//marker")
assert len(markers) == 1
async def test_cut_shortens_the_timeline(self, project):
from server import handle_apply_voice_actions
before = safe_parse(str(project)).getroot().find(".//asset-clip").get("duration")
result = await handle_apply_voice_actions({
"filepath": str(project),
"actions": [{"kind": "cut", "start": 5.0, "end": 10.0, "reason": "digressão"}],
})
assert "Cuts applied" in result[0].text
clips = safe_parse(str(_out(project))).getroot().findall(".//asset-clip")
total = sum(
int(c.get("duration").split("/")[0]) / int(c.get("duration").split("/")[1].rstrip("s"))
for c in clips
)
original = int(before.split("/")[0]) / int(before.split("/")[1].rstrip("s"))
assert total < original
async def test_action_inside_a_cut_is_dropped_and_reported(self, project):
from server import handle_apply_voice_actions
result = await handle_apply_voice_actions({
"filepath": str(project),
"actions": [
{"kind": "cut", "start": 4.0, "end": 12.0},
{"kind": "zoom", "start": 6.0, "end": 7.0, "reason": "some no material cortado"},
],
})
text = result[0].text
assert "Dropped" in text
assert "some no material cortado" in text
assert safe_parse(str(_out(project))).getroot().findall(".//adjust-transform") == []
async def test_action_beyond_the_media_is_reported_not_silent(self, project):
from server import handle_apply_voice_actions
result = await handle_apply_voice_actions({
"filepath": str(project),
"actions": [{"kind": "zoom", "start": 500.0, "end": 501.0}],
})
assert "Not placed" in result[0].text
assert "outside the edited timeline" in result[0].text
async def test_original_file_is_untouched(self, project):
from server import handle_apply_voice_actions
original = project.read_text()
await handle_apply_voice_actions({
"filepath": str(project),
"actions": [{"kind": "zoom", "start": 5.0, "end": 6.0}],
})
assert project.read_text() == original
async def test_mixed_valid_and_invalid_applies_the_valid_ones(self, project):
from server import handle_apply_voice_actions
result = await handle_apply_voice_actions({
"filepath": str(project),
"actions": [
{"kind": "zoom", "start": 5.0, "end": 6.0},
{"kind": "zoom", "start": 8.0, "end": 9.0, "params": {"scale": 99}},
],
})
assert "Rejected" in result[0].text
assert len(safe_parse(str(_out(project))).getroot().findall(".//adjust-transform")) == 1
async def test_respects_explicit_output_path(self, project, tmp_path):
from server import handle_apply_voice_actions
target = tmp_path / "custom.fcpxml"
await handle_apply_voice_actions({
"filepath": str(project),
"actions": [{"kind": "marker", "start": 1.0, "end": 2.0}],
"output_path": str(target),
})
assert target.exists()
async def test_output_is_valid_parseable_fcpxml(self, project):
from server import handle_apply_voice_actions
await handle_apply_voice_actions({
"filepath": str(project),
"actions": [
{"kind": "cut", "start": 2.0, "end": 4.0},
{"kind": "zoom", "start": 10.0, "end": 11.0},
{"kind": "text", "start": 12.0, "end": 13.0, "params": {"content": "OK"}},
],
})
root = safe_parse(str(_out(project))).getroot()
assert root.tag == "fcpxml"
assert root.find(".//spine") is not None
class TestSampleFixtureStillParses:
"""The applier must not corrupt a real-world document."""
async def test_real_sample_survives_a_zoom(self, tmp_path):
from server import handle_apply_voice_actions
src = "examples/sample.fcpxml"
target = tmp_path / "sample.fcpxml"
shutil.copy(src, target)
result = await handle_apply_voice_actions({
"filepath": str(target),
"actions": [{"kind": "marker", "start": 1.0, "end": 2.0, "reason": "teste"}],
})
assert "Voice Actions Applied" in result[0].text
class TestPlacementsLandOnTheRightPieceAfterCuts:
"""Cutting splits a clip into same-named pieces. Placing before cutting
duplicated the zoom onto every piece and lost markers outright; a
name-based lookup afterwards would always resolve to the first piece.
Both bugs shipped past the suite and only showed up on real footage."""
async def test_zoom_lands_on_exactly_one_piece(self, project):
from server import handle_apply_voice_actions
await handle_apply_voice_actions({
"filepath": str(project),
"actions": [
{"kind": "cut", "start": 2.0, "end": 5.0},
{"kind": "cut", "start": 8.0, "end": 12.0},
{"kind": "zoom", "start": 15.0, "end": 16.0, "params": {"scale": 1.2}},
],
})
root = safe_parse(str(_out(project))).getroot()
assert len(root.findall(".//spine/asset-clip")) == 3
assert len(root.findall(".//adjust-transform")) == 1
async def test_zoom_lands_on_the_last_piece_not_the_first(self, project):
from server import handle_apply_voice_actions
await handle_apply_voice_actions({
"filepath": str(project),
"actions": [
{"kind": "cut", "start": 2.0, "end": 5.0},
{"kind": "zoom", "start": 15.0, "end": 16.0},
],
})
clips = safe_parse(str(_out(project))).getroot().findall(".//spine/asset-clip")
# the zoom is at 15s source -> 12s after a 3s cut, i.e. the 2nd piece
assert clips[0].find("adjust-transform") is None
assert clips[1].find("adjust-transform") is not None
async def test_markers_survive_the_cut(self, project):
from server import handle_apply_voice_actions
result = await handle_apply_voice_actions({
"filepath": str(project),
"actions": [
{"kind": "cut", "start": 5.0, "end": 10.0},
{"kind": "marker", "start": 4.9, "end": 5.05, "reason": "emenda"},
{"kind": "marker", "start": 15.0, "end": 15.2, "reason": "depois"},
],
})
assert "Dropped" not in result[0].text
markers = safe_parse(str(_out(project))).getroot().findall(".//marker")
assert len(markers) == 2
+111
View File
@@ -0,0 +1,111 @@
"""Tests for the Voice Analysis settings — persistence + MCP config tools.
The config file (~/.fcp-mcp-server/config.json) is redirected to a tmp_path
so these never touch the developer's real settings.
"""
import json
import pytest
from fcpxml import model_manager
@pytest.fixture(autouse=True)
def isolated_config(tmp_path, monkeypatch):
"""Point model_manager's config file at a throwaway directory."""
monkeypatch.setattr(model_manager, "_CONFIG_DIR", tmp_path)
monkeypatch.setattr(model_manager, "_CONFIG_FILE", tmp_path / "config.json")
return tmp_path / "config.json"
class TestLoadVoiceAnalysisConfig:
def test_defaults_when_nothing_stored(self):
cfg = model_manager.load_voice_analysis_config()
assert cfg == model_manager.DEFAULT_VOICE_ANALYSIS_CONFIG
def test_defaults_are_not_shared_mutable_state(self):
cfg = model_manager.load_voice_analysis_config()
cfg["emphasis_weights"]["energy"] = 0.99
fresh = model_manager.load_voice_analysis_config()
assert fresh["emphasis_weights"]["energy"] == 0.30
def test_malformed_stored_values_fall_back_to_defaults(self, isolated_config):
isolated_config.write_text(json.dumps({"voice_analysis": {"energy_threshold": "loud"}}))
cfg = model_manager.load_voice_analysis_config()
assert cfg["energy_threshold"] == 0.5
def test_non_dict_stored_value_falls_back(self, isolated_config):
isolated_config.write_text(json.dumps({"voice_analysis": "nonsense"}))
assert model_manager.load_voice_analysis_config() == model_manager.DEFAULT_VOICE_ANALYSIS_CONFIG
def test_thresholds_are_clamped_to_unit_range(self, isolated_config):
isolated_config.write_text(
json.dumps({"voice_analysis": {"energy_threshold": 5.0, "emphasis_floor": -2.0}})
)
cfg = model_manager.load_voice_analysis_config()
assert cfg["energy_threshold"] == 1.0
assert cfg["emphasis_floor"] == 0.0
class TestSaveVoiceAnalysisConfig:
def test_saves_and_reloads(self):
model_manager.save_voice_analysis_config(energy_threshold=0.7, emotion_enabled=True)
cfg = model_manager.load_voice_analysis_config()
assert cfg["energy_threshold"] == 0.7
assert cfg["emotion_enabled"] is True
def test_omitted_fields_keep_current_value(self):
model_manager.save_voice_analysis_config(energy_threshold=0.7)
model_manager.save_voice_analysis_config(emphasis_floor=0.9)
cfg = model_manager.load_voice_analysis_config()
assert cfg["energy_threshold"] == 0.7
assert cfg["emphasis_floor"] == 0.9
def test_partial_weight_update_keeps_other_weights(self):
model_manager.save_voice_analysis_config(emphasis_weights={"energy": 0.55})
weights = model_manager.load_voice_analysis_config()["emphasis_weights"]
assert weights["energy"] == 0.55
assert weights["pitch_variation"] == 0.25
def test_unknown_weight_key_is_ignored(self):
model_manager.save_voice_analysis_config(emphasis_weights={"loudness": 9.0})
weights = model_manager.load_voice_analysis_config()["emphasis_weights"]
assert "loudness" not in weights
def test_does_not_clobber_unrelated_config_keys(self):
model_manager.save_hf_token("tok123")
model_manager.save_voice_analysis_config(energy_threshold=0.7)
assert model_manager.load_hf_token() == "tok123"
def test_returns_full_merged_config(self):
returned = model_manager.save_voice_analysis_config(energy_threshold=0.7)
assert returned == model_manager.load_voice_analysis_config()
class TestVoiceAnalysisConfigTools:
async def test_get_reports_current_settings(self):
from server import handle_get_voice_analysis_config
result = await handle_get_voice_analysis_config({})
text = result[0].text
assert "Voice Analysis Settings" in text
assert "Emphasis Weights" in text
async def test_save_persists_and_echoes_back(self):
from server import handle_save_voice_analysis_config
result = await handle_save_voice_analysis_config(
{"energy_threshold": 0.8, "emotion_enabled": True}
)
assert "0.80" in result[0].text
cfg = model_manager.load_voice_analysis_config()
assert cfg["energy_threshold"] == 0.8
assert cfg["emotion_enabled"] is True
async def test_save_with_no_arguments_is_a_noop(self):
from server import handle_save_voice_analysis_config
before = model_manager.load_voice_analysis_config()
await handle_save_voice_analysis_config({})
assert model_manager.load_voice_analysis_config() == before
+137
View File
@@ -0,0 +1,137 @@
"""Tests for fcpxml/voice_features.py — acoustic features.
The pure helpers (speech rate, pauses, window averaging) need no audio.
The librosa-backed extractors are skipped when the optional [intelligence]
extra is absent, matching the pattern in test_media_intel.py.
"""
import math
import struct
import wave
import pytest
from fcpxml.voice_features import (
compute_pauses,
compute_speech_rate,
extract_energy,
extract_pitch,
features_capability,
word_pitch_energy,
)
try:
import librosa # noqa: F401
LIBROSA = True
except ImportError:
LIBROSA = False
def _write_tone_wav(path: str, hz: float = 220.0, seconds: float = 2.0, rate: int = 22050) -> None:
n = int(rate * seconds)
frames = [int(20000 * math.sin(2 * math.pi * hz * i / rate)) for i in range(n)]
with wave.open(path, "w") as f:
f.setnchannels(1)
f.setsampwidth(2)
f.setframerate(rate)
f.writeframes(struct.pack("<%dh" % n, *frames))
class TestComputePauses:
def test_first_word_pause_is_time_from_zero(self):
words = [{"start": 1.5, "end": 2.0}]
assert compute_pauses(words) == [1.5]
def test_gap_between_words(self):
words = [{"start": 0.0, "end": 1.0}, {"start": 2.5, "end": 3.0}]
assert compute_pauses(words) == [0.0, 1.5]
def test_overlapping_words_clamp_to_zero(self):
words = [{"start": 0.0, "end": 2.0}, {"start": 1.0, "end": 3.0}]
assert compute_pauses(words) == [0.0, 0.0]
def test_empty_words(self):
assert compute_pauses([]) == []
class TestComputeSpeechRate:
def test_rate_counts_words_in_trailing_window(self):
# 3 words within a 3s window -> 1.0 word/sec at the last one
words = [{"start": 0.0}, {"start": 1.0}, {"start": 2.0}]
rates = compute_speech_rate(words, window_seconds=3.0)
assert rates[-1] == pytest.approx(1.0)
def test_old_words_fall_out_of_window(self):
words = [{"start": 0.0}, {"start": 100.0}]
rates = compute_speech_rate(words, window_seconds=3.0)
# only the word itself is in range at t=100
assert rates[-1] == pytest.approx(1 / 3.0)
def test_zero_window_is_not_a_division_error(self):
assert compute_speech_rate([{"start": 0.0}], window_seconds=0.0) == [0.0]
def test_empty_words(self):
assert compute_speech_rate([]) == []
class TestWordPitchEnergy:
def test_averages_track_values_within_word_span(self):
words = [{"word": "a", "start": 0.0, "end": 1.0}]
pitch = [(0.0, 100.0), (0.5, 200.0), (5.0, 999.0)]
energy = [(0.0, 0.2), (1.0, 0.4)]
out = word_pitch_energy(words, pitch, energy)
assert out[0]["pitch_hz"] == pytest.approx(150.0)
assert out[0]["energy"] == pytest.approx(0.3)
def test_none_when_no_frames_in_span(self):
words = [{"word": "a", "start": 10.0, "end": 11.0}]
out = word_pitch_energy(words, [(0.0, 100.0)], [(0.0, 0.5)])
assert out[0]["pitch_hz"] is None
assert out[0]["energy"] is None
def test_none_tracks_degrade_gracefully(self):
out = word_pitch_energy([{"word": "a", "start": 0.0, "end": 1.0}], None, None)
assert out[0]["pitch_hz"] is None
assert out[0]["energy"] is None
def test_does_not_mutate_input(self):
words = [{"word": "a", "start": 0.0, "end": 1.0}]
word_pitch_energy(words, [(0.0, 100.0)], None)
assert "pitch_hz" not in words[0]
def test_word_shorter_than_hop_gets_none_not_crash(self):
"""A word briefer than the frame spacing may contain no frame at all."""
words = [{"word": "a", "start": 0.501, "end": 0.502}]
out = word_pitch_energy(words, [(0.0, 100.0), (1.0, 200.0)], None)
assert out[0]["pitch_hz"] is None
class TestExtractorsDegradeGracefully:
def test_missing_file_returns_none(self):
assert extract_pitch("/nonexistent/audio.wav") is None
assert extract_energy("/nonexistent/audio.wav") is None
@pytest.mark.skipif(not LIBROSA, reason="librosa not installed")
class TestExtractorsWithLibrosa:
def test_capability_is_available(self):
ok, _msg = features_capability()
assert ok is True
def test_extracts_pitch_of_known_tone(self, tmp_path):
wav = tmp_path / "tone.wav"
_write_tone_wav(str(wav), hz=220.0, seconds=2.0)
track = extract_pitch(str(wav))
assert track is not None and len(track) > 0
hz_values = sorted(hz for _t, hz in track)
median = hz_values[len(hz_values) // 2]
assert median == pytest.approx(220.0, rel=0.1)
def test_extracts_energy_track(self, tmp_path):
wav = tmp_path / "tone.wav"
_write_tone_wav(str(wav), seconds=1.0)
track = extract_energy(str(wav))
assert track is not None and len(track) > 0
assert all(rms >= 0 for _t, rms in track)
assert max(rms for _t, rms in track) > 0
+147
View File
@@ -0,0 +1,147 @@
"""Tests for the analyze_voice_features MCP tool.
librosa and Whisper are monkeypatched so these run without the optional
extras, matching the pattern used by TestDetectBeatsHandler.
"""
import json
import struct
import wave
import pytest
def _write_silent_wav(path: str, seconds: float = 2.0) -> None:
n = int(44100 * seconds)
with wave.open(path, "w") as f:
f.setnchannels(1)
f.setsampwidth(2)
f.setframerate(44100)
f.writeframes(struct.pack("<%dh" % n, *([0] * n)))
_FAKE_TRANSCRIPT = {
"language": "pt",
"duration": 3.0,
"text": "isso e seguranca",
"segments": [{"text": "isso e seguranca", "start": 0.0, "end": 3.0}],
"words": [
{"word": "isso", "start": 0.0, "end": 0.4, "confidence": 0.9},
{"word": "e", "start": 0.5, "end": 0.7, "confidence": 0.9},
{"word": "seguranca", "start": 2.0, "end": 2.9, "confidence": 0.9},
],
}
@pytest.fixture
def wav(tmp_path):
path = tmp_path / "clip.wav"
_write_silent_wav(str(path))
return path
@pytest.fixture
def patched_analysis(monkeypatch):
"""Make the tool's transcription + librosa extractors deterministic."""
import server_tools._shared as _shared_mod
import server_tools.voice as server_mod
monkeypatch.setattr(_shared_mod, "transcribe", lambda *a, **k: _FAKE_TRANSCRIPT)
monkeypatch.setattr(server_mod, "features_capability", lambda: (True, "ok"))
# "seguranca" (2.0-2.9s) is the loud, high-pitched, emphatic word
monkeypatch.setattr(
server_mod,
"extract_pitch",
lambda *a, **k: [(0.2, 120.0), (0.6, 118.0), (2.4, 260.0)],
)
monkeypatch.setattr(
server_mod,
"extract_energy",
lambda *a, **k: [(0.2, 0.10), (0.6, 0.12), (2.4, 0.95)],
)
class TestAnalyzeVoiceFeaturesHandler:
async def test_reports_when_librosa_unavailable(self, wav, monkeypatch):
import server_tools.voice as server_mod
from server import handle_analyze_voice_features
monkeypatch.setattr(
server_mod, "features_capability", lambda: (False, "componente librosa ausente.")
)
result = await handle_analyze_voice_features({"media_path": str(wav)})
assert "librosa" in result[0].text.lower()
async def test_rejects_disallowed_extension(self, tmp_path):
from server import handle_analyze_voice_features
bad = tmp_path / "clip.txt"
bad.write_text("not audio")
with pytest.raises(ValueError):
await handle_analyze_voice_features({"media_path": str(bad)})
async def test_writes_features_json_with_emphasis_per_word(self, wav, patched_analysis):
from server import handle_analyze_voice_features
result = await handle_analyze_voice_features({"media_path": str(wav)})
text = result[0].text
json_path = wav.parent / "clip_voice_features.json"
assert str(json_path) in text
data = json.loads(json_path.read_text())
assert len(data["words"]) == 3
assert all("emphasis" in w for w in data["words"])
assert all(0.0 <= w["emphasis"] <= 1.0 for w in data["words"])
async def test_loudest_word_scores_highest_emphasis(self, wav, patched_analysis):
from server import handle_analyze_voice_features
await handle_analyze_voice_features({"media_path": str(wav)})
data = json.loads((wav.parent / "clip_voice_features.json").read_text())
by_word = {w["word"]: w["emphasis"] for w in data["words"]}
assert by_word["seguranca"] > by_word["isso"]
assert by_word["seguranca"] > by_word["e"]
async def test_persisted_config_is_embedded_in_output(self, wav, patched_analysis):
from server import handle_analyze_voice_features
await handle_analyze_voice_features({"media_path": str(wav)})
data = json.loads((wav.parent / "clip_voice_features.json").read_text())
assert "energy_threshold" in data["config"]
assert "emphasis_weights" in data["config"]
async def test_empty_transcript_reports_instead_of_crashing(self, wav, monkeypatch):
import server_tools._shared as _shared_mod
import server_tools.voice as server_mod
from server import handle_analyze_voice_features
monkeypatch.setattr(server_mod, "features_capability", lambda: (True, "ok"))
monkeypatch.setattr(
_shared_mod, "transcribe", lambda *a, **k: {**_FAKE_TRANSCRIPT, "words": []}
)
result = await handle_analyze_voice_features({"media_path": str(wav)})
assert "no words" in result[0].text.lower()
async def test_untranscribable_media_reports_install_hint(self, wav, monkeypatch):
import server_tools._shared as _shared_mod
import server_tools.voice as server_mod
from server import handle_analyze_voice_features
monkeypatch.setattr(server_mod, "features_capability", lambda: (True, "ok"))
monkeypatch.setattr(_shared_mod, "transcribe", lambda *a, **k: None)
result = await handle_analyze_voice_features({"media_path": str(wav)})
assert "faster-whisper" in result[0].text
async def test_missing_pitch_track_degrades_without_crashing(self, wav, monkeypatch):
import server_tools._shared as _shared_mod
import server_tools.voice as server_mod
from server import handle_analyze_voice_features
monkeypatch.setattr(_shared_mod, "transcribe", lambda *a, **k: _FAKE_TRANSCRIPT)
monkeypatch.setattr(server_mod, "features_capability", lambda: (True, "ok"))
monkeypatch.setattr(server_mod, "extract_pitch", lambda *a, **k: None)
monkeypatch.setattr(server_mod, "extract_energy", lambda *a, **k: None)
result = await handle_analyze_voice_features({"media_path": str(wav)})
assert "Voice Feature Analysis" in result[0].text
data = json.loads((wav.parent / "clip_voice_features.json").read_text())
assert all(w["emphasis"] >= 0.0 for w in data["words"])
+338
View File
@@ -0,0 +1,338 @@
"""Tests for fcpxml/voice_timeline.py — the consolidated AI-readable timeline.
The document's shape is the contract downstream consumers (rules engine, a
model reading the JSON) rely on, so these tests pin the shape as much as
the values — including that it survives every analysis layer being absent.
"""
import json
import pytest
from fcpxml.voice_timeline import (
VOICE_TIMELINE_VERSION,
build_voice_timeline,
enrich_words,
load_voice_timeline,
save_voice_timeline,
voice_timeline_path,
)
_TRANSCRIPT = {
"language": "pt",
"duration": 4.0,
"text": "isso e seguranca total",
"segments": [
{"text": "isso e", "start": 0.0, "end": 1.0},
{"text": "seguranca total", "start": 2.0, "end": 4.0},
],
"words": [
{"word": "isso", "start": 0.0, "end": 0.4, "confidence": 0.9},
{"word": "e", "start": 0.5, "end": 0.7, "confidence": 0.9},
{"word": "seguranca", "start": 2.0, "end": 2.9, "confidence": 0.9},
{"word": "total", "start": 3.0, "end": 3.5, "confidence": 0.9},
],
}
# "seguranca" is the loud, high-pitched moment
_PITCH = [(0.2, 120.0), (0.6, 118.0), (2.4, 260.0), (3.2, 130.0)]
_ENERGY = [(0.2, 0.10), (0.6, 0.12), (2.4, 0.95), (3.2, 0.20)]
class TestEnrichWords:
def test_normalizes_energy_against_loudest_word(self):
enriched = enrich_words(_TRANSCRIPT["words"], _PITCH, _ENERGY)
loudest = max(enriched, key=lambda w: w["energy_norm"])
assert loudest["word"] == "seguranca"
assert loudest["energy_norm"] == pytest.approx(1.0)
def test_all_values_stay_within_unit_range(self):
enriched = enrich_words(_TRANSCRIPT["words"], _PITCH, _ENERGY)
for w in enriched:
for key in ("energy_norm", "pitch_delta", "rate_delta", "emphasis"):
assert 0.0 <= w[key] <= 1.0, f"{key} out of range on {w['word']}"
def test_empty_words_returns_empty(self):
assert enrich_words([], _PITCH, _ENERGY) == []
def test_missing_tracks_give_zero_not_crash(self):
enriched = enrich_words(_TRANSCRIPT["words"], None, None)
assert all(w["energy_norm"] == 0.0 for w in enriched)
assert all(w["pitch_delta"] == 0.0 for w in enriched)
class TestBuildVoiceTimeline:
@pytest.fixture
def timeline(self, monkeypatch):
import fcpxml.voice_timeline as vt
monkeypatch.setattr(vt, "extract_pitch", lambda *a, **k: _PITCH)
monkeypatch.setattr(vt, "extract_energy", lambda *a, **k: _ENERGY)
return build_voice_timeline("/tmp/clip.wav", _TRANSCRIPT)
def test_document_has_all_top_level_layers(self, timeline):
for key in ("version", "source", "language", "scales", "summary", "speakers", "segments"):
assert key in timeline
assert timeline["version"] == VOICE_TIMELINE_VERSION
def test_scales_document_every_word_metric(self, timeline):
word = timeline["segments"][0]["words"][0]
for metric in timeline["scales"]["word"]:
assert metric in word, f"{metric} documented in scales but absent from words"
def test_scales_document_every_segment_metric(self, timeline):
segment = timeline["segments"][0]
for metric in timeline["scales"]["segment"]:
assert metric in segment, f"{metric} documented in scales but absent from segments"
def test_take_boundary_flags_a_long_gap(self, timeline):
# the fixture has a 1s gap between its two segments -> not a boundary
assert timeline["segments"][1]["gap_before"] > 0
assert timeline["segments"][1]["take_boundary"] is False
def test_summary_counts_match_the_detail(self, timeline):
summary = timeline["summary"]
assert summary["segment_count"] == len(timeline["segments"])
total_words = sum(len(s["words"]) for s in timeline["segments"])
assert summary["word_count"] == total_words
def test_words_are_grouped_under_their_segment(self, timeline):
first, second = timeline["segments"]
assert [w["text"] for w in first["words"]] == ["isso", "e"]
assert [w["text"] for w in second["words"]] == ["seguranca", "total"]
def test_segment_aggregates_reflect_their_words(self, timeline):
loud_segment = timeline["segments"][1]
quiet_segment = timeline["segments"][0]
assert loud_segment["avg_energy"] > quiet_segment["avg_energy"]
assert loud_segment["peak_emphasis"] >= max(w["emphasis"] for w in loud_segment["words"])
def test_peak_moments_are_sorted_by_emphasis(self, timeline):
peaks = timeline["summary"]["peak_moments"]
assert peaks == sorted(peaks, key=lambda m: m["emphasis"], reverse=True)
def test_defaults_to_single_speaker_without_token(self, timeline):
assert timeline["summary"]["speaker_count"] == 1
assert all(w["speaker"] == "SPEAKER_00" for s in timeline["segments"] for w in s["words"])
def test_is_json_serializable(self, timeline):
# the whole point is handing this to a model / writing it to disk
assert json.loads(json.dumps(timeline, ensure_ascii=False))["version"]
class TestDegradesWithoutAnalysisLayers:
def test_shape_survives_with_no_acoustics(self, monkeypatch):
import fcpxml.voice_timeline as vt
monkeypatch.setattr(vt, "extract_pitch", lambda *a, **k: None)
monkeypatch.setattr(vt, "extract_energy", lambda *a, **k: None)
timeline = build_voice_timeline("/tmp/clip.wav", _TRANSCRIPT)
assert timeline["summary"]["word_count"] == 4
assert timeline["summary"]["avg_emphasis"] >= 0.0
def test_empty_transcript_still_yields_valid_document(self, monkeypatch):
import fcpxml.voice_timeline as vt
monkeypatch.setattr(vt, "extract_pitch", lambda *a, **k: None)
monkeypatch.setattr(vt, "extract_energy", lambda *a, **k: None)
timeline = build_voice_timeline(
"/tmp/clip.wav", {"duration": 0.0, "segments": [], "words": []}
)
assert timeline["segments"] == []
assert timeline["summary"]["word_count"] == 0
assert timeline["summary"]["avg_emphasis"] == 0.0
def test_progress_callback_is_reported(self, monkeypatch):
import fcpxml.voice_timeline as vt
monkeypatch.setattr(vt, "extract_pitch", lambda *a, **k: None)
monkeypatch.setattr(vt, "extract_energy", lambda *a, **k: None)
seen: list[tuple[float, str]] = []
build_voice_timeline("/tmp/clip.wav", _TRANSCRIPT, progress_cb=lambda f, s: seen.append((f, s)))
assert seen and all(0.0 <= f <= 1.0 for f, _ in seen)
class TestPersistence:
def test_round_trip(self, tmp_path):
path = tmp_path / "clip_voice_timeline.json"
timeline = {"version": "1.0", "segments": [], "summary": {}}
save_voice_timeline(timeline, path)
assert load_voice_timeline(path) == timeline
def test_accented_text_stays_readable(self, tmp_path):
path = tmp_path / "t.json"
save_voice_timeline({"segments": [{"text": "segurança"}]}, path)
assert "segurança" in path.read_text(encoding="utf-8")
def test_missing_file_returns_none(self, tmp_path):
assert load_voice_timeline(tmp_path / "absent.json") is None
def test_malformed_json_returns_none(self, tmp_path):
path = tmp_path / "bad.json"
path.write_text("{not json")
assert load_voice_timeline(path) is None
def test_wrong_shape_returns_none(self, tmp_path):
path = tmp_path / "other.json"
path.write_text('{"something": "else"}')
assert load_voice_timeline(path) is None
def test_path_next_to_media_by_default(self):
assert voice_timeline_path("/media/clip.mov").name == "clip_voice_timeline.json"
def test_path_honours_output_dir(self, tmp_path):
path = voice_timeline_path("/media/clip.mov", output_dir=str(tmp_path))
assert path.parent == tmp_path
_RAW_WORDS = [
# a loud outlier that will be cut, plus quieter material that survives
{"text": "GRITO", "start": 1.0, "end": 1.5, "speaker": "SPEAKER_00",
"energy": 0.5, "pitch_delta": 0.5, "rate_delta": 0.0, "pause_before": 0.0,
"emphasis": 0.5, "energy_raw": 1.0, "pitch_hz": 300.0},
{"text": "mastopexia", "start": 10.0, "end": 10.8, "speaker": "SPEAKER_00",
"energy": 0.2, "pitch_delta": 0.1, "rate_delta": 0.0, "pause_before": 0.0,
"emphasis": 0.1, "energy_raw": 0.4, "pitch_hz": 190.0},
{"text": "a", "start": 11.0, "end": 11.1, "speaker": "SPEAKER_00",
"energy": 0.15, "pitch_delta": 0.05, "rate_delta": 0.0, "pause_before": 0.0,
"emphasis": 0.08, "energy_raw": 0.3, "pitch_hz": 185.0},
]
_RESTRICT_TIMELINE = {
"version": "1.0", "source": "x.mp4", "speakers": [],
"segments": [
{"start": 1.0, "end": 1.5, "speaker": "SPEAKER_00", "text": "GRITO",
"gap_before": 0.0, "take_boundary": False, "avg_energy": 0.5,
"peak_emphasis": 0.5, "words": [_RAW_WORDS[0]]},
{"start": 10.0, "end": 11.1, "speaker": "SPEAKER_00",
"text": "mastopexia a", "gap_before": 8.5, "take_boundary": True,
"avg_energy": 0.17, "peak_emphasis": 0.1, "words": _RAW_WORDS[1:]},
],
}
class TestRestrictToKept:
"""Emphasis is relative. Cut the loudest moment out and everything left
is still scored against something the viewer will never see, so the
surviving material has to be re-normalized on its own."""
def test_cut_words_are_dropped(self):
from fcpxml.voice_timeline import restrict_to_kept
r = restrict_to_kept(_RESTRICT_TIMELINE, [(0.0, 5.0)])
texts = [w["text"] for s in r["segments"] for w in s["words"]]
assert "GRITO" not in texts
assert "mastopexia" in texts
def test_survivors_are_rescored_against_each_other(self):
from fcpxml.voice_timeline import restrict_to_kept
r = restrict_to_kept(_RESTRICT_TIMELINE, [(0.0, 5.0)])
word = next(w for s in r["segments"] for w in s["words"] if w["text"] == "mastopexia")
# was 0.2 against the shout's 1.0; alone it becomes the loudest
assert word["energy"] == pytest.approx(1.0)
def test_empty_segments_are_removed(self):
from fcpxml.voice_timeline import restrict_to_kept
r = restrict_to_kept(_RESTRICT_TIMELINE, [(0.0, 5.0)])
assert len(r["segments"]) == 1
def test_no_cuts_keeps_everything(self):
from fcpxml.voice_timeline import restrict_to_kept
r = restrict_to_kept(_RESTRICT_TIMELINE, [])
assert sum(len(s["words"]) for s in r["segments"]) == 3
def test_times_stay_in_original_source_seconds(self):
from fcpxml.voice_timeline import restrict_to_kept
r = restrict_to_kept(_RESTRICT_TIMELINE, [(0.0, 5.0)])
assert r["segments"][0]["start"] == 10.0
class TestSuggestZoomWindows:
def test_skips_function_words(self):
from fcpxml.voice_timeline import restrict_to_kept, suggest_zoom_windows
r = restrict_to_kept(_RESTRICT_TIMELINE, [(0.0, 5.0)])
zooms = suggest_zoom_windows(r)
assert zooms and all(z["word"] != "a" for z in zooms)
def test_window_runs_from_the_word_to_the_end_of_its_line(self):
from fcpxml.voice_timeline import restrict_to_kept, suggest_zoom_windows
r = restrict_to_kept(_RESTRICT_TIMELINE, [(0.0, 5.0)])
z = suggest_zoom_windows(r)[0]
assert z["start"] == 10.0 and z["end"] == 11.1
def test_min_gap_keeps_zooms_apart(self):
from fcpxml.voice_timeline import suggest_zoom_windows
timeline = {"segments": [
{"start": t, "end": t + 1.0, "text": "linha",
"words": [{"text": "palavra", "start": t, "end": t + 0.5, "emphasis": 0.5 - i * 0.01}]}
for i, t in enumerate([0.0, 1.0, 2.0, 30.0])
]}
zooms = suggest_zoom_windows(timeline, min_gap=8.0)
assert len(zooms) == 2
def test_max_zooms_caps_the_result(self):
from fcpxml.voice_timeline import suggest_zoom_windows
timeline = {"segments": [
{"start": t, "end": t + 1.0, "text": "linha",
"words": [{"text": "palavra", "start": t, "end": t + 0.5, "emphasis": 0.5}]}
for t in [0.0, 20.0, 40.0, 60.0]
]}
assert len(suggest_zoom_windows(timeline, min_gap=8.0, max_zooms=2)) == 2
def test_results_are_in_chronological_order(self):
from fcpxml.voice_timeline import suggest_zoom_windows
timeline = {"segments": [
{"start": t, "end": t + 1.0, "text": "linha",
"words": [{"text": "palavra", "start": t, "end": t + 0.5, "emphasis": e}]}
for t, e in [(60.0, 0.9), (0.0, 0.5), (30.0, 0.7)]
]}
zooms = suggest_zoom_windows(timeline, min_gap=8.0)
assert [z["start"] for z in zooms] == sorted(z["start"] for z in zooms)
class TestSentenceEnd:
"""Transcription segments break on breath, not grammar — a sentence
routinely spans several. A zoom ending on a segment boundary releases
mid-thought, which is what makes a punch-in feel arbitrary."""
SEGS = [
{"start": 0.0, "end": 5.0, "text": "Aquela mama com um formato, que dá aquele ar",
"take_boundary": False},
{"start": 5.0, "end": 10.7, "text": "de elegância, isso é desejo de muitas mulheres, né?",
"take_boundary": False},
{"start": 11.0, "end": 14.0, "text": "Com o tempo, o corpo muda.", "take_boundary": False},
]
def test_extends_past_a_segment_that_does_not_end_a_sentence(self):
from fcpxml.voice_timeline import sentence_end
assert sentence_end(self.SEGS, 0) == 10.7
def test_stops_at_terminal_punctuation(self):
from fcpxml.voice_timeline import sentence_end
assert sentence_end(self.SEGS, 2) == 14.0
def test_never_runs_past_a_take_boundary(self):
from fcpxml.voice_timeline import sentence_end
segs = [
{"start": 0.0, "end": 5.0, "text": "frase sem fim", "take_boundary": False},
{"start": 12.0, "end": 15.0, "text": "outra tomada", "take_boundary": True},
]
assert sentence_end(segs, 0) == 5.0
def test_last_segment_without_punctuation_ends_at_itself(self):
from fcpxml.voice_timeline import sentence_end
segs = [{"start": 0.0, "end": 4.0, "text": "sem ponto final", "take_boundary": False}]
assert sentence_end(segs, 0) == 4.0
+107
View File
@@ -0,0 +1,107 @@
"""Tests for the build_voice_timeline MCP tool."""
import json
import pytest
from tests.test_voice_features_tool import _write_silent_wav
_TRANSCRIPT = {
"language": "pt",
"duration": 4.0,
"text": "isso e seguranca total",
"segments": [
{"text": "isso e", "start": 0.0, "end": 1.0},
{"text": "seguranca total", "start": 2.0, "end": 4.0},
],
"words": [
{"word": "isso", "start": 0.0, "end": 0.4, "confidence": 0.9},
{"word": "e", "start": 0.5, "end": 0.7, "confidence": 0.9},
{"word": "seguranca", "start": 2.0, "end": 2.9, "confidence": 0.9},
{"word": "total", "start": 3.0, "end": 3.5, "confidence": 0.9},
],
}
@pytest.fixture
def wav(tmp_path):
path = tmp_path / "clip.wav"
_write_silent_wav(str(path))
return path
@pytest.fixture
def patched(monkeypatch):
"""Deterministic transcription + acoustics, no optional extras needed."""
import fcpxml.voice_timeline as vt
import server_tools._shared as _shared_mod
monkeypatch.setattr(_shared_mod, "transcribe", lambda *a, **k: _TRANSCRIPT)
monkeypatch.setattr(vt, "extract_pitch", lambda *a, **k: [(2.4, 260.0), (0.2, 120.0)])
monkeypatch.setattr(vt, "extract_energy", lambda *a, **k: [(2.4, 0.95), (0.2, 0.10)])
class TestBuildVoiceTimelineHandler:
async def test_writes_timeline_json(self, wav, patched):
from server import handle_build_voice_timeline
result = await handle_build_voice_timeline({"media_path": str(wav)})
text = result[0].text
json_path = wav.parent / "clip_voice_timeline.json"
assert str(json_path) in text
data = json.loads(json_path.read_text(encoding="utf-8"))
assert data["summary"]["word_count"] == 4
assert len(data["segments"]) == 2
async def test_reports_which_layers_ran(self, wav, patched):
from server import handle_build_voice_timeline
result = await handle_build_voice_timeline({"media_path": str(wav)})
text = result[0].text
assert "Analysis Layers" in text
assert "Transcript" in text
async def test_rejects_disallowed_extension(self, tmp_path):
from server import handle_build_voice_timeline
bad = tmp_path / "clip.txt"
bad.write_text("not audio")
with pytest.raises(ValueError):
await handle_build_voice_timeline({"media_path": str(bad)})
async def test_untranscribable_media_reports_hint(self, wav, monkeypatch):
import server_tools._shared as _shared_mod
from server import handle_build_voice_timeline
monkeypatch.setattr(_shared_mod, "transcribe", lambda *a, **k: None)
result = await handle_build_voice_timeline({"media_path": str(wav)})
assert "faster-whisper" in result[0].text
async def test_uses_persisted_peak_settings(self, wav, patched, monkeypatch):
"""A wider percentile must surface more peak moments."""
import server_tools.voice as server_mod
from server import handle_build_voice_timeline
def config(percentile):
return {
"energy_threshold": 0.5,
"peak_percentile": percentile,
"emphasis_floor": 0.0,
"emphasis_weights": {
"energy": 0.30, "pitch_variation": 0.25, "rate_variation": 0.20,
"pause_before": 0.15, "duration": 0.10,
},
"emotion_enabled": False,
"emotion_sensitivity": 0.5,
}
monkeypatch.setattr(server_mod, "load_voice_analysis_config", lambda: config(0.01))
await handle_build_voice_timeline({"media_path": str(wav)})
strict = json.loads((wav.parent / "clip_voice_timeline.json").read_text())
monkeypatch.setattr(server_mod, "load_voice_analysis_config", lambda: config(1.0))
await handle_build_voice_timeline({"media_path": str(wav)})
loose = json.loads((wav.parent / "clip_voice_timeline.json").read_text())
assert loose["summary"]["peak_count"] > strict["summary"]["peak_count"]
+275 -4
View File
@@ -652,6 +652,7 @@ def test_add_zoom_creates_keyframed_transform(temp_fcpxml):
"""4 keyframes: 100% -> scale -> scale -> 100%, all within [start, end]."""
modifier = FCPXMLModifier(temp_fcpxml)
clip = modifier.add_zoom(clip_id='Broll_Studio', start=1.0, end=3.0, scale=1.3, ease=0.5)
frame = float(modifier.frame_duration_fraction())
transform = clip.find('adjust-transform')
assert transform is not None
@@ -660,9 +661,23 @@ def test_add_zoom_creates_keyframed_transform(temp_fcpxml):
keyframes = param.find('keyframeAnimation').findall('keyframe')
assert len(keyframes) == 4
assert [kf.get('value') for kf in keyframes] == ['1 1', '1.3 1.3', '1.3 1.3', '1 1']
assert all(kf.get('interp') == 'ease' for kf in keyframes)
# Bare keyframes — only time and value — matching a zoom exported from
# FCP itself. It rejects 'interp' on this vector param (discarding the
# whole <param>), and its own export writes no 'curve' either.
assert not any(kf.get('interp') for kf in keyframes)
assert not any(kf.get('curve') for kf in keyframes)
assert all(set(kf.attrib) == {'time', 'value'} for kf in keyframes)
# Keyframe times are anchored in the clip's SOURCE timebase (its own
# `start`), not clip-relative. This fixture starts at 10s, so a zoom over
# clip seconds 1-3 must be written at 11-13s. Writing 1-3s here would put
# the animation outside the clip and FCP imports it as nothing.
origin = modifier._parse_time(clip.get('start', '0s')).to_seconds()
times = [modifier._parse_time(kf.get('time')).to_seconds() for kf in keyframes]
assert times == pytest.approx([1.0, 1.5, 2.5, 3.0], abs=0.05)
assert origin > 0, "fixture must start off zero, or this asserts nothing"
# in over 0.5s, hold, then snap back on the very next frame
assert times == pytest.approx(
[origin + 1.0, origin + 1.5, origin + 3.0 - frame, origin + 3.0], abs=0.05
)
assert times == sorted(times)
@@ -699,8 +714,9 @@ def test_add_zoom_window_outside_clip_duration_raises(temp_fcpxml):
def test_add_zoom_ease_too_long_for_window_raises(temp_fcpxml):
modifier = FCPXMLModifier(temp_fcpxml)
with pytest.raises(ValueError, match="doesn't fit"):
modifier.add_zoom(clip_id='Broll_Studio', start=0.0, end=1.0, ease=1.0)
# start away from the clip head so the ramp-in is actually written
with pytest.raises(ValueError, match="don't fit"):
modifier.add_zoom(clip_id='Broll_Studio', start=1.0, end=2.0, ease=1.5)
def test_change_speed_twice_no_duplicate_elements(temp_fcpxml):
@@ -2067,3 +2083,258 @@ def test_remove_trailing_gaps_noop_without_gap():
assert len(children) == 1
assert children[0].tag == 'asset-clip'
Path(f.name).unlink(missing_ok=True)
class TestNTSCFrameAlignment:
"""23.976/29.97 timebases must not be reported as misaligned.
Regression: the check used int(fps), so an exactly frame-aligned NTSC
duration (a whole multiple of 1001/24000s) was flagged as broken —
every NTSC project produced spurious warnings that buried real ones.
"""
NTSC_DOC = """<?xml version="1.0" encoding="UTF-8"?>
<fcpxml version="1.13">
<resources>
<format id="r1" name="FFVideoFormat1080p2398" frameDuration="1001/24000s" width="1920" height="1080"/>
<asset id="a1" name="v" start="0s" duration="24437413/24000s" hasVideo="1" format="r1">
<media-rep kind="original-media" src="file:///v.mov"/>
</asset>
</resources>
<library><event name="E"><project name="P">
<sequence format="r1" duration="24437413/24000s" tcStart="0s">
<spine><asset-clip name="v" ref="a1" offset="0s" start="0s" duration="24437413/24000s"/></spine>
</sequence>
</project></event></library>
</fcpxml>
"""
def _issues(self, xml):
from fcpxml.safe_xml import safe_fromstring
from fcpxml.writer import validate_fcpxml
root = safe_fromstring(xml)
return [
i for i in (validate_fcpxml(root) or [])
if "frame" in str(i.issue_type.value)
]
def test_aligned_ntsc_duration_is_not_flagged(self):
# 24437413/24000s is exactly 24413 frames of 1001/24000s
assert self._issues(self.NTSC_DOC) == []
def test_genuinely_misaligned_duration_is_still_flagged(self):
broken = self.NTSC_DOC.replace(
'<asset-clip name="v" ref="a1" offset="0s" start="0s" duration="24437413/24000s"/>',
'<asset-clip name="v" ref="a1" offset="0s" start="0s" duration="500/24000s"/>',
)
assert len(self._issues(broken)) == 1
def test_message_names_the_real_rate_not_a_rounded_one(self):
broken = self.NTSC_DOC.replace('duration="24437413/24000s"/>', 'duration="500/24000s"/>')
issues = self._issues(broken)
assert issues and "23.976fps" in issues[0].message
class TestZoomPreservesExistingFraming:
"""A clip may already carry the editor's reframe — rotation for footage
shot sideways, position, a base scale. add_zoom used to delete it, which
on real footage brought the zoomed section back rotated."""
FRAMED = """<?xml version="1.0" encoding="UTF-8"?>
<fcpxml version="1.13">
<resources>
<format id="r1" frameDuration="100/3000s" width="1920" height="1080"/>
<asset id="a1" name="v" start="0s" duration="600/30s" hasVideo="1" format="r1">
<media-rep kind="original-media" src="file:///v.mov"/>
</asset>
</resources>
<library><event name="E"><project name="P">
<sequence format="r1" duration="600/30s" tcStart="0s">
<spine>
<asset-clip name="v" ref="a1" offset="0s" start="0s" duration="600/30s">
<adjust-transform position="0.16 0.66" rotation="90.1" scale="1.77311 1.77311"/>
</asset-clip>
</spine>
</sequence>
</project></event></library>
</fcpxml>
"""
def _zoomed(self, tmp_path, scale=1.2):
from fcpxml.writer import FCPXMLModifier
path = tmp_path / "framed.fcpxml"
path.write_text(self.FRAMED)
m = FCPXMLModifier(str(path))
m.add_zoom(clip_id="v", start=1.0, end=5.0, scale=scale)
return m.root.find(".//adjust-transform")
def test_rotation_and_position_survive(self, tmp_path):
t = self._zoomed(tmp_path)
assert t.get("rotation") == "90.1"
assert t.get("position") == "0.16 0.66"
def test_animation_rests_at_the_existing_scale(self, tmp_path):
kfs = self._zoomed(tmp_path).findall(".//keyframe")
# first and last keyframe return to the clip's own framing, not to 1
assert kfs[0].get("value").split()[0].startswith("1.77")
assert kfs[-1].get("value").split()[0].startswith("1.77")
def test_peak_multiplies_the_existing_scale(self, tmp_path):
kfs = self._zoomed(tmp_path, scale=2.0).findall(".//keyframe")
peak = float(kfs[1].get("value").split()[0])
assert peak == pytest.approx(1.77311 * 2.0, rel=1e-4)
def test_only_one_transform_remains(self, tmp_path):
from fcpxml.writer import FCPXMLModifier
path = tmp_path / "framed.fcpxml"
path.write_text(self.FRAMED)
m = FCPXMLModifier(str(path))
m.add_zoom(clip_id="v", start=1.0, end=5.0, scale=1.2)
m.add_zoom(clip_id="v", start=6.0, end=9.0, scale=1.4)
assert len(m.root.findall(".//adjust-transform")) == 1
def test_second_zoom_on_same_clip_keeps_the_real_base_scale(self, tmp_path):
"""Found on real footage: two zoom actions landing on disjoint
windows of the same post-cut clip. The second add_zoom call used to
see the already-animated <param name="scale"> from the first zoom
instead of a static attribute, read that as "no framing", and
default the base to 1.0 — silently shrinking the shot back to its
unframed size for the whole clip wherever no keyframe applied, and
discarding the first zoom's animation in the process."""
from fcpxml.writer import FCPXMLModifier
path = tmp_path / "framed.fcpxml"
path.write_text(self.FRAMED)
m = FCPXMLModifier(str(path))
m.add_zoom(clip_id="v", start=1.0, end=5.0, scale=1.2)
m.add_zoom(clip_id="v", start=6.0, end=9.0, scale=1.4)
kfs = m.root.findall(".//keyframe")
# Every rest keyframe returns to the clip's real base scale, never 1.0.
rest_values = {kf.get("value") for kf in (kfs[0], kfs[3], kfs[4], kfs[-1])}
assert rest_values == {"1.77311 1.77311"}
# Both peaks survive — the second call didn't erase the first.
peaks = sorted(float(kf.get("value").split()[0]) for kf in (kfs[1], kfs[5]))
assert peaks[0] == pytest.approx(1.77311 * 1.2, rel=1e-4)
assert peaks[1] == pytest.approx(1.77311 * 1.4, rel=1e-4)
def test_overlapping_zoom_on_same_clip_replaces_instead_of_stacking(self, tmp_path):
"""Two windows that OVERLAP mean "redo this zoom", not "add another
one" — the old keyframes are stale and all of them go, matching
test_add_zoom_replaces_existing_zoom's contract."""
from fcpxml.writer import FCPXMLModifier
path = tmp_path / "framed.fcpxml"
path.write_text(self.FRAMED)
m = FCPXMLModifier(str(path))
m.add_zoom(clip_id="v", start=1.0, end=5.0, scale=1.2)
m.add_zoom(clip_id="v", start=3.0, end=7.0, scale=1.5)
values = [kf.get("value") for kf in m.root.findall(".//keyframe")]
assert not any(v.startswith("2.1277") for v in values) # 1.77311*1.2 gone
assert any(v.startswith("2.6596") for v in values) # 1.77311*1.5 present
def test_unframed_clip_still_rests_at_one(self, tmp_path):
from fcpxml.writer import FCPXMLModifier
path = tmp_path / "plain.fcpxml"
path.write_text(self.FRAMED.replace(
'<adjust-transform position="0.16 0.66" rotation="90.1" scale="1.77311 1.77311"/>', ''))
m = FCPXMLModifier(str(path))
m.add_zoom(clip_id="v", start=1.0, end=5.0, scale=1.3)
kfs = m.root.findall(".//keyframe")
assert kfs[0].get("value") == "1 1"
assert kfs[1].get("value") == "1.3 1.3"
class TestZoomShapeIsAsymmetric:
"""The editorial shape: ramp in fast to land with the emphasised word,
hold through the impact phrase, then snap back in a single frame so the
video resumes its normal framing without a drift that draws the eye."""
def _times(self, temp_fcpxml, **kw):
from fcpxml.writer import FCPXMLModifier
m = FCPXMLModifier(temp_fcpxml)
# Broll_Studio is 5s long; end=3.0 keeps clear of the hold-at-cut
# margin so these exercise the ordinary return-to-framing shape.
kw.setdefault('end', 3.0)
clip = m.add_zoom(clip_id='Broll_Studio', start=1.0, **kw)
origin = m._parse_time(clip.get('start', '0s')).to_seconds()
kfs = clip.find('adjust-transform').find('param').find('keyframeAnimation')
return m, origin, [m._parse_time(k.get('time')).to_seconds() - origin
for k in kfs.findall('keyframe')]
def test_return_takes_a_single_frame(self, temp_fcpxml):
m, _, t = self._times(temp_fcpxml, scale=1.2)
frame = float(m.frame_duration_fraction())
assert t[3] - t[2] == pytest.approx(frame, abs=0.005)
def test_ramp_in_is_quick(self, temp_fcpxml):
"""Fast enough to land with the emphasised word rather than drift."""
_, _, t = self._times(temp_fcpxml, scale=1.2)
assert t[1] - t[0] == pytest.approx(0.25, abs=0.05)
def test_peak_is_held_until_the_return(self, temp_fcpxml):
_, _, t = self._times(temp_fcpxml, scale=1.2)
# hold spans from the top of the ramp to one frame before the end
assert t[2] - t[1] > 1.4
def test_zoom_opening_at_a_cut_starts_already_zoomed(self, temp_fcpxml):
"""The cut is the transition — ramping up from it reads as the shot
settling rather than as emphasis."""
from fcpxml.writer import FCPXMLModifier
m = FCPXMLModifier(temp_fcpxml)
clip = m.add_zoom(clip_id='Broll_Studio', start=0.1, end=3.0, scale=1.2)
kfs = clip.find('adjust-transform').find('param').find('keyframeAnimation')
values = [k.get('value') for k in kfs.findall('keyframe')]
assert values[0] != values[-1] # opens zoomed, returns to framing
assert values[0] == values[-2] # ...and was at the peak from frame one
def test_opening_at_peak_can_be_forced_off(self, temp_fcpxml):
from fcpxml.writer import FCPXMLModifier
m = FCPXMLModifier(temp_fcpxml)
clip = m.add_zoom(
clip_id='Broll_Studio', start=0.1, end=3.0, scale=1.2, start_at_peak=False
)
kfs = clip.find('adjust-transform').find('param').find('keyframeAnimation')
assert len(kfs.findall('keyframe')) == 4
def test_ease_out_can_be_made_gradual(self, temp_fcpxml):
_, _, t = self._times(temp_fcpxml, scale=1.2, ease_out=1.0)
assert t[3] - t[2] == pytest.approx(1.0, abs=0.05)
def test_zoom_reaching_the_cut_holds_instead_of_returning(self, temp_fcpxml):
"""Returning right before a cut is wasted motion — the next clip
opens on its own framing, so the move back reads as a twitch."""
_, _, t = self._times(temp_fcpxml, scale=1.2, end=5.0)
assert len(t) == 3 # rest, peak, still peak at the cut
def test_hold_can_be_forced_off_at_a_cut(self, temp_fcpxml):
_, _, t = self._times(temp_fcpxml, scale=1.2, end=5.0, hold_at_end=False)
assert len(t) == 4
def test_hold_can_be_forced_on_mid_clip(self, temp_fcpxml):
_, _, t = self._times(temp_fcpxml, scale=1.2, end=3.0, hold_at_end=True)
assert len(t) == 3
def test_held_zoom_stays_at_the_peak(self, temp_fcpxml):
from fcpxml.writer import FCPXMLModifier
m = FCPXMLModifier(temp_fcpxml)
clip = m.add_zoom(clip_id='Broll_Studio', start=1.0, end=5.0, scale=1.2)
kfs = clip.find('adjust-transform').find('param').find('keyframeAnimation')
values = [k.get('value') for k in kfs.findall('keyframe')]
assert values[-1] == values[-2] != values[0]
def test_window_too_short_for_the_ramp_is_rejected(self, temp_fcpxml):
from fcpxml.writer import FCPXMLModifier
m = FCPXMLModifier(temp_fcpxml)
with pytest.raises(ValueError, match="don't fit"):
m.add_zoom(clip_id='Broll_Studio', start=1.0, end=1.2, scale=1.2, ease=1.5)