diff --git a/.claude/skills/editar-por-voz/SKILL.md b/.claude/skills/editar-por-voz/SKILL.md new file mode 100644 index 0000000..32af24e --- /dev/null +++ b/.claude/skills/editar-por-voz/SKILL.md @@ -0,0 +1,93 @@ +--- +name: editar-por-voz +description: Edita um vídeo a partir da análise de voz — lê a timeline de voz (JSON) de uma gravação, separa o roteiro da conversa de bastidor, escolhe a melhor tomada de cada frase, decide cortes/zooms/textos e devolve a lista de decisões em JSON. Use quando o usuário pedir para editar por voz, montar um corte automático, limpar tomadas repetidas, escolher as melhores tomadas, ou marcar os momentos de ênfase de uma fala. +--- + +# Editar por voz + +Você recebe um JSON com o que foi dito, por quem e **como**; devolve um JSON +com **o que fazer**. Quem aplica é o programa — você nunca escreve XML. + +## Princípio + +O sistema já mede *como* a pessoa falou, de forma reprodutível. Não +recalcule nada disso nem reestime tempos "no olho". + +Seu trabalho é o que nenhum limiar resolve: **o que aquilo significa.** O +índice diz que uma palavra foi dita com força; só você sabe se ela é o +argumento central ou uma piada com o câmera. + +## Entrada e saída + +**Entrada:** `_voice_timeline.json` (de `build_voice_timeline`). Se +não existir, rode a ferramenta; se existir, leia direto — a análise leva +minutos. + +**Saída:** JSON com a lista de ações → `criterios/08-formato-de-saida.md` + +`apply_voice_actions` aplica direto, mas é para **teste**. O produto do seu +trabalho é a lista de decisões. + +## Ordem de trabalho + +Siga nesta ordem. Pular a Fase 2 ou a 3 leva a decisões erradas. + +| Fase | O que fazer | Critérios | +|---|---|---| +| **0** | Ler `layers` — saber o que rodou | `criterios/09-analise-incompleta.md` | +| **1** | Ler o JSON em camadas | `criterios/01-leitura-do-json.md` | +| **2** | Separar roteiro de conversa de bastidor | `criterios/02-triagem-roteiro-vs-conversa.md` | +| **3** | Escolher a melhor tomada de cada frase | `criterios/03-escolha-da-melhor-tomada.md` | +| **4** | **Reanalisar** o material que sobrou | `criterios/04-reanalise-do-material-restante.md` | +| **5** | Decidir zooms | `criterios/05-zoom.md` | +| **6** | Decidir textos, cortes e marcadores | `criterios/06-texto-corte-marcador.md` | +| **7** | Cortar a lista pelo ritmo | `criterios/07-ritmo.md` | +| **8** | Montar o JSON de saída | `criterios/08-formato-de-saida.md` | + +## As três armadilhas + +Cada uma já causou erro silencioso em material real: + +1. **A conversa de bastidor tem a maior ênfase do vídeo.** O índice acústico + favorece a fala solta sobre o texto decorado. Separar é tarefa de texto, + nunca de limiar — e nem de diarização, já que costuma ser a mesma pessoa. + +2. **Ênfase é relativa ao conjunto analisado.** Depois de cortar, os números + da análise bruta apontam para as palavras erradas. Sempre reanalise + (Fase 4) antes de escolher zooms. + +3. **Tempos sempre na mídia original.** Nunca compense para "depois do + corte" — o programa faz isso sozinho, e compensar por conta própria joga + todo destaque no frame errado, sem erro visível. + +## Ordem no sistema + +Sua edição roda **primeiro, na timeline intacta**. Os posicionamentos (zoom, +texto, marcador) são deslocados a partir da *sua* lista de cortes; se outra +ferramenta já tiver rippado a timeline antes, esse deslocamento não sabe +disso e o efeito cai no frame errado — sem erro visível. + +``` +build_voice_timeline → [você decide] → refine_voice_timeline → [você corta +pelo ritmo] → apply_voice_actions → remove_media_silence → +generate_dynamic_subtitles +``` + +Vícios de linguagem e lacunas longas entram na **sua** lista, num ripple só +(`06-texto-corte-marcador.md`). Silêncio fino e legendas vêm depois. + +`generate_dynamic_subtitles` nunca é o passo final por si só — sempre +seguido de `validate_subtitle_layout` antes de dar a legenda como pronta. +Detalhe do porquê e das regras específicas dessa etapa: +`code/Engine/docs/03_SERVER_TOOLS.md`, seção "Legendas dinâmicas". + +## Ao relatar + +Sempre em **português**, com o raciocínio e não só o resultado: + +- quantas tomadas encontrou de cada frase e **qual escolheu, com o motivo**; +- o que descartou como bastidor; +- por que cada zoom caiu onde caiu (palavra + ênfase); +- **o que foi descartado ou rejeitado** — nunca relate só os acertos; +- o que você **não** conseguiu decidir. Em dúvida entre duas tomadas, + marque as duas e deixe para o editor. diff --git a/.claude/skills/editar-por-voz/criterios/01-leitura-do-json.md b/.claude/skills/editar-por-voz/criterios/01-leitura-do-json.md new file mode 100644 index 0000000..0576d93 --- /dev/null +++ b/.claude/skills/editar-por-voz/criterios/01-leitura-do-json.md @@ -0,0 +1,71 @@ +# 01 — Leitura do JSON + +O arquivo `_voice_timeline.json` é a entrada de todo o trabalho. +Leia em camadas, de cima para baixo, e só desça quando precisar. + +## Camadas + +| Camada | O que traz | Para quê | +|---|---|---| +| `layers` | o que de fato rodou na análise | **leia primeiro** — ver `07-analise-incompleta.md` | +| `summary` | forma da peça, `peak_moments`, contagens | visão geral em poucos números | +| `speakers` | quem fala, % do tempo, frases de exemplo | identificar papéis | +| `segments` | cada fala com seus agregados | **onde você mais trabalha** | +| `segments[].words` | detalhe por palavra | achar o instante exato de um destaque | +| `scales` | o que cada número significa | documentação dentro do próprio arquivo | + +## Campos que decidem quase tudo + +**`gap_before`** — silêncio antes da fala, em segundos. É o mapa estrutural +da gravação: acima de ~3s (`take_boundary: true`) a câmera parou ou a +tomada recomeçou. Num material real de 3min17s isso identificou 6 +fronteiras, todas exatamente onde a pessoa recomeçava o roteiro. + +**`take_boundary`** — booleano derivado do `gap_before`. Use para agrupar +tomadas. + +**`emphasis`** (0–1) — índice combinado de energia, variação de tom, +variação de ritmo, pausa anterior e duração. **É relativo ao material +analisado**, nunca uma medida absoluta. Ver `02-enfase-e-reanalise.md`. + +**`energy`** (0–1) — intensidade relativa ao trecho mais alto da gravação. + +**`pitch_delta`** (0–1) — quanto o tom se afasta da média do falante. + +**`peak_emphasis`** e **`avg_energy`** (por segmento) — permitem julgar uma +frase inteira sem ler palavra por palavra. É por aqui que você avalia o +arco narrativo. + +**`energy_raw`** e **`pitch_hz`** — valores brutos, sem normalização. Não +use para decidir; existem para permitir a reanálise da Fase 2. + +## O que NÃO fazer + +- **Não recalcule** energia, tom ou ênfase. O sistema mede melhor e de + forma reprodutível. +- **Não reestime tempos "no olho".** Use os timestamps do JSON. +- **Não trate `emphasis` como valor absoluto.** Um 0,35 pode ser o pico de + uma gravação e ruído em outra. + +## O timestamp por palavra tem um viés conhecido + +O início de cada palavra vem sistematicamente **adiantado em ~0,3-0,5s** em +relação ao ataque real da fala — medido em material real com ffmpeg +(`astats`), consistente em 6 pontos do mesmo vídeo. O fim da palavra não +tem esse problema (erro de poucos centésimos). Causa: `word_timestamps` do +faster-whisper deriva por atenção cruzada, sem alinhamento forçado — ver +`05_EXPERIENCIAS.md`, entrada de 2026-08-19. + +Isso não é "reestimar no olho" — é um bug de medição na fonte, não um +julgamento seu. Na prática: + +- Ao posicionar um `zoom` cujo `start` precisa cair exatamente na palavra + (não uma frase inteira), some **+0,3 a +0,4s** ao timestamp do JSON antes + de decidir, ou confira com `ffmpeg -af astats` se a precisão importar + para o frame. +- **Não aplique essa correção a `gap_before` para decidir corte** — a régua + de silêncio (`06-texto-corte-marcador.md`) já é conservadora o bastante + para absorver esse erro; corrigir os dois ao mesmo tempo é redundante. +- Se um dia o pipeline ganhar alinhamento forçado (WhisperX), este aviso + perde a razão de existir — confira se `layers` ou a versão do documento + já indicam isso antes de aplicar o offset manualmente. diff --git a/.claude/skills/editar-por-voz/criterios/02-triagem-roteiro-vs-conversa.md b/.claude/skills/editar-por-voz/criterios/02-triagem-roteiro-vs-conversa.md new file mode 100644 index 0000000..be785dc --- /dev/null +++ b/.claude/skills/editar-por-voz/criterios/02-triagem-roteiro-vs-conversa.md @@ -0,0 +1,51 @@ +# 02 — Triagem: roteiro vs. conversa de bastidor + +**Primeira coisa a fazer, antes de qualquer decisão de efeito.** + +Material bruto de gravação quase nunca é uma tomada só. A pessoa lê o +roteiro, erra, conversa com a equipe e recomeça. + +## Por que isso é tarefa sua, e não do sistema + +No áudio essa separação é **invisível** — e pior: o índice de ênfase +*favorece* a conversa, que é mais solta e mais alta que o texto decorado. + +Caso real: a fala mais enfática de um vídeo inteiro (energia **1,00**, o +topo absoluto da gravação) era *"Amor, eu tô intacto!"*, dita para o marido +fora de quadro. Três das sete palavras de maior ênfase do vídeo vinham +dessa única frase de bastidor. + +Nenhum limiar acústico separa isso. O **texto** separa sem erro. + +> Atenção: isso também **não é diarização**. Num caso real, a pessoa da +> equipe estava fora do microfone — a diarização a ouvia, mas o Whisper não +> a transcrevia. As falas a descartar eram da **própria protagonista**: +> mesma voz, contexto diferente. "Quem fala" e "isso é tomada válida" são +> perguntas diferentes. + +## Descartar — conversa com a equipe + +Reconhece-se pelo **conteúdo**: + +- **vocativo para alguém da sala** — *"Amor, eu tô intacto!"* +- **pergunta operacional** — *"Posso começar da mastopexia?"*, + *"E aí, continua?"*, *"Mas eu vou ter que falar tudo de novo?"* +- **instrução técnica** — *"Só clica aí agora na tela."*, *"Aumenta."* +- **comentário sobre a própria gravação** — *"Vou falar só a última frase, + só um pouquinho, não pegou?"* + +## Descartar — frases interrompidas + +Texto que morre no meio, tipicamente em reticências ou emendando numa +pergunta: + +- *"E tudo isso associado à medida..."* +- *"Aquela mama com um formato mais estruturado, com o colo que..."* +- *"de pele..."* + +## Sinais estruturais que ajudam + +Use `take_boundary` para achar onde cada tomada recomeça. Num material +real, as fronteiras (gaps de 3,6s a 19,8s) caíam exatamente nos pontos onde +a médica reiniciava o roteiro — inclusive nas duas retomadas da frase de +abertura. diff --git a/.claude/skills/editar-por-voz/criterios/03-escolha-da-melhor-tomada.md b/.claude/skills/editar-por-voz/criterios/03-escolha-da-melhor-tomada.md new file mode 100644 index 0000000..4c74910 --- /dev/null +++ b/.claude/skills/editar-por-voz/criterios/03-escolha-da-melhor-tomada.md @@ -0,0 +1,60 @@ +# 03 — Escolha da melhor tomada + +A mesma frase costuma aparecer 2, 3, 4 vezes. Seu trabalho é ficar com +**uma**. + +## Como agrupar + +1. Use `take_boundary` para localizar onde cada tomada recomeça. +2. Agrupe as repetições **pelo texto**, não pelo tempo — a mesma frase + reaparece em pontos distantes da gravação. Num caso real, a abertura + *"Aquela mama com um formato mais estruturado"* apareceu aos 2,0s, 64,9s + e 86,4s. + +## Critérios, nesta ordem + +### 1. Completa +Não morre no meio, não emenda numa pergunta. Uma tomada incompleta está +descartada por definição, mesmo que a dicção seja ótima. + +### 2. Dicção limpa +Sem tropeço, sem repetição de palavra, sem vício de linguagem. **Compare os +textos lado a lado:** + +| Tomada 1 | Tomada 3 | Escolha | +|---|---|---| +| *"isso é desejo de muitas mulheres"* | *"**aí** isso é desejo de muitas mulheres"* | Tomada 1 | + +### 3. Formulação melhor +Quando as duas estão limpas, prefira a mais direta — normalmente a última, +porque é onde a pessoa já se ajustou: + +| Antes | Depois | Escolha | +|---|---|---| +| *"a gente **faz a inserção de** próteses"* | *"a gente **insere** próteses"* | a segunda | +| *"reestrutura a mama"* | *"reestrutura a **sua** mama"* | a segunda | + +### 4. Entrega +**Só então** desempate por `avg_energy` / `peak_emphasis`. + +Quando duas tomadas têm texto **idêntico palavra por palavra**, aí a +energia decide sozinha — é o único sinal disponível. Caso real: o fecho +tinha duas tomadas iguais, energia **0,38** e **0,19**. A de 0,38 é a boa, +e o texto sozinho jamais diria isso. + +## Regra de ouro + +A **última** tomada costuma ser a melhor — é onde a pessoa acertou. Mas +**confirme lendo o texto**; nunca assuma. + +## Quando estiver em dúvida + +Não decida no escuro. Coloque um `marker` nas duas candidatas, explique a +dúvida no `reason`, e deixe a escolha para o editor humano. + +## Continuidade + +Ao montar o corte final você pode misturar blocos de tomadas diferentes — +abertura da tomada 1, corpo da tomada 3. Isso é normal. Mas **avise nas +emendas**: coloque um `marker` em cada junção para o editor conferir se o +enquadramento e a posição da pessoa combinam. diff --git a/.claude/skills/editar-por-voz/criterios/04-reanalise-do-material-restante.md b/.claude/skills/editar-por-voz/criterios/04-reanalise-do-material-restante.md new file mode 100644 index 0000000..f12dd1f --- /dev/null +++ b/.claude/skills/editar-por-voz/criterios/04-reanalise-do-material-restante.md @@ -0,0 +1,70 @@ +# 04 — Reanálise do material que sobrou + +**Não escolha zooms com os números da análise bruta.** + +## O problema + +Ênfase e energia são **relativas ao conjunto analisado**. Energia é +normalizada contra o momento mais alto da gravação; ênfase deriva dela. + +Se esse momento mais alto foi cortado — uma piada, um grito, uma conversa +de bastidor — tudo que sobrou continua pontuado contra uma referência que o +espectador **nunca verá**. As notas do corte final ficam artificialmente +comprimidas, e o ranking aponta para as palavras erradas. + +Caso real: o pico do vídeo era *"Amor, eu tô intacto!"* (energia 1,00), +descartado na triagem. Todo o material restante estava sendo medido contra +ele. + +## A solução + +Depois de definir os cortes, renormalize sobre os sobreviventes. Chame a +ferramenta **`refine_voice_timeline`**, passando a mídia e a lista de +cortes que você já decidiu: + +``` +refine_voice_timeline(media_path, cuts=[{start, end}, ...], min_gap=8.0) +``` + +Ela devolve, numa chamada só, a comparação bruto × sobreviventes, os picos +re-ranqueados e os candidatos a zoom. É barata: renormaliza os números já +medidos, sem reabrir o áudio. + +Efeito medido no mesmo material: + +| | Bruto | Só o que sobrou | +|---|---|---| +| Ênfase média | 0,179 | **0,197** | +| *"Aquela"* | 0,39 | **0,42** | +| *"mastopexia"* | 0,26 | **0,34** | +| *"devolver"* | — | **0,35** | + +*"mastopexia"* só virou candidata legítima depois da reanálise. + +## As janelas que ela propõe + +A seção **Zoom Candidates** da resposta já vem com três coisas resolvidas: + +1. **Pega a palavra de conteúdo mais enfática de cada frase.** Artigos e + conectivos são filtrados — um *"a"* falado alto continua sendo um artigo. + Sem esse filtro, o ranking bruto apontava para "o", "a", "eu": picos de + *entrega*, não de *sentido*. +2. **Estende a janela até o fim da frase**, não do segmento (ver + `05-zoom.md`). +3. **Mantém distância mínima** entre zooms. + +São **candidatos, não obrigações.** Corte a lista pelo ritmo +(`07-ritmo.md`). Os tempos continuam na mídia original — vão direto para +`apply_voice_actions`. + +`max_zooms` limita a lista, mas prefira cortá-la você mesmo: o corte por +ritmo é decisão editorial, não um teto numérico. + +## Princípio geral + +> "Qual o momento mais forte da **gravação**?" e "qual o momento mais forte +> do **vídeo final**?" são perguntas diferentes sempre que a métrica for +> relativa. + +Toda métrica normalizada precisa ser recalculada quando o conjunto muda — +senão ela responde a pergunta errada, silenciosamente. diff --git a/.claude/skills/editar-por-voz/criterios/05-zoom.md b/.claude/skills/editar-por-voz/criterios/05-zoom.md new file mode 100644 index 0000000..e533290 --- /dev/null +++ b/.claude/skills/editar-por-voz/criterios/05-zoom.md @@ -0,0 +1,71 @@ +# 05 — Zoom (punch-in) + +## Quando usar + +No momento em que o argumento vira. Um pico acústico só merece zoom se for +também um pico **de sentido**. + +Palavra gritada sem peso narrativo não ganha nada — e isso inclui os picos +que caem em artigos e conectivos, que são picos de entrega, não de conteúdo. + +## A janela + +**`start`** — na palavra de ênfase. + +**`end`** — no **fim da frase**. A frase inteira, não o fim do segmento da +transcrição. + +O Whisper corta frases no meio, por respiração e não por gramática: + +> *"Aquela mama com um formato mais estruturado, que valoriza o seu colo, +> que dá aquele ar"* **|** *"de elegância, isso é desejo de muitas mulheres, +> né?"* + +Soltar o zoom no fim do primeiro segmento libera **no meio do pensamento** — +é o que faz um punch-in parecer arbitrário. `suggest_zoom_windows` já +estende até a pontuação final (`.` `!` `?` `…`), e nunca atravessa uma +fronteira de tomada. + +## A forma — o programa decide sozinho + +Você escolhe `start` e `end`; a forma sai da posição da janela dentro do +trecho: + +| Situação | Comportamento | Por quê | +|---|---|---| +| Começa a **>0,5s** do início do trecho | entrada rápida (~0,25s) | o movimento chega junto com a palavra | +| Começa a **≤0,5s** do início | **entra já ampliado, sem transição** | o corte já foi a transição; uma rampa ali lê como a imagem se acomodando | +| Frase termina no meio do trecho | **saída seca**, 1 frame | volta ao enquadramento sem chamar atenção | +| Frase termina a **≤1s** do corte | **não volta** — segura até o corte | o próximo trecho já abre no enquadramento dele; voltar antes é movimento desperdiçado | + +Os limiares são diferentes de propósito: no fim o corte esconde um retorno +inacabado, mas no início a rampa é visível desde o primeiro frame. + +Para forçar manualmente, existem `start_at_peak` e `hold_at_end` — mas o +automático acerta na quase totalidade dos casos. + +## Escala + +| Valor | Uso | +|---|---| +| 1,15 | sutil | +| 1,18 – 1,3 | padrão | +| 1,5 | forte | + +Em vídeo institucional, fique na faixa baixa. Acima de 3,0 é rejeitado. + +O zoom é **relativo ao enquadramento existente**: se o clipe já tem escala +1,77 (material gravado de lado e reenquadrado), um zoom 1,18 anima de 1,77 +para 2,09 e preserva rotação e posição. + +## Dois zooms no mesmo clipe + +Depois do corte, dois picos que você escolheu podem cair no **mesmo** +trecho sobrevivente (nenhum corte os separou em clipes distintos) — é +comum quando a corrida limpa de uma tomada é longa. O sistema resolve isso +sozinho, e a regra é a mesma que rege o resto: janelas **distantes** +empilham (os dois zooms convivem, cada um voltando ao enquadramento real +entre um e outro); janelas que **se sobrepõem** substituem (é o mesmo +evento sendo reajustado, não dois). Você não precisa calcular isso na +hora de decidir — só respeitar o `min_gap` de `07-ritmo.md`, que já +garante que dois zooms escolhidos por você nunca se sobrepõem. diff --git a/.claude/skills/editar-por-voz/criterios/06-texto-corte-marcador.md b/.claude/skills/editar-por-voz/criterios/06-texto-corte-marcador.md new file mode 100644 index 0000000..ad9dd58 --- /dev/null +++ b/.claude/skills/editar-por-voz/criterios/06-texto-corte-marcador.md @@ -0,0 +1,88 @@ +# 06 — Texto, corte e marcador + +## Texto + +Para fixar um **conceito, número ou nome** que o espectador precisa reter. + +- Use a palavra **dita**, não uma paráfrase. +- Curta, em caixa alta. Até 120 caracteres (é truncado além disso). +- Uma por frase, no máximo. + +**Não legende a frase inteira.** Para isso existe +`generate_dynamic_subtitles`, que é outra ferramenta e outro propósito. + +Boas candidatas são as palavras-chave que sobram depois de filtrar as +funcionais — num caso real: *mastopexia*, *flacidez*, *próteses*, +*devolver*, *desejo*. + +## Corte + +Digressão, repetição, frase abandonada, conversa de bastidor, tomada pior — +e mais duas coisas que **são** seu trabalho, ao contrário do que parece. + +### Vícios de linguagem entram na sua lista + +Não delegue para `remove_filler_words`. Você já está percorrendo palavra por +palavra na triagem; marcar as muletas é uma linha a mais, sem custo. E você +tem o que a lista fixa não tem: **contexto**. + +Um *"tipo"* em *"tipo assim, sabe"* é muleta. Em *"esse tipo de cirurgia"* +é a palavra principal. Um *"não não não"* pode ser gagueira ou ênfase. A +lista fixa não distingue; você distingue. + +### Lacunas longas entram na sua lista — curtas, nunca + +**O tamanho da lacuna muda o que ela é.** A régua está medida em material +real (`pause_weight()` em `emphasis.py`, `TAKE_BOUNDARY_GAP` em +`voice_timeline.py`): + +| `pause_before` | O que é | O que fazer | +|---|---|---| +| até ~1,5s | o falante montando a frase — **isso É a ênfase** | **nunca cortar** | +| 1,5–3s | zona cinza | julgue pela frase | +| acima de 3s | troca de tomada, ar morto, outra pessoa falando | **cortar** | + +Cortar a pausa curta é o erro grave: ela é uma das cinco entradas do índice +de ênfase, então você estaria apagando justamente a batida que faz a palavra +seguinte pontuar alto. Uma frase fluida não se aperta. + +Acima de 3s a pausa deixa de contar como ênfase por construção — medido em +material real, lacunas de 6–9s rankeavam como os momentos mais enfáticos da +gravação só porque a escala saturava. + +### O que continua NÃO sendo seu trabalho + +| Tarefa | Ferramenta | Por quê | +|---|---|---| +| Apertar o ar **dentro** da fala | `remove_media_silence` | Lê o áudio real com ffmpeg; você só tem os intervalos entre palavras transcritas | + +E cuidado: **ausência de fala não é ausência de som.** Respiração, riso, +suspiro, a reação depois da frase — nada disso vira palavra, então aparece +como lacuna, e às vezes é o melhor frame do vídeo. Lacuna longa é candidata +a corte, não corte automático. + +`remove_media_silence` roda **depois** de `apply_voice_actions`, como +acabamento opcional, sobre o material que sobrou. + +**Antes de rodar `remove_media_silence` sobre o corte final, sempre rode a +detecção primeiro** (sem aplicar) e leia os spans um a um contra a régua +acima. O detector corta por limiar de dB — ele não sabe distinguir "batida +de 0,8s entre duas frases", que a régua protege, de "ar morto de emenda", +que deveria ser apertado. Aplicar direto, sem essa checagem, é o mesmo erro +de cortar pausa curta, só que por outra ferramenta. + +## Marcador + +Quando você quer **sinalizar para o editor humano decidir**, em vez de +decidir por ele. + +Use em: + +- **Emendas entre tomadas** — sempre. O editor precisa conferir se o + enquadramento e a posição da pessoa combinam na junção. +- **Dúvida entre duas tomadas** — marque as duas, explique no `reason`. +- **Momentos que talvez mereçam efeito** mas que você não tem confiança + para decidir. + +Marcador é um **ponto**, não um trecho: sobrevive mesmo encostado na borda +de um corte, o que é justamente o caso das emendas. diff --git a/.claude/skills/editar-por-voz/criterios/07-ritmo.md b/.claude/skills/editar-por-voz/criterios/07-ritmo.md new file mode 100644 index 0000000..1b0741e --- /dev/null +++ b/.claude/skills/editar-por-voz/criterios/07-ritmo.md @@ -0,0 +1,37 @@ +# 07 — Ritmo + +**O erro mais comum é efeito demais.** Cansa mais que efeito de menos, e +denuncia edição automática. + +## Limites + +| Regra | Valor | +|---|---| +| Distância mínima entre dois zooms | **8–10 segundos** | +| Zooms por minuto de vídeo | **2 a 4** (teto) | +| Zoom + texto no mesmo instante | só com motivo claro | + +Se dois picos estiverem colados, **escolha o mais forte e abra mão do +outro**. Não tente encaixar os dois. + +## Candidatos ≠ obrigações + +`suggest_zoom_windows` devolve uma lista de candidatos. Normalmente você usa +uma **fração** dela. + +Caso real: num corte de 47,6s a ferramenta sugeriu **5** janelas. O certo +foram **3** — 5 violaria o teto de 2–4 por minuto. Ficaram a abertura, o +termo central e o fecho; as duas descartadas eram frases de apoio. + +O mesmo vale para `peak_moments` no `summary`: é lista de candidatos. + +## Como escolher quais manter + +Quando precisar cortar a lista, priorize por **função narrativa**, não por +nota: + +1. **A abertura** — prende o espectador. +2. **O conceito central** — o termo que o vídeo existe para explicar. +3. **O fecho** — a frase que fica. + +Só depois disso, as frases de apoio, por ordem de ênfase. diff --git a/.claude/skills/editar-por-voz/criterios/08-formato-de-saida.md b/.claude/skills/editar-por-voz/criterios/08-formato-de-saida.md new file mode 100644 index 0000000..5e1ebd4 --- /dev/null +++ b/.claude/skills/editar-por-voz/criterios/08-formato-de-saida.md @@ -0,0 +1,65 @@ +# 08 — Formato de saída + +O produto do seu trabalho é **este JSON**. É ele que vai para o programa +gerar o FCPXML. Você nunca escreve XML. + +## Estrutura + +```json +{ + "source": "0E6A8290.mp4", + "actions": [ + {"kind": "cut", "start": 21.9, "end": 127.6, + "reason": "tomadas descartadas, frases interrompidas e conversa com a equipe"}, + {"kind": "zoom", "start": 2.0, "end": 10.7, + "params": {"scale": 1.15}, "reason": "abertura: \"Aquela mama\" (ênfase 0.42)"}, + {"kind": "text", "start": 127.7, "end": 129.0, + "params": {"content": "MASTOPEXIA"}, "reason": "fixa o termo central"}, + {"kind": "marker", "start": 21.85, "end": 22.0, + "reason": "EMENDA 1 — conferir junção entre tomadas"} + ] +} +``` + +## Regras + +### 1. Tempos em segundos da mídia ORIGINAL +Exatamente como aparecem no `voice_timeline.json`. + +**Nunca compense para "depois do corte".** O programa faz esse deslocamento +sozinho: ele resolve os cortes primeiro e reposiciona todo o resto. Se você +compensar por conta própria, **todo destaque cai no frame errado** — e o +erro é silencioso. + +### 2. `end` sempre maior que `start` +Ambos ≥ 0. Um `end <= start` é rejeitado. + +### 3. Tipos +`cut` · `zoom` · `text` · `marker` + +### 4. Parâmetros por tipo + +| Tipo | `params` | +|---|---| +| `cut` | nenhum | +| `zoom` | `scale` entre 1.0 e 3.0 (padrão 1.3 se omitido) | +| `text` | `content` **obrigatório**, até 120 caracteres | +| `marker` | opcional: `content` vira o nome do marcador | + +### 5. `reason` — sempre preencha +É o que o usuário lê para revisar sua decisão, e o que te obriga a **ter** +uma. Um `reason` vazio é sinal de decisão sem critério. + +Inclua o dado que embasou: *"abertura: 'Aquela mama' (ênfase 0.42)"* é útil; +*"zoom"* não é. + +## Como o programa trata erros + +- **Ação inválida** → rejeitada e reportada **individualmente**. Uma linha + malformada nunca derruba as outras. +- **Ação apontando para material cortado** → descartada e reportada, nunca + deslizada para o conteúdo vizinho. +- **Ação fora da mídia** → reportada como não colocada. + +Você recebe o relatório dos três casos. **Repasse ao usuário** — nunca +relate só os acertos. diff --git a/.claude/skills/editar-por-voz/criterios/09-analise-incompleta.md b/.claude/skills/editar-por-voz/criterios/09-analise-incompleta.md new file mode 100644 index 0000000..9cef960 --- /dev/null +++ b/.claude/skills/editar-por-voz/criterios/09-analise-incompleta.md @@ -0,0 +1,48 @@ +# 09 — Quando a análise veio incompleta + +O bloco `layers` no topo do JSON diz **o que de fato rodou**. Leia antes de +qualquer outra coisa. + +```json +"layers": {"transcript": true, "acoustics": false, "speakers": false} +``` + +## Por que esse bloco existe + +Fala monótona e acústica que não carregou deixam **os mesmos zeros** nos +dados. Sem o `layers`, é impossível distinguir "esta pessoa fala de forma +uniforme" de "a análise acústica falhou". + +## Os casos + +### `acoustics: false` +Todos os valores acústicos são 0. **Você não tem ênfase real.** + +- Decida só pelo texto. +- **Avise o usuário** explicitamente. +- Prefira `marker` a `zoom` — sinalize em vez de decidir. + +Causa comum: o componente librosa não está instalado, ou o ffmpeg não +conseguiu extrair o áudio do container. + +### `speakers: false` num vídeo com várias pessoas +A diarização não rodou — falta o token do HuggingFace (aba Modelos do app). + +- Avise antes de tratar tudo como uma voz só. +- Lembre que isso **não impede** a triagem roteiro/conversa, que é feita + pelo texto (ver `02-triagem-roteiro-vs-conversa.md`). + +### `peak_count: 0` +Nada cruzou o piso de ênfase. Duas causas possíveis: + +1. A fala é uniforme mesmo — material sem picos. +2. O limiar está alto para esse material. + +Sugira ajustar em **Análise de Voz** no app. **Não force destaques +inexistentes** só para entregar alguma coisa. + +## Regra geral + +Não finja precisão que você não tem. Uma edição entregue com a ressalva +certa é útil; uma entregue como se estivesse completa, quando metade dos +dados faltou, custa a confiança do usuário no sistema inteiro. diff --git a/CLAUDE.md b/CLAUDE.md index 59c1bfa..b3e91e7 100755 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -9,7 +9,7 @@ normalmente; a regra é sobre a comunicação com o usuário. ## What This Is -MCP server that reads/writes Final Cut Pro XML (FCPXML) files. 62 tools for timeline analysis, batch editing, QC, generation, multi-track support, media relink, NLE export, transcript-based editing (local Whisper), and LIVE FCP control (push_to_fcp / list_fcp_libraries via Apple events). Reads FCPXML 1.8–1.14 (incl. `.fcpxmld` bundles with sidecar preservation), writes 1.13 by default. Dual-mode (XML + Live) direction: `code/docs/CAPABILITY-AUDIT-2026-06.md`. +MCP server that reads/writes Final Cut Pro XML (FCPXML) files. 73 tools for timeline analysis, batch editing, QC, generation, multi-track support, media relink, NLE export, transcript-based editing (local Whisper), and LIVE FCP control (push_to_fcp / list_fcp_libraries via Apple events). Reads FCPXML 1.8–1.14 (incl. `.fcpxmld` bundles with sidecar preservation), writes 1.13 by default. Dual-mode (XML + Live) direction: `code/docs/CAPABILITY-AUDIT-2026-06.md`. ## Architecture @@ -88,7 +88,7 @@ CI runs both on every push to main. If either fails, the commit gets an X on Git ## Testing -1032 tests across 24 files. `test_models.py` covers TimeValue arithmetic, Timecode parsing/formatting, Clip properties, validation models, and Timeline helpers. `test_writer.py` covers insert_clip, add_marker (all types), trim_clip, delete_clip, split_clip, and change_speed operations. `test_server.py` covers MCP tool handlers, parsers, and dispatch. `test_rough_cut.py` covers RoughCutGenerator. `test_features_v05.py` covers connected clips, roles, timeline diff, reformat, silence detection, export, and backward compatibility. `test_marker_pipeline.py` covers build_marker_element shared builder, batch auto-modes, clip index duplicate-name behavior, and write_fcpxml output format. `test_refactored_helpers.py` covers _index_elements, _iter_spine_clips, _find_spine_clip_at_seconds, _resolve_clip_duration, _make_asset_clip, _format_batch_result, and serialize_xml edge cases. `test_transcribe.py` covers phrase/filler span matching, range merge/invert algebra, whisper graceful degradation, and transcript-driven handler cuts against cached transcripts. `test_media_intel.py` covers silencedetect stderr parsing, source-to-timeline mapping, parameter bounds, and real-WAV ffmpeg integration (skips without ffmpeg; CI installs it). Tests use `examples/sample.fcpxml` as fixture data and inline XML fixtures. Tests create temp files and clean up after. +1342 tests across 34 files. `test_models.py` covers TimeValue arithmetic, Timecode parsing/formatting, Clip properties, validation models, and Timeline helpers. `test_writer.py` covers insert_clip, add_marker (all types), trim_clip, delete_clip, split_clip, and change_speed operations. `test_server.py` covers MCP tool handlers, parsers, and dispatch. `test_rough_cut.py` covers RoughCutGenerator. `test_features_v05.py` covers connected clips, roles, timeline diff, reformat, silence detection, export, and backward compatibility. `test_marker_pipeline.py` covers build_marker_element shared builder, batch auto-modes, clip index duplicate-name behavior, and write_fcpxml output format. `test_refactored_helpers.py` covers _index_elements, _iter_spine_clips, _find_spine_clip_at_seconds, _resolve_clip_duration, _make_asset_clip, _format_batch_result, and serialize_xml edge cases. `test_transcribe.py` covers phrase/filler span matching, range merge/invert algebra, whisper graceful degradation, and transcript-driven handler cuts against cached transcripts. `test_media_intel.py` covers silencedetect stderr parsing, source-to-timeline mapping, parameter bounds, and real-WAV ffmpeg integration (skips without ffmpeg; CI installs it). Tests use `examples/sample.fcpxml` as fixture data and inline XML fixtures. Tests create temp files and clean up after. ## FCPXML Gotchas diff --git a/admin/models_api.py b/admin/models_api.py index b004e4b..3052586 100644 --- a/admin/models_api.py +++ b/admin/models_api.py @@ -41,6 +41,34 @@ Commands: clips, cuts, connected, markers}]} or {"ok": false, "error": "..."} + analyze_voice {"path": "...", "output_dir": "...", "model": "...", + "language": "pt"|"auto"|null, "hf_token": "..."|null, + "num_speakers": ""|null} + Build the voice timeline (transcript+diarization+acoustics) for + every unique source media — analysis only, writes _voice_timeline.json + next to each media, `path` passes through unchanged. Meant as one + entry in the batch operations list (see processBatchStep), so + `refine_voice_timeline` never has to reopen the audio later. + -> {"ok": true, "path": "...", "message": "..."} or {"ok": false, "error": "..."} + + dynamic_subtitle_config {} + -> {"ok": true, "band_height", "block_center_y", "line_gap", "font", + "font_size", "emphasis_font", "emphasis_face", "emphasis_size", + "active_color", "emphasis_color", "text_scale"} + + set_dynamic_subtitle_config {} + Persists only the given fields to ~/.fcp-mcp-server/config.json. + generate_dynamic_subtitles reads this as its own fallback default. + -> {"ok": true, } + + silence_config {} + -> {"ok": true, "noise_db": -30.0, "min_silence": 0.5, "padding": 0.05} + + set_silence_config {"noise_db": -30.0, "min_silence": 0.5, "padding": 0.05} + Persists only the given fields. detect_media_silence and + remove_media_silence read this as their own fallback default. + -> {"ok": true, } + transcribe {"path": "...", "model": "small", "language": "pt"|null, "hf_token": "..."|null, "num_speakers": ""|null} -> JSON-lines: @@ -76,7 +104,9 @@ Commands: generate_dynamic_subtitles {"path": "...", "clip_name": "..."|null, "band_height": 0.22, "block_center_y": -167, "font": "Helvetica Neue", "font_size": 128, - "active_color": "1 1 1 1", "inactive_color": "0.7 0.7 0.7 1", + "emphasis_font": "Playfair Display", + "emphasis_face": "Medium Italic", "emphasis_size": 265, + "active_color": "1 1 1 1", "emphasis_color": "1 1 1 1", "model": "small", "language": "pt"|null} -> {"ok": true, "path": "..._dynamic_subtitles.fcpxml", "message": "..."} or {"ok": false, "error": "..."} @@ -89,6 +119,16 @@ Commands: -> {"ok": true, "diarization": bool, "diarization_message": "...", "num_speakers": "..."} + voice_analysis + -> {"ok": true, "energy_threshold": 0.5, "emphasis_threshold": 0.85, + "emphasis_weights": {...}, "emotion_enabled": false, + "emotion_sensitivity": 0.5} + + set_voice_analysis {"energy_threshold": 0.6, "emphasis_threshold": 0.9, + "emphasis_weights": {"energy": 0.4}|null, + "emotion_enabled": true, "emotion_sensitivity": 0.5} + -> same shape as voice_analysis (only given fields change) + Exit code 0 on success, 1 on error. """ @@ -122,16 +162,24 @@ from fcpxml.model_manager import ( # noqa: E402 is_model_downloaded, list_installed_models, load_catalog, + load_dynamic_subtitle_config, load_hf_token, load_num_speakers, + load_project_config, load_selected_model, + load_silence_config, load_transcript_language, + load_voice_analysis_config, model_cache_dir, + save_dynamic_subtitle_config, save_hf_token, save_models_dir, save_num_speakers, + save_project_config, save_selected_model, + save_silence_config, save_transcript_language, + save_voice_analysis_config, ) from fcpxml.parser import parse_fcpxml # noqa: E402 from fcpxml.transcribe import transcribe # noqa: E402 @@ -600,14 +648,23 @@ def cmd_transcribe(args: dict) -> int: total = len(media_paths) results: list[dict] = [] for i, mp in enumerate(media_paths, 1): - _emit({"type": "progress", "fraction": i / total, "stage": f"Transcrevendo {Path(mp).name} ({i}/{total})…"}) + stage = f"Transcrevendo {Path(mp).name} ({i}/{total})…" + _emit({"type": "progress", "fraction": (i - 1) / total, "stage": stage}) json_path = _transcript_json_path(mp, output_dir) cached = _load_cached_transcript(json_path) if cached is not None: + _emit({"type": "progress", "fraction": i / total, "stage": stage}) results.append(_result_row(mp, cached)) continue - data = transcribe(mp, model_size=model, language=language) + def _on_progress(file_fraction: float, _i: int = i, _stage: str = stage) -> None: + # Blend this file's own progress into the overall fraction so a + # single-media project doesn't jump straight to 100% before the + # actual (slow) decoding work has even started. + overall = (_i - 1 + file_fraction) / total + _emit({"type": "progress", "fraction": overall, "stage": _stage}) + + data = transcribe(mp, model_size=model, language=language, progress_cb=_on_progress) if data is None: _emit({"type": "error", "message": f"Não foi possível transcrever: {Path(mp).name}"}) return 1 @@ -638,6 +695,65 @@ def cmd_transcribe(args: dict) -> int: return 0 +def cmd_analyze_voice(args: dict) -> int: + """Build the voice timeline (transcript+diarization+acoustics -> emphasis) + for every unique source media in the project, so `refine_voice_timeline` + and friends have something to read without ever reopening the audio. + + Analysis only — writes _voice_timeline.json next to each media, doesn't + touch the project XML. `path` passes through unchanged so it composes + with the other batch steps (silence removal, captions) regardless of + where in the list it runs. + """ + path = str(args.get("path", "")) + if not path or not Path(path).exists(): + _emit({"ok": False, "error": "Arquivo de projeto não encontrado."}) + return 1 + + model = str(args.get("model", "") or load_selected_model() or "") + language = args.get("language") + if language is None: + language = load_transcript_language() + if language == "auto": + language = None + token = str(args.get("hf_token") or load_hf_token() or "") + num_speakers = str(args.get("num_speakers") or load_num_speakers() or "") + + try: + proj = parse_fcpxml(path) + except Exception as exc: + _emit({"ok": False, "error": f"Erro ao ler o projeto: {exc}"}) + return 1 + tl = proj.primary_timeline or (proj.timelines[0] if proj.timelines else None) + media_paths: list[str] = [] + if tl is not None: + for clip in getattr(tl, "clips", []): + mp = media_src_to_path(clip.media_path or "") + if mp and Path(mp).is_file() and mp not in media_paths: + media_paths.append(mp) + if not media_paths: + _emit({"ok": False, "error": "Nenhum arquivo de mídia acessível encontrado."}) + return 1 + + from server import handle_build_voice_timeline + + messages: list[str] = [] + for mp in media_paths: + try: + contents = asyncio.run(handle_build_voice_timeline({ + "media_path": mp, "model": model, "language": language, + "hf_token": token, "num_speakers": num_speakers, + "output_dir": args.get("output_dir"), + })) + except Exception as exc: + _emit({"ok": False, "error": f"Falha analisando {Path(mp).name}: {exc}"}) + return 1 + messages.append("\n".join(getattr(c, "text", str(c)) for c in contents)) + + _emit({"ok": True, "path": path, "message": "\n\n---\n\n".join(messages)}) + return 0 + + def cmd_export_srt(args: dict) -> int: """Write a captions .srt synced to the edited timeline. @@ -822,6 +938,135 @@ def cmd_set_diarization(args: dict) -> int: return 0 +def cmd_voice_analysis(args: dict) -> int: + """Read the persisted voice-analysis settings (energy/emphasis/emotion).""" + _emit({"ok": True, **load_voice_analysis_config()}) + return 0 + + +def cmd_set_voice_analysis(args: dict) -> int: + """Persist voice-analysis settings. Only the given fields change.""" + weights = args.get("emphasis_weights") + config = save_voice_analysis_config( + energy_threshold=args.get("energy_threshold"), + emphasis_weights=weights if isinstance(weights, dict) else None, + emphasis_threshold=args.get("emphasis_threshold"), + emotion_enabled=args.get("emotion_enabled"), + emotion_sensitivity=args.get("emotion_sensitivity"), + ) + _emit({"ok": True, **config}) + return 0 + + +def cmd_dynamic_subtitle_config(args: dict) -> int: + """Read the persisted dynamic-subtitle style (font, size, color, layout).""" + _emit({"ok": True, **load_dynamic_subtitle_config()}) + return 0 + + +def cmd_set_dynamic_subtitle_config(args: dict) -> int: + """Persist dynamic-subtitle style fields. Only the given fields change.""" + config = save_dynamic_subtitle_config(**{ + k: args.get(k) for k in ( + "band_height", "block_center_y", "line_gap", "font", "font_size", + "emphasis_font", "emphasis_face", "emphasis_size", + "active_color", "emphasis_color", "text_scale", + ) + }) + _emit({"ok": True, **config}) + return 0 + + +def cmd_silence_config(args: dict) -> int: + """Read the persisted silence thresholds (noise floor, duration, padding).""" + _emit({"ok": True, **load_silence_config()}) + return 0 + + +def cmd_set_silence_config(args: dict) -> int: + """Persist silence thresholds. Only the given fields change.""" + config = save_silence_config( + noise_db=args.get("noise_db"), + min_silence=args.get("min_silence"), + padding=args.get("padding"), + ) + _emit({"ok": True, **config}) + return 0 + + +def cmd_apply_voice_actions(args: dict) -> int: + """Apply a decision list (cuts/zooms/texts/markers) to the project XML. + + The list is produced by a model reading the _voice_timeline.json — this + is the step that turns those decisions into an edit, and the one the + batch chain was missing: without it the app could measure the voice and + caption the result, but never cut by it. + + `actions_path` points at the JSON; either a bare list or the + ``{"actions": [...]}`` wrapper the skill emits is accepted. Times stay in + ORIGINAL source seconds — the handler resolves cuts first and shifts + everything else itself. + """ + path = str(args.get("path", "")) + if not path or not Path(path).exists(): + _emit({"ok": False, "error": "Arquivo de projeto não encontrado."}) + return 1 + + actions = args.get("actions") + if actions is None: + actions_path = str(args.get("actions_path", "")) + if not actions_path or not Path(actions_path).exists(): + _emit({"ok": False, "error": "Arquivo de decisões (JSON) não encontrado."}) + return 1 + try: + with open(actions_path, encoding="utf-8") as fh: + loaded = json.load(fh) + except (OSError, ValueError) as exc: + _emit({"ok": False, "error": f"Erro ao ler as decisões: {exc}"}) + return 1 + actions = loaded.get("actions") if isinstance(loaded, dict) else loaded + + if not isinstance(actions, list) or not actions: + _emit({"ok": False, "error": "A lista de decisões está vazia ou malformada."}) + return 1 + + from server import handle_apply_voice_actions + + try: + contents = asyncio.run(handle_apply_voice_actions({ + "filepath": path, + "actions": actions, + "output_dir": args.get("output_dir"), + })) + except Exception as exc: + _emit({"ok": False, "error": f"Falha ao aplicar as decisões: {exc}"}) + return 1 + + message = "\n".join(getattr(c, "text", str(c)) for c in contents) + # The handler reports dropped/rejected actions individually; hand the + # whole report back so the app can surface them instead of only the count. + out_path = path + for line in message.splitlines(): + if line.startswith("- **Saved to**:"): + out_path = line.split("`")[1] if "`" in line else path + break + _emit({"ok": True, "path": out_path, "message": message}) + return 0 + + +def cmd_project_config(args: dict) -> int: + """Read the last project folder/file the app was working on.""" + _emit({"ok": True, **load_project_config()}) + return 0 + + +def cmd_set_project_config(args: dict) -> int: + """Persist the last project folder/file. Only the given fields change.""" + config = save_project_config(folder=args.get("folder"), file=args.get("file")) + _emit({"ok": True, **config}) + return 0 + + def _result_row(mp: str, data: dict) -> dict: words = data.get("words", []) preview = (data.get("text", "") or "")[:160] @@ -875,6 +1120,16 @@ def main() -> int: "zoom_segments": cmd_zoom_segments, "rename_speakers": cmd_rename_speakers, "set_diarization": cmd_set_diarization, + "voice_analysis": cmd_voice_analysis, + "set_voice_analysis": cmd_set_voice_analysis, + "analyze_voice": cmd_analyze_voice, + "dynamic_subtitle_config": cmd_dynamic_subtitle_config, + "set_dynamic_subtitle_config": cmd_set_dynamic_subtitle_config, + "apply_voice_actions": cmd_apply_voice_actions, + "project_config": cmd_project_config, + "set_project_config": cmd_set_project_config, + "silence_config": cmd_silence_config, + "set_silence_config": cmd_set_silence_config, } handler = handlers.get(command) if handler is None: diff --git a/code/Engine/docs/01_ARCHITECTURE.md b/code/Engine/docs/01_ARCHITECTURE.md index a61e178..02522b8 100644 --- a/code/Engine/docs/01_ARCHITECTURE.md +++ b/code/Engine/docs/01_ARCHITECTURE.md @@ -16,7 +16,7 @@ O sistema é um **servidor MCP em Python** que lê/analisa/reescreve arquivos │ graphify.sh/.md Pipeline de graphify do código │ ├─────────────────────────────────────────────────────────────┤ │ server.py — CAMADA MCP / TRANSPORTE (NÃO tem lógica) │ -│ 62 tools, handlers, prompts, resources, dispatch │ +│ 73 tools, handlers, prompts, resources, dispatch │ │ Só valida entrada/saída e traduz JSON-RPC → chamadas │ ├─────────────────────────────────────────────────────────────┤ │ fcpxml/ — "ENGINE" = NÚCLEO PURO Python (desacoplado) │ @@ -90,7 +90,7 @@ programático — round-trips voltam pelas ferramentas XML. | Controle Live do FCP | `fcpxml/live.py` | | Segurança XML (`defusedxml`, `serialize_xml`) | `fcpxml/safe_xml.py` | | Validação contra DTDs da Apple | `fcpxml/dtd.py` | -| Transporte MCP (62 tools) | `server.py` | +| Transporte MCP (73 tools) | `server.py` | ## 6. Mapa de dependências (você está aqui se for mexer no X → quem tocar) diff --git a/code/Engine/docs/03_SERVER_TOOLS.md b/code/Engine/docs/03_SERVER_TOOLS.md index 5f4d531..472a4d8 100644 --- a/code/Engine/docs/03_SERVER_TOOLS.md +++ b/code/Engine/docs/03_SERVER_TOOLS.md @@ -1,4 +1,4 @@ -# 03 — Camada MCP (`server.py`) — 62 ferramentas +# 03 — Camada MCP (`server.py`) — 73 ferramentas `server.py` (3824 linhas) é a camada de transporte. Não tem lógica de timeline — mapeia nome → handler e delega ao Engine. O dispatch é um dicionário @@ -20,7 +20,7 @@ mapeia nome → handler e delega ao Engine. O dispatch é um dicionário | `_parse_timestamp_parts()` | 433 | Parse de timestamps (min:seg, H:MM:SS, SMPTE) | | `_detect_flash_frames/gaps/duplicate_groups()` | 1667+ | Detectores de QC | -## As 62 ferramentas por categoria +## As 73 ferramentas por categoria ### Timeline & análise (Projeto) `list_projects`, `analyze_timeline`, `list_clips`, `list_markers`, `list_connected_clips`, @@ -57,6 +57,72 @@ mapeia nome → handler e delega ao Engine. O dispatch é um dicionário ### Reformat `reformat_timeline`. +### Voz (análise → decisão → aplicação) +`analyze_voice_features`, `build_voice_timeline`, `refine_voice_timeline`, +`remove_speakers`, `apply_voice_actions`, `get_voice_analysis_config`, +`save_voice_analysis_config`. + +O fluxo é sempre o mesmo: `build_voice_timeline` mede (caro, roda uma vez) → +o modelo decide os cortes → **`refine_voice_timeline` renormaliza sobre o que +sobrou** (barato, sem reabrir áudio) e propõe as janelas de zoom → o modelo +corta a lista pelo ritmo → `apply_voice_actions` aplica. Pular a renormalização +faz o ranking de ênfase apontar para as palavras erradas (ver +`05_EXPERIENCIAS.md`). + +### Legendas dinâmicas (geração → validação → aplicação) +`generate_dynamic_subtitles`, `validate_subtitle_layout`, `transcript_markers`. + +**Sempre gere e depois valide — nunca dê a geração como pronta sem +`validate_subtitle_layout`.** A composição garante "sem sobreposição" só +*por construção* dentro do que ela mesma sabe medir; um título editado à +mão, uma palavra fora do alcance do que foi calibrado, ou conteúdo antigo +no mesmo arquivo escapam dessa garantia. Fluxo: + +``` +generate_dynamic_subtitles(filepath) + ↓ +validate_subtitle_layout(output_path) ← sempre, mesmo quando "parece certo" + ↓ +severidade none/warning → entregar +severidade probable/severe → investigar CADA colisão pela fração exata do +XML antes de mudar código (ver checklist abaixo) +``` + +**Antes de atribuir uma colisão ao gerador, confirme que é o gerador.** +Um `` de nome estranho (`ref` diferente, params tipo `Auto-Shrink`/ +`Left Margin` que `_make_text_title_clip` nunca escreve) é conteúdo humano +ou de outra ferramenta, não um bug — comparar contra o arquivo original +(`grep` pelo texto) resolve em segundos. Caso real: uma colisão "severa" +era um título manual feito no FCP que sobrou no arquivo reaproveitado como +base de teste (`05_EXPERIENCIAS.md`, 2026-08-19). + +**Antes de atribuir uma colisão a uma sobreposição real, confirme pela +fração exata do FCPXML, não pelo float arredondado.** Dois títulos que só +se tocam na borda (um bloco some exatamente quando o próximo começa, por +design) podem imprimir tempos "iguais" e ainda assim colidir no relatório +por ruído de ponto flutuante — `float(a+b) != float(c)` mesmo quando as +frações `a+b` e `c` são idênticas. `temporal_overlap()` já tem uma +tolerância (`_BOUNDARY_EPSILON = 1e-6`, muitas ordens abaixo de um frame) +para absorver isso; se uma colisão nova parecer nascer do nada, comparar +`m._parse_time(...)` dos dois títulos por igualdade exata antes de +suspeitar de sobreposição de verdade. + +**Constantes que resolvem os três bugs já encontrados nesta área** (todas +em `fcpxml/text_layout.py`, exceto a última): + +| Constante | O que resolve | Por quê | +|---|---|---| +| `TEXT_TEMPLATE_FONT_SCALE = 2.0` | Posição e tamanho de fonte dessincronizados | O template "Text" do FCP posiciona no espaço do **frame** (2160×3840), mas o layout mede em pontos de meia-escala (1080×1920). Escalar só o tamanho da fonte e não a posição espalha o texto errado — os dois têm que ser convertidos pelo mesmo fator na saída. | +| `_EMPHASIS_ITALIC_CUSHION_RATIO = 0.06` | Linha de corpo lendo apertada sob a linha de ênfase | O itálico da Playfair inclina as hastes além da caixa de tinta que a métrica mede; ~14pt de respiro extra só nessa fronteira corrige sem tocar no `line_gap` do resto. | +| `fit_emphasis()` (função, não constante) | Palavra de ênfase estourando o frame inteiro | Só linhas de corpo faziam wrap contra `box.width`; a linha de ênfase (sempre uma palavra só) nunca foi checada. Uma palavra longa ou toda maiúscula podia medir mais que o frame inteiro sozinha. Encolhe `font_size`+`kerning` pelo mesmo fator até caber — nunca abaixo do tamanho do corpo, senão ênfase deixa de ser ênfase. | +| `_BOUNDARY_EPSILON = 1e-6` (`fcpxml/collision.py`) | Falso positivo de colisão em títulos que só se tocam | Ver parágrafo acima. | + +Detalhe de implementação e efeito medido de cada um: `05_EXPERIENCIAS.md`, +entradas de 2026-08-19 (#15 zoom, #16 colisão por float, #17 auto-fit da +ênfase — a #15 é do módulo de voz, não de legendas, mas mesma causa-raiz +de fundo: um mecanismo que lê a própria saída anterior precisa continuar +sendo fonte de verdade legível, não só efeito colateral write-only). + ### Live (macOS) `push_to_fcp`, `list_fcp_libraries`. diff --git a/code/Engine/docs/04_TESTS_AND_WORKFLOW.md b/code/Engine/docs/04_TESTS_AND_WORKFLOW.md index f27b32b..6264b6f 100644 --- a/code/Engine/docs/04_TESTS_AND_WORKFLOW.md +++ b/code/Engine/docs/04_TESTS_AND_WORKFLOW.md @@ -63,7 +63,7 @@ uv run --extra dev pytest tests/ -v # Testes com extra de dev ## 4. Estado atual do sistema (resumo "até agora") - **v0.6.35** — núcleo FCPXML completo em Python (`fcpxml/`). -- **62 ferramentas MCP** em `server.py`, organizadas por dispatch `TOOL_HANDLERS`. +- **73 ferramentas MCP** em `server.py`, organizadas por dispatch `TOOL_HANDLERS`. - **Suporte FCPXML 1.8–1.14** (`.fcpxml` e bundles `.fcpxmld` com sidecars), escrita padrão 1.13. - **Dual-mode:** XML (principal) + Live (push_to_fcp / list_fcp_libraries via Apple events). diff --git a/code/Engine/docs/05_EXPERIENCIAS.md b/code/Engine/docs/05_EXPERIENCIAS.md index df0749c..f72fc25 100644 --- a/code/Engine/docs/05_EXPERIENCIAS.md +++ b/code/Engine/docs/05_EXPERIENCIAS.md @@ -33,6 +33,407 @@ Use o bloco abaixo como modelo. Uma entrada = um problema resolvido/reconhecido. ## Registro de Experiências +### 2026-08-19 — Linha de ênfase das legendas dinâmicas sem limite de largura: auto-fit implementado + +- **Contexto:** validando `generate_dynamic_subtitles` sobre o corte real do Mastopexia (ver entrada anterior sobre `validate_subtitle_layout`), sobraram 15 títulos `outside_frame` mesmo depois de eliminados os falsos positivos de colisão. +- **Causa raiz:** `compose_sentence()` (`fcpxml/text_layout.py`) faz wrap das linhas de corpo contra `box.width`, mas a linha de ênfase — sempre uma palavra só, a "key word" em itálico grande — nunca era checada contra largura nenhuma, porque uma palavra sozinha não tem como quebrar em duas linhas. Uma palavra longa (`estruturado,`, `sustentação`, `proporcional`) ou toda maiúscula media mais que o frame inteiro sozinha: `estruturado,` a 460pt (ponto emitido, após `TEXT_TEMPLATE_FONT_SCALE`) mediu **2538px de largura contra 2160px de frame**, estourando os dois lados centrada. +- **Onde:** `fcpxml/text_layout.py::compose_sentence`. +- **Solução adotada:** `fit_emphasis()` mede a linha de ênfase e, se ultrapassar `box.width`, encolhe `font_size` e `kerning` pelo mesmo fator — a largura é linear nesses dois parâmetros juntos, então o fator exato é `box.width / width_medida`, sem iteração. Nunca encolhe abaixo do tamanho do corpo (`body_look.font_size`): ênfase do mesmo tamanho que o texto normal deixa de ser ênfase. Como a extensão vertical de tinta (`ink_extent`) também é linear em `font_size`, encolher a largura encolhe a altura usada no empilhamento junto — restaurando de brinde a garantia "sem sobreposição por construção" que o resto da função já tinha, sem precisar de lógica extra para isso. +- **Efeito medido:** revalidando o mesmo corte, `outside_frame` caiu de 15 para 0, severidade de `severe` para `warning` (restam avisos de área segura, não erros de frame). +- **Achado colateral:** a única colisão que sobrou depois da correção (`MASTOPEXIA` × `é a cirurgia que`) não era bug nenhum — era um título manual (`Auto-Shrink`, template "Text" nativo do FCP, sem relação com `_make_text_title_clip`) presente no arquivo base reutilizado para o teste, provavelmente de uma edição manual no FCP, não gerado por nenhuma chamada da sessão. Revalidando sobre uma base limpa, sem esse título estranho: **0 colisões**. Vale sempre revalidar sobre uma base conhecida antes de atribuir um achado ao código. +- **Estado:** `resolvido` — corrigido em `text_layout.py`, suíte completa e lint passando, revalidado sobre o corte real. + +--- + +### 2026-08-19 — `validate_subtitle_layout` acusava colisão em 7 de 8 casos por ruído de ponto flutuante, não por sobreposição real + +- **Contexto:** primeiro uso real de `validate_subtitle_layout` (ferramenta nova) sobre o corte do Mastopexia com legendas dinâmicas geradas. Relatou severidade `severe`: 8 colisões, 15 títulos fora do frame, 10 fora da área segura. +- **Sintoma:** ao rastrear cada colisão reportada pelas frações exatas do FCPXML, 7 dos 8 pares eram títulos **consecutivos que terminam exatamente quando o próximo começa** — o design pretendido ("cada bloco some quando o próximo aparece", `writer.py`'s `block_ends[i] = block_starts[i+1]`) funcionando corretamente. O oitavo (`MASTOPEXIA` × `é a cirurgia que`) era uma sobreposição real de ~0,46s. +- **Causa raiz:** `temporal_overlap()` em `fcpxml/collision.py` compara `end = start.to_seconds() + duration.to_seconds()` (soma de dois floats já arredondados) contra `start.to_seconds()` de outro título (uma única divisão) — mesmo quando a fração exata subjacente é bit-idêntica nos dois casos, a soma de dois floats arredondados não bate com uma única divisão da soma exata dos numeradores (não-associatividade de ponto flutuante). Medido: diferença de ~4,5×10⁻¹³s — treze ordens de grandeza menor que um frame (~0,04s) — suficiente para inverter `start_b < end_a` de `False` para `True` e disparar uma colisão `severe` fantasma. Contradizia o próprio comentário do código ("a title ending exactly as the next begins is never flagged"). +- **Onde:** `fcpxml/collision.py::temporal_overlap`. +- **Solução adotada:** tolerância `_BOUNDARY_EPSILON = 1e-6` subtraída de ambos os lados da comparação — muitas ordens de grandeza abaixo de qualquer fronteira de frame real, então não mascara nenhuma sobreposição genuína, só absorve o ruído de arredondamento entre dois caminhos de cálculo do mesmo instante. +- **Aprendizado:** comparar dois floats derivados do MESMO valor exato por caminhos aritméticos diferentes (soma vs. divisão direta) nunca deve usar igualdade/desigualdade estrita — vale para qualquer checagem "toca a borda mas não deveria contar", não só tempo de título. O sintoma (severidade `severe` sem nenhuma sobreposição visível no material) é o sinal de alerta: sempre rastrear a colisão até as frações exatas do XML antes de aceitar o relatório da ferramenta de validação como verdade. +- **Estado:** `resolvido` — corrigido em `collision.py`, teste de regressão com os números reais do caso (`test_boundary_survives_float_noise_from_the_writer`), suíte completa (1369 testes) e lint passando, revalidado sobre o corte real: 8 colisões → 1. + +--- + +### 2026-08-19 — `add_zoom` perdia o enquadramento real quando dois zooms caíam no mesmo clipe pós-corte + +- **Sintoma:** no mesmo teste real (Mastopexia), o clipe de abertura do corte apareceu "achatado" no FCP — Scale 100% em vez do enquadramento real (~177%) que o projeto original já tinha, enquanto os clipes seguintes apareciam corretos. Rotação e posição estavam certas; só a escala quebrava, e só no primeiro trecho. +- **Causa raiz:** `add_zoom()` (`fcpxml/writer.py`) lê o enquadramento-base de um clipe só de um jeito: o atributo estático `scale="X Y"` em `<adjust-transform>`. Isso funciona na primeira chamada. Mas quando dois zooms editoriais caem dentro do **mesmo** clipe sobrevivente (dois picos de ênfase que o corte não separou em clipes distintos), a segunda chamada de `add_zoom` encontra não mais um atributo estático, e sim um `<param name="scale">` já **animado** pela primeira — e o código só sabia ler atributo. `stale.get('scale')` voltava `None`, o base virava `1.0` por padrão, e a linha seguinte (`clip.remove(stale)`) **apagava a animação da primeira chamada inteira**, substituindo por uma segunda com base errada. +- **Onde:** `fcpxml/writer.py::add_zoom` (base scale + merge de keyframes); reproduzido isolando `cut_clip_ranges` + duas chamadas de `add_zoom` no mesmo elemento. +- **Por que passou despercebido:** o teste existente (`test_only_one_transform_remains`) já chamava `add_zoom` duas vezes no mesmo clipe, mas só checava que sobrava **um** `<adjust-transform>` na árvore — nunca verificou se a base da segunda chamada estava certa. A suíte cobria a estrutura, não o valor. +- **Solução adotada (duas partes):** + 1. Quando não há atributo `scale` estático, `add_zoom` agora lê o `<param name="scale">` existente e recupera a base como o **menor** valor entre as keyframes — válido porque `MIN_ZOOM_SCALE == 1.0` garante que todo pico é `>= base`, então o menor valor keyframeado é sempre o resting scale, seja ele o de abertura, o de fecho ou qualquer um no meio. + 2. Duas janelas de zoom no mesmo clipe agora só se **substituem** quando as janelas de tempo se sobrepõem (é o mesmo evento sendo reajustado); quando são **disjuntas** (dois picos editoriais distintos que um corte não separou), as keyframes são **empilhadas** no mesmo `keyframeAnimation` em vez de uma apagar a outra — FCPXML aceita quantas keyframes forem necessárias num único `<param>`. +- **Aprendizado:** "preservar o enquadramento existente" precisa valer em **toda** leitura subsequente do mesmo clipe, não só na primeira. Um mecanismo que lê corretamente da fonte original mas degrada ao reler sua própria saída anterior é o mesmo bug de fundo da entrada #13 (renormalização) por outro ângulo: qualquer estado que o sistema regrava precisa continuar sendo uma fonte de verdade legível, não só um efeito colateral write-only. Vale desconfiar de qualquer `findall()`/leitura de atributo que tenha um "senão assume 1.0/padrão" — é aí que a segunda chamada perde o que a primeira escreveu. +- **Estado:** `resolvido` — corrigido em `writer.py`, dois testes de regressão adicionados (`test_second_zoom_on_same_clip_keeps_the_real_base_scale`, `test_overlapping_zoom_on_same_clip_replaces_instead_of_stacking`), suíte completa (1344 testes) e lint passando, corte real do Mastopexia regravado e conferido. + +--- + +### 2026-08-19 — Teste real fechou o ciclo, e revelou um offset sistemático de ~0,4s no timing por palavra + +- **Contexto:** primeiro teste ponta a ponta de `editar-por-voz` num projeto real (Mastopexia, 196,6s de gravação de roteiro com 6 tomadas). Fluxo completo: `build_voice_timeline` → triagem manual (tomada/bastidor/frase abandonada) → `refine_voice_timeline` sobre os sobreviventes → escolha de zoom por função narrativa → `apply_voice_actions`. Corte final: 196,6s → ~50s, 3 clipes. +- **Sintoma:** antes de rodar `remove_media_silence`, medi manualmente o RMS do áudio real nas emendas propostas pelo corte e achei folgas de ~0,4-0,6s onde o JSON dizia que a fala começava/terminava. Comparando timestamp da transcrição contra o ataque real medido em 6 pontos do vídeo (ffmpeg `astats`), o erro era **sistemático, sempre no início da palavra**, entre +0,35s e +0,51s — os finais de palavra batiam certo (+0,01 a +0,15s). +- **Causa raiz:** `transcribe.py` usa `word_timestamps=True` do faster-whisper, que deriva os tempos por atenção cruzada — aproximado por natureza, sem alinhamento forçado. O submódulo WHISPERX existe no repositório mas **não é usado** em nenhum ponto do código; não há etapa de alinhamento fonético. +- **Por que isso importa mais do que parece:** o erro contamina toda decisão temporal a jusante — zoom disparava ~0,4s antes da palavra-alvo, `gap_before` subestimava pausas reais na mesma medida (o que afeta diretamente a régua de silêncio recém-adotada), e as folgas de corte saíam erradas nas emendas. +- **Decisão tomada:** não rodei `remove_media_silence` bruto sobre o corte. A detecção (ffmpeg, limiar -30dB/0,5s) não distingue "batida entre frases dentro da régua de 1,5s" de "ar morto de emenda" — cortar ambos teria apertado frases fluidas. Corrigi os tempos manualmente medindo o ataque real nos pontos críticos (cabeça, 2 emendas, cauda, 3 zooms) e refiz o corte numa passada só. +- **Solução adotada (paliativa, aplicada manualmente neste teste):** medir o RMS real com `ffmpeg -af astats=metadata=1:reset=1:length=0.05,ametadata=print` em janelas curtas ao redor de cada ponto crítico antes de fixar um corte ou zoom que dependa de precisão de frame. Não é o padrão do sistema — é o que cobre a lacuna até o alinhamento forçado existir. +- **Solução estrutural ainda pendente:** ligar o WhisperX (ou alinhamento forçado equivalente) em `transcribe.py`, o que levaria o erro de ~400ms para ~30ms e corrigiria zoom, corte e `gap_before` de uma vez, sem paliativo por projeto. Não implementado ainda — é mudança de pipeline, exige regerar todos os `_transcript.json`/`_voice_timeline.json` existentes. +- **Aprendizado:** "não reestime tempos no olho" (critério 01) continua certo para decisão *editorial* — mas não cobre erro sistemático de *medição* na fonte dos tempos. Um offset constante e na mesma direção, em vários pontos do material, é sinal de bug no pipeline de transcrição, não de julgamento errado sobre o material. Vale conferir com uma amostra de áudio real antes de confiar cegamente em timestamp de word-level de qualquer fonte nova. +- **Estado:** `parcialmente resolvido` — paliativo documentado e aplicado neste teste; correção estrutural (WhisperX) pendente de implementação. + +--- + +### 2026-08-19 — Capacidade existente sem porta de entrada: a Fase 4 da skill era letra morta + +- **Sintoma:** a skill `editar-por-voz` manda, como fase obrigatória, reanalisar o material sobrevivente antes de escolher zooms — e o modelo não tinha como cumprir isso. `restrict_to_kept()` e `suggest_zoom_windows()` existiam, estavam testadas e documentadas, mas **nenhuma ferramenta MCP as expunha**. Na prática, todo zoom continuava sendo escolhido com o ranking bruto, exatamente o erro que a entrada anterior descreve. +- **Causa raiz:** a implementação parou na camada Engine. O critério foi escrito descrevendo chamadas Python, que só os testes conseguiam fazer — a distância entre "existe no `fcpxml/`" e "o modelo consegue chamar" passou despercebida porque a suíte cobria a função, não o caminho. +- **Onde:** `server.py` (nova tool `refine_voice_timeline` + handler + dispatch), `tests/test_refine_voice_timeline_tool.py`, `.claude/skills/editar-por-voz/criterios/04-reanalise-do-material-restante.md`. +- **Solução adotada:** ferramenta `refine_voice_timeline(media_path, cuts, min_gap, max_zooms, save)`, que numa chamada devolve a comparação bruto × sobreviventes, os picos re-ranqueados e os candidatos a zoom. Handler fino: nada de lógica nova, só o caminho até o que já existia. O critério 04 passou a citar a ferramenta em vez das funções Python. +- **Aprendizado:** função coberta por teste unitário **não** é capacidade entregue. Toda vez que um critério de skill mandar "rode X", verifique que X é chamável pelo modelo — senão o critério vira instrução impossível, e o modelo segue em frente sem erro visível. Vale um teste de registro (`a tool está em list_tools` + `está no TOOL_HANDLERS`) para cada ferramenta nova. +- **Estado:** `resolvido` + +--- + +### 2026-08-19 — Ênfase é relativa: analisar o bruto e editar o corte final são perguntas diferentes + +- **Problema:** os zooms estavam sendo escolhidos a partir da análise do material **bruto**. Mas energia é normalizada contra o momento mais alto da gravação — que era `"Amor, eu tô intacto!"` (energia 1,00), justamente uma das falas **cortadas**. Todo o material que sobrou estava pontuado contra uma referência que o espectador nunca veria, comprimindo artificialmente as notas do corte final. +- **Solução:** `restrict_to_kept()` filtra a timeline pelos ranges removidos e **re-normaliza sobre os sobreviventes**. Efeito medido: ênfase média subiu de 0,179 (bruto) para 0,197 (só o que ficou) — o material restante passou a usar a escala inteira. E o ranking mudou de figura: `"Aquela"` 0,39→0,42, `"mastopexia"` 0,26→0,34, `"devolver"` entrando com 0,35. +- **Pré-requisito que virou bug na hora:** re-normalizar exige os valores **brutos**, e o JSON só guardava os normalizados. Adicionados `energy_raw` e `pitch_hz` em cada palavra. A primeira execução saiu com todas as notas caindo — sintoma de estar lendo campo ausente num JSON gerado antes da mudança. Regerar o JSON resolveu; vale lembrar que mudança de esquema exige regerar os caches antes de interpretar qualquer resultado. +- **Armadilha de API evitada:** `enrich_words` chamava `word_pitch_energy`, que **sobrescreve** `energy`/`pitch_hz` com `None` quando não há frame tracks — então re-analisar um subconjunto zerava tudo. Em vez de remendar restaurando os valores depois (que foi a primeira tentativa, e ficou ilegível), entrou o parâmetro `already_measured`. +- **`suggest_zoom_windows()`:** propõe uma janela por frase, a partir da palavra de **conteúdo** mais enfática (artigos e conectivos filtrados por `_FUNCTION_WORDS` — um "a" falado alto continua sendo um artigo), indo até o fim da frase; `min_gap` mantém os zooms afastados. +- **Aprendizado:** "qual o momento mais forte da gravação?" e "qual o momento mais forte do vídeo final?" são perguntas distintas sempre que a métrica for relativa. Toda métrica normalizada precisa ser recalculada quando o conjunto muda — caso contrário ela responde a pergunta errada, silenciosamente. +- **Estado:** `resolvido` (1325 testes verdes; 3 zooms escolhidos pela re-análise aplicados em clipes distintos, DTD 1.14 válido). + +--- + +### 2026-08-19 — Forma do punch-in é assimétrica: entrada de 0,5s, saída de 1 frame + +- **Regra editorial (do usuário):** na palavra de ênfase, zoom in **rápido** (~meio segundo); segura durante a frase de impacto; e no fim **volta de um quadro para o outro, sem transição nenhuma** — o vídeo simplesmente retoma o enquadramento e segue o fluxo. +- **O que havia:** `add_zoom` tinha um único `ease` (padrão 0,3s) aplicado **simetricamente** na entrada e na saída, produzindo um retorno lento que chama atenção para si. +- **Solução:** `ease` passou a valer só para a entrada (padrão **0,5s**) e a saída virou **um frame**, calculado do `frameDuration` real da sequência (`ease_out` opcional para quem quiser retorno gradual). Medido no material: entrada 0,501s, hold 3,378s, saída 0,042s = 1 frame a 23,976fps. +- **Detalhe que quase passou:** o handler em `server.py` forçava `ease=float(action.params.get("ease", 0.3))`, então o padrão novo do writer nunca chegava a valer — o zoom saía com 0,33s de entrada. Um default duplicado em duas camadas é sempre o errado das duas; o handler passou a repassar `ease` **só quando explicitamente informado**, deixando o writer ser o dono do padrão. +- **Aprendizado:** ao mudar um default, procurar quem já o repassa. Um `params.get("x", <default>)` numa camada acima anula silenciosamente o default da camada que de fato conhece o assunto. +- **Estado:** `resolvido` (1311 testes verdes, `TestZoomShapeIsAsymmetric` fixa a forma; DTD 1.14 válido) — pendente de conferência visual no FCP. + +--- + +### 2026-08-19 — Export de calibração do FCP fecha o zoom: keyframe só com `time` e `value` + +- **Como veio:** depois de o zoom continuar não aparecendo, o usuário fez o zoom **à mão no FCP** sobre o mesmo material e exportou o FCPXML (`Mastopexia - exemplo de zoom.fcpxmld`) — o padrão de calibração já registrado em 2026-08-15 e 2026-08-18, agora aplicado a keyframes. +- **O que o export real mostrou:** + ```xml + <adjust-transform position="0.160319 0.663249" rotation="90.1008"> + <param name="scale"> + <keyframeAnimation> + <keyframe time="2329601280/720000s" value="1.77311 1.77311"/> + ``` + 1. `position`/`rotation` mantidos como atributos e o atributo `scale` **removido** quando a escala é animada — confirmou a correção de preservação de enquadramento; + 2. o primeiro keyframe cai exatamente no `start` do clipe (3235,557s) — **confirmou** a correção de timebase de origem; + 3. **`<keyframe>` carrega apenas `time` e `value`** — sem `interp` e **sem `curve`**. +- **Correção final:** removido o `curve="smooth"` que eu havia adicionado ao trocar o `interp`. O DTD permite `curve` (default `smooth`), mas como o importador já havia rejeitado `interp` neste mesmo param vetorial, não há razão para apostar que `curve` sobrevive — o export real do FCP não escreve nenhum dos dois, então passamos a escrever nenhum dos dois. Estrutura agora idêntica à do FCP. +- **Aprendizado:** ao corrigir um atributo rejeitado pelo importador, **não basta trocar por outro plausível do DTD** — foi o que fiz (`interp` → `curve`) e ficou uma segunda aposta não verificada em cima da primeira. A resposta certa era pedir um export de calibração e copiar. Vale a regra: diante de qualquer incerteza sobre o que o FCP aceita, o caminho mais curto é um export real, não uma segunda leitura do DTD. +- **Estado:** `resolvido` (1306 testes verdes, estrutura conferida atributo a atributo contra o export do usuário, DTD 1.14 válido) — pendente de confirmação de importação. + +--- + +### 2026-08-19 — Zoom importava como nada: keyframe em tempo relativo, não no timebase de origem + +- **Sintoma:** usuário importou o FCPXML no FCP e relatou "não tem zoom, não tem corte, não tem nada". +- **Causa 1 (real) — o zoom:** os `<keyframe>` do `adjust-transform` eram escritos em segundos **relativos ao clipe** (0,08s a 4,80s), mas o clipe tem `start="74637363/24000s"` = **3109,9s** (timecode de origem). O FCP procura a animação no timebase do próprio clipe, não encontra keyframe nenhum na janela dele, e importa o zoom como **nada** — sem erro, sem aviso. Corrigido somando o `start` do clipe (`media_origin + tempo relativo`), exatamente o que `add_text_title` já fazia e **documentava**: *"Anchored in SOURCE media coordinates… so the title lands on screen instead of at ~0s of the media (which FCP silently drops)"*. O `add_zoom` nunca recebeu o mesmo tratamento. +- **Por que escapou:** todo fixture sintético e o projeto da Erika têm clipe com `start="0s"`, onde relativo e absoluto coincidem. Pior: o teste `test_add_zoom_creates_keyframed_transform` usava uma fixture com `start="10s"` — tinha tudo para pegar o bug — mas afirmava `times == [1.0, 1.5, 2.5, 3.0]`, ou seja, **fixava o comportamento errado**. Reescrito para ancorar em `origin + relativo`, com `assert origin > 0` garantindo que a fixture continue exercitando o caso. +- **Causa 2 (percepção) — o corte:** os cortes **estavam** no arquivo (3 clipes, 47,6s contra 196,7s do bruto). Mas o XML mantinha `<event name="17-08-2026">` e `<project name="Mastopexia">` idênticos ao original, então a importação criava um projeto homônimo no mesmo evento e o usuário abriu o antigo. Passou-se a renomear o projeto para `<nome> — corte por voz`. +- **Aprendizado 1:** "não fez nada" pode ser duas coisas muito diferentes — não gerou, ou gerou e o usuário não achou. Vale sempre inspecionar o arquivo antes de concluir, e nunca deixar a saída indistinguível da entrada dentro do app de destino. +- **Aprendizado 2 (repetição do padrão de 2026-08-19/`interp`):** quando um módulo já resolve um problema de coordenadas e **documenta** a solução no docstring, procurar os irmãos que fazem operação equivalente. `add_text_title` sabia ancorar em coordenadas de origem; `add_zoom` e qualquer outro futuro escritor de keyframes precisam da mesma regra. +- **Estado:** `resolvido` (1306 testes verdes, keyframes conferidos dentro da janela do clipe: 3297,74–3302,46s para um clipe de 3297,7–3302,6s; DTD 1.14 válido) — pendente de nova importação no FCP. + +--- + +### 2026-08-19 — Primeira edição real ponta a ponta: três bugs de ordem/identidade que a suíte não pegava + +Triagem editorial completa de um vídeo institucional (196,7s → 47,6s, 76% de redução), +feita a partir do `_voice_timeline.json`. A geração do FCPXML expôs três bugs, todos +invisíveis em fixture sintética porque dependem de um projeto **já editado**. + +**1. `add_zoom` destruía o enquadramento do editor.** O clipe original trazia +`<adjust-transform position="0.160319 0.663249" rotation="90.1008" scale="1.77311 1.77311"/>` +— material gravado de lado e reenquadrado à mão. `add_zoom` removia qualquer +`adjust-transform` existente ("Replace rather than stack a prior zoom") e criava o seu do +zero, então o trecho com zoom voltava **girado 90°**. Corrigido: os atributos estáticos +(position/rotation/anchor) são preservados e a escala passa a ser animada **relativa** à +base (1,77311 → 1,77311 × 1,18 → 1,77311). Comentário "replace rather than stack" estava +certo na intenção e errado no alcance — nem todo `adjust-transform` é um zoom anterior. + +**2. Posicionar antes de cortar espalhava o zoom e apagava marcadores.** +`cut_clip_ranges` divide o clipe e reescreve a spine; o `adjust-transform` era **copiado +para os 3 pedaços** e os `<marker>` sumiam. Invertida a ordem: **cortes primeiro**, +posicionamentos depois — `resolve_actions` já converte os tempos para a timeline +pós-corte, então continuam apontando para o mesmo instante. + +**3. Depois de cortar, todos os pedaços têm o MESMO nome.** `add_zoom`/`add_marker` +resolviam o clipe por nome (`_require_clip`), então toda edição caía no **primeiro** +pedaço. `_require_clip` passou a aceitar um `Element` direto (como `add_text_title` já +fazia), e o handler passa o elemento exato — mapeando pelo **offset na timeline**, não +mais pela janela de origem. + +**Bônus — marcador é ponto, não trecho.** `resolve_actions` exigia que início *e* fim +sobrevivessem ao corte, descartando justamente os marcadores úteis: os que sinalizam uma +emenda e por definição encostam na borda do corte. Marcadores passaram a resolver só pelo +início. + +- **Aprendizado:** os três bugs são a mesma família — **ordem de operações e identidade de + elemento** depois de uma operação que reestrutura a árvore. Fixture sintética tem um + clipe limpo, sem transformações prévias e sem nomes duplicados, então nada disso + aparece. Testar contra um projeto **real já editado** é categoricamente diferente de + testar contra XML gerado por nós. +- **Regressões adicionadas:** `TestZoomPreservesExistingFraming` (5), + `TestPlacementsLandOnTheRightPieceAfterCuts` (3), `TestMarkersSurviveCutEdges` (4). +- **Estado:** `resolvido` (1301 testes verdes, DTD 1.14 válido, enquadramento conferido no + XML) — pendente de importação real no FCP pelo usuário. + +--- + +### 2026-08-19 — Diarização quebrada pelo torchcodec + descoberta: separar tomada de conversa NÃO é diarização + +**Parte 1 — a falha técnica.** Com token e termos válidos, `diarize()` retornava `None` e o log dizia só "diarization failed" (o `except Exception:` amplo engolia a causa). Rodando o pyannote direto, a causa apareceu: `pyannote.audio` 4.x decodifica áudio via **torchcodec**, que linka contra uma versão específica do FFmpeg — `dlopen(libtorchcodec_core4.dylib): Library not loaded: @rpath/libavutil.56.dylib`. O FFmpeg instalado é outro major, e a diarização caía inteira numa máquina em que tudo o mais funcionava. + +- **Solução:** `_load_waveform()` em `diarize.py` decodifica o áudio por conta própria (reusando `decodable_audio()` do `voice_features.py`, que já extrai WAV mono 16 kHz via ffmpeg) e passa a `{"waveform": tensor, "sample_rate": sr}` que o pyannote aceita — pulando o torchcodec por completo. Bônus: containers de vídeo passam a funcionar direto. Resultado: 33 turnos, 2 participantes, ~1min40s para 3min17s de áudio. +- **Aprendizado:** `except Exception` sem registrar a exceção transforma falha diagnosticável em mistério. O log deveria carregar a causa; sem isso, foi preciso reexecutar a biblioteca à mão para ver o erro real. + +**Parte 2 — a descoberta de produto, mais importante.** Com a diarização funcionando, o `remove_speakers` encontrou só **uma** fala do SPEAKER_01 — e ainda por cima uma atribuição errada. Cruzando os turnos com a transcrição, a explicação apareceu: o pyannote detecta a segunda voz em 44,1–46,0s e 51,4–54,8s, mas a transcrição **não tem nada** nesses intervalos (vãos de 43,8→47,5 e 48,8→55,4). A pessoa da equipe está **fora do microfone**: a diarização a ouve, o Whisper não a transcreve. + +- **Consequência:** as falas que o usuário quer descartar (*"Amor, eu tô intacto!"*, *"Só clica aí agora na tela."*, *"Posso começar da mastopexia?"*) são **da própria protagonista** — mesma voz, contexto diferente. Cortar por participante não resolve esse caso. +- **Aprendizado:** "quem fala" e "isso é tomada válida?" são perguntas **diferentes**, e é tentador confundi-las porque ambas soam como "separar as partes do vídeo". Diarização resolve a primeira; só a linguagem resolve a segunda. `remove_speakers` continua válido para o caso em que o entrevistador está microfonado (ex. depoimento da Erika), mas não é a ferramenta para triar tomada de conversa. +- **Estado:** `resolvido` (diarização funcional, 1294 testes verdes); triagem tomada/conversa fica na camada de linguagem, critérios em `.claude/skills/editar-por-voz/SKILL.md`. + +--- + +### 2026-08-19 — Pausa longa: ruído para ênfase, sinal para estrutura (o mesmo dado, dois usos opostos) + +- **Sintoma:** no material real, palavras de energia baixíssima lideravam o ranking de ênfase. No Mastopexia, 4 dos 7 picos eram assim: `"mastopexia"` (energia 0,25, pausa 6,2s), `"Aquela"` (0,32, 8,7s), `"Aumenta."` (0,25, 5,7s). O mesmo padrão aparecia no depoimento da Erika. +- **Causa raiz:** `compute_emphasis` normalizava a pausa contra `max_pause=1,5s` **saturando** — ou seja, uma pausa de 8,7s e uma de 1,5s recebiam nota idêntica (1,0). Mas gap de 6–9s não é ênfase dramática: é troca de tomada, inserção de B-roll ou a outra pessoa falando. O índice estava premiando corte de cena como se fosse entrega enfática. +- **Solução adotada:** `pause_weight()` — a contribuição sobe até `max_pause` e **cai a zero** acima de `pause_ignore_above` (3s), em vez de saturar. Resultado imediato no mesmo vídeo: o topo passou a ser `"eu"` (energia 1,00), `"o"` (0,87), `"intacto!"` (0,70), `"tô"` (0,74) — todas da mesma frase, que é de fato a fala de impacto. +- **A virada:** o mesmo dado que era ruído virou o sinal mais útil do documento. Os gaps descartados marcam **onde a tomada recomeçou**. Adicionados `gap_before` e `take_boundary` (>= 3s) em cada segmento; num material de 3min17s isso detectou 6 fronteiras, exatamente onde a médica recomeçava o roteiro. +- **Descoberta de produto:** o bruto de consultório não é uma tomada — é o mesmo roteiro gravado 3–4 vezes, entremeado de conversa com a equipe (*"Amor, eu tô intacto!"*, *"Posso começar da mastopexia?"*, *"Só clica aí agora na tela."*). Separar tomada válida de conversa é **tarefa de linguagem, não de acústica**: no áudio a conversa é mais solta e mais alta que o texto decorado, então qualquer limiar acústico erra o alvo por construção. Critérios registrados em `.claude/skills/editar-por-voz/SKILL.md`. +- **Aprendizado:** antes de descartar um sinal por estar poluindo uma métrica, perguntar **para que outra pergunta ele é a resposta**. Aqui a mesma pausa respondia mal "isso foi enfático?" e otimamente "a tomada recomeçou aqui?". +- **Bug pego pelo próprio teste:** ao adicionar `gap_before`, o teste `test_scales_document_every_word_metric` quebrou — `scales` documentava tudo como métrica de palavra, e `gap_before` é de segmento. `VALUE_SCALES` passou a ser aninhado (`word`/`segment`), com um teste por nível. Um contrato auto-descritivo só vale se um teste garantir que ele não mente. +- **Estado:** `resolvido` (1294 testes verdes, validado nos dois vídeos reais). + +--- + +### 2026-08-19 — FCP descartava TODO zoom gerado: `interp` em param vetorial (DTD-válido ≠ FCP-aceito) + +- **Sintoma:** ao importar o FCPXML no Final Cut, o aviso `This param element was ignored because it does not support the interpolation attribute on its keyframes (.../adjust-transform[1]/param[1])`. O `<param name="scale">` inteiro era **descartado** — ou seja, o zoom simplesmente não existia no projeto importado, sem erro nem falha visível. +- **Causa raiz:** `add_zoom` escrevia `interp="ease"` em cada `<keyframe>` do parâmetro `scale`. O importador do FCP só aceita `interp` em parâmetros **escalares** (opacidade, volume); `scale` é vetorial (`value="1.25 1.25"`) e admite apenas `curve`. +- **Por que passou por tudo:** o DTD oficial da Apple declara `<!ATTLIST keyframe interp (linear|ease|easeIn|easeOut) "linear">` — ou seja, `interp` é **DTD-válido em qualquer keyframe**. A validação contra o DTD passava com 100% de sucesso, e o teste `test_add_zoom_creates_keyframed_transform` **afirmava** `interp == 'ease'`, travando o comportamento errado. Só a importação real no FCP revelou. +- **Solução adotada:** `curve="smooth"` no lugar de `interp` (o `curve` já é `smooth` por padrão no DTD, mas explícito documenta a intenção e protege contra mudança de default). Teste invertido: agora exige `curve == 'smooth'` **e** ausência de `interp`. +- **Atenção — não confundir com `timept`:** o `<timept>` do `timeMap` (usado em `change_speed`) aceita `interp` normalmente e **não** foi alterado. A restrição é do `<keyframe>` em param vetorial. +- **Aprendizado (o mais importante desta série):** **DTD-válido ≠ aceito pelo Final Cut.** O DTD descreve a gramática, não as regras semânticas do importador. Para qualquer construção nova de XML, validar contra o DTD é o piso, não o teto — só a importação real fecha a verificação. E um teste escrito a partir do próprio código gerado (em vez de um export real do FCP) apenas congela o erro: reforça o padrão já registrado em 2026-08-15 e 2026-08-18 — **extrair a verdade de um export real do FCP, nunca do que nós mesmos geramos.** +- **Estado:** `resolvido` no XML (1286 testes verdes, DTD 1.14 válido) — **pendente de nova confirmação de importação no FCP pelo usuário.** + +--- + +### 2026-08-19 — Primeiro teste em material real: quatro bugs que só apareceram fora dos testes + +Rodar o pipeline completo num depoimento real (17 min, 4K, 2320 palavras) expôs quatro +problemas que a suíte inteira, verde, não pegava. Todos vinham de premissas +que só material sintético sustentava. + +**1. Teto de 100 MB rejeitava a mídia (14,7 GB).** `MAX_FILE_SIZE` existe para +documentos que lemos **inteiros na memória** (FCPXML, JSON) — onde um arquivo +gigante é o próprio ataque. Mídia nunca é carregada assim: ffmpeg e librosa +leem em fluxo, com timeout e limite de duração próprios. Criado +`MAX_MEDIA_FILE_SIZE` (32 GB) e `_validate_filepath(..., max_size=)`. Afetava +também o `detect_beats`, que já rejeitava qualquer WAV acima de ~10 minutos. + +**2. librosa não lia `.mp4` — faltava a extração de áudio.** O PDF previa +"FFmpeg para extração do áudio" e eu pulei essa etapa, analisando o container +direto. Criado `decodable_audio()` em `voice_features.py`: passa adiante +arquivos de áudio nativos e extrai um WAV mono 16 kHz temporário via ffmpeg +para containers de vídeo (16 kHz basta — o teto de pitch é 1 kHz). + +**3. O relatório MENTIA sobre o que rodou.** A tabela dizia "Acoustics: yes" +enquanto as duas extrações falhavam, porque reportava `features_capability()` +— se a *biblioteca está instalada* — e não se a *análise funcionou*. Agora o +JSON carrega um bloco `layers` com o que de fato executou. Fundamental porque +"fala monótona" e "acústica não carregou" deixam **os mesmos zeros** nos dados: +sem esse bloco, nem o usuário nem a IA que lê o arquivo conseguem distinguir. + +**4. Limiar absoluto de ênfase não generaliza — trocado por percentil.** O +0,85 do PDF eu já havia recalibrado para 0,60 usando dados sintéticos; no +material real o índice **nunca passou de 0,544** (mediana 0,127), então 0,60 +ainda selecionava nada. Corrigido de vez trocando o mecanismo: `select_peaks` +pega o **top N%** (padrão 2%), com um piso mínimo apenas como guarda para +áudio genuinamente plano. Qualquer corte fixo ou inunda um material ou zera +o outro; percentil entrega um punhado útil nos dois casos. + +- **Aprendizado central:** limiar calibrado em dado sintético é chute. Duas + recalibrações erradas seguidas (0,85 → 0,60, ambas inúteis) só pararam + quando a régua virou **relativa à distribuição do próprio material**. + Sempre que um número governar seleção, prefira percentil a valor absoluto. +- **Observação de qualidade ainda aberta:** no top de ênfase real aparecem + palavras com energia baixíssima (`"No"`, energia 0,07) pontuando alto só + por virem depois de pausa longa. Em entrevista, pausa longa costuma ser o + entrevistador falando — não ênfase. O peso `pause_before` (0,15) com + saturação em 1,5s recompensa o sinal errado; avaliar reduzir o peso ou + ignorar pausas acima de ~3s. +- **Estado:** `resolvido` (1286 testes verdes; saída **validada contra o DTD + oficial FCPXML 1.14 da Apple**) — pendente de importação real no FCP. + +--- + +### 2026-08-19 — Validador acusava desalinhamento de frame em todo projeto NTSC (falso positivo) + +- **Sintoma:** o FCPXML gerado a partir de um projeto real 23,976fps acusava `Duration ... is not frame-aligned at 24fps` em clipes que estavam perfeitamente alinhados. Conferido na mão: `1200199/12000s ÷ 1001/24000s = 2398` frames exatos — inteiro, sem resto. O aviso do *arquivo original*, intocado, também era falso. +- **Causa raiz:** `_check_frame_alignment` fazia `fps_int = int(fps)` e multiplicava os segundos por esse inteiro. A 23,976 (`1001/24000s`), uma duração exatamente alinhada **não** é múltiplo inteiro de "24fps" — então todo projeto NTSC (23,976 / 29,97 / 59,94, ou seja, a maioria) era reportado como quebrado. Detalhe irônico: o docstring de `serialize_xml` já alertava para passar a taxa real "so NTSC projects don't get spurious warnings", mas o `int()` logo adiante destruía a correção. +- **Solução adotada:** o validador passou a ler o `frameDuration` exato do formato que a `<sequence>` referencia (`_document_frame_duration`) e a comparar com aritmética de `Fraction` — alinhado é quando `duração / frameDuration` tem denominador 1. A mensagem também passou a nomear a taxa real (`23.976fps`), não uma arredondada. +- **Aprendizado:** nunca converter timebase para inteiro/float para checar alinhamento — o projeto inteiro é construído sobre tempo racional justamente por isso (`TimeValue`), e a validação precisa seguir a mesma regra que a escrita. Falso positivo em validador é pior que ausência de validação: ensina o usuário a ignorar avisos, e aí o aviso verdadeiro passa batido. +- **Estado:** `resolvido` (3 testes de regressão em `TestNTSCFrameAlignment`, incluindo um que garante que desalinhamento **real** continua sendo detectado). + +--- + +### 2026-08-18 — Divisão IA × sistema: a IA devolve DECISÕES, nunca XML; e tudo em tempo de origem + +- **Contexto:** definido como a inteligência entra no editor automático. A ideia inicial era mandar o JSON para um serviço externo que devolveria o material já editado; evoluiu para fazer a decisão aqui dentro, com um skill versionado no repo (`.claude/skills/editar-por-voz/SKILL.md`). +- **Decisão 1 — o que a IA NÃO faz:** silêncio, vícios de linguagem, extração acústica, índice de ênfase, diarização e **geração de FCPXML** continuam determinísticos. Tudo que tem resposta objetiva (um limiar decide) não ganha nada indo para um modelo — só custo, latência e perda de reprodutibilidade. Geração de XML em particular é matemática de tempo racional frame a frame: modelo gerando XML produz arquivo sutilmente quebrado. +- **Decisão 2 — a IA devolve uma lista de ações, não mídia editada.** Contrato em `fcpxml/voice_actions.py` (`kind`/`start`/`end`/`params`/`reason`). Motivos: dá para **validar** antes de aplicar; é **reprodutível** (mesma lista → mesmo FCPXML); e o usuário **revisa** antes de qualquer coisa tocar a timeline. `parse_actions` trata a lista como entrada não confiável — uma linha malformada é reportada e pulada, nunca derruba a edição inteira. +- **Decisão 3 (a que evita a pior classe de bug) — todos os tempos em segundos da mídia ORIGINAL.** Cortes deslocam tudo que vem depois: se as decisões viessem em tempo pós-corte, cada destaque cairia silenciosamente no frame errado assim que um corte fosse adicionado. `resolve_actions`/`shift_after_cuts` resolvem o deslocamento na hora de aplicar, e ação que aponta para material removido é **descartada e reportada**, nunca deslizada para o conteúdo vizinho. +- **Bug pego pelo próprio relatório:** título e marcador não apareciam no XML. A causa era `str(TimeValue)` devolvendo o `__repr__` (`TimeValue(3/1s = 3.000s)`) em vez da string racional — o método certo é `to_fcpxml()`. Só foi visível na hora porque o handler reporta o que **não** conseguiu colocar, com a exceção real, em vez de aplicar em silêncio. +- **Aprendizado:** todo handler que aplica uma lista de operações deve relatar as três categorias — aplicadas, descartadas e rejeitadas. Um handler que só conta sucessos transforma bug em "não aconteceu nada" e some do radar. Nunca converter `TimeValue` para string com `str()`: sempre `to_fcpxml()`. +- **Estado:** `resolvido` (1276 testes verdes; aplicação validada contra FCPXML real, incluindo o `examples/sample.fcpxml`) — **pendente de confirmação de importação real no FCP pelo usuário.** + +--- + +### 2026-08-18 — Limiar de ênfase de 0,85 do PDF era inalcançável: média ponderada não chega lá + +- **Sintoma:** com a timeline de voz montada, o campo `peak_count` vinha **sempre 0**. Nem a palavra mais alta e mais aguda de um trecho de demonstração ("segurança", energia normalizada 1,0 e pico de tom) era marcada como candidata a punch-in. +- **Investigação:** medido o teto real da fórmula. `compute_emphasis` é uma **média ponderada** de cinco fatores normalizados (energia 0,30 / tom 0,25 / ritmo 0,20 / pausa 0,15 / duração 0,10). Para o resultado passar de 0,85 seria preciso que quase todos os cinco estivessem no máximo **simultaneamente** — o que a fala real não produz: uma palavra com pausa dramática antes dela quase por definição não tem desvio de ritmo alto. Valores medidos: todos os fatores no máximo = 1,00; pico realista (energia e tom máximos, pausa longa, palavra longa, ritmo normal) = **0,80**; pico comum = **0,66**. +- **Causa raiz:** o 0,85 veio literalmente da especificação do PDF (`SE emphasis > 0.85 ENTÃO aplicar punch-in`), que pressupunha outra normalização — provavelmente um índice de máximo, não de média. Copiar a constante sem conferir a distribuição da nossa fórmula tornou o recurso inerte. +- **Solução adotada:** padrão recalibrado para **0,60** (`DEFAULT_VOICE_ANALYSIS_CONFIG` em `fcpxml/model_manager.py`), com o porquê comentado no próprio código. O texto da tela e a descrição da tool passaram a dizer que picos reais ficam na faixa 0,55–0,80 — para que ninguém volte a subir o valor achando que "quanto maior, mais seletivo" sem saber onde fica o teto. +- **Aprendizado:** constante numérica herdada de especificação externa precisa ser **validada contra a distribuição real da fórmula implementada** antes de virar padrão. O sintoma aqui foi silencioso (nenhum erro, nenhum teste vermelho — só um recurso que nunca disparava), e só apareceu porque a saída de demonstração foi inspecionada com dados realistas. Vale gerar uma amostra de verdade e olhar os números sempre que um limiar governar um comportamento. +- **Estado:** `resolvido` (padrão 0,60 verificado: o mesmo trecho passou a marcar corretamente 1 pico). + +--- + +### 2026-08-18 — Configurações de análise de voz: uma fonte de verdade só (config.json do backend), não UserDefaults + +- **Contexto:** Fases 2–3 da arquitetura de análise de voz (features acústicas + índice de ênfase) e a tela de configurações pedida pelo usuário para regular limiares de energia, ênfase e emoção. +- **Decisão:** os parâmetros de análise ficam **só** em `~/.fcp-mcp-server/config.json` (via `model_manager.load/save_voice_analysis_config`), diferente do padrão `@AppStorage`/UserDefaults usado por `CaptionsView.swift` para estilo de legenda. Motivo: estilo de legenda é preferência de UI que só é lida na hora de montar os argumentos de uma chamada; já os limiares de análise são lidos **pelo próprio motor** (`handle_analyze_voice_features`) mesmo quando a análise é disparada fora do app (tool MCP direta, script). Duplicar em UserDefaults criaria duas verdades divergentes — a tela mostraria um valor e a análise usaria outro. +- **Onde:** `fcpxml/model_manager.py` (`DEFAULT_VOICE_ANALYSIS_CONFIG`, `load/save_voice_analysis_config`), comandos `voice_analysis`/`set_voice_analysis` em `admin/models_api.py`, tools MCP `get/save_voice_analysis_config`, tela `MacApp/Sources/VoiceAnalysisView.swift`. +- **Cuidado que rendeu teste:** `load_voice_analysis_config` precisa devolver uma **cópia** dos defaults — a primeira versão devolvia o dict aninhado `emphasis_weights` por referência, e quem mutasse o resultado corrompia o default do módulo para o resto do processo. Coberto por `test_defaults_are_not_shared_mutable_state`. +- **Aprendizado:** ao adicionar configuração nova, perguntar "quem lê esse valor?" — se for o motor Python, ele mora no config.json do backend; se for só a montagem de argumentos na UI, UserDefaults serve. E todo default composto (dict/lista) devolvido de um `load_*` precisa ser cópia, nunca a constante do módulo. +- **Estado:** `resolvido` (1201 testes verdes, lint zero erros, ciclo salvar→reler validado pelo bridge) — **a renderização visual da tela no app não pôde ser confirmada por captura de tela** (janela do app em outro Space); a aba foi confirmada via árvore de acessibilidade. + +--- + +### 2026-08-18 — Diarização de locutor já existia pronta e testada, mas órfã (nenhuma tool MCP a expunha) + +- **Contexto:** início da implementação da "Arquitetura de Análise de Voz para Editor Automático" (locutor, energia, pitch, ênfase, motor de regras → FCPXML), especificada num PDF trazido pelo usuário. Plano salvo em `~/.claude/plans/volumes-merongo-downloads-arquitetura-a-mossy-micali.md`. +- **Descoberta:** `fcpxml/diarize.py` (diarização via `pyannote/speaker-diarization-3.1`, com `diarization_capability`, `diarize`, `assign_speakers`, `build_speakers`) e `tests/test_diarize.py` já existiam completos e passando, mas nenhuma tool em `server.py` chamava esse módulo — código morto do ponto de vista de uso real. `model_manager.py` também já tinha `load_hf_token`/`save_hf_token` prontos para o token do HuggingFace exigido pelo pyannote. +- **Decisão de arquitetura:** usar pyannote (já é dependência declarada em `pyproject.toml` como extra `diarization`) para diarização bruta por turno, em vez de treinar/rodar SpeechBrain ECAPA-TDNN do zero como o PDF sugeria em primeiro lugar. ECAPA-TDNN fica reservado para uma fase futura (reconhecimento de pessoa cadastrada por cima dos turnos já diarizados), evitando duas libs pesadas resolvendo o mesmo problema. +- **Onde:** nova tool `diarize_media` em `server.py` (handler `handle_diarize_media`), reaproveitando `diarize.py` sem alterá-lo; cache em `_diarization.json` ao lado da mídia, seguindo exatamente o padrão de `_transcript.json`/`_beats.json` já usados por `transcribe_media`/`detect_beats`. +- **Aprendizado:** antes de implementar uma fase "do zero" a partir de uma spec externa, vale sempre grepar o `fcpxml/` por nomes prováveis (`diarize`, `speaker`, etc.) — pode já existir motor pronto e testado, só faltando a camada de exposição via MCP tool. +- **Estado:** `resolvido` (tool nova + testes, 1154 testes verdes, lint zero erros). + +--- + +### 2026-08-18 — Terceira linha "colando" na linha de ênfase: o gap simétrico não bastava para o itálico + +- **Sintoma:** no bloco de composição "phrase" (uma palavra de ênfase em itálico grande, cercada por linhas de corpo), a linha logo abaixo da ênfase aparecia quase tocando o texto — mesmo com o slider "Espaçamento entre linhas" da tela de legendas dinâmicas configurado. +- **Investigação:** reproduzido o cálculo de `compose_sentence` fora do FCP com a frase exata do usuário ("de" / "encontrar" / "roupa,") — o gap entre as caixas de tinta dava **exatamente 8pt nos dois lados** (acima e abaixo da ênfase), confirmando que o valor do slider chega corretamente até o layout (`CaptionsView.swift` → `admin/models_api.py` → `server.py` → `compose_sentence`). Não era bug de configuração não aplicada. +- **Causa raiz:** a caixa de tinta medida (`ink_extent`, `fcpxml/text_layout.py`) é vertical e simétrica, mas a inclinação itálica do Playfair Display faz os traços "vazarem" visualmente para baixo além do que a métrica vertical mede — então o mesmo gap numérico lê como mais apertado abaixo da linha de ênfase do que acima dela. +- **Onde:** `fcpxml/text_layout.py::compose_sentence` (função `pair_gap` nova) e constante `_EMPHASIS_ITALIC_CUSHION_RATIO`. +- **Solução adotada:** gap por par de linhas em vez de um valor único para todo o bloco — quando a linha anterior é a de ênfase, soma-se uma folga extra proporcional ao seu `font_size` (`ratio = 0.06`, ~14pt a 230pt) só naquele par; todos os outros pares continuam usando exatamente o `line_gap` do usuário. A folga entra tanto no teste de "cabe na banda" quanto na centralização da pilha, senão o bloco vazaria do box.height ou ficaria descentrado. +- **Aprendizado:** medir a caixa de tinta (ascendente/descendente reais) resolve colisão entre glifos retos, mas não captura o "peso visual" da inclinação itálica — para faces itálicas grandes ao lado de corpo reto, a folga simétrica por ink-box ainda pode ler como assimétrica no render final. Um cushion proporcional ao tamanho da fonte, aplicado só no lado que precisa, corrige sem inflar o espaçamento nos pares que já estavam certos. +- **Estado:** `resolvido` no cálculo (1150 testes verdes, valores conferidos numericamente) — **pendente de confirmação visual real no FCP pelo usuário**. + +--- + +### 2026-08-18 — Build Out desligado libera toda a janela de "Per Object" para o Build In terminar de revelar + +- **Sintoma:** em blocos de palavras curtos, a animação de entrada do título ("Text"/Basic Text template) às vezes cortava antes de terminar de revelar a palavra — o corte pro próximo bloco acontecia no meio do reveal. +- **Causa raiz:** `Apply Speed = "2 (Per Object)"` (já presente em `_TEXT_TITLE_PARAMS`) faz o FCP comprimir/esticar a animação **inteira** do template (build in + build out) para caber exatamente na duração real do `<title>`. Com as duas fases ativas, build in e build out disputam a mesma janela comprimida — em clipes curtos, build in não tinha tempo suficiente. +- **Onde:** `fcpxml/writer.py::_TEXT_TITLE_PARAMS` (`FCPXMLModifier._make_text_title_clip`). +- **Descoberta do `key`:** não havia como adivinhar — o usuário desmarcou manualmente "Build Out" no Inspector de um título "Text" isolado no FCP e exportou o FCPXML. O override só aparece no XML quando o valor difere do default do template (por isso um export sem a alteração real não mostra o `param` nenhum). Valor capturado: `<param name="Build Out" key="9999/10000/2/102" value="0"/>`. +- **Solução adotada:** `Build Out` adicionado como primeiro item de `_TEXT_TITLE_PARAMS`, sempre `"0"` (desligado) em todo título gerado. Não foi necessário nenhum parâmetro extra de velocidade — desligar o build out já entrega toda a janela "Per Object" comprimida ao build in, que é o efeito de "sempre acelerado" pedido pelo usuário. +- **Aprendizado:** pra descobrir o `key` de um checkbox/param publicado num template Motion, o export de calibração **precisa** ter o valor realmente alterado no Inspector antes de exportar — reexportar o projeto sem mexer em nada não revela nada (o FCP só escreve params que divergem do default). Reforça o padrão já registrado em 2026-08-15: nunca adivinhar `key`, sempre extrair de um export real. +- **Estado:** `resolvido` no XML (1 novo param verificado no writer) — **pendente de confirmação de importação real no FCP pelo usuário**. + +--- + +### 2026-08-18 — Espaço de coordenadas do modelo de título: tamanho E posição + +- **Sintoma (1ª metade):** o bloco caía exatamente onde o preview mostrava, mas + o texto renderizava cerca de **metade** do tamanho configurado — com 213pt a + ênfase deveria ocupar ~87% da largura do quadro e ocupava ~35%. +- **Sintoma (2ª metade, causado pela primeira correção):** ao dobrar só o + `fontSize`, o tamanho ficou certo e as **linhas passaram a se sobrepor** — o + bloco mantinha o espalhamento antigo com o dobro de letra dentro. +- **Causa raiz:** a calibração de tamanhos veio do export manual feito com o + modelo **"Essencial - Título"**, cujo espaço de coordenadas é o canvas de + pontos (metade do quadro). Esse modelo nunca renderizou quando gerado por nós + (entrada de 2026-08-17), então o writer passou a emitir o **"Basic Text > + Text" (Text.moti)** — cujo espaço é o **quadro inteiro** (2160×3840). Tudo o + que esse modelo lê está nesse espaço: `fontSize`, `kerning` **e** `Position`. +- **Onde:** `fcpxml/text_layout.py` (`TEXT_TEMPLATE_FONT_SCALE`, + `position_param`), `fcpxml/writer.py`, `fcpxml/models.py` + (`DynamicSubtitleConfig.text_scale`), `server.py`, + `MacApp/Sources/CaptionsView.swift`. +- **Tentativas que falharam:** (a) procurar a diferença nos params do título + (`Auto-Shrink`, margens, `Layout Method`) — todos idênticos ao export manual; + (b) **converter só o `fontSize`** — corrigiu o tamanho e quebrou o + espaçamento, que é o erro registrado aqui como aprendizado principal. +- **Solução adotada:** um único fator, `TEXT_TEMPLATE_FONT_SCALE = 2.0`, + aplicado ao `fontSize`, ao `kerning` **e** à `Position` na saída. O layout + continua medindo em pontos de canvas — toda constante calibrada depende + disso — e a conversão acontece só na emissão, que é a única forma de os dois + andarem juntos. Exposto como `text_scale`. +- **Aprendizado:** um espaço de coordenadas é indivisível. Converter metade das + grandezas que vivem nele é **pior** do que não converter nenhuma: sem + conversão o erro é uniforme e parece "só um ajuste de tamanho"; pela metade, + tipo e espaçamento se descolam e o defeito muda de cara. Ao trocar o modelo + de título, toda constante calibrada contra o modelo antigo vira suspeita — a + posição foi re-verificada em 2026-08-17 e o tamanho não, e o bug ficou + invisível porque "está no lugar certo" parece "está certo". +- **Estado:** `resolvido` + +--- + +### 2026-08-18 — Preview das legendas dinâmicas desproporcional ao render do FCP + +- **Sintoma:** o painel "Legendas Dinâmicas" mostrava um preview que não batia + com o resultado no Final Cut: linhas de apoio coladas nas bordas do quadro, + espaçamento entre linhas errado, a banda do bloco invisível e as cores + aplicadas de forma trocada. O formulário de controles também estava confuso, + com blocos de texto explicativo a cada slider. +- **Causa raiz:** `SubtitlePreviewView` era um desenho aproximado feito à mão + (VStack + Spacer, gap fixo de 14pt, canvas mapeado só na altura) e não + reproduzia `compose_sentence` de `fcpxml/text_layout.py`. Além disso, o app + mandava `inactive_color`, que em `granularity="phrase"` o backend **nunca + usa** — o preview pintava a ênfase com uma cor que o FCP ignoraria. +- **Onde:** `MacApp/Sources/SubtitlePreviewView.swift`, + `MacApp/Sources/CaptionsView.swift`, `server.py` + (`handle_generate_dynamic_subtitles`), `admin/models_api.py` (docstring). +- **Tentativas que falharam:** apenas re-escalar as fontes do preview — a + posição continuava errada, porque o desalinhamento vinha do *stagger* e do + gap, não do tamanho. +- **Solução adotada:** o preview passou a espelhar a geometria do backend — + canvas 1080×1920 pt (largura inclusa), margem lateral de 4%, + `REFERENCE_BLOCK_LINE_GAP` (8 pt), `REFERENCE_STAGGER_RATIO` (0.8) com lados + alternados a partir da esquerda, corpo em Bold, empilhamento sobre a tinta + (cap-height + descida) e compensação do centro do frame do `Text`. O + backend ganhou `emphasis_color` (padrão = `active_color`), e a UI foi + reagrupada em "Linhas de apoio" / "Palavra de ênfase" com sliders em + `LabeledContent` e explicação em tooltip. +- **Aprendizado:** um preview só é útil se for derivado das MESMAS constantes + do gerador. Quando o preview é redesenhado "de olho", ele vira uma segunda + fonte de verdade que diverge silenciosamente. E todo controle exposto na UI + precisa existir de fato no caminho de código que ele diz configurar. +- **Estado:** `resolvido` + +--- + ### 2026-08-17 — Garantir que dois blocos nunca se sobreponham: empilhar pela TINTA real, não pela cap-height - **Sintoma:** na composição progressiva, a cedilha de "começar" (Playfair @@ -676,6 +1077,42 @@ Use o bloco abaixo como modelo. Uma entrada = um problema resolvido/reconhecido. <!-- NOVAS ENTRADAS DEVEM SER ADICIONADAS ACIMA DESTA LINHA, SEMPRE NO TOPO DA LISTA, PARA QUE A MAIS RECENTE FIQUE EM PRIMEIRO LUGAR. --> +### 2026-08-19 — Legendas dinâmicas geradas com `bold="0" fontFace="Bold"` não renderizam no FCP + +- **Sintoma:** no corte real da Mastopexia, as legendas dinâmicas (composição + progressiva) não apareciam no Final Cut — só as primeiras linhas de cada + bloco surgiam e o restante sumia. O arquivo que o usuário re-exportou do FCP + ("legendas dinamicas.fcpxmld") renderizava normalmente. +- **Causa raiz:** o corpo das legendas era definido como + `EDITORIAL_BODY_LOOK = WordLook(88, ..., face="Bold")` e o gravador emitia + `bold="0"` + `fontFace="Bold"` (pois `WordStyle.bold` é `False` por padrão). + Essa combinação é contraditória: no FCPXML negrito é o **atributo** `bold="1"` + (nunca um `fontFace="Bold"`), e itálico é `fontFace="... Italic"` **mais** + `italic="1"`. O FCP re-exporta `bold="1"` (sem `fontFace`) e + `fontFace="Medium Italic"` + `italic="1"`, provando o formato correto. +- **Onde:** `fcpxml/writer.py::_make_text_title_clip` (emissão do `text-style`); + o estilo em si em `fcpxml/models.py::EDITORIAL_BODY_LOOK`. +- **Tentativas que falharam:** corrigir manualmente o XML gerado trocando + `bold="0" fontFace="Bold"` por `bold="1"` — resolvia só aquele arquivo e o + bug reaparecia a cada geração. Também tentei "corrigir" os offsets dos + títulos (achando que estavam fora da realidade por estarem em coordenadas de + source) e quebrei o arquivo com timebases errados (24000, 30000) — os offsets + em source coords estavam corretos o tempo todo (ver entrada de 2026-08-17 + sobre "anchored in SOURCE media coordinates"). +- **Solução adotada:** em `_make_text_title_clip`, traduzir a face "bold" para + `bold="1"` sem `fontFace`; emitir `italic="1"` quando a face contém "italic"; + e não mais emitir `bold="0"` junto de uma face. Agora a saída bate com a + re-exportação do FCP (corpo `bold="1"`, palavra-chave `fontFace` + `italic="1"`). +- **Aprendizado:** o FCPXML do template "Text" usa `bold` (atributo) para peso e + `fontFace`+`italic` para a face itálica; "Bold" não é um valor válido de + `fontFace`. Ao duvidar de um formato, confiar na re-exportação do FCP (saída + canônica) e nunca "corrigir" offsets/times que já seguem a convenção do + gerador. Também: comparar a saída gerada contra o FCP byte a byte por campo + (bold/fontFace/italic) antes de assumir o problema em outro lugar. +- **Estado:** `resolvido` + +--- + ### 2026-08-14 — Início do registro de experiências - **Sintoma:** não havia um local centralizado para registrar erros/estruturas @@ -692,6 +1129,60 @@ Use o bloco abaixo como modelo. Uma entrada = um problema resolvido/reconhecido. --- +## 19 — `output_dir` aplicado só como cerca, nunca como destino + +- **Data:** 2026-08-19 +- **Sintoma:** toda chamada com `output_dir` diferente da pasta do arquivo de + entrada morria com `Output path escapes allowed directory`, apontando para um + caminho que a própria função tinha acabado de montar. Na prática o ajuste + "Pasta do projeto" do app só funcionava quando apontava para a pasta onde o + arquivo já ia cair sozinho — ou seja, nunca fazia nada. +- **Causa raiz:** em `_resolve_io_paths` (`server_tools/_shared.py`) o + `output_dir` virava apenas `anchor_dir` da validação, enquanto o nome do + arquivo continuava saindo de `generate_output_path(filepath, suffix)`, que + preserva o diretório da ENTRADA. Cerca em um lugar, destino em outro: o + caminho gerado ficava fora da própria cerca. Afetava os 18+ handlers de + escrita, não só as legendas onde o erro apareceu. +- **Solução adotada:** quando `output_dir` é passado, o destino padrão passa a + ser `<output_dir>/<nome derivado>`; sem ele, mantém-se o comportamento antigo + (ao lado da entrada). Um `output_path` explícito continua vencendo e continua + obrigado a ficar dentro da âncora. `build_voice_timeline` e + `refine_voice_timeline` passaram a aceitar e repassar `output_dir`; os + leitores procuram na pasta do projeto primeiro e caem para o lado da mídia, + para não perder timelines geradas antes da mudança. +- **Aprendizado:** validação e destino não podem ser derivados de fontes + diferentes. Quando um parâmetro tem dois papéis (permissão e endereço), + aplicar só um dos dois produz um erro que acusa o próprio código — e some da + vista porque o caso que funciona é justamente o caso trivial. +- **Estado:** `resolvido` + +--- + +## 20 — Cadeia de processamento sem o passo que corta + +- **Data:** 2026-08-19 +- **Sintoma:** o encadeamento do app ia de `analyze_voice` direto para + `remove_silences`/legendas. Dava para medir a voz e legendar o resultado, mas + não para aplicar as decisões de edição — o corte por voz tinha que ser rodado + à mão, fora do app, e era fácil parar no primeiro passo achando que o arquivo + estava pronto. +- **Causa raiz:** `apply_voice_actions` existia como handler MCP mas nunca foi + exposto na ponte `admin/models_api.py`, então o batch não tinha como chamá-lo. +- **Solução adotada:** comando `apply_voice_actions` na ponte (aceita + `actions_path` apontando para o JSON de decisões, com ou sem o embrulho + `{"actions": [...]}`), e a etapa correspondente no batch do app, posicionada + logo após a análise e **antes** de qualquer passo que faça ripple — os + zooms/textos/marcadores são posicionados deslocando a partir da própria lista + de cortes, então rodar depois de outro corte os joga no frame errado sem erro + visível. O botão fica bloqueado se a etapa estiver ligada sem arquivo + escolhido, para a cadeia não quebrar no meio. +- **Aprendizado:** um passo que só existe como ferramenta MCP não existe para + quem usa o app. Vale conferir se toda etapa documentada no fluxo tem + representação na cadeia que o usuário de fato executa. +- **Estado:** `resolvido` + +--- + ## Resumo rápido (índice) | # | Data | Problema | Estado | @@ -704,5 +1195,15 @@ Use o bloco abaixo como modelo. Uma entrada = um problema resolvido/reconhecido. | 8 | 2026-08-17 | Importação recusada: `id` de `<text-style-def>` derivado do texto (acentos/espaços/dígito inicial) não é XML Name válido | `resolvido` | | 9 | 2026-08-17 | Legendas palavra a palavra centradas em vez da composição progressiva diagramada (bloco por trecho, palavra-chave em display italic) | `resolvido` | | 10 | 2026-08-17 | Cedilha/acentos da display italic invadindo a linha vizinha: empilhamento passou a usar a tinta real por classe de glifo | `resolvido` | +| 11 | 2026-08-18 | Preview das legendas dinâmicas desproporcional ao render do FCP (stagger/gap/canvas divergentes) e `inactive_color` exposto sem efeito | `resolvido` | +| 12 | 2026-08-18 | Espaço de coordenadas do modelo "Text": `fontSize`, `kerning` e `Position` no espaço do quadro — converter só o tamanho descolou o espaçamento | `resolvido` | +| 13 | 2026-08-19 | Reanálise de ênfase implementada no Engine mas sem ferramenta MCP — Fase 4 da skill era inexecutável | `resolvido` | +| 14 | 2026-08-19 | Offset sistemático de ~0,4s no timing por palavra (faster-whisper sem alinhamento forçado) — corrigido manualmente no teste, WhisperX pendente | `parcialmente resolvido` | +| 15 | 2026-08-19 | `add_zoom` perdia o enquadramento real (voltava a 100%) quando dois zooms caiam no mesmo clipe pós-corte; agora empilha ou substitui conforme as janelas se sobrepõem | `resolvido` | +| 16 | 2026-08-19 | `validate_subtitle_layout` acusava colisão severa em títulos que só se tocam na borda, por não-associatividade de float; 7 de 8 colisões reportadas no teste real eram falso positivo | `resolvido` | +| 17 | 2026-08-19 | Linha de ênfase das legendas dinâmicas sem limite de largura — palavra longa/maiúscula estourava o frame inteiro; auto-fit encolhe até caber, nunca abaixo do corpo | `resolvido` | +| 18 | 2026-08-19 | Legendas dinâmicas geradas com `bold="0" fontFace="Bold"` não renderizam no FCP — negrito deve ser `bold="1"` (atributo) e itálico `fontFace`+`italic="1"` | `resolvido` | +| 19 | 2026-08-19 | `output_dir` usado só como cerca de validação e nunca como destino — toda chamada entre pastas falhava acusando o caminho que ela mesma gerou | `resolvido` | +| 20 | 2026-08-19 | `apply_voice_actions` ausente da ponte e do encadeamento do app — dava para analisar e legendar, não para cortar | `resolvido` | > Mantenha o índice acima sempre sincronizado com as entradas mais recentes. diff --git a/code/MacApp/Sources/App.swift b/code/MacApp/Sources/App.swift index 001ebc3..23cef88 100644 --- a/code/MacApp/Sources/App.swift +++ b/code/MacApp/Sources/App.swift @@ -14,6 +14,7 @@ struct GArtApp: App { enum ActiveTab: Hashable { case project case captions + case voiceAnalysis case models case about } @@ -26,8 +27,10 @@ struct ContentView: View { List(selection: $activeTab) { Label("Projeto", systemImage: "film") .tag(ActiveTab.project) - Label("Legendas", systemImage: "captions.bubble") + Label("Legendas Dinâmicas", systemImage: "captions.bubble") .tag(ActiveTab.captions) + Label("Análise de Voz", systemImage: "waveform") + .tag(ActiveTab.voiceAnalysis) Label("Modelos", systemImage: "tray.and.arrow.down") .tag(ActiveTab.models) Label("Sobre", systemImage: "info.circle") @@ -42,7 +45,10 @@ struct ContentView: View { .navigationTitle("Projeto") case .captions: CaptionsView().id(UUID()) - .navigationTitle("Legendas") + .navigationTitle("Legendas Dinâmicas") + case .voiceAnalysis: + VoiceAnalysisView().id(UUID()) + .navigationTitle("Análise de Voz") case .models: ModelDownloadView().id(UUID()) .navigationTitle("Modelos") diff --git a/code/MacApp/Sources/CaptionsView.swift b/code/MacApp/Sources/CaptionsView.swift index 691ad28..ea959ea 100644 --- a/code/MacApp/Sources/CaptionsView.swift +++ b/code/MacApp/Sources/CaptionsView.swift @@ -1,195 +1,416 @@ import SwiftUI import UniformTypeIdentifiers -/// Guia "Legendas" — configura e gera legendas dinâmicas como clipes de título -/// editáveis no Final Cut Pro (template "Essencial - Título"). Cada palavra -/// vira um clipe de título posicionado: as palavras da frase vão surgindo -/// conforme são faladas, se acumulam num bloco centralizado, e somem todas -/// juntas no fim da frase. Sem Compound Clip. +/// Guia "Legendas Dinâmicas" — apenas configuração de estilo. Nenhum +/// processamento acontece aqui: a geração das legendas roda na aba +/// "Projeto". +/// +/// Os valores ficam em `~/.fcp-mcp-server/config.json` (via +/// `model_manager.save_dynamic_subtitle_config`), os mesmos lidos por +/// `generate_dynamic_subtitles` como padrão — não em UserDefaults/ +/// `@AppStorage`, para que a geração renderize exatamente o que esta tela +/// mostra, e para que qualquer chamador (app, MCP, uma sessão de IA) veja o +/// mesmo estilo sem precisar repassar os 11 campos a cada chamada. Mesmo +/// desenho de `VoiceAnalysisView`. +/// +/// Layout em duas colunas: à esquerda, um **preview vertical 9:16** fixo que +/// reproduz a composição real (posição, escala, quebra de linhas e o +/// escalonamento das linhas de apoio); à direita, os controles agrupados por +/// assunto. struct CaptionsView: View { - @State private var projectPath: String? - @State private var clipName = "" - - // Posicionamento do bloco - @State private var bandHeight: Double = 0.22 - @State private var blockCenterY: Double = -167 - - // Estilo - @State private var font = "Helvetica Neue" - @State private var fontSize: Double = 90 - @State private var activeColor = Color.white - @State private var inactiveColor = Color(white: 0.7) - - @State private var isGenerating = false - @State private var resultPath: String? + @State private var config = CaptionStyleConfig.defaults + @State private var isLoading = true @State private var errorMessage: String? + // Frase de amostra do preview — só conveniência local, não afeta a + // geração real (que usa as palavras de verdade da transcrição), então + // continua em @AppStorage em vez do config compartilhado. + @AppStorage("capSampleBefore") private var sampleBefore = "que vão" + @AppStorage("capSampleEmphasis") private var sampleEmphasis = "melhorar" + @AppStorage("capSampleAfter") private var sampleAfter = "sua legenda" + @AppStorage("capShowGuides") private var showsGuides = true + private let fontChoices = [ "Helvetica Neue", "Helvetica", "Arial", "Avenir Next", "Futura", "SF Pro Display", "Georgia", "Impact", ] + private let emphasisFontChoices = [ + "Playfair Display", "Georgia", "Didot", "Futura", + "Avenir Next", "Times New Roman", "Helvetica Neue", "Impact", + ] + + private let emphasisFaceChoices = [ + "Medium Italic", "Italic", "Bold Italic", "Bold", "Regular", "Light Italic", + ] + + /// Salva no arquivo a cada mudança e devolve um Binding, para os controles + /// continuarem simples (`$config.x` viraria só memória local). + private func bound<T>(_ keyPath: WritableKeyPath<CaptionStyleConfig, T>) -> Binding<T> { + Binding( + get: { config[keyPath: keyPath] }, + set: { config[keyPath: keyPath] = $0; save() } + ) + } + + private func colorBound(_ keyPath: WritableKeyPath<CaptionStyleConfig, String>) -> Binding<Color> { + Binding( + get: { Color(rgbaString: config[keyPath: keyPath]) }, + set: { config[keyPath: keyPath] = $0.fcpxmlColorString; save() } + ) + } + var body: some View { + HSplitView { + previewColumn + .frame(minWidth: 250, idealWidth: 300, maxWidth: 380) + controlsColumn + .frame(minWidth: 380, idealWidth: 460) + } + .task { await load() } + } + + // MARK: - Coluna da pré-visualização + + private var previewColumn: some View { + VStack(alignment: .leading, spacing: 14) { + HStack { + Label("Pré-visualização", systemImage: "rectangle.on.rectangle.angled") + .font(.headline) + Spacer() + Toggle("Guias", isOn: $showsGuides) + .toggleStyle(.switch) + .controlSize(.mini) + .labelsHidden() + .help("Mostra a faixa do bloco e a linha de centro do quadro.") + } + + SubtitlePreviewView( + bodyFont: config.font, + bodySize: config.fontSize, + bodyColor: Color(rgbaString: config.activeColor), + emphasisFont: config.emphasisFont, + emphasisFace: config.emphasisFace, + emphasisSize: config.emphasisSize, + emphasisColor: Color(rgbaString: config.emphasisColor), + bandHeight: config.bandHeight, + blockCenterY: config.blockCenterY, + lineGap: config.lineGap, + beforeText: $sampleBefore, + emphasisText: $sampleEmphasis, + afterText: $sampleAfter, + showsGuides: showsGuides + ) + .frame(maxWidth: .infinity, maxHeight: 420) + + GroupBox("Frase de amostra") { + VStack(spacing: 6) { + LabeledContent("Antes") { + TextField("", text: $sampleBefore).textFieldStyle(.roundedBorder) + } + LabeledContent("Ênfase") { + TextField("", text: $sampleEmphasis).textFieldStyle(.roundedBorder) + } + LabeledContent("Depois") { + TextField("", text: $sampleAfter).textFieldStyle(.roundedBorder) + } + } + .padding(.vertical, 4) + } + + if emphasisOverflows { + Label( + "A palavra de ênfase é mais larga que o quadro nesse tamanho — reduza o tamanho da ênfase para não sair cortada.", + systemImage: "exclamationmark.triangle.fill" + ) + .font(.caption) + .foregroundStyle(.orange) + .fixedSize(horizontal: false, vertical: true) + } + + Text("Quadro vertical 9:16 na mesma geometria do Final Cut: a ênfase fica centrada e as linhas de apoio se deslocam para os lados alternados. Larguras são estimadas — a posição e a escala são reais.") + .font(.caption2) + .foregroundStyle(.secondary) + .fixedSize(horizontal: false, vertical: true) + + Spacer(minLength: 0) + } + .padding(20) + .frame(maxHeight: .infinity, alignment: .top) + } + + // MARK: - Coluna de controles + + private var controlsColumn: some View { Form { - Section("Projeto do Final Cut Pro") { - HStack { - Text(projectName) - .foregroundStyle(projectPath == nil ? .secondary : .primary) - .lineLimit(1) - Spacer() - Button("Escolher…") { pickProjectFile() } + if isLoading { + Section { + ProgressView().controlSize(.small) + .frame(maxWidth: .infinity, alignment: .center) } - if let projectPath { - HStack { - Image(systemName: "doc.text").foregroundStyle(.secondary) - Text(projectPath).font(.caption).foregroundStyle(.secondary).lineLimit(1) - } - } - TextField("Nome do clipe (opcional — vazio = todos os clipes)", text: $clipName) - .textFieldStyle(.roundedBorder) + } else { + positionSection + bodySection + emphasisSection + calibrationSection } - - Section { - VStack(alignment: .leading, spacing: 6) { - HStack { - Text("Altura do bloco").font(.callout.weight(.medium)) - Spacer() - Text("\(Int(bandHeight * 100))%").font(.caption).foregroundStyle(.secondary).monospacedDigit() - } - Slider(value: $bandHeight, in: 0.10...0.50, step: 0.01) - Text("Quanto da altura do quadro a frase pode ocupar antes de quebrar em outro bloco. Maior = mais palavras juntas na tela.") - .font(.caption2).foregroundStyle(.secondary) - } - .padding(.vertical, 4) - - VStack(alignment: .leading, spacing: 6) { - HStack { - Text("Altura na tela").font(.callout.weight(.medium)) - Spacer() - Text("\(Int(blockCenterY))").font(.caption).foregroundStyle(.secondary).monospacedDigit() - } - Slider(value: $blockCenterY, in: -700...300, step: 1) - Text("Posição vertical do bloco. 0 é o centro do quadro; valores negativos descem.") - .font(.caption2).foregroundStyle(.secondary) - } - .padding(.vertical, 4) - } header: { - Text("Posicionamento do Bloco") - } footer: { - Text("Cada palavra vira um clipe de título solto na timeline (sem Compound Clip), posicionado para não sobrepor as outras palavras da frase.") - .font(.caption2).foregroundStyle(.secondary) - } - - Section("Estilo do Texto") { - Picker("Fonte", selection: $font) { - ForEach(fontChoices, id: \.self) { Text($0).tag($0) } - } - - VStack(alignment: .leading, spacing: 6) { - HStack { - Text("Tamanho da fonte").font(.callout.weight(.medium)) - Spacer() - Text("\(Int(fontSize))pt").font(.caption).foregroundStyle(.secondary).monospacedDigit() - } - Slider(value: $fontSize, in: 20...200, step: 1) - } - .padding(.vertical, 4) - - ColorPicker("Cor A", selection: $activeColor, supportsOpacity: true) - ColorPicker("Cor B", selection: $inactiveColor, supportsOpacity: true) - - Text("Tamanho, cor e estilo variam por palavra seguindo um ritmo fixo, calibrado a partir de um projeto real do Final Cut. As cores acima entram nesse ritmo.") - .font(.caption2).foregroundStyle(.secondary) - } - - Section { - Button { - generate() - } label: { - if isGenerating { ProgressView().controlSize(.small) } - Label("Gerar Legendas Dinâmicas", systemImage: "captions.bubble.fill") - } - .disabled(isGenerating || projectPath == nil) - .frame(maxWidth: .infinity) - .buttonStyle(.borderedProminent) - - if let errorMessage { + if let errorMessage { + Section { Label(errorMessage, systemImage: "exclamationmark.triangle.fill") .foregroundStyle(.red) } - if let resultPath { - Text(resultPath).font(.caption).lineLimit(1) - HStack { - Button("Abrir no Final Cut Pro") { NSWorkspace.shared.open(URL(fileURLWithPath: resultPath)) } - Button("Mostrar no Finder") { - NSWorkspace.shared.activateFileViewerSelecting([URL(fileURLWithPath: resultPath)]) - } - } - } - } footer: { - Text("Usa a transcrição local (Whisper) já feita na aba Projeto/Transcrição. Se ainda não houver transcrição salva, ela é gerada automaticamente.") - .font(.caption2).foregroundStyle(.secondary) } } .formStyle(.grouped) } - private var projectName: String { - guard let p = projectPath else { return "Nenhum projeto selecionado" } - return URL(fileURLWithPath: p).lastPathComponent + private var positionSection: some View { + Section { + slider( + "Altura do bloco", + value: bound(\.bandHeight), in: 0.10...0.50, step: 0.01, + readout: "\(Int(config.bandHeight * 100))%", + help: "Quanto da altura do quadro a frase pode ocupar antes de quebrar em outro bloco." + ) + slider( + "Posição vertical", + value: bound(\.blockCenterY), in: -700...300, step: 1, + readout: "\(Int(config.blockCenterY))", + help: "0 é o centro do quadro; negativos descem, positivos sobem." + ) + slider( + "Espaço entre linhas", + value: bound(\.lineGap), in: -80...120, step: 1, + readout: "\(Int(config.lineGap))pt", + help: "Distância entre uma linha e a outra, além das próprias letras. Negativo sobrepõe." + ) + } header: { + Text("Posicionamento no Quadro") + } footer: { + Text("Medidas em pontos do canvas de referência 2160×3840. O espaço entre linhas é contado a partir da tinta real de cada linha — em 0 elas se encostam, e em negativo uma entra na outra.") + .font(.caption2).foregroundStyle(.secondary) + } } - private func pickProjectFile() { - let panel = NSOpenPanel() - panel.canChooseFiles = true - panel.canChooseDirectories = false - panel.allowsMultipleSelection = false - panel.prompt = "Selecionar" - panel.message = "Selecione o arquivo de projeto (.fcpxml) exportado pelo Final Cut Pro." - if panel.runModal() == .OK, let url = panel.url { - let ext = url.pathExtension.lowercased() - if ext == "fcpxml" || ext == "xml" || ext == "fcpxmld" { - projectPath = url.path - } else { - errorMessage = "Selecione um arquivo .fcpxml ou .xml do Final Cut Pro." + private var bodySection: some View { + Section("Linhas de apoio") { + Picker("Fonte", selection: bound(\.font)) { + ForEach(fontChoices, id: \.self) { Text($0).tag($0) } + } + slider( + "Tamanho", + value: bound(\.fontSize), in: 20...200, step: 1, + readout: "\(Int(config.fontSize))pt", + help: "Tamanho das linhas que acompanham a palavra-chave." + ) + ColorPicker("Cor", selection: colorBound(\.activeColor), supportsOpacity: true) + } + } + + private var emphasisSection: some View { + Section("Palavra de ênfase") { + Picker("Fonte", selection: bound(\.emphasisFont)) { + ForEach(emphasisFontChoices, id: \.self) { Text($0).tag($0) } + } + Picker("Estilo", selection: bound(\.emphasisFace)) { + ForEach(emphasisFaceChoices, id: \.self) { Text($0).tag($0) } + } + slider( + "Tamanho", + value: bound(\.emphasisSize), in: 60...400, step: 1, + readout: "\(Int(config.emphasisSize))pt", + help: "A palavra-chave da frase, sozinha em sua linha e maior." + ) + ColorPicker("Cor", selection: colorBound(\.emphasisColor), supportsOpacity: true) + } + } + + private var calibrationSection: some View { + Section { + slider( + "Escala no Final Cut", + value: bound(\.textScale), in: 0.5...3.0, step: 0.05, + readout: String(format: "%.2f×", config.textScale), + help: "Converte o tamanho escolhido para o espaço em que o modelo de título do Final Cut desenha o texto." + ) + + HStack { + Spacer() + Button("Restaurar padrão") { + config = CaptionStyleConfig.defaults + save() + } + .buttonStyle(.link) + } + } header: { + Text("Calibração") + } footer: { + VStack(alignment: .leading, spacing: 6) { + Text("O modelo \"Text\" do Final Cut trabalha no espaço do quadro inteiro (2160×3840), enquanto o layout é calculado em pontos — metade disso. Por isso o padrão é 2,00×, aplicado ao tamanho E à posição juntos. Ajuste só se o seu modelo usar outra proporção.") + Text("Este estilo fica salvo em ~/.fcp-mcp-server/config.json e é usado automaticamente ao gerar legendas dinâmicas — nada é processado aqui.") + } + .font(.caption2).foregroundStyle(.secondary) + } + } + + /// Slider com rótulo, leitura numérica alinhada e explicação em tooltip — + /// mantém as linhas do formulário com a mesma altura. + private func slider( + _ title: String, + value: Binding<Double>, + in range: ClosedRange<Double>, + step: Double, + readout: String, + help: String + ) -> some View { + LabeledContent(title) { + HStack(spacing: 10) { + Slider(value: value, in: range, step: step) + Text(readout) + .font(.caption).monospacedDigit() + .foregroundStyle(.secondary) + .frame(width: 46, alignment: .trailing) } } + .help(help) } - private func generate() { - guard let projectPath else { return } - isGenerating = true - errorMessage = nil - resultPath = nil + /// Largura útil do canvas de referência (1080 pt menos 4% de cada lado). + private static let usableCanvasWidth: CGFloat = 1080 * 0.92 - var args: [String: Any] = [ - "path": projectPath, - "band_height": bandHeight, - "block_center_y": blockCenterY, - "font": font, - "font_size": Int(fontSize), - "active_color": activeColor.fcpxmlColorString, - "inactive_color": inactiveColor.fcpxmlColorString, - ] - let trimmedClip = clipName.trimmingCharacters(in: .whitespaces) - if !trimmedClip.isEmpty { - args["clip_name"] = trimmedClip + /// A palavra de ênfase da amostra passa da largura do quadro no tamanho + /// escolhido? Medida no mesmo canvas que o backend usa. + private var emphasisOverflows: Bool { + let text = sampleEmphasis.trimmingCharacters(in: .whitespaces) + guard !text.isEmpty else { return false } + var descriptor = NSFontDescriptor(fontAttributes: [.family: config.emphasisFont]) + if !config.emphasisFace.isEmpty { + descriptor = descriptor.addingAttributes([.face: config.emphasisFace]) } + let nsFont = NSFont(descriptor: descriptor, size: CGFloat(config.emphasisSize)) + ?? NSFont.systemFont(ofSize: CGFloat(config.emphasisSize)) + let width = (text as NSString).size(withAttributes: [.font: nsFont]).width + return width > Self.usableCanvasWidth + } - PythonBridge.call(command: "generate_dynamic_subtitles", arguments: args) { result, err in - DispatchQueue.main.async { - isGenerating = false - if result?["ok"] as? Bool == true { - resultPath = result?["path"] as? String - } else { - errorMessage = result?["error"] as? String ?? err ?? "Falha ao gerar legendas dinâmicas." + // MARK: - Backend + + @MainActor + private func load() async { + await withCheckedContinuation { continuation in + PythonBridge.call(command: "dynamic_subtitle_config") { result, error in + DispatchQueue.main.async { + if let result { + config = CaptionStyleConfig(from: result) + } else if let error { + errorMessage = error + } + isLoading = false + continuation.resume() } } } } + + private func save() { + PythonBridge.call(command: "set_dynamic_subtitle_config", arguments: config.arguments()) { _, error in + DispatchQueue.main.async { errorMessage = error } + } + } +} + +/// O estilo das legendas dinâmicas, no formato que a tela edita e o bridge +/// (`admin/models_api.py` → `set_dynamic_subtitle_config`) persiste. +struct CaptionStyleConfig { + var bandHeight: Double + var blockCenterY: Double + var lineGap: Double + var font: String + var fontSize: Double + var emphasisFont: String + var emphasisFace: String + var emphasisSize: Double + var activeColor: String + var emphasisColor: String + var textScale: Double + + static let defaults = CaptionStyleConfig( + bandHeight: 0.22, + blockCenterY: -167, + lineGap: 8, + font: "Helvetica Neue", + fontSize: 104, + emphasisFont: "Playfair Display", + emphasisFace: "Medium Italic", + emphasisSize: 265, + activeColor: "1 1 1 1", + emphasisColor: "1 1 1 1", + textScale: 2.0 + ) + + /// Lê a resposta do bridge, caindo no padrão para qualquer campo ausente. + init(from json: [String: Any]) { + let d = CaptionStyleConfig.defaults + self.init( + bandHeight: json["band_height"] as? Double ?? d.bandHeight, + blockCenterY: json["block_center_y"] as? Double ?? d.blockCenterY, + lineGap: json["line_gap"] as? Double ?? d.lineGap, + font: json["font"] as? String ?? d.font, + fontSize: (json["font_size"] as? NSNumber)?.doubleValue ?? d.fontSize, + emphasisFont: json["emphasis_font"] as? String ?? d.emphasisFont, + emphasisFace: json["emphasis_face"] as? String ?? d.emphasisFace, + emphasisSize: (json["emphasis_size"] as? NSNumber)?.doubleValue ?? d.emphasisSize, + activeColor: json["active_color"] as? String ?? d.activeColor, + emphasisColor: json["emphasis_color"] as? String ?? d.emphasisColor, + textScale: json["text_scale"] as? Double ?? d.textScale + ) + } + + init( + bandHeight: Double, blockCenterY: Double, lineGap: Double, + font: String, fontSize: Double, + emphasisFont: String, emphasisFace: String, emphasisSize: Double, + activeColor: String, emphasisColor: String, textScale: Double + ) { + self.bandHeight = bandHeight + self.blockCenterY = blockCenterY + self.lineGap = lineGap + self.font = font + self.fontSize = fontSize + self.emphasisFont = emphasisFont + self.emphasisFace = emphasisFace + self.emphasisSize = emphasisSize + self.activeColor = activeColor + self.emphasisColor = emphasisColor + self.textScale = textScale + } + + func arguments() -> [String: Any] { + [ + "band_height": bandHeight, + "block_center_y": blockCenterY, + "line_gap": lineGap, + "font": font, + "font_size": Int(fontSize), + "emphasis_font": emphasisFont, + "emphasis_face": emphasisFace, + "emphasis_size": Int(emphasisSize), + "active_color": activeColor, + "emphasis_color": emphasisColor, + "text_scale": textScale, + ] + } } extension Color { + /// Parses an FCPXML "R G B A" space-separated 0-1 string into a Color. + init(rgbaString: String) { + let parts = rgbaString.split(separator: " ").compactMap { Double($0) } + guard parts.count >= 3 else { self = .white; return } + let r = parts[0], g = parts[1], b = parts[2], a = parts.count >= 4 ? parts[3] : 1.0 + self.init(.sRGB, red: r, green: g, blue: b, opacity: a) + } + /// Converts to FCPXML's "R G B A" space-separated 0-1 string (sRGB). var fcpxmlColorString: String { let ns = NSColor(self).usingColorSpace(.sRGB) ?? NSColor(self) diff --git a/code/MacApp/Sources/ModelDownloadView.swift b/code/MacApp/Sources/ModelDownloadView.swift index 596d873..6fb4cf2 100644 --- a/code/MacApp/Sources/ModelDownloadView.swift +++ b/code/MacApp/Sources/ModelDownloadView.swift @@ -105,25 +105,64 @@ struct ModelDownloadView: View { .font(.caption) .foregroundStyle(catalog?.diarization == true ? Color.secondary : Color.orange) } - SecureField("Token HuggingFace (diarização)", text: $hfTokenText) - .textFieldStyle(.roundedBorder) - .help("Token com acesso aos modelos gated pyannote (segmentation + speaker-diarization)") + tokenField HStack { TextField("Nº de participantes (vazio = automático)", text: $numSpeakersText) .textFieldStyle(.roundedBorder) Button("Salvar") { saveDiarization() } .disabled(isLoading) } + setupSteps } } header: { Text("Diarização (participantes)") } footer: { - Text("Opcional. Sem token, cada fala é atribuída a um único participante padrão.") + Text("Opcional. Sem token, cada fala é atribuída a um único participante padrão. O token só é usado para baixar o modelo uma vez — depois disso a análise roda offline, nesta máquina.") .font(.caption) .foregroundStyle(.secondary) } } + /// Campo do token. Quando já existe um salvo, mostra o estado em vez de um + /// campo vazio ambíguo — e oferece a remoção, já que salvar vazio não apaga. + @ViewBuilder + private var tokenField: some View { + if catalog?.hfTokenSet == true { + HStack { + Label("Token salvo nesta máquina", systemImage: "key.fill") + .font(.caption) + .foregroundStyle(.secondary) + Spacer() + Button("Remover") { removeToken() } + .buttonStyle(.link) + .disabled(isLoading) + } + } + SecureField( + catalog?.hfTokenSet == true ? "Substituir token…" : "Token HuggingFace (diarização)", + text: $hfTokenText + ) + .textFieldStyle(.roundedBorder) + .help("Token com acesso aos modelos gated pyannote (segmentation + speaker-diarization)") + } + + /// O token sozinho não basta: os dois modelos pyannote são "gated" e exigem + /// aceitar os termos na conta antes do download funcionar. + private var setupSteps: some View { + VStack(alignment: .leading, spacing: 4) { + Text("Para ativar, uma vez só:") + .font(.caption).bold() + .foregroundStyle(.secondary) + Link("1. Aceitar os termos do modelo de segmentação", + destination: URL(string: "https://huggingface.co/pyannote/segmentation-3.0")!) + Link("2. Aceitar os termos do modelo de diarização", + destination: URL(string: "https://huggingface.co/pyannote/speaker-diarization-3.1")!) + Link("3. Criar um token de acesso e colar acima", + destination: URL(string: "https://huggingface.co/settings/tokens")!) + } + .font(.caption) + } + private func saveDiarization() { var args: [String: Any] = ["num_speakers": numSpeakersText] if !hfTokenText.isEmpty { @@ -138,6 +177,15 @@ struct ModelDownloadView: View { } } + private func removeToken() { + PythonBridge.call(command: "set_diarization", arguments: ["token": ""]) { _, _ in + DispatchQueue.main.async { + hfTokenText = "" + Task { await refresh() } + } + } + } + // MARK: - Storage private var storageSection: some View { diff --git a/code/MacApp/Sources/ProjectView.swift b/code/MacApp/Sources/ProjectView.swift index 96c466f..84f7b51 100644 --- a/code/MacApp/Sources/ProjectView.swift +++ b/code/MacApp/Sources/ProjectView.swift @@ -101,6 +101,25 @@ struct ProjectView: View { } } .formStyle(.grouped) + .task { restoreLastProject() } + } + + /// Reabre o último projeto salvo em ~/.fcp-mcp-server/config.json, para o + /// app voltar onde parou em vez de pedir o arquivo de novo a cada abertura. + /// É aqui que a restauração precisa morar: a TranscriptionView embutida só + /// existe depois que há um projeto carregado, então ela não consegue se + /// restaurar sozinha. Caminhos que sumiram do disco voltam vazios do + /// Python, e nesse caso a tela abre limpa como antes. + private func restoreLastProject() { + guard project == nil else { return } + PythonBridge.call(command: "project_config") { result, _ in + DispatchQueue.main.async { + guard project == nil, + let result, result["ok"] as? Bool == true, + let file = result["file"] as? String, !file.isEmpty else { return } + inspect(file) + } + } } /// Presents an NSOpenPanel configured to select a single file (not a @@ -156,6 +175,23 @@ struct ProjectView: View { private func handleDrop(_ providers: [NSItemProvider]) -> Bool { guard let provider = providers.first else { return false } + + // `loadObject(ofClass: URL.self)` is Foundation's own bridge for a + // dropped file URL and handles every representation Finder/FCP may + // hand back (NSURL via secure coding, a bookmark, a plain path). + // The item can ALSO be manually pulled as raw `Data` — but that only + // works if the bytes are exactly `URL.dataRepresentation`'s format + // (a UTF-8 absolute-string encoding), which a `.fcpxmld` *bundle* + // (a package macOS treats as a directory, not a plain file) does not + // always arrive as: the decode silently returns nil, so the drop + // looks like it does nothing. Try the robust path first. + if provider.canLoadObject(ofClass: URL.self) { + _ = provider.loadObject(ofClass: URL.self) { url, _ in + self.handleDroppedURL(url) + } + return true + } + provider.loadItem(forTypeIdentifier: UTType.fileURL.identifier, options: nil) { item, _ in var url: URL? if let data = item as? Data { @@ -163,17 +199,21 @@ struct ProjectView: View { } else if let u = item as? URL { url = u } - if let url, isAccepted(url) { - DispatchQueue.main.async { inspect(url.path) } - } else { - DispatchQueue.main.async { - errorMessage = "Este arquivo não parece ser um projeto do Final Cut Pro." - } - } + self.handleDroppedURL(url) } return true } + private func handleDroppedURL(_ url: URL?) { + DispatchQueue.main.async { + if let url, isAccepted(url) { + inspect(url.path) + } else { + errorMessage = "Este arquivo não parece ser um projeto do Final Cut Pro." + } + } + } + private func isAccepted(_ url: URL) -> Bool { let ext = url.pathExtension.lowercased() return ext == "fcpxml" || ext == "fcpxmld" || ext == "xml" @@ -242,6 +282,8 @@ struct ProjectView: View { if let result, result["ok"] as? Bool == true { project = ProjectInfo(json: result) showTranscription = true + PythonBridge.call(command: "set_project_config", + arguments: ["file": path]) { _, _ in } } else { errorMessage = result?["error"] as? String ?? err ?? "Falha ao ler o projeto." } diff --git a/code/MacApp/Sources/SubtitlePreviewView.swift b/code/MacApp/Sources/SubtitlePreviewView.swift new file mode 100644 index 0000000..0757abe --- /dev/null +++ b/code/MacApp/Sources/SubtitlePreviewView.swift @@ -0,0 +1,218 @@ +import SwiftUI + +/// Simulação ao vivo da composição `phrase` das legendas dinâmicas como um +/// **quadro vertical 9:16**, na mesma geometria que `fcpxml/text_layout.py` +/// usa para posicionar os títulos no Final Cut. +/// +/// Mapeamento fiel ao backend (`compose_sentence`): +/// * o canvas de referência é **1080×1920 pt** (2160×3840 a `POINT_SCALE = 0.5`); +/// * `bodySize`/`emphasisSize` e `blockCenterY` são pontos nesse canvas — o +/// preview multiplica tudo por ``altura do preview / 1920``; +/// * a linha de ênfase fica centrada e as linhas de apoio "penduram" nas +/// bordas dela, alternando lados a partir da esquerda, deslocadas por +/// `REFERENCE_STAGGER_RATIO` (0.8) da folga em relação à linha mais larga; +/// * as linhas empilham com `lineGap` pontos de tinta a tinta entre elas — 0 +/// as encosta e negativo sobrepõe — e o bloco inteiro é centrado em +/// `blockCenterY`; +/// * o corpo é sempre **Bold**, a ênfase usa a face escolhida. +/// +/// As larguras são estimadas pelo próprio layout de texto do macOS, então o +/// resultado é uma **aproximação** do render do FCP — mas a posição relativa, +/// as proporções e o escalonamento são reais. +struct SubtitlePreviewView: View { + var bodyFont: String + var bodySize: Double + var bodyColor: Color + var emphasisFont: String + var emphasisFace: String + var emphasisSize: Double + var emphasisColor: Color + var bandHeight: Double + var blockCenterY: Double + var lineGap: Double + + @Binding var beforeText: String + @Binding var emphasisText: String + @Binding var afterText: String + + /// Mostra a faixa (banda) e a linha de centro por cima do quadro. + var showsGuides: Bool = true + + /// Canvas de referência do FCP: 2160×3840 px a POINT_SCALE 0.5. + private let canvasHeight: CGFloat = 1920 + private let canvasWidth: CGFloat = 1080 + /// REFERENCE_STAGGER_RATIO. + private let staggerRatio: CGFloat = 0.8 + /// `side_margin` de LayoutBox.for_frame. + private let sideMargin: CGFloat = 0.04 + + var body: some View { + GeometryReader { geo in + let scale = geo.size.height / canvasHeight + let lines = composedLines(scale: scale, usableWidth: geo.size.width * (1 - 2 * sideMargin)) + let bandPixels = geo.size.height * CGFloat(bandHeight) + // y cresce para cima no FCP; na tela cresce para baixo. + let centerY = geo.size.height * 0.5 - CGFloat(blockCenterY) * scale + + ZStack { + LinearGradient( + colors: [Color(white: 0.14), Color(white: 0.03)], + startPoint: .top, endPoint: .bottom + ) + + if showsGuides { + Rectangle() + .fill(Color.accentColor.opacity(0.10)) + .frame(height: bandPixels) + .overlay(alignment: .top) { guideRule } + .overlay(alignment: .bottom) { guideRule } + .position(x: geo.size.width / 2, y: centerY) + + Rectangle() + .fill(Color.white.opacity(0.16)) + .frame(height: 1) + .position(x: geo.size.width / 2, y: geo.size.height * 0.5) + } + + ForEach(Array(lines.enumerated()), id: \.offset) { _, line in + Text(line.text) + .font(line.font) + .foregroundColor(line.color) + .lineLimit(1) + .fixedSize() + .position( + x: geo.size.width / 2 + line.x, + y: centerY + line.y + ) + } + } + .clipShape(RoundedRectangle(cornerRadius: 10)) + .overlay( + RoundedRectangle(cornerRadius: 10) + .strokeBorder(Color.white.opacity(0.15), lineWidth: 1) + ) + } + .aspectRatio(canvasWidth / canvasHeight, contentMode: .fit) + } + + private var guideRule: some View { + Rectangle().fill(Color.accentColor.opacity(0.45)).frame(height: 1) + } + + // MARK: - Composição + + private struct Line { + var text: String + var font: Font + var nsFont: NSFont + var color: Color + var isEmphasis: Bool + var width: CGFloat + var height: CGFloat + var x: CGFloat = 0 + var y: CGFloat = 0 + } + + /// Reproduz `compose_sentence`: quebra as linhas de apoio na largura útil, + /// empilha os blocos com `lineGap` e desloca as linhas de apoio para os + /// lados alternados da linha mais larga. + private func composedLines(scale: CGFloat, usableWidth: CGFloat) -> [Line] { + let bodyPt = max(3, CGFloat(bodySize) * scale) + let emphasisPt = max(3, CGFloat(emphasisSize) * scale) + + var lines: [Line] = [] + lines += bodyLines(before(), size: bodyPt, maxWidth: usableWidth) + let emphasis = emphasisText.trimmingCharacters(in: .whitespaces) + lines.append(makeLine( + emphasis.isEmpty ? " " : emphasis, + nsFont: resolvedFont(emphasisFont, size: emphasisPt, face: emphasisFace), + color: emphasisColor, + isEmphasis: true + )) + lines += bodyLines(after(), size: bodyPt, maxWidth: usableWidth) + + guard !lines.isEmpty else { return [] } + + let gap = CGFloat(lineGap) * scale + let stackHeight = lines.reduce(0) { $0 + $1.height } + gap * CGFloat(lines.count - 1) + let anchor = lines.map(\.width).max() ?? 0 + + // Topo do bloco relativo ao seu próprio centro. + var edge = -stackHeight / 2 + var side: CGFloat = -1 + for index in lines.indices { + if index > 0 { edge += gap } + // `.position` centra o frame do Text, cujo centro fica acima do + // centro da tinta — corrige para a tinta cair onde o FCP a põe. + let f = lines[index].nsFont + lines[index].y = edge + lines[index].height / 2 + (f.capHeight - f.ascender) / 2 + edge += lines[index].height + if !lines[index].isEmphasis { + lines[index].x = side * (anchor - lines[index].width) / 2 * staggerRatio + side = -side + } + } + return lines + } + + private func before() -> String { beforeText.trimmingCharacters(in: .whitespaces) } + private func after() -> String { afterText.trimmingCharacters(in: .whitespaces) } + + /// Quebra um trecho de apoio em linhas que cabem em `maxWidth`, palavra a + /// palavra — o mesmo critério de `body_lines` no backend. + private func bodyLines(_ run: String, size: CGFloat, maxWidth: CGFloat) -> [Line] { + guard !run.isEmpty else { return [] } + let nsFont = resolvedFont(bodyFont, size: size, face: "Bold") + var out: [Line] = [] + var current = "" + for word in run.split(separator: " ").map(String.init) { + let trial = current.isEmpty ? word : current + " " + word + if !current.isEmpty, textWidth(trial, nsFont) > maxWidth { + out.append(makeLine(current, nsFont: nsFont, color: bodyColor, isEmphasis: false)) + current = word + } else { + current = trial + } + } + if !current.isEmpty { + out.append(makeLine(current, nsFont: nsFont, color: bodyColor, isEmphasis: false)) + } + return out + } + + private func makeLine(_ text: String, nsFont: NSFont, color: Color, isEmphasis: Bool) -> Line { + Line( + text: text, + font: Font(nsFont as CTFont), + nsFont: nsFont, + color: color, + isEmphasis: isEmphasis, + // Empilha sobre a tinta real (cap-height + descida), como ink_extent. + width: textWidth(text, nsFont), + height: nsFont.capHeight + abs(nsFont.descender) + ) + } + + private func textWidth(_ text: String, _ font: NSFont) -> CGFloat { + (text as NSString).size(withAttributes: [.font: font]).width + } + + /// Resolve família + face para uma `NSFont`, caindo para a fonte do sistema + /// quando a família não está instalada — assim o preview nunca some. + private func resolvedFont(_ family: String, size: CGFloat, face: String?) -> NSFont { + var descriptor = NSFontDescriptor(fontAttributes: [.family: family]) + if let face, !face.isEmpty { + descriptor = descriptor.addingAttributes([.face: face]) + } + if let font = NSFont(descriptor: descriptor, size: size) { + return font + } + var traits: NSFontDescriptor.SymbolicTraits = [] + if face?.localizedCaseInsensitiveContains("italic") == true { traits.insert(.italic) } + if face?.localizedCaseInsensitiveContains("bold") == true { traits.insert(.bold) } + let fallback = NSFontDescriptor(fontAttributes: [.family: family]) + .withSymbolicTraits(traits) + return NSFont(descriptor: fallback, size: size) + ?? NSFont.systemFont(ofSize: size) + } +} diff --git a/code/MacApp/Sources/TranscriptionView.swift b/code/MacApp/Sources/TranscriptionView.swift index 504f611..feee256 100644 --- a/code/MacApp/Sources/TranscriptionView.swift +++ b/code/MacApp/Sources/TranscriptionView.swift @@ -16,6 +16,8 @@ struct TranscriptionView: View { @State private var processedPath: String? @State private var isRemovingSilences = false @State private var silencePadding: Double = 0.05 + @State private var silenceNoiseDb: Double = -30 + @State private var silenceMinDuration: Double = 0.5 @State private var isRemovingFillers = false @State private var fillerRemovedPath: String? @State private var phraseInput = "" @@ -39,27 +41,19 @@ struct TranscriptionView: View { @State private var selectedZoomEndID: Int? @State private var zoomMessage = "" @State private var outputFolder: String? + @State private var batchVoiceAnalysis = false + @State private var batchVoiceEdit = false + @State private var voiceActionsPath: String? @State private var batchSilences = true @State private var batchFillers = false @State private var batchPhrases = false @State private var batchMarkers = false @State private var batchSubtitles = true @State private var batchDynamicSubtitles = false - @State private var dynamicSubtitlesBandHeight: Double = 0.22 - @State private var dynamicSubtitlesBlockCenterY: Double = -167 - @State private var dynamicSubtitlesFont = "Helvetica Neue" - @State private var dynamicSubtitlesFontSize: Double = 90 - @State private var dynamicSubtitlesActiveColor = Color.white - @State private var dynamicSubtitlesInactiveColor = Color(white: 0.7) @State private var dynamicSubtitlesPath: String? @State private var isBatchProcessing = false @State private var batchStatus = "" - private let dynamicSubtitlesFontChoices = [ - "Helvetica Neue", "Helvetica", "Arial", "Avenir Next", - "Futura", "SF Pro Display", "Georgia", "Impact", - ] - init(projectPath: String? = nil, embedded: Bool = false) { self.embedded = embedded _projectPath = State(initialValue: projectPath) @@ -133,6 +127,19 @@ struct TranscriptionView: View { .disabled(isRunning || projectPath == nil || outputFolder == nil || hasNoInstalledModel) .frame(maxWidth: .infinity) .buttonStyle(.borderedProminent) + + if isRunning { + VStack(alignment: .leading, spacing: 6) { + ProgressView(value: progress) + HStack { + Text(stage.isEmpty ? "Processando o áudio…" : stage) + Spacer() + Text("\(Int(progress * 100))%").monospacedDigit() + } + .font(.caption).foregroundStyle(.secondary) + } + .padding(.top, 4) + } } .disabled(outputFolder == nil) @@ -143,15 +150,6 @@ struct TranscriptionView: View { } } - if isRunning { - Section { - VStack(alignment: .leading, spacing: 8) { - ProgressView(value: progress) - Text("\(Int(progress * 100))%").font(.caption).foregroundStyle(.secondary) - } - } - } - if !results.isEmpty { Section("Transcrição Concluída") { ForEach(results, id: \.media) { r in @@ -176,105 +174,174 @@ struct TranscriptionView: View { if outputFolder == nil { Text("Selecione a pasta do projeto para habilitar o processamento.") .font(.caption2).foregroundStyle(.secondary) + } else if results.isEmpty { + Label( + isRunning ? "Aguardando a transcrição terminar…" : "Transcreva o projeto acima para liberar o processamento.", + systemImage: "lock.fill" + ) + .font(.caption).foregroundStyle(.orange) } - GroupBox("Processar em lote") { - VStack(alignment: .leading, spacing: 6) { - Toggle("Remover silêncios do áudio", isOn: $batchSilences) - VStack(alignment: .leading, spacing: 4) { - HStack { - Text("Tolerância do corte") - .foregroundStyle(batchSilences ? .primary : .secondary) - Spacer() - Text(String(format: "%.2fs", silencePadding)) - .font(.caption).foregroundStyle(.secondary).monospacedDigit() - } - Slider(value: $silencePadding, in: 0...2, step: 0.01) - .disabled(!batchSilences) - Text("Quanto de silêncio sobra em volta de cada corte.") - .font(.caption2).foregroundStyle(.secondary) + GroupBox { + VStack(alignment: .leading, spacing: 18) { + batchOptionRow( + toggle: Toggle("Analisar voz (transcrição, locutor, ênfase)", isOn: $batchVoiceAnalysis), + expanded: batchVoiceAnalysis + ) { + Text("Gera o JSON com transcrição, diarização e intensidade (pitch/energia/ritmo) por palavra — a base que o corte por voz usa para decidir tomadas e zooms sem reabrir o áudio depois. Só análise: não corta nada.") + .font(.caption).foregroundStyle(.secondary) } - .padding(.leading, 20) + + Divider() + batchOptionRow( + toggle: Toggle("Aplicar edição por voz (lista de decisões)", isOn: $batchVoiceEdit), + expanded: batchVoiceEdit + ) { + VStack(alignment: .leading, spacing: 6) { + Text("Aplica um JSON de decisões — cortes, zooms, textos e marcadores — gerado a partir da análise de voz. É o passo que descarta bastidor e tomadas repetidas.") + .font(.caption).foregroundStyle(.secondary) + HStack { + Image(systemName: "doc.text") + Text(voiceActionsPath.map { URL(fileURLWithPath: $0).lastPathComponent } + ?? "Nenhum arquivo de decisões") + .foregroundStyle(voiceActionsPath == nil ? .secondary : .primary) + .lineLimit(1).truncationMode(.middle) + Spacer() + Button("Escolher…") { pickVoiceActions() } + } + } + } + + Divider() + batchOptionRow( + toggle: Toggle("Remover silêncios do áudio", isOn: $batchSilences), + expanded: true + ) { + VStack(alignment: .leading, spacing: 6) { + HStack { + Text("Tolerância do corte") + .font(.callout) + .foregroundStyle(.secondary) + Spacer() + Text(String(format: "%.2fs", silencePadding)) + .font(.callout).foregroundStyle(.secondary).monospacedDigit() + } + Slider(value: $silencePadding, in: 0...2, step: 0.01) + .disabled(!batchSilences) + .onChange(of: silencePadding) { _, _ in saveSilenceConfig() } + HStack { + Text("Limiar de silêncio") + .font(.callout) + .foregroundStyle(.secondary) + Spacer() + Text(String(format: "%.0f dB", silenceNoiseDb)) + .font(.callout).foregroundStyle(.secondary).monospacedDigit() + } + Slider(value: $silenceNoiseDb, in: -60...(-10), step: 1) + .disabled(!batchSilences) + .onChange(of: silenceNoiseDb) { _, _ in saveSilenceConfig() } + HStack { + Text("Duração mínima") + .font(.callout) + .foregroundStyle(.secondary) + Spacer() + Text(String(format: "%.2fs", silenceMinDuration)) + .font(.callout).foregroundStyle(.secondary).monospacedDigit() + } + Slider(value: $silenceMinDuration, in: 0.1...5, step: 0.05) + .disabled(!batchSilences) + .onChange(of: silenceMinDuration) { _, _ in saveSilenceConfig() } + Text("Quanto de silêncio sobra em volta de cada corte, abaixo de que volume conta como silêncio, e quanto tempo ele precisa durar. Fica salvo e vale também fora do app.") + .font(.caption).foregroundStyle(.secondary) + } + } + + Divider() Toggle("Remover palavras de preenchimento", isOn: $batchFillers) - Toggle("Cortar frases ditas", isOn: $batchPhrases) - if batchPhrases { + + Divider() + batchOptionRow( + toggle: Toggle("Cortar frases ditas", isOn: $batchPhrases), + expanded: batchPhrases + ) { TextField("Frases separadas por vírgula", text: $phraseInput) .textFieldStyle(.roundedBorder) - .padding(.leading, 20) } + + Divider() Toggle("Marcar o que foi dito na timeline", isOn: $batchMarkers) + + Divider() Toggle("Exportar legendas SRT", isOn: $batchSubtitles) - Toggle("Gerar legendas dinâmicas (cascata, editáveis no FCP)", isOn: $batchDynamicSubtitles) - if batchDynamicSubtitles { - VStack(alignment: .leading, spacing: 6) { - Picker("Fonte", selection: $dynamicSubtitlesFont) { - ForEach(dynamicSubtitlesFontChoices, id: \.self) { Text($0).tag($0) } - } - HStack { - Text("Tamanho da fonte") - Spacer() - Text("\(Int(dynamicSubtitlesFontSize))pt") - .font(.caption).foregroundStyle(.secondary).monospacedDigit() - } - Slider(value: $dynamicSubtitlesFontSize, in: 20...200, step: 1) - ColorPicker("Cor A", selection: $dynamicSubtitlesActiveColor, supportsOpacity: true) - ColorPicker("Cor B", selection: $dynamicSubtitlesInactiveColor, supportsOpacity: true) - HStack { - Text("Altura do bloco") - Spacer() - Text("\(Int(dynamicSubtitlesBandHeight * 100))%") - .font(.caption).foregroundStyle(.secondary).monospacedDigit() - } - Slider(value: $dynamicSubtitlesBandHeight, in: 0.10...0.50, step: 0.01) - HStack { - Text("Altura na tela") - Spacer() - Text("\(Int(dynamicSubtitlesBlockCenterY))") - .font(.caption).foregroundStyle(.secondary).monospacedDigit() - } - Slider(value: $dynamicSubtitlesBlockCenterY, in: -700...300, step: 1) - Text("Roda por último, depois dos outros passos marcados acima, usando o timing já cortado da timeline.") - .font(.caption2).foregroundStyle(.secondary) - } - .padding(.leading, 20) - } - Button { - processBatch() - } label: { - if isBatchProcessing { ProgressView().controlSize(.small) } - Label("Processar selecionados", systemImage: "play.fill") - } - .disabled(isBatchProcessing || isRunning || projectPath == nil || outputFolder == nil - || !batchSilences && !batchFillers && !batchPhrases && !batchMarkers - && !batchSubtitles && !batchDynamicSubtitles) - if !batchStatus.isEmpty { - Text(batchStatus).font(.caption).foregroundStyle(.secondary) - } - HStack { - Button("Abrir pasta selecionada") { - if let outputFolder { - NSWorkspace.shared.open(URL(fileURLWithPath: outputFolder)) - } - } - .disabled(outputFolder == nil) - Button("Abrir no Final Cut Pro") { - if let path = dynamicSubtitlesPath ?? processedPath ?? transcriptEditPath ?? transcriptMarkersPath { - NSWorkspace.shared.open(URL(fileURLWithPath: path)) - } - } - .disabled(dynamicSubtitlesPath == nil && processedPath == nil - && transcriptEditPath == nil && transcriptMarkersPath == nil) + + Divider() + batchOptionRow( + toggle: Toggle("Gerar legendas dinâmicas (cascata, editáveis no FCP)", isOn: $batchDynamicSubtitles), + expanded: batchDynamicSubtitles + ) { + Text("Usa o estilo configurado na aba \"Legendas Dinâmicas\". Roda por último, depois dos outros passos marcados acima, usando o timing já cortado da timeline.") + .font(.caption).foregroundStyle(.secondary) } } + .padding(.vertical, 8) } - if results.isEmpty { - Text("Transcreva primeiro para habilitar os cortes por texto.") - .font(.caption2).foregroundStyle(.secondary) + VStack(alignment: .leading, spacing: 12) { + Button { + processBatch() + } label: { + if isBatchProcessing { + HStack { + ProgressView().controlSize(.small) + Text("Processando…") + } + .frame(maxWidth: .infinity) + } else { + Label("Processar selecionados", systemImage: "play.fill") + .frame(maxWidth: .infinity) + } + } + .buttonStyle(.borderedProminent) + .controlSize(.large) + .disabled(isBatchProcessing || isRunning || projectPath == nil || outputFolder == nil + || !batchVoiceAnalysis && !batchVoiceEdit && !batchSilences && !batchFillers + && !batchPhrases && !batchMarkers && !batchSubtitles && !batchDynamicSubtitles + // Ligada sem arquivo, a etapa falharia no meio da + // cadeia e interromperia tudo o que vem depois. + || batchVoiceEdit && voiceActionsPath == nil) + + if !batchStatus.isEmpty { + Text(batchStatus).font(.caption).foregroundStyle(.secondary) + } + + HStack(spacing: 12) { + Button("Abrir pasta selecionada") { + if let outputFolder { + NSWorkspace.shared.open(URL(fileURLWithPath: outputFolder)) + } + } + .disabled(outputFolder == nil) + Button("Abrir no Final Cut Pro") { + if let path = dynamicSubtitlesPath ?? processedPath ?? transcriptEditPath ?? transcriptMarkersPath { + NSWorkspace.shared.open(URL(fileURLWithPath: path)) + } + } + .disabled(dynamicSubtitlesPath == nil && processedPath == nil + && transcriptEditPath == nil && transcriptMarkersPath == nil) + Spacer() + } } + .padding(.top, 4) } - .disabled(outputFolder == nil) + .disabled(outputFolder == nil || results.isEmpty) Section("Zoom (aproximar em um trecho)") { + if outputFolder != nil && results.isEmpty { + Label( + isRunning ? "Aguardando a transcrição terminar…" : "Transcreva o projeto acima para liberar o zoom.", + systemImage: "lock.fill" + ) + .font(.caption).foregroundStyle(.orange) + } if zoomClips.isEmpty { Text("Carregando clipes do projeto…") .font(.caption).foregroundStyle(.secondary) @@ -347,12 +414,47 @@ struct TranscriptionView: View { resultRow(zoomResultPath) } } - .disabled(outputFolder == nil) + .disabled(outputFolder == nil || results.isEmpty) } .formStyle(.grouped) - .task { await loadCatalog(); loadZoomClips() } - .onChange(of: projectPath) { _, _ in loadZoomClips() } - .onChange(of: outputFolder) { _, _ in loadZoomClips() } + .task { await loadCatalog(); loadProjectConfig(); loadZoomClips(); loadSilenceConfig() } + .onChange(of: projectPath) { _, newValue in + loadZoomClips() + saveProjectConfig(file: newValue ?? "") + } + .onChange(of: outputFolder) { _, newValue in + loadZoomClips() + saveProjectConfig(folder: newValue ?? "") + } + } + + /// Reabre o último projeto salvo em ~/.fcp-mcp-server/config.json — pasta + /// de saída e arquivo — para a tela não começar vazia a cada abertura. + /// Só preenche o que ainda está vazio, então o projeto aberto pela aba + /// Projeto (embedded) nunca é sobrescrito pelo que ficou salvo. Caminhos + /// que sumiram do disco voltam vazios do Python e são ignorados aqui. + private func loadProjectConfig() { + PythonBridge.call(command: "project_config") { result, _ in + DispatchQueue.main.async { + guard let result, result["ok"] as? Bool == true else { return } + if outputFolder == nil, let folder = result["folder"] as? String, !folder.isEmpty { + outputFolder = folder + } + if projectPath == nil, let file = result["file"] as? String, !file.isEmpty { + projectPath = file + } + } + } + } + + /// Grava a pasta e/ou o arquivo do projeto no config compartilhado. Campos + /// omitidos ficam como estão; string vazia limpa o campo. + private func saveProjectConfig(folder: String? = nil, file: String? = nil) { + var arguments: [String: Any] = [:] + if let folder { arguments["folder"] = folder } + if let file { arguments["file"] = file } + guard !arguments.isEmpty else { return } + PythonBridge.call(command: "set_project_config", arguments: arguments) { _, _ in } } /// Presents an NSOpenPanel configured to select a single file (not a @@ -374,6 +476,45 @@ struct TranscriptionView: View { } } + /// Lê os limiares de silêncio salvos em ~/.fcp-mcp-server/config.json — + /// os mesmos que remove_media_silence usa por padrão — para os controles + /// abrirem já mostrando o que de fato vai rodar. + private func loadSilenceConfig() { + PythonBridge.call(command: "silence_config") { result, _ in + DispatchQueue.main.async { + guard let result, result["ok"] as? Bool == true else { return } + silencePadding = result["padding"] as? Double ?? silencePadding + silenceNoiseDb = result["noise_db"] as? Double ?? silenceNoiseDb + silenceMinDuration = result["min_silence"] as? Double ?? silenceMinDuration + } + } + } + + private func saveSilenceConfig() { + PythonBridge.call(command: "set_silence_config", arguments: [ + "padding": silencePadding, + "noise_db": silenceNoiseDb, + "min_silence": silenceMinDuration, + ]) { _, _ in } + } + + /// Escolhe o JSON de decisões (cortes/zooms/textos/marcadores) que + /// `apply_voice_actions` vai aplicar. Pré-seleciona a pasta do projeto, + /// que é onde o arquivo costuma ser gravado. + private func pickVoiceActions() { + let panel = NSOpenPanel() + panel.canChooseFiles = true + panel.canChooseDirectories = false + panel.allowsMultipleSelection = false + panel.allowedContentTypes = [.json] + panel.prompt = "Usar este arquivo" + panel.message = "Selecione o JSON com a lista de decisões da edição por voz." + if let outputFolder { panel.directoryURL = URL(fileURLWithPath: outputFolder) } + if panel.runModal() == .OK, let url = panel.url { + voiceActionsPath = url.path + } + } + private func pickOutputFolder() { let panel = NSOpenPanel() panel.canChooseFiles = false @@ -438,6 +579,28 @@ struct TranscriptionView: View { } } + /// Renders a toggle plus its optional expanded detail block, indented and + /// visually tied together so batch options don't collapse into one wall + /// of controls with no breathing room. + @ViewBuilder + private func batchOptionRow<Toggle: View, Detail: View>( + toggle: Toggle, expanded: Bool, @ViewBuilder detail: () -> Detail + ) -> some View { + VStack(alignment: .leading, spacing: 12) { + toggle + if expanded { + detail() + .padding(.leading, 20) + .padding(.vertical, 10) + .padding(.trailing, 8) + .background( + RoundedRectangle(cornerRadius: 8) + .fill(Color.secondary.opacity(0.06)) + ) + } + } + } + @ViewBuilder private func resultRow(_ path: String) -> some View { Text(path).font(.caption).lineLimit(1) @@ -596,9 +759,11 @@ struct TranscriptionView: View { guard let projectPath else { return } isRemovingSilences = true errorMessage = nil + // Sem repassar limiares: remove_silences lê os valores salvos em + // ~/.fcp-mcp-server/config.json (os mesmos que os controles acima + // gravam), então há uma só fonte da verdade. PythonBridge.call(command: "remove_silences", arguments: [ "path": projectPath, - "padding": silencePadding, ]) { result, err in DispatchQueue.main.async { isRemovingSilences = false @@ -614,6 +779,16 @@ struct TranscriptionView: View { private func processBatch() { guard let projectPath, let outputFolder else { return } var operations: [String] = [] + // Analysis first: it only writes a JSON sidecar next to the source + // media and never touches the project XML, so its position relative + // to the other steps doesn't change what they do — but running it + // before any cut keeps the mental model simple (measure, then edit). + if batchVoiceAnalysis { operations.append("analyze_voice") } + // The edit itself, and it has to run FIRST on the intact timeline: + // every zoom/text/marker it places is positioned by shifting from its + // own cut list, so a step that already rippled the timeline would put + // them on the wrong frame — with no visible error. + if batchVoiceEdit { operations.append("apply_voice_actions") } if batchSilences { operations.append("remove_silences") } if batchFillers { operations.append("remove_filler_words") } if batchPhrases { operations.append("edit_by_transcript") } @@ -638,19 +813,17 @@ struct TranscriptionView: View { let operation = operations[index] batchStatus = "Processando: \(operation)…" var arguments: [String: Any] = ["path": currentPath, "output_dir": outputFolder] - if operation == "remove_silences" { arguments["padding"] = silencePadding } + if operation == "apply_voice_actions", let voiceActionsPath { + arguments["actions_path"] = voiceActionsPath + } if operation == "edit_by_transcript" { let phrases = phraseInput.split(separator: ",").map { $0.trimmingCharacters(in: .whitespaces) }.filter { !$0.isEmpty } arguments["phrases"] = phrases } - if operation == "generate_dynamic_subtitles" { - arguments["band_height"] = dynamicSubtitlesBandHeight - arguments["block_center_y"] = dynamicSubtitlesBlockCenterY - arguments["font"] = dynamicSubtitlesFont - arguments["font_size"] = Int(dynamicSubtitlesFontSize) - arguments["active_color"] = dynamicSubtitlesActiveColor.fcpxmlColorString - arguments["inactive_color"] = dynamicSubtitlesInactiveColor.fcpxmlColorString - } + // O estilo das legendas não é repassado aqui: generate_dynamic_subtitles + // lê ~/.fcp-mcp-server/config.json (o mesmo arquivo que a aba "Legendas + // Dinâmicas" grava) como seu próprio padrão, então há uma só fonte da + // verdade em vez de duas cópias podendo divergir. PythonBridge.call(command: operation, arguments: arguments) { result, err in DispatchQueue.main.async { guard result?["ok"] as? Bool == true else { diff --git a/code/MacApp/Sources/VoiceAnalysisView.swift b/code/MacApp/Sources/VoiceAnalysisView.swift new file mode 100644 index 0000000..51b3842 --- /dev/null +++ b/code/MacApp/Sources/VoiceAnalysisView.swift @@ -0,0 +1,258 @@ +import SwiftUI + +/// Guia "Análise de Voz" — parâmetros do motor de análise acústica da fala. +/// +/// Controla os limiares que decidem quais palavras são candidatas a +/// punch-in/destaque: energia (intensidade da fala), o índice de ênfase e +/// seus pesos, e a camada opcional de emoção. Os valores ficam em +/// `~/.fcp-mcp-server/config.json` (via `model_manager.save_voice_analysis_config`), +/// os mesmos lidos pelas ferramentas de análise — não em UserDefaults, para +/// que a análise rode com exatamente o que a tela mostra. +struct VoiceAnalysisView: View { + @State private var config = VoiceAnalysisConfig.defaults + @State private var isLoading = true + @State private var errorMessage: String? + + var body: some View { + Form { + if isLoading { + Section { + ProgressView().controlSize(.small) + .frame(maxWidth: .infinity, alignment: .center) + } + } else { + energySection + emphasisSection + weightsSection + emotionSection + resetSection + } + if let errorMessage { + Section { + Label(errorMessage, systemImage: "exclamationmark.triangle.fill") + .foregroundStyle(.red) + } + } + } + .formStyle(.grouped) + .task { await load() } + } + + // MARK: - Energia + + private var energySection: some View { + Section { + sliderRow( + title: "Limiar de energia", + value: $config.energyThreshold, + help: "Acima deste valor a fala conta como \"alta energia\"." + ) + } header: { + Text("Energia da Fala") + } footer: { + Text("A energia de cada palavra é normalizada (0–1) pelo trecho mais alto do áudio. Valores mais baixos marcam mais palavras como intensas.") + .font(.caption) + .foregroundStyle(.secondary) + } + } + + // MARK: - Ênfase + + private var emphasisSection: some View { + Section { + sliderRow( + title: "Limiar de ênfase", + value: $config.emphasisThreshold, + help: "Acima deste valor a palavra vira candidata a punch-in/destaque." + ) + } header: { + Text("Índice de Ênfase") + } footer: { + Text("O índice é a média ponderada de energia, variação de tom, variação de ritmo, pausa anterior e duração — com os pesos abaixo. Por ser média, picos reais ficam entre 0,55 e 0,80: acima de 0,85 quase nada é selecionado.") + .font(.caption) + .foregroundStyle(.secondary) + } + } + + private var weightsSection: some View { + Section { + sliderRow(title: "Energia", value: $config.weightEnergy, range: 0...1) + sliderRow(title: "Variação de tom", value: $config.weightPitch, range: 0...1) + sliderRow(title: "Variação de ritmo", value: $config.weightRate, range: 0...1) + sliderRow(title: "Pausa anterior", value: $config.weightPause, range: 0...1) + sliderRow(title: "Duração da palavra", value: $config.weightDuration, range: 0...1) + } header: { + Text("Pesos do Índice de Ênfase") + } footer: { + Text("Não precisam somar 1 — são normalizados internamente. O que importa é a proporção entre eles.") + .font(.caption) + .foregroundStyle(.secondary) + } + } + + // MARK: - Emoção + + private var emotionSection: some View { + Section { + Toggle("Detectar emoção durante a fala", isOn: $config.emotionEnabled) + .onChange(of: config.emotionEnabled) { _, _ in save() } + if config.emotionEnabled { + sliderRow( + title: "Sensibilidade", + value: $config.emotionSensitivity, + help: "Confiança mínima para aceitar uma emoção detectada." + ) + } + } header: { + Text("Emoção") + } footer: { + Text("A emoção nunca decide um corte sozinha — entra combinada com energia, tom e ênfase. Exige o componente opcional de emoção instalado; sem ele, a análise segue normalmente sem essa camada.") + .font(.caption) + .foregroundStyle(.secondary) + } + } + + private var resetSection: some View { + Section { + Button("Restaurar padrões") { + config = .defaults + save() + } + } + } + + // MARK: - Componentes + + /// Slider com rótulo à esquerda e valor numérico à direita, salvando no + /// backend só quando o arraste termina (evita uma escrita por quadro). + private func sliderRow( + title: String, + value: Binding<Double>, + range: ClosedRange<Double> = 0...1, + help: String? = nil + ) -> some View { + VStack(alignment: .leading, spacing: 2) { + HStack { + Text(title) + Spacer() + Text(String(format: "%.2f", value.wrappedValue)) + .monospacedDigit() + .foregroundStyle(.secondary) + } + Slider(value: value, in: range) { editing in + if !editing { save() } + } + if let help { + Text(help) + .font(.caption) + .foregroundStyle(.secondary) + } + } + } + + // MARK: - Backend + + @MainActor + private func load() async { + await withCheckedContinuation { continuation in + PythonBridge.call(command: "voice_analysis") { result, error in + DispatchQueue.main.async { + if let result { + config = VoiceAnalysisConfig(from: result) + } else if let error { + errorMessage = error + } + isLoading = false + continuation.resume() + } + } + } + } + + private func save() { + PythonBridge.call(command: "set_voice_analysis", arguments: config.arguments()) { _, error in + DispatchQueue.main.async { errorMessage = error } + } + } +} + +/// Os parâmetros de análise de voz, no formato que a tela edita e o bridge +/// (`admin/models_api.py` → `set_voice_analysis`) persiste. +struct VoiceAnalysisConfig { + var energyThreshold: Double + var emphasisThreshold: Double + var weightEnergy: Double + var weightPitch: Double + var weightRate: Double + var weightPause: Double + var weightDuration: Double + var emotionEnabled: Bool + var emotionSensitivity: Double + + static let defaults = VoiceAnalysisConfig( + energyThreshold: 0.5, + emphasisThreshold: 0.60, + weightEnergy: 0.30, + weightPitch: 0.25, + weightRate: 0.20, + weightPause: 0.15, + weightDuration: 0.10, + emotionEnabled: false, + emotionSensitivity: 0.5 + ) + + init( + energyThreshold: Double, + emphasisThreshold: Double, + weightEnergy: Double, + weightPitch: Double, + weightRate: Double, + weightPause: Double, + weightDuration: Double, + emotionEnabled: Bool, + emotionSensitivity: Double + ) { + self.energyThreshold = energyThreshold + self.emphasisThreshold = emphasisThreshold + self.weightEnergy = weightEnergy + self.weightPitch = weightPitch + self.weightRate = weightRate + self.weightPause = weightPause + self.weightDuration = weightDuration + self.emotionEnabled = emotionEnabled + self.emotionSensitivity = emotionSensitivity + } + + /// Lê a resposta do bridge, caindo no padrão para qualquer campo ausente. + init(from json: [String: Any]) { + let defaults = VoiceAnalysisConfig.defaults + let weights = json["emphasis_weights"] as? [String: Any] ?? [:] + self.init( + energyThreshold: json["energy_threshold"] as? Double ?? defaults.energyThreshold, + emphasisThreshold: json["emphasis_threshold"] as? Double ?? defaults.emphasisThreshold, + weightEnergy: weights["energy"] as? Double ?? defaults.weightEnergy, + weightPitch: weights["pitch_variation"] as? Double ?? defaults.weightPitch, + weightRate: weights["rate_variation"] as? Double ?? defaults.weightRate, + weightPause: weights["pause_before"] as? Double ?? defaults.weightPause, + weightDuration: weights["duration"] as? Double ?? defaults.weightDuration, + emotionEnabled: json["emotion_enabled"] as? Bool ?? defaults.emotionEnabled, + emotionSensitivity: json["emotion_sensitivity"] as? Double ?? defaults.emotionSensitivity + ) + } + + func arguments() -> [String: Any] { + [ + "energy_threshold": energyThreshold, + "emphasis_threshold": emphasisThreshold, + "emphasis_weights": [ + "energy": weightEnergy, + "pitch_variation": weightPitch, + "rate_variation": weightRate, + "pause_before": weightPause, + "duration": weightDuration, + ], + "emotion_enabled": emotionEnabled, + "emotion_sensitivity": emotionSensitivity, + ] + } +} diff --git a/code/README.md b/code/README.md index 0963229..53f43c4 100755 --- a/code/README.md +++ b/code/README.md @@ -1,6 +1,6 @@ # FCPXML MCP -**The bridge between Final Cut Pro and AI. 62 tools that turn timeline XML into structured data Claude can read, edit, and generate.** +**The bridge between Final Cut Pro and AI. 73 tools that turn timeline XML into structured data Claude can read, edit, and generate.** [![CI](https://github.com/DareDev256/fcp-mcp-server/actions/workflows/test.yml/badge.svg)](https://github.com/DareDev256/fcp-mcp-server/actions) [![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg)](https://opensource.org/licenses/MIT) @@ -410,7 +410,7 @@ Select these from Claude's prompt menu (⌘/) — they chain multiple tools auto ``` fcp-mcp-server/ ~9.4k lines Python -├── server.py MCP entry point — 62 tools, 5 prompts, resource discovery +├── server.py MCP entry point — 73 tools, 5 prompts, resource discovery │ _resolve_io_paths() / _setup_modifier() / _setup_generator() │ _format_clip_table() / _markdown_table() / _format_batch_result() │ _raw_markers_to_batch() @@ -430,7 +430,7 @@ fcp-mcp-server/ ~9.4k lines Python │ ├── safe_xml.py Centralized defusedxml wrappers (XXE/entity-bomb protection) + serialize_xml() │ ├── dtd.py Validate output against Apple's official DTDs (located in the FCP app bundle) │ └── templates.py Template system (intro/outro, lower thirds, music video) -├── tests/ 1032 tests across 24 suites +├── tests/ 1342 tests across 34 suites │ ├── test_models.py TimeValue math, Timecode formatting, MarkerType contracts │ ├── test_parser.py FCPXML parsing, connected clips, edge cases │ ├── test_writer.py Clip editing, marker writing, speed changes @@ -568,7 +568,7 @@ uv run --extra dev pytest tests/ -v # or: python3 -m pytest tests/ -v ruff check . --exclude docs/ # lint — must pass before committing ``` -1032 tests across 24 suites covering models, parser, writer, FCPXMLWriter generation, server handlers, rough cut generation, speed cutting & pacing curves, marker pipeline, refactored helper functions, regression fixes, security hardening (XXE, entity expansion, path traversal, sandbox boundaries, minidom defense-in-depth, JSON depth limits, input validation, ffmpeg bounds, write-handler sandboxing), connected clips, roles, diff, export, compound clip flattening, audio track generation, templates, effects, `.fcpxmld` bundles with sidecar preservation, bulk media relink, real media silence detection (parser, timeline mapping, real-WAV ffmpeg integration), and DTD validation against Apple's official DTDs (auto-skipped on machines without Final Cut Pro). +1342 tests across 34 suites covering models, parser, writer, FCPXMLWriter generation, server handlers, rough cut generation, speed cutting & pacing curves, marker pipeline, refactored helper functions, regression fixes, security hardening (XXE, entity expansion, path traversal, sandbox boundaries, minidom defense-in-depth, JSON depth limits, input validation, ffmpeg bounds, write-handler sandboxing), connected clips, roles, diff, export, compound clip flattening, audio track generation, templates, effects, `.fcpxmld` bundles with sidecar preservation, bulk media relink, real media silence detection (parser, timeline mapping, real-WAV ffmpeg integration), and DTD validation against Apple's official DTDs (auto-skipped on machines without Final Cut Pro). --- diff --git a/code/docs/WORKFLOWS.md b/code/docs/WORKFLOWS.md index 90e7a8e..551a273 100755 --- a/code/docs/WORKFLOWS.md +++ b/code/docs/WORKFLOWS.md @@ -168,6 +168,29 @@ Mark mode adds markers instead of deleting — safer for first pass. --- +## Dynamic Subtitles: Generate, Then Always Validate + +**Scenario:** Word-by-word progressive-composition subtitles (the diagrammed look — small supporting words, one key word large in a display italic) need to go on a cut before delivery. + +``` +"Generate dynamic subtitles for /path/to/project.fcpxml" +``` + +**Tool chain:** `generate_dynamic_subtitles` → `validate_subtitle_layout` + +Run `generate_dynamic_subtitles` on the *final* cut, after cuts/zooms are already applied — a connected title anchors to its parent clip's source-media coordinates, so re-cutting the timeline afterward can silently detach captions from the words they were built for. + +**Never treat generation as done without the second call.** The layout only guarantees non-overlap *by construction* for what it itself lays out — it cannot see a hand-edited title, stray content left over in a reused base file, or a word long/uppercase enough to have needed shrinking. `validate_subtitle_layout` re-measures every `<title>` independently and reports a severity (`none`/`warning`/`probable`/`severe`) plus per-issue suggested corrections: + +``` +"Validate the subtitle layout in /path/to/project_dynamic_subtitles.fcpxmld" +``` + +- `none`/`warning` (only `outside_safe_area`, no `outside_frame` or `collision`) — safe to deliver; the emphasis word sitting close to the 5% margin is expected on the diagrammed look. +- `probable`/`severe` — investigate before touching code. Read the issue's exact FCPXML fraction times (not the rounded float) before deciding whether it's a real overlap; two titles that are only touching at a shared boundary can still round to "equal-looking but not bit-identical" floats and misreport. See `Engine/docs/03_SERVER_TOOLS.md`'s "Legendas dinâmicas" section and `Engine/docs/05_EXPERIENCIAS.md` (2026-08-19 entries) for the concrete bugs already found and fixed this way, and the checklist for the next one. + +--- + ## Composing Tools in AI Agent Workflows Each tool in this MCP server follows the same pattern: read FCPXML → process → write modified FCPXML. This makes them composable — the output of one tool is valid input for the next. diff --git a/code/fcpxml/collision.py b/code/fcpxml/collision.py new file mode 100644 index 0000000..9d7efcb --- /dev/null +++ b/code/fcpxml/collision.py @@ -0,0 +1,472 @@ +"""Collision detection and layout validation for dynamic-subtitle titles. + +Pure functions — no I/O, no FCPXML parsing — that answer one question over and +over: given the boxes a set of titles occupy on screen, do any two titles that +are on screen at the same time intersect? And are they inside the frame, inside +the safe area, and using a font the layout actually measured? + +This is the post-generation guarantee the layout engine only provides *by +construction* (``text_layout.compose_sentence`` stacks lines so their ink boxes +never touch). Re-running it over already-emitted titles catches the cases the +layout cannot see: a hand-edited position, a template whose type scales +differently than ``text_scale`` assumed, a font that fell back to an estimate, +or a word pushed off frame by a long emphasis line. + +Boxes are measured in the *emitted* template space (frame pixels) — the same +numbers the writer wrote to the FCPXML (``fontSize``, ``kerning`` and +``Position`` are all already scaled by ``text_scale``), so validation re-measures +with ``measure_text``/``ink_extent`` against those same numbers and never +re-applies the scale factor. See ``writer.validate_subtitle_layout``. +""" + +from dataclasses import dataclass +from math import hypot +from typing import Dict, List, Optional, Sequence + +from .text_layout import ( + ink_extent, + measure_text, + metrics_for, + vertical_metrics_for, +) + +# Severity buckets for a spatial overlap, ordered from harmless to blocking. +# ``render_tolerance`` is the 5px the renderer can round off; ``severe`` is a +# real collision that must be fixed before export. +OVERLAP_NONE = "none" +OVERLAP_RENDER_TOLERANCE = "render_tolerance" +OVERLAP_WARNING = "warning" +OVERLAP_PROBABLE = "probable" +OVERLAP_SEVERE = "severe" + +# Max fraction of the smaller box a severe collision may cover (spec 7.2). +SEVERE_OVERLAP_RATIO = 0.15 + +# issue types (spec 16) +SPATIAL_COLLISION = "spatial_collision" +OUTSIDE_FRAME = "outside_frame" +OUTSIDE_SAFE_AREA = "outside_safe_area" +INSUFFICIENT_SPACING = "insufficient_spacing" +EXCESSIVE_SPACING = "excessive_spacing" +FONT_MISSING = "font_missing" +FONT_TOO_SMALL = "font_too_small" +INVALID_ANCHOR = "invalid_anchor" +UNRESOLVED_TRANSFORM = "unresolved_transform" + + +@dataclass +class Box: + """An axis-aligned rectangle in frame coordinates, y growing upward.""" + + left: float + right: float + bottom: float + top: float + + @property + def width(self) -> float: + return self.right - self.left + + @property + def height(self) -> float: + return self.top - self.bottom + + @property + def area(self) -> float: + return self.width * self.height + + def overlaps(self, other: "Box") -> bool: + """True if the two boxes intersect (strict — touching edges do not).""" + return ( + self.left < other.right + and other.left < self.right + and self.bottom < other.top + and other.bottom < self.top + ) + + +# A boundary the writer places deliberately exact — one block's title +# duration set to literally equal the next block's start (see writer.py's +# ``block_ends``) — can still land a few float-ULPs apart by the time it +# gets here: an ``end`` re-derived as ``start + duration`` from two already- +# rounded floats isn't bit-identical to a ``start`` read as one division of +# the same exact fraction, even though both trace back to one FCPXML value. +# Found on real footage: 7 of 8 "severe" collisions in one clip were exactly +# this — same instant, off by ~1e-13s, nowhere near a real frame boundary +# (~0.04s). A tolerance many orders below one frame absorbs the artifact +# without hiding a genuine overlap. +_BOUNDARY_EPSILON = 1e-6 + + +def temporal_overlap( + start_a: float, end_a: float, start_b: float, end_b: float +) -> bool: + """Whether the half-open intervals ``[start, end)`` intersect (spec 7.1). + + Strict on both sides, so a title that ends exactly when the next begins is + never treated as simultaneous — see ``_BOUNDARY_EPSILON`` for why "exactly" + needs a tolerance rather than bare float comparison. + """ + return ( + start_a < end_b - _BOUNDARY_EPSILON + and start_b < end_a - _BOUNDARY_EPSILON + ) + + +def overlap_metrics(a: Box, b: Box) -> Dict[str, float]: + """Width, height, area and ratio of the intersection of ``a`` and ``b``. + + ``overlap_ratio`` is the shared area over the *smaller* box's area, so a + small box swallowed by a big one reads as the severe case it is. + """ + overlap_width = min(a.right, b.right) - max(a.left, b.left) + overlap_height = min(a.top, b.top) - max(a.bottom, b.bottom) + overlap_area = max(0.0, overlap_width) * max(0.0, overlap_height) + smaller = min(a.area, b.area) + ratio = overlap_area / smaller if smaller > 0 else 0.0 + return { + "overlap_width": overlap_width, + "overlap_height": overlap_height, + "overlap_area": overlap_area, + "overlap_ratio": ratio, + } + + +def classify_overlap(metrics: Dict[str, float]) -> str: + """Severity bucket for an overlap, following spec 7.2. + + Zero area is no conflict at all; a ratio above ``SEVERE_OVERLAP_RATIO`` is + severe regardless of absolute size; otherwise the vertical penetration is + bucketed into tolerance / warning / probable / severe. + """ + height = metrics["overlap_height"] + if metrics["overlap_area"] <= 0: + return OVERLAP_NONE + if metrics["overlap_ratio"] > SEVERE_OVERLAP_RATIO: + return OVERLAP_SEVERE + if height <= 5: + return OVERLAP_RENDER_TOLERANCE + if height <= 20: + return OVERLAP_WARNING + if height <= 50: + return OVERLAP_PROBABLE + return OVERLAP_SEVERE + + +def distance_between(a: Box, b: Box) -> Dict[str, float]: + """Gap between two non-overlapping boxes, per axis and euclidean (spec 8).""" + if a.right < b.left: + distance_x = b.left - a.right + elif b.right < a.left: + distance_x = a.left - b.right + else: + distance_x = 0.0 + + if a.top < b.bottom: + distance_y = b.bottom - a.top + elif b.top < a.bottom: + distance_y = a.bottom - b.top + else: + distance_y = 0.0 + + return { + "distance_x": distance_x, + "distance_y": distance_y, + "distance": hypot(distance_x, distance_y), + } + + +def separation_suggestion( + a: Box, b: Box, min_gap: float = 0.0 +) -> Dict[str, float]: + """The minimum translation that separates two overlapping boxes (spec 10). + + Picks the smallest of the four penetrations (move left/right/up/down) and + reports that axis plus the required movement (penetration + ``min_gap``). + """ + move_left = a.right - b.left + move_right = b.right - a.left + move_down = a.top - b.bottom + move_up = b.top - a.bottom + candidates = [ + ("horizontal", move_left), + ("horizontal", move_right), + ("vertical", move_down), + ("vertical", move_up), + ] + axis, penetration = min(candidates, key=lambda kv: kv[1]) + return { + "axis": axis, + "minimum_movement": max(0.0, penetration + min_gap), + } + + +def measure_title_box( + text: str, + font_size: float, + *, + x: float, + y: float, + font: Optional[str] = None, + face: Optional[str] = None, + kerning: float = 0.0, +) -> Box: + """The on-screen box of one title, measured in the emitted template space. + + ``font_size``/``kerning``/``x``/``y`` are the values the writer put into the + FCPXML, so the box is comparable across every title in the document without + any further scaling. Width comes from the real advance table, vertical + extent from the real ink (accents and descenders included); the anchor is + the title's centre. + """ + width = measure_text( + text, font_size, kerning=kerning, font=font, face=face + ) + top, bottom = ink_extent(text, font_size, font=font, face=face) + return Box( + left=x - width / 2, + right=x + width / 2, + bottom=y + bottom, + top=y + top, + ) + + +def _is_font_measured(font: Optional[str], face: Optional[str]) -> bool: + """Whether both advance and vertical metrics for ``font``/``face`` exist.""" + if metrics_for(font, face) is None: + return False + _, measured = vertical_metrics_for(font, face) + return measured + + +def _box_within(box: Box, limits: Box) -> bool: + return ( + box.left >= limits.left + and box.right <= limits.right + and box.bottom >= limits.bottom + and box.top <= limits.top + ) + + +def _issue(severity: str, type_: str, **fields) -> Dict: + return {"severity": severity, "type": type_, **fields} + + +def validate_titles( + titles: Sequence[Dict], + frame_width: float, + frame_height: float, + *, + safe_margin_x: float = 0.05, + safe_margin_y: float = 0.05, + min_font_size: Optional[float] = None, + min_distance: Optional[float] = None, + max_distance: Optional[float] = None, +) -> Dict: + """Validate a set of already-positioned titles and return a report. + + Each title dict must carry the emitted values: + + - ``text`` (str) + - ``font_size`` (float), ``kerning`` (float), ``x``/``y`` (floats) + - ``font`` (str) and ``face`` (str|None) + - ``start``/``end`` (seconds) for temporal overlap + - ``group`` (hashable) for spacing checks: titles sharing a group are one + block, expected to sit near each other (spec 14). Optional. + + Returns ``{"severity", "issues", "summary"}`` where ``severity`` is the + worst bucket seen and ``issues`` are the spec-16-shaped occurrences. + """ + frame_left = -frame_width / 2 + frame_right = frame_width / 2 + frame_bottom = -frame_height / 2 + frame_top = frame_height / 2 + frame_box = Box(frame_left, frame_right, frame_bottom, frame_top) + + safe_box = Box( + left=frame_left + safe_margin_x * frame_width, + right=frame_right - safe_margin_x * frame_width, + bottom=frame_bottom + safe_margin_y * frame_height, + top=frame_top - safe_margin_y * frame_height, + ) + + issues: List[Dict] = [] + boxes: List[Box] = [] + measured_flags: List[bool] = [] + for title in titles: + text = str(title.get("text", "") or "") + font = title.get("font") or None + face = title.get("face") or None + box = measure_title_box( + text, + float(title.get("font_size", 0.0)), + x=float(title.get("x", 0.0)), + y=float(title.get("y", 0.0)), + font=font, + face=face, + kerning=float(title.get("kerning", 0.0)), + ) + boxes.append(box) + measured_flags.append(_is_font_measured(font, face)) + + if not _is_font_measured(font, face): + issues.append( + _issue( + "warning", + FONT_MISSING, + title=text, + font=font, + face=face, + message=( + f"Font '{font or '?'}" + + (f" {face}" if face else "") + + "' has no embedded metrics; widths are estimated" + ), + ) + ) + if min_font_size is not None and float(title.get("font_size", 0.0)) < min_font_size: + issues.append( + _issue( + "warning", + FONT_TOO_SMALL, + title=text, + font_size=float(title.get("font_size", 0.0)), + minimum=min_font_size, + ) + ) + if not _box_within(box, frame_box): + issues.append( + _issue( + "error", + OUTSIDE_FRAME, + title=text, + left=box.left, + right=box.right, + bottom=box.bottom, + top=box.top, + ) + ) + elif not _box_within(box, safe_box): + issues.append( + _issue( + "warning", + OUTSIDE_SAFE_AREA, + title=text, + left=box.left, + right=box.right, + bottom=box.bottom, + top=box.top, + ) + ) + + # Spatial collisions between temporally overlapping titles. + for i in range(len(titles)): + for j in range(i + 1, len(titles)): + a, b = titles[i], titles[j] + if not temporal_overlap( + float(a.get("start", 0.0)), float(a.get("end", 0.0)), + float(b.get("start", 0.0)), float(b.get("end", 0.0)), + ): + continue + box_a, box_b = boxes[i], boxes[j] + if not box_a.overlaps(box_b): + continue + metrics = overlap_metrics(box_a, box_b) + severity = classify_overlap(metrics) + issues.append( + _issue( + severity, + SPATIAL_COLLISION, + first_title=str(a.get("text", "")), + second_title=str(b.get("text", "")), + time_start=float(a.get("start", 0.0)), + time_end=float(b.get("end", 0.0)), + overlap_width=metrics["overlap_width"], + overlap_height=metrics["overlap_height"], + overlap_area=metrics["overlap_area"], + overlap_ratio=metrics["overlap_ratio"], + suggested_correction=separation_suggestion(box_a, box_b), + ) + ) + + # Spacing within a block (spec 8/14). Only when the caller asked for it — + # a generic minimum can fire on the reference look's own tight stacking. + if min_distance is not None or max_distance is not None: + groups: Dict = {} + for index, title in enumerate(titles): + groups.setdefault(title.get("group", index), []).append(index) + for members in groups.values(): + for m in range(len(members)): + for n in range(m + 1, len(members)): + i, j = members[m], members[n] + box_a, box_b = boxes[i], boxes[j] + if box_a.overlaps(box_b): + continue + gap = distance_between(box_a, box_b)["distance"] + if min_distance is not None and gap < min_distance: + issues.append( + _issue( + "warning", + INSUFFICIENT_SPACING, + first_title=str(titles[i].get("text", "")), + second_title=str(titles[j].get("text", "")), + distance=gap, + minimum=min_distance, + ) + ) + if max_distance is not None and gap > max_distance: + issues.append( + _issue( + "warning", + EXCESSIVE_SPACING, + first_title=str(titles[i].get("text", "")), + second_title=str(titles[j].get("text", "")), + distance=gap, + maximum=max_distance, + ) + ) + + _rank = { + OVERLAP_NONE: 0, + OVERLAP_RENDER_TOLERANCE: 1, + OVERLAP_WARNING: 2, + OVERLAP_PROBABLE: 3, + OVERLAP_SEVERE: 4, + } + severities = [issue["severity"] for issue in issues] + worst = max(severities, key=lambda s: _rank.get(s, 0), default=OVERLAP_NONE) + + return { + "severity": worst, + "issues": issues, + "summary": { + "title_count": len(titles), + "issue_count": len(issues), + "spatial_collision": sum( + 1 for i in issues if i["type"] == SPATIAL_COLLISION + ), + "outside_frame": sum( + 1 for i in issues if i["type"] == OUTSIDE_FRAME + ), + "outside_safe_area": sum( + 1 for i in issues if i["type"] == OUTSIDE_SAFE_AREA + ), + "font_missing": sum( + 1 for i in issues if i["type"] == FONT_MISSING + ), + "font_too_small": sum( + 1 for i in issues if i["type"] == FONT_TOO_SMALL + ), + "insufficient_spacing": sum( + 1 for i in issues if i["type"] == INSUFFICIENT_SPACING + ), + "excessive_spacing": sum( + 1 for i in issues if i["type"] == EXCESSIVE_SPACING + ), + }, + } + + +def blocking(severity: str) -> bool: + """Whether a validation severity should block export (spec 16).""" + return severity in (OVERLAP_SEVERE, OVERLAP_PROBABLE) diff --git a/code/fcpxml/diarize.py b/code/fcpxml/diarize.py index ff060e9..8b49efe 100644 --- a/code/fcpxml/diarize.py +++ b/code/fcpxml/diarize.py @@ -37,6 +37,36 @@ def diarization_capability(token: Optional[str]) -> Tuple[bool, str]: return True, "Identificação de participantes disponível." +def _load_waveform(path: str) -> Optional[dict]: + """Decode ``path`` ourselves into the waveform dict pyannote accepts. + + pyannote 4.x decodes audio through torchcodec, which links against a + specific FFmpeg major version and fails outright when the installed one + differs (``libavutil.56.dylib`` not found) — taking diarization down on + an otherwise working machine. Handing it an already-decoded waveform + skips that path entirely and reuses the ffmpeg extraction the acoustic + analysis already relies on, so video containers work too. + + Returns ``None`` when decoding is not possible, letting the caller fall + back to passing the path and whatever pyannote can do with it. + """ + try: + import soundfile + import torch + + from .voice_features import decodable_audio + + with decodable_audio(path) as audio_path: + if audio_path is None: + return None + data, sample_rate = soundfile.read(audio_path, dtype="float32", always_2d=True) + # soundfile gives (samples, channels); pyannote wants (channels, samples) + return {"waveform": torch.from_numpy(data.T), "sample_rate": int(sample_rate)} + except Exception: + logger.info("could not pre-decode %s for diarization", path) + return None + + def diarize( path: str, token: Optional[str], @@ -67,7 +97,7 @@ def diarize( n = str(num_speakers or "").strip() if n.isdigit() and int(n) > 0: kwargs["num_speakers"] = int(n) - result = pipe(path, **kwargs) + result = pipe(_load_waveform(path) or path, **kwargs) # pyannote.audio >= 4.0 wraps the annotation; normalize to the raw one. if hasattr(result, "exclusive_speaker_diarization"): result = result.exclusive_speaker_diarization diff --git a/code/fcpxml/emphasis.py b/code/fcpxml/emphasis.py new file mode 100644 index 0000000..e24761e --- /dev/null +++ b/code/fcpxml/emphasis.py @@ -0,0 +1,133 @@ +"""Emphasis index — how much a spoken word "pops" acoustically. + +Pure functions over already-extracted per-word features (energy, pitch +delta, rate delta, pause before, duration); no I/O, no external dependency. +Combines them into a single ``[0, 1]`` score, configurable via +:class:`EmphasisWeights` so the weighting can be tuned (and persisted, +see ``model_manager.load_voice_analysis_config``) without touching code. +""" + +from dataclasses import dataclass +from typing import List, Sequence + +_FIELDS = ("energy", "pitch_variation", "rate_variation", "pause_before", "duration") + + +@dataclass +class EmphasisWeights: + energy: float = 0.30 + pitch_variation: float = 0.25 + rate_variation: float = 0.20 + pause_before: float = 0.15 + duration: float = 0.10 + + def as_dict(self) -> dict: + return {field: getattr(self, field) for field in _FIELDS} + + @classmethod + def from_dict(cls, data: dict) -> "EmphasisWeights": + defaults = cls() + return cls(**{field: float(data.get(field, getattr(defaults, field))) for field in _FIELDS}) + + +def _clamp01(x: float) -> float: + return max(0.0, min(1.0, x)) + + +def pause_weight( + pause_before: float, max_pause: float = 1.5, ignore_above: float = 3.0 +) -> float: + """How much a preceding silence counts as emphasis, in ``[0, 1]``. + + A short beat before a word is real emphasis: the speaker is setting it + up. A *long* gap is not — it is an edit point, a B-roll insert, or the + other person in the room talking. Measured on real footage, gaps of + 6-9s were scoring as the most emphatic moments in the recording purely + because the scale saturated, ranking a scene change above a word the + speaker actually hit hard. + + So the contribution rises up to ``max_pause`` and then drops to zero + past ``ignore_above``, instead of saturating. Set ``ignore_above`` to + ``0`` to disable the cutoff and keep the old saturating behaviour. + """ + if pause_before <= 0 or max_pause <= 0: + return 0.0 + if ignore_above > 0 and pause_before > ignore_above: + return 0.0 + return _clamp01(pause_before / max_pause) + + +def compute_emphasis( + energy: float, + pitch_delta: float, + rate_delta: float, + pause_before: float, + word_duration: float, + weights: EmphasisWeights = EmphasisWeights(), + *, + max_pause: float = 1.5, + max_duration: float = 1.0, + pause_ignore_above: float = 3.0, +) -> float: + """Emphasis score in ``[0, 1]`` for one word. + + ``energy``/``pitch_delta``/``rate_delta`` are expected already + normalized to roughly ``[0, 1]`` (deltas may be negative — only their + magnitude counts as emphasis). ``pause_before``/``word_duration`` are + raw seconds; duration saturates at ``max_duration``, while the pause + contribution is shaped by :func:`pause_weight`. + """ + energy_n = _clamp01(energy) + pitch_n = _clamp01(abs(pitch_delta)) + rate_n = _clamp01(abs(rate_delta)) + pause_n = pause_weight(pause_before, max_pause, pause_ignore_above) + duration_n = _clamp01(word_duration / max_duration) if max_duration > 0 else 0.0 + + total_weight = sum(getattr(weights, field) for field in _FIELDS) + if total_weight <= 0: + return 0.0 + + score = ( + weights.energy * energy_n + + weights.pitch_variation * pitch_n + + weights.rate_variation * rate_n + + weights.pause_before * pause_n + + weights.duration * duration_n + ) + return _clamp01(score / total_weight) + + +def annotate_emphasis( + words: Sequence[dict], + weights: EmphasisWeights = EmphasisWeights(), + *, + max_pause: float = 1.5, + max_duration: float = 1.0, + pause_ignore_above: float = 3.0, +) -> List[dict]: + """Return copies of ``words`` with an ``"emphasis"`` key added. + + Each word dict is expected to carry ``energy``, ``pitch_delta``, + ``rate_delta``, ``pause_before`` (all pre-computed, e.g. by + ``voice_features.py``), plus ``start``/``end`` — or an explicit + ``duration`` — to derive word length. + """ + out: List[dict] = [] + for w in words: + ww = dict(w) + duration = ww.get("duration") + if duration is None: + duration = max(0.0, float(ww.get("end", 0.0)) - float(ww.get("start", 0.0))) + ww["emphasis"] = compute_emphasis( + energy=float(ww.get("energy") or 0.0), + pitch_delta=float(ww.get("pitch_delta") or 0.0), + rate_delta=float(ww.get("rate_delta") or 0.0), + pause_before=float(ww.get("pause_before") or 0.0), + word_duration=float(duration), + weights=weights, + max_pause=max_pause, + max_duration=max_duration, + pause_ignore_above=pause_ignore_above, + ) + out.append(ww) + return out diff --git a/code/fcpxml/model_manager.py b/code/fcpxml/model_manager.py index 2e3b144..2f73468 100644 --- a/code/fcpxml/model_manager.py +++ b/code/fcpxml/model_manager.py @@ -359,3 +359,291 @@ def save_num_speakers(num: str) -> str: data["num_speakers"] = val _write_config(data) return val + + +DEFAULT_VOICE_ANALYSIS_CONFIG: dict = { + "energy_threshold": 0.5, + "emphasis_weights": { + "energy": 0.30, + "pitch_variation": 0.25, + "rate_variation": 0.20, + "pause_before": 0.15, + "duration": 0.10, + }, + # Peaks are selected RELATIVELY — the top slice of the distribution — + # because the emphasis index is a weighted average whose real range + # depends on the material. Measured on a 17-minute interview the index + # never passed 0.55, so any absolute cutoff near the spec's 0.85 selects + # nothing; on punchier material the same cutoff would flood the edit. + # 2% of words is roughly one highlight every 50 words. + "peak_percentile": 0.02, + # Guard for genuinely flat audio, where even the top of the distribution + # carries no emphasis worth cutting on. + "emphasis_floor": 0.25, + "emotion_enabled": False, + "emotion_sensitivity": 0.5, +} + + +def load_voice_analysis_config() -> dict: + """The persisted voice-analysis thresholds/weights, merged over defaults. + + Backs the "Análise de Voz" settings screen: energy threshold (how loud + counts as "high energy"), the emphasis-index weights (see + ``emphasis.EmphasisWeights``), the punch-in emphasis cutoff, and the + emotion-detection toggle/sensitivity. Unknown/malformed stored values + fall back to the default rather than raising, so a hand-edited or + partially-written config.json never breaks the settings screen. + """ + cfg = { + **DEFAULT_VOICE_ANALYSIS_CONFIG, + "emphasis_weights": dict(DEFAULT_VOICE_ANALYSIS_CONFIG["emphasis_weights"]), + } + stored = _load_config().get("voice_analysis") + if not isinstance(stored, dict): + return cfg + for key in ("energy_threshold", "peak_percentile", "emphasis_floor", "emotion_sensitivity"): + if key in stored: + try: + cfg[key] = max(0.0, min(1.0, float(stored[key]))) + except (TypeError, ValueError): + pass + if "emotion_enabled" in stored: + cfg["emotion_enabled"] = bool(stored["emotion_enabled"]) + weights = stored.get("emphasis_weights") + if isinstance(weights, dict): + for key in cfg["emphasis_weights"]: + if key in weights: + try: + cfg["emphasis_weights"][key] = max(0.0, float(weights[key])) + except (TypeError, ValueError): + pass + return cfg + + +def save_voice_analysis_config( + energy_threshold: float | None = None, + emphasis_weights: dict | None = None, + peak_percentile: float | None = None, + emphasis_floor: float | None = None, + emotion_enabled: bool | None = None, + emotion_sensitivity: float | None = None, +) -> dict: + """Persist voice-analysis thresholds/weights. Only given fields change. + + Returns the full merged config (same shape as + :func:`load_voice_analysis_config`) so callers can render it back + immediately without a second round-trip. + """ + cfg = load_voice_analysis_config() + if energy_threshold is not None: + cfg["energy_threshold"] = max(0.0, min(1.0, float(energy_threshold))) + if peak_percentile is not None: + cfg["peak_percentile"] = max(0.0, min(1.0, float(peak_percentile))) + if emphasis_floor is not None: + cfg["emphasis_floor"] = max(0.0, min(1.0, float(emphasis_floor))) + if emotion_enabled is not None: + cfg["emotion_enabled"] = bool(emotion_enabled) + if emotion_sensitivity is not None: + cfg["emotion_sensitivity"] = max(0.0, min(1.0, float(emotion_sensitivity))) + if emphasis_weights is not None: + for key, value in emphasis_weights.items(): + if key in cfg["emphasis_weights"] and value is not None: + cfg["emphasis_weights"][key] = max(0.0, float(value)) + data = _load_config() + data["voice_analysis"] = cfg + _write_config(data) + return cfg + + +# Mirrors the "Legendas Dinâmicas" tab's own defaults (MacApp/Sources/ +# CaptionsView.swift), so a fresh install shows the same look in the UI and +# in what generate_dynamic_subtitles renders when no override is passed. +DEFAULT_DYNAMIC_SUBTITLE_CONFIG: dict = { + "band_height": 0.22, + "block_center_y": -167.0, + "line_gap": 8.0, + "font": "Helvetica Neue", + "font_size": 104, + "emphasis_font": "Playfair Display", + "emphasis_face": "Medium Italic", + "emphasis_size": 265, + "active_color": "1 1 1 1", + "emphasis_color": "1 1 1 1", + "text_scale": 2.0, +} + + +def load_dynamic_subtitle_config() -> dict: + """The persisted dynamic-subtitle style, merged over defaults. + + Backs the "Legendas Dinâmicas" settings screen AND is the fallback + ``generate_dynamic_subtitles`` reads for any field the caller doesn't + explicitly override — so the style configured in the UI is what actually + renders, without the app having to thread every field through each call. + Unknown/malformed stored values fall back to the default, same as + :func:`load_voice_analysis_config`. + """ + cfg = dict(DEFAULT_DYNAMIC_SUBTITLE_CONFIG) + stored = _load_config().get("dynamic_subtitles") + if not isinstance(stored, dict): + return cfg + for key in ("band_height", "block_center_y", "line_gap", "text_scale"): + if key in stored: + try: + cfg[key] = float(stored[key]) + except (TypeError, ValueError): + pass + for key in ("font_size", "emphasis_size"): + if key in stored: + try: + cfg[key] = int(stored[key]) + except (TypeError, ValueError): + pass + for key in ("font", "emphasis_font", "emphasis_face", "active_color", "emphasis_color"): + if key in stored and isinstance(stored[key], str) and stored[key]: + cfg[key] = stored[key] + return cfg + + +def save_dynamic_subtitle_config(**fields) -> dict: + """Persist dynamic-subtitle style fields. Only given fields change. + + Accepts the same keys as :data:`DEFAULT_DYNAMIC_SUBTITLE_CONFIG`; unknown + keys are ignored so a newer app talking to an older config shape degrades + quietly. Returns the full merged config, mirroring + :func:`save_voice_analysis_config`. + """ + cfg = load_dynamic_subtitle_config() + for key, value in fields.items(): + if key not in DEFAULT_DYNAMIC_SUBTITLE_CONFIG or value is None: + continue + if isinstance(DEFAULT_DYNAMIC_SUBTITLE_CONFIG[key], float): + try: + cfg[key] = float(value) + except (TypeError, ValueError): + continue + elif isinstance(DEFAULT_DYNAMIC_SUBTITLE_CONFIG[key], int): + try: + cfg[key] = int(value) + except (TypeError, ValueError): + continue + else: + cfg[key] = str(value) + data = _load_config() + data["dynamic_subtitles"] = cfg + _write_config(data) + return cfg + + +# Mirrors the silence thresholds the detection/removal handlers use when no +# argument is passed (server_tools/qc.py). Persisted so the app's slider and +# any later run agree without threading three fields through every call. +DEFAULT_SILENCE_CONFIG: dict = { + # dBFS below which audio counts as silence. + "noise_db": -30.0, + # Seconds a quiet stretch must last before it's a cut candidate. + "min_silence": 0.5, + # Seconds left inside each cut so speech never gets clipped at the edges. + "padding": 0.05, +} + + +def load_silence_config() -> dict: + """The persisted silence-detection thresholds, merged over defaults. + + Read by ``detect_media_silence``/``remove_media_silence`` as their + fallback, so the tolerance chosen in the app is what actually runs. + Malformed stored values fall back to the default rather than raising, + matching :func:`load_voice_analysis_config`. + """ + cfg = dict(DEFAULT_SILENCE_CONFIG) + stored = _load_config().get("silence") + if not isinstance(stored, dict): + return cfg + for key in cfg: + if key in stored: + try: + cfg[key] = float(stored[key]) + except (TypeError, ValueError): + pass + return cfg + + +def save_silence_config( + noise_db: float | None = None, + min_silence: float | None = None, + padding: float | None = None, +) -> dict: + """Persist silence thresholds. Only the given fields change. + + Values are clamped to the same ranges the handlers validate against, so + a bad write here can't produce a config the tools would later reject. + """ + cfg = load_silence_config() + if noise_db is not None: + try: + cfg["noise_db"] = max(-120.0, min(0.0, float(noise_db))) + except (TypeError, ValueError): + pass + if min_silence is not None: + try: + cfg["min_silence"] = max(0.01, min(3600.0, float(min_silence))) + except (TypeError, ValueError): + pass + if padding is not None: + try: + cfg["padding"] = max(0.0, min(5.0, float(padding))) + except (TypeError, ValueError): + pass + data = _load_config() + data["silence"] = cfg + _write_config(data) + return cfg + + +# Last project worked on, so the app reopens where the user left off instead of +# making them pick the folder again every launch. Only paths that still exist +# are handed back — a project on an unmounted volume degrades to "none selected" +# rather than to a dead path the tools would later fail on. +DEFAULT_PROJECT_CONFIG: dict = { + # Folder every generated file (transcript .json, XML, SRT) is written to. + "folder": "", + # The .fcpxml/.fcpxmld that was loaded from it. + "file": "", +} + + +def load_project_config() -> dict: + """The persisted last project (folder + file), merged over defaults. + + Paths that no longer exist on disk come back empty, matching what the app + shows for "nothing selected". Malformed stored values fall back to the + default rather than raising, same as :func:`load_voice_analysis_config`. + """ + cfg = dict(DEFAULT_PROJECT_CONFIG) + stored = _load_config().get("project") + if not isinstance(stored, dict): + return cfg + for key in cfg: + value = stored.get(key) + if isinstance(value, str) and value and Path(value).exists(): + cfg[key] = value + return cfg + + +def save_project_config(folder: str | None = None, file: str | None = None) -> dict: + """Persist the last project folder/file. Only the given fields change. + + Passing an empty string clears a field (the app does this when the user + deselects), while ``None`` leaves it untouched. + """ + cfg = load_project_config() + for key, value in (("folder", folder), ("file", file)): + if value is None: + continue + cfg[key] = str(Path(value).expanduser()) if str(value).strip() else "" + data = _load_config() + data["project"] = cfg + _write_config(data) + return cfg diff --git a/code/fcpxml/models.py b/code/fcpxml/models.py index e7623f7..f38371a 100755 --- a/code/fcpxml/models.py +++ b/code/fcpxml/models.py @@ -13,6 +13,8 @@ from functools import total_ordering from math import gcd from typing import Any, Callable, Dict, List, Optional, Tuple +from .text_layout import REFERENCE_BLOCK_LINE_GAP, TEXT_TEMPLATE_FONT_SCALE + # ============================================================================ # ENUMS # ============================================================================ @@ -1071,3 +1073,19 @@ class DynamicSubtitleConfig: # grouped, the key word alone and large (the reference look). "word": one # title per word, the earlier rhythm. granularity: str = "phrase" + # Ratio between the template's fontSize space and the canvas-point space + # its Position uses. See text_layout.TEXT_TEMPLATE_FONT_SCALE: the "Text" + # (Text.moti) template sizes type in frame pixels, so a size chosen in + # points renders half as large unless it is converted on the way out. + text_scale: float = TEXT_TEMPLATE_FONT_SCALE + # Vertical air between stacked lines, in canvas points. Negative values + # deliberately overlap the lines — the display italic tucking under the + # line above is a real editorial look, and the stacking arithmetic places + # ink boxes edge to edge, so a negative gap moves them by exactly that + # much rather than colliding unpredictably. + line_gap: float = REFERENCE_BLOCK_LINE_GAP + # Run the post-generation collision validation (collision.validate_titles) + # and refuse to emit when it reports a blocking overlap. Off by default so + # generation stays byte-identical to before this flag existed; flip it on + # for a guaranteed no-collision export. + validate: bool = False diff --git a/code/fcpxml/text_layout.py b/code/fcpxml/text_layout.py index 4cc2d9d..75544e3 100644 --- a/code/fcpxml/text_layout.py +++ b/code/fcpxml/text_layout.py @@ -122,6 +122,33 @@ REFERENCE_BLOCK_LINE_GAP = 8.0 # emphasis line's edge; the reference leaves a little air. REFERENCE_STAGGER_RATIO = 0.8 +# Extra gap, as a fraction of the emphasis line's font size, added only to +# the boundary right below it. The display italic's slant leans its stems +# past the vertical ink box the metrics measure, so a body line directly +# under the emphasis line reads tighter than the same nominal gap anywhere +# else in the stack — this cushion (~14pt at the 230pt reference size) +# closes that optical gap without touching the user's `line_gap` elsewhere. +_EMPHASIS_ITALIC_CUSHION_RATIO = 0.06 + +# The numbers above were read off a hand export that used the "Essencial - +# Título" template. That template never rendered when we generated it (see +# Engine/docs/05_EXPERIENCIAS.md, 2026-08-17), so the writer switched to FCP's +# own "Basic Text > Text" (Text.moti) — whose coordinate space is the FRAME +# ITSELF (2160x3840), not the half-scale point canvas the numbers above were +# measured in. Everything the template reads is in that space: fontSize, +# kerning AND Position alike. +# +# Getting this half-right is worse than getting it wrong. Scaling only the type +# left the block at the old spread with twice the type in it, so the lines +# collided; scaling only the positions would spread a block of half-size type +# across the frame. The layout keeps measuring in canvas points — every +# constant above depends on that — and this single factor converts the whole +# result on the way out, which is the only way the two stay in step. +# +# Exposed as `text_scale` on DynamicSubtitleConfig for a template authored +# against a different space. +TEXT_TEMPLATE_FONT_SCALE = 2.0 + def metrics_for(font: Optional[str], face: Optional[str] = None) -> Optional[Dict]: """The embedded advance table for *font*/*face*, or None if uncovered. @@ -296,9 +323,14 @@ class PlacedWord: and other.bottom < self.top ) - def position_param(self) -> str: - """The value for the title's "Posição" param, as FCP writes it.""" - return f"{self.x:g} {self.y:g}" + def position_param(self, scale: float = 1.0) -> str: + """The value for the title's "Posição" param, as FCP writes it. + + *scale* converts from canvas points to the template's own space; see + TEXT_TEMPLATE_FONT_SCALE. It must be the same factor the emitted + fontSize uses, or the type and the spacing drift apart. + """ + return f"{self.x * scale:g} {self.y * scale:g}" # The rest of this block is the interface a placed unit shares with # PlacedBlock, so the writer emits titles from either without caring @@ -649,9 +681,14 @@ class PlacedBlock: """When this block finishes being spoken, in seconds.""" return max(float(w.get('end', 0.0)) for w in self.words) - def position_param(self) -> str: - """The value for the title's "Posição" param, as FCP writes it.""" - return f"{self.x:g} {self.y:g}" + def position_param(self, scale: float = 1.0) -> str: + """The value for the title's "Posição" param, as FCP writes it. + + *scale* converts from canvas points to the template's own space; see + TEXT_TEMPLATE_FONT_SCALE. It must be the same factor the emitted + fontSize uses, or the type and the spacing drift apart. + """ + return f"{self.x * scale:g} {self.y * scale:g}" def overlaps(self, other: 'PlacedBlock') -> bool: return ( @@ -742,10 +779,33 @@ def compose_sentence( for run in body_lines(entries[emphasis_index + 1:]): lines.append((run, body_look, False)) + # A body run wraps onto a new line when it doesn't fit — but the + # emphasis line is always exactly one word, so it can't wrap, and + # nothing capped its size against the box. A long or all-caps word (an + # emphasis pass sometimes upper-cases its pick) could run past both + # edges of the frame — found on real footage, wide enough to spill off + # BOTH sides while centred. Shrinking it back to the box scales its + # font_size and kerning by the same factor, so the ink height used for + # stacking below shrinks with it too — restoring the vertical + # non-overlap the rest of this function already guarantees by + # construction. Never shrunk below the body size: emphasis smaller + # than body text isn't emphasis anymore, it's just a different font. + def fit_emphasis(text: str, look) -> tuple: + size, kerning, width = measure(text, look) + if width <= box.width: + return size, kerning, width + floor = float(body_look.font_size) * scale + fit = max(box.width / width, floor / size) if size > 0 else 1.0 + fit = min(fit, 1.0) + return size * fit, kerning * fit, width * fit + measured = [] for run, look, is_emphasis in lines: text = ' '.join(t for _, t in run) - size, kerning, width = measure(text, look) + if is_emphasis: + size, kerning, width = fit_emphasis(text, look) + else: + size, kerning, width = measure(text, look) # Stack on the real ink each line contains, not on a nominal # cap-height: the display italic's accents and descenders run well # past it, and a nominal box lets them collide with the neighbour. @@ -764,14 +824,29 @@ def compose_sentence( # loses the whole point of the look — so if it does not fit, everything # from the emphasis on overflows together. gap = line_gap * scale + + # The emphasis line's italic slant carries visual weight below its own + # ink box — Playfair's stems lean past what the vertical metrics measure + # — so a body line sitting right under it reads tighter than the same + # nominal gap elsewhere, even though the ink boxes themselves never + # touch. Add a size-proportional cushion only to the boundary right + # after the emphasis line; every other pair keeps exactly the caller's + # ``line_gap``. + def pair_gap(prev_line: dict) -> float: + if prev_line['emphasis']: + return gap + _EMPHASIS_ITALIC_CUSHION_RATIO * prev_line['font_size'] + return gap + kept = 0 total = 0.0 + prev = None for line in measured: - advance = line['height'] if not kept else line['height'] + gap - if kept and total + advance > box.height: + advance = line['height'] if prev is None else line['height'] + pair_gap(prev) + if prev is not None and total + advance > box.height: break total += advance kept += 1 + prev = line kept = max(kept, 1) if not any(line['emphasis'] for line in measured[:kept]): kept = min(kept, next( @@ -783,12 +858,12 @@ def compose_sentence( result.overflow.extend(w for w, _ in line['run']) visible = measured[:kept] - # Stack the ink boxes edge to edge with exactly *gap* between them, then - # centre the whole stack on the band. Because the boxes are the real ink, - # "no overlap" is a property of the arithmetic, not of a safety factor. - stack_height = ( - sum(line['height'] for line in visible) + gap * (len(visible) - 1) - ) + # Stack the ink boxes edge to edge with exactly *gap* between them (plus + # the emphasis cushion where it applies), then centre the whole stack on + # the band. Because the boxes are the real ink, "no overlap" is a + # property of the arithmetic, not of a safety factor. + gaps = [pair_gap(visible[i - 1]) for i in range(1, len(visible))] + stack_height = sum(line['height'] for line in visible) + sum(gaps) edge = box.center_y + stack_height / 2 # Body lines hang off the emphasis line's edges, alternating sides in @@ -797,7 +872,7 @@ def compose_sentence( side = -1 for index, line in enumerate(visible): if index: - edge -= gap + edge -= gaps[index - 1] cursor_y = edge - line['ink_top'] edge = cursor_y + line['ink_bottom'] if line['emphasis']: diff --git a/code/fcpxml/transcribe.py b/code/fcpxml/transcribe.py index b9e15a8..a14ba6b 100755 --- a/code/fcpxml/transcribe.py +++ b/code/fcpxml/transcribe.py @@ -16,7 +16,7 @@ import logging import os import re from pathlib import Path -from typing import List, Optional, Sequence, Tuple +from typing import Callable, List, Optional, Sequence, Tuple logger = logging.getLogger(__name__) @@ -118,7 +118,10 @@ def invert_ranges( def transcribe( - path: str, model_size: str = "base", language: Optional[str] = None + path: str, + model_size: str = "base", + language: Optional[str] = None, + progress_cb: Optional[Callable[[float], None]] = None, ) -> Optional[dict]: """Transcribe an audio/video file locally with word-level timestamps. @@ -170,6 +173,10 @@ def transcribe( ) segments: List[dict] = [] words: List[dict] = [] + # `info.duration` is known upfront (from the container), so each + # segment's end time — yielded lazily as faster-whisper decodes — + # gives real, granular progress instead of a single before/after step. + total_duration = float(info.duration) if info.duration else 0.0 for seg in segments_iter: start = float(seg.start) end = float(seg.end) @@ -182,6 +189,8 @@ def transcribe( "end_fmt": format_timestamp(end), } ) + if progress_cb is not None and total_duration > 0: + progress_cb(min(end / total_duration, 1.0)) for w in seg.words or []: ws = float(w.start) we = float(w.end) diff --git a/code/fcpxml/voice_actions.py b/code/fcpxml/voice_actions.py new file mode 100644 index 0000000..42c5bdc --- /dev/null +++ b/code/fcpxml/voice_actions.py @@ -0,0 +1,248 @@ +"""Voice actions — the editing decisions produced from a voice timeline. + +This is the contract between *deciding* and *applying*. Whoever makes the +editorial call — the deterministic rules engine, or a model reading the +voice timeline JSON — emits the same list of actions, and one applier turns +it into FCPXML. Nothing that produces actions ever touches XML. + +Every action's ``start``/``end`` is in **original source seconds**, matching +the voice timeline. That matters: cuts shift everything after them, so if +decisions were expressed in post-cut time they would silently land in the +wrong place the moment a cut was added. Keeping one origin and resolving the +shift at apply time (:func:`shift_after_cuts`) removes that whole class of bug. + +Actions arriving from a model are untrusted input: :func:`parse_actions` +validates and reports what it rejected rather than raising, so one malformed +row never discards a whole edit. +""" + +from dataclasses import dataclass, field +from typing import Any, List, Optional, Sequence, Tuple + +# What an action can ask for. Deliberately small — each maps onto one +# existing writer capability, so no new XML knowledge lives here. +ACTION_KINDS = ("cut", "zoom", "text", "marker") + +# Bounds for a zoom's scale factor. Below 1.0 is a pull-back, not a punch-in; +# above 3x the image falls apart on any normal footage. +MIN_ZOOM_SCALE = 1.0 +MAX_ZOOM_SCALE = 3.0 + +MAX_TEXT_LENGTH = 120 + + +@dataclass +class VoiceAction: + """One editing decision, in original source time.""" + + kind: str + start: float + end: float + params: dict = field(default_factory=dict) + reason: str = "" + speaker: str = "" + + @property + def duration(self) -> float: + return max(0.0, self.end - self.start) + + def as_dict(self) -> dict: + return { + "kind": self.kind, + "start": round(self.start, 3), + "end": round(self.end, 3), + "params": self.params, + "reason": self.reason, + "speaker": self.speaker, + } + + +def _validate_one(raw: Any, index: int) -> Tuple[Optional[VoiceAction], str]: + """Turn one raw row into a VoiceAction, or explain why it can't be.""" + where = f"action[{index}]" + if not isinstance(raw, dict): + return None, f"{where}: expected an object, got {type(raw).__name__}" + + kind = str(raw.get("kind", "")).strip().lower() + if kind not in ACTION_KINDS: + return None, f"{where}: unknown kind {raw.get('kind')!r} (expected one of {', '.join(ACTION_KINDS)})" + + try: + start = float(raw.get("start")) + end = float(raw.get("end")) + except (TypeError, ValueError): + return None, f"{where}: start/end must be numbers (seconds)" + + if start < 0: + return None, f"{where}: start is negative ({start})" + if end <= start: + return None, f"{where}: end ({end}) must be after start ({start})" + + params = raw.get("params") + params = dict(params) if isinstance(params, dict) else {} + + if kind == "zoom": + try: + scale = float(params.get("scale", 1.3)) + except (TypeError, ValueError): + return None, f"{where}: zoom scale must be a number" + if not (MIN_ZOOM_SCALE <= scale <= MAX_ZOOM_SCALE): + return None, ( + f"{where}: zoom scale {scale} outside {MIN_ZOOM_SCALE}-{MAX_ZOOM_SCALE}" + ) + params["scale"] = scale + + if kind == "text": + content = str(params.get("content", "")).strip() + if not content: + return None, f"{where}: text action needs params.content" + params["content"] = content[:MAX_TEXT_LENGTH] + + return ( + VoiceAction( + kind=kind, + start=start, + end=end, + params=params, + reason=str(raw.get("reason", "")), + speaker=str(raw.get("speaker", "")), + ), + "", + ) + + +def parse_actions(data: Any) -> Tuple[List[VoiceAction], List[str]]: + """Validate a decision list into actions, collecting rejections. + + Accepts either a bare list of actions or ``{"actions": [...]}`` — the + shape a model is most likely to return. Returns ``(actions, errors)``; + a row that fails validation is reported and skipped, never fatal. + """ + if isinstance(data, dict): + data = data.get("actions", []) + if not isinstance(data, Sequence) or isinstance(data, (str, bytes)): + return [], ["expected a list of actions, or an object with an 'actions' list"] + + actions: List[VoiceAction] = [] + errors: List[str] = [] + for i, raw in enumerate(data): + action, error = _validate_one(raw, i) + if action is not None: + actions.append(action) + else: + errors.append(error) + return actions, errors + + +def speaker_cut_actions( + timeline: dict, + speaker_ids: Sequence[str], + padding: float = 0.15, +) -> List[VoiceAction]: + """Cut actions removing everything the given speakers say. + + The everyday case on a testimonial shoot: an interviewer or a crew + member talks over the take, and only the subject should survive the + edit. ``padding`` trims slightly *inside* each segment rather than + around it — speech boundaries from a transcript are approximate, and + eating into the neighbouring silence is far safer than clipping the + first syllable of the person being kept. + """ + wanted = {str(s) for s in speaker_ids} + actions: List[VoiceAction] = [] + for segment in timeline.get("segments", []): + if str(segment.get("speaker", "")) not in wanted: + continue + start = float(segment.get("start", 0.0)) + padding + end = float(segment.get("end", 0.0)) - padding + if end <= start: + continue + actions.append( + VoiceAction( + kind="cut", + start=start, + end=end, + reason=f"fala de {segment.get('speaker')}", + speaker=str(segment.get("speaker", "")), + ) + ) + return actions + + +def merge_cut_ranges(actions: Sequence[VoiceAction]) -> List[Tuple[float, float]]: + """The cut actions as merged, sorted, non-overlapping source ranges.""" + cuts = sorted((a.start, a.end) for a in actions if a.kind == "cut") + merged: List[Tuple[float, float]] = [] + for start, end in cuts: + if merged and start <= merged[-1][1]: + merged[-1] = (merged[-1][0], max(merged[-1][1], end)) + else: + merged.append((start, end)) + return merged + + +def shift_after_cuts( + time: float, cuts: Sequence[Tuple[float, float]] +) -> Optional[float]: + """Where source ``time`` lands once ``cuts`` are removed. + + Returns ``None`` when the time falls *inside* a cut — the material it + referred to no longer exists, so the action that pointed at it must be + dropped rather than silently slid onto neighbouring content. + ``cuts`` must be merged and sorted (see :func:`merge_cut_ranges`). + """ + shift = 0.0 + for start, end in cuts: + if time < start: + break + if time < end: + return None + shift += end - start + return time - shift + + +def resolve_actions( + actions: Sequence[VoiceAction], +) -> Tuple[List[Tuple[float, float]], List[VoiceAction], List[VoiceAction]]: + """Split a decision list into what to cut and what to place afterwards. + + Returns ``(cut_ranges, placed, dropped)``. Non-cut actions are moved onto + their post-cut times; any that pointed into removed material land in + ``dropped`` so the caller can report them instead of losing them quietly. + """ + cut_ranges = merge_cut_ranges(actions) + placed: List[VoiceAction] = [] + dropped: List[VoiceAction] = [] + + for action in actions: + if action.kind == "cut": + continue + new_start = shift_after_cuts(action.start, cut_ranges) + if new_start is None: + dropped.append(action) + continue + if action.kind == "marker": + # A marker is a point, not a span: it survives as long as its own + # instant does. Requiring its nominal end to survive too would + # drop exactly the markers worth keeping — the ones flagging a + # join, which sit right against a cut edge by definition. + new_end = new_start + action.duration + else: + new_end = shift_after_cuts(action.end, cut_ranges) + if new_end is None: + dropped.append(action) + continue + if new_end <= new_start: + dropped.append(action) + continue + placed.append( + VoiceAction( + kind=action.kind, + start=new_start, + end=new_end, + params=action.params, + reason=action.reason, + speaker=action.speaker, + ) + ) + return cut_ranges, placed, dropped diff --git a/code/fcpxml/voice_features.py b/code/fcpxml/voice_features.py new file mode 100644 index 0000000..d6901ca --- /dev/null +++ b/code/fcpxml/voice_features.py @@ -0,0 +1,220 @@ +"""Acoustic features for voice analysis — pitch, energy, rate, pauses. + +Mirrors the ``media_intel.py`` contract: librosa is an optional dependency +(``pip install 'fcp-mcp-server[intelligence]'``, already required by beat +detection), imported lazily, and every extractor degrades to ``None`` when +the library is missing or the file cannot be analyzed — never crashes. + +``compute_speech_rate``/``compute_pauses`` are pure functions over +word-timestamp dicts (the shape ``transcribe.py`` already produces) and need +no audio file at all. +""" + +import contextlib +import logging +import shutil +import subprocess +import tempfile +from pathlib import Path +from typing import Iterator, List, Optional, Sequence, Tuple + +logger = logging.getLogger(__name__) + +# Human voice fundamental frequency range (covers low male to high female/child). +PITCH_FMIN_HZ = 65.0 +PITCH_FMAX_HZ = 1000.0 + +# Formats librosa reads directly through soundfile. Anything else — notably +# the .mov/.mp4 that source footage actually arrives in — must be decoded by +# ffmpeg first, or analysis fails outright. +NATIVE_AUDIO_SUFFIXES = {".wav", ".aif", ".aiff", ".flac"} + +# Voice analysis only needs the speech band: 16 kHz mono is well above the +# Nyquist limit for our 1 kHz pitch ceiling, and keeps the extracted file +# small and fast to decode even for hour-long footage. +EXTRACT_SAMPLE_RATE = 16000 +EXTRACT_TIMEOUT_SECONDS = 600 + + +@contextlib.contextmanager +def decodable_audio(path: str) -> Iterator[Optional[str]]: + """Yield a path librosa can read, extracting the audio track if needed. + + Audio files pass straight through. Video containers are decoded to a + temporary mono WAV with ffmpeg and cleaned up on exit. Yields ``None`` + when the audio cannot be obtained (no ffmpeg, no audio track, failure), + keeping the graceful-degradation contract of this module. + """ + file_path = Path(path) + if file_path.suffix.lower() in NATIVE_AUDIO_SUFFIXES: + yield str(file_path) + return + + if shutil.which("ffmpeg") is None: + logger.info("ffmpeg not found on PATH; cannot extract audio from %s", file_path) + yield None + return + + tmp_dir = tempfile.mkdtemp(prefix="fcp_voice_") + wav_path = Path(tmp_dir) / "audio.wav" + try: + result = subprocess.run( + [ + "ffmpeg", "-hide_banner", "-nostdin", "-y", + "-i", str(file_path), + "-vn", # audio only: decoding video would dominate the runtime + "-ac", "1", + "-ar", str(EXTRACT_SAMPLE_RATE), + str(wav_path), + ], + capture_output=True, + text=True, + timeout=EXTRACT_TIMEOUT_SECONDS, + ) + if result.returncode != 0 or not wav_path.is_file(): + logger.warning("ffmpeg could not extract audio from %s", file_path) + yield None + else: + yield str(wav_path) + except (OSError, subprocess.TimeoutExpired): + logger.warning("audio extraction failed for %s", file_path) + yield None + finally: + shutil.rmtree(tmp_dir, ignore_errors=True) + + +def features_capability() -> Tuple[bool, str]: + """Whether pitch/energy extraction is available (librosa installed).""" + try: + import librosa # noqa: F401 + except Exception: + return False, "Análise acústica indisponível: componente librosa ausente." + return True, "Análise acústica disponível." + + +def extract_pitch( + path: str, hop_length: int = 512, max_analysis_seconds: float = 1200.0 +) -> Optional[List[Tuple[float, float]]]: + """Frame-level pitch (F0) track via librosa's ``pyin``. + + Returns ``[(time_seconds, hz), ...]`` for voiced frames only (unvoiced + frames, where ``pyin`` reports no pitch, are dropped), or ``None`` when + librosa is unavailable or the file cannot be analyzed. + """ + file_path = Path(path) + if not file_path.is_file(): + return None + try: + import librosa + except ImportError: + logger.info("librosa not installed; pitch extraction unavailable") + return None + try: + with decodable_audio(str(file_path)) as audio_path: + if audio_path is None: + return None + y, sr = librosa.load(audio_path, sr=None, mono=True, duration=max_analysis_seconds) + f0, voiced_flag, _voiced_prob = librosa.pyin( + y, fmin=PITCH_FMIN_HZ, fmax=PITCH_FMAX_HZ, sr=sr, hop_length=hop_length + ) + times = librosa.times_like(f0, sr=sr, hop_length=hop_length) + except Exception: + logger.warning("librosa pitch analysis failed for %s", file_path) + return None + return [ + (float(t), float(hz)) + for t, hz, voiced in zip(times, f0, voiced_flag) + if voiced and hz == hz # ``hz == hz`` filters NaN without importing math/numpy here + ] + + +def extract_energy( + path: str, hop_length: int = 512, max_analysis_seconds: float = 1200.0 +) -> Optional[List[Tuple[float, float]]]: + """Frame-level RMS energy track via librosa. + + Returns ``[(time_seconds, rms), ...]``, or ``None`` when librosa is + unavailable or the file cannot be analyzed. + """ + file_path = Path(path) + if not file_path.is_file(): + return None + try: + import librosa + except ImportError: + logger.info("librosa not installed; energy extraction unavailable") + return None + try: + with decodable_audio(str(file_path)) as audio_path: + if audio_path is None: + return None + y, sr = librosa.load(audio_path, sr=None, mono=True, duration=max_analysis_seconds) + rms = librosa.feature.rms(y=y, hop_length=hop_length)[0] + times = librosa.times_like(rms, sr=sr, hop_length=hop_length) + except Exception: + logger.warning("librosa energy analysis failed for %s", file_path) + return None + return [(float(t), float(r)) for t, r in zip(times, rms)] + + +def _window_average(track: Sequence[Tuple[float, float]], start: float, end: float) -> Optional[float]: + """Average of ``track`` values whose timestamp falls in ``[start, end]``.""" + values = [v for t, v in track if start <= t <= end] + if not values: + return None + return sum(values) / len(values) + + +def word_pitch_energy( + words: Sequence[dict], + pitch_track: Optional[Sequence[Tuple[float, float]]], + energy_track: Optional[Sequence[Tuple[float, float]]], +) -> List[dict]: + """Attach average pitch/energy over each word's ``[start, end]`` span. + + Words carry ``pitch_hz``/``energy`` (``None`` when the span has no + voiced frames or a track is unavailable). Both tracks are the output of + :func:`extract_pitch`/:func:`extract_energy`. + """ + out: List[dict] = [] + for w in words: + ww = dict(w) + start = float(w.get("start", 0.0)) + end = float(w.get("end", start)) + ww["pitch_hz"] = _window_average(pitch_track, start, end) if pitch_track else None + ww["energy"] = _window_average(energy_track, start, end) if energy_track else None + out.append(ww) + return out + + +def compute_speech_rate(words: Sequence[dict], window_seconds: float = 3.0) -> List[float]: + """Local speech rate (words/second) around each word. + + For word *i*, counts every word whose start falls within + ``[start_i - window_seconds, start_i]`` and divides by + ``window_seconds`` — a trailing local rate, cheap to compute and stable + against a single long/short word skewing the whole utterance's average. + """ + starts = [float(w.get("start", 0.0)) for w in words] + rates: List[float] = [] + for i, s in enumerate(starts): + lo = s - window_seconds + count = sum(1 for t in starts[: i + 1] if t >= lo) + rates.append(count / window_seconds if window_seconds > 0 else 0.0) + return rates + + +def compute_pauses(words: Sequence[dict]) -> List[float]: + """Silence (seconds) immediately before each word. + + The first word's "pause before" is the time from the start of the audio + to its own start; every other word measures the gap since the previous + word's end (clamped to ``0`` for overlapping/adjacent words). + """ + pauses: List[float] = [] + prev_end = 0.0 + for w in words: + start = float(w.get("start", 0.0)) + pauses.append(max(0.0, start - prev_end)) + prev_end = float(w.get("end", start)) + return pauses diff --git a/code/fcpxml/voice_timeline.py b/code/fcpxml/voice_timeline.py new file mode 100644 index 0000000..6cd3a81 --- /dev/null +++ b/code/fcpxml/voice_timeline.py @@ -0,0 +1,505 @@ +"""Voice timeline — the consolidated, AI-readable view of how a video is spoken. + +This is the *source of truth* between analysis and editing: it merges what +was said (transcript), who said it (diarization), and how it was said +(pitch/energy/rate/pauses → emphasis) into one JSON document, decoupling the +audio analysis from FCPXML generation entirely. + +The shape is designed to be handed to a language model so it can reason about +the narrative — which beats carry weight, where a speaker changes, where the +delivery peaks — and decide how to direct the edit. Two design choices serve +that goal: + +* **Layered, not flat.** A ``summary`` gives the whole picture in a few + numbers, ``segments`` group words into utterances with their own + aggregates, and ``words`` hold the fine detail. A model can reason from + the top layer and only descend where it matters, instead of parsing + thousands of word rows to find the shape of the piece. +* **Normalized, self-describing values.** Every acoustic value is 0–1 and + relative to *this* recording (a quiet podcast and a shouted ad both use + the full range), and ``scales`` documents that contract inline, so the + numbers are interpretable without external context. +""" + +import json +import logging +from pathlib import Path +from typing import Callable, List, Optional, Sequence, Tuple + +from .diarize import DEFAULT_SPEAKER, assign_speakers, build_speakers, diarize +from .emphasis import EmphasisWeights, annotate_emphasis +from .voice_features import ( + compute_pauses, + compute_speech_rate, + extract_energy, + extract_pitch, + word_pitch_energy, +) + +logger = logging.getLogger(__name__) + +VOICE_TIMELINE_VERSION = "1.0" + +# Silence long enough to mean the take stopped rather than the speaker paused. +# On real footage, boundaries between retakes showed gaps of 3.6-19.8s while +# dramatic beats inside a delivered line stayed under ~2s. +TAKE_BOUNDARY_GAP = 3.0 + +# How to read the values in this document, split by the level they live on. +# Embedded in the output so a model consuming the JSON needs no external +# documentation — and kept honest: a metric listed under "word" must exist on +# every word row, and one under "segment" on every segment row. +VALUE_SCALES = { + "word": { + "energy": "0-1, loudness relative to the loudest moment of this recording", + "pitch_delta": "0-1, how far this word's pitch sits from the speaker's average", + "rate_delta": "0-1, how much the local speaking rate departs from the average", + "pause_before": "seconds of silence immediately before the word", + "emphasis": "0-1 combined index; high values are punch-in/highlight candidates", + }, + "segment": { + "gap_before": "seconds of silence before this line", + "take_boundary": "true when the gap is long enough that the take likely restarted here", + "avg_energy": "0-1 mean loudness across the line", + "peak_emphasis": "0-1 highest emphasis of any word in the line", + }, +} + + +def _normalize(value: Optional[float], maximum: float) -> float: + """Scale ``value`` into 0-1 against ``maximum`` (0.0 when unavailable).""" + if value is None or maximum <= 0: + return 0.0 + return max(0.0, min(1.0, value / maximum)) + + +def _round_word(word: dict) -> dict: + """One word row, rounded to a size a model can read without noise. + + The raw ``energy_raw``/``pitch_hz`` ride along beside the normalized + values so the document can be re-analyzed over a subset later. That + matters after cutting: every normalized value is relative to the + loudest moment of the *whole* recording, and if that moment gets cut + the survivors are scored against something that no longer exists. + """ + return { + "text": word.get("word", ""), + "start": round(float(word.get("start", 0.0)), 3), + "end": round(float(word.get("end", 0.0)), 3), + "speaker": word.get("speaker_id", DEFAULT_SPEAKER), + "energy": round(word.get("energy_norm", 0.0), 3), + "pitch_delta": round(word.get("pitch_delta", 0.0), 3), + "rate_delta": round(word.get("rate_delta", 0.0), 3), + "pause_before": round(word.get("pause_before", 0.0), 3), + "emphasis": round(word.get("emphasis", 0.0), 3), + "energy_raw": word.get("energy"), + "pitch_hz": word.get("pitch_hz"), + } + + +def enrich_words( + words: Sequence[dict], + pitch_track: Optional[Sequence] = None, + energy_track: Optional[Sequence] = None, + weights: EmphasisWeights = EmphasisWeights(), + already_measured: bool = False, +) -> List[dict]: + """Attach normalized acoustic features + the emphasis index to each word. + + Normalization is per-recording: energy against the loudest word, pitch + against the spread around this recording's average, rate against the + largest local departure. That makes the numbers comparable within a + piece regardless of how it was recorded. + + Set ``already_measured`` when the words already carry ``energy`` and + ``pitch_hz`` from a previous pass — re-analyzing a subset, say. The + frame tracks are then unnecessary, and sampling them again would + overwrite good values with ``None``. + """ + if not words: + return [] + + if not already_measured: + words = word_pitch_energy(words, pitch_track, energy_track) + rates = compute_speech_rate(words) + pauses = compute_pauses(words) + + energies = [w["energy"] for w in words if w.get("energy") is not None] + max_energy = max(energies) if energies else 0.0 + pitches = [w["pitch_hz"] for w in words if w.get("pitch_hz") is not None] + avg_pitch = sum(pitches) / len(pitches) if pitches else 0.0 + pitch_span = (max(pitches) - min(pitches)) if len(pitches) > 1 else 0.0 + avg_rate = sum(rates) / len(rates) if rates else 0.0 + max_rate = max(rates) if rates else 0.0 + + enriched: List[dict] = [] + for i, w in enumerate(words): + ww = dict(w) + ww["energy_norm"] = _normalize(w.get("energy"), max_energy) + pitch = w.get("pitch_hz") + ww["pitch_delta"] = ( + _normalize(abs(pitch - avg_pitch), pitch_span) if pitch is not None else 0.0 + ) + ww["rate_delta"] = _normalize(abs(rates[i] - avg_rate), max_rate) + ww["pause_before"] = pauses[i] + enriched.append(ww) + + annotated = annotate_emphasis( + [{**w, "energy": w["energy_norm"]} for w in enriched], weights=weights + ) + for word, scored in zip(enriched, annotated): + word["emphasis"] = scored["emphasis"] + return enriched + + +def _segment_rows(segments: Sequence[dict], words: Sequence[dict]) -> List[dict]: + """Group enriched words under their segment, with per-segment aggregates. + + The aggregates are what let a model judge a whole utterance ("this line + is delivered hot, that one trails off") without reading every word. + """ + rows: List[dict] = [] + previous_end = 0.0 + for seg in segments: + start = float(seg.get("start", 0.0)) + end = float(seg.get("end", 0.0)) + in_seg = [w for w in words if start <= float(w.get("start", 0.0)) < end] + energies = [w["energy_norm"] for w in in_seg] + emphases = [w["emphasis"] for w in in_seg] + gap = max(0.0, start - previous_end) + rows.append( + { + "start": round(start, 3), + "end": round(end, 3), + "speaker": seg.get("speaker_id", DEFAULT_SPEAKER), + "text": (seg.get("text") or "").strip(), + # Silence before this line. Long gaps are where the camera + # stopped or the take restarted, so this is the structural + # hint for splitting a recording into takes — the same signal + # that is *noise* for emphasis (see emphasis.pause_weight). + "gap_before": round(gap, 3), + "take_boundary": gap >= TAKE_BOUNDARY_GAP, + "avg_energy": round(sum(energies) / len(energies), 3) if energies else 0.0, + "peak_emphasis": round(max(emphases), 3) if emphases else 0.0, + "words": [_round_word(w) for w in in_seg], + } + ) + previous_end = end + return rows + + +# Words too common to ever be the point of a punch-in. A zoom lands on what a +# sentence is *about*, and an article spoken loudly is still an article. +_FUNCTION_WORDS = { + "a", "o", "e", "de", "da", "do", "que", "é", "em", "um", "uma", "as", "os", + "no", "na", "com", "pra", "para", "por", "se", "mais", "isso", "aí", "tudo", + "ao", "à", "dos", "das", "nos", "nas", "ou", "mas", "já", "ele", "ela", + "eu", "você", "seu", "sua", "meu", "minha", "esse", "essa", "aquele", +} + + +def _survives(start: float, end: float, cuts: Sequence[Tuple[float, float]]) -> bool: + """Whether a span lies entirely outside every removed range.""" + return all(end <= cut_start or start >= cut_end for cut_start, cut_end in cuts) + + +def restrict_to_kept( + timeline: dict, + cut_ranges: Sequence[Tuple[float, float]], + weights: EmphasisWeights = EmphasisWeights(), + peak_percentile: float = 0.02, + emphasis_floor: float = 0.25, +) -> dict: + """Re-analyze a timeline over only the material that survives ``cut_ranges``. + + Emphasis is *relative*: energy is scored against the loudest word, + pitch against the spread of the recording. Cut the loudest moment out — + a laugh, an aside to the crew — and every remaining score is measured + against something the viewer will never see. Re-running the + normalization over just the survivors is what makes "the most emphatic + line of the final video" a meaningful question. + + Returns a timeline of the same shape, with times still in original + source seconds so the result can be fed straight back as actions. + """ + kept_words = [ + w + for segment in timeline.get("segments", []) + for w in segment.get("words", []) + if _survives(w["start"], w["end"], cut_ranges) + ] + # enrich_words expects the raw analysis keys, not the normalized ones. + raw = [ + { + "word": w["text"], + "start": w["start"], + "end": w["end"], + "speaker_id": w.get("speaker", DEFAULT_SPEAKER), + "energy": w.get("energy_raw"), + "pitch_hz": w.get("pitch_hz"), + } + for w in kept_words + ] + enriched = enrich_words(raw, weights=weights, already_measured=True) + + kept_segments = [ + {**s, "words": [w for w in s.get("words", []) if _survives(w["start"], w["end"], cut_ranges)]} + for s in timeline.get("segments", []) + ] + kept_segments = [s for s in kept_segments if s["words"]] + rows = _segment_rows( + [{"text": s["text"], "start": s["start"], "end": s["end"], + "speaker_id": s.get("speaker", DEFAULT_SPEAKER)} for s in kept_segments], + enriched, + ) + duration = sum(s["end"] - s["start"] for s in rows) + return { + **timeline, + "summary": _summary(enriched, rows, timeline.get("speakers", []), + duration, peak_percentile, emphasis_floor), + "segments": rows, + } + + +def sentence_end(segments: Sequence[dict], index: int) -> float: + """Where the sentence starting at ``segments[index]`` actually finishes. + + Transcription segments break on breath and timing, not on grammar — a + sentence routinely spans two or three of them ("…que dá aquele ar" / + "de elegância, isso é desejo de muitas mulheres, né?"). A zoom that + ends on a segment boundary would therefore release mid-thought, so the + window is extended until a segment closes with terminal punctuation. + """ + last = float(segments[index]["end"]) + for offset, segment in enumerate(segments[index:]): + # A long gap means the take stopped; never run a zoom across that. + # Checked before adopting the end, or the boundary segment's own + # end would already have been taken. + if offset > 0 and segment.get("take_boundary"): + break + last = float(segment["end"]) + if (segment.get("text") or "").strip().endswith((".", "!", "?", "…")): + break + return last + + +def suggest_zoom_windows( + timeline: dict, + min_gap: float = 8.0, + max_zooms: Optional[int] = None, +) -> List[dict]: + """Propose punch-in windows over a timeline's strongest lines. + + One zoom per line at most, taken from the line's most emphatic + *content* word — a loudly spoken "a" is still an article, so function + words are skipped. The window runs from that word to the end of its + line, which is the shape the edit wants: the move lands with the word + and holds through the rest of the phrase. + + ``min_gap`` keeps successive zooms apart; effects stacked close + together read as nervous editing rather than emphasis. + """ + segments = timeline.get("segments", []) + candidates: List[dict] = [] + for i, segment in enumerate(segments): + content = [ + w for w in segment.get("words", []) + if w["text"].strip(",.!?;:").lower() not in _FUNCTION_WORDS + ] + if not content: + continue + best = max(content, key=lambda w: w["emphasis"]) + candidates.append({ + "start": best["start"], + # Hold through to the end of the sentence, not of the segment — + # releasing mid-thought is what makes a punch-in feel arbitrary. + "end": sentence_end(segments, i), + "word": best["text"], + "emphasis": best["emphasis"], + "line": segment["text"], + }) + + chosen: List[dict] = [] + for candidate in sorted(candidates, key=lambda c: c["emphasis"], reverse=True): + if max_zooms is not None and len(chosen) >= max_zooms: + break + if any(abs(candidate["start"] - c["start"]) < min_gap for c in chosen): + continue + chosen.append(candidate) + return sorted(chosen, key=lambda c: c["start"]) + + +def speaker_profiles(segments: Sequence[dict], duration: float) -> List[dict]: + """Per-speaker statistics and sample lines, so a person can tell who is who. + + A bare ``SPEAKER_00`` label is useless for deciding whose audio to cut. + What identifies a role is *how* someone participates: an interviewer or + a crew member asks short questions and holds little of the runtime, + while the subject speaks in long stretches. ``avg_segment`` and + ``share`` capture exactly that contrast, and the sample lines confirm + it in the person's own words. + """ + by_speaker: dict = {} + for seg in segments: + sid = seg.get("speaker", seg.get("speaker_id", DEFAULT_SPEAKER)) + length = max(0.0, float(seg.get("end", 0.0)) - float(seg.get("start", 0.0))) + entry = by_speaker.setdefault(sid, {"seconds": 0.0, "segments": [], "words": 0}) + entry["seconds"] += length + entry["words"] += len(seg.get("words", [])) + entry["segments"].append(seg) + + profiles: List[dict] = [] + for i, (sid, entry) in enumerate( + sorted(by_speaker.items(), key=lambda kv: kv[1]["seconds"], reverse=True) + ): + count = len(entry["segments"]) + # Longest lines identify a role far better than the first ones: a + # question and an answer look alike at the start of a recording. + longest = sorted( + entry["segments"], + key=lambda s: float(s.get("end", 0)) - float(s.get("start", 0)), + reverse=True, + )[:3] + profiles.append({ + "id": sid, + "name": f"Speaker {i + 1}", + "speaking_seconds": round(entry["seconds"], 2), + "share": round(entry["seconds"] / duration, 3) if duration > 0 else 0.0, + "segment_count": count, + "avg_segment": round(entry["seconds"] / count, 2) if count else 0.0, + "word_count": entry["words"], + "samples": [(s.get("text") or "").strip()[:160] for s in longest], + }) + return profiles + + +def select_peaks( + words: Sequence[dict], percentile: float, floor: float +) -> List[dict]: + """The most emphatic words: the top ``percentile`` fraction, above ``floor``. + + Selection is relative on purpose. The emphasis index is a weighted + average whose real range depends entirely on the material — a measured + interview peaks around 0.5 while an energetic ad reaches much higher — + so any fixed cutoff either floods one and selects nothing in the other. + Asking for "the top 2%" instead yields a usable handful either way. + + ``floor`` is only a sanity guard for genuinely flat audio, where even + the top of the distribution carries no emphasis worth cutting on. + """ + ranked = sorted(words, key=lambda w: w["emphasis"], reverse=True) + keep = max(1, round(len(ranked) * percentile)) if ranked else 0 + return [w for w in ranked[:keep] if w["emphasis"] >= floor] + + +def _summary(words: Sequence[dict], segments: Sequence[dict], speakers: Sequence[dict], + duration: float, peak_percentile: float, emphasis_floor: float) -> dict: + """The top layer: the shape of the piece in a handful of numbers.""" + emphases = [w["emphasis"] for w in words] + peaks = select_peaks(words, peak_percentile, emphasis_floor) + return { + "duration": round(duration, 3), + "speaker_count": len(speakers), + "segment_count": len(segments), + "word_count": len(words), + "avg_emphasis": round(sum(emphases) / len(emphases), 3) if emphases else 0.0, + "peak_selection": f"top {peak_percentile:.0%} of words, minimum emphasis {emphasis_floor:.2f}", + "peak_count": len(peaks), + "peak_moments": [ + { + "time": round(float(w.get("start", 0.0)), 3), + "text": w.get("word", ""), + "speaker": w.get("speaker_id", DEFAULT_SPEAKER), + "emphasis": round(w["emphasis"], 3), + } + for w in sorted(peaks, key=lambda w: w["emphasis"], reverse=True)[:20] + ], + } + + +def build_voice_timeline( + media_path: str, + transcript: dict, + hf_token: Optional[str] = None, + num_speakers: str = "", + weights: EmphasisWeights = EmphasisWeights(), + peak_percentile: float = 0.02, + emphasis_floor: float = 0.25, + progress_cb: Optional[Callable[[float, str], None]] = None, +) -> dict: + """Build the consolidated voice timeline for one media file. + + Every analysis layer is optional and degrades independently: without + librosa the acoustic values are ``0.0``; without a diarization token + every word belongs to ``SPEAKER_00``. The document's shape never + changes, so downstream consumers (the rules engine, or a model reading + the JSON) can rely on it. + """ + def report(fraction: float, stage: str) -> None: + if progress_cb: + progress_cb(fraction, stage) + + report(0.1, "Analisando tom e energia...") + pitch_track = extract_pitch(media_path) + energy_track = extract_energy(media_path) + + report(0.5, "Calculando ênfase...") + words = enrich_words(transcript.get("words", []), pitch_track, energy_track, weights) + + report(0.7, "Identificando participantes...") + tracks = diarize(media_path, hf_token, num_speakers) if hf_token else None + segments, words = assign_speakers(transcript.get("segments", []), words, tracks) + speakers = build_speakers(segments) + + report(0.9, "Montando linha do tempo...") + duration = float(transcript.get("duration", 0.0)) + segment_rows = _segment_rows(segments, words) + return { + "version": VOICE_TIMELINE_VERSION, + "source": Path(media_path).name, + "language": transcript.get("language", ""), + # What actually ran, not what was installed — a consumer must be able + # to tell "this speech is flat" from "the acoustics never loaded", + # since both leave the same zeros in the data. + "layers": { + "transcript": bool(transcript.get("words")), + "acoustics": pitch_track is not None or energy_track is not None, + "speakers": tracks is not None, + }, + "scales": VALUE_SCALES, + "summary": _summary( + words, segments, speakers, duration, peak_percentile, emphasis_floor + ), + "speakers": speaker_profiles(segment_rows, duration), + "segments": segment_rows, + } + + +def voice_timeline_path(media_path: str, output_dir: Optional[str] = None) -> Path: + """Where the ``_voice_timeline.json`` for ``media_path`` lives. + + Mirrors ``_transcript.json``: next to the media, or in the chosen + project folder when one is set. + """ + p = Path(media_path) + if output_dir: + directory = Path(output_dir).expanduser() + directory.mkdir(parents=True, exist_ok=True) + return directory / f"{p.stem}_voice_timeline.json" + return p.with_name(p.stem + "_voice_timeline.json") + + +def save_voice_timeline(timeline: dict, path: Path) -> None: + """Write the timeline as UTF-8 JSON (accented transcripts stay readable).""" + with open(path, "w", encoding="utf-8") as f: + json.dump(timeline, f, ensure_ascii=False, indent=2) + + +def load_voice_timeline(path: Path) -> Optional[dict]: + """Read a cached voice timeline, or ``None`` when absent/unreadable.""" + try: + with open(path, encoding="utf-8") as f: + data = json.load(f) + except (OSError, json.JSONDecodeError, UnicodeDecodeError): + return None + return data if isinstance(data, dict) and "segments" in data else None diff --git a/code/fcpxml/writer.py b/code/fcpxml/writer.py index b11564e..c853df2 100755 --- a/code/fcpxml/writer.py +++ b/code/fcpxml/writer.py @@ -35,9 +35,11 @@ import unicodedata import uuid import xml.etree.ElementTree as ET from datetime import datetime +from fractions import Fraction from pathlib import Path from typing import Any, Dict, List, Optional, Tuple +from .collision import blocking, validate_titles from .models import ( _FCPXML_STANDARD_TIMEBASES, DynamicSubtitleConfig, @@ -50,7 +52,12 @@ from .models import ( ValidationIssue, ValidationIssueType, ) -from .text_layout import LayoutBox, compose_sentence, layout_sentence +from .text_layout import ( + TEXT_TEMPLATE_FONT_SCALE, + LayoutBox, + compose_sentence, + layout_sentence, +) from .transcribe import group_words_by_segment # Maximum lengths for XML attribute values to prevent memory abuse @@ -154,6 +161,23 @@ _ASSET_CLIP_CHILD_ORDER = [ _CHILD_ORDER_INDEX = {tag: i for i, tag in enumerate(_ASSET_CLIP_CHILD_ORDER)} +# How close to the end of a clip a zoom must finish for the return to be +# skipped. Within this margin the cut arrives before the eye registers the +# move back, so the return reads as a twitch rather than a resolution. +HOLD_AT_CUT_THRESHOLD = 1.0 + +# How close to the start of a clip a zoom must begin for the ramp-in to be +# skipped and the shot to simply open already zoomed. Tighter than the end +# margin on purpose: at the end the cut hides an unfinished return, but at +# the start a ramp is visible from frame one and reads as the shot settling. +START_AT_CUT_THRESHOLD = 0.5 + + +def _fmt_scale(value: float) -> str: + """Format a scale factor without trailing float noise (1.0 -> "1").""" + return f"{value:.6f}".rstrip("0").rstrip(".") or "0" + + def _dtd_insert(parent: ET.Element, child: ET.Element) -> ET.Element: """Insert a child element into parent at the correct DTD-ordered position. @@ -541,10 +565,43 @@ def _check_timebases(root: ET.Element) -> List[ValidationIssue]: return issues +def _document_frame_duration(root: ET.Element) -> Optional[Fraction]: + """The sequence's exact ``frameDuration`` as a fraction, if declared. + + Read from the format the ``<sequence>`` references (falling back to the + first declared format), so the value is the document's own timebase + rather than an assumed rate. + """ + formats = {f.get('id'): f for f in root.findall('.//format') if f.get('id')} + sequence = root.find('.//sequence') + fmt = formats.get(sequence.get('format')) if sequence is not None else None + if fmt is None: + fmt = next(iter(formats.values()), None) + if fmt is None: + return None + raw = fmt.get('frameDuration', '') + if not (raw.endswith('s') and '/' in raw): + return None + numerator, denominator = raw[:-1].split('/', 1) + try: + value = Fraction(int(numerator), int(denominator)) + except (ValueError, ZeroDivisionError): + return None + return value if value > 0 else None + + def _check_frame_alignment(root: ET.Element, fps: float = 24.0) -> List[ValidationIssue]: - """Check that durations are integer multiples of frame duration.""" + """Check that durations are integer multiples of the frame duration. + + Uses the document's exact ``frameDuration`` fraction and rational + arithmetic. Comparing against an integer fps instead would flag every + NTSC project as broken: at 1001/24000s (23.976fps) a perfectly aligned + duration is not an integer number of "24fps" frames, so whole timelines + would be reported misaligned when nothing is wrong. + """ issues = [] - fps_int = int(fps) + frame_duration = _document_frame_duration(root) + label = f"{1 / float(frame_duration):.3f}".rstrip('0').rstrip('.') if frame_duration else str(fps) for elem in root.iter(): dur_str = elem.get('duration') if not dur_str or not dur_str.endswith('s'): @@ -553,14 +610,19 @@ def _check_frame_alignment(root: ET.Element, fps: float = 24.0) -> List[Validati continue try: tv = TimeValue.from_timecode(dur_str) - frames = tv.to_seconds() * fps_int - if abs(frames - round(frames)) > 0.01: + if frame_duration is not None: + frames = Fraction(tv.numerator, tv.denominator) / frame_duration + aligned = frames.denominator == 1 + else: + approx = tv.to_seconds() * fps + aligned = abs(approx - round(approx)) <= 0.01 + if not aligned: issues.append(ValidationIssue( issue_type=ValidationIssueType.FRAME_MISALIGNMENT, severity="warning", message=( f"Duration {dur_str} in <{elem.tag}> " - f"'{elem.get('name', '')}' is not frame-aligned at {fps_int}fps." + f"'{elem.get('name', '')}' is not frame-aligned at {label}fps." ), clip_name=elem.get('name'), )) @@ -1030,13 +1092,21 @@ class FCPXMLModifier: elem.set(attr, val) return elem - def _require_clip(self, clip_id: str) -> ET.Element: + def _require_clip(self, clip_id: 'str | ET.Element') -> ET.Element: """Look up a clip by ID/name, raising if not found. Centralises the get-or-raise pattern used by every clip-mutating method so the error message stays consistent and future enhancements (fuzzy matching, suggestions) only need one site. + + An Element is returned as-is. That matters after ``split_clip`` or + ``cut_clip_ranges``: the resulting pieces all carry the *same* name, + so a name lookup would always resolve to the first one and silently + put the edit on the wrong piece. Callers holding the exact element + pass it directly. """ + if isinstance(clip_id, ET.Element): + return clip_id clip = self.clips.get(clip_id) if clip is None: raise ValueError(f"Clip not found: {clip_id}") @@ -1421,7 +1491,7 @@ class FCPXMLModifier: def add_marker( self, - clip_id: str, + clip_id: 'str | ET.Element', timecode: str, name: str, marker_type: "MarkerType | str" = MarkerType.STANDARD, @@ -1937,35 +2007,50 @@ class FCPXMLModifier: def add_zoom( self, - clip_id: str, + clip_id: 'str | ET.Element', start: float, end: float, scale: float = 1.3, - ease: float = 0.3, + ease: float = 0.25, position: str = "0 0", + ease_out: Optional[float] = None, + hold_at_end: Optional[bool] = None, + start_at_peak: Optional[bool] = None, ) -> ET.Element: - """Add a smooth ease-in/ease-out punch-in zoom to a clip. + """Add a punch-in zoom to a clip, snapping back to its framing at the end. - Animates ``<adjust-transform>``'s ``scale`` param (per the FCPXML - DTD: ``<param>`` + ``<keyframeAnimation>`` of ``<keyframe>`` - elements, ``interp="ease"``) from 100% up to *scale* and back down - to 100%, entirely within ``[start, end]`` — clip-relative seconds - (seconds from the clip's own head, same convention as - ``cut_clip_ranges``). The ease portions each last *ease* seconds; - the zoom holds at *scale* in between. + Animates ``<adjust-transform>``'s ``scale`` param (``<param>`` + + ``<keyframeAnimation>`` of ``<keyframe>``) from the clip's current + scale up to *scale* times it, holds, then returns — all within + ``[start, end]`` — clip-relative seconds (same convention as + ``cut_clip_ranges``). + + The two ends are deliberately asymmetric. *ease* ramps the zoom + **in** over half a second by default, fast enough to land with the + emphasised word. The way **out** is instant — a single frame — so + the moment the impact phrase ends the shot is simply back to its + normal framing and the video resumes its flow, with no drift + drawing attention to itself. Pass *ease_out* to ramp the return + gradually instead. + + *hold_at_end* keeps the peak instead of returning, and + *start_at_peak* opens already zoomed with no ramp. Left as ``None`` + both decide on their own from how close the window sits to the + clip's edges: a cut is itself the transition, so ramping away from + one — or back toward one — is motion the viewer reads as a wobble + rather than as emphasis. """ if end <= start: raise ValueError(f"end ({end}) must be greater than start ({start})") if ease <= 0: raise ValueError(f"ease must be positive, got {ease}") - if ease * 2 > (end - start): - raise ValueError( - f"ease ({ease}s x2 = {ease * 2}s) doesn't fit in the zoom " - f"window ({end - start}s) — shorten ease or widen start/end" - ) if scale <= 0: raise ValueError(f"scale must be positive, got {scale}") + frame = float(self.frame_duration_fraction()) + ramp_out = frame if ease_out is None else ease_out + if ramp_out <= 0: + raise ValueError(f"ease_out must be positive, got {ease_out}") clip = self._require_clip(clip_id) clip_duration = self._parse_time(clip.get('duration', '0s')).to_seconds() if start < 0 or end > clip_duration: @@ -1974,26 +2059,146 @@ class FCPXMLModifier: f"duration (0 to {clip_duration:.3f}s)" ) - # Replace rather than stack a prior zoom on the same clip. + # Replace a prior zoom, but never the clip's framing. A clip can + # already carry an <adjust-transform> holding the editor's own + # reframe — rotation for footage shot sideways, position, a scale + # that makes the shot work at all. Dropping it outright (the old + # behaviour) silently destroyed that framing; on real footage the + # zoomed section came back rotated. So: keep the static attributes, + # and animate *relative to* the existing scale. + base_x, base_y = 1.0, 1.0 + carried: dict = {} + old_keyframes: list = [] for stale in clip.findall('adjust-transform'): + carried = {k: v for k, v in stale.attrib.items() if k != 'scale'} + parts = (stale.get('scale') or '').split() + if len(parts) == 2: + try: + base_x, base_y = float(parts[0]), float(parts[1]) + except ValueError: + base_x, base_y = 1.0, 1.0 + else: + # No static attribute — a PRIOR zoom on this same clip left + # an animated <param name="scale"> instead, and the true + # resting framing lives in its keyframes, not in 1.0. + # Reading it as 1.0 here doesn't just miss the framing: it + # replaces the earlier zoom's whole animation with a wrong + # one, since this loop unconditionally removes `stale` + # right after. The rest value is recoverable without + # knowing which keyframe it is: MIN_ZOOM_SCALE == 1.0 means + # every keyframed value is >= the rest scale, so the + # smallest one keyframed is the rest value, peak or not. + for old_param in stale.findall("param[@name='scale']"): + xs, ys = [], [] + for kf in old_param.findall('.//keyframe'): + kv = (kf.get('value') or '').split() + if len(kv) == 2: + try: + xs.append(float(kv[0])) + ys.append(float(kv[1])) + except ValueError: + pass + # Kept for merging: a second zoom on the same clip + # (two emphatic beats a cut didn't separate) should + # stack alongside the first, not erase it — the + # earlier peak is still a real editorial decision. + old_keyframes.append((kf.get('time', '0s'), kf.get('value', ''))) + if xs and ys: + base_x, base_y = min(xs), min(ys) clip.remove(stale) transform = ET.Element('adjust-transform') + for key, value in carried.items(): + transform.set(key, value) scale_param = ET.SubElement(transform, 'param') scale_param.set('name', 'scale') anim = ET.SubElement(scale_param, 'keyframeAnimation') - scale_value = f"{scale} {scale}" - for seconds, value in ( - (start, "1 1"), - (start + ease, scale_value), - (end - ease, scale_value), - (end, "1 1"), - ): + # Keyframe times live in the clip's SOURCE timebase — the same origin + # as its own ``start`` — not in clip-relative seconds. A clip whose + # media starts at, say, 3109.9s of timecode looks for the animation + # there; keyframes written at 0-5s land outside the clip entirely and + # Final Cut imports the zoom as nothing at all, silently. Matches what + # add_text_title already does, and only shows up on footage whose + # start isn't 0s — every synthetic fixture starts at 0s and hides it. + media_origin = self._parse_time(clip.get('start', '0s')) + + rest_value = f"{_fmt_scale(base_x)} {_fmt_scale(base_y)}" + scale_value = f"{_fmt_scale(base_x * scale)} {_fmt_scale(base_y * scale)}" + + # A return that lands right before a cut is wasted motion: the next + # clip begins on its own framing anyway, so all the viewer sees is a + # twitch on the way out. When the zoom runs to the end of the clip, + # hold the peak and let the cut do the resetting. + holds_to_cut = ( + hold_at_end + if hold_at_end is not None + else (clip_duration - end) <= HOLD_AT_CUT_THRESHOLD + ) + opens_at_peak = ( + start_at_peak + if start_at_peak is not None + else start <= START_AT_CUT_THRESHOLD + ) + + # Only the ramps actually written have to fit in the window: a zoom + # that opens at the peak spends no time ramping in, and one held to + # the cut spends none ramping out. + needed = (0.0 if opens_at_peak else ease) + (0.0 if holds_to_cut else ramp_out) + if needed > (end - start): + raise ValueError( + f"the ramps ({needed}s) don't fit in the zoom window " + f"({end - start}s) — shorten them or widen start/end" + ) + + if opens_at_peak: + # The cut already delivered the change of framing; ramping up + # from it just looks like the shot settling. + keyframes = [(start, scale_value)] + else: + keyframes = [(start, rest_value), (start + ease, scale_value)] + if holds_to_cut: + keyframes.append((end, scale_value)) + else: + # Hold the peak right up to the end, then drop back on the very + # next frame — the snap-back the edit wants, not a slow drift. + keyframes.append((end - ramp_out, scale_value)) + keyframes.append((end, rest_value)) + + new_entries = [ + ((media_origin + self.snap_seconds_to_frame(seconds)), value) + for seconds, value in keyframes + ] + new_start_time = new_entries[0][0] + new_end_time = new_entries[-1][0] + + # Two calls on the same clip mean two different things depending on + # whether their windows overlap. Overlapping = redoing the *same* + # zoom with new numbers — the old keyframes are stale and all of + # them go. Disjoint = a second, separate beat that a cut didn't + # separate onto its own clip — that one stacks alongside the first + # instead of erasing it, since both are real editorial decisions. + old_times = [self._parse_time(t) for t, _ in old_keyframes] + old_span_overlaps_new = bool(old_times) and not ( + max(old_times) < new_start_time or min(old_times) > new_end_time + ) + if old_span_overlaps_new: + surviving_old: list = [] + else: + surviving_old = [(self._parse_time(t), v) for t, v in old_keyframes] + all_entries = sorted(surviving_old + new_entries, key=lambda e: e[0]) + + for time_value, value in all_entries: kf = ET.SubElement(anim, 'keyframe') - kf.set('time', self.snap_seconds_to_frame(seconds).to_fcpxml()) + kf.set('time', time_value.to_fcpxml()) kf.set('value', value) - kf.set('interp', 'ease') + # Only 'time' and 'value' — no 'interp', no 'curve'. The DTD allows + # both, but Final Cut rejected 'interp' on this vector param + # ("does not support the interpolation attribute") and discarded + # the whole <param>. A hand-made zoom exported from FCP itself + # writes bare keyframes and relies on the DTD default + # (curve="smooth"), so we match that export exactly rather than + # guess which attributes survive its importer. if position != "0 0": pos_param = ET.SubElement(transform, 'param') @@ -2707,7 +2912,18 @@ class FCPXMLModifier: _TEXT_POSITION_KEY = '9999/10003/13260/3296672360/1/100/101' # Layout params the "Text" template ships with. These keys are the # template's own defaults and never vary between instances. + # + # "Build Out" is the one deliberate override: with "Apply Speed" set to + # "2 (Per Object)" below, the template's whole built-in animation (build + # in + build out) is always compressed to exactly fill the title's own + # on-screen duration — so on a short word-length clip, build out was + # eating time that build in needed to finish revealing the text before + # the cut. Disabling build out hands that entire compressed window to + # build in alone, which is what "sempre acelerado" turned out to mean: + # no separate speed knob needed. Value captured from a real FCP export + # with "Build Out" unchecked in the Inspector (see chat, 2026-08-18). _TEXT_TITLE_PARAMS = ( + ('Build Out', '9999/10000/2/102', '0'), ('Layout Method', '9999/10003/13260/3296672360/2/314', '1 (Paragraph)'), ('Left Margin', '9999/10003/13260/3296672360/2/323', '-1210'), ('Right Margin', '9999/10003/13260/3296672360/2/324', '1210'), @@ -2795,6 +3011,7 @@ class FCPXMLModifier: bold: bool = True, face: Optional[str] = None, kerning: Optional[float] = None, + font_scale: float = TEXT_TEMPLATE_FONT_SCALE, ) -> ET.Element: """Build a standalone ``<title>`` clip from the "Text" (Basic Text) template. @@ -2849,13 +3066,31 @@ class FCPXMLModifier: style_def.set('id', ts_id) text_style = ET.SubElement(style_def, 'text-style') text_style.set('font', font) - text_style.set('fontSize', str(font_size)) + # Text.moti sizes type in frame pixels but positions in canvas points. + # See TEXT_TEMPLATE_FONT_SCALE: layout measures in points, so only the + # emitted size (and its kerning, to keep the same letter spacing) is + # converted here. + scale = float(font_scale) or 1.0 + text_style.set('fontSize', f"{float(font_size) * scale:g}") text_style.set('fontColor', font_color) - text_style.set('bold', '1' if bold else '0') - if face: + # FCP represents bold weight as the bold attribute — never as a + # fontFace. Writing ``bold="0" fontFace="Bold"`` (the previous + # behaviour) is contradictory and FCP refuses to render the text. + # Italic, by contrast, IS a face: FCP writes both ``fontFace`` and + # ``italic="1"``. See Engine/docs/05_EXPERIENCIAS.md, entry 2026-08-19. + face_lower = (face or '').strip().lower() + if face_lower == 'bold': + text_style.set('bold', '1') + elif 'italic' in face_lower: text_style.set('fontFace', face) + text_style.set('italic', '1') + else: + if bold: + text_style.set('bold', '1') + if face: + text_style.set('fontFace', face) if kerning: - text_style.set('kerning', f"{float(kerning):g}") + text_style.set('kerning', f"{float(kerning) * scale:g}") text_style.set('alignment', 'center') text_style.set('lineSpacing', '-19') @@ -3005,7 +3240,9 @@ class FCPXMLModifier: def lay_out(pending: List[Dict]): """Place what fits; return (units, still-unplaced words).""" if phrase_mode: - composition = compose_sentence(pending, config.style, box) + composition = compose_sentence( + pending, config.style, box, line_gap=config.line_gap, + ) return composition.blocks, composition.overflow layout = layout_sentence(pending, config.style, box) return layout.placed, layout.overflow @@ -3083,19 +3320,98 @@ class FCPXMLModifier: duration, lane=lane, name=f"caption_{uuid.uuid4().hex[:8]}", - position=unit.position_param(), + position=unit.position_param(config.text_scale), font=unit.font or config.style.font, font_size=int(round(unit.font_size)), font_color=unit.color or config.style.active_color, bold=config.style.bold, face=unit.face, kerning=unit.kerning, + font_scale=config.text_scale, ) _dtd_insert(parent, title) created.append(title) + if getattr(config, 'validate', False): + report = self.validate_subtitle_layout() + if blocking(report["severity"]): + raise ValueError( + "Subtitle layout validation failed: " + + str(report["summary"]) + ) + return created + def validate_subtitle_layout( + self, + *, + safe_margin_x: float = 0.05, + safe_margin_y: float = 0.05, + min_font_size: Optional[float] = None, + min_distance: Optional[float] = None, + max_distance: Optional[float] = None, + ) -> dict: + """Re-measure every ``<title>`` in the document and report collisions. + + Reconstructs each title's on-screen box from the values the writer + emitted (``fontSize``/``kerning``/``Position`` are already in template + space), then checks for temporal+spatial collisions, frame/safe-area + containment, and font fallbacks. This is the spec-16 validation pass the + layout engine does not do on its own — it only guarantees non-overlap + *by construction* while composing, and cannot see a hand-edited title. + + Returns the ``collision.validate_titles`` report: ``severity`` (worst + bucket), ``issues`` (spec-16 occurrences) and ``summary`` (counts). + """ + titles = [] + for elem in self.root.iter('title'): + text_el = elem.find('text/text-style') + text = (text_el.text or '').strip() if text_el is not None else '' + style = elem.find('text-style-def/text-style') + font = style.get('font') if style is not None else None + face = style.get('fontFace') if style is not None else None + font_size = ( + float(style.get('fontSize', '0')) if style is not None else 0.0 + ) + kerning = ( + float(style.get('kerning', '0') or 0) + if style is not None else 0.0 + ) + + x = y = 0.0 + for param in elem.findall('param'): + if param.get('name') == 'Position' and param.get('value'): + parts = param.get('value').split() + if len(parts) >= 2: + x, y = float(parts[0]), float(parts[1]) + + start = self._parse_time(elem.get('offset', '0s')).to_seconds() + duration = self._parse_time(elem.get('duration', '0s')).to_seconds() + + titles.append({ + 'text': text, + 'font': font, + 'face': face, + 'font_size': font_size, + 'kerning': kerning, + 'x': x, + 'y': y, + 'start': start, + 'end': start + duration, + 'group': start + duration, + }) + + return validate_titles( + titles, + self.frame_width(), + self.frame_height(), + safe_margin_x=safe_margin_x, + safe_margin_y=safe_margin_y, + min_font_size=min_font_size, + min_distance=min_distance, + max_distance=max_distance, + ) + # ======================================================================== # AUDIO CLIP OPERATIONS (v0.6.0) # ======================================================================== diff --git a/code/server.py b/code/server.py old mode 100755 new mode 100644 index 2f16ce6..0215603 --- a/code/server.py +++ b/code/server.py @@ -2,17 +2,17 @@ """ FCPXML MCP Server — Batch operations and analysis for Final Cut Pro XML files. -Provides 53 tools, MCP resources for file discovery, and pre-built prompt -workflows for common editing tasks. +Composition root: wires MCP resources/prompts, concatenates each +server_tools/*.py module's TOOLS and HANDLERS into list_tools() and +TOOL_HANDLERS, and dispatches. No tool logic lives here — see +Engine/docs/03_SERVER_TOOLS.md for the catalog and server_tools/ for the +73 tool definitions + handlers, grouped by category. Author: DareDev256 (https://github.com/DareDev256) """ from __future__ import annotations -import json -import os -import re from pathlib import Path from typing import Any, Sequence @@ -25,533 +25,317 @@ from mcp.types import ( PromptMessage, Resource, TextContent, - Tool, ) -from fcpxml.diff import compare_timelines -from fcpxml.export import DaVinciExporter -from fcpxml.media_intel import ( - detect_beats, - detect_silence, - map_silence_to_timeline, - media_src_to_path, +from server_tools import ( + editing, + export, + generation, + live, + markers_import, + qc, + roles, + subtitles, + timeline, + transcript, + voice, ) -from fcpxml.models import ( - DuplicateGroup, - DynamicSubtitleConfig, - FlashFrame, - FlashFrameSeverity, - GapInfo, - MarkerType, - SegmentSpec, - Timecode, - TimeValue, - WordLook, - WordStyle, +from server_tools._shared import ( + _DIARIZATION_INSTALL_HINT, + _FEATURES_INSTALL_HINT, + _MAX_JSON_DEPTH, + _SANDBOX_ENABLED, + _TRANSCRIBE_INSTALL_HINT, + AUDIO_MEDIA_EXTENSIONS, + MAX_FILE_SIZE, + MAX_MEDIA_FILE_SIZE, + PROJECTS_DIR, + TRANSCRIBE_MAX_MEDIA, + _apply_placed_action, + _check_json_depth, + _cut_transcript_spans, + _detect_duplicate_groups, + _detect_flash_frames, + _detect_gaps, + _extract_subtitle_blocks, + _fmt_suggestions, + _format_batch_result, + _format_clip_table, + _load_or_transcribe, + _markdown_table, + _no_timeline, + _NoTimelineError, + _parse_project, + _parse_timestamp_parts, + _raw_markers_to_batch, + _require_timeline, + _resolve_io_paths, + _setup_generator, + _setup_modifier, + _speaker_table, + _text_result, + _transcript_cut_report, + _transcript_json_path, + _validate_directory, + _validate_filepath, + _validate_output_path, + _voice_analysis_config_text, + find_fcpxml_files, + format_duration, + format_timecode, + generate_output_path, + parse_srt, + parse_transcript_timestamps, + parse_vtt, ) -from fcpxml.parser import FCPXMLParser -from fcpxml.rough_cut import RoughCutGenerator -from fcpxml.templates import ClipSpec, apply_template, list_templates -from fcpxml.transcribe import ( - DEFAULT_FILLERS, - find_filler_spans, - find_phrase_spans, - invert_ranges, - merge_ranges, - segments_to_srt, - transcribe, +from server_tools.editing import ( + handle_add_audio, + handle_add_connected_clip, + handle_add_marker, + handle_add_transition, + handle_add_zoom, + handle_batch_add_markers, + handle_change_speed, + handle_create_compound_clip, + handle_delete_clips, + handle_fill_gaps, + handle_fix_flash_frames, + handle_flatten_compound_clip, + handle_insert_clip, + handle_rapid_trim, + handle_reformat_timeline, + handle_reorder_clips, + handle_split_clip, + handle_trim_clip, ) -from fcpxml.writer import FCPXMLModifier, list_effects +from server_tools.export import ( + handle_export_csv, + handle_export_edl, + handle_export_fcp7_xml, + handle_export_resolve_xml, + handle_relink_media, +) +from server_tools.generation import ( + handle_apply_template, + handle_auto_rough_cut, + handle_generate_ab_roll, + handle_generate_montage, + handle_list_templates, +) +from server_tools.live import ( + handle_list_fcp_libraries, + handle_push_to_fcp, +) +from server_tools.markers_import import ( + handle_import_beat_markers, + handle_import_srt_markers, + handle_import_transcript_markers, + handle_snap_to_beats, + handle_transcript_markers, +) +from server_tools.qc import ( + handle_analyze_pacing, + handle_detect_beats, + handle_detect_duplicates, + handle_detect_flash_frames, + handle_detect_gaps, + handle_detect_media_silence, + handle_detect_silence_candidates, + handle_find_long_clips, + handle_find_short_cuts, + handle_remove_media_silence, + handle_remove_silence_candidates, + handle_validate_timeline, +) +from server_tools.roles import ( + handle_assign_role, + handle_export_role_stems, + handle_filter_by_role, +) +from server_tools.subtitles import ( + handle_generate_dynamic_subtitles, + handle_validate_subtitle_layout, +) +from server_tools.timeline import ( + handle_analyze_timeline, + handle_diff_timelines, + handle_list_clips, + handle_list_compound_clips, + handle_list_connected_clips, + handle_list_effects, + handle_list_keywords, + handle_list_library_clips, + handle_list_markers, + handle_list_projects, + handle_list_roles, +) +from server_tools.transcript import ( + handle_edit_by_transcript, + handle_remove_filler_words, + handle_transcribe_media, +) +from server_tools.voice import ( + handle_analyze_voice_features, + handle_apply_voice_actions, + handle_build_voice_timeline, + handle_diarize_media, + handle_get_voice_analysis_config, + handle_refine_voice_timeline, + handle_remove_speakers, + handle_save_voice_analysis_config, +) + +# Re-exported below (via __all__) so `from server import handle_x` / `_helper` +# keeps working for tests and any external caller that imported the old +# monolith directly — this module is the stable public surface, +# server_tools/ is the implementation. __all__ tells the linter these +# imports are the point, not dead code. + +__all__ = [ + "server", + "main", + "main_sync", + "call_tool", + "list_tools", + "list_resources", + "read_resource", + "list_prompts", + "get_prompt", + "TOOL_HANDLERS", + "PROJECTS_DIR", + "_SANDBOX_ENABLED", + "MAX_FILE_SIZE", + "MAX_MEDIA_FILE_SIZE", + "_MAX_JSON_DEPTH", + "_check_json_depth", + "_validate_filepath", + "_validate_output_path", + "_validate_directory", + "find_fcpxml_files", + "format_timecode", + "format_duration", + "_format_clip_table", + "_markdown_table", + "_format_batch_result", + "_fmt_suggestions", + "generate_output_path", + "_parse_project", + "_text_result", + "_no_timeline", + "_require_timeline", + "_NoTimelineError", + "_resolve_io_paths", + "_setup_modifier", + "_setup_generator", + "_parse_timestamp_parts", + "_raw_markers_to_batch", + "_extract_subtitle_blocks", + "parse_srt", + "parse_vtt", + "parse_transcript_timestamps", + "_detect_flash_frames", + "_detect_gaps", + "_detect_duplicate_groups", + "AUDIO_MEDIA_EXTENSIONS", + "_DIARIZATION_INSTALL_HINT", + "_FEATURES_INSTALL_HINT", + "_voice_analysis_config_text", + "_apply_placed_action", + "_speaker_table", + "TRANSCRIBE_MAX_MEDIA", + "_TRANSCRIBE_INSTALL_HINT", + "_transcript_json_path", + "_load_or_transcribe", + "_cut_transcript_spans", + "_transcript_cut_report", + "handle_list_projects", + "handle_analyze_timeline", + "handle_list_clips", + "handle_list_markers", + "handle_list_keywords", + "handle_list_library_clips", + "handle_list_connected_clips", + "handle_list_compound_clips", + "handle_list_roles", + "handle_diff_timelines", + "handle_list_effects", + "handle_find_short_cuts", + "handle_find_long_clips", + "handle_analyze_pacing", + "handle_detect_flash_frames", + "handle_detect_duplicates", + "handle_detect_gaps", + "handle_validate_timeline", + "handle_detect_media_silence", + "handle_detect_beats", + "handle_remove_media_silence", + "handle_detect_silence_candidates", + "handle_remove_silence_candidates", + "handle_add_marker", + "handle_batch_add_markers", + "handle_trim_clip", + "handle_reorder_clips", + "handle_add_transition", + "handle_change_speed", + "handle_add_zoom", + "handle_delete_clips", + "handle_split_clip", + "handle_insert_clip", + "handle_fix_flash_frames", + "handle_rapid_trim", + "handle_fill_gaps", + "handle_add_connected_clip", + "handle_reformat_timeline", + "handle_add_audio", + "handle_create_compound_clip", + "handle_flatten_compound_clip", + "handle_auto_rough_cut", + "handle_generate_montage", + "handle_generate_ab_roll", + "handle_list_templates", + "handle_apply_template", + "handle_import_beat_markers", + "handle_snap_to_beats", + "handle_import_srt_markers", + "handle_import_transcript_markers", + "handle_transcript_markers", + "handle_assign_role", + "handle_filter_by_role", + "handle_export_role_stems", + "handle_transcribe_media", + "handle_edit_by_transcript", + "handle_remove_filler_words", + "handle_export_edl", + "handle_export_csv", + "handle_export_resolve_xml", + "handle_export_fcp7_xml", + "handle_relink_media", + "handle_diarize_media", + "handle_analyze_voice_features", + "handle_build_voice_timeline", + "handle_remove_speakers", + "handle_refine_voice_timeline", + "handle_apply_voice_actions", + "handle_get_voice_analysis_config", + "handle_save_voice_analysis_config", + "handle_validate_subtitle_layout", + "handle_generate_dynamic_subtitles", + "handle_push_to_fcp", + "handle_list_fcp_libraries", +] __version__ = "0.13.1" server = Server("fcp-mcp-server", version=__version__) -PROJECTS_DIR = os.environ.get("FCP_PROJECTS_DIR", os.path.expanduser("~/Movies")) -# When set explicitly via env var, enforce sandbox boundaries on list_projects. -_SANDBOX_ENABLED = "FCP_PROJECTS_DIR" in os.environ -# Maximum file size for parsing (100 MB). -MAX_FILE_SIZE = 100 * 1024 * 1024 +# The category modules composing this server's tool catalog, in the order +# Engine/docs/03_SERVER_TOOLS.md lists them. +_CATEGORY_MODULES = ( + timeline, qc, editing, generation, markers_import, roles, + transcript, export, voice, subtitles, live, +) -# ============================================================================ -# SECURITY UTILITIES -# ============================================================================ - -# Maximum nesting depth for JSON deserialization (beat markers, configs). -# Prevents stack overflow / memory exhaustion from deeply nested payloads. -_MAX_JSON_DEPTH = 50 - - -def _check_json_depth(obj: object, _depth: int = 0) -> None: - """Reject JSON structures nested beyond _MAX_JSON_DEPTH. - - Prevents denial-of-service via deeply nested objects that exhaust the - call stack or memory during downstream processing. Called after - json.load() since Python's json module has no built-in depth limit. - """ - if _depth > _MAX_JSON_DEPTH: - raise ValueError( - f"JSON nesting depth exceeds {_MAX_JSON_DEPTH} — " - "file may be malformed or adversarial" - ) - if isinstance(obj, dict): - for v in obj.values(): - _check_json_depth(v, _depth + 1) - elif isinstance(obj, list): - for item in obj: - _check_json_depth(item, _depth + 1) - - -def _validate_filepath(filepath: str, allowed_extensions: tuple[str, ...] | None = None) -> str: - """Validate a user-provided file path against traversal and size attacks. - - Resolves symlinks, blocks null bytes, enforces extension whitelist, and - checks file size before any parsing takes place. - - Raises: - ValueError: For invalid paths (null bytes, bad extensions, oversized). - FileNotFoundError: When the resolved path does not exist. - """ - if '\x00' in filepath: - raise ValueError("Invalid file path: null byte detected") - - resolved = Path(filepath).resolve() - - if not resolved.exists(): - raise FileNotFoundError(f"File not found: {filepath}") - - # .fcpxmld bundles are directories (a package wrapping Info.fcpxml plus - # sidecar data files for object tracking / Cinematic mode). The size - # check applies to the inner Info.fcpxml, which is what gets parsed. - if resolved.is_dir(): - if resolved.suffix.lower() != '.fcpxmld': - raise ValueError(f"Not a regular file: {filepath}") - inner = resolved / 'Info.fcpxml' - if not inner.is_file(): - raise ValueError(f"Invalid bundle (no Info.fcpxml): {filepath}") - size_target = inner - elif not resolved.is_file(): - raise ValueError(f"Not a regular file: {filepath}") - else: - size_target = resolved - - if allowed_extensions and resolved.suffix.lower() not in allowed_extensions: - raise ValueError( - f"Invalid file type '{resolved.suffix}'. " - f"Allowed: {', '.join(allowed_extensions)}" - ) - - if size_target.stat().st_size > MAX_FILE_SIZE: - size_mb = size_target.stat().st_size / (1024 * 1024) - raise ValueError(f"File too large ({size_mb:.1f} MB). Maximum: {MAX_FILE_SIZE // (1024 * 1024)} MB") - - return str(resolved) - - -def _validate_output_path(output_path: str, *, anchor_dir: str | None = None) -> str: - """Validate an output path with optional sandbox enforcement. - - Resolves traversal, blocks null bytes, ensures parent exists, and — when - *anchor_dir* is provided — verifies the resolved output lives under that - directory. This prevents LLM-generated tool calls from writing to - arbitrary filesystem locations (e.g. ``/etc/cron.d/backdoor``). - - Args: - output_path: The raw output path to validate. - anchor_dir: If set, the resolved output must be a child of this - directory. Typically the parent directory of the input file so - outputs stay co-located with their sources. - - Raises: - ValueError: For null bytes, missing parent, or sandbox escape. - """ - if '\x00' in output_path: - raise ValueError("Invalid output path: null byte detected") - - resolved = Path(output_path).resolve() - - if not resolved.parent.exists(): - raise ValueError(f"Output directory does not exist: {resolved.parent}") - - if anchor_dir is not None: - anchor = Path(anchor_dir).resolve() - try: - resolved.relative_to(anchor) - except ValueError: - raise ValueError( - f"Output path escapes allowed directory: " - f"{resolved} is not under {anchor}" - ) - - return str(resolved) - - -def _validate_directory(directory: str, *, allowed_root: str | None = None) -> str: - """Validate a user-provided directory path against traversal and injection. - - Resolves symlinks, blocks null bytes, and verifies the path is a real - directory. When *allowed_root* is given, the resolved path must be a - descendant of (or equal to) that root — preventing filesystem enumeration - beyond the project workspace. - - Raises: - ValueError: For invalid paths (null bytes, not a directory, sandbox escape). - """ - if '\x00' in directory: - raise ValueError("Invalid directory path: null byte detected") - - resolved = Path(directory).resolve() - - if not resolved.is_dir(): - raise ValueError(f"Not a valid directory: {directory}") - - if allowed_root is not None: - root = Path(allowed_root).resolve() - try: - resolved.relative_to(root) - except ValueError: - raise ValueError( - f"Directory escapes allowed root: " - f"{resolved} is not under {root}" - ) - - return str(resolved) - - -# ============================================================================ -# UTILITIES -# ============================================================================ - -def find_fcpxml_files(directory: str) -> list[str]: - """Find all FCPXML files in a directory.""" - path = Path(directory) - files = list(str(f) for f in path.rglob("*.fcpxml")) - files.extend(str(f) for f in path.rglob("*.fcpxmld")) - return sorted(files) - - -def format_timecode(tc) -> str: - """Format a Timecode object to SMPTE string.""" - return tc.to_smpte() if tc else "00:00:00:00" - - -def format_duration(seconds: float) -> str: - """Format seconds into human-readable duration.""" - if seconds < 1: - return f"{seconds*1000:.0f}ms" - elif seconds < 60: - return f"{seconds:.2f}s" - return f"{int(seconds // 60)}m {seconds % 60:.1f}s" - - -def _format_clip_table(clips: list, header: str) -> str: - """Render a list of clips as a markdown table with timecodes and durations. - - Shared by handlers that filter clips by duration threshold - (find_short_cuts, find_long_clips). - """ - result = f"{header}\n\n| Name | TC | Duration |\n|------|----|---------|\n" - result += "\n".join( - f"| {c.name} | {format_timecode(c.start)} | {format_duration(c.duration_seconds)} |" - for c in clips - ) - return result - - -def _markdown_table(headers: list[str], rows: list[list[str]]) -> str: - """Build a markdown table from headers and rows. - - Returns header row, separator row, and data rows as a single string. - Callers avoid repeating the ``| H1 | H2 |\\n|---|---|`` boilerplate - that appears in 15+ handlers. - """ - header_line = "| " + " | ".join(headers) + " |" - sep_line = "|" + "|".join("------" for _ in headers) + "|" - data_lines = "\n".join( - "| " + " | ".join(str(c) for c in row) + " |" for row in rows - ) - return f"{header_line}\n{sep_line}\n{data_lines}" - - -def _format_batch_result( - title: str, - summary: dict[str, str], - headers: list[str], - rows: list[list[str]], - output_path: str, -) -> str: - """Build a standard batch-operation result with summary, table, and save footer. - - Used by batch fix handlers (flash frames, rapid trim, fill gaps) that all - share the same markdown structure: ``# Title → ## Summary → ## Details table - → Saved to`` footer. - """ - summary_lines = "\n".join(f"- **{k}**: {v}" for k, v in summary.items()) - table = _markdown_table(headers, rows) - return ( - f"# {title}\n\n" - f"## Summary\n{summary_lines}\n\n" - f"## Details\n{table}\n\n" - f"Saved to: `{output_path}`" - ) - - -def _fmt_suggestions(suggestions: list[str]) -> str: - """Format pacing suggestions as markdown list (Python 3.10 compatible).""" - if not suggestions: - return "- Pacing looks good!" - nl = "\n" - return nl.join(f"- {s}" for s in suggestions) - - -def generate_output_path(input_path: str, suffix: str = "_modified") -> str: - """Generate output path from input path. - - The suffix is sanitized to prevent path-component injection — only - alphanumeric, hyphen, underscore, and dot characters survive. - """ - # Strip anything that could inject path separators or traversal sequences - clean_suffix = re.sub(r'[^a-zA-Z0-9._-]', '', suffix) - if not clean_suffix: - clean_suffix = "_modified" - p = Path(input_path) - return str(p.parent / f"{p.stem}{clean_suffix}{p.suffix}") - - -def _parse_project(filepath: str): - """Parse an FCPXML file and return the project with its primary timeline.""" - filepath = _validate_filepath(filepath, ('.fcpxml', '.fcpxmld')) - project = FCPXMLParser().parse_file(filepath) - if not project.timelines: - return None, None - return project, project.primary_timeline - - -def _text_result(text: str) -> list[TextContent]: - """Wrap a string in the MCP TextContent list that every tool handler returns.""" - return [TextContent(type="text", text=text)] - - -def _no_timeline(): - """Standard response when no timelines are found.""" - return _text_result("No timelines found") - - -def _require_timeline(filepath: str): - """Parse FCPXML and return (project, timeline), raising if no timeline exists. - - Centralises the repeated _parse_project + _no_timeline guard that - appears in every read-only timeline handler. Returns a tuple so - callers can destructure directly:: - - project, tl = _require_timeline(arguments["filepath"]) - """ - project, tl = _parse_project(filepath) - if not tl: - raise _NoTimelineError() - return project, tl - - -class _NoTimelineError(Exception): - """Sentinel raised by _require_timeline when no timelines exist.""" - - -def _resolve_io_paths( - arguments: dict, - suffix: str = "_modified", -) -> tuple[str, str]: - """Validate input filepath and resolve the output path. - - Shared foundation for every handler that reads an FCPXML and writes - a derived file. Validates the input, falls back to a suffixed - output name when ``output_path`` is not supplied, and sandbox-checks - the result. - - Args: - arguments: Tool arguments dict (must contain ``filepath``; may - contain ``output_path``). - suffix: Default output filename suffix when ``output_path`` is - not provided (e.g. ``"_modified"``, ``"_beats"``). - - Returns: - ``(filepath, output_path)`` tuple with both paths validated. - """ - filepath = _validate_filepath(arguments["filepath"], ('.fcpxml', '.fcpxmld')) - # Anchor write operations to the input file's directory so LLM-generated - # tool calls cannot write to arbitrary filesystem locations (e.g. - # /etc/cron.d/backdoor). When the explicit sandbox is off, the anchor - # still prevents writes outside the source directory tree. - output_dir = arguments.get("output_dir") - anchor = _validate_directory(str(output_dir)) if output_dir else str(Path(filepath).resolve().parent) - output_path = _validate_output_path( - arguments.get("output_path") or generate_output_path(filepath, suffix), - anchor_dir=anchor, - ) - return filepath, output_path - - -def _setup_modifier( - arguments: dict, - suffix: str = "_modified", -) -> tuple[str, str, "FCPXMLModifier"]: - """Common setup for write handlers: validate paths and create modifier. - - Consolidates the repeated validate-filepath → resolve-output-path → - create-modifier boilerplate shared by 18+ write handlers. - - Args: - arguments: Tool arguments dict (must contain ``filepath``; may - contain ``output_path``). - suffix: Default output filename suffix when ``output_path`` is - not provided (e.g. ``"_modified"``, ``"_flash_fixed"``). - - Returns: - ``(filepath, output_path, modifier)`` tuple ready for the - handler's domain-specific operation. - """ - filepath, output_path = _resolve_io_paths(arguments, suffix) - modifier = FCPXMLModifier(filepath) - return filepath, output_path, modifier - - -def _setup_generator( - arguments: dict, - suffix: str = "_roughcut", -) -> tuple[str, str, "RoughCutGenerator"]: - """Common setup for generation handlers: validate paths and create generator. - - Args: - arguments: Tool arguments dict (must contain ``filepath`` and - ``output_path``). - suffix: Default output filename suffix. - - Returns: - ``(filepath, output_path, generator)`` tuple. - """ - filepath, output_path = _resolve_io_paths(arguments, suffix) - generator = RoughCutGenerator(filepath) - return filepath, output_path, generator - - -def _parse_timestamp_parts( - parts: list[str], *, frame_rate: float = 24.0 -) -> float | None: - """Convert colon-separated timestamp parts to total seconds. - - Handles 2-part (M:SS), 3-part (H:MM:SS / HH:MM:SS.ms), and - 4-part (HH:MM:SS:FF SMPTE) formats. Returns ``None`` when the - part count is unrecognised so callers can skip. - - Args: - parts: Colon-split timestamp components. - frame_rate: FPS used to convert the frame component of SMPTE - timecodes into fractional seconds (default 24.0). - """ - if len(parts) == 2: - return int(parts[0]) * 60 + float(parts[1]) - elif len(parts) == 3: - return int(parts[0]) * 3600 + int(parts[1]) * 60 + float(parts[2]) - elif len(parts) == 4: - # SMPTE: HH:MM:SS:FF — convert frames to fractional seconds - base = int(parts[0]) * 3600 + int(parts[1]) * 60 + float(parts[2]) - frames = int(parts[3]) - return base + (frames / frame_rate) if frame_rate > 0 else base - return None - - -def _raw_markers_to_batch( - raw_markers: list[dict], - marker_type: str = "chapter", - max_label: int | None = None, -) -> list[dict]: - """Convert raw {seconds, text} marker dicts to batch_add_markers format. - - Shared by import_srt_markers and import_transcript_markers. - """ - batch = [] - for m in raw_markers: - label = m["text"] - if max_label and len(label) > max_label: - label = label[:max_label] - batch.append({ - "timecode": f"{m['seconds']}s", - "name": label, - "marker_type": marker_type.upper(), - }) - return batch - - -def _extract_subtitle_blocks(text: str, *, strip_vtt_tags: bool = False) -> list[dict]: - """Extract timestamp/text pairs from subtitle cue blocks (SRT or VTT). - - Both SRT and VTT use the same ``start --> end`` cue syntax with - text lines underneath; only header stripping and tag cleaning differ. - """ - markers = [] - blocks = re.split(r'\n\s*\n', text.strip()) - for block in blocks: - lines = block.strip().split('\n') - if len(lines) < 2: - continue - ts_line = None - text_lines = [] - for line in lines: - if '-->' in line: - ts_line = line - elif ts_line is not None: - if strip_vtt_tags: - line = re.sub(r'<[^>]+>', '', line) - cleaned = line.strip() - if cleaned: - text_lines.append(cleaned) - if not ts_line or not text_lines: - continue - start_str = ts_line.split('-->')[0].strip().replace(',', '.') - seconds = _parse_timestamp_parts(start_str.split(':')) - if seconds is not None: - markers.append({'seconds': seconds, 'text': ' '.join(text_lines)}) - return markers - - -def parse_srt(text: str) -> list[dict]: - """Parse SRT subtitle format into timestamp/text pairs.""" - return _extract_subtitle_blocks(text) - - -def parse_vtt(text: str) -> list[dict]: - """Parse WebVTT subtitle format into timestamp/text pairs.""" - text = re.sub(r'^WEBVTT.*?\n', '', text, flags=re.MULTILINE) - text = re.sub(r'NOTE\n.*?\n\n', '', text, flags=re.DOTALL) - return _extract_subtitle_blocks(text, strip_vtt_tags=True) - - -def parse_transcript_timestamps(text: str) -> list[dict]: - """Parse timestamped text (YouTube description format) into markers. - - Supports formats like: - 0:00 Introduction - 00:01:30 Main Topic - 1:05:30 Conclusion - 00:00:00:00 SMPTE timecode - """ - markers = [] - for line in text.strip().split('\n'): - line = line.strip() - if not line: - continue - match = re.match(r'^(\d{1,2}:\d{2}(?::\d{2}){0,2})\s+(.+)$', line) - if match: - seconds = _parse_timestamp_parts(match.group(1).split(':')) - if seconds is not None: - markers.append({'seconds': seconds, 'text': match.group(2).strip()}) - return markers - - -# ============================================================================ -# MCP RESOURCES — File discovery -# ============================================================================ - @server.list_resources() async def list_resources() -> list[Resource]: """Expose discovered FCPXML files as MCP resources.""" @@ -750,3360 +534,24 @@ Please: raise ValueError(f"Unknown prompt: {name}") + # ============================================================================ -# TOOL DEFINITIONS +# TOOL CATALOG & DISPATCH # ============================================================================ @server.list_tools() -async def list_tools() -> list[Tool]: - return [ - # ===== READ TOOLS ===== - Tool( - name="list_projects", - description="List all FCPXML projects in directory", - inputSchema={ - "type": "object", - "properties": { - "directory": {"type": "string", "description": "Directory to search (default: ~/Movies)"} - } - } - ), - Tool( - name="analyze_timeline", - description="Get comprehensive timeline statistics including duration, resolution, clip count, pacing metrics", - inputSchema={ - "type": "object", - "properties": {"filepath": {"type": "string", "description": "Path to FCPXML file"}}, - "required": ["filepath"] - } - ), - Tool( - name="list_clips", - description="List all clips with timecodes, durations, and metadata", - inputSchema={ - "type": "object", - "properties": { - "filepath": {"type": "string"}, - "limit": {"type": "integer", "description": "Max clips to return"} - }, - "required": ["filepath"] - } - ), - Tool( - name="list_markers", - description="Extract markers (chapter, todo, standard) with timestamps", - inputSchema={ - "type": "object", - "properties": { - "filepath": {"type": "string"}, - "marker_type": {"type": "string", "enum": ["all", "chapter", "todo", "standard", "completed"]}, - "format": {"type": "string", "enum": ["detailed", "youtube", "simple"]} - }, - "required": ["filepath"] - } - ), - Tool( - name="find_short_cuts", - description="Find clips shorter than threshold (flash frame detection)", - inputSchema={ - "type": "object", - "properties": { - "filepath": {"type": "string"}, - "threshold_seconds": {"type": "number", "default": 0.5} - }, - "required": ["filepath"] - } - ), - Tool( - name="find_long_clips", - description="Find clips longer than threshold", - inputSchema={ - "type": "object", - "properties": { - "filepath": {"type": "string"}, - "threshold_seconds": {"type": "number", "default": 10.0} - }, - "required": ["filepath"] - } - ), - Tool( - name="list_keywords", - description="Extract all keywords/tags from project", - inputSchema={ - "type": "object", - "properties": {"filepath": {"type": "string"}}, - "required": ["filepath"] - } - ), - Tool( - name="export_edl", - description="Generate EDL (Edit Decision List) from timeline", - inputSchema={ - "type": "object", - "properties": {"filepath": {"type": "string"}}, - "required": ["filepath"] - } - ), - Tool( - name="export_csv", - description="Export timeline data to CSV format", - inputSchema={ - "type": "object", - "properties": { - "filepath": {"type": "string"}, - "include": {"type": "array", "items": {"type": "string"}} - }, - "required": ["filepath"] - } - ), - Tool( - name="analyze_pacing", - description="Analyze edit pacing with suggestions for improvements", - inputSchema={ - "type": "object", - "properties": {"filepath": {"type": "string"}}, - "required": ["filepath"] - } - ), - Tool( - name="list_library_clips", - description="List all available clips in the library (source media, not yet on timeline)", - inputSchema={ - "type": "object", - "properties": { - "filepath": {"type": "string", "description": "Path to FCPXML file"}, - "keywords": {"type": "array", "items": {"type": "string"}, "description": "Filter by keywords"}, - "limit": {"type": "integer", "description": "Max clips to return"} - }, - "required": ["filepath"] - } - ), - - # ===== QC / VALIDATION TOOLS ===== - Tool( - name="detect_flash_frames", - description="Find ultra-short clips (flash frames) that are likely errors, with severity categorization", - inputSchema={ - "type": "object", - "properties": { - "filepath": {"type": "string", "description": "Path to FCPXML file"}, - "critical_threshold_frames": {"type": "integer", "default": 2, "description": "Frames below this = critical (default: 2)"}, - "warning_threshold_frames": {"type": "integer", "default": 6, "description": "Frames below this = warning (default: 6)"} - }, - "required": ["filepath"] - } - ), - Tool( - name="detect_duplicates", - description="Find clips using the same source media (potential duplicates)", - inputSchema={ - "type": "object", - "properties": { - "filepath": {"type": "string", "description": "Path to FCPXML file"}, - "mode": {"type": "string", "enum": ["same_source", "overlapping_ranges", "identical"], "default": "same_source", "description": "Detection mode"} - }, - "required": ["filepath"] - } - ), - Tool( - name="detect_gaps", - description="Find unintentional gaps in the timeline", - inputSchema={ - "type": "object", - "properties": { - "filepath": {"type": "string", "description": "Path to FCPXML file"}, - "min_gap_frames": {"type": "integer", "default": 1, "description": "Minimum gap size to detect (default: 1 frame)"} - }, - "required": ["filepath"] - } - ), - - # ===== WRITE TOOLS ===== - Tool( - name="add_marker", - description="Add a marker at a specific timecode", - inputSchema={ - "type": "object", - "properties": { - "filepath": {"type": "string", "description": "Path to FCPXML file"}, - "timecode": {"type": "string", "description": "Position (00:00:10:00 or 10s)"}, - "name": {"type": "string", "description": "Marker label"}, - "marker_type": {"type": "string", "enum": ["standard", "chapter", "todo", "completed"], "default": "standard"}, - "note": {"type": "string", "description": "Optional note"}, - "output_path": {"type": "string", "description": "Output path (default: adds _modified suffix)"} - }, - "required": ["filepath", "timecode", "name"] - } - ), - Tool( - name="batch_add_markers", - description="Add multiple markers at once, or auto-generate at cuts/intervals", - inputSchema={ - "type": "object", - "properties": { - "filepath": {"type": "string"}, - "markers": { - "type": "array", - "items": { - "type": "object", - "properties": { - "timecode": {"type": "string"}, - "name": {"type": "string"}, - "marker_type": {"type": "string"}, - "note": {"type": "string"} - } - }, - "description": "List of markers to add" - }, - "auto_at_cuts": {"type": "boolean", "description": "Add marker at every cut"}, - "auto_at_intervals": {"type": "string", "description": "Add markers every N seconds (e.g., '30s')"}, - "output_path": {"type": "string"} - }, - "required": ["filepath"] - } - ), - Tool( - name="trim_clip", - description="Trim a clip's in-point and/or out-point", - inputSchema={ - "type": "object", - "properties": { - "filepath": {"type": "string"}, - "clip_id": {"type": "string", "description": "Clip name or ID"}, - "trim_start": {"type": "string", "description": "New in-point or delta (+1s, -10f)"}, - "trim_end": {"type": "string", "description": "New out-point or delta"}, - "ripple": {"type": "boolean", "default": True, "description": "Shift subsequent clips"}, - "output_path": {"type": "string"} - }, - "required": ["filepath", "clip_id"] - } - ), - Tool( - name="reorder_clips", - description="Move clips to a new position in the timeline", - inputSchema={ - "type": "object", - "properties": { - "filepath": {"type": "string"}, - "clip_ids": {"type": "array", "items": {"type": "string"}, "description": "Clips to move"}, - "target_position": {"type": "string", "description": "'start', 'end', timecode, or 'after:clip_id'"}, - "ripple": {"type": "boolean", "default": True}, - "output_path": {"type": "string"} - }, - "required": ["filepath", "clip_ids", "target_position"] - } - ), - Tool( - name="add_transition", - description="Add a transition between clips", - inputSchema={ - "type": "object", - "properties": { - "filepath": {"type": "string"}, - "clip_id": {"type": "string", "description": "Clip to add transition to"}, - "position": {"type": "string", "enum": ["start", "end", "both"], "default": "end"}, - "transition_type": {"type": "string", "enum": ["cross-dissolve", "fade-to-black", "fade-from-black", "wipe"], "default": "cross-dissolve"}, - "duration": {"type": "string", "default": "00:00:00:15"}, - "output_path": {"type": "string"} - }, - "required": ["filepath", "clip_id"] - } - ), - Tool( - name="change_speed", - description="Change clip playback speed (slow motion or speed up)", - inputSchema={ - "type": "object", - "properties": { - "filepath": {"type": "string"}, - "clip_id": {"type": "string"}, - "speed": {"type": "number", "description": "Speed multiplier (0.5 = half, 2.0 = double)"}, - "preserve_pitch": {"type": "boolean", "default": True}, - "output_path": {"type": "string"} - }, - "required": ["filepath", "clip_id", "speed"] - } - ), - Tool( - name="add_zoom", - description="Add a smooth ease-in/ease-out punch-in zoom to a clip, animating <adjust-transform>'s scale param via keyframes (100% -> scale -> 100%) entirely within [start, end] (clip-relative seconds, i.e. seconds from the clip's own head). The ease portions each last `ease` seconds; the zoom holds at `scale` in between. Replaces any existing zoom on the same clip rather than stacking.", - inputSchema={ - "type": "object", - "properties": { - "filepath": {"type": "string"}, - "clip_id": {"type": "string", "description": "Name/ID of the clip to zoom"}, - "start": {"type": "number", "description": "Clip-relative seconds where the ease-in begins"}, - "end": {"type": "number", "description": "Clip-relative seconds where the ease-out ends (back to 100%)"}, - "scale": {"type": "number", "default": 1.3, "description": "Zoom scale, e.g. 1.3 = 130%"}, - "ease": {"type": "number", "default": 0.3, "description": "Seconds for each of the ease-in/ease-out portions (must fit: 2*ease <= end-start)"}, - "position": {"type": "string", "default": "0 0", "description": "Optional pan offset \"x y\" applied for the duration of the transform"}, - "output_path": {"type": "string"} - }, - "required": ["filepath", "clip_id", "start", "end"] - } - ), - Tool( - name="delete_clips", - description="Delete clips from timeline", - inputSchema={ - "type": "object", - "properties": { - "filepath": {"type": "string"}, - "clip_ids": {"type": "array", "items": {"type": "string"}}, - "ripple": {"type": "boolean", "default": True, "description": "Close gaps after deletion"}, - "output_path": {"type": "string"} - }, - "required": ["filepath", "clip_ids"] - } - ), - Tool( - name="split_clip", - description="Split a clip at specified timecodes", - inputSchema={ - "type": "object", - "properties": { - "filepath": {"type": "string"}, - "clip_id": {"type": "string"}, - "split_points": {"type": "array", "items": {"type": "string"}, "description": "Timecodes to split at"}, - "output_path": {"type": "string"} - }, - "required": ["filepath", "clip_id", "split_points"] - } - ), - Tool( - name="insert_clip", - description="Insert a library clip onto the timeline at a specific position", - inputSchema={ - "type": "object", - "properties": { - "filepath": {"type": "string", "description": "Path to FCPXML file"}, - "asset_id": {"type": "string", "description": "Asset reference ID (e.g., 'r3')"}, - "asset_name": {"type": "string", "description": "Asset name (alternative to asset_id)"}, - "position": {"type": "string", "description": "'start', 'end', timecode, or 'after:clip_name'"}, - "duration": {"type": "string", "description": "Clip duration (if not using in/out points)"}, - "in_point": {"type": "string", "description": "Source in-point for subclip"}, - "out_point": {"type": "string", "description": "Source out-point for subclip"}, - "ripple": {"type": "boolean", "default": True, "description": "Shift subsequent clips"}, - "output_path": {"type": "string", "description": "Output path (default: adds _modified suffix)"} - }, - "required": ["filepath", "position"] - } - ), - - # ===== BATCH FIX TOOLS ===== - Tool( - name="fix_flash_frames", - description="Automatically fix detected flash frames by extending neighbors or deleting", - inputSchema={ - "type": "object", - "properties": { - "filepath": {"type": "string", "description": "Path to FCPXML file"}, - "mode": {"type": "string", "enum": ["extend_previous", "extend_next", "delete", "auto"], "default": "auto", "description": "How to fix: extend previous/next clip, delete, or auto"}, - "threshold_frames": {"type": "integer", "default": 6, "description": "Frames below this threshold are flash frames"}, - "output_path": {"type": "string", "description": "Output path (default: adds _modified suffix)"} - }, - "required": ["filepath"] - } - ), - Tool( - name="rapid_trim", - description="Batch trim clips to a maximum duration for fast-paced montages", - inputSchema={ - "type": "object", - "properties": { - "filepath": {"type": "string", "description": "Path to FCPXML file"}, - "max_duration": {"type": "string", "description": "Maximum clip duration (e.g., '2s', '00:00:02:00')"}, - "min_duration": {"type": "string", "description": "Minimum clip duration (optional)"}, - "keywords": {"type": "array", "items": {"type": "string"}, "description": "Only trim clips with these keywords"}, - "trim_from": {"type": "string", "enum": ["start", "end", "center"], "default": "end", "description": "Where to trim from"}, - "output_path": {"type": "string", "description": "Output path (default: adds _modified suffix)"} - }, - "required": ["filepath", "max_duration"] - } - ), - Tool( - name="fill_gaps", - description="Automatically fill gaps in the timeline by extending adjacent clips", - inputSchema={ - "type": "object", - "properties": { - "filepath": {"type": "string", "description": "Path to FCPXML file"}, - "mode": {"type": "string", "enum": ["extend_previous", "extend_next", "delete"], "default": "extend_previous", "description": "How to fill gaps"}, - "max_gap": {"type": "string", "description": "Only fill gaps smaller than this (e.g., '1s')"}, - "output_path": {"type": "string", "description": "Output path (default: adds _modified suffix)"} - }, - "required": ["filepath"] - } - ), - Tool( - name="validate_timeline", - description="Comprehensive timeline health check for flash frames, gaps, duplicates, and issues", - inputSchema={ - "type": "object", - "properties": { - "filepath": {"type": "string", "description": "Path to FCPXML file"}, - "checks": {"type": "array", "items": {"type": "string", "enum": ["all", "flash_frames", "gaps", "duplicates", "offsets"]}, "default": ["all"], "description": "Which checks to run"} - }, - "required": ["filepath"] - } - ), - - # ===== GENERATION TOOLS ===== - Tool( - name="auto_rough_cut", - description="Generate a rough cut from source clips based on keywords, duration, and pacing", - inputSchema={ - "type": "object", - "properties": { - "filepath": {"type": "string", "description": "Source FCPXML with clips"}, - "output_path": {"type": "string", "description": "Where to save rough cut"}, - "target_duration": {"type": "string", "description": "Target length (3m, 00:03:00:00)"}, - "pacing": {"type": "string", "enum": ["slow", "medium", "fast", "dynamic"], "default": "medium"}, - "keywords": {"type": "array", "items": {"type": "string"}, "description": "Filter clips by keywords"}, - "segments": { - "type": "array", - "items": { - "type": "object", - "properties": { - "name": {"type": "string"}, - "keywords": {"type": "array", "items": {"type": "string"}}, - "duration": {"type": "number"} - } - }, - "description": "Segment structure [{name, keywords, duration_seconds}]" - }, - "priority": {"type": "string", "enum": ["best", "favorites", "longest", "shortest", "random"], "default": "best"}, - "favorites_only": {"type": "boolean", "default": False}, - "add_transitions": {"type": "boolean", "default": False} - }, - "required": ["filepath", "output_path", "target_duration"] - } - ), - Tool( - name="generate_montage", - description="Create rapid-fire montages with pacing curves (accelerating, decelerating, pyramid)", - inputSchema={ - "type": "object", - "properties": { - "filepath": {"type": "string", "description": "Source FCPXML with clips"}, - "output_path": {"type": "string", "description": "Where to save montage"}, - "target_duration": {"type": "string", "description": "Total montage length (e.g., '30s', '00:00:30:00')"}, - "pacing_curve": {"type": "string", "enum": ["accelerating", "decelerating", "pyramid", "constant"], "default": "accelerating", "description": "How clip duration changes over time"}, - "start_duration": {"type": "number", "default": 2.0, "description": "Clip duration at start (seconds)"}, - "end_duration": {"type": "number", "default": 0.5, "description": "Clip duration at end (seconds)"}, - "keywords": {"type": "array", "items": {"type": "string"}, "description": "Filter clips by keywords"}, - "add_transitions": {"type": "boolean", "default": False, "description": "Add quick dissolves"} - }, - "required": ["filepath", "output_path", "target_duration"] - } - ), - Tool( - name="generate_ab_roll", - description="Create documentary-style A/B roll edits alternating between main content and cutaways", - inputSchema={ - "type": "object", - "properties": { - "filepath": {"type": "string", "description": "Source FCPXML with clips"}, - "output_path": {"type": "string", "description": "Where to save A/B roll edit"}, - "target_duration": {"type": "string", "description": "Total duration (e.g., '3m', '00:03:00:00')"}, - "a_keywords": {"type": "array", "items": {"type": "string"}, "description": "Keywords for A-roll (main content, interviews)"}, - "b_keywords": {"type": "array", "items": {"type": "string"}, "description": "Keywords for B-roll (cutaways, visuals)"}, - "a_duration": {"type": "string", "default": "5s", "description": "Duration of each A-roll segment"}, - "b_duration": {"type": "string", "default": "3s", "description": "Duration of each B-roll cutaway"}, - "start_with": {"type": "string", "enum": ["a", "b"], "default": "a", "description": "Which roll to start with"}, - "add_transitions": {"type": "boolean", "default": True, "description": "Add cross-dissolves"} - }, - "required": ["filepath", "output_path", "target_duration", "a_keywords", "b_keywords"] - } - ), - - # ===== BEAT SYNC TOOLS ===== - Tool( - name="import_beat_markers", - description="Import beat markers from external audio analysis (JSON format)", - inputSchema={ - "type": "object", - "properties": { - "filepath": {"type": "string", "description": "Path to FCPXML file"}, - "beats_path": {"type": "string", "description": "Path to beats JSON file"}, - "marker_type": {"type": "string", "enum": ["standard", "chapter"], "default": "standard"}, - "beat_filter": {"type": "string", "enum": ["all", "downbeat", "measure"], "default": "all", "description": "Which beats to import"}, - "output_path": {"type": "string", "description": "Output path (default: adds _beats suffix)"} - }, - "required": ["filepath", "beats_path"] - } - ), - Tool( - name="snap_to_beats", - description="Align cuts to nearest beat markers for music-synced edits", - inputSchema={ - "type": "object", - "properties": { - "filepath": {"type": "string", "description": "Path to FCPXML file with beat markers"}, - "max_shift_frames": {"type": "integer", "default": 6, "description": "Maximum frames to shift a cut"}, - "prefer": {"type": "string", "enum": ["earlier", "later", "nearest"], "default": "nearest", "description": "Which beat to prefer when equidistant"}, - "output_path": {"type": "string", "description": "Output path (default: adds _synced suffix)"} - }, - "required": ["filepath"] - } - ), - - # ===== SUBTITLE / TRANSCRIPT TOOLS ===== - Tool( - name="import_srt_markers", - description="Import SRT or VTT subtitles as chapter markers on the timeline", - inputSchema={ - "type": "object", - "properties": { - "filepath": {"type": "string", "description": "Path to FCPXML file"}, - "srt_path": {"type": "string", "description": "Path to SRT or VTT subtitle file"}, - "mode": {"type": "string", "enum": ["all", "first_per_minute", "scene_changes"], "default": "first_per_minute", "description": "How to create markers: every subtitle, first per minute, or on text changes"}, - "marker_type": {"type": "string", "enum": ["standard", "chapter"], "default": "chapter"}, - "max_label_length": {"type": "integer", "default": 50, "description": "Truncate marker labels to this length"}, - "output_path": {"type": "string", "description": "Output path (default: adds _subtitled suffix)"} - }, - "required": ["filepath", "srt_path"] - } - ), - Tool( - name="import_transcript_markers", - description="Import timestamped transcript (YouTube chapter format) as markers. Supports '0:00 Title' and 'HH:MM:SS Title' formats", - inputSchema={ - "type": "object", - "properties": { - "filepath": {"type": "string", "description": "Path to FCPXML file"}, - "transcript": {"type": "string", "description": "Timestamped text (one per line: '0:00 Introduction')"}, - "transcript_path": {"type": "string", "description": "Path to text file with timestamps (alternative to inline transcript)"}, - "marker_type": {"type": "string", "enum": ["standard", "chapter"], "default": "chapter"}, - "output_path": {"type": "string", "description": "Output path (default: adds _chapters suffix)"} - }, - "required": ["filepath"] - } - ), - - # ===== CONNECTED CLIPS & COMPOUND CLIPS (v0.5.0) ===== - Tool( - name="list_connected_clips", - description="List all connected clips (B-roll, titles, audio) with their lanes and parent clips", - inputSchema={ - "type": "object", - "properties": { - "filepath": {"type": "string", "description": "Path to FCPXML file"}, - "lane": {"type": "integer", "description": "Filter by lane number (positive=above, negative=below)"}, - }, - "required": ["filepath"] - } - ), - Tool( - name="add_connected_clip", - description="Connect a library clip to an existing timeline clip (B-roll overlay, audio, title)", - inputSchema={ - "type": "object", - "properties": { - "filepath": {"type": "string", "description": "Path to FCPXML file"}, - "parent_clip_id": {"type": "string", "description": "Name/ID of the clip to attach to"}, - "asset_id": {"type": "string", "description": "Asset reference ID"}, - "asset_name": {"type": "string", "description": "Asset name (alternative to asset_id)"}, - "offset": {"type": "string", "default": "0s", "description": "Position relative to parent clip start"}, - "duration": {"type": "string", "description": "Duration (default: full asset)"}, - "lane": {"type": "integer", "default": 1, "description": "Lane number (positive=above, negative=below)"}, - "output_path": {"type": "string", "description": "Output path (default: adds _modified suffix)"} - }, - "required": ["filepath", "parent_clip_id"] - } - ), - Tool( - name="list_compound_clips", - description="List compound clips (ref-clips) and their nested content", - inputSchema={ - "type": "object", - "properties": { - "filepath": {"type": "string", "description": "Path to FCPXML file"}, - }, - "required": ["filepath"] - } - ), - - # ===== ROLES MANAGEMENT (v0.5.0) ===== - Tool( - name="list_roles", - description="List all audio/video roles used in the timeline with clip counts", - inputSchema={ - "type": "object", - "properties": { - "filepath": {"type": "string", "description": "Path to FCPXML file"}, - }, - "required": ["filepath"] - } - ), - Tool( - name="assign_role", - description="Set the audio or video role on a clip (dialogue, music, effects, titles, etc.)", - inputSchema={ - "type": "object", - "properties": { - "filepath": {"type": "string", "description": "Path to FCPXML file"}, - "clip_id": {"type": "string", "description": "Clip name or ID"}, - "audio_role": {"type": "string", "description": "Audio role (e.g., dialogue, music, effects)"}, - "video_role": {"type": "string", "description": "Video role (e.g., video, titles)"}, - "output_path": {"type": "string", "description": "Output path (default: adds _modified suffix)"} - }, - "required": ["filepath", "clip_id"] - } - ), - Tool( - name="filter_by_role", - description="List all clips matching a specific audio or video role", - inputSchema={ - "type": "object", - "properties": { - "filepath": {"type": "string", "description": "Path to FCPXML file"}, - "role": {"type": "string", "description": "Role name to filter by"}, - "role_type": {"type": "string", "enum": ["audio", "video", "any"], "default": "any", "description": "Which role type to search"}, - }, - "required": ["filepath", "role"] - } - ), - Tool( - name="export_role_stems", - description="Export clip list grouped by role for audio mixing stem planning", - inputSchema={ - "type": "object", - "properties": { - "filepath": {"type": "string", "description": "Path to FCPXML file"}, - }, - "required": ["filepath"] - } - ), - - # ===== TIMELINE DIFF (v0.5.0) ===== - Tool( - name="diff_timelines", - description="Compare two FCPXML files and report differences in clips, markers, transitions, and format", - inputSchema={ - "type": "object", - "properties": { - "filepath_a": {"type": "string", "description": "Path to first FCPXML file (baseline)"}, - "filepath_b": {"type": "string", "description": "Path to second FCPXML file (comparison)"}, - }, - "required": ["filepath_a", "filepath_b"] - } - ), - - # ===== SOCIAL MEDIA REFORMAT (v0.5.0) ===== - Tool( - name="reformat_timeline", - description="Create new FCPXML with different resolution/aspect ratio (9:16 for TikTok, 1:1 for Instagram, etc.)", - inputSchema={ - "type": "object", - "properties": { - "filepath": {"type": "string", "description": "Path to FCPXML file"}, - "format": {"type": "string", "enum": ["9:16", "1:1", "4:5", "16:9", "4:3", "custom"], "description": "Target format preset"}, - "width": {"type": "integer", "description": "Custom width (only with format='custom')"}, - "height": {"type": "integer", "description": "Custom height (only with format='custom')"}, - "output_path": {"type": "string", "description": "Output path (default: adds _reformatted suffix)"} - }, - "required": ["filepath", "format"] - } - ), - - # ===== MEDIA INTELLIGENCE (v0.10.0) ===== - Tool( - name="detect_media_silence", - description="Detect REAL silence by analyzing each clip's source audio with ffmpeg silencedetect, mapped into timeline time. Unlike detect_silence_candidates (XML-only heuristics), this reads the actual media files referenced by the timeline. Requires ffmpeg; clips whose media is missing or unreadable are reported, not failed.", - inputSchema={ - "type": "object", - "properties": { - "filepath": {"type": "string", "description": "Path to FCPXML file"}, - "noise_db": {"type": "number", "default": -30.0, "description": "Silence threshold in dBFS, -120 to 0 (default -30)"}, - "min_silence": {"type": "number", "default": 0.5, "description": "Minimum silence duration in seconds to report (default 0.5)"}, - "clip_name": {"type": "string", "description": "Only analyze the clip with this name"}, - }, - "required": ["filepath"] - } - ), - - Tool( - name="detect_beats", - description="Detect musical beats and tempo in an audio/video file (librosa beat tracker). Writes a beats JSON next to the media file that plugs directly into import_beat_markers + snap_to_beats for beat-synced editing. Requires the optional [intelligence] extra (librosa); degrades to an install hint without it.", - inputSchema={ - "type": "object", - "properties": { - "media_path": {"type": "string", "description": "Path to audio/video file (.wav, .mp3, .m4a, .aac, .aif, .flac, .mov, .mp4)"}, - }, - "required": ["media_path"] - } - ), - Tool( - name="remove_media_silence", - description="Detect REAL silence in each clip's source audio (ffmpeg) and CUT it out of the timeline with ripple. Clips are split around silence; the silent middles are removed and everything after shifts earlier. Non-destructive: writes a _silence_removed copy. Preview with detect_media_silence first.", - inputSchema={ - "type": "object", - "properties": { - "filepath": {"type": "string", "description": "Path to FCPXML file"}, - "noise_db": {"type": "number", "default": -30.0, "description": "Silence threshold in dBFS, -120 to 0 (default -30)"}, - "min_silence": {"type": "number", "default": 0.5, "description": "Minimum silence duration in seconds to cut (default 0.5)"}, - "padding": {"type": "number", "default": 0.05, "description": "Seconds of silence to keep on each side of a cut so edits breathe (default 0.05, max 5)"}, - "clip_name": {"type": "string", "description": "Only cut silence in the clip with this name"}, - "output_path": {"type": "string", "description": "Output path (default: adds _silence_removed suffix)"}, - }, - "required": ["filepath"] - } - ), - - # ===== TRANSCRIPT INTELLIGENCE (v0.13.1) ===== - Tool( - name="transcribe_media", - description="Transcribe each clip's source media locally with word-level timestamps (faster-whisper). Writes a _transcript.json next to each media file (reused by edit_by_transcript / remove_filler_words so media is only transcribed once) and optionally an SRT for captions. Requires the optional [transcribe] extra; degrades to an install hint without it.", - inputSchema={ - "type": "object", - "properties": { - "filepath": {"type": "string", "description": "Path to FCPXML file"}, - "clip_name": {"type": "string", "description": "Only transcribe the clip with this name"}, - "model": {"type": "string", "default": "base", "description": "Whisper model size: tiny, base, small, medium, large-v3 (default base; larger = slower + more accurate)"}, - "language": {"type": "string", "description": "ISO language code hint (e.g. 'en'); auto-detected if omitted"}, - "write_srt": {"type": "boolean", "default": False, "description": "Also write a _transcript.srt next to each media file (plugs into import_srt_markers)"}, - }, - "required": ["filepath"] - } - ), - Tool( - name="edit_by_transcript", - description="Text-based editing: cut timeline content by what was SAID. mode=remove cuts every occurrence of the given phrases (with ripple); mode=keep_only keeps only the matched phrases and cuts everything else in each matched clip (clips with no matches are left untouched). Uses each media file's _transcript.json (auto-transcribes if missing). Non-destructive: writes a _transcript_edit copy.", - inputSchema={ - "type": "object", - "properties": { - "filepath": {"type": "string", "description": "Path to FCPXML file"}, - "phrases": {"type": "array", "items": {"type": "string"}, "description": "Spoken phrases to match (case/punctuation-insensitive)"}, - "mode": {"type": "string", "enum": ["remove", "keep_only"], "default": "remove", "description": "remove=cut matches out; keep_only=keep only matches"}, - "clip_name": {"type": "string", "description": "Only edit the clip with this name"}, - "model": {"type": "string", "default": "base", "description": "Whisper model size if transcription is needed"}, - "padding": {"type": "number", "default": 0.0, "description": "Seconds to widen each cut on both sides (0-2, default 0)"}, - "output_path": {"type": "string", "description": "Output path (default: adds _transcript_edit suffix)"}, - }, - "required": ["filepath", "phrases"] - } - ), - Tool( - name="remove_filler_words", - description="Cut filler words (um, uh, erm...) out of the timeline with ripple, using word-level transcripts of the real source audio. Conservative default filler list — words like 'like' and 'so' are only cut if you pass them explicitly. Uses each media file's _transcript.json (auto-transcribes if missing). Non-destructive: writes a _defillered copy.", - inputSchema={ - "type": "object", - "properties": { - "filepath": {"type": "string", "description": "Path to FCPXML file"}, - "fillers": {"type": "array", "items": {"type": "string"}, "description": "Filler words/phrases to cut (default: um, uh, uhh, umm, erm, ehm, mmm, hmm, mhm)"}, - "clip_name": {"type": "string", "description": "Only clean the clip with this name"}, - "model": {"type": "string", "default": "base", "description": "Whisper model size if transcription is needed"}, - "padding": {"type": "number", "default": 0.02, "description": "Seconds to widen each cut on both sides (0-2, default 0.02)"}, - "output_path": {"type": "string", "description": "Output path (default: adds _defillered suffix)"}, - }, - "required": ["filepath"] - } - ), - Tool( - name="transcript_markers", - description="Add a marker at the start of every transcribed segment (sentence-level), using each media file's local Whisper transcript. Maps each segment's source-media timestamp to its correct timeline position per clip, so it stays accurate across multiple clips/trims — unlike import_transcript_markers (plain timestamp text) or import_srt_markers (a caption track already synced to the whole export). Uses each media file's _transcript.json (auto-transcribes if missing). Non-destructive: writes a _transcript_markers copy.", - inputSchema={ - "type": "object", - "properties": { - "filepath": {"type": "string", "description": "Path to FCPXML file"}, - "clip_name": {"type": "string", "description": "Only mark the clip with this name"}, - "marker_type": {"type": "string", "default": "chapter", "description": "Marker type: standard, chapter, todo, completed"}, - "max_label_length": {"type": "integer", "default": 50, "description": "Truncate marker labels to this many characters (0 = no truncation)"}, - "model": {"type": "string", "default": "base", "description": "Whisper model size if transcription is needed"}, - "output_path": {"type": "string", "description": "Output path (default: adds _transcript_markers suffix)"}, - }, - "required": ["filepath"] - } - ), - Tool( - name="generate_dynamic_subtitles", - description="Generate progressive-composition subtitles as real, editable FCPXML title clips (the 'Text'/Basic Text template). Whisper's segments become sentences; each sentence is diagrammed as stacked blocks — supporting words grouped small in a grotesque, the sentence's key word alone and large in a display italic, body lines staggered to opposite edges. One <title> per block: each enters as its own words are spoken and stays on screen, so the sentence assembles itself, and every block clears at the same instant. Set granularity='word' for the older one-title-per-word rhythm. A sentence too tall for the band splits into successive compositions. These are TITLES, not captions: no subtitles role, so they render over the video without enabling caption display. Uses each media file's local Whisper word-level transcript (_transcript.json, auto-transcribes if missing). Non-destructive: writes a _dynamic_subtitles copy.", - inputSchema={ - "type": "object", - "properties": { - "filepath": {"type": "string", "description": "Path to FCPXML file"}, - "clip_name": {"type": "string", "description": "Only caption the clip with this name (default: all spine clips with matched source media)"}, - "model": {"type": "string", "default": "base", "description": "Whisper model size if transcription is needed"}, - "language": {"type": "string", "description": "ISO language code hint (e.g. 'en'); auto-detected if omitted"}, - "band_height": {"type": "number", "default": 0.22, "description": "Fraction of frame height the sentence block may fill before splitting into another block (default 0.22 — about three lines)"}, - "block_center_y": {"type": "number", "default": -167.0, "description": "Vertical centre of the block in canvas points; negative sits below frame centre (default -167, just under centre)"}, - "granularity": {"type": "string", "enum": ["phrase", "word"], "default": "phrase", "description": "'phrase': one title per LINE of the composition, key word set large (the reference look). 'word': one title per word."}, - "emphasis_font": {"type": "string", "default": "Playfair Display", "description": "Family for the key word (phrase mode). Must be installed on the editing Mac; unmeasured families fall back to estimated widths"}, - "emphasis_face": {"type": "string", "default": "Medium Italic", "description": "Face for the key word, e.g. 'Medium Italic' or a script/calligraphic face"}, - "emphasis_size": {"type": "integer", "default": 230, "description": "Key-word size in canvas points, at the 2160x3840 reference frame"}, - "font": {"type": "string", "default": "Helvetica Neue", "description": "Title font family (supporting lines in phrase mode)"}, - "font_size": {"type": "integer", "default": 88, "description": "Supporting-line font size in canvas points, at the 2160x3840 reference frame"}, - "active_color": {"type": "string", "default": "1 1 1 1", "description": "RGBA (0-1, space-separated) for even-indexed lines"}, - "inactive_color": {"type": "string", "default": "0.7 0.7 0.7 1", "description": "RGBA (0-1, space-separated) for odd-indexed lines — alternates with active_color for visual variety between stacked lines"}, - "output_path": {"type": "string", "description": "Output path (default: adds _dynamic_subtitles suffix)"}, - }, - "required": ["filepath"] - } - ), - - # ===== SILENCE DETECTION (v0.5.0) ===== - Tool( - name="detect_silence_candidates", - description="Detect potential silence/dead air using timeline heuristics (gaps, ultra-short clips, name patterns, duration anomalies)", - inputSchema={ - "type": "object", - "properties": { - "filepath": {"type": "string", "description": "Path to FCPXML file"}, - "min_gap_seconds": {"type": "number", "default": 0.5, "description": "Minimum gap duration to flag"}, - "patterns": {"type": "array", "items": {"type": "string"}, "description": "Name patterns to match (default: gap, silence, room tone)"}, - }, - "required": ["filepath"] - } - ), - Tool( - name="remove_silence_candidates", - description="Remove or mark detected silence candidates from timeline", - inputSchema={ - "type": "object", - "properties": { - "filepath": {"type": "string", "description": "Path to FCPXML file"}, - "mode": {"type": "string", "enum": ["delete", "mark"], "default": "mark", "description": "delete=remove clips/gaps, mark=add red markers"}, - "min_gap_seconds": {"type": "number", "default": 0.5}, - "min_confidence": {"type": "number", "default": 0.7, "description": "Only act on candidates above this confidence"}, - "output_path": {"type": "string", "description": "Output path (default: adds _silence_cleaned suffix)"} - }, - "required": ["filepath"] - } - ), - - # ===== NLE EXPORT (v0.5.0) ===== - Tool( - name="export_resolve_xml", - description="Export timeline as DaVinci Resolve compatible FCPXML (simplified v1.9)", - inputSchema={ - "type": "object", - "properties": { - "filepath": {"type": "string", "description": "Path to FCPXML file"}, - "flatten_compounds": {"type": "boolean", "default": True, "description": "Flatten compound clips for compatibility"}, - "output_path": {"type": "string", "description": "Output path (default: adds _resolve suffix)"}, - }, - "required": ["filepath"] - } - ), - Tool( - name="export_fcp7_xml", - description="Export timeline as FCP7 XML (XMEML) for Premiere Pro, DaVinci Resolve, and Avid compatibility", - inputSchema={ - "type": "object", - "properties": { - "filepath": {"type": "string", "description": "Path to FCPXML file"}, - "output_path": {"type": "string", "description": "Output path (default: adds _fcp7.xml suffix)"}, - }, - "required": ["filepath"] - } - ), - - # ===== v0.6.0 TOOLS ===== - Tool( - name="list_effects", - description="List all available FCP transition effects with slugs and UUIDs", - inputSchema={ - "type": "object", - "properties": {}, - } - ), - Tool( - name="add_audio", - description="Add an audio clip or music bed to the timeline", - inputSchema={ - "type": "object", - "properties": { - "filepath": {"type": "string", "description": "Path to FCPXML file"}, - "parent_clip_id": {"type": "string", "description": "Clip to attach audio to (omit for music bed spanning full timeline)"}, - "asset_id": {"type": "string", "description": "Existing asset reference ID"}, - "src": {"type": "string", "description": "Path to audio file (creates new asset)"}, - "offset": {"type": "string", "description": "Position relative to parent clip start", "default": "0s"}, - "duration": {"type": "string", "description": "Duration of audio clip"}, - "role": {"type": "string", "description": "Audio role (dialogue, music, effects, etc.)", "default": "dialogue"}, - "lane": {"type": "integer", "description": "Lane number (negative = below)", "default": -1}, - "output_path": {"type": "string", "description": "Output path"}, - }, - "required": ["filepath"] - } - ), - Tool( - name="create_compound_clip", - description="Group spine clips into a compound clip", - inputSchema={ - "type": "object", - "properties": { - "filepath": {"type": "string", "description": "Path to FCPXML file"}, - "clip_ids": {"type": "array", "items": {"type": "string"}, "description": "Clip IDs to group"}, - "name": {"type": "string", "description": "Name for the compound clip", "default": "Compound Clip"}, - "output_path": {"type": "string", "description": "Output path"}, - }, - "required": ["filepath", "clip_ids"] - } - ), - Tool( - name="flatten_compound_clip", - description="Flatten a compound clip back into individual clips in the spine", - inputSchema={ - "type": "object", - "properties": { - "filepath": {"type": "string", "description": "Path to FCPXML file"}, - "ref_clip_id": {"type": "string", "description": "ID of the ref-clip to flatten"}, - "output_path": {"type": "string", "description": "Output path"}, - }, - "required": ["filepath", "ref_clip_id"] - } - ), - Tool( - name="list_templates", - description="List available timeline templates with slot definitions", - inputSchema={ - "type": "object", - "properties": {}, - } - ), - Tool( - name="apply_template", - description="Fill a timeline template with clips and generate FCPXML", - inputSchema={ - "type": "object", - "properties": { - "template_name": {"type": "string", "description": "Template name (intro_outro, lower_thirds, music_video)"}, - "clips": {"type": "object", "description": "Map of slot_name -> {src, name, duration} or {asset_id, name, duration}"}, - "output_path": {"type": "string", "description": "Output FCPXML path"}, - "fps": {"type": "number", "description": "Frame rate", "default": 24}, - }, - "required": ["template_name", "clips", "output_path"] - } - ), - - # ===== v0.8.0 TOOLS ===== - Tool( - name="relink_media", - description="Bulk-rewrite media source paths (asset/media-rep src URLs) to relink moved or renamed media folders without opening FCP. Prefix-based: find='/Volumes/OldDrive/Media' replace='/Volumes/NewDrive/Media'. Use dry_run to preview.", - inputSchema={ - "type": "object", - "properties": { - "filepath": {"type": "string", "description": "Path to FCPXML file or .fcpxmld bundle"}, - "find": {"type": "string", "description": "Old path prefix to match (plain path or file:// URL)"}, - "replace": {"type": "string", "description": "New path prefix to substitute"}, - "dry_run": {"type": "boolean", "description": "Preview changes without writing", "default": False}, - "output_path": {"type": "string", "description": "Output path (default: adds _relinked suffix)"}, - }, - "required": ["filepath", "find", "replace"] - } - ), - - # ===== v0.9.0 LIVE MODE (macOS + Final Cut Pro required) ===== - Tool( - name="push_to_fcp", - description="LIVE: send an FCPXML file into the running Final Cut Pro with zero clicks (official Open Document Apple event). Creates/targets a library via import-options. Launches FCP if needed. macOS-only; first use triggers an Automation permission prompt. For true zero-click, pass a library_location ending in .fcpbundle (a new path is auto-created); omitting it makes FCP show a modal library picker.", - inputSchema={ - "type": "object", - "properties": { - "filepath": {"type": "string", "description": "Path to FCPXML file or .fcpxmld bundle to import"}, - "library_location": {"type": "string", "description": "Target .fcpbundle library path (auto-created if it doesn't exist; the extension is normalized to .fcpbundle). Omit to import into the active library, but note FCP then shows a modal 'Open Library' picker that blocks until answered"}, - "suppress_warnings": {"type": "boolean", "description": "Suppress non-fatal import warning dialogs", "default": True}, - "copy_assets": {"type": "boolean", "description": "Copy media into the library (true) or link in place (false). Omit for FCP default"}, - }, - "required": ["filepath"] - } - ), - Tool( - name="list_fcp_libraries", - description="LIVE: enumerate the running Final Cut Pro's open libraries, events, and projects via Apple's read-only scripting dictionary. Refuses to launch FCP unless allow_launch is true. macOS-only.", - inputSchema={ - "type": "object", - "properties": { - "allow_launch": {"type": "boolean", "description": "Launch FCP if it isn't running", "default": False}, - }, - } - ), - ] - - -# ============================================================================ -# QC DETECTION HELPERS — Pure detection logic, reusable across handlers -# ============================================================================ - - -def _detect_flash_frames( - tl: Any, *, critical_threshold: int = 2, warning_threshold: int = 6, -) -> list: - """Find clips shorter than *warning_threshold* frames. - - Returns a list of ``FlashFrame`` objects sorted by severity. Shared by - ``handle_detect_flash_frames`` and ``handle_validate_timeline`` so the - detection logic lives in exactly one place. - """ - fps = tl.frame_rate - flash_frames: list[FlashFrame] = [] - for clip in tl.clips: - duration_frames = int(clip.duration_seconds * fps) - if duration_frames < warning_threshold: - severity = ( - FlashFrameSeverity.CRITICAL - if duration_frames < critical_threshold - else FlashFrameSeverity.WARNING - ) - flash_frames.append(FlashFrame( - clip_name=clip.name, clip_id=clip.name, - start=clip.start, duration_frames=duration_frames, - duration_seconds=clip.duration_seconds, severity=severity, - )) - return flash_frames - - -def _detect_gaps(tl: Any, *, min_gap_frames: int = 1) -> list: - """Find inter-clip gaps of at least *min_gap_frames* length. - - Returns a list of ``GapInfo`` objects. Shared by ``handle_detect_gaps`` - and ``handle_validate_timeline``. - """ - fps = tl.frame_rate - min_gap_seconds = min_gap_frames / fps - gaps: list[GapInfo] = [] - sorted_clips = sorted(tl.clips, key=lambda c: c.start.seconds) - for i in range(len(sorted_clips) - 1): - current_end = sorted_clips[i].end.seconds - next_start = sorted_clips[i + 1].start.seconds - gap_duration = next_start - current_end - if gap_duration >= min_gap_seconds: - gaps.append(GapInfo( - start=Timecode(frames=int(current_end * fps), frame_rate=fps), - duration_frames=int(gap_duration * fps), - duration_seconds=gap_duration, - previous_clip=sorted_clips[i].name, - next_clip=sorted_clips[i + 1].name, - )) - return gaps - - -def _detect_duplicate_groups(tl: Any, *, mode: str = "same_source") -> list: - """Group clips that share a source media reference. - - Returns a list of ``DuplicateGroup`` objects. Shared by - ``handle_detect_duplicates`` and ``handle_validate_timeline``. - """ - source_groups: dict[str, list[dict]] = {} - for clip in tl.clips: - source_key = clip.media_path or clip.name - if source_key not in source_groups: - source_groups[source_key] = [] - source_groups[source_key].append({ - 'name': clip.name, - 'start': clip.start.seconds, - 'duration': clip.duration_seconds, - 'source_start': clip.source_start.seconds if clip.source_start else 0, - 'source_duration': clip.duration_seconds, - 'timecode': format_timecode(clip.start), - }) - - duplicates: list[DuplicateGroup] = [] - for source_key, clips in source_groups.items(): - if len(clips) <= 1: - continue - group = DuplicateGroup( - source_ref=source_key, - source_name=source_key.split('/')[-1] if '/' in source_key else source_key, - clips=clips, - ) - if mode == "same_source": - duplicates.append(group) - elif mode == "overlapping_ranges" and group.has_overlapping_ranges: - duplicates.append(group) - elif mode == "identical": - seen_ranges: set[tuple] = set() - identical_clips = [] - for c in clips: - range_key = (c['source_start'], c['source_duration']) - if range_key in seen_ranges: - identical_clips.append(c) - seen_ranges.add(range_key) - if identical_clips: - group.clips = identical_clips - duplicates.append(group) - return duplicates - - -# ============================================================================ -# TOOL HANDLERS — Each tool gets its own function -# ============================================================================ - -# ----- READ HANDLERS ----- - -async def handle_list_projects(arguments: dict) -> Sequence[TextContent]: - directory = arguments.get("directory", PROJECTS_DIR) - resolved_dir = _validate_directory( - directory, allowed_root=PROJECTS_DIR if _SANDBOX_ENABLED else None - ) - files = find_fcpxml_files(resolved_dir) - if not files: - return _text_result(f"No FCPXML files found in {directory}") - return _text_result(f"Found {len(files)} FCPXML file(s):\n" + "\n".join(f" - {f}" for f in files)) - - -async def handle_analyze_timeline(arguments: dict) -> Sequence[TextContent]: - project, tl = _require_timeline(arguments["filepath"]) - durs = [c.duration_seconds for c in tl.clips] - avg, med, mn, mx = (0, 0, 0, 0) if not durs else ( - sum(durs)/len(durs), sorted(durs)[len(durs)//2], min(durs), max(durs)) - return _text_result(f"""# Timeline Analysis: {tl.name} - -## Overview -- **Duration**: {format_duration(tl.duration.seconds)} -- **Resolution**: {tl.width}x{tl.height} @ {tl.frame_rate}fps - -## Clip Statistics -- **Total Clips**: {tl.total_clips} -- **Total Cuts**: {tl.total_cuts} -- **Transitions**: {len(tl.transitions)} - -## Pacing -- **Average**: {format_duration(avg)} -- **Median**: {format_duration(med)} -- **Shortest**: {format_duration(mn)} -- **Longest**: {format_duration(mx)} -- **Cuts/Minute**: {tl.cuts_per_minute:.1f} - -## Markers -- **Total**: {len(tl.markers)} -- **Chapters**: {len([m for m in tl.markers if m.marker_type == MarkerType.CHAPTER])} -""") - - -async def handle_list_clips(arguments: dict) -> Sequence[TextContent]: - project, tl = _require_timeline(arguments["filepath"]) - limit = arguments.get("limit") - clips = tl.clips[:limit] if limit else tl.clips - result = f"# Clips in {tl.name}\n\n| # | Name | Start | Duration | Keywords |\n|---|------|-------|----------|----------|\n" - for i, c in enumerate(clips, 1): - kws = ", ".join(k.value for k in c.keywords) if c.keywords else "-" - result += f"| {i} | {c.name} | {format_timecode(c.start)} | {format_duration(c.duration_seconds)} | {kws} |\n" - return _text_result(result) - - -async def handle_list_markers(arguments: dict) -> Sequence[TextContent]: - project, tl = _require_timeline(arguments["filepath"]) - markers = list(tl.markers) - for clip in tl.clips: - markers.extend(clip.markers) - marker_type = arguments.get("marker_type", "all") - if marker_type != "all": - markers = [m for m in markers if m.marker_type == MarkerType.from_string(marker_type)] - markers.sort(key=lambda m: m.start.frames) - fmt = arguments.get("format", "detailed") - if fmt == "youtube": - result = "# YouTube Chapters\n\n" + "\n".join(f"{m.to_youtube_timestamp()} {m.name}" for m in markers) - elif fmt == "simple": - result = "\n".join(f"{format_timecode(m.start)} - {m.name}" for m in markers) - else: - result = f"# Markers ({len(markers)})\n\n| TC | Name | Type |\n|---|------|------|\n" - result += "\n".join(f"| {format_timecode(m.start)} | {m.name} | {m.marker_type.value} |" for m in markers) - return _text_result(result) - - -async def handle_find_short_cuts(arguments: dict) -> Sequence[TextContent]: - project, tl = _require_timeline(arguments["filepath"]) - threshold = arguments.get("threshold_seconds", 0.5) - short = tl.get_clips_shorter_than(threshold) - if not short: - return _text_result(f"No clips shorter than {threshold}s") - return _text_result(_format_clip_table( - short, f"# Short Clips (< {threshold}s) - {len(short)} found", - )) - - -async def handle_find_long_clips(arguments: dict) -> Sequence[TextContent]: - project, tl = _require_timeline(arguments["filepath"]) - threshold = arguments.get("threshold_seconds", 10.0) - long = tl.get_clips_longer_than(threshold) - if not long: - return _text_result(f"No clips longer than {threshold}s") - return _text_result(_format_clip_table( - long, f"# Long Clips (> {threshold}s) - {len(long)} found", - )) - - -async def handle_list_keywords(arguments: dict) -> Sequence[TextContent]: - project, tl = _require_timeline(arguments["filepath"]) - keywords = {} - for clip in tl.clips: - for kw in clip.keywords: - keywords.setdefault(kw.value, []).append(clip.name) - if not keywords: - return _text_result("No keywords found") - result = f"# Keywords ({len(keywords)})\n\n" - for kw, clips in sorted(keywords.items()): - result += f"**{kw}** ({len(clips)} clips)\n" - return _text_result(result) - - -async def handle_export_edl(arguments: dict) -> Sequence[TextContent]: - project, tl = _require_timeline(arguments["filepath"]) - edl = f"TITLE: {tl.name}\nFCM: NON-DROP FRAME\n\n" - for i, c in enumerate(tl.clips, 1): - edl += f"{i:03d} AX V C {format_timecode(c.source_start)} {format_timecode(c.end)} {format_timecode(c.start)} {format_timecode(c.end)}\n" - edl += f"* FROM CLIP NAME: {c.name}\n\n" - return _text_result(f"```edl\n{edl}```") - - -async def handle_export_csv(arguments: dict) -> Sequence[TextContent]: - project, tl = _require_timeline(arguments["filepath"]) - csv = "Name,Start,End,Duration,Keywords\n" - for c in tl.clips: - kws = "|".join(k.value for k in c.keywords) - csv += f'"{c.name}",{format_timecode(c.start)},{format_timecode(c.end)},{c.duration_seconds:.3f},"{kws}"\n' - return _text_result(f"```csv\n{csv}```") - - -async def handle_analyze_pacing(arguments: dict) -> Sequence[TextContent]: - project, tl = _require_timeline(arguments["filepath"]) - if not tl.clips: - return _text_result("No clips to analyze") - durs = [c.duration_seconds for c in tl.clips] - avg = sum(durs) / len(durs) - q_len = len(durs) // 4 or 1 - segments = [durs[i:i+q_len] for i in range(0, len(durs), q_len)][:4] - seg_avgs = [sum(s)/len(s) if s else 0 for s in segments] - suggestions = [] - flash = [c for c in tl.clips if c.duration_seconds < 0.2] - if flash: - suggestions.append(f" {len(flash)} potential flash frames (< 0.2s)") - long = [c for c in tl.clips if c.duration_seconds > 30] - if long: - suggestions.append(f" {len(long)} long takes (> 30s) - consider trimming") - if len(seg_avgs) >= 4 and seg_avgs[3] < seg_avgs[0] * 0.7: - suggestions.append(" Pacing accelerates toward end - good for building energy") - elif len(seg_avgs) >= 4 and seg_avgs[3] > seg_avgs[0] * 1.3: - suggestions.append(" Pacing slows toward end - consider tightening") - return _text_result(f"""# Pacing Analysis: {tl.name} - -## Overall -- **Avg Cut**: {format_duration(avg)} -- **Cuts/Min**: {tl.cuts_per_minute:.1f} - -## By Section -| Q1 | Q2 | Q3 | Q4 | -|----|----|----|----| -| {format_duration(seg_avgs[0]) if len(seg_avgs) > 0 else 'N/A'} | {format_duration(seg_avgs[1]) if len(seg_avgs) > 1 else 'N/A'} | {format_duration(seg_avgs[2]) if len(seg_avgs) > 2 else 'N/A'} | {format_duration(seg_avgs[3]) if len(seg_avgs) > 3 else 'N/A'} | - -## Suggestions -{_fmt_suggestions(suggestions)} -""") - - -async def handle_list_library_clips(arguments: dict) -> Sequence[TextContent]: - filepath = _validate_filepath(arguments["filepath"], ('.fcpxml', '.fcpxmld')) - parser = FCPXMLParser() - parser.parse_file(filepath) - keywords = arguments.get("keywords") - library_clips = parser.get_library_clips(keywords=keywords) - limit = arguments.get("limit") - if limit: - library_clips = library_clips[:limit] - if not library_clips: - return _text_result("No library clips found") - result = f"# Library Clips ({len(library_clips)} available)\n\n" - result += "| ID | Name | Duration | Has Video | Has Audio |\n" - result += "|----|------|----------|-----------|----------|\n" - for c in library_clips: - result += f"| {c['asset_id']} | {c['name']} | {format_duration(c['duration_seconds'])} | {'Y' if c['has_video'] else 'N'} | {'Y' if c['has_audio'] else 'N'} |\n" - result += "\n*Use `insert_clip` to add these to your timeline.*" - return _text_result(result) - - -# ----- QC / VALIDATION HANDLERS ----- - -async def handle_detect_flash_frames(arguments: dict) -> Sequence[TextContent]: - project, tl = _require_timeline(arguments["filepath"]) - critical_threshold = arguments.get("critical_threshold_frames", 2) - warning_threshold = arguments.get("warning_threshold_frames", 6) - - flash_frames = _detect_flash_frames( - tl, critical_threshold=critical_threshold, warning_threshold=warning_threshold, - ) - - if not flash_frames: - return _text_result(f"No flash frames detected (threshold: {warning_threshold} frames)") - - critical = [f for f in flash_frames if f.severity == FlashFrameSeverity.CRITICAL] - warnings = [f for f in flash_frames if f.severity == FlashFrameSeverity.WARNING] - - result = f"""# Flash Frame Detection - -## Summary -- **Critical** (< {critical_threshold} frames): {len(critical)} found -- **Warning** (< {warning_threshold} frames): {len(warnings)} found -- **Total**: {len(flash_frames)} flash frames - -## Critical Flash Frames -""" - flash_headers = ["Clip", "Timecode", "Frames", "Duration"] - if critical: - result += _markdown_table(flash_headers, [ - [f.clip_name, format_timecode(f.start), f"{f.duration_frames}f", format_duration(f.duration_seconds)] - for f in critical - ]) + "\n" - else: - result += "_None_\n" - - result += "\n## Warning Flash Frames\n" - if warnings: - result += _markdown_table(flash_headers, [ - [f.clip_name, format_timecode(f.start), f"{f.duration_frames}f", format_duration(f.duration_seconds)] - for f in warnings - ]) + "\n" - else: - result += "_None_\n" - - result += "\n*Use `fix_flash_frames` to automatically resolve these issues.*" - return _text_result(result) - - -async def handle_detect_duplicates(arguments: dict) -> Sequence[TextContent]: - project, tl = _require_timeline(arguments["filepath"]) - mode = arguments.get("mode", "same_source") - - duplicates = _detect_duplicate_groups(tl, mode=mode) - - if not duplicates: - return _text_result(f"No duplicate clips found (mode: {mode})") - - result = f"""# Duplicate Clip Detection - -## Summary -- **Mode**: {mode} -- **Duplicate Groups**: {len(duplicates)} -- **Total Duplicate Clips**: {sum(g.count for g in duplicates)} - -## Duplicate Groups -""" - for group in duplicates: - result += f"\n### {group.source_name} ({group.count} uses)\n" - result += "| Clip Name | Timeline Position | Duration |\n|-----------|-------------------|----------|\n" - for c in group.clips: - result += f"| {c['name']} | {c['timecode']} | {format_duration(c['duration'])} |\n" - - return _text_result(result) - - -async def handle_detect_gaps(arguments: dict) -> Sequence[TextContent]: - project, tl = _require_timeline(arguments["filepath"]) - min_gap_frames = arguments.get("min_gap_frames", 1) - - gaps = _detect_gaps(tl, min_gap_frames=min_gap_frames) - - if not gaps: - return _text_result(f"No gaps detected (minimum: {min_gap_frames} frame(s))") - - result = f"""# Gap Detection - -## Summary -- **Gaps Found**: {len(gaps)} -- **Total Gap Duration**: {format_duration(sum(g.duration_seconds for g in gaps))} -- **Minimum Detection**: {min_gap_frames} frame(s) - -## Gaps -""" - result += _markdown_table( - ["Position", "Duration", "Between"], - [[gap.timecode, f"{gap.duration_frames}f ({format_duration(gap.duration_seconds)})", - f"{gap.previous_clip} -> {gap.next_clip}"] for gap in gaps], - ) + "\n" - - result += "\n*Use `fill_gaps` to automatically close these gaps.*" - return _text_result(result) - - -# ----- WRITE HANDLERS ----- - -async def handle_add_marker(arguments: dict) -> Sequence[TextContent]: - filepath, output_path, modifier = _setup_modifier(arguments) - marker_type = MarkerType.from_string(arguments.get("marker_type", "standard")) - modifier.add_marker_at_timeline( - timecode=arguments["timecode"], name=arguments["name"], - marker_type=marker_type, note=arguments.get("note"), - ) - modifier.save(output_path) - return _text_result(f"Added marker '{arguments['name']}' at {arguments['timecode']}\n\nSaved to: {output_path}") - - -async def handle_batch_add_markers(arguments: dict) -> Sequence[TextContent]: - filepath, output_path, modifier = _setup_modifier(arguments) - markers_added = modifier.batch_add_markers( - markers=arguments.get("markers", []), - auto_at_cuts=arguments.get("auto_at_cuts", False), - auto_at_intervals=arguments.get("auto_at_intervals"), - ) - modifier.save(output_path) - return _text_result(f"Added {len(markers_added)} markers\n\nSaved to: {output_path}") - - -async def handle_trim_clip(arguments: dict) -> Sequence[TextContent]: - filepath, output_path, modifier = _setup_modifier(arguments) - modifier.trim_clip( - clip_id=arguments["clip_id"], - trim_start=arguments.get("trim_start"), - trim_end=arguments.get("trim_end"), - ripple=arguments.get("ripple", True), - ) - modifier.save(output_path) - return _text_result(f"Trimmed clip '{arguments['clip_id']}'\n\nSaved to: {output_path}") - - -async def handle_reorder_clips(arguments: dict) -> Sequence[TextContent]: - filepath, output_path, modifier = _setup_modifier(arguments) - modifier.reorder_clips( - clip_ids=arguments["clip_ids"], - target_position=arguments["target_position"], - ripple=arguments.get("ripple", True), - ) - modifier.save(output_path) - clips_moved = ", ".join(arguments["clip_ids"]) - return _text_result(f"Moved clips [{clips_moved}] to {arguments['target_position']}\n\nSaved to: {output_path}") - - -async def handle_add_transition(arguments: dict) -> Sequence[TextContent]: - filepath, output_path, modifier = _setup_modifier(arguments) - modifier.add_transition( - clip_id=arguments["clip_id"], - position=arguments.get("position", "end"), - transition_type=arguments.get("transition_type", "cross-dissolve"), - duration=arguments.get("duration", "00:00:00:15"), - ) - modifier.save(output_path) - return _text_result(f"Added {arguments.get('transition_type', 'cross-dissolve')} to '{arguments['clip_id']}'\n\nSaved to: {output_path}") - - -async def handle_change_speed(arguments: dict) -> Sequence[TextContent]: - speed = arguments["speed"] - if not isinstance(speed, (int, float)) or speed <= 0 or speed > 100: - raise ValueError( - f"Speed must be a positive number between 0 (exclusive) and 100, got {speed!r}" - ) - filepath, output_path, modifier = _setup_modifier(arguments) - modifier.change_speed( - clip_id=arguments["clip_id"], - speed=speed, - preserve_pitch=arguments.get("preserve_pitch", True), - ) - modifier.save(output_path) - speed_desc = f"{speed}x" if speed >= 1 else f"{int(1/speed)}x slow motion" - return _text_result(f"Changed speed of '{arguments['clip_id']}' to {speed_desc}\n\nSaved to: {output_path}") - - -async def handle_add_zoom(arguments: dict) -> Sequence[TextContent]: - start = float(arguments["start"]) - end = float(arguments["end"]) - scale = float(arguments.get("scale", 1.3)) - ease = float(arguments.get("ease", 0.3)) - position = arguments.get("position", "0 0") - - filepath, output_path, modifier = _setup_modifier(arguments) - modifier.add_zoom( - clip_id=arguments["clip_id"], start=start, end=end, - scale=scale, ease=ease, position=position, - ) - modifier.save(output_path) - return _text_result( - f"# Zoom Added\n\n" - f"- **Clip**: {arguments['clip_id']}\n" - f"- **Window**: {start}s → {end}s (clip-relative)\n" - f"- **Scale**: {int(scale * 100)}%\n" - f"- **Ease**: {ease}s in/out\n\n" - f"Saved to: {output_path}" - ) - - -async def handle_delete_clips(arguments: dict) -> Sequence[TextContent]: - filepath, output_path, modifier = _setup_modifier(arguments) - modifier.delete_clip( - clip_ids=arguments["clip_ids"], - ripple=arguments.get("ripple", True), - ) - modifier.save(output_path) - return _text_result(f"Deleted {len(arguments['clip_ids'])} clip(s)\n\nSaved to: {output_path}") - - -async def handle_split_clip(arguments: dict) -> Sequence[TextContent]: - filepath, output_path, modifier = _setup_modifier(arguments) - new_clips = modifier.split_clip( - clip_id=arguments["clip_id"], - split_points=arguments["split_points"], - ) - modifier.save(output_path) - return _text_result(f"Split '{arguments['clip_id']}' into {len(new_clips)} clips\n\nSaved to: {output_path}") - - -async def handle_insert_clip(arguments: dict) -> Sequence[TextContent]: - filepath, output_path, modifier = _setup_modifier(arguments) - new_clip = modifier.insert_clip( - asset_id=arguments.get("asset_id"), - asset_name=arguments.get("asset_name"), - position=arguments["position"], - duration=arguments.get("duration"), - in_point=arguments.get("in_point"), - out_point=arguments.get("out_point"), - ripple=arguments.get("ripple", True), - ) - modifier.save(output_path) - clip_name = new_clip.get('name', 'Unknown') - pos = arguments["position"] - return _text_result(f"Inserted '{clip_name}' at position '{pos}'\n\nSaved to: {output_path}") - - -# ----- BATCH FIX HANDLERS ----- - -async def handle_fix_flash_frames(arguments: dict) -> Sequence[TextContent]: - filepath, output_path, modifier = _setup_modifier(arguments, "_flash_fixed") - fixed = modifier.fix_flash_frames( - mode=arguments.get("mode", "auto"), - threshold_frames=arguments.get("threshold_frames", 6), - ) - modifier.save(output_path) - - if not fixed: - return _text_result("No flash frames found to fix.") - - result = _format_batch_result( - title="Flash Frames Fixed", - summary={"Fixed": f"{len(fixed)} flash frames", "Mode": arguments.get('mode', 'auto')}, - headers=["Clip", "Frames", "Action", "Result"], - rows=[ - [f['clip_name'], f"{f['duration_frames']}f", f['action'], f"Extended: {f.get('extended_clip', 'N/A')}"] - for f in fixed - ], - output_path=output_path, - ) - return _text_result(result) - - -async def handle_rapid_trim(arguments: dict) -> Sequence[TextContent]: - filepath, output_path, modifier = _setup_modifier(arguments, "_rapid_trim") - trimmed = modifier.rapid_trim( - max_duration=arguments["max_duration"], - min_duration=arguments.get("min_duration"), - keywords=arguments.get("keywords"), - trim_from=arguments.get("trim_from", "end"), - ) - modifier.save(output_path) - - if not trimmed: - return _text_result(f"No clips exceeded {arguments['max_duration']} - nothing trimmed.") - - total_before = sum(t['original_duration'] for t in trimmed) - total_after = sum(t['new_duration'] for t in trimmed) - - result = _format_batch_result( - title="Rapid Trim Complete", - summary={ - "Clips Trimmed": str(len(trimmed)), - "Max Duration": str(arguments['max_duration']), - "Trim From": arguments.get('trim_from', 'end'), - "Time Saved": format_duration(total_before - total_after), - }, - headers=["Clip", "Before", "After"], - rows=[ - [t['clip_name'], format_duration(t['original_duration']), format_duration(t['new_duration'])] - for t in trimmed - ], - output_path=output_path, - ) - return _text_result(result) - - -async def handle_fill_gaps(arguments: dict) -> Sequence[TextContent]: - filepath, output_path, modifier = _setup_modifier(arguments, "_gaps_filled") - filled = modifier.fill_gaps( - mode=arguments.get("mode", "extend_previous"), - max_gap=arguments.get("max_gap"), - ) - modifier.save(output_path) - - if not filled: - return _text_result("No gaps found to fill.") - - result = _format_batch_result( - title="Gaps Filled", - summary={"Gaps Filled": str(len(filled)), "Mode": arguments.get('mode', 'extend_previous')}, - headers=["Position", "Duration", "Action"], - rows=[[g['timecode'], f"{g['duration_frames']}f", g['action']] for g in filled], - output_path=output_path, - ) - return _text_result(result) - - -async def handle_validate_timeline(arguments: dict) -> Sequence[TextContent]: - project, tl = _require_timeline(arguments["filepath"]) - checks = arguments.get("checks", ["all"]) - run_all = "all" in checks - - issues: list[str] = [] - flash_count = 0 - gap_count = 0 - duplicate_count = 0 - - if run_all or "flash_frames" in checks: - flashes = _detect_flash_frames(tl) - flash_count = len(flashes) - for f in flashes: - severity = "error" if f.severity == FlashFrameSeverity.CRITICAL else "warning" - issues.append( - f"- [{severity.upper()}] Flash frame: {f.clip_name} " - f"({f.duration_frames}f) at {format_timecode(f.start)}" - ) - - if run_all or "gaps" in checks: - detected_gaps = _detect_gaps(tl) - gap_count = len(detected_gaps) - for g in detected_gaps: - issues.append(f"- [WARNING] Gap: {g.duration_frames}f at {g.timecode}") - - if run_all or "duplicates" in checks: - dup_groups = _detect_duplicate_groups(tl) - for group in dup_groups: - duplicate_count += group.count - issues.append( - f"- [INFO] Duplicate source: {group.source_name} ({group.count} uses)" - ) - - error_weight = 10 - warning_weight = 3 - info_weight = 1 - errors = len([i for i in issues if "[ERROR]" in i]) - warnings = len([i for i in issues if "[WARNING]" in i]) - infos = len([i for i in issues if "[INFO]" in i]) - penalty = (errors * error_weight) + (warnings * warning_weight) + (infos * info_weight) - health_score = max(0, 100 - penalty) - - result = f"""# Timeline Validation: {tl.name} - -## Health Score: {health_score}% - -## Summary -| Check | Count | Status | -|-------|-------|--------| -| Flash Frames | {flash_count} | {'PASS' if flash_count == 0 else 'FAIL'} | -| Gaps | {gap_count} | {'PASS' if gap_count == 0 else 'WARN'} | -| Duplicate Sources | {duplicate_count} | {'PASS' if duplicate_count == 0 else 'INFO'} | - -## Issues ({len(issues)}) -""" - if issues: - result += "\n".join(issues[:20]) - if len(issues) > 20: - result += f"\n... and {len(issues) - 20} more issues" - else: - result += "_No issues found!_" - - result += "\n\n*Use `fix_flash_frames` and `fill_gaps` to automatically resolve issues.*" - return _text_result(result) - - -# ----- GENERATION HANDLERS ----- - -async def handle_auto_rough_cut(arguments: dict) -> Sequence[TextContent]: - filepath, output_path, generator = _setup_generator(arguments, "_roughcut") - - segments = None - if arguments.get("segments"): - segments = [ - SegmentSpec( - name=s.get("name", "Segment"), - keywords=s.get("keywords", []), - duration_seconds=s.get("duration", 0), - priority=s.get("priority", "best"), - ) - for s in arguments["segments"] - ] - result = generator.generate( - output_path=output_path, - target_duration=arguments["target_duration"], - pacing=arguments.get("pacing", "medium"), - keywords=arguments.get("keywords"), - segments=segments, - priority=arguments.get("priority", "best"), - favorites_only=arguments.get("favorites_only", False), - add_transitions=arguments.get("add_transitions", False), - ) - - return _text_result(f"""# Rough Cut Generated - -## Summary -- **Clips Used**: {result.clips_used} of {result.clips_available} available -- **Target Duration**: {format_duration(result.target_duration)} -- **Actual Duration**: {format_duration(result.actual_duration)} -- **Average Clip**: {format_duration(result.average_clip_duration)} - -## Output -Saved to: `{result.output_path}` - -**Next step**: Import this FCPXML into Final Cut Pro (File > Import > XML) -""") - - -async def handle_generate_montage(arguments: dict) -> Sequence[TextContent]: - filepath, output_path, generator = _setup_generator(arguments, "_montage") - result = generator.generate_montage( - output_path=output_path, - target_duration=arguments["target_duration"], - pacing_curve=arguments.get("pacing_curve", "accelerating"), - start_duration=arguments.get("start_duration", 2.0), - end_duration=arguments.get("end_duration", 0.5), - keywords=arguments.get("keywords"), - add_transitions=arguments.get("add_transitions", False), - ) - - curve_desc = { - 'accelerating': 'slow to fast (builds energy)', - 'decelerating': 'fast to slow (winds down)', - 'pyramid': 'slow to fast to slow (dramatic arc)', - 'constant': 'same duration throughout', - } - - return _text_result(f"""# Montage Generated - -## Summary -- **Clips Used**: {result['clips_used']} of {result['clips_available']} available -- **Target Duration**: {format_duration(result['target_duration'])} -- **Actual Duration**: {format_duration(result['actual_duration'])} -- **Pacing Curve**: {result['pacing_curve']} - {curve_desc.get(result['pacing_curve'], '')} - -## Pacing -- **Start Clip Duration**: {format_duration(result['start_clip_duration'])} -- **End Clip Duration**: {format_duration(result['end_clip_duration'])} - -## Output -Saved to: `{result['output_path']}` -""") - - -async def handle_generate_ab_roll(arguments: dict) -> Sequence[TextContent]: - filepath, output_path, generator = _setup_generator(arguments, "_ab_roll") - result = generator.generate_ab_roll( - output_path=output_path, - target_duration=arguments["target_duration"], - a_keywords=arguments["a_keywords"], - b_keywords=arguments["b_keywords"], - a_duration=arguments.get("a_duration", "5s"), - b_duration=arguments.get("b_duration", "3s"), - start_with=arguments.get("start_with", "a"), - add_transitions=arguments.get("add_transitions", True), - ) - - return _text_result(f"""# A/B Roll Edit Generated - -## Summary -- **A-Roll Segments**: {result['a_segments']} (from {result['a_clips_available']} available) -- **B-Roll Segments**: {result['b_segments']} (from {result['b_clips_available']} available) -- **Total Clips**: {result['clips_used']} - -## Timing -- **Target Duration**: {format_duration(result['target_duration'])} -- **Actual Duration**: {format_duration(result['actual_duration'])} -- **A-Roll Duration**: {result['a_duration_setting']} per segment -- **B-Roll Duration**: {result['b_duration_setting']} per cutaway - -## Output -Saved to: `{result['output_path']}` - -**Next step**: Import this FCPXML into Final Cut Pro (File > Import > XML) -""") - - -# ----- BEAT SYNC HANDLERS ----- - -async def handle_import_beat_markers(arguments: dict) -> Sequence[TextContent]: - filepath, output_path = _resolve_io_paths(arguments, "_beats") - beats_path = _validate_filepath(arguments["beats_path"], ('.json',)) - - with open(beats_path, 'r') as f: - beats_data = json.load(f) - _check_json_depth(beats_data) - - beat_times = [] - if isinstance(beats_data, list): - beat_times = beats_data - elif isinstance(beats_data, dict): - beat_times = beats_data.get('beats', beats_data.get('times', beats_data.get('markers', []))) - - beat_filter = arguments.get("beat_filter", "all") - if beat_filter == "downbeat" and isinstance(beats_data, dict): - beat_times = beats_data.get('downbeats', beat_times[::4]) - elif beat_filter == "measure" and isinstance(beats_data, dict): - beat_times = beats_data.get('measures', beat_times[::4]) - - markers = [] - marker_type = arguments.get("marker_type", "standard") - for i, beat_time in enumerate(beat_times): - if isinstance(beat_time, (int, float)): - markers.append({ - 'timecode': f"{beat_time}s", - 'name': f"Beat {i+1}", - 'marker_type': marker_type.upper(), - }) - elif isinstance(beat_time, dict): - markers.append({ - 'timecode': f"{beat_time.get('time', beat_time.get('position', 0))}s", - 'name': beat_time.get('label', f"Beat {i+1}"), - 'marker_type': marker_type.upper(), - }) - - modifier = FCPXMLModifier(filepath) - - # Songs routinely run longer than the edit — beats past the timeline's - # end are skipped (add_marker_at_timeline would raise on them). - timeline_end = modifier._timeline_duration().to_seconds() - in_range = [m for m in markers if float(m['timecode'].rstrip('s')) < timeline_end] - skipped_count = len(markers) - len(in_range) - - added = modifier.batch_add_markers(markers=in_range) - modifier.save(output_path) - - skipped_note = ( - f"- **Skipped**: {skipped_count} beat(s) beyond the timeline end " - f"({format_duration(timeline_end)})\n" if skipped_count else "" - ) - return _text_result(f"""# Beat Markers Imported - -## Summary -- **Beats Found**: {len(beat_times)} -- **Markers Added**: {len(added)} -{skipped_note}- **Filter**: {beat_filter} -- **Marker Type**: {marker_type} - -## Output -Saved to: `{output_path}` - -*Use `snap_to_beats` to align your cuts to these markers.* -""") - - -async def handle_snap_to_beats(arguments: dict) -> Sequence[TextContent]: - filepath, output_path = _resolve_io_paths(arguments, "_synced") - max_shift = arguments.get("max_shift_frames", 6) - prefer = arguments.get("prefer", "nearest") - - parser = FCPXMLParser() - project = parser.parse_file(filepath) - if not project.timelines: - return _no_timeline() - - tl = project.primary_timeline - fps = tl.frame_rate - - markers = list(tl.markers) - for clip in tl.clips: - markers.extend(clip.markers) - - if not markers: - return _text_result("No markers found. Use `import_beat_markers` first.") - - marker_times = sorted([m.start.seconds for m in markers]) - - modifier = FCPXMLModifier(filepath) - spine = modifier._get_spine() - adjusted_count = 0 - total_shift = 0 - - clips_list = [c for c in spine if c.tag in ('clip', 'asset-clip', 'video', 'ref-clip')] - - for i, clip in enumerate(clips_list[1:], 1): - cut_offset = modifier._parse_time(clip.get('offset', '0s')) - cut_seconds = cut_offset.to_seconds() - - best_marker = None - best_distance = float('inf') - - for marker_time in marker_times: - distance = abs(marker_time - cut_seconds) - distance_frames = distance * fps - - if distance_frames <= max_shift: - if prefer == "earlier" and marker_time <= cut_seconds: - if distance < best_distance: - best_distance = distance - best_marker = marker_time - elif prefer == "later" and marker_time >= cut_seconds: - if distance < best_distance: - best_distance = distance - best_marker = marker_time - elif prefer == "nearest": - if distance < best_distance: - best_distance = distance - best_marker = marker_time - - if best_marker is not None and best_distance > 0.001: - shift = best_marker - cut_seconds - shift_frames = int(shift * fps) - - prev_clip = clips_list[i - 1] - prev_dur = modifier._parse_time(prev_clip.get('duration', '0s')) - new_prev_dur = prev_dur + modifier._parse_time(f"{shift}s") - prev_clip.set('duration', new_prev_dur.to_fcpxml()) - - new_offset = modifier._parse_time(f"{best_marker}s") - clip.set('offset', new_offset.to_fcpxml()) - - adjusted_count += 1 - total_shift += abs(shift_frames) - - modifier.save(output_path) - avg_shift = total_shift / adjusted_count if adjusted_count > 0 else 0 - - return _text_result(f"""# Cuts Snapped to Beats - -## Summary -- **Cuts Adjusted**: {adjusted_count} -- **Max Shift Allowed**: {max_shift} frames -- **Preference**: {prefer} -- **Average Shift**: {avg_shift:.1f} frames - -## Output -Saved to: `{output_path}` - -Your edits are now synced to the beat! -""") - - -# ----- SUBTITLE / TRANSCRIPT HANDLERS ----- - -async def handle_import_srt_markers(arguments: dict) -> Sequence[TextContent]: - filepath, output_path = _resolve_io_paths(arguments, "_subtitled") - srt_path = _validate_filepath(arguments["srt_path"], ('.srt', '.vtt')) - mode = arguments.get("mode", "first_per_minute") - marker_type = arguments.get("marker_type", "chapter") - max_label = arguments.get("max_label_length", 50) - - text = Path(srt_path).read_text(encoding='utf-8') - - # Detect format and parse - if srt_path.endswith('.vtt') or text.strip().startswith('WEBVTT'): - raw_markers = parse_vtt(text) - fmt_name = "WebVTT" - else: - raw_markers = parse_srt(text) - fmt_name = "SRT" - - if not raw_markers: - return _text_result(f"No subtitles found in {srt_path}") - - # Apply mode filtering - filtered = [] - if mode == "all": - filtered = raw_markers - elif mode == "first_per_minute": - seen_minutes = set() - for m in raw_markers: - minute = int(m['seconds'] // 60) - if minute not in seen_minutes: - seen_minutes.add(minute) - filtered.append(m) - elif mode == "scene_changes": - # Group by similar text, take first occurrence of each unique line - seen_texts = set() - for m in raw_markers: - # Normalize: lowercase, strip punctuation - normalized = re.sub(r'[^\w\s]', '', m['text'].lower()).strip() - words = normalized.split()[:3] # First 3 words as key - key = ' '.join(words) - if key and key not in seen_texts: - seen_texts.add(key) - filtered.append(m) - - markers = _raw_markers_to_batch(filtered, marker_type, max_label=max_label) - - modifier = FCPXMLModifier(filepath) - added = modifier.batch_add_markers(markers=markers) - modifier.save(output_path) - - return _text_result(f"""# Subtitle Markers Imported - -## Summary -- **Format**: {fmt_name} -- **Subtitles Parsed**: {len(raw_markers)} -- **Mode**: {mode} -- **Markers Added**: {len(added)} -- **Marker Type**: {marker_type} - -## Output -Saved to: `{output_path}` -""") - - -async def handle_import_transcript_markers(arguments: dict) -> Sequence[TextContent]: - filepath, output_path = _resolve_io_paths(arguments, "_chapters") - marker_type = arguments.get("marker_type", "chapter") - - # Get transcript text from inline or file - transcript = arguments.get("transcript") - transcript_path = arguments.get("transcript_path") - - if not transcript and not transcript_path: - return _text_result("Provide either 'transcript' (inline text) or 'transcript_path' (path to file)") - - if transcript_path: - # .txt only: the parser below understands "0:00 Title" lines, not real - # SRT/VTT cue syntax — that's import_srt_markers (parse_srt/parse_vtt). - transcript_path = _validate_filepath(transcript_path, ('.txt',)) - transcript = Path(transcript_path).read_text(encoding='utf-8') - - raw_markers = parse_transcript_timestamps(transcript or "") - - if not raw_markers: - return _text_result("No timestamps found. Expected format: '0:00 Title' or 'HH:MM:SS Title', one per line.") - - markers = _raw_markers_to_batch(raw_markers, marker_type) - - modifier = FCPXMLModifier(filepath) - added = modifier.batch_add_markers(markers=markers) - modifier.save(output_path) - - return _text_result(f"""# Transcript Markers Imported - -## Summary -- **Timestamps Found**: {len(raw_markers)} -- **Markers Added**: {len(added)} -- **Marker Type**: {marker_type} - -## Markers -""" + "\n".join(f"- `{m['timecode']}` {m['name']}" for m in markers) + f""" - -## Output -Saved to: `{output_path}` -""") - - -# ----- CONNECTED CLIPS & COMPOUND CLIPS HANDLERS (v0.5.0) ----- - -async def handle_list_connected_clips(arguments: dict) -> Sequence[TextContent]: - project, tl = _require_timeline(arguments["filepath"]) - - lane_filter = arguments.get("lane") - clips = tl.connected_clips - if lane_filter is not None: - clips = [c for c in clips if c.lane == lane_filter] - - if not clips: - return _text_result("No connected clips found in timeline.") - - result = f"# Connected Clips in {tl.name}\n\n**Total**: {len(clips)}\n\n" - result += "| # | Name | Lane | Type | Duration | Parent | Role |\n" - result += "|---|------|------|------|----------|--------|------|\n" - for i, c in enumerate(clips, 1): - result += ( - f"| {i} | {c.name} | {c.lane} | {c.clip_type} | " - f"{format_duration(c.duration_seconds)} | {c.parent_clip_name} | " - f"{c.role or '-'} |\n" - ) - return _text_result(result) - - -async def handle_add_connected_clip(arguments: dict) -> Sequence[TextContent]: - filepath, output_path, modifier = _setup_modifier(arguments) - modifier.add_connected_clip( - parent_clip_id=arguments["parent_clip_id"], - asset_id=arguments.get("asset_id"), - asset_name=arguments.get("asset_name"), - offset=arguments.get("offset", "0s"), - duration=arguments.get("duration"), - lane=arguments.get("lane", 1), - ) - modifier.save(output_path) - return _text_result(( - f"Connected clip added to '{arguments['parent_clip_id']}' on lane {arguments.get('lane', 1)}\n\n" - f"Saved to: `{output_path}`" - )) - - -async def handle_list_compound_clips(arguments: dict) -> Sequence[TextContent]: - project, tl = _require_timeline(arguments["filepath"]) - - if not tl.compound_clips: - return _text_result("No compound clips found in timeline.") - - result = f"# Compound Clips in {tl.name}\n\n" - for i, cc in enumerate(tl.compound_clips, 1): - result += f"### {i}. {cc.name}\n" - result += f"- **Ref ID**: {cc.ref_id}\n" - result += f"- **Duration**: {format_duration(cc.duration_seconds)}\n" - result += f"- **Clips inside**: {len(cc.clips)}\n\n" - return _text_result(result) - - -# ----- ROLES HANDLERS (v0.5.0) ----- - -async def handle_list_roles(arguments: dict) -> Sequence[TextContent]: - project, tl = _require_timeline(arguments["filepath"]) - - audio_roles: dict[str, int] = {} - video_roles: dict[str, int] = {} - - for clip in tl.clips: - if clip.audio_role: - audio_roles[clip.audio_role] = audio_roles.get(clip.audio_role, 0) + 1 - if clip.video_role: - video_roles[clip.video_role] = video_roles.get(clip.video_role, 0) + 1 - - for cc in tl.connected_clips: - if cc.role: - # Determine type from clip_type - if cc.clip_type in ('audio', 'audio-clip'): - audio_roles[cc.role] = audio_roles.get(cc.role, 0) + 1 - else: - video_roles[cc.role] = video_roles.get(cc.role, 0) + 1 - - result = f"# Roles in {tl.name}\n\n" - if audio_roles: - result += "## Audio Roles\n\n| Role | Clips |\n|------|-------|\n" - for role, count in sorted(audio_roles.items()): - result += f"| {role} | {count} |\n" - else: - result += "## Audio Roles\n\nNo audio roles assigned.\n" - - result += "\n" - if video_roles: - result += "## Video Roles\n\n| Role | Clips |\n|------|-------|\n" - for role, count in sorted(video_roles.items()): - result += f"| {role} | {count} |\n" - else: - result += "## Video Roles\n\nNo video roles assigned.\n" - - return _text_result(result) - - -async def handle_assign_role(arguments: dict) -> Sequence[TextContent]: - filepath, output_path, modifier = _setup_modifier(arguments) - modifier.assign_role( - clip_id=arguments["clip_id"], - audio_role=arguments.get("audio_role"), - video_role=arguments.get("video_role"), - ) - modifier.save(output_path) - - roles_set = [] - if arguments.get("audio_role"): - roles_set.append(f"audioRole={arguments['audio_role']}") - if arguments.get("video_role"): - roles_set.append(f"videoRole={arguments['video_role']}") - - return _text_result(( - f"Set {', '.join(roles_set)} on '{arguments['clip_id']}'\n\n" - f"Saved to: `{output_path}`" - )) - - -async def handle_filter_by_role(arguments: dict) -> Sequence[TextContent]: - project, tl = _require_timeline(arguments["filepath"]) - - role = arguments["role"].lower() - role_type = arguments.get("role_type", "any") - matches = [] - - for clip in tl.clips: - if role_type in ("audio", "any") and clip.audio_role.lower() == role: - matches.append((clip.name, "audio", clip.audio_role, format_duration(clip.duration_seconds))) - if role_type in ("video", "any") and clip.video_role.lower() == role: - matches.append((clip.name, "video", clip.video_role, format_duration(clip.duration_seconds))) - - if not matches: - return _text_result(f"No clips found with role '{role}'.") - - result = f"# Clips with role '{role}'\n\n" - result += "| Clip | Type | Role | Duration |\n|------|------|------|----------|\n" - for name, rtype, rval, dur in matches: - result += f"| {name} | {rtype} | {rval} | {dur} |\n" - return _text_result(result) - - -async def handle_export_role_stems(arguments: dict) -> Sequence[TextContent]: - project, tl = _require_timeline(arguments["filepath"]) - - stems: dict[str, list] = {} - for clip in tl.clips: - role = clip.audio_role or "unassigned" - stems.setdefault(role, []).append(clip) - - for cc in tl.connected_clips: - role = cc.role or "unassigned" - stems.setdefault(role, []).append(cc) - - result = f"# Audio Stem Plan for {tl.name}\n\n" - for role, clips in sorted(stems.items()): - total_dur = sum(c.duration_seconds for c in clips) - result += f"## {role.title()} ({len(clips)} clips, {format_duration(total_dur)})\n\n" - for c in clips: - result += f"- {c.name} ({format_duration(c.duration_seconds)})\n" - result += "\n" - - return _text_result(result) - - -# ----- TIMELINE DIFF HANDLER (v0.5.0) ----- - -async def handle_diff_timelines(arguments: dict) -> Sequence[TextContent]: - filepath_a = _validate_filepath(arguments["filepath_a"], ('.fcpxml', '.fcpxmld')) - filepath_b = _validate_filepath(arguments["filepath_b"], ('.fcpxml', '.fcpxmld')) - - diff = compare_timelines(filepath_a, filepath_b) - - if not diff.has_changes: - return _text_result(( - f"# Timeline Diff: No Changes\n\n" - f"**{diff.timeline_a_name}** vs **{diff.timeline_b_name}** are identical." - )) - - result = ( - f"# Timeline Diff\n\n" - f"**Baseline**: {diff.timeline_a_name}\n" - f"**Comparison**: {diff.timeline_b_name}\n" - f"**Total changes**: {diff.total_changes}\n\n" - ) - - if diff.format_changes: - result += "## Format Changes\n\n" - for change in diff.format_changes: - result += f"- {change}\n" - result += "\n" - - clip_changes = [d for d in diff.clip_diffs if d.action != "unchanged"] - if clip_changes: - result += "## Clip Changes\n\n| Action | Clip | Details |\n|--------|------|--------|\n" - for d in clip_changes: - result += f"| {d.action.upper()} | {d.clip_name} | {d.details} |\n" - result += "\n" - - if diff.marker_diffs: - result += "## Marker Changes\n\n| Action | Marker | Details |\n|--------|--------|--------|\n" - for d in diff.marker_diffs: - result += f"| {d.action.upper()} | {d.marker_name} | {d.details} |\n" - result += "\n" - - if diff.transition_diffs: - result += "## Transition Changes\n\n" - for change in diff.transition_diffs: - result += f"- {change}\n" - - return _text_result(result) - - -# ----- SOCIAL MEDIA REFORMAT HANDLER (v0.5.0) ----- - -async def handle_reformat_timeline(arguments: dict) -> Sequence[TextContent]: - filepath, output_path = _resolve_io_paths(arguments, "_reformatted") - - fmt = arguments["format"] - if fmt == "custom": - width = arguments.get("width") - height = arguments.get("height") - if not width or not height: - return _text_result("Custom format requires both 'width' and 'height' parameters.") - else: - formats = FCPXMLModifier.SOCIAL_FORMATS - if fmt not in formats: - return _text_result(f"Unknown format: {fmt}. Valid: {', '.join(formats.keys())}") - width, height = formats[fmt] - - modifier = FCPXMLModifier(filepath) - modifier.reformat_resolution(width, height) - modifier.save(output_path) - - return _text_result(( - f"# Timeline Reformatted\n\n" - f"- **Format**: {fmt} ({width}x{height})\n" - f"- **Aspect ratio**: {width}:{height}\n\n" - f"Saved to: `{output_path}`\n\n" - f"**Next step**: Import into FCP (File > Import > XML). " - f"FCP will handle spatial conforming automatically." - )) - - -# ----- SILENCE DETECTION HANDLERS (v0.5.0) ----- - -async def handle_detect_media_silence(arguments: dict) -> Sequence[TextContent]: - noise_db = float(arguments.get("noise_db", -30.0)) - min_silence = float(arguments.get("min_silence", 0.5)) - # Same bounds detect_silence() enforces — validated here so a bad request - # fails before any media file is opened. - if not (-120.0 <= noise_db <= 0.0): - raise ValueError(f"noise_db must be between -120 and 0 dB, got {noise_db}") - if not (0 < min_silence <= 3600): - raise ValueError(f"min_silence must be between 0 and 3600 seconds, got {min_silence}") - - filepath = _validate_filepath(arguments["filepath"], ('.fcpxml', '.fcpxmld')) - modifier = FCPXMLModifier(filepath) - clip_filter = arguments.get("clip_name") - - max_media_probes = 100 - findings: list[tuple[str, float, float]] = [] - skipped: list[tuple[str, str]] = [] - probe_cache: dict[str, list | None] = {} - for el in [el for _, el in modifier._iter_spine_clips()]: - name = el.get("name", "") - if clip_filter and name != clip_filter: - continue - src = modifier.resources.get(el.get("ref", ""), {}).get("src", "") - media_path = media_src_to_path(src) - if not media_path or not Path(media_path).is_file(): - skipped.append((name, "media file missing")) - continue - if media_path not in probe_cache: - if len(probe_cache) >= max_media_probes: - skipped.append((name, f"probe cap reached ({max_media_probes} media files)")) - continue - probe_cache[media_path] = detect_silence( - media_path, noise_db=noise_db, min_duration=min_silence - ) - silences = probe_cache[media_path] - if silences is None: - skipped.append((name, "unanalyzable (ffmpeg missing or media unreadable)")) - continue - source_start = modifier.source_file_start(el).to_seconds() - clip_duration = modifier._parse_time(el.get("duration", "0s")).to_seconds() - timeline_offset = modifier._parse_time(el.get("offset", "0s")).to_seconds() - mapped = map_silence_to_timeline( - silences, source_start, clip_duration, timeline_offset - ) - findings.extend((name, start, end) for start, end in mapped) - - total_silence = sum(end - start for _, start, end in findings) - result = f"""# Media Silence Detection (real audio analysis) - -## Summary -- **Threshold**: {noise_db} dB for >= {min_silence}s -- **Media Files Probed**: {len(probe_cache)} -- **Silence Spans Found**: {len(findings)} ({format_duration(total_silence)} total) -""" - if findings: - result += "\n## Silence Spans (timeline time)\n" - result += _markdown_table( - ["Clip", "Start", "End", "Duration"], - [[name, f"{start:.2f}s", f"{end:.2f}s", f"{end - start:.2f}s"] - for name, start, end in findings], - ) + "\n" - result += "\n*To remove: `split_clip` at each boundary, then `delete_clips` with ripple.*" - if skipped: - result += "\n## Skipped Clips\n" - result += _markdown_table( - ["Clip", "Reason"], [[name, reason] for name, reason in skipped] - ) + "\n" - if not findings and not skipped: - result += "\nNo silence detected in any clip's source audio." - return _text_result(result) - - -async def handle_remove_media_silence(arguments: dict) -> Sequence[TextContent]: - noise_db = float(arguments.get("noise_db", -30.0)) - min_silence = float(arguments.get("min_silence", 0.5)) - padding = float(arguments.get("padding", 0.05)) - if not (-120.0 <= noise_db <= 0.0): - raise ValueError(f"noise_db must be between -120 and 0 dB, got {noise_db}") - if not (0 < min_silence <= 3600): - raise ValueError(f"min_silence must be between 0 and 3600 seconds, got {min_silence}") - if not (0 <= padding <= 5): - raise ValueError(f"padding must be between 0 and 5 seconds, got {padding}") - - filepath, output_path, modifier = _setup_modifier(arguments, "_silence_removed") - clip_filter = arguments.get("clip_name") - to_frame_timevalue = modifier.snap_seconds_to_frame - - max_media_probes = 100 - cuts_made: list[tuple[str, int, float]] = [] - skipped: list[tuple[str, str]] = [] - probe_cache: dict[str, list | None] = {} - spine_clips = [el for _, el in modifier._iter_spine_clips()] - for el in spine_clips: - name = el.get("name", "") - if clip_filter and name != clip_filter: - continue - src = modifier.resources.get(el.get("ref", ""), {}).get("src", "") - media_path = media_src_to_path(src) - if not media_path or not Path(media_path).is_file(): - skipped.append((name, "media file missing")) - continue - if media_path not in probe_cache: - if len(probe_cache) >= max_media_probes: - skipped.append((name, f"probe cap reached ({max_media_probes} media files)")) - continue - probe_cache[media_path] = detect_silence( - media_path, noise_db=noise_db, min_duration=min_silence - ) - silences = probe_cache[media_path] - if silences is None: - skipped.append((name, "unanalyzable (ffmpeg missing or media unreadable)")) - continue - - clip_source_start = modifier.source_file_start(el).to_seconds() - clip_duration = modifier._parse_time(el.get("duration", "0s")).to_seconds() - cut_ranges = [] - for sil_start, sil_end in silences: - # Source time -> clip-relative, padded so cuts breathe. - cut_start = max(sil_start, clip_source_start) - clip_source_start + padding - cut_end = min(sil_end, clip_source_start + clip_duration) - clip_source_start - padding - if cut_end > cut_start: - cut_ranges.append((to_frame_timevalue(cut_start), to_frame_timevalue(cut_end))) - if not cut_ranges: - continue - removed = modifier.cut_clip_ranges(el, cut_ranges) - if removed > TimeValue.zero(): - cuts_made.append((name, len(cut_ranges), removed.to_seconds())) - - if not cuts_made: - text = "# Media Silence Removal\n\nNo silence found to remove — file unchanged (nothing saved)." - if skipped: - text += "\n\n## Skipped Clips\n" + _markdown_table( - ["Clip", "Reason"], [[name, reason] for name, reason in skipped] - ) - return _text_result(text) - - modifier.remove_trailing_gaps() - modifier.save(output_path) - total_removed = sum(seconds for _, _, seconds in cuts_made) - result = f"""# Media Silence Removal (real audio analysis) - -## Summary -- **Threshold**: {noise_db} dB for >= {min_silence}s, padding {padding}s -- **Clips Cut**: {len(cuts_made)} -- **Total Removed**: {format_duration(total_removed)} - -## Cuts -""" - result += _markdown_table( - ["Clip", "Silence Spans Cut", "Removed"], - [[name, str(count), f"{seconds:.2f}s"] for name, count, seconds in cuts_made], - ) + "\n" - if skipped: - result += "\n## Skipped Clips\n" + _markdown_table( - ["Clip", "Reason"], [[name, reason] for name, reason in skipped] - ) + "\n" - result += f"\nSaved to: {output_path}\n\n*Preview first next time with `detect_media_silence`. Original file untouched.*" - return _text_result(result) - - -AUDIO_MEDIA_EXTENSIONS = ( - '.wav', '.aif', '.aiff', '.mp3', '.m4a', '.aac', '.flac', '.mov', '.mp4', -) - - -async def handle_detect_beats(arguments: dict) -> Sequence[TextContent]: - media_path = _validate_filepath(arguments["media_path"], AUDIO_MEDIA_EXTENSIONS) - - result = detect_beats(media_path) - if result is None: - return _text_result( - "Beat detection unavailable — librosa is not installed or the file " - "could not be analyzed.\n\nInstall the optional media-intelligence " - "extra:\n\n pip install 'fcp-mcp-server[intelligence]'" - ) - - bpm, beats = result["bpm"], result["beats"] - beats_data = { - "source": str(Path(media_path).name), - "bpm": round(bpm, 2), - "beats": [round(b, 4) for b in beats], - "downbeats": [round(b, 4) for b in beats[::4]], - } - json_path = _validate_output_path( - str(Path(media_path).with_name(Path(media_path).stem + "_beats.json")), - anchor_dir=str(Path(media_path).parent), - ) - with open(json_path, "w") as f: - json.dump(beats_data, f, indent=2) - - preview = beats[:16] - result_text = f"""# Beat Detection - -## Summary -- **Source**: {Path(media_path).name} -- **Estimated Tempo**: {bpm:.1f} BPM -- **Beats Detected**: {len(beats)} ({format_duration(beats[-1]) if beats else '0s'} span) -- **Beats JSON**: {json_path} - -## First Beats -""" - result_text += _markdown_table( - ["#", "Time"], - [[str(i + 1), f"{b:.3f}s"] for i, b in enumerate(preview)], - ) + "\n" - result_text += ( - f"\n*Next: `import_beat_markers` with beats_path=\"{json_path}\" to place " - "markers, then `snap_to_beats` to align your cuts.*" - ) - return _text_result(result_text) - - -# ===== TRANSCRIPT INTELLIGENCE (v0.13.1) ===== - -TRANSCRIBE_MAX_MEDIA = 10 - -_TRANSCRIBE_INSTALL_HINT = ( - "\n\nInstall the optional transcription extra:\n\n" - " pip install 'fcp-mcp-server[transcribe]'\n\n" - "or run via uvx:\n\n" - " uvx --from \"fcp-mcp-server[transcribe]\" fcp-mcp-server" -) - - -def _transcript_json_path(media_path: str, output_dir: str | None = None) -> Path: - """Where the ``_transcript.json`` for ``media_path`` lives. - - When ``output_dir`` (the user-selected project folder) is set, the - transcript is saved/read there instead of next to the source media. - """ - p = Path(media_path) - if output_dir: - directory = Path(output_dir).expanduser() - directory.mkdir(parents=True, exist_ok=True) - return directory / f"{p.stem}_transcript.json" - return p.with_name(p.stem + "_transcript.json") - - -def _load_or_transcribe( - media_path: str, model: str, language: str | None, output_dir: str | None = None -) -> tuple[dict | None, str]: - """Load a cached ``_transcript.json`` for a media file, else transcribe and cache it. - - Returns ``(transcript, "")`` or ``(None, reason)``. The cache makes - transcription a one-time cost per media file across all transcript tools. - """ - json_path = _transcript_json_path(media_path, output_dir) - if json_path.is_file(): - try: - with open(json_path) as f: - data = json.load(f) - if isinstance(data, dict) and isinstance(data.get("words"), list): - return data, "" - except (OSError, json.JSONDecodeError, UnicodeDecodeError): - pass # unreadable cache falls through to re-transcribe - result = transcribe(media_path, model_size=model, language=language) - if result is None: - return None, "untranscribable (faster-whisper not installed or media unreadable)" - anchor = str(Path(output_dir).expanduser()) if output_dir else str(Path(media_path).parent) - out_path = _validate_output_path(str(json_path), anchor_dir=anchor) - with open(out_path, "w") as f: - json.dump({"source": Path(media_path).name, **result}, f, indent=2) - return result, "" - - -def _cut_transcript_spans(modifier, clip_filter, model, language, padding, spans_fn, keep_only=False, output_dir=None): - """Shared cut engine for transcript-driven editing. - - ``spans_fn(words) -> [(start, end), ...]`` in source seconds. Spans are - padded, clamped to each clip's used source window, optionally inverted - (keep_only), snapped to the frame grid, and cut with ripple. - """ - to_frame = modifier.snap_seconds_to_frame - - cache: dict[str, tuple] = {} - cuts_made: list[tuple[str, int, float]] = [] - skipped: list[tuple[str, str]] = [] - spine_clips = [el for _, el in modifier._iter_spine_clips()] - for el in spine_clips: - name = el.get("name", "") - if clip_filter and name != clip_filter: - continue - src = modifier.resources.get(el.get("ref", ""), {}).get("src", "") - media_path = media_src_to_path(src) - if not media_path or not Path(media_path).is_file(): - skipped.append((name, "media file missing")) - continue - if media_path not in cache: - if len(cache) >= TRANSCRIBE_MAX_MEDIA: - skipped.append((name, f"transcription cap reached ({TRANSCRIBE_MAX_MEDIA} media files)")) - continue - cache[media_path] = _load_or_transcribe(media_path, model, language, output_dir) - data, reason = cache[media_path] - if data is None: - skipped.append((name, reason)) - continue - - clip_source_start = modifier.source_file_start(el).to_seconds() - clip_duration = modifier._parse_time(el.get("duration", "0s")).to_seconds() - window_start = clip_source_start - window_end = clip_source_start + clip_duration - - spans = spans_fn(data.get("words", [])) - padded = merge_ranges([(s - padding, e + padding) for s, e in spans]) - clamped = [ - (max(s, window_start), min(e, window_end)) - for s, e in padded - if min(e, window_end) > max(s, window_start) - ] - if keep_only: - if not clamped: - # Never delete a whole clip just because nothing matched in it. - skipped.append((name, "no phrase matches — left untouched (keep_only)")) - continue - cut_source = invert_ranges(clamped, window_start, window_end) - else: - cut_source = clamped - cut_ranges = [ - (to_frame(s - clip_source_start), to_frame(e - clip_source_start)) - for s, e in cut_source - ] - cut_ranges = [(a, b) for a, b in cut_ranges if b > a] - if not cut_ranges: - continue - removed = modifier.cut_clip_ranges(el, cut_ranges) - if removed > TimeValue.zero(): - cuts_made.append((name, len(cut_ranges), removed.to_seconds())) - return cuts_made, skipped - - -def _transcript_cut_report(title, summary_lines, cuts_made, skipped, output_path, footer): - if not cuts_made: - text = f"# {title}\n\nNo cuts to make — file unchanged (nothing saved)." - if skipped: - text += "\n\n## Skipped Clips\n" + _markdown_table( - ["Clip", "Reason"], [[name, reason] for name, reason in skipped] - ) - if any("faster-whisper" in reason for _, reason in skipped): - text += _TRANSCRIBE_INSTALL_HINT - return _text_result(text) - total_removed = sum(seconds for _, _, seconds in cuts_made) - result = f"# {title}\n\n## Summary\n" - result += "\n".join(summary_lines) + "\n" - result += f"- **Clips Cut**: {len(cuts_made)}\n- **Total Removed**: {format_duration(total_removed)}\n" - result += "\n## Cuts\n" - result += _markdown_table( - ["Clip", "Ranges Cut", "Removed"], - [[name, str(count), f"{seconds:.2f}s"] for name, count, seconds in cuts_made], - ) + "\n" - if skipped: - result += "\n## Skipped Clips\n" + _markdown_table( - ["Clip", "Reason"], [[name, reason] for name, reason in skipped] - ) + "\n" - result += f"\nSaved to: {output_path}\n\n{footer}" - return _text_result(result) - - -async def handle_transcribe_media(arguments: dict) -> Sequence[TextContent]: - model = arguments.get("model", "base") - language = arguments.get("language") - output_dir = arguments.get("output_dir") - write_srt = bool(arguments.get("write_srt", False)) - _, tl = _require_timeline(arguments["filepath"]) - clip_filter = arguments.get("clip_name") - - done: dict[str, dict | None] = {} - skipped: list[tuple[str, str]] = [] - rows: list[list[str]] = [] - srt_paths: list[str] = [] - for clip in tl.clips: - if clip_filter and clip.name != clip_filter: - continue - media_path = media_src_to_path(clip.media_path or "") - if not media_path or not Path(media_path).is_file(): - skipped.append((clip.name, "media file missing")) - continue - if media_path in done: - continue - if len(done) >= TRANSCRIBE_MAX_MEDIA: - skipped.append((clip.name, f"transcription cap reached ({TRANSCRIBE_MAX_MEDIA} media files)")) - continue - data, reason = _load_or_transcribe(media_path, model, language, output_dir) - done[media_path] = data - if data is None: - skipped.append((clip.name, reason)) - continue - if write_srt and data.get("segments"): - srt_name = Path(media_path).stem + "_transcript.srt" - srt_anchor = str(Path(output_dir).expanduser()) if output_dir else str(Path(media_path).parent) - srt_path = _validate_output_path( - str(Path(srt_anchor) / srt_name), - anchor_dir=srt_anchor, - ) - with open(srt_path, "w") as f: - f.write(segments_to_srt(data["segments"])) - srt_paths.append(srt_path) - preview = data.get("text", "")[:160] - rows.append([ - Path(media_path).name, - data.get("language", "?"), - str(len(data.get("words", []))), - format_duration(float(data.get("duration", 0.0))), - preview + ("…" if len(data.get("text", "")) > 160 else ""), - ]) - - result = f"""# Media Transcription (local Whisper) - -## Summary -- **Model**: {model} -- **Media Files Transcribed**: {len(rows)} -""" - if rows: - result += "\n## Transcripts (saved as _transcript.json next to each media file)\n" - result += _markdown_table( - ["Media", "Language", "Words", "Duration", "Preview"], rows - ) + "\n" - result += ( - "\n*Next: `edit_by_transcript` to cut by what was said, or " - "`remove_filler_words` to clean ums/uhs. Transcripts are cached — " - "media is only transcribed once.*" - ) - if srt_paths: - result += "\n\n## SRT Files\n" + "\n".join(f"- {p}" for p in srt_paths) - if skipped: - result += "\n## Skipped Clips\n" + _markdown_table( - ["Clip", "Reason"], [[name, reason] for name, reason in skipped] - ) + "\n" - if not rows and any("faster-whisper" in reason for _, reason in skipped): - result += _TRANSCRIBE_INSTALL_HINT - return _text_result(result) - - -async def handle_edit_by_transcript(arguments: dict) -> Sequence[TextContent]: - phrases = arguments.get("phrases") or [] - if not isinstance(phrases, list) or not all(isinstance(p, str) for p in phrases): - raise ValueError("phrases must be a list of strings") - phrases = [p for p in phrases if p.strip()] - if not phrases: - raise ValueError("phrases must contain at least one non-empty string") - mode = arguments.get("mode", "remove") - if mode not in ("remove", "keep_only"): - raise ValueError(f"mode must be 'remove' or 'keep_only', got {mode!r}") - padding = float(arguments.get("padding", 0.0)) - if not (0 <= padding <= 2): - raise ValueError(f"padding must be between 0 and 2 seconds, got {padding}") - model = arguments.get("model", "base") - language = arguments.get("language") - output_dir = arguments.get("output_dir") - - filepath, output_path, modifier = _setup_modifier(arguments, "_transcript_edit") - - def spans_fn(words): - return merge_ranges( - [span for phrase in phrases for span in find_phrase_spans(words, phrase)] - ) - - cuts_made, skipped = _cut_transcript_spans( - modifier, arguments.get("clip_name"), model, language, padding, - spans_fn, keep_only=(mode == "keep_only"), output_dir=output_dir, - ) - if cuts_made: - modifier.save(output_path) - verb = "kept only" if mode == "keep_only" else "removed" - return _transcript_cut_report( - "Transcript Edit", - [f"- **Mode**: {mode} ({verb} the matched phrases)", - f"- **Phrases**: {', '.join(repr(p) for p in phrases)}", - f"- **Padding**: {padding}s"], - cuts_made, skipped, output_path, - "*Transcripts are cached as _transcript.json. Original file untouched.*", - ) - - -async def handle_remove_filler_words(arguments: dict) -> Sequence[TextContent]: - fillers = arguments.get("fillers") or list(DEFAULT_FILLERS) - if not isinstance(fillers, list) or not all(isinstance(f, str) for f in fillers): - raise ValueError("fillers must be a list of strings") - padding = float(arguments.get("padding", 0.02)) - if not (0 <= padding <= 2): - raise ValueError(f"padding must be between 0 and 2 seconds, got {padding}") - model = arguments.get("model", "base") - language = arguments.get("language") - output_dir = arguments.get("output_dir") - - filepath, output_path, modifier = _setup_modifier(arguments, "_defillered") - - cuts_made, skipped = _cut_transcript_spans( - modifier, arguments.get("clip_name"), model, language, padding, - lambda words: merge_ranges(find_filler_spans(words, fillers)), - output_dir=output_dir, - ) - if cuts_made: - modifier.save(output_path) - return _transcript_cut_report( - "Filler Word Removal", - [f"- **Fillers**: {', '.join(fillers)}", f"- **Padding**: {padding}s"], - cuts_made, skipped, output_path, - "*Transcripts are cached as _transcript.json. Original file untouched.*", - ) - - -async def handle_transcript_markers(arguments: dict) -> Sequence[TextContent]: - """Add a marker at the start of each transcribed segment, using each - media's cached (or freshly transcribed) local Whisper transcript. - - Unlike ``import_transcript_markers`` (plain "0:00 Title" text) or - ``import_srt_markers`` (a caption track already synced to the whole - exported video), this maps each segment's SOURCE-media timestamp to its - TIMELINE position per spine clip — the same source->timeline mapping - ``detect_media_silence`` uses — so it stays correct across multiple - clips built from different (and differently-trimmed) source files. - """ - marker_type = arguments.get("marker_type", "chapter") - max_label = int(arguments.get("max_label_length", 50)) - model = arguments.get("model", "base") - language = arguments.get("language") - output_dir = arguments.get("output_dir") - clip_filter = arguments.get("clip_name") - - filepath, output_path, modifier = _setup_modifier(arguments, "_transcript_markers") - - added: list[tuple[str, float, str]] = [] - skipped: list[tuple[str, str]] = [] - spine_clips = [el for _, el in modifier._iter_spine_clips()] - for el in spine_clips: - name = el.get("name", "") - if clip_filter and name != clip_filter: - continue - src = modifier.resources.get(el.get("ref", ""), {}).get("src", "") - media_path = media_src_to_path(src) - if not media_path or not Path(media_path).is_file(): - skipped.append((name, "media file missing")) - continue - data, reason = _load_or_transcribe(media_path, model, language, output_dir) - if data is None: - skipped.append((name, reason)) - continue - - clip_source_start = modifier.source_file_start(el).to_seconds() - clip_duration = modifier._parse_time(el.get("duration", "0s")).to_seconds() - clip_offset = modifier._parse_time(el.get("offset", "0s")).to_seconds() - window_end = clip_source_start + clip_duration - - for seg in data.get("segments", []): - seg_start = float(seg.get("start", 0.0)) - if seg_start < clip_source_start or seg_start >= window_end: - continue - label = seg.get("text", "").strip() - if not label: - continue - if max_label and len(label) > max_label: - label = label[:max_label] - timeline_seconds = clip_offset + (seg_start - clip_source_start) - modifier.add_marker_at_timeline( - timecode=f"{timeline_seconds}s", name=label, marker_type=marker_type, - ) - added.append((name, seg_start, label)) - - if not added: - text = "# Transcript Markers\n\nNo segments to mark — file unchanged (nothing saved)." - if skipped: - text += "\n\n## Skipped Clips\n" + _markdown_table( - ["Clip", "Reason"], [[n, r] for n, r in skipped] - ) - return _text_result(text) - - modifier.save(output_path) - result = "# Transcript Markers Imported (local Whisper)\n\n## Summary\n" - result += f"- **Markers Added**: {len(added)}\n- **Marker Type**: {marker_type}\n\n" - result += _markdown_table( - ["Clip", "Start", "Label"], [[n, f"{s:.2f}s", label] for n, s, label in added] - ) - if skipped: - result += "\n## Skipped Clips\n" + _markdown_table( - ["Clip", "Reason"], [[n, r] for n, r in skipped] - ) - result += f"\n\nSaved to: `{output_path}`\n\n*Transcripts are cached as _transcript.json.*" - return _text_result(result) - - -async def handle_generate_dynamic_subtitles(arguments: dict) -> Sequence[TextContent]: - """Generate per-word subtitle titles laid out as a block per sentence. - - Whisper's segments become sentences; each word becomes its own positioned - <title> connected clip, appearing as it is spoken and accumulating on - screen until the whole block clears at once. No compound clip. - - Reuses the same SOURCE-media -> TIMELINE mapping as ``transcript_markers`` - (``modifier.source_file_start`` per spine clip) so word timestamps land - at the correct position even across trimmed/multiple clips. - """ - model = arguments.get("model", "base") - language = arguments.get("language") - output_dir = arguments.get("output_dir") - clip_filter = arguments.get("clip_name") - - body_color = arguments.get("active_color", "1 1 1 1") - config = DynamicSubtitleConfig( - style=WordStyle( - font=arguments.get("font", "Helvetica Neue"), - font_size=int(arguments.get("font_size", 88)), - active_color=body_color, - inactive_color=arguments.get("inactive_color", "0.7 0.7 0.7 1"), - emphasis_look=WordLook( - int(arguments.get("emphasis_size", 230)), - body_color, - font=arguments.get("emphasis_font", "Playfair Display"), - face=arguments.get("emphasis_face", "Medium Italic"), - kerning=0.0, - ), - body_look=WordLook( - int(arguments.get("font_size", 88)), - body_color, - font=arguments.get("font", "Helvetica Neue"), - face="Bold", - kerning=1.2, - ), - ), - band_height=float(arguments.get("band_height", 0.22)), - block_center_y=float(arguments.get("block_center_y", -167.0)), - granularity=arguments.get("granularity", "phrase"), - ) - - filepath, output_path, modifier = _setup_modifier(arguments, "_dynamic_subtitles") - - added: list[tuple[str, int, int]] = [] - skipped: list[tuple[str, str]] = [] - spine_clips = [el for _, el in modifier._iter_spine_clips()] - for el in spine_clips: - name = el.get("name", "") - if clip_filter and name != clip_filter: - continue - src = modifier.resources.get(el.get("ref", ""), {}).get("src", "") - media_path = media_src_to_path(src) - if not media_path or not Path(media_path).is_file(): - skipped.append((name, "media file missing")) - continue - data, reason = _load_or_transcribe(media_path, model, language, output_dir) - if data is None: - skipped.append((name, reason)) - continue - - clip_source_start = modifier.source_file_start(el).to_seconds() - clip_duration = modifier._parse_time(el.get("duration", "0s")).to_seconds() - window_end = clip_source_start + clip_duration - - clip_words = [ - { - "word": w.get("word", ""), - "start": float(w.get("start", 0.0)) - clip_source_start, - "end": float(w.get("end", 0.0)) - clip_source_start, - } - for w in data.get("words", []) - if clip_source_start <= float(w.get("start", 0.0)) < window_end - ] - if not clip_words: - skipped.append((name, "no words in clip's source range")) - continue - - # Sentence boundaries, rebased the same way, so each sentence becomes - # its own block of titles that builds up and then clears together. - # Overlap rather than containment: a segment straddling the clip's - # in-point still governs the words that made the cut. - clip_segments = [ - { - "start": float(s.get("start", 0.0)) - clip_source_start, - "end": float(s.get("end", 0.0)) - clip_source_start, - } - for s in data.get("segments", []) - if float(s.get("end", 0.0)) > clip_source_start - and float(s.get("start", 0.0)) < window_end - ] - - # Pass the element itself, not `name` — after ripple-cut/silence - # removal every fragment of an originally-named clip keeps the same - # `name`, so a name lookup here would resolve every clip in this - # loop to whichever one `self.clips` last indexed, stacking every - # clip's captions onto a single wrong spine element instead of each - # clip's own. See Engine/docs/05_EXPERIENCIAS.md, entry 2026-08-17. - lines = modifier.generate_dynamic_subtitles( - el, clip_words, config, segments=clip_segments - ) - added.append((name, len(lines), len(clip_words))) - - if not added: - text = "# Dynamic Subtitles\n\nNo captions generated — file unchanged (nothing saved)." - if skipped: - text += "\n\n## Skipped Clips\n" + _markdown_table( - ["Clip", "Reason"], [[n, r] for n, r in skipped] - ) - return _text_result(text) - - modifier.save(output_path) - total_lines = sum(lines for _, lines, _ in added) - total_words = sum(words for _, _, words in added) - result = "# Dynamic Subtitles Generated (local Whisper)\n\n## Summary\n" - result += ( - f"- **Clips Captioned**: {len(added)}\n" - f"- **Caption Lines (Title Clips)**: {total_lines}\n" - f"- **Total Words**: {total_words}\n\n" - ) - result += _markdown_table( - ["Clip", "Caption Lines", "Words"], - [[n, str(lines), str(words)] for n, lines, words in added], - ) - if skipped: - result += "\n## Skipped Clips\n" + _markdown_table( - ["Clip", "Reason"], [[n, r] for n, r in skipped] - ) - result += f"\n\nSaved to: `{output_path}`\n\n*Transcripts are cached as _transcript.json.*" - return _text_result(result) - - -async def handle_detect_silence_candidates(arguments: dict) -> Sequence[TextContent]: - filepath = _validate_filepath(arguments["filepath"], ('.fcpxml', '.fcpxmld')) - modifier = FCPXMLModifier(filepath) - candidates = modifier.detect_silence_candidates( - min_gap_seconds=arguments.get("min_gap_seconds", 0.5), - patterns=arguments.get("patterns"), - ) - - if not candidates: - return _text_result("No silence candidates detected.") - - result = f"# Silence Candidates Detected\n\n**Found**: {len(candidates)}\n\n" - result += "| # | Timecode | Duration | Reason | Confidence | Clip |\n" - result += "|---|----------|----------|--------|------------|------|\n" - for i, c in enumerate(candidates, 1): - result += ( - f"| {i} | {c['start_timecode']} | {format_duration(c['duration_seconds'])} | " - f"{c['reason']} | {c['confidence']:.0%} | {c.get('clip_name') or '-'} |\n" - ) - result += ( - "\n**Note**: Detection uses timeline heuristics (gaps, ultra-short clips, name patterns). " - "Review candidates before removing — some may be intentional." - ) - return _text_result(result) - - -async def handle_remove_silence_candidates(arguments: dict) -> Sequence[TextContent]: - filepath, output_path, modifier = _setup_modifier(arguments, "_silence_cleaned") - actions = modifier.remove_silence_candidates( - mode=arguments.get("mode", "mark"), - min_gap_seconds=arguments.get("min_gap_seconds", 0.5), - min_confidence=arguments.get("min_confidence", 0.7), - ) - modifier.save(output_path) - - if not actions: - return _text_result("No silence candidates met the confidence threshold.") - - mode = arguments.get("mode", "mark") - result = f"# Silence Candidates {'Marked' if mode == 'mark' else 'Removed'}\n\n" - result += f"**Actions taken**: {len(actions)}\n\n" - for a in actions: - result += f"- **{a['action']}** {a.get('clip_name', 'gap')} ({a['reason']})\n" - result += f"\nSaved to: `{output_path}`" - return _text_result(result) - - -# ----- NLE EXPORT HANDLERS (v0.5.0) ----- - -async def handle_export_resolve_xml(arguments: dict) -> Sequence[TextContent]: - filepath, output_path = _resolve_io_paths(arguments, "_resolve") - exporter = DaVinciExporter(filepath) - exporter.export_simplified_fcpxml( - output_path, - flatten_compounds=arguments.get("flatten_compounds", True), - ) - return _text_result(( - f"# Exported for DaVinci Resolve\n\n" - f"- **Format**: Simplified FCPXML v1.9\n" - f"- **Compound clips flattened**: {arguments.get('flatten_compounds', True)}\n\n" - f"Saved to: `{output_path}`\n\n" - f"**Next step**: In DaVinci Resolve, go to File > Import > Timeline > Import AAF/EDL/XML" - )) - - -async def handle_export_fcp7_xml(arguments: dict) -> Sequence[TextContent]: - filepath, output_path = _resolve_io_paths(arguments, "_fcp7") - exporter = DaVinciExporter(filepath) - exporter.export_xmeml(output_path) - return _text_result(( - f"# Exported as FCP7 XML (XMEML)\n\n" - f"- **Format**: XMEML v5\n" - f"- **Compatible with**: Premiere Pro, DaVinci Resolve, Avid Media Composer\n\n" - f"Saved to: `{output_path}`\n\n" - f"**Next step**: Import via File > Import in your target NLE" - )) - - -# ----- v0.6.0 HANDLERS ----- - -async def handle_list_effects(arguments: dict) -> Sequence[TextContent]: - effects = list_effects() - lines = ["# Available FCP Transition Effects\n"] - for eff in effects: - lines.append(f"- **{eff['slug']}**: {eff['name']} (`{eff['uuid']}`)") - return _text_result("\n".join(lines)) - - -async def handle_add_audio(arguments: dict) -> Sequence[TextContent]: - filepath, output_path, modifier = _setup_modifier(arguments, "_audio") - - parent_clip_id = arguments.get("parent_clip_id") - if parent_clip_id: - modifier.add_audio_clip( - parent_clip_id=parent_clip_id, - asset_id=arguments.get("asset_id"), - offset=arguments.get("offset", "0s"), - duration=arguments.get("duration"), - role=arguments.get("role", "dialogue"), - lane=arguments.get("lane", -1), - src=arguments.get("src"), - ) - action = f"Added audio clip to '{parent_clip_id}'" - else: - modifier.add_music_bed( - asset_id=arguments.get("asset_id"), - duration=arguments.get("duration"), - role=arguments.get("role", "music"), - src=arguments.get("src"), - ) - action = "Added music bed spanning full timeline" - - modifier.save(output_path) - return _text_result(f"{action}\nSaved to: `{output_path}`") - - -async def handle_create_compound_clip(arguments: dict) -> Sequence[TextContent]: - filepath, output_path, modifier = _setup_modifier(arguments, "_compound") - clip_ids = arguments["clip_ids"] - name = arguments.get("name", "Compound Clip") - modifier.create_compound_clip(clip_ids, name) - modifier.save(output_path) - return _text_result(( - f"Created compound clip '{name}' from {len(clip_ids)} clips.\n" - f"Saved to: `{output_path}`" - )) - - -async def handle_flatten_compound_clip(arguments: dict) -> Sequence[TextContent]: - filepath, output_path, modifier = _setup_modifier(arguments, "_flattened") - ref_clip_id = arguments["ref_clip_id"] - extracted = modifier.flatten_compound_clip(ref_clip_id) - modifier.save(output_path) - return _text_result(( - f"Flattened compound clip '{ref_clip_id}' into {len(extracted)} clips.\n" - f"Saved to: `{output_path}`" - )) - - -async def handle_list_templates(arguments: dict) -> Sequence[TextContent]: - templates = list_templates() - lines = ["# Available Timeline Templates\n"] - for tmpl in templates: - lines.append(f"## {tmpl['name']}") - lines.append(f"{tmpl['description']}\n") - lines.append("| Slot | Type | Default Duration | Lane | Required |") - lines.append("|------|------|-----------------|------|----------|") - for s in tmpl['slots']: - lines.append( - f"| {s['name']} | {s['slot_type']} | {s['default_duration']}s " - f"| {s['lane']} | {'Yes' if s['required'] else 'No'} |" - ) - lines.append("") - return _text_result("\n".join(lines)) - - -async def handle_apply_template(arguments: dict) -> Sequence[TextContent]: - template_name = arguments["template_name"] - clips_raw = arguments["clips"] - output_path = _validate_output_path(arguments["output_path"], anchor_dir=PROJECTS_DIR) - fps = arguments.get("fps", 24) - - # Convert raw clips dict to ClipSpec objects - clips_map = {} - for slot_name, spec_data in clips_raw.items(): - if isinstance(spec_data, dict): - clips_map[slot_name] = ClipSpec( - asset_id=spec_data.get("asset_id"), - src=spec_data.get("src"), - name=spec_data.get("name", slot_name), - duration=spec_data.get("duration"), - ) - - result_path = apply_template(template_name, clips_map, output_path, fps) - return _text_result(( - f"Applied template '{template_name}' with {len(clips_map)} clips.\n" - f"Saved to: `{result_path}`" - )) - - -async def handle_relink_media(arguments: dict) -> Sequence[TextContent]: - dry_run = arguments.get("dry_run", False) - if dry_run: - filepath = _validate_filepath(arguments["filepath"], ('.fcpxml', '.fcpxmld')) - modifier = FCPXMLModifier(filepath) - result = modifier.relink_media( - arguments["find"], arguments["replace"], dry_run=True - ) - footer = "Dry run — no file written." - else: - filepath, output_path, modifier = _setup_modifier(arguments, "_relinked") - result = modifier.relink_media(arguments["find"], arguments["replace"]) - saved = modifier.save(output_path) - footer = f"Saved to: {saved}" - - if not result["relinked"]: - return _text_result( - f"No media paths matched prefix '{arguments['find']}' " - f"({result['total_assets']} assets scanned). Nothing to relink." - ) - - lines = [ - f"{'Would relink' if dry_run else 'Relinked'} " - f"{result['relinked']} media reference(s) " - f"across {result['total_assets']} asset(s):", - "", - ] - missing = 0 - for change in result["changes"]: - mark = "✓" if change["target_exists"] else "⚠ target missing" - if not change["target_exists"]: - missing += 1 - lines.append(f" {change['asset']}: {change['new']} [{mark}]") - if missing: - lines.append("") - lines.append( - f"⚠ {missing} new path(s) do not exist on this machine — " - f"FCP will show those clips as missing until the media is present." - ) - lines.append("") - lines.append(footer) - return _text_result("\n".join(lines)) - - -async def handle_push_to_fcp(arguments: dict) -> Sequence[TextContent]: - from fcpxml.live import push_to_fcp - - filepath = _validate_filepath(arguments["filepath"], ('.fcpxml', '.fcpxmld')) - - # Flat files get an options-injected sibling copy (never touch the - # original); the copy path goes through the same write sandbox as - # every other derived output. - import_copy = None - if Path(filepath).suffix.lower() == '.fcpxml': - anchor = str(Path(filepath).resolve().parent) - import_copy = _validate_output_path( - generate_output_path(filepath, "_import"), anchor_dir=anchor - ) - - result = push_to_fcp( - filepath, - library_location=arguments.get("library_location"), - suppress_warnings=arguments.get("suppress_warnings", True), - copy_assets=arguments.get("copy_assets"), - import_copy_path=import_copy, - ) - lines = [ - f"Sent to Final Cut Pro: {result['sent']}", - f"FCP {'was launched' if result['launched_fcp'] else 'was already running'} — " - f"import happens in-app (libraries/events are created or merged per import-options).", - ] - if arguments.get("library_location"): - lines.append(f"Target library: {arguments['library_location']}") - lines.append( - "Note: Apple offers no programmatic export — to round-trip edits " - "back, use File > Export XML in FCP." - ) - return _text_result("\n".join(lines)) - - -async def handle_list_fcp_libraries(arguments: dict) -> Sequence[TextContent]: - from fcpxml.live import list_fcp_libraries - - try: - libraries = list_fcp_libraries( - allow_launch=arguments.get("allow_launch", False) - ) - except RuntimeError as exc: - return _text_result(str(exc)) - - if not libraries: - return _text_result("Final Cut Pro is running but reports no open libraries.") - - lines = [f"Open libraries in Final Cut Pro ({len(libraries)}):", ""] - for lib in libraries: - lines.append(f"📚 {lib['name']}") - for event in lib["events"]: - lines.append(f" └─ {event['name']}") - for proj in event["projects"]: - lines.append(f" • {proj}") - return _text_result("\n".join(lines)) - - -# ============================================================================ -# TOOL DISPATCH -# ============================================================================ - -TOOL_HANDLERS = { - # Read - "list_projects": handle_list_projects, - "analyze_timeline": handle_analyze_timeline, - "list_clips": handle_list_clips, - "list_markers": handle_list_markers, - "find_short_cuts": handle_find_short_cuts, - "find_long_clips": handle_find_long_clips, - "list_keywords": handle_list_keywords, - "export_edl": handle_export_edl, - "export_csv": handle_export_csv, - "analyze_pacing": handle_analyze_pacing, - "list_library_clips": handle_list_library_clips, - # QC - "detect_flash_frames": handle_detect_flash_frames, - "detect_duplicates": handle_detect_duplicates, - "detect_gaps": handle_detect_gaps, - # Write - "add_marker": handle_add_marker, - "batch_add_markers": handle_batch_add_markers, - "trim_clip": handle_trim_clip, - "reorder_clips": handle_reorder_clips, - "add_transition": handle_add_transition, - "change_speed": handle_change_speed, - "add_zoom": handle_add_zoom, - "delete_clips": handle_delete_clips, - "split_clip": handle_split_clip, - "insert_clip": handle_insert_clip, - # Batch Fix - "fix_flash_frames": handle_fix_flash_frames, - "rapid_trim": handle_rapid_trim, - "fill_gaps": handle_fill_gaps, - "validate_timeline": handle_validate_timeline, - # Generation - "auto_rough_cut": handle_auto_rough_cut, - "generate_montage": handle_generate_montage, - "generate_ab_roll": handle_generate_ab_roll, - # Beat Sync - "import_beat_markers": handle_import_beat_markers, - "snap_to_beats": handle_snap_to_beats, - # SRT / Transcript - "import_srt_markers": handle_import_srt_markers, - "import_transcript_markers": handle_import_transcript_markers, - # Connected Clips & Compound Clips (v0.5.0) - "list_connected_clips": handle_list_connected_clips, - "add_connected_clip": handle_add_connected_clip, - "list_compound_clips": handle_list_compound_clips, - # Roles (v0.5.0) - "list_roles": handle_list_roles, - "assign_role": handle_assign_role, - "filter_by_role": handle_filter_by_role, - "export_role_stems": handle_export_role_stems, - # Timeline Diff (v0.5.0) - "diff_timelines": handle_diff_timelines, - # Social Media Reformat (v0.5.0) - "reformat_timeline": handle_reformat_timeline, - # Silence Detection (v0.5.0) - "detect_media_silence": handle_detect_media_silence, - "remove_media_silence": handle_remove_media_silence, - "transcribe_media": handle_transcribe_media, - "edit_by_transcript": handle_edit_by_transcript, - "remove_filler_words": handle_remove_filler_words, - "transcript_markers": handle_transcript_markers, - "generate_dynamic_subtitles": handle_generate_dynamic_subtitles, - "detect_beats": handle_detect_beats, - "detect_silence_candidates": handle_detect_silence_candidates, - "remove_silence_candidates": handle_remove_silence_candidates, - # NLE Export (v0.5.0) - "export_resolve_xml": handle_export_resolve_xml, - "export_fcp7_xml": handle_export_fcp7_xml, - # v0.6.0 - "list_effects": handle_list_effects, - "add_audio": handle_add_audio, - "create_compound_clip": handle_create_compound_clip, - "flatten_compound_clip": handle_flatten_compound_clip, - "list_templates": handle_list_templates, - "apply_template": handle_apply_template, - # v0.8.0 - "relink_media": handle_relink_media, - # v0.9.0 — Live mode - "push_to_fcp": handle_push_to_fcp, - "list_fcp_libraries": handle_list_fcp_libraries, -} +async def list_tools() -> list: + """Concatenate every category module's Tool() schemas.""" + tools = [] + for module in _CATEGORY_MODULES: + tools.extend(module.TOOLS) + return tools + + +TOOL_HANDLERS = {} +for _module in _CATEGORY_MODULES: + TOOL_HANDLERS.update(_module.HANDLERS) +del _module @server.call_tool() diff --git a/code/server_tools/__init__.py b/code/server_tools/__init__.py new file mode 100644 index 0000000..179fecb --- /dev/null +++ b/code/server_tools/__init__.py @@ -0,0 +1,8 @@ +"""Tool handlers and schemas for the FCPXML MCP server, split by category. + +server.py is the composition root: it imports each module's TOOLS/HANDLERS +and concatenates them for list_tools()/TOOL_HANDLERS. Each module here owns +one category (see Engine/docs/03_SERVER_TOOLS.md) — its Tool() schemas and +handle_<name> functions live together, so a tool's contract and its +implementation are never in different files. +""" diff --git a/code/server_tools/_shared.py b/code/server_tools/_shared.py new file mode 100644 index 0000000..cf359c3 --- /dev/null +++ b/code/server_tools/_shared.py @@ -0,0 +1,829 @@ +"""Shared internal helpers used by tool handlers across categories. + +Extracted from server.py — validation, formatting, and small parsing utilities +that more than one server_tools/*.py module needs. +""" + +from __future__ import annotations + +import json +import os +import re +from pathlib import Path +from typing import Any, Sequence + +from mcp.types import TextContent + +from fcpxml.media_intel import media_src_to_path +from fcpxml.models import ( + DuplicateGroup, + FlashFrame, + FlashFrameSeverity, + GapInfo, + Timecode, + TimeValue, +) +from fcpxml.parser import FCPXMLParser +from fcpxml.rough_cut import RoughCutGenerator +from fcpxml.transcribe import invert_ranges, merge_ranges, transcribe +from fcpxml.writer import FCPXMLModifier + +PROJECTS_DIR = os.environ.get("FCP_PROJECTS_DIR", os.path.expanduser("~/Movies")) + +_SANDBOX_ENABLED = "FCP_PROJECTS_DIR" in os.environ + +MAX_FILE_SIZE = 100 * 1024 * 1024 + +MAX_MEDIA_FILE_SIZE = 32 * 1024 * 1024 * 1024 + +_MAX_JSON_DEPTH = 50 + +def _check_json_depth(obj: object, _depth: int = 0) -> None: + """Reject JSON structures nested beyond _MAX_JSON_DEPTH. + + Prevents denial-of-service via deeply nested objects that exhaust the + call stack or memory during downstream processing. Called after + json.load() since Python's json module has no built-in depth limit. + """ + if _depth > _MAX_JSON_DEPTH: + raise ValueError( + f"JSON nesting depth exceeds {_MAX_JSON_DEPTH} — " + "file may be malformed or adversarial" + ) + if isinstance(obj, dict): + for v in obj.values(): + _check_json_depth(v, _depth + 1) + elif isinstance(obj, list): + for item in obj: + _check_json_depth(item, _depth + 1) + +def _validate_filepath( + filepath: str, + allowed_extensions: tuple[str, ...] | None = None, + max_size: int = MAX_FILE_SIZE, +) -> str: + """Validate a user-provided file path against traversal and size attacks. + + Resolves symlinks, blocks null bytes, enforces extension whitelist, and + checks file size before any parsing takes place. + + ``max_size`` defaults to the document limit; callers handling source + media pass ``MAX_MEDIA_FILE_SIZE``, since media is streamed rather than + parsed into memory (see the constant for why). + + Raises: + ValueError: For invalid paths (null bytes, bad extensions, oversized). + FileNotFoundError: When the resolved path does not exist. + """ + if '\x00' in filepath: + raise ValueError("Invalid file path: null byte detected") + + resolved = Path(filepath).resolve() + + if not resolved.exists(): + raise FileNotFoundError(f"File not found: {filepath}") + + # .fcpxmld bundles are directories (a package wrapping Info.fcpxml plus + # sidecar data files for object tracking / Cinematic mode). The size + # check applies to the inner Info.fcpxml, which is what gets parsed. + if resolved.is_dir(): + if resolved.suffix.lower() != '.fcpxmld': + raise ValueError(f"Not a regular file: {filepath}") + inner = resolved / 'Info.fcpxml' + if not inner.is_file(): + raise ValueError(f"Invalid bundle (no Info.fcpxml): {filepath}") + size_target = inner + elif not resolved.is_file(): + raise ValueError(f"Not a regular file: {filepath}") + else: + size_target = resolved + + if allowed_extensions and resolved.suffix.lower() not in allowed_extensions: + raise ValueError( + f"Invalid file type '{resolved.suffix}'. " + f"Allowed: {', '.join(allowed_extensions)}" + ) + + if size_target.stat().st_size > max_size: + size_mb = size_target.stat().st_size / (1024 * 1024) + raise ValueError(f"File too large ({size_mb:.1f} MB). Maximum: {max_size // (1024 * 1024)} MB") + + return str(resolved) + +def _validate_output_path(output_path: str, *, anchor_dir: str | None = None) -> str: + """Validate an output path with optional sandbox enforcement. + + Resolves traversal, blocks null bytes, ensures parent exists, and — when + *anchor_dir* is provided — verifies the resolved output lives under that + directory. This prevents LLM-generated tool calls from writing to + arbitrary filesystem locations (e.g. ``/etc/cron.d/backdoor``). + + Args: + output_path: The raw output path to validate. + anchor_dir: If set, the resolved output must be a child of this + directory. Typically the parent directory of the input file so + outputs stay co-located with their sources. + + Raises: + ValueError: For null bytes, missing parent, or sandbox escape. + """ + if '\x00' in output_path: + raise ValueError("Invalid output path: null byte detected") + + resolved = Path(output_path).resolve() + + if not resolved.parent.exists(): + raise ValueError(f"Output directory does not exist: {resolved.parent}") + + if anchor_dir is not None: + anchor = Path(anchor_dir).resolve() + try: + resolved.relative_to(anchor) + except ValueError: + raise ValueError( + f"Output path escapes allowed directory: " + f"{resolved} is not under {anchor}" + ) + + return str(resolved) + +def _validate_directory(directory: str, *, allowed_root: str | None = None) -> str: + """Validate a user-provided directory path against traversal and injection. + + Resolves symlinks, blocks null bytes, and verifies the path is a real + directory. When *allowed_root* is given, the resolved path must be a + descendant of (or equal to) that root — preventing filesystem enumeration + beyond the project workspace. + + Raises: + ValueError: For invalid paths (null bytes, not a directory, sandbox escape). + """ + if '\x00' in directory: + raise ValueError("Invalid directory path: null byte detected") + + resolved = Path(directory).resolve() + + if not resolved.is_dir(): + raise ValueError(f"Not a valid directory: {directory}") + + if allowed_root is not None: + root = Path(allowed_root).resolve() + try: + resolved.relative_to(root) + except ValueError: + raise ValueError( + f"Directory escapes allowed root: " + f"{resolved} is not under {root}" + ) + + return str(resolved) + +def find_fcpxml_files(directory: str) -> list[str]: + """Find all FCPXML files in a directory.""" + path = Path(directory) + files = list(str(f) for f in path.rglob("*.fcpxml")) + files.extend(str(f) for f in path.rglob("*.fcpxmld")) + return sorted(files) + +def format_timecode(tc) -> str: + """Format a Timecode object to SMPTE string.""" + return tc.to_smpte() if tc else "00:00:00:00" + +def format_duration(seconds: float) -> str: + """Format seconds into human-readable duration.""" + if seconds < 1: + return f"{seconds*1000:.0f}ms" + elif seconds < 60: + return f"{seconds:.2f}s" + return f"{int(seconds // 60)}m {seconds % 60:.1f}s" + +def _format_clip_table(clips: list, header: str) -> str: + """Render a list of clips as a markdown table with timecodes and durations. + + Shared by handlers that filter clips by duration threshold + (find_short_cuts, find_long_clips). + """ + result = f"{header}\n\n| Name | TC | Duration |\n|------|----|---------|\n" + result += "\n".join( + f"| {c.name} | {format_timecode(c.start)} | {format_duration(c.duration_seconds)} |" + for c in clips + ) + return result + +def _markdown_table(headers: list[str], rows: list[list[str]]) -> str: + """Build a markdown table from headers and rows. + + Returns header row, separator row, and data rows as a single string. + Callers avoid repeating the ``| H1 | H2 |\\n|---|---|`` boilerplate + that appears in 15+ handlers. + """ + header_line = "| " + " | ".join(headers) + " |" + sep_line = "|" + "|".join("------" for _ in headers) + "|" + data_lines = "\n".join( + "| " + " | ".join(str(c) for c in row) + " |" for row in rows + ) + return f"{header_line}\n{sep_line}\n{data_lines}" + +def _format_batch_result( + title: str, + summary: dict[str, str], + headers: list[str], + rows: list[list[str]], + output_path: str, +) -> str: + """Build a standard batch-operation result with summary, table, and save footer. + + Used by batch fix handlers (flash frames, rapid trim, fill gaps) that all + share the same markdown structure: ``# Title → ## Summary → ## Details table + → Saved to`` footer. + """ + summary_lines = "\n".join(f"- **{k}**: {v}" for k, v in summary.items()) + table = _markdown_table(headers, rows) + return ( + f"# {title}\n\n" + f"## Summary\n{summary_lines}\n\n" + f"## Details\n{table}\n\n" + f"Saved to: `{output_path}`" + ) + +def _fmt_suggestions(suggestions: list[str]) -> str: + """Format pacing suggestions as markdown list (Python 3.10 compatible).""" + if not suggestions: + return "- Pacing looks good!" + nl = "\n" + return nl.join(f"- {s}" for s in suggestions) + +def generate_output_path(input_path: str, suffix: str = "_modified") -> str: + """Generate output path from input path. + + The suffix is sanitized to prevent path-component injection — only + alphanumeric, hyphen, underscore, and dot characters survive. + """ + # Strip anything that could inject path separators or traversal sequences + clean_suffix = re.sub(r'[^a-zA-Z0-9._-]', '', suffix) + if not clean_suffix: + clean_suffix = "_modified" + p = Path(input_path) + return str(p.parent / f"{p.stem}{clean_suffix}{p.suffix}") + +def _parse_project(filepath: str): + """Parse an FCPXML file and return the project with its primary timeline.""" + filepath = _validate_filepath(filepath, ('.fcpxml', '.fcpxmld')) + project = FCPXMLParser().parse_file(filepath) + if not project.timelines: + return None, None + return project, project.primary_timeline + +def _text_result(text: str) -> list[TextContent]: + """Wrap a string in the MCP TextContent list that every tool handler returns.""" + return [TextContent(type="text", text=text)] + +def _no_timeline(): + """Standard response when no timelines are found.""" + return _text_result("No timelines found") + +def _require_timeline(filepath: str): + """Parse FCPXML and return (project, timeline), raising if no timeline exists. + + Centralises the repeated _parse_project + _no_timeline guard that + appears in every read-only timeline handler. Returns a tuple so + callers can destructure directly:: + + project, tl = _require_timeline(arguments["filepath"]) + """ + project, tl = _parse_project(filepath) + if not tl: + raise _NoTimelineError() + return project, tl + +class _NoTimelineError(Exception): + """Sentinel raised by _require_timeline when no timelines exist.""" + +def _resolve_io_paths( + arguments: dict, + suffix: str = "_modified", +) -> tuple[str, str]: + """Validate input filepath and resolve the output path. + + Shared foundation for every handler that reads an FCPXML and writes + a derived file. Validates the input, falls back to a suffixed + output name when ``output_path`` is not supplied, and sandbox-checks + the result. + + Args: + arguments: Tool arguments dict (must contain ``filepath``; may + contain ``output_path``). + suffix: Default output filename suffix when ``output_path`` is + not provided (e.g. ``"_modified"``, ``"_beats"``). + + Returns: + ``(filepath, output_path)`` tuple with both paths validated. + """ + filepath = _validate_filepath(arguments["filepath"], ('.fcpxml', '.fcpxmld')) + # Anchor write operations to the input file's directory so LLM-generated + # tool calls cannot write to arbitrary filesystem locations (e.g. + # /etc/cron.d/backdoor). When the explicit sandbox is off, the anchor + # still prevents writes outside the source directory tree. + # `output_dir` is where the caller wants the file written, not merely a + # sandbox boundary: the app's "Pasta do projeto" promises that everything + # generated lands there. Deriving the name from the input but keeping the + # input's directory made every cross-directory call fail its own anchor + # check ("output path escapes allowed directory"), so the setting silently + # only worked when it pointed at the directory the file was already going + # to. An explicit `output_path` still wins, and still has to sit inside + # the anchor. + output_dir = arguments.get("output_dir") + if output_dir: + anchor = _validate_directory(str(output_dir)) + default_output = str(Path(anchor) / Path(generate_output_path(filepath, suffix)).name) + else: + anchor = str(Path(filepath).resolve().parent) + default_output = generate_output_path(filepath, suffix) + output_path = _validate_output_path( + arguments.get("output_path") or default_output, + anchor_dir=anchor, + ) + return filepath, output_path + +def _setup_modifier( + arguments: dict, + suffix: str = "_modified", +) -> tuple[str, str, "FCPXMLModifier"]: + """Common setup for write handlers: validate paths and create modifier. + + Consolidates the repeated validate-filepath → resolve-output-path → + create-modifier boilerplate shared by 18+ write handlers. + + Args: + arguments: Tool arguments dict (must contain ``filepath``; may + contain ``output_path``). + suffix: Default output filename suffix when ``output_path`` is + not provided (e.g. ``"_modified"``, ``"_flash_fixed"``). + + Returns: + ``(filepath, output_path, modifier)`` tuple ready for the + handler's domain-specific operation. + """ + filepath, output_path = _resolve_io_paths(arguments, suffix) + modifier = FCPXMLModifier(filepath) + return filepath, output_path, modifier + +def _setup_generator( + arguments: dict, + suffix: str = "_roughcut", +) -> tuple[str, str, "RoughCutGenerator"]: + """Common setup for generation handlers: validate paths and create generator. + + Args: + arguments: Tool arguments dict (must contain ``filepath`` and + ``output_path``). + suffix: Default output filename suffix. + + Returns: + ``(filepath, output_path, generator)`` tuple. + """ + filepath, output_path = _resolve_io_paths(arguments, suffix) + generator = RoughCutGenerator(filepath) + return filepath, output_path, generator + +def _parse_timestamp_parts( + parts: list[str], *, frame_rate: float = 24.0 +) -> float | None: + """Convert colon-separated timestamp parts to total seconds. + + Handles 2-part (M:SS), 3-part (H:MM:SS / HH:MM:SS.ms), and + 4-part (HH:MM:SS:FF SMPTE) formats. Returns ``None`` when the + part count is unrecognised so callers can skip. + + Args: + parts: Colon-split timestamp components. + frame_rate: FPS used to convert the frame component of SMPTE + timecodes into fractional seconds (default 24.0). + """ + if len(parts) == 2: + return int(parts[0]) * 60 + float(parts[1]) + elif len(parts) == 3: + return int(parts[0]) * 3600 + int(parts[1]) * 60 + float(parts[2]) + elif len(parts) == 4: + # SMPTE: HH:MM:SS:FF — convert frames to fractional seconds + base = int(parts[0]) * 3600 + int(parts[1]) * 60 + float(parts[2]) + frames = int(parts[3]) + return base + (frames / frame_rate) if frame_rate > 0 else base + return None + +def _raw_markers_to_batch( + raw_markers: list[dict], + marker_type: str = "chapter", + max_label: int | None = None, +) -> list[dict]: + """Convert raw {seconds, text} marker dicts to batch_add_markers format. + + Shared by import_srt_markers and import_transcript_markers. + """ + batch = [] + for m in raw_markers: + label = m["text"] + if max_label and len(label) > max_label: + label = label[:max_label] + batch.append({ + "timecode": f"{m['seconds']}s", + "name": label, + "marker_type": marker_type.upper(), + }) + return batch + +def _extract_subtitle_blocks(text: str, *, strip_vtt_tags: bool = False) -> list[dict]: + """Extract timestamp/text pairs from subtitle cue blocks (SRT or VTT). + + Both SRT and VTT use the same ``start --> end`` cue syntax with + text lines underneath; only header stripping and tag cleaning differ. + """ + markers = [] + blocks = re.split(r'\n\s*\n', text.strip()) + for block in blocks: + lines = block.strip().split('\n') + if len(lines) < 2: + continue + ts_line = None + text_lines = [] + for line in lines: + if '-->' in line: + ts_line = line + elif ts_line is not None: + if strip_vtt_tags: + line = re.sub(r'<[^>]+>', '', line) + cleaned = line.strip() + if cleaned: + text_lines.append(cleaned) + if not ts_line or not text_lines: + continue + start_str = ts_line.split('-->')[0].strip().replace(',', '.') + seconds = _parse_timestamp_parts(start_str.split(':')) + if seconds is not None: + markers.append({'seconds': seconds, 'text': ' '.join(text_lines)}) + return markers + +def parse_srt(text: str) -> list[dict]: + """Parse SRT subtitle format into timestamp/text pairs.""" + return _extract_subtitle_blocks(text) + +def parse_vtt(text: str) -> list[dict]: + """Parse WebVTT subtitle format into timestamp/text pairs.""" + text = re.sub(r'^WEBVTT.*?\n', '', text, flags=re.MULTILINE) + text = re.sub(r'NOTE\n.*?\n\n', '', text, flags=re.DOTALL) + return _extract_subtitle_blocks(text, strip_vtt_tags=True) + +def parse_transcript_timestamps(text: str) -> list[dict]: + """Parse timestamped text (YouTube description format) into markers. + + Supports formats like: + 0:00 Introduction + 00:01:30 Main Topic + 1:05:30 Conclusion + 00:00:00:00 SMPTE timecode + """ + markers = [] + for line in text.strip().split('\n'): + line = line.strip() + if not line: + continue + match = re.match(r'^(\d{1,2}:\d{2}(?::\d{2}){0,2})\s+(.+)$', line) + if match: + seconds = _parse_timestamp_parts(match.group(1).split(':')) + if seconds is not None: + markers.append({'seconds': seconds, 'text': match.group(2).strip()}) + return markers + +def _detect_flash_frames( + tl: Any, *, critical_threshold: int = 2, warning_threshold: int = 6, +) -> list: + """Find clips shorter than *warning_threshold* frames. + + Returns a list of ``FlashFrame`` objects sorted by severity. Shared by + ``handle_detect_flash_frames`` and ``handle_validate_timeline`` so the + detection logic lives in exactly one place. + """ + fps = tl.frame_rate + flash_frames: list[FlashFrame] = [] + for clip in tl.clips: + duration_frames = int(clip.duration_seconds * fps) + if duration_frames < warning_threshold: + severity = ( + FlashFrameSeverity.CRITICAL + if duration_frames < critical_threshold + else FlashFrameSeverity.WARNING + ) + flash_frames.append(FlashFrame( + clip_name=clip.name, clip_id=clip.name, + start=clip.start, duration_frames=duration_frames, + duration_seconds=clip.duration_seconds, severity=severity, + )) + return flash_frames + +def _detect_gaps(tl: Any, *, min_gap_frames: int = 1) -> list: + """Find inter-clip gaps of at least *min_gap_frames* length. + + Returns a list of ``GapInfo`` objects. Shared by ``handle_detect_gaps`` + and ``handle_validate_timeline``. + """ + fps = tl.frame_rate + min_gap_seconds = min_gap_frames / fps + gaps: list[GapInfo] = [] + sorted_clips = sorted(tl.clips, key=lambda c: c.start.seconds) + for i in range(len(sorted_clips) - 1): + current_end = sorted_clips[i].end.seconds + next_start = sorted_clips[i + 1].start.seconds + gap_duration = next_start - current_end + if gap_duration >= min_gap_seconds: + gaps.append(GapInfo( + start=Timecode(frames=int(current_end * fps), frame_rate=fps), + duration_frames=int(gap_duration * fps), + duration_seconds=gap_duration, + previous_clip=sorted_clips[i].name, + next_clip=sorted_clips[i + 1].name, + )) + return gaps + +def _detect_duplicate_groups(tl: Any, *, mode: str = "same_source") -> list: + """Group clips that share a source media reference. + + Returns a list of ``DuplicateGroup`` objects. Shared by + ``handle_detect_duplicates`` and ``handle_validate_timeline``. + """ + source_groups: dict[str, list[dict]] = {} + for clip in tl.clips: + source_key = clip.media_path or clip.name + if source_key not in source_groups: + source_groups[source_key] = [] + source_groups[source_key].append({ + 'name': clip.name, + 'start': clip.start.seconds, + 'duration': clip.duration_seconds, + 'source_start': clip.source_start.seconds if clip.source_start else 0, + 'source_duration': clip.duration_seconds, + 'timecode': format_timecode(clip.start), + }) + + duplicates: list[DuplicateGroup] = [] + for source_key, clips in source_groups.items(): + if len(clips) <= 1: + continue + group = DuplicateGroup( + source_ref=source_key, + source_name=source_key.split('/')[-1] if '/' in source_key else source_key, + clips=clips, + ) + if mode == "same_source": + duplicates.append(group) + elif mode == "overlapping_ranges" and group.has_overlapping_ranges: + duplicates.append(group) + elif mode == "identical": + seen_ranges: set[tuple] = set() + identical_clips = [] + for c in clips: + range_key = (c['source_start'], c['source_duration']) + if range_key in seen_ranges: + identical_clips.append(c) + seen_ranges.add(range_key) + if identical_clips: + group.clips = identical_clips + duplicates.append(group) + return duplicates + +AUDIO_MEDIA_EXTENSIONS = ( + '.wav', '.aif', '.aiff', '.mp3', '.m4a', '.aac', '.flac', '.mov', '.mp4', +) + +_DIARIZATION_INSTALL_HINT = ( + "\n\nInstall the optional diarization extra:\n\n" + " pip install 'fcp-mcp-server[diarization]'\n\n" + "and set a HuggingFace token with access to " + "pyannote/speaker-diarization-3.1 (pass hf_token= or persist one via " + "save_hf_token)." +) + +_FEATURES_INSTALL_HINT = ( + "\n\nInstall the optional media-intelligence extra:\n\n" + " pip install 'fcp-mcp-server[intelligence]'" +) + +def _voice_analysis_config_text(config: dict) -> str: + w = config["emphasis_weights"] + text = "# Voice Analysis Settings\n\n" + text += _markdown_table( + ["Setting", "Value"], + [ + ["Energy threshold", f"{config['energy_threshold']:.2f}"], + ["Peak selection", f"top {config['peak_percentile']:.1%} of words"], + ["Emphasis floor", f"{config['emphasis_floor']:.2f}"], + ["Emotion detection", "on" if config["emotion_enabled"] else "off"], + ["Emotion sensitivity", f"{config['emotion_sensitivity']:.2f}"], + ], + ) + "\n\n## Emphasis Weights\n" + text += _markdown_table( + ["Factor", "Weight"], + [[k.replace("_", " ").title(), f"{v:.2f}"] for k, v in w.items()], + ) + return text + +def _apply_placed_action(modifier, clip_el, action, clip_start: float) -> str: + """Apply one non-cut action to the clip that hosts it. + + ``clip_start`` is where that clip begins on the timeline; the writer + wants times relative to the clip's own head, so the rebase happens here + — the single place that knows about the conversion. The clip *element* + is passed through rather than its name: after a cut the pieces share a + name, and a name lookup would land every edit on the first piece. + """ + rel_start = action.start - clip_start + rel_end = action.end - clip_start + + if action.kind == "zoom": + # Only forward an explicit ease — otherwise add_zoom's own default + # (a fast ramp in, instant snap back out) is what should apply. + zoom_args = {} + if action.params.get("ease") is not None: + zoom_args["ease"] = float(action.params["ease"]) + if action.params.get("ease_out") is not None: + zoom_args["ease_out"] = float(action.params["ease_out"]) + modifier.add_zoom( + clip_id=clip_el, + start=rel_start, + end=rel_end, + scale=float(action.params.get("scale", 1.3)), + **zoom_args, + ) + return f"zoom {action.params.get('scale', 1.3):.2f}x" + + if action.kind == "text": + modifier.add_text_title( + clip_el, + action.params["content"], + offset=modifier.snap_seconds_to_frame(rel_start).to_fcpxml(), + duration=modifier.snap_seconds_to_frame(action.duration).to_fcpxml(), + ) + return f"text \"{action.params['content'][:24]}\"" + + # marker + modifier.add_marker( + clip_id=clip_el, + timecode=modifier.snap_seconds_to_frame(rel_start).to_fcpxml(), + name=action.params.get("content") or action.reason or "Voice action", + note=action.reason or None, + ) + return "marker" + +def _speaker_table(profiles: Sequence[dict]) -> str: + """Who was detected, ordered by how much of the runtime each holds.""" + return _markdown_table( + ["ID", "Name", "Share", "Speaking", "Lines", "Avg line"], + [ + [ + p["id"], + p.get("name", ""), + f"{p['share']:.0%}", + format_duration(p["speaking_seconds"]), + str(p["segment_count"]), + f"{p['avg_segment']:.1f}s", + ] + for p in profiles + ], + ) + +TRANSCRIBE_MAX_MEDIA = 10 + +_TRANSCRIBE_INSTALL_HINT = ( + "\n\nInstall the optional transcription extra:\n\n" + " pip install 'fcp-mcp-server[transcribe]'\n\n" + "or run via uvx:\n\n" + " uvx --from \"fcp-mcp-server[transcribe]\" fcp-mcp-server" +) + +def _transcript_json_path(media_path: str, output_dir: str | None = None) -> Path: + """Where the ``_transcript.json`` for ``media_path`` lives. + + When ``output_dir`` (the user-selected project folder) is set, the + transcript is saved/read there instead of next to the source media. + """ + p = Path(media_path) + if output_dir: + directory = Path(output_dir).expanduser() + directory.mkdir(parents=True, exist_ok=True) + return directory / f"{p.stem}_transcript.json" + return p.with_name(p.stem + "_transcript.json") + +def _load_or_transcribe( + media_path: str, model: str, language: str | None, output_dir: str | None = None +) -> tuple[dict | None, str]: + """Load a cached ``_transcript.json`` for a media file, else transcribe and cache it. + + Returns ``(transcript, "")`` or ``(None, reason)``. The cache makes + transcription a one-time cost per media file across all transcript tools. + """ + json_path = _transcript_json_path(media_path, output_dir) + if json_path.is_file(): + try: + with open(json_path) as f: + data = json.load(f) + if isinstance(data, dict) and isinstance(data.get("words"), list): + return data, "" + except (OSError, json.JSONDecodeError, UnicodeDecodeError): + pass # unreadable cache falls through to re-transcribe + result = transcribe(media_path, model_size=model, language=language) + if result is None: + return None, "untranscribable (faster-whisper not installed or media unreadable)" + anchor = str(Path(output_dir).expanduser()) if output_dir else str(Path(media_path).parent) + out_path = _validate_output_path(str(json_path), anchor_dir=anchor) + with open(out_path, "w") as f: + json.dump({"source": Path(media_path).name, **result}, f, indent=2) + return result, "" + +def _cut_transcript_spans(modifier, clip_filter, model, language, padding, spans_fn, keep_only=False, output_dir=None): + """Shared cut engine for transcript-driven editing. + + ``spans_fn(words) -> [(start, end), ...]`` in source seconds. Spans are + padded, clamped to each clip's used source window, optionally inverted + (keep_only), snapped to the frame grid, and cut with ripple. + """ + to_frame = modifier.snap_seconds_to_frame + + cache: dict[str, tuple] = {} + cuts_made: list[tuple[str, int, float]] = [] + skipped: list[tuple[str, str]] = [] + spine_clips = [el for _, el in modifier._iter_spine_clips()] + for el in spine_clips: + name = el.get("name", "") + if clip_filter and name != clip_filter: + continue + src = modifier.resources.get(el.get("ref", ""), {}).get("src", "") + media_path = media_src_to_path(src) + if not media_path or not Path(media_path).is_file(): + skipped.append((name, "media file missing")) + continue + if media_path not in cache: + if len(cache) >= TRANSCRIBE_MAX_MEDIA: + skipped.append((name, f"transcription cap reached ({TRANSCRIBE_MAX_MEDIA} media files)")) + continue + cache[media_path] = _load_or_transcribe(media_path, model, language, output_dir) + data, reason = cache[media_path] + if data is None: + skipped.append((name, reason)) + continue + + clip_source_start = modifier.source_file_start(el).to_seconds() + clip_duration = modifier._parse_time(el.get("duration", "0s")).to_seconds() + window_start = clip_source_start + window_end = clip_source_start + clip_duration + + spans = spans_fn(data.get("words", [])) + padded = merge_ranges([(s - padding, e + padding) for s, e in spans]) + clamped = [ + (max(s, window_start), min(e, window_end)) + for s, e in padded + if min(e, window_end) > max(s, window_start) + ] + if keep_only: + if not clamped: + # Never delete a whole clip just because nothing matched in it. + skipped.append((name, "no phrase matches — left untouched (keep_only)")) + continue + cut_source = invert_ranges(clamped, window_start, window_end) + else: + cut_source = clamped + cut_ranges = [ + (to_frame(s - clip_source_start), to_frame(e - clip_source_start)) + for s, e in cut_source + ] + cut_ranges = [(a, b) for a, b in cut_ranges if b > a] + if not cut_ranges: + continue + removed = modifier.cut_clip_ranges(el, cut_ranges) + if removed > TimeValue.zero(): + cuts_made.append((name, len(cut_ranges), removed.to_seconds())) + return cuts_made, skipped + +def _transcript_cut_report(title, summary_lines, cuts_made, skipped, output_path, footer): + if not cuts_made: + text = f"# {title}\n\nNo cuts to make — file unchanged (nothing saved)." + if skipped: + text += "\n\n## Skipped Clips\n" + _markdown_table( + ["Clip", "Reason"], [[name, reason] for name, reason in skipped] + ) + if any("faster-whisper" in reason for _, reason in skipped): + text += _TRANSCRIBE_INSTALL_HINT + return _text_result(text) + total_removed = sum(seconds for _, _, seconds in cuts_made) + result = f"# {title}\n\n## Summary\n" + result += "\n".join(summary_lines) + "\n" + result += f"- **Clips Cut**: {len(cuts_made)}\n- **Total Removed**: {format_duration(total_removed)}\n" + result += "\n## Cuts\n" + result += _markdown_table( + ["Clip", "Ranges Cut", "Removed"], + [[name, str(count), f"{seconds:.2f}s"] for name, count, seconds in cuts_made], + ) + "\n" + if skipped: + result += "\n## Skipped Clips\n" + _markdown_table( + ["Clip", "Reason"], [[name, reason] for name, reason in skipped] + ) + "\n" + result += f"\nSaved to: {output_path}\n\n{footer}" + return _text_result(result) diff --git a/code/server_tools/editing.py b/code/server_tools/editing.py new file mode 100644 index 0000000..2fa1225 --- /dev/null +++ b/code/server_tools/editing.py @@ -0,0 +1,649 @@ +"""Edição — tool schemas and handlers. + +Extracted from server.py; see Engine/docs/03_SERVER_TOOLS.md for the tool catalog. +""" + +from __future__ import annotations + +from typing import Sequence + +from mcp.types import TextContent, Tool + +from fcpxml.models import MarkerType +from fcpxml.writer import FCPXMLModifier +from server_tools._shared import ( + _format_batch_result, + _resolve_io_paths, + _setup_modifier, + _text_result, + format_duration, +) + +TOOLS = [ + Tool( + name="add_marker", + description="Add a marker at a specific timecode", + inputSchema={ + "type": "object", + "properties": { + "filepath": {"type": "string", "description": "Path to FCPXML file"}, + "timecode": {"type": "string", "description": "Position (00:00:10:00 or 10s)"}, + "name": {"type": "string", "description": "Marker label"}, + "marker_type": {"type": "string", "enum": ["standard", "chapter", "todo", "completed"], "default": "standard"}, + "note": {"type": "string", "description": "Optional note"}, + "output_path": {"type": "string", "description": "Output path (default: adds _modified suffix)"} + }, + "required": ["filepath", "timecode", "name"] + } + ), + Tool( + name="batch_add_markers", + description="Add multiple markers at once, or auto-generate at cuts/intervals", + inputSchema={ + "type": "object", + "properties": { + "filepath": {"type": "string"}, + "markers": { + "type": "array", + "items": { + "type": "object", + "properties": { + "timecode": {"type": "string"}, + "name": {"type": "string"}, + "marker_type": {"type": "string"}, + "note": {"type": "string"} + } + }, + "description": "List of markers to add" + }, + "auto_at_cuts": {"type": "boolean", "description": "Add marker at every cut"}, + "auto_at_intervals": {"type": "string", "description": "Add markers every N seconds (e.g., '30s')"}, + "output_path": {"type": "string"} + }, + "required": ["filepath"] + } + ), + Tool( + name="trim_clip", + description="Trim a clip's in-point and/or out-point", + inputSchema={ + "type": "object", + "properties": { + "filepath": {"type": "string"}, + "clip_id": {"type": "string", "description": "Clip name or ID"}, + "trim_start": {"type": "string", "description": "New in-point or delta (+1s, -10f)"}, + "trim_end": {"type": "string", "description": "New out-point or delta"}, + "ripple": {"type": "boolean", "default": True, "description": "Shift subsequent clips"}, + "output_path": {"type": "string"} + }, + "required": ["filepath", "clip_id"] + } + ), + Tool( + name="reorder_clips", + description="Move clips to a new position in the timeline", + inputSchema={ + "type": "object", + "properties": { + "filepath": {"type": "string"}, + "clip_ids": {"type": "array", "items": {"type": "string"}, "description": "Clips to move"}, + "target_position": {"type": "string", "description": "'start', 'end', timecode, or 'after:clip_id'"}, + "ripple": {"type": "boolean", "default": True}, + "output_path": {"type": "string"} + }, + "required": ["filepath", "clip_ids", "target_position"] + } + ), + Tool( + name="add_transition", + description="Add a transition between clips", + inputSchema={ + "type": "object", + "properties": { + "filepath": {"type": "string"}, + "clip_id": {"type": "string", "description": "Clip to add transition to"}, + "position": {"type": "string", "enum": ["start", "end", "both"], "default": "end"}, + "transition_type": {"type": "string", "enum": ["cross-dissolve", "fade-to-black", "fade-from-black", "wipe"], "default": "cross-dissolve"}, + "duration": {"type": "string", "default": "00:00:00:15"}, + "output_path": {"type": "string"} + }, + "required": ["filepath", "clip_id"] + } + ), + Tool( + name="change_speed", + description="Change clip playback speed (slow motion or speed up)", + inputSchema={ + "type": "object", + "properties": { + "filepath": {"type": "string"}, + "clip_id": {"type": "string"}, + "speed": {"type": "number", "description": "Speed multiplier (0.5 = half, 2.0 = double)"}, + "preserve_pitch": {"type": "boolean", "default": True}, + "output_path": {"type": "string"} + }, + "required": ["filepath", "clip_id", "speed"] + } + ), + Tool( + name="add_zoom", + description="Add a smooth ease-in/ease-out punch-in zoom to a clip, animating <adjust-transform>'s scale param via keyframes (100% -> scale -> 100%) entirely within [start, end] (clip-relative seconds, i.e. seconds from the clip's own head). The ease portions each last `ease` seconds; the zoom holds at `scale` in between. Replaces any existing zoom on the same clip rather than stacking.", + inputSchema={ + "type": "object", + "properties": { + "filepath": {"type": "string"}, + "clip_id": {"type": "string", "description": "Name/ID of the clip to zoom"}, + "start": {"type": "number", "description": "Clip-relative seconds where the ease-in begins"}, + "end": {"type": "number", "description": "Clip-relative seconds where the ease-out ends (back to 100%)"}, + "scale": {"type": "number", "default": 1.3, "description": "Zoom scale, e.g. 1.3 = 130%"}, + "ease": {"type": "number", "default": 0.3, "description": "Seconds for each of the ease-in/ease-out portions (must fit: 2*ease <= end-start)"}, + "position": {"type": "string", "default": "0 0", "description": "Optional pan offset \"x y\" applied for the duration of the transform"}, + "output_path": {"type": "string"} + }, + "required": ["filepath", "clip_id", "start", "end"] + } + ), + Tool( + name="delete_clips", + description="Delete clips from timeline", + inputSchema={ + "type": "object", + "properties": { + "filepath": {"type": "string"}, + "clip_ids": {"type": "array", "items": {"type": "string"}}, + "ripple": {"type": "boolean", "default": True, "description": "Close gaps after deletion"}, + "output_path": {"type": "string"} + }, + "required": ["filepath", "clip_ids"] + } + ), + Tool( + name="split_clip", + description="Split a clip at specified timecodes", + inputSchema={ + "type": "object", + "properties": { + "filepath": {"type": "string"}, + "clip_id": {"type": "string"}, + "split_points": {"type": "array", "items": {"type": "string"}, "description": "Timecodes to split at"}, + "output_path": {"type": "string"} + }, + "required": ["filepath", "clip_id", "split_points"] + } + ), + Tool( + name="insert_clip", + description="Insert a library clip onto the timeline at a specific position", + inputSchema={ + "type": "object", + "properties": { + "filepath": {"type": "string", "description": "Path to FCPXML file"}, + "asset_id": {"type": "string", "description": "Asset reference ID (e.g., 'r3')"}, + "asset_name": {"type": "string", "description": "Asset name (alternative to asset_id)"}, + "position": {"type": "string", "description": "'start', 'end', timecode, or 'after:clip_name'"}, + "duration": {"type": "string", "description": "Clip duration (if not using in/out points)"}, + "in_point": {"type": "string", "description": "Source in-point for subclip"}, + "out_point": {"type": "string", "description": "Source out-point for subclip"}, + "ripple": {"type": "boolean", "default": True, "description": "Shift subsequent clips"}, + "output_path": {"type": "string", "description": "Output path (default: adds _modified suffix)"} + }, + "required": ["filepath", "position"] + } + ), + Tool( + name="fix_flash_frames", + description="Automatically fix detected flash frames by extending neighbors or deleting", + inputSchema={ + "type": "object", + "properties": { + "filepath": {"type": "string", "description": "Path to FCPXML file"}, + "mode": {"type": "string", "enum": ["extend_previous", "extend_next", "delete", "auto"], "default": "auto", "description": "How to fix: extend previous/next clip, delete, or auto"}, + "threshold_frames": {"type": "integer", "default": 6, "description": "Frames below this threshold are flash frames"}, + "output_path": {"type": "string", "description": "Output path (default: adds _modified suffix)"} + }, + "required": ["filepath"] + } + ), + Tool( + name="rapid_trim", + description="Batch trim clips to a maximum duration for fast-paced montages", + inputSchema={ + "type": "object", + "properties": { + "filepath": {"type": "string", "description": "Path to FCPXML file"}, + "max_duration": {"type": "string", "description": "Maximum clip duration (e.g., '2s', '00:00:02:00')"}, + "min_duration": {"type": "string", "description": "Minimum clip duration (optional)"}, + "keywords": {"type": "array", "items": {"type": "string"}, "description": "Only trim clips with these keywords"}, + "trim_from": {"type": "string", "enum": ["start", "end", "center"], "default": "end", "description": "Where to trim from"}, + "output_path": {"type": "string", "description": "Output path (default: adds _modified suffix)"} + }, + "required": ["filepath", "max_duration"] + } + ), + Tool( + name="fill_gaps", + description="Automatically fill gaps in the timeline by extending adjacent clips", + inputSchema={ + "type": "object", + "properties": { + "filepath": {"type": "string", "description": "Path to FCPXML file"}, + "mode": {"type": "string", "enum": ["extend_previous", "extend_next", "delete"], "default": "extend_previous", "description": "How to fill gaps"}, + "max_gap": {"type": "string", "description": "Only fill gaps smaller than this (e.g., '1s')"}, + "output_path": {"type": "string", "description": "Output path (default: adds _modified suffix)"} + }, + "required": ["filepath"] + } + ), + Tool( + name="add_connected_clip", + description="Connect a library clip to an existing timeline clip (B-roll overlay, audio, title)", + inputSchema={ + "type": "object", + "properties": { + "filepath": {"type": "string", "description": "Path to FCPXML file"}, + "parent_clip_id": {"type": "string", "description": "Name/ID of the clip to attach to"}, + "asset_id": {"type": "string", "description": "Asset reference ID"}, + "asset_name": {"type": "string", "description": "Asset name (alternative to asset_id)"}, + "offset": {"type": "string", "default": "0s", "description": "Position relative to parent clip start"}, + "duration": {"type": "string", "description": "Duration (default: full asset)"}, + "lane": {"type": "integer", "default": 1, "description": "Lane number (positive=above, negative=below)"}, + "output_path": {"type": "string", "description": "Output path (default: adds _modified suffix)"} + }, + "required": ["filepath", "parent_clip_id"] + } + ), + Tool( + name="reformat_timeline", + description="Create new FCPXML with different resolution/aspect ratio (9:16 for TikTok, 1:1 for Instagram, etc.)", + inputSchema={ + "type": "object", + "properties": { + "filepath": {"type": "string", "description": "Path to FCPXML file"}, + "format": {"type": "string", "enum": ["9:16", "1:1", "4:5", "16:9", "4:3", "custom"], "description": "Target format preset"}, + "width": {"type": "integer", "description": "Custom width (only with format='custom')"}, + "height": {"type": "integer", "description": "Custom height (only with format='custom')"}, + "output_path": {"type": "string", "description": "Output path (default: adds _reformatted suffix)"} + }, + "required": ["filepath", "format"] + } + ), + Tool( + name="add_audio", + description="Add an audio clip or music bed to the timeline", + inputSchema={ + "type": "object", + "properties": { + "filepath": {"type": "string", "description": "Path to FCPXML file"}, + "parent_clip_id": {"type": "string", "description": "Clip to attach audio to (omit for music bed spanning full timeline)"}, + "asset_id": {"type": "string", "description": "Existing asset reference ID"}, + "src": {"type": "string", "description": "Path to audio file (creates new asset)"}, + "offset": {"type": "string", "description": "Position relative to parent clip start", "default": "0s"}, + "duration": {"type": "string", "description": "Duration of audio clip"}, + "role": {"type": "string", "description": "Audio role (dialogue, music, effects, etc.)", "default": "dialogue"}, + "lane": {"type": "integer", "description": "Lane number (negative = below)", "default": -1}, + "output_path": {"type": "string", "description": "Output path"}, + }, + "required": ["filepath"] + } + ), + Tool( + name="create_compound_clip", + description="Group spine clips into a compound clip", + inputSchema={ + "type": "object", + "properties": { + "filepath": {"type": "string", "description": "Path to FCPXML file"}, + "clip_ids": {"type": "array", "items": {"type": "string"}, "description": "Clip IDs to group"}, + "name": {"type": "string", "description": "Name for the compound clip", "default": "Compound Clip"}, + "output_path": {"type": "string", "description": "Output path"}, + }, + "required": ["filepath", "clip_ids"] + } + ), + Tool( + name="flatten_compound_clip", + description="Flatten a compound clip back into individual clips in the spine", + inputSchema={ + "type": "object", + "properties": { + "filepath": {"type": "string", "description": "Path to FCPXML file"}, + "ref_clip_id": {"type": "string", "description": "ID of the ref-clip to flatten"}, + "output_path": {"type": "string", "description": "Output path"}, + }, + "required": ["filepath", "ref_clip_id"] + } + ), +] + + +async def handle_add_marker(arguments: dict) -> Sequence[TextContent]: + filepath, output_path, modifier = _setup_modifier(arguments) + marker_type = MarkerType.from_string(arguments.get("marker_type", "standard")) + modifier.add_marker_at_timeline( + timecode=arguments["timecode"], name=arguments["name"], + marker_type=marker_type, note=arguments.get("note"), + ) + modifier.save(output_path) + return _text_result(f"Added marker '{arguments['name']}' at {arguments['timecode']}\n\nSaved to: {output_path}") + + +async def handle_batch_add_markers(arguments: dict) -> Sequence[TextContent]: + filepath, output_path, modifier = _setup_modifier(arguments) + markers_added = modifier.batch_add_markers( + markers=arguments.get("markers", []), + auto_at_cuts=arguments.get("auto_at_cuts", False), + auto_at_intervals=arguments.get("auto_at_intervals"), + ) + modifier.save(output_path) + return _text_result(f"Added {len(markers_added)} markers\n\nSaved to: {output_path}") + + +async def handle_trim_clip(arguments: dict) -> Sequence[TextContent]: + filepath, output_path, modifier = _setup_modifier(arguments) + modifier.trim_clip( + clip_id=arguments["clip_id"], + trim_start=arguments.get("trim_start"), + trim_end=arguments.get("trim_end"), + ripple=arguments.get("ripple", True), + ) + modifier.save(output_path) + return _text_result(f"Trimmed clip '{arguments['clip_id']}'\n\nSaved to: {output_path}") + + +async def handle_reorder_clips(arguments: dict) -> Sequence[TextContent]: + filepath, output_path, modifier = _setup_modifier(arguments) + modifier.reorder_clips( + clip_ids=arguments["clip_ids"], + target_position=arguments["target_position"], + ripple=arguments.get("ripple", True), + ) + modifier.save(output_path) + clips_moved = ", ".join(arguments["clip_ids"]) + return _text_result(f"Moved clips [{clips_moved}] to {arguments['target_position']}\n\nSaved to: {output_path}") + + +async def handle_add_transition(arguments: dict) -> Sequence[TextContent]: + filepath, output_path, modifier = _setup_modifier(arguments) + modifier.add_transition( + clip_id=arguments["clip_id"], + position=arguments.get("position", "end"), + transition_type=arguments.get("transition_type", "cross-dissolve"), + duration=arguments.get("duration", "00:00:00:15"), + ) + modifier.save(output_path) + return _text_result(f"Added {arguments.get('transition_type', 'cross-dissolve')} to '{arguments['clip_id']}'\n\nSaved to: {output_path}") + + +async def handle_change_speed(arguments: dict) -> Sequence[TextContent]: + speed = arguments["speed"] + if not isinstance(speed, (int, float)) or speed <= 0 or speed > 100: + raise ValueError( + f"Speed must be a positive number between 0 (exclusive) and 100, got {speed!r}" + ) + filepath, output_path, modifier = _setup_modifier(arguments) + modifier.change_speed( + clip_id=arguments["clip_id"], + speed=speed, + preserve_pitch=arguments.get("preserve_pitch", True), + ) + modifier.save(output_path) + speed_desc = f"{speed}x" if speed >= 1 else f"{int(1/speed)}x slow motion" + return _text_result(f"Changed speed of '{arguments['clip_id']}' to {speed_desc}\n\nSaved to: {output_path}") + + +async def handle_add_zoom(arguments: dict) -> Sequence[TextContent]: + start = float(arguments["start"]) + end = float(arguments["end"]) + scale = float(arguments.get("scale", 1.3)) + ease = float(arguments.get("ease", 0.3)) + position = arguments.get("position", "0 0") + + filepath, output_path, modifier = _setup_modifier(arguments) + modifier.add_zoom( + clip_id=arguments["clip_id"], start=start, end=end, + scale=scale, ease=ease, position=position, + ) + modifier.save(output_path) + return _text_result( + f"# Zoom Added\n\n" + f"- **Clip**: {arguments['clip_id']}\n" + f"- **Window**: {start}s → {end}s (clip-relative)\n" + f"- **Scale**: {int(scale * 100)}%\n" + f"- **Ease**: {ease}s in/out\n\n" + f"Saved to: {output_path}" + ) + + +async def handle_delete_clips(arguments: dict) -> Sequence[TextContent]: + filepath, output_path, modifier = _setup_modifier(arguments) + modifier.delete_clip( + clip_ids=arguments["clip_ids"], + ripple=arguments.get("ripple", True), + ) + modifier.save(output_path) + return _text_result(f"Deleted {len(arguments['clip_ids'])} clip(s)\n\nSaved to: {output_path}") + + +async def handle_split_clip(arguments: dict) -> Sequence[TextContent]: + filepath, output_path, modifier = _setup_modifier(arguments) + new_clips = modifier.split_clip( + clip_id=arguments["clip_id"], + split_points=arguments["split_points"], + ) + modifier.save(output_path) + return _text_result(f"Split '{arguments['clip_id']}' into {len(new_clips)} clips\n\nSaved to: {output_path}") + + +async def handle_insert_clip(arguments: dict) -> Sequence[TextContent]: + filepath, output_path, modifier = _setup_modifier(arguments) + new_clip = modifier.insert_clip( + asset_id=arguments.get("asset_id"), + asset_name=arguments.get("asset_name"), + position=arguments["position"], + duration=arguments.get("duration"), + in_point=arguments.get("in_point"), + out_point=arguments.get("out_point"), + ripple=arguments.get("ripple", True), + ) + modifier.save(output_path) + clip_name = new_clip.get('name', 'Unknown') + pos = arguments["position"] + return _text_result(f"Inserted '{clip_name}' at position '{pos}'\n\nSaved to: {output_path}") + + +async def handle_fix_flash_frames(arguments: dict) -> Sequence[TextContent]: + filepath, output_path, modifier = _setup_modifier(arguments, "_flash_fixed") + fixed = modifier.fix_flash_frames( + mode=arguments.get("mode", "auto"), + threshold_frames=arguments.get("threshold_frames", 6), + ) + modifier.save(output_path) + + if not fixed: + return _text_result("No flash frames found to fix.") + + result = _format_batch_result( + title="Flash Frames Fixed", + summary={"Fixed": f"{len(fixed)} flash frames", "Mode": arguments.get('mode', 'auto')}, + headers=["Clip", "Frames", "Action", "Result"], + rows=[ + [f['clip_name'], f"{f['duration_frames']}f", f['action'], f"Extended: {f.get('extended_clip', 'N/A')}"] + for f in fixed + ], + output_path=output_path, + ) + return _text_result(result) + + +async def handle_rapid_trim(arguments: dict) -> Sequence[TextContent]: + filepath, output_path, modifier = _setup_modifier(arguments, "_rapid_trim") + trimmed = modifier.rapid_trim( + max_duration=arguments["max_duration"], + min_duration=arguments.get("min_duration"), + keywords=arguments.get("keywords"), + trim_from=arguments.get("trim_from", "end"), + ) + modifier.save(output_path) + + if not trimmed: + return _text_result(f"No clips exceeded {arguments['max_duration']} - nothing trimmed.") + + total_before = sum(t['original_duration'] for t in trimmed) + total_after = sum(t['new_duration'] for t in trimmed) + + result = _format_batch_result( + title="Rapid Trim Complete", + summary={ + "Clips Trimmed": str(len(trimmed)), + "Max Duration": str(arguments['max_duration']), + "Trim From": arguments.get('trim_from', 'end'), + "Time Saved": format_duration(total_before - total_after), + }, + headers=["Clip", "Before", "After"], + rows=[ + [t['clip_name'], format_duration(t['original_duration']), format_duration(t['new_duration'])] + for t in trimmed + ], + output_path=output_path, + ) + return _text_result(result) + + +async def handle_fill_gaps(arguments: dict) -> Sequence[TextContent]: + filepath, output_path, modifier = _setup_modifier(arguments, "_gaps_filled") + filled = modifier.fill_gaps( + mode=arguments.get("mode", "extend_previous"), + max_gap=arguments.get("max_gap"), + ) + modifier.save(output_path) + + if not filled: + return _text_result("No gaps found to fill.") + + result = _format_batch_result( + title="Gaps Filled", + summary={"Gaps Filled": str(len(filled)), "Mode": arguments.get('mode', 'extend_previous')}, + headers=["Position", "Duration", "Action"], + rows=[[g['timecode'], f"{g['duration_frames']}f", g['action']] for g in filled], + output_path=output_path, + ) + return _text_result(result) + + +async def handle_add_connected_clip(arguments: dict) -> Sequence[TextContent]: + filepath, output_path, modifier = _setup_modifier(arguments) + modifier.add_connected_clip( + parent_clip_id=arguments["parent_clip_id"], + asset_id=arguments.get("asset_id"), + asset_name=arguments.get("asset_name"), + offset=arguments.get("offset", "0s"), + duration=arguments.get("duration"), + lane=arguments.get("lane", 1), + ) + modifier.save(output_path) + return _text_result(( + f"Connected clip added to '{arguments['parent_clip_id']}' on lane {arguments.get('lane', 1)}\n\n" + f"Saved to: `{output_path}`" + )) + + +async def handle_reformat_timeline(arguments: dict) -> Sequence[TextContent]: + filepath, output_path = _resolve_io_paths(arguments, "_reformatted") + + fmt = arguments["format"] + if fmt == "custom": + width = arguments.get("width") + height = arguments.get("height") + if not width or not height: + return _text_result("Custom format requires both 'width' and 'height' parameters.") + else: + formats = FCPXMLModifier.SOCIAL_FORMATS + if fmt not in formats: + return _text_result(f"Unknown format: {fmt}. Valid: {', '.join(formats.keys())}") + width, height = formats[fmt] + + modifier = FCPXMLModifier(filepath) + modifier.reformat_resolution(width, height) + modifier.save(output_path) + + return _text_result(( + f"# Timeline Reformatted\n\n" + f"- **Format**: {fmt} ({width}x{height})\n" + f"- **Aspect ratio**: {width}:{height}\n\n" + f"Saved to: `{output_path}`\n\n" + f"**Next step**: Import into FCP (File > Import > XML). " + f"FCP will handle spatial conforming automatically." + )) + + +async def handle_add_audio(arguments: dict) -> Sequence[TextContent]: + filepath, output_path, modifier = _setup_modifier(arguments, "_audio") + + parent_clip_id = arguments.get("parent_clip_id") + if parent_clip_id: + modifier.add_audio_clip( + parent_clip_id=parent_clip_id, + asset_id=arguments.get("asset_id"), + offset=arguments.get("offset", "0s"), + duration=arguments.get("duration"), + role=arguments.get("role", "dialogue"), + lane=arguments.get("lane", -1), + src=arguments.get("src"), + ) + action = f"Added audio clip to '{parent_clip_id}'" + else: + modifier.add_music_bed( + asset_id=arguments.get("asset_id"), + duration=arguments.get("duration"), + role=arguments.get("role", "music"), + src=arguments.get("src"), + ) + action = "Added music bed spanning full timeline" + + modifier.save(output_path) + return _text_result(f"{action}\nSaved to: `{output_path}`") + + +async def handle_create_compound_clip(arguments: dict) -> Sequence[TextContent]: + filepath, output_path, modifier = _setup_modifier(arguments, "_compound") + clip_ids = arguments["clip_ids"] + name = arguments.get("name", "Compound Clip") + modifier.create_compound_clip(clip_ids, name) + modifier.save(output_path) + return _text_result(( + f"Created compound clip '{name}' from {len(clip_ids)} clips.\n" + f"Saved to: `{output_path}`" + )) + + +async def handle_flatten_compound_clip(arguments: dict) -> Sequence[TextContent]: + filepath, output_path, modifier = _setup_modifier(arguments, "_flattened") + ref_clip_id = arguments["ref_clip_id"] + extracted = modifier.flatten_compound_clip(ref_clip_id) + modifier.save(output_path) + return _text_result(( + f"Flattened compound clip '{ref_clip_id}' into {len(extracted)} clips.\n" + f"Saved to: `{output_path}`" + )) + + +HANDLERS = { + "add_marker": handle_add_marker, + "batch_add_markers": handle_batch_add_markers, + "trim_clip": handle_trim_clip, + "reorder_clips": handle_reorder_clips, + "add_transition": handle_add_transition, + "change_speed": handle_change_speed, + "add_zoom": handle_add_zoom, + "delete_clips": handle_delete_clips, + "split_clip": handle_split_clip, + "insert_clip": handle_insert_clip, + "fix_flash_frames": handle_fix_flash_frames, + "rapid_trim": handle_rapid_trim, + "fill_gaps": handle_fill_gaps, + "add_connected_clip": handle_add_connected_clip, + "reformat_timeline": handle_reformat_timeline, + "add_audio": handle_add_audio, + "create_compound_clip": handle_create_compound_clip, + "flatten_compound_clip": handle_flatten_compound_clip, +} diff --git a/code/server_tools/export.py b/code/server_tools/export.py new file mode 100644 index 0000000..e8bbf66 --- /dev/null +++ b/code/server_tools/export.py @@ -0,0 +1,185 @@ +"""Export / relink — tool schemas and handlers. + +Extracted from server.py; see Engine/docs/03_SERVER_TOOLS.md for the tool catalog. +""" + +from __future__ import annotations + +from typing import Sequence + +from mcp.types import TextContent, Tool + +from fcpxml.export import DaVinciExporter +from fcpxml.writer import FCPXMLModifier +from server_tools._shared import ( + _require_timeline, + _resolve_io_paths, + _setup_modifier, + _text_result, + _validate_filepath, + format_timecode, +) + +TOOLS = [ + Tool( + name="export_edl", + description="Generate EDL (Edit Decision List) from timeline", + inputSchema={ + "type": "object", + "properties": {"filepath": {"type": "string"}}, + "required": ["filepath"] + } + ), + Tool( + name="export_csv", + description="Export timeline data to CSV format", + inputSchema={ + "type": "object", + "properties": { + "filepath": {"type": "string"}, + "include": {"type": "array", "items": {"type": "string"}} + }, + "required": ["filepath"] + } + ), + Tool( + name="export_resolve_xml", + description="Export timeline as DaVinci Resolve compatible FCPXML (simplified v1.9)", + inputSchema={ + "type": "object", + "properties": { + "filepath": {"type": "string", "description": "Path to FCPXML file"}, + "flatten_compounds": {"type": "boolean", "default": True, "description": "Flatten compound clips for compatibility"}, + "output_path": {"type": "string", "description": "Output path (default: adds _resolve suffix)"}, + }, + "required": ["filepath"] + } + ), + Tool( + name="export_fcp7_xml", + description="Export timeline as FCP7 XML (XMEML) for Premiere Pro, DaVinci Resolve, and Avid compatibility", + inputSchema={ + "type": "object", + "properties": { + "filepath": {"type": "string", "description": "Path to FCPXML file"}, + "output_path": {"type": "string", "description": "Output path (default: adds _fcp7.xml suffix)"}, + }, + "required": ["filepath"] + } + ), + Tool( + name="relink_media", + description="Bulk-rewrite media source paths (asset/media-rep src URLs) to relink moved or renamed media folders without opening FCP. Prefix-based: find='/Volumes/OldDrive/Media' replace='/Volumes/NewDrive/Media'. Use dry_run to preview.", + inputSchema={ + "type": "object", + "properties": { + "filepath": {"type": "string", "description": "Path to FCPXML file or .fcpxmld bundle"}, + "find": {"type": "string", "description": "Old path prefix to match (plain path or file:// URL)"}, + "replace": {"type": "string", "description": "New path prefix to substitute"}, + "dry_run": {"type": "boolean", "description": "Preview changes without writing", "default": False}, + "output_path": {"type": "string", "description": "Output path (default: adds _relinked suffix)"}, + }, + "required": ["filepath", "find", "replace"] + } + ), +] + + +async def handle_export_edl(arguments: dict) -> Sequence[TextContent]: + project, tl = _require_timeline(arguments["filepath"]) + edl = f"TITLE: {tl.name}\nFCM: NON-DROP FRAME\n\n" + for i, c in enumerate(tl.clips, 1): + edl += f"{i:03d} AX V C {format_timecode(c.source_start)} {format_timecode(c.end)} {format_timecode(c.start)} {format_timecode(c.end)}\n" + edl += f"* FROM CLIP NAME: {c.name}\n\n" + return _text_result(f"```edl\n{edl}```") + + +async def handle_export_csv(arguments: dict) -> Sequence[TextContent]: + project, tl = _require_timeline(arguments["filepath"]) + csv = "Name,Start,End,Duration,Keywords\n" + for c in tl.clips: + kws = "|".join(k.value for k in c.keywords) + csv += f'"{c.name}",{format_timecode(c.start)},{format_timecode(c.end)},{c.duration_seconds:.3f},"{kws}"\n' + return _text_result(f"```csv\n{csv}```") + + +async def handle_export_resolve_xml(arguments: dict) -> Sequence[TextContent]: + filepath, output_path = _resolve_io_paths(arguments, "_resolve") + exporter = DaVinciExporter(filepath) + exporter.export_simplified_fcpxml( + output_path, + flatten_compounds=arguments.get("flatten_compounds", True), + ) + return _text_result(( + f"# Exported for DaVinci Resolve\n\n" + f"- **Format**: Simplified FCPXML v1.9\n" + f"- **Compound clips flattened**: {arguments.get('flatten_compounds', True)}\n\n" + f"Saved to: `{output_path}`\n\n" + f"**Next step**: In DaVinci Resolve, go to File > Import > Timeline > Import AAF/EDL/XML" + )) + + +async def handle_export_fcp7_xml(arguments: dict) -> Sequence[TextContent]: + filepath, output_path = _resolve_io_paths(arguments, "_fcp7") + exporter = DaVinciExporter(filepath) + exporter.export_xmeml(output_path) + return _text_result(( + f"# Exported as FCP7 XML (XMEML)\n\n" + f"- **Format**: XMEML v5\n" + f"- **Compatible with**: Premiere Pro, DaVinci Resolve, Avid Media Composer\n\n" + f"Saved to: `{output_path}`\n\n" + f"**Next step**: Import via File > Import in your target NLE" + )) + + +async def handle_relink_media(arguments: dict) -> Sequence[TextContent]: + dry_run = arguments.get("dry_run", False) + if dry_run: + filepath = _validate_filepath(arguments["filepath"], ('.fcpxml', '.fcpxmld')) + modifier = FCPXMLModifier(filepath) + result = modifier.relink_media( + arguments["find"], arguments["replace"], dry_run=True + ) + footer = "Dry run — no file written." + else: + filepath, output_path, modifier = _setup_modifier(arguments, "_relinked") + result = modifier.relink_media(arguments["find"], arguments["replace"]) + saved = modifier.save(output_path) + footer = f"Saved to: {saved}" + + if not result["relinked"]: + return _text_result( + f"No media paths matched prefix '{arguments['find']}' " + f"({result['total_assets']} assets scanned). Nothing to relink." + ) + + lines = [ + f"{'Would relink' if dry_run else 'Relinked'} " + f"{result['relinked']} media reference(s) " + f"across {result['total_assets']} asset(s):", + "", + ] + missing = 0 + for change in result["changes"]: + mark = "✓" if change["target_exists"] else "⚠ target missing" + if not change["target_exists"]: + missing += 1 + lines.append(f" {change['asset']}: {change['new']} [{mark}]") + if missing: + lines.append("") + lines.append( + f"⚠ {missing} new path(s) do not exist on this machine — " + f"FCP will show those clips as missing until the media is present." + ) + lines.append("") + lines.append(footer) + return _text_result("\n".join(lines)) + + +HANDLERS = { + "export_edl": handle_export_edl, + "export_csv": handle_export_csv, + "export_resolve_xml": handle_export_resolve_xml, + "export_fcp7_xml": handle_export_fcp7_xml, + "relink_media": handle_relink_media, +} diff --git a/code/server_tools/generation.py b/code/server_tools/generation.py new file mode 100644 index 0000000..f70d060 --- /dev/null +++ b/code/server_tools/generation.py @@ -0,0 +1,271 @@ +"""Geração — tool schemas and handlers. + +Extracted from server.py; see Engine/docs/03_SERVER_TOOLS.md for the tool catalog. +""" + +from __future__ import annotations + +from typing import Sequence + +from mcp.types import TextContent, Tool + +from fcpxml.models import SegmentSpec +from fcpxml.templates import ClipSpec, apply_template, list_templates +from server_tools._shared import ( + PROJECTS_DIR, + _setup_generator, + _text_result, + _validate_output_path, + format_duration, +) + +TOOLS = [ + Tool( + name="auto_rough_cut", + description="Generate a rough cut from source clips based on keywords, duration, and pacing", + inputSchema={ + "type": "object", + "properties": { + "filepath": {"type": "string", "description": "Source FCPXML with clips"}, + "output_path": {"type": "string", "description": "Where to save rough cut"}, + "target_duration": {"type": "string", "description": "Target length (3m, 00:03:00:00)"}, + "pacing": {"type": "string", "enum": ["slow", "medium", "fast", "dynamic"], "default": "medium"}, + "keywords": {"type": "array", "items": {"type": "string"}, "description": "Filter clips by keywords"}, + "segments": { + "type": "array", + "items": { + "type": "object", + "properties": { + "name": {"type": "string"}, + "keywords": {"type": "array", "items": {"type": "string"}}, + "duration": {"type": "number"} + } + }, + "description": "Segment structure [{name, keywords, duration_seconds}]" + }, + "priority": {"type": "string", "enum": ["best", "favorites", "longest", "shortest", "random"], "default": "best"}, + "favorites_only": {"type": "boolean", "default": False}, + "add_transitions": {"type": "boolean", "default": False} + }, + "required": ["filepath", "output_path", "target_duration"] + } + ), + Tool( + name="generate_montage", + description="Create rapid-fire montages with pacing curves (accelerating, decelerating, pyramid)", + inputSchema={ + "type": "object", + "properties": { + "filepath": {"type": "string", "description": "Source FCPXML with clips"}, + "output_path": {"type": "string", "description": "Where to save montage"}, + "target_duration": {"type": "string", "description": "Total montage length (e.g., '30s', '00:00:30:00')"}, + "pacing_curve": {"type": "string", "enum": ["accelerating", "decelerating", "pyramid", "constant"], "default": "accelerating", "description": "How clip duration changes over time"}, + "start_duration": {"type": "number", "default": 2.0, "description": "Clip duration at start (seconds)"}, + "end_duration": {"type": "number", "default": 0.5, "description": "Clip duration at end (seconds)"}, + "keywords": {"type": "array", "items": {"type": "string"}, "description": "Filter clips by keywords"}, + "add_transitions": {"type": "boolean", "default": False, "description": "Add quick dissolves"} + }, + "required": ["filepath", "output_path", "target_duration"] + } + ), + Tool( + name="generate_ab_roll", + description="Create documentary-style A/B roll edits alternating between main content and cutaways", + inputSchema={ + "type": "object", + "properties": { + "filepath": {"type": "string", "description": "Source FCPXML with clips"}, + "output_path": {"type": "string", "description": "Where to save A/B roll edit"}, + "target_duration": {"type": "string", "description": "Total duration (e.g., '3m', '00:03:00:00')"}, + "a_keywords": {"type": "array", "items": {"type": "string"}, "description": "Keywords for A-roll (main content, interviews)"}, + "b_keywords": {"type": "array", "items": {"type": "string"}, "description": "Keywords for B-roll (cutaways, visuals)"}, + "a_duration": {"type": "string", "default": "5s", "description": "Duration of each A-roll segment"}, + "b_duration": {"type": "string", "default": "3s", "description": "Duration of each B-roll cutaway"}, + "start_with": {"type": "string", "enum": ["a", "b"], "default": "a", "description": "Which roll to start with"}, + "add_transitions": {"type": "boolean", "default": True, "description": "Add cross-dissolves"} + }, + "required": ["filepath", "output_path", "target_duration", "a_keywords", "b_keywords"] + } + ), + Tool( + name="list_templates", + description="List available timeline templates with slot definitions", + inputSchema={ + "type": "object", + "properties": {}, + } + ), + Tool( + name="apply_template", + description="Fill a timeline template with clips and generate FCPXML", + inputSchema={ + "type": "object", + "properties": { + "template_name": {"type": "string", "description": "Template name (intro_outro, lower_thirds, music_video)"}, + "clips": {"type": "object", "description": "Map of slot_name -> {src, name, duration} or {asset_id, name, duration}"}, + "output_path": {"type": "string", "description": "Output FCPXML path"}, + "fps": {"type": "number", "description": "Frame rate", "default": 24}, + }, + "required": ["template_name", "clips", "output_path"] + } + ), +] + + +async def handle_auto_rough_cut(arguments: dict) -> Sequence[TextContent]: + filepath, output_path, generator = _setup_generator(arguments, "_roughcut") + + segments = None + if arguments.get("segments"): + segments = [ + SegmentSpec( + name=s.get("name", "Segment"), + keywords=s.get("keywords", []), + duration_seconds=s.get("duration", 0), + priority=s.get("priority", "best"), + ) + for s in arguments["segments"] + ] + result = generator.generate( + output_path=output_path, + target_duration=arguments["target_duration"], + pacing=arguments.get("pacing", "medium"), + keywords=arguments.get("keywords"), + segments=segments, + priority=arguments.get("priority", "best"), + favorites_only=arguments.get("favorites_only", False), + add_transitions=arguments.get("add_transitions", False), + ) + + return _text_result(f"""# Rough Cut Generated + +## Summary +- **Clips Used**: {result.clips_used} of {result.clips_available} available +- **Target Duration**: {format_duration(result.target_duration)} +- **Actual Duration**: {format_duration(result.actual_duration)} +- **Average Clip**: {format_duration(result.average_clip_duration)} + +## Output +Saved to: `{result.output_path}` + +**Next step**: Import this FCPXML into Final Cut Pro (File > Import > XML) +""") + + +async def handle_generate_montage(arguments: dict) -> Sequence[TextContent]: + filepath, output_path, generator = _setup_generator(arguments, "_montage") + result = generator.generate_montage( + output_path=output_path, + target_duration=arguments["target_duration"], + pacing_curve=arguments.get("pacing_curve", "accelerating"), + start_duration=arguments.get("start_duration", 2.0), + end_duration=arguments.get("end_duration", 0.5), + keywords=arguments.get("keywords"), + add_transitions=arguments.get("add_transitions", False), + ) + + curve_desc = { + 'accelerating': 'slow to fast (builds energy)', + 'decelerating': 'fast to slow (winds down)', + 'pyramid': 'slow to fast to slow (dramatic arc)', + 'constant': 'same duration throughout', + } + + return _text_result(f"""# Montage Generated + +## Summary +- **Clips Used**: {result['clips_used']} of {result['clips_available']} available +- **Target Duration**: {format_duration(result['target_duration'])} +- **Actual Duration**: {format_duration(result['actual_duration'])} +- **Pacing Curve**: {result['pacing_curve']} - {curve_desc.get(result['pacing_curve'], '')} + +## Pacing +- **Start Clip Duration**: {format_duration(result['start_clip_duration'])} +- **End Clip Duration**: {format_duration(result['end_clip_duration'])} + +## Output +Saved to: `{result['output_path']}` +""") + + +async def handle_generate_ab_roll(arguments: dict) -> Sequence[TextContent]: + filepath, output_path, generator = _setup_generator(arguments, "_ab_roll") + result = generator.generate_ab_roll( + output_path=output_path, + target_duration=arguments["target_duration"], + a_keywords=arguments["a_keywords"], + b_keywords=arguments["b_keywords"], + a_duration=arguments.get("a_duration", "5s"), + b_duration=arguments.get("b_duration", "3s"), + start_with=arguments.get("start_with", "a"), + add_transitions=arguments.get("add_transitions", True), + ) + + return _text_result(f"""# A/B Roll Edit Generated + +## Summary +- **A-Roll Segments**: {result['a_segments']} (from {result['a_clips_available']} available) +- **B-Roll Segments**: {result['b_segments']} (from {result['b_clips_available']} available) +- **Total Clips**: {result['clips_used']} + +## Timing +- **Target Duration**: {format_duration(result['target_duration'])} +- **Actual Duration**: {format_duration(result['actual_duration'])} +- **A-Roll Duration**: {result['a_duration_setting']} per segment +- **B-Roll Duration**: {result['b_duration_setting']} per cutaway + +## Output +Saved to: `{result['output_path']}` + +**Next step**: Import this FCPXML into Final Cut Pro (File > Import > XML) +""") + + +async def handle_list_templates(arguments: dict) -> Sequence[TextContent]: + templates = list_templates() + lines = ["# Available Timeline Templates\n"] + for tmpl in templates: + lines.append(f"## {tmpl['name']}") + lines.append(f"{tmpl['description']}\n") + lines.append("| Slot | Type | Default Duration | Lane | Required |") + lines.append("|------|------|-----------------|------|----------|") + for s in tmpl['slots']: + lines.append( + f"| {s['name']} | {s['slot_type']} | {s['default_duration']}s " + f"| {s['lane']} | {'Yes' if s['required'] else 'No'} |" + ) + lines.append("") + return _text_result("\n".join(lines)) + + +async def handle_apply_template(arguments: dict) -> Sequence[TextContent]: + template_name = arguments["template_name"] + clips_raw = arguments["clips"] + output_path = _validate_output_path(arguments["output_path"], anchor_dir=PROJECTS_DIR) + fps = arguments.get("fps", 24) + + # Convert raw clips dict to ClipSpec objects + clips_map = {} + for slot_name, spec_data in clips_raw.items(): + if isinstance(spec_data, dict): + clips_map[slot_name] = ClipSpec( + asset_id=spec_data.get("asset_id"), + src=spec_data.get("src"), + name=spec_data.get("name", slot_name), + duration=spec_data.get("duration"), + ) + + result_path = apply_template(template_name, clips_map, output_path, fps) + return _text_result(( + f"Applied template '{template_name}' with {len(clips_map)} clips.\n" + f"Saved to: `{result_path}`" + )) + + +HANDLERS = { + "auto_rough_cut": handle_auto_rough_cut, + "generate_montage": handle_generate_montage, + "generate_ab_roll": handle_generate_ab_roll, + "list_templates": handle_list_templates, + "apply_template": handle_apply_template, +} diff --git a/code/server_tools/live.py b/code/server_tools/live.py new file mode 100644 index 0000000..16e5938 --- /dev/null +++ b/code/server_tools/live.py @@ -0,0 +1,110 @@ +"""Live (macOS) — tool schemas and handlers. + +Extracted from server.py; see Engine/docs/03_SERVER_TOOLS.md for the tool catalog. +""" + +from __future__ import annotations + +from pathlib import Path +from typing import Sequence + +from mcp.types import TextContent, Tool + +from server_tools._shared import ( + _text_result, + _validate_filepath, + _validate_output_path, + generate_output_path, +) + +TOOLS = [ + Tool( + name="push_to_fcp", + description="LIVE: send an FCPXML file into the running Final Cut Pro with zero clicks (official Open Document Apple event). Creates/targets a library via import-options. Launches FCP if needed. macOS-only; first use triggers an Automation permission prompt. For true zero-click, pass a library_location ending in .fcpbundle (a new path is auto-created); omitting it makes FCP show a modal library picker.", + inputSchema={ + "type": "object", + "properties": { + "filepath": {"type": "string", "description": "Path to FCPXML file or .fcpxmld bundle to import"}, + "library_location": {"type": "string", "description": "Target .fcpbundle library path (auto-created if it doesn't exist; the extension is normalized to .fcpbundle). Omit to import into the active library, but note FCP then shows a modal 'Open Library' picker that blocks until answered"}, + "suppress_warnings": {"type": "boolean", "description": "Suppress non-fatal import warning dialogs", "default": True}, + "copy_assets": {"type": "boolean", "description": "Copy media into the library (true) or link in place (false). Omit for FCP default"}, + }, + "required": ["filepath"] + } + ), + Tool( + name="list_fcp_libraries", + description="LIVE: enumerate the running Final Cut Pro's open libraries, events, and projects via Apple's read-only scripting dictionary. Refuses to launch FCP unless allow_launch is true. macOS-only.", + inputSchema={ + "type": "object", + "properties": { + "allow_launch": {"type": "boolean", "description": "Launch FCP if it isn't running", "default": False}, + }, + } + ), +] + + +async def handle_push_to_fcp(arguments: dict) -> Sequence[TextContent]: + from fcpxml.live import push_to_fcp + + filepath = _validate_filepath(arguments["filepath"], ('.fcpxml', '.fcpxmld')) + + # Flat files get an options-injected sibling copy (never touch the + # original); the copy path goes through the same write sandbox as + # every other derived output. + import_copy = None + if Path(filepath).suffix.lower() == '.fcpxml': + anchor = str(Path(filepath).resolve().parent) + import_copy = _validate_output_path( + generate_output_path(filepath, "_import"), anchor_dir=anchor + ) + + result = push_to_fcp( + filepath, + library_location=arguments.get("library_location"), + suppress_warnings=arguments.get("suppress_warnings", True), + copy_assets=arguments.get("copy_assets"), + import_copy_path=import_copy, + ) + lines = [ + f"Sent to Final Cut Pro: {result['sent']}", + f"FCP {'was launched' if result['launched_fcp'] else 'was already running'} — " + f"import happens in-app (libraries/events are created or merged per import-options).", + ] + if arguments.get("library_location"): + lines.append(f"Target library: {arguments['library_location']}") + lines.append( + "Note: Apple offers no programmatic export — to round-trip edits " + "back, use File > Export XML in FCP." + ) + return _text_result("\n".join(lines)) + + +async def handle_list_fcp_libraries(arguments: dict) -> Sequence[TextContent]: + from fcpxml.live import list_fcp_libraries + + try: + libraries = list_fcp_libraries( + allow_launch=arguments.get("allow_launch", False) + ) + except RuntimeError as exc: + return _text_result(str(exc)) + + if not libraries: + return _text_result("Final Cut Pro is running but reports no open libraries.") + + lines = [f"Open libraries in Final Cut Pro ({len(libraries)}):", ""] + for lib in libraries: + lines.append(f"📚 {lib['name']}") + for event in lib["events"]: + lines.append(f" └─ {event['name']}") + for proj in event["projects"]: + lines.append(f" • {proj}") + return _text_result("\n".join(lines)) + + +HANDLERS = { + "push_to_fcp": handle_push_to_fcp, + "list_fcp_libraries": handle_list_fcp_libraries, +} diff --git a/code/server_tools/markers_import.py b/code/server_tools/markers_import.py new file mode 100644 index 0000000..cbbe77d --- /dev/null +++ b/code/server_tools/markers_import.py @@ -0,0 +1,458 @@ +"""Beats / markers importados — tool schemas and handlers. + +Extracted from server.py; see Engine/docs/03_SERVER_TOOLS.md for the tool catalog. +""" + +from __future__ import annotations + +import json +import re +from pathlib import Path +from typing import Sequence + +from mcp.types import TextContent, Tool + +from fcpxml.media_intel import media_src_to_path +from fcpxml.parser import FCPXMLParser +from fcpxml.writer import FCPXMLModifier +from server_tools._shared import ( + _check_json_depth, + _load_or_transcribe, + _markdown_table, + _no_timeline, + _raw_markers_to_batch, + _resolve_io_paths, + _setup_modifier, + _text_result, + _validate_filepath, + format_duration, + parse_srt, + parse_transcript_timestamps, + parse_vtt, +) + +TOOLS = [ + Tool( + name="import_beat_markers", + description="Import beat markers from external audio analysis (JSON format)", + inputSchema={ + "type": "object", + "properties": { + "filepath": {"type": "string", "description": "Path to FCPXML file"}, + "beats_path": {"type": "string", "description": "Path to beats JSON file"}, + "marker_type": {"type": "string", "enum": ["standard", "chapter"], "default": "standard"}, + "beat_filter": {"type": "string", "enum": ["all", "downbeat", "measure"], "default": "all", "description": "Which beats to import"}, + "output_path": {"type": "string", "description": "Output path (default: adds _beats suffix)"} + }, + "required": ["filepath", "beats_path"] + } + ), + Tool( + name="snap_to_beats", + description="Align cuts to nearest beat markers for music-synced edits", + inputSchema={ + "type": "object", + "properties": { + "filepath": {"type": "string", "description": "Path to FCPXML file with beat markers"}, + "max_shift_frames": {"type": "integer", "default": 6, "description": "Maximum frames to shift a cut"}, + "prefer": {"type": "string", "enum": ["earlier", "later", "nearest"], "default": "nearest", "description": "Which beat to prefer when equidistant"}, + "output_path": {"type": "string", "description": "Output path (default: adds _synced suffix)"} + }, + "required": ["filepath"] + } + ), + Tool( + name="import_srt_markers", + description="Import SRT or VTT subtitles as chapter markers on the timeline", + inputSchema={ + "type": "object", + "properties": { + "filepath": {"type": "string", "description": "Path to FCPXML file"}, + "srt_path": {"type": "string", "description": "Path to SRT or VTT subtitle file"}, + "mode": {"type": "string", "enum": ["all", "first_per_minute", "scene_changes"], "default": "first_per_minute", "description": "How to create markers: every subtitle, first per minute, or on text changes"}, + "marker_type": {"type": "string", "enum": ["standard", "chapter"], "default": "chapter"}, + "max_label_length": {"type": "integer", "default": 50, "description": "Truncate marker labels to this length"}, + "output_path": {"type": "string", "description": "Output path (default: adds _subtitled suffix)"} + }, + "required": ["filepath", "srt_path"] + } + ), + Tool( + name="import_transcript_markers", + description="Import timestamped transcript (YouTube chapter format) as markers. Supports '0:00 Title' and 'HH:MM:SS Title' formats", + inputSchema={ + "type": "object", + "properties": { + "filepath": {"type": "string", "description": "Path to FCPXML file"}, + "transcript": {"type": "string", "description": "Timestamped text (one per line: '0:00 Introduction')"}, + "transcript_path": {"type": "string", "description": "Path to text file with timestamps (alternative to inline transcript)"}, + "marker_type": {"type": "string", "enum": ["standard", "chapter"], "default": "chapter"}, + "output_path": {"type": "string", "description": "Output path (default: adds _chapters suffix)"} + }, + "required": ["filepath"] + } + ), + Tool( + name="transcript_markers", + description="Add a marker at the start of every transcribed segment (sentence-level), using each media file's local Whisper transcript. Maps each segment's source-media timestamp to its correct timeline position per clip, so it stays accurate across multiple clips/trims — unlike import_transcript_markers (plain timestamp text) or import_srt_markers (a caption track already synced to the whole export). Uses each media file's _transcript.json (auto-transcribes if missing). Non-destructive: writes a _transcript_markers copy.", + inputSchema={ + "type": "object", + "properties": { + "filepath": {"type": "string", "description": "Path to FCPXML file"}, + "clip_name": {"type": "string", "description": "Only mark the clip with this name"}, + "marker_type": {"type": "string", "default": "chapter", "description": "Marker type: standard, chapter, todo, completed"}, + "max_label_length": {"type": "integer", "default": 50, "description": "Truncate marker labels to this many characters (0 = no truncation)"}, + "model": {"type": "string", "default": "base", "description": "Whisper model size if transcription is needed"}, + "output_path": {"type": "string", "description": "Output path (default: adds _transcript_markers suffix)"}, + }, + "required": ["filepath"] + } + ), +] + + +async def handle_import_beat_markers(arguments: dict) -> Sequence[TextContent]: + filepath, output_path = _resolve_io_paths(arguments, "_beats") + beats_path = _validate_filepath(arguments["beats_path"], ('.json',)) + + with open(beats_path, 'r') as f: + beats_data = json.load(f) + _check_json_depth(beats_data) + + beat_times = [] + if isinstance(beats_data, list): + beat_times = beats_data + elif isinstance(beats_data, dict): + beat_times = beats_data.get('beats', beats_data.get('times', beats_data.get('markers', []))) + + beat_filter = arguments.get("beat_filter", "all") + if beat_filter == "downbeat" and isinstance(beats_data, dict): + beat_times = beats_data.get('downbeats', beat_times[::4]) + elif beat_filter == "measure" and isinstance(beats_data, dict): + beat_times = beats_data.get('measures', beat_times[::4]) + + markers = [] + marker_type = arguments.get("marker_type", "standard") + for i, beat_time in enumerate(beat_times): + if isinstance(beat_time, (int, float)): + markers.append({ + 'timecode': f"{beat_time}s", + 'name': f"Beat {i+1}", + 'marker_type': marker_type.upper(), + }) + elif isinstance(beat_time, dict): + markers.append({ + 'timecode': f"{beat_time.get('time', beat_time.get('position', 0))}s", + 'name': beat_time.get('label', f"Beat {i+1}"), + 'marker_type': marker_type.upper(), + }) + + modifier = FCPXMLModifier(filepath) + + # Songs routinely run longer than the edit — beats past the timeline's + # end are skipped (add_marker_at_timeline would raise on them). + timeline_end = modifier._timeline_duration().to_seconds() + in_range = [m for m in markers if float(m['timecode'].rstrip('s')) < timeline_end] + skipped_count = len(markers) - len(in_range) + + added = modifier.batch_add_markers(markers=in_range) + modifier.save(output_path) + + skipped_note = ( + f"- **Skipped**: {skipped_count} beat(s) beyond the timeline end " + f"({format_duration(timeline_end)})\n" if skipped_count else "" + ) + return _text_result(f"""# Beat Markers Imported + +## Summary +- **Beats Found**: {len(beat_times)} +- **Markers Added**: {len(added)} +{skipped_note}- **Filter**: {beat_filter} +- **Marker Type**: {marker_type} + +## Output +Saved to: `{output_path}` + +*Use `snap_to_beats` to align your cuts to these markers.* +""") + + +async def handle_snap_to_beats(arguments: dict) -> Sequence[TextContent]: + filepath, output_path = _resolve_io_paths(arguments, "_synced") + max_shift = arguments.get("max_shift_frames", 6) + prefer = arguments.get("prefer", "nearest") + + parser = FCPXMLParser() + project = parser.parse_file(filepath) + if not project.timelines: + return _no_timeline() + + tl = project.primary_timeline + fps = tl.frame_rate + + markers = list(tl.markers) + for clip in tl.clips: + markers.extend(clip.markers) + + if not markers: + return _text_result("No markers found. Use `import_beat_markers` first.") + + marker_times = sorted([m.start.seconds for m in markers]) + + modifier = FCPXMLModifier(filepath) + spine = modifier._get_spine() + adjusted_count = 0 + total_shift = 0 + + clips_list = [c for c in spine if c.tag in ('clip', 'asset-clip', 'video', 'ref-clip')] + + for i, clip in enumerate(clips_list[1:], 1): + cut_offset = modifier._parse_time(clip.get('offset', '0s')) + cut_seconds = cut_offset.to_seconds() + + best_marker = None + best_distance = float('inf') + + for marker_time in marker_times: + distance = abs(marker_time - cut_seconds) + distance_frames = distance * fps + + if distance_frames <= max_shift: + if prefer == "earlier" and marker_time <= cut_seconds: + if distance < best_distance: + best_distance = distance + best_marker = marker_time + elif prefer == "later" and marker_time >= cut_seconds: + if distance < best_distance: + best_distance = distance + best_marker = marker_time + elif prefer == "nearest": + if distance < best_distance: + best_distance = distance + best_marker = marker_time + + if best_marker is not None and best_distance > 0.001: + shift = best_marker - cut_seconds + shift_frames = int(shift * fps) + + prev_clip = clips_list[i - 1] + prev_dur = modifier._parse_time(prev_clip.get('duration', '0s')) + new_prev_dur = prev_dur + modifier._parse_time(f"{shift}s") + prev_clip.set('duration', new_prev_dur.to_fcpxml()) + + new_offset = modifier._parse_time(f"{best_marker}s") + clip.set('offset', new_offset.to_fcpxml()) + + adjusted_count += 1 + total_shift += abs(shift_frames) + + modifier.save(output_path) + avg_shift = total_shift / adjusted_count if adjusted_count > 0 else 0 + + return _text_result(f"""# Cuts Snapped to Beats + +## Summary +- **Cuts Adjusted**: {adjusted_count} +- **Max Shift Allowed**: {max_shift} frames +- **Preference**: {prefer} +- **Average Shift**: {avg_shift:.1f} frames + +## Output +Saved to: `{output_path}` + +Your edits are now synced to the beat! +""") + + +async def handle_import_srt_markers(arguments: dict) -> Sequence[TextContent]: + filepath, output_path = _resolve_io_paths(arguments, "_subtitled") + srt_path = _validate_filepath(arguments["srt_path"], ('.srt', '.vtt')) + mode = arguments.get("mode", "first_per_minute") + marker_type = arguments.get("marker_type", "chapter") + max_label = arguments.get("max_label_length", 50) + + text = Path(srt_path).read_text(encoding='utf-8') + + # Detect format and parse + if srt_path.endswith('.vtt') or text.strip().startswith('WEBVTT'): + raw_markers = parse_vtt(text) + fmt_name = "WebVTT" + else: + raw_markers = parse_srt(text) + fmt_name = "SRT" + + if not raw_markers: + return _text_result(f"No subtitles found in {srt_path}") + + # Apply mode filtering + filtered = [] + if mode == "all": + filtered = raw_markers + elif mode == "first_per_minute": + seen_minutes = set() + for m in raw_markers: + minute = int(m['seconds'] // 60) + if minute not in seen_minutes: + seen_minutes.add(minute) + filtered.append(m) + elif mode == "scene_changes": + # Group by similar text, take first occurrence of each unique line + seen_texts = set() + for m in raw_markers: + # Normalize: lowercase, strip punctuation + normalized = re.sub(r'[^\w\s]', '', m['text'].lower()).strip() + words = normalized.split()[:3] # First 3 words as key + key = ' '.join(words) + if key and key not in seen_texts: + seen_texts.add(key) + filtered.append(m) + + markers = _raw_markers_to_batch(filtered, marker_type, max_label=max_label) + + modifier = FCPXMLModifier(filepath) + added = modifier.batch_add_markers(markers=markers) + modifier.save(output_path) + + return _text_result(f"""# Subtitle Markers Imported + +## Summary +- **Format**: {fmt_name} +- **Subtitles Parsed**: {len(raw_markers)} +- **Mode**: {mode} +- **Markers Added**: {len(added)} +- **Marker Type**: {marker_type} + +## Output +Saved to: `{output_path}` +""") + + +async def handle_import_transcript_markers(arguments: dict) -> Sequence[TextContent]: + filepath, output_path = _resolve_io_paths(arguments, "_chapters") + marker_type = arguments.get("marker_type", "chapter") + + # Get transcript text from inline or file + transcript = arguments.get("transcript") + transcript_path = arguments.get("transcript_path") + + if not transcript and not transcript_path: + return _text_result("Provide either 'transcript' (inline text) or 'transcript_path' (path to file)") + + if transcript_path: + # .txt only: the parser below understands "0:00 Title" lines, not real + # SRT/VTT cue syntax — that's import_srt_markers (parse_srt/parse_vtt). + transcript_path = _validate_filepath(transcript_path, ('.txt',)) + transcript = Path(transcript_path).read_text(encoding='utf-8') + + raw_markers = parse_transcript_timestamps(transcript or "") + + if not raw_markers: + return _text_result("No timestamps found. Expected format: '0:00 Title' or 'HH:MM:SS Title', one per line.") + + markers = _raw_markers_to_batch(raw_markers, marker_type) + + modifier = FCPXMLModifier(filepath) + added = modifier.batch_add_markers(markers=markers) + modifier.save(output_path) + + return _text_result(f"""# Transcript Markers Imported + +## Summary +- **Timestamps Found**: {len(raw_markers)} +- **Markers Added**: {len(added)} +- **Marker Type**: {marker_type} + +## Markers +""" + "\n".join(f"- `{m['timecode']}` {m['name']}" for m in markers) + f""" + +## Output +Saved to: `{output_path}` +""") + + +async def handle_transcript_markers(arguments: dict) -> Sequence[TextContent]: + """Add a marker at the start of each transcribed segment, using each + media's cached (or freshly transcribed) local Whisper transcript. + + Unlike ``import_transcript_markers`` (plain "0:00 Title" text) or + ``import_srt_markers`` (a caption track already synced to the whole + exported video), this maps each segment's SOURCE-media timestamp to its + TIMELINE position per spine clip — the same source->timeline mapping + ``detect_media_silence`` uses — so it stays correct across multiple + clips built from different (and differently-trimmed) source files. + """ + marker_type = arguments.get("marker_type", "chapter") + max_label = int(arguments.get("max_label_length", 50)) + model = arguments.get("model", "base") + language = arguments.get("language") + output_dir = arguments.get("output_dir") + clip_filter = arguments.get("clip_name") + + filepath, output_path, modifier = _setup_modifier(arguments, "_transcript_markers") + + added: list[tuple[str, float, str]] = [] + skipped: list[tuple[str, str]] = [] + spine_clips = [el for _, el in modifier._iter_spine_clips()] + for el in spine_clips: + name = el.get("name", "") + if clip_filter and name != clip_filter: + continue + src = modifier.resources.get(el.get("ref", ""), {}).get("src", "") + media_path = media_src_to_path(src) + if not media_path or not Path(media_path).is_file(): + skipped.append((name, "media file missing")) + continue + data, reason = _load_or_transcribe(media_path, model, language, output_dir) + if data is None: + skipped.append((name, reason)) + continue + + clip_source_start = modifier.source_file_start(el).to_seconds() + clip_duration = modifier._parse_time(el.get("duration", "0s")).to_seconds() + clip_offset = modifier._parse_time(el.get("offset", "0s")).to_seconds() + window_end = clip_source_start + clip_duration + + for seg in data.get("segments", []): + seg_start = float(seg.get("start", 0.0)) + if seg_start < clip_source_start or seg_start >= window_end: + continue + label = seg.get("text", "").strip() + if not label: + continue + if max_label and len(label) > max_label: + label = label[:max_label] + timeline_seconds = clip_offset + (seg_start - clip_source_start) + modifier.add_marker_at_timeline( + timecode=f"{timeline_seconds}s", name=label, marker_type=marker_type, + ) + added.append((name, seg_start, label)) + + if not added: + text = "# Transcript Markers\n\nNo segments to mark — file unchanged (nothing saved)." + if skipped: + text += "\n\n## Skipped Clips\n" + _markdown_table( + ["Clip", "Reason"], [[n, r] for n, r in skipped] + ) + return _text_result(text) + + modifier.save(output_path) + result = "# Transcript Markers Imported (local Whisper)\n\n## Summary\n" + result += f"- **Markers Added**: {len(added)}\n- **Marker Type**: {marker_type}\n\n" + result += _markdown_table( + ["Clip", "Start", "Label"], [[n, f"{s:.2f}s", label] for n, s, label in added] + ) + if skipped: + result += "\n## Skipped Clips\n" + _markdown_table( + ["Clip", "Reason"], [[n, r] for n, r in skipped] + ) + result += f"\n\nSaved to: `{output_path}`\n\n*Transcripts are cached as _transcript.json.*" + return _text_result(result) + + +HANDLERS = { + "import_beat_markers": handle_import_beat_markers, + "snap_to_beats": handle_snap_to_beats, + "import_srt_markers": handle_import_srt_markers, + "import_transcript_markers": handle_import_transcript_markers, + "transcript_markers": handle_transcript_markers, +} diff --git a/code/server_tools/qc.py b/code/server_tools/qc.py new file mode 100644 index 0000000..47b4e69 --- /dev/null +++ b/code/server_tools/qc.py @@ -0,0 +1,696 @@ +"""QC e detecção — tool schemas and handlers. + +Extracted from server.py; see Engine/docs/03_SERVER_TOOLS.md for the tool catalog. +""" + +from __future__ import annotations + +import json +from pathlib import Path +from typing import Sequence + +from mcp.types import TextContent, Tool + +from fcpxml.media_intel import ( + detect_beats, + detect_silence, + map_silence_to_timeline, + media_src_to_path, +) +from fcpxml.model_manager import load_silence_config +from fcpxml.models import FlashFrameSeverity, TimeValue +from fcpxml.writer import FCPXMLModifier +from server_tools._shared import ( + AUDIO_MEDIA_EXTENSIONS, + MAX_MEDIA_FILE_SIZE, + _detect_duplicate_groups, + _detect_flash_frames, + _detect_gaps, + _fmt_suggestions, + _format_clip_table, + _markdown_table, + _require_timeline, + _setup_modifier, + _text_result, + _validate_filepath, + _validate_output_path, + format_duration, + format_timecode, +) + +TOOLS = [ + Tool( + name="find_short_cuts", + description="Find clips shorter than threshold (flash frame detection)", + inputSchema={ + "type": "object", + "properties": { + "filepath": {"type": "string"}, + "threshold_seconds": {"type": "number", "default": 0.5} + }, + "required": ["filepath"] + } + ), + Tool( + name="find_long_clips", + description="Find clips longer than threshold", + inputSchema={ + "type": "object", + "properties": { + "filepath": {"type": "string"}, + "threshold_seconds": {"type": "number", "default": 10.0} + }, + "required": ["filepath"] + } + ), + Tool( + name="analyze_pacing", + description="Analyze edit pacing with suggestions for improvements", + inputSchema={ + "type": "object", + "properties": {"filepath": {"type": "string"}}, + "required": ["filepath"] + } + ), + Tool( + name="detect_flash_frames", + description="Find ultra-short clips (flash frames) that are likely errors, with severity categorization", + inputSchema={ + "type": "object", + "properties": { + "filepath": {"type": "string", "description": "Path to FCPXML file"}, + "critical_threshold_frames": {"type": "integer", "default": 2, "description": "Frames below this = critical (default: 2)"}, + "warning_threshold_frames": {"type": "integer", "default": 6, "description": "Frames below this = warning (default: 6)"} + }, + "required": ["filepath"] + } + ), + Tool( + name="detect_duplicates", + description="Find clips using the same source media (potential duplicates)", + inputSchema={ + "type": "object", + "properties": { + "filepath": {"type": "string", "description": "Path to FCPXML file"}, + "mode": {"type": "string", "enum": ["same_source", "overlapping_ranges", "identical"], "default": "same_source", "description": "Detection mode"} + }, + "required": ["filepath"] + } + ), + Tool( + name="detect_gaps", + description="Find unintentional gaps in the timeline", + inputSchema={ + "type": "object", + "properties": { + "filepath": {"type": "string", "description": "Path to FCPXML file"}, + "min_gap_frames": {"type": "integer", "default": 1, "description": "Minimum gap size to detect (default: 1 frame)"} + }, + "required": ["filepath"] + } + ), + Tool( + name="validate_timeline", + description="Comprehensive timeline health check for flash frames, gaps, duplicates, and issues", + inputSchema={ + "type": "object", + "properties": { + "filepath": {"type": "string", "description": "Path to FCPXML file"}, + "checks": {"type": "array", "items": {"type": "string", "enum": ["all", "flash_frames", "gaps", "duplicates", "offsets"]}, "default": ["all"], "description": "Which checks to run"} + }, + "required": ["filepath"] + } + ), + Tool( + name="detect_media_silence", + description="Detect REAL silence by analyzing each clip's source audio with ffmpeg silencedetect, mapped into timeline time. Unlike detect_silence_candidates (XML-only heuristics), this reads the actual media files referenced by the timeline. Requires ffmpeg; clips whose media is missing or unreadable are reported, not failed.", + inputSchema={ + "type": "object", + "properties": { + "filepath": {"type": "string", "description": "Path to FCPXML file"}, + "noise_db": {"type": "number", "description": "Silence threshold in dBFS, -120 to 0. Falls back to the saved silence settings (default -30)"}, + "min_silence": {"type": "number", "description": "Minimum silence duration in seconds to report. Falls back to the saved silence settings (default 0.5)"}, + "clip_name": {"type": "string", "description": "Only analyze the clip with this name"}, + }, + "required": ["filepath"] + } + ), + Tool( + name="detect_beats", + description="Detect musical beats and tempo in an audio/video file (librosa beat tracker). Writes a beats JSON next to the media file that plugs directly into import_beat_markers + snap_to_beats for beat-synced editing. Requires the optional [intelligence] extra (librosa); degrades to an install hint without it.", + inputSchema={ + "type": "object", + "properties": { + "media_path": {"type": "string", "description": "Path to audio/video file (.wav, .mp3, .m4a, .aac, .aif, .flac, .mov, .mp4)"}, + }, + "required": ["media_path"] + } + ), + Tool( + name="remove_media_silence", + description="Detect REAL silence in each clip's source audio (ffmpeg) and CUT it out of the timeline with ripple. Clips are split around silence; the silent middles are removed and everything after shifts earlier. Non-destructive: writes a _silence_removed copy. Preview with detect_media_silence first.", + inputSchema={ + "type": "object", + "properties": { + "filepath": {"type": "string", "description": "Path to FCPXML file"}, + "noise_db": {"type": "number", "description": "Silence threshold in dBFS, -120 to 0. Falls back to the saved silence settings (default -30)"}, + "min_silence": {"type": "number", "description": "Minimum silence duration in seconds to cut. Falls back to the saved silence settings (default 0.5)"}, + "padding": {"type": "number", "description": "Seconds of silence to keep on each side of a cut so edits breathe (max 5). Falls back to the saved silence settings (default 0.05)"}, + "clip_name": {"type": "string", "description": "Only cut silence in the clip with this name"}, + "output_path": {"type": "string", "description": "Output path (default: adds _silence_removed suffix)"}, + }, + "required": ["filepath"] + } + ), + Tool( + name="detect_silence_candidates", + description="Detect potential silence/dead air using timeline heuristics (gaps, ultra-short clips, name patterns, duration anomalies)", + inputSchema={ + "type": "object", + "properties": { + "filepath": {"type": "string", "description": "Path to FCPXML file"}, + "min_gap_seconds": {"type": "number", "default": 0.5, "description": "Minimum gap duration to flag"}, + "patterns": {"type": "array", "items": {"type": "string"}, "description": "Name patterns to match (default: gap, silence, room tone)"}, + }, + "required": ["filepath"] + } + ), + Tool( + name="remove_silence_candidates", + description="Remove or mark detected silence candidates from timeline", + inputSchema={ + "type": "object", + "properties": { + "filepath": {"type": "string", "description": "Path to FCPXML file"}, + "mode": {"type": "string", "enum": ["delete", "mark"], "default": "mark", "description": "delete=remove clips/gaps, mark=add red markers"}, + "min_gap_seconds": {"type": "number", "default": 0.5}, + "min_confidence": {"type": "number", "default": 0.7, "description": "Only act on candidates above this confidence"}, + "output_path": {"type": "string", "description": "Output path (default: adds _silence_cleaned suffix)"} + }, + "required": ["filepath"] + } + ), +] + + +async def handle_find_short_cuts(arguments: dict) -> Sequence[TextContent]: + project, tl = _require_timeline(arguments["filepath"]) + threshold = arguments.get("threshold_seconds", 0.5) + short = tl.get_clips_shorter_than(threshold) + if not short: + return _text_result(f"No clips shorter than {threshold}s") + return _text_result(_format_clip_table( + short, f"# Short Clips (< {threshold}s) - {len(short)} found", + )) + + +async def handle_find_long_clips(arguments: dict) -> Sequence[TextContent]: + project, tl = _require_timeline(arguments["filepath"]) + threshold = arguments.get("threshold_seconds", 10.0) + long = tl.get_clips_longer_than(threshold) + if not long: + return _text_result(f"No clips longer than {threshold}s") + return _text_result(_format_clip_table( + long, f"# Long Clips (> {threshold}s) - {len(long)} found", + )) + + +async def handle_analyze_pacing(arguments: dict) -> Sequence[TextContent]: + project, tl = _require_timeline(arguments["filepath"]) + if not tl.clips: + return _text_result("No clips to analyze") + durs = [c.duration_seconds for c in tl.clips] + avg = sum(durs) / len(durs) + q_len = len(durs) // 4 or 1 + segments = [durs[i:i+q_len] for i in range(0, len(durs), q_len)][:4] + seg_avgs = [sum(s)/len(s) if s else 0 for s in segments] + suggestions = [] + flash = [c for c in tl.clips if c.duration_seconds < 0.2] + if flash: + suggestions.append(f" {len(flash)} potential flash frames (< 0.2s)") + long = [c for c in tl.clips if c.duration_seconds > 30] + if long: + suggestions.append(f" {len(long)} long takes (> 30s) - consider trimming") + if len(seg_avgs) >= 4 and seg_avgs[3] < seg_avgs[0] * 0.7: + suggestions.append(" Pacing accelerates toward end - good for building energy") + elif len(seg_avgs) >= 4 and seg_avgs[3] > seg_avgs[0] * 1.3: + suggestions.append(" Pacing slows toward end - consider tightening") + return _text_result(f"""# Pacing Analysis: {tl.name} + +## Overall +- **Avg Cut**: {format_duration(avg)} +- **Cuts/Min**: {tl.cuts_per_minute:.1f} + +## By Section +| Q1 | Q2 | Q3 | Q4 | +|----|----|----|----| +| {format_duration(seg_avgs[0]) if len(seg_avgs) > 0 else 'N/A'} | {format_duration(seg_avgs[1]) if len(seg_avgs) > 1 else 'N/A'} | {format_duration(seg_avgs[2]) if len(seg_avgs) > 2 else 'N/A'} | {format_duration(seg_avgs[3]) if len(seg_avgs) > 3 else 'N/A'} | + +## Suggestions +{_fmt_suggestions(suggestions)} +""") + + +async def handle_detect_flash_frames(arguments: dict) -> Sequence[TextContent]: + project, tl = _require_timeline(arguments["filepath"]) + critical_threshold = arguments.get("critical_threshold_frames", 2) + warning_threshold = arguments.get("warning_threshold_frames", 6) + + flash_frames = _detect_flash_frames( + tl, critical_threshold=critical_threshold, warning_threshold=warning_threshold, + ) + + if not flash_frames: + return _text_result(f"No flash frames detected (threshold: {warning_threshold} frames)") + + critical = [f for f in flash_frames if f.severity == FlashFrameSeverity.CRITICAL] + warnings = [f for f in flash_frames if f.severity == FlashFrameSeverity.WARNING] + + result = f"""# Flash Frame Detection + +## Summary +- **Critical** (< {critical_threshold} frames): {len(critical)} found +- **Warning** (< {warning_threshold} frames): {len(warnings)} found +- **Total**: {len(flash_frames)} flash frames + +## Critical Flash Frames +""" + flash_headers = ["Clip", "Timecode", "Frames", "Duration"] + if critical: + result += _markdown_table(flash_headers, [ + [f.clip_name, format_timecode(f.start), f"{f.duration_frames}f", format_duration(f.duration_seconds)] + for f in critical + ]) + "\n" + else: + result += "_None_\n" + + result += "\n## Warning Flash Frames\n" + if warnings: + result += _markdown_table(flash_headers, [ + [f.clip_name, format_timecode(f.start), f"{f.duration_frames}f", format_duration(f.duration_seconds)] + for f in warnings + ]) + "\n" + else: + result += "_None_\n" + + result += "\n*Use `fix_flash_frames` to automatically resolve these issues.*" + return _text_result(result) + + +async def handle_detect_duplicates(arguments: dict) -> Sequence[TextContent]: + project, tl = _require_timeline(arguments["filepath"]) + mode = arguments.get("mode", "same_source") + + duplicates = _detect_duplicate_groups(tl, mode=mode) + + if not duplicates: + return _text_result(f"No duplicate clips found (mode: {mode})") + + result = f"""# Duplicate Clip Detection + +## Summary +- **Mode**: {mode} +- **Duplicate Groups**: {len(duplicates)} +- **Total Duplicate Clips**: {sum(g.count for g in duplicates)} + +## Duplicate Groups +""" + for group in duplicates: + result += f"\n### {group.source_name} ({group.count} uses)\n" + result += "| Clip Name | Timeline Position | Duration |\n|-----------|-------------------|----------|\n" + for c in group.clips: + result += f"| {c['name']} | {c['timecode']} | {format_duration(c['duration'])} |\n" + + return _text_result(result) + + +async def handle_detect_gaps(arguments: dict) -> Sequence[TextContent]: + project, tl = _require_timeline(arguments["filepath"]) + min_gap_frames = arguments.get("min_gap_frames", 1) + + gaps = _detect_gaps(tl, min_gap_frames=min_gap_frames) + + if not gaps: + return _text_result(f"No gaps detected (minimum: {min_gap_frames} frame(s))") + + result = f"""# Gap Detection + +## Summary +- **Gaps Found**: {len(gaps)} +- **Total Gap Duration**: {format_duration(sum(g.duration_seconds for g in gaps))} +- **Minimum Detection**: {min_gap_frames} frame(s) + +## Gaps +""" + result += _markdown_table( + ["Position", "Duration", "Between"], + [[gap.timecode, f"{gap.duration_frames}f ({format_duration(gap.duration_seconds)})", + f"{gap.previous_clip} -> {gap.next_clip}"] for gap in gaps], + ) + "\n" + + result += "\n*Use `fill_gaps` to automatically close these gaps.*" + return _text_result(result) + + +async def handle_validate_timeline(arguments: dict) -> Sequence[TextContent]: + project, tl = _require_timeline(arguments["filepath"]) + checks = arguments.get("checks", ["all"]) + run_all = "all" in checks + + issues: list[str] = [] + flash_count = 0 + gap_count = 0 + duplicate_count = 0 + + if run_all or "flash_frames" in checks: + flashes = _detect_flash_frames(tl) + flash_count = len(flashes) + for f in flashes: + severity = "error" if f.severity == FlashFrameSeverity.CRITICAL else "warning" + issues.append( + f"- [{severity.upper()}] Flash frame: {f.clip_name} " + f"({f.duration_frames}f) at {format_timecode(f.start)}" + ) + + if run_all or "gaps" in checks: + detected_gaps = _detect_gaps(tl) + gap_count = len(detected_gaps) + for g in detected_gaps: + issues.append(f"- [WARNING] Gap: {g.duration_frames}f at {g.timecode}") + + if run_all or "duplicates" in checks: + dup_groups = _detect_duplicate_groups(tl) + for group in dup_groups: + duplicate_count += group.count + issues.append( + f"- [INFO] Duplicate source: {group.source_name} ({group.count} uses)" + ) + + error_weight = 10 + warning_weight = 3 + info_weight = 1 + errors = len([i for i in issues if "[ERROR]" in i]) + warnings = len([i for i in issues if "[WARNING]" in i]) + infos = len([i for i in issues if "[INFO]" in i]) + penalty = (errors * error_weight) + (warnings * warning_weight) + (infos * info_weight) + health_score = max(0, 100 - penalty) + + result = f"""# Timeline Validation: {tl.name} + +## Health Score: {health_score}% + +## Summary +| Check | Count | Status | +|-------|-------|--------| +| Flash Frames | {flash_count} | {'PASS' if flash_count == 0 else 'FAIL'} | +| Gaps | {gap_count} | {'PASS' if gap_count == 0 else 'WARN'} | +| Duplicate Sources | {duplicate_count} | {'PASS' if duplicate_count == 0 else 'INFO'} | + +## Issues ({len(issues)}) +""" + if issues: + result += "\n".join(issues[:20]) + if len(issues) > 20: + result += f"\n... and {len(issues) - 20} more issues" + else: + result += "_No issues found!_" + + result += "\n\n*Use `fix_flash_frames` and `fill_gaps` to automatically resolve issues.*" + return _text_result(result) + + +async def handle_detect_media_silence(arguments: dict) -> Sequence[TextContent]: + # Unpassed thresholds come from the persisted silence settings (the app's + # own slider), not a hardcoded constant, so detection previews exactly + # what removal would cut. + saved = load_silence_config() + noise_db = float(arguments.get("noise_db", saved["noise_db"])) + min_silence = float(arguments.get("min_silence", saved["min_silence"])) + # Same bounds detect_silence() enforces — validated here so a bad request + # fails before any media file is opened. + if not (-120.0 <= noise_db <= 0.0): + raise ValueError(f"noise_db must be between -120 and 0 dB, got {noise_db}") + if not (0 < min_silence <= 3600): + raise ValueError(f"min_silence must be between 0 and 3600 seconds, got {min_silence}") + + filepath = _validate_filepath(arguments["filepath"], ('.fcpxml', '.fcpxmld')) + modifier = FCPXMLModifier(filepath) + clip_filter = arguments.get("clip_name") + + max_media_probes = 100 + findings: list[tuple[str, float, float]] = [] + skipped: list[tuple[str, str]] = [] + probe_cache: dict[str, list | None] = {} + for el in [el for _, el in modifier._iter_spine_clips()]: + name = el.get("name", "") + if clip_filter and name != clip_filter: + continue + src = modifier.resources.get(el.get("ref", ""), {}).get("src", "") + media_path = media_src_to_path(src) + if not media_path or not Path(media_path).is_file(): + skipped.append((name, "media file missing")) + continue + if media_path not in probe_cache: + if len(probe_cache) >= max_media_probes: + skipped.append((name, f"probe cap reached ({max_media_probes} media files)")) + continue + probe_cache[media_path] = detect_silence( + media_path, noise_db=noise_db, min_duration=min_silence + ) + silences = probe_cache[media_path] + if silences is None: + skipped.append((name, "unanalyzable (ffmpeg missing or media unreadable)")) + continue + source_start = modifier.source_file_start(el).to_seconds() + clip_duration = modifier._parse_time(el.get("duration", "0s")).to_seconds() + timeline_offset = modifier._parse_time(el.get("offset", "0s")).to_seconds() + mapped = map_silence_to_timeline( + silences, source_start, clip_duration, timeline_offset + ) + findings.extend((name, start, end) for start, end in mapped) + + total_silence = sum(end - start for _, start, end in findings) + result = f"""# Media Silence Detection (real audio analysis) + +## Summary +- **Threshold**: {noise_db} dB for >= {min_silence}s +- **Media Files Probed**: {len(probe_cache)} +- **Silence Spans Found**: {len(findings)} ({format_duration(total_silence)} total) +""" + if findings: + result += "\n## Silence Spans (timeline time)\n" + result += _markdown_table( + ["Clip", "Start", "End", "Duration"], + [[name, f"{start:.2f}s", f"{end:.2f}s", f"{end - start:.2f}s"] + for name, start, end in findings], + ) + "\n" + result += "\n*To remove: `split_clip` at each boundary, then `delete_clips` with ripple.*" + if skipped: + result += "\n## Skipped Clips\n" + result += _markdown_table( + ["Clip", "Reason"], [[name, reason] for name, reason in skipped] + ) + "\n" + if not findings and not skipped: + result += "\nNo silence detected in any clip's source audio." + return _text_result(result) + + +async def handle_detect_beats(arguments: dict) -> Sequence[TextContent]: + media_path = _validate_filepath( + arguments["media_path"], AUDIO_MEDIA_EXTENSIONS, max_size=MAX_MEDIA_FILE_SIZE + ) + + result = detect_beats(media_path) + if result is None: + return _text_result( + "Beat detection unavailable — librosa is not installed or the file " + "could not be analyzed.\n\nInstall the optional media-intelligence " + "extra:\n\n pip install 'fcp-mcp-server[intelligence]'" + ) + + bpm, beats = result["bpm"], result["beats"] + beats_data = { + "source": str(Path(media_path).name), + "bpm": round(bpm, 2), + "beats": [round(b, 4) for b in beats], + "downbeats": [round(b, 4) for b in beats[::4]], + } + json_path = _validate_output_path( + str(Path(media_path).with_name(Path(media_path).stem + "_beats.json")), + anchor_dir=str(Path(media_path).parent), + ) + with open(json_path, "w") as f: + json.dump(beats_data, f, indent=2) + + preview = beats[:16] + result_text = f"""# Beat Detection + +## Summary +- **Source**: {Path(media_path).name} +- **Estimated Tempo**: {bpm:.1f} BPM +- **Beats Detected**: {len(beats)} ({format_duration(beats[-1]) if beats else '0s'} span) +- **Beats JSON**: {json_path} + +## First Beats +""" + result_text += _markdown_table( + ["#", "Time"], + [[str(i + 1), f"{b:.3f}s"] for i, b in enumerate(preview)], + ) + "\n" + result_text += ( + f"\n*Next: `import_beat_markers` with beats_path=\"{json_path}\" to place " + "markers, then `snap_to_beats` to align your cuts.*" + ) + return _text_result(result_text) + + +async def handle_remove_media_silence(arguments: dict) -> Sequence[TextContent]: + saved = load_silence_config() + noise_db = float(arguments.get("noise_db", saved["noise_db"])) + min_silence = float(arguments.get("min_silence", saved["min_silence"])) + padding = float(arguments.get("padding", saved["padding"])) + if not (-120.0 <= noise_db <= 0.0): + raise ValueError(f"noise_db must be between -120 and 0 dB, got {noise_db}") + if not (0 < min_silence <= 3600): + raise ValueError(f"min_silence must be between 0 and 3600 seconds, got {min_silence}") + if not (0 <= padding <= 5): + raise ValueError(f"padding must be between 0 and 5 seconds, got {padding}") + + filepath, output_path, modifier = _setup_modifier(arguments, "_silence_removed") + clip_filter = arguments.get("clip_name") + to_frame_timevalue = modifier.snap_seconds_to_frame + + max_media_probes = 100 + cuts_made: list[tuple[str, int, float]] = [] + skipped: list[tuple[str, str]] = [] + probe_cache: dict[str, list | None] = {} + spine_clips = [el for _, el in modifier._iter_spine_clips()] + for el in spine_clips: + name = el.get("name", "") + if clip_filter and name != clip_filter: + continue + src = modifier.resources.get(el.get("ref", ""), {}).get("src", "") + media_path = media_src_to_path(src) + if not media_path or not Path(media_path).is_file(): + skipped.append((name, "media file missing")) + continue + if media_path not in probe_cache: + if len(probe_cache) >= max_media_probes: + skipped.append((name, f"probe cap reached ({max_media_probes} media files)")) + continue + probe_cache[media_path] = detect_silence( + media_path, noise_db=noise_db, min_duration=min_silence + ) + silences = probe_cache[media_path] + if silences is None: + skipped.append((name, "unanalyzable (ffmpeg missing or media unreadable)")) + continue + + clip_source_start = modifier.source_file_start(el).to_seconds() + clip_duration = modifier._parse_time(el.get("duration", "0s")).to_seconds() + cut_ranges = [] + for sil_start, sil_end in silences: + # Source time -> clip-relative, padded so cuts breathe. + cut_start = max(sil_start, clip_source_start) - clip_source_start + padding + cut_end = min(sil_end, clip_source_start + clip_duration) - clip_source_start - padding + if cut_end > cut_start: + cut_ranges.append((to_frame_timevalue(cut_start), to_frame_timevalue(cut_end))) + if not cut_ranges: + continue + removed = modifier.cut_clip_ranges(el, cut_ranges) + if removed > TimeValue.zero(): + cuts_made.append((name, len(cut_ranges), removed.to_seconds())) + + if not cuts_made: + text = "# Media Silence Removal\n\nNo silence found to remove — file unchanged (nothing saved)." + if skipped: + text += "\n\n## Skipped Clips\n" + _markdown_table( + ["Clip", "Reason"], [[name, reason] for name, reason in skipped] + ) + return _text_result(text) + + modifier.remove_trailing_gaps() + modifier.save(output_path) + total_removed = sum(seconds for _, _, seconds in cuts_made) + result = f"""# Media Silence Removal (real audio analysis) + +## Summary +- **Threshold**: {noise_db} dB for >= {min_silence}s, padding {padding}s +- **Clips Cut**: {len(cuts_made)} +- **Total Removed**: {format_duration(total_removed)} + +## Cuts +""" + result += _markdown_table( + ["Clip", "Silence Spans Cut", "Removed"], + [[name, str(count), f"{seconds:.2f}s"] for name, count, seconds in cuts_made], + ) + "\n" + if skipped: + result += "\n## Skipped Clips\n" + _markdown_table( + ["Clip", "Reason"], [[name, reason] for name, reason in skipped] + ) + "\n" + result += f"\nSaved to: {output_path}\n\n*Preview first next time with `detect_media_silence`. Original file untouched.*" + return _text_result(result) + + +async def handle_detect_silence_candidates(arguments: dict) -> Sequence[TextContent]: + filepath = _validate_filepath(arguments["filepath"], ('.fcpxml', '.fcpxmld')) + modifier = FCPXMLModifier(filepath) + candidates = modifier.detect_silence_candidates( + min_gap_seconds=arguments.get("min_gap_seconds", 0.5), + patterns=arguments.get("patterns"), + ) + + if not candidates: + return _text_result("No silence candidates detected.") + + result = f"# Silence Candidates Detected\n\n**Found**: {len(candidates)}\n\n" + result += "| # | Timecode | Duration | Reason | Confidence | Clip |\n" + result += "|---|----------|----------|--------|------------|------|\n" + for i, c in enumerate(candidates, 1): + result += ( + f"| {i} | {c['start_timecode']} | {format_duration(c['duration_seconds'])} | " + f"{c['reason']} | {c['confidence']:.0%} | {c.get('clip_name') or '-'} |\n" + ) + result += ( + "\n**Note**: Detection uses timeline heuristics (gaps, ultra-short clips, name patterns). " + "Review candidates before removing — some may be intentional." + ) + return _text_result(result) + + +async def handle_remove_silence_candidates(arguments: dict) -> Sequence[TextContent]: + filepath, output_path, modifier = _setup_modifier(arguments, "_silence_cleaned") + actions = modifier.remove_silence_candidates( + mode=arguments.get("mode", "mark"), + min_gap_seconds=arguments.get("min_gap_seconds", 0.5), + min_confidence=arguments.get("min_confidence", 0.7), + ) + modifier.save(output_path) + + if not actions: + return _text_result("No silence candidates met the confidence threshold.") + + mode = arguments.get("mode", "mark") + result = f"# Silence Candidates {'Marked' if mode == 'mark' else 'Removed'}\n\n" + result += f"**Actions taken**: {len(actions)}\n\n" + for a in actions: + result += f"- **{a['action']}** {a.get('clip_name', 'gap')} ({a['reason']})\n" + result += f"\nSaved to: `{output_path}`" + return _text_result(result) + + +HANDLERS = { + "find_short_cuts": handle_find_short_cuts, + "find_long_clips": handle_find_long_clips, + "analyze_pacing": handle_analyze_pacing, + "detect_flash_frames": handle_detect_flash_frames, + "detect_duplicates": handle_detect_duplicates, + "detect_gaps": handle_detect_gaps, + "validate_timeline": handle_validate_timeline, + "detect_media_silence": handle_detect_media_silence, + "detect_beats": handle_detect_beats, + "remove_media_silence": handle_remove_media_silence, + "detect_silence_candidates": handle_detect_silence_candidates, + "remove_silence_candidates": handle_remove_silence_candidates, +} diff --git a/code/server_tools/roles.py b/code/server_tools/roles.py new file mode 100644 index 0000000..7ac4a67 --- /dev/null +++ b/code/server_tools/roles.py @@ -0,0 +1,133 @@ +"""Roles — tool schemas and handlers. + +Extracted from server.py; see Engine/docs/03_SERVER_TOOLS.md for the tool catalog. +""" + +from __future__ import annotations + +from typing import Sequence + +from mcp.types import TextContent, Tool + +from server_tools._shared import ( + _require_timeline, + _setup_modifier, + _text_result, + format_duration, +) + +TOOLS = [ + Tool( + name="assign_role", + description="Set the audio or video role on a clip (dialogue, music, effects, titles, etc.)", + inputSchema={ + "type": "object", + "properties": { + "filepath": {"type": "string", "description": "Path to FCPXML file"}, + "clip_id": {"type": "string", "description": "Clip name or ID"}, + "audio_role": {"type": "string", "description": "Audio role (e.g., dialogue, music, effects)"}, + "video_role": {"type": "string", "description": "Video role (e.g., video, titles)"}, + "output_path": {"type": "string", "description": "Output path (default: adds _modified suffix)"} + }, + "required": ["filepath", "clip_id"] + } + ), + Tool( + name="filter_by_role", + description="List all clips matching a specific audio or video role", + inputSchema={ + "type": "object", + "properties": { + "filepath": {"type": "string", "description": "Path to FCPXML file"}, + "role": {"type": "string", "description": "Role name to filter by"}, + "role_type": {"type": "string", "enum": ["audio", "video", "any"], "default": "any", "description": "Which role type to search"}, + }, + "required": ["filepath", "role"] + } + ), + Tool( + name="export_role_stems", + description="Export clip list grouped by role for audio mixing stem planning", + inputSchema={ + "type": "object", + "properties": { + "filepath": {"type": "string", "description": "Path to FCPXML file"}, + }, + "required": ["filepath"] + } + ), +] + + +async def handle_assign_role(arguments: dict) -> Sequence[TextContent]: + filepath, output_path, modifier = _setup_modifier(arguments) + modifier.assign_role( + clip_id=arguments["clip_id"], + audio_role=arguments.get("audio_role"), + video_role=arguments.get("video_role"), + ) + modifier.save(output_path) + + roles_set = [] + if arguments.get("audio_role"): + roles_set.append(f"audioRole={arguments['audio_role']}") + if arguments.get("video_role"): + roles_set.append(f"videoRole={arguments['video_role']}") + + return _text_result(( + f"Set {', '.join(roles_set)} on '{arguments['clip_id']}'\n\n" + f"Saved to: `{output_path}`" + )) + + +async def handle_filter_by_role(arguments: dict) -> Sequence[TextContent]: + project, tl = _require_timeline(arguments["filepath"]) + + role = arguments["role"].lower() + role_type = arguments.get("role_type", "any") + matches = [] + + for clip in tl.clips: + if role_type in ("audio", "any") and clip.audio_role.lower() == role: + matches.append((clip.name, "audio", clip.audio_role, format_duration(clip.duration_seconds))) + if role_type in ("video", "any") and clip.video_role.lower() == role: + matches.append((clip.name, "video", clip.video_role, format_duration(clip.duration_seconds))) + + if not matches: + return _text_result(f"No clips found with role '{role}'.") + + result = f"# Clips with role '{role}'\n\n" + result += "| Clip | Type | Role | Duration |\n|------|------|------|----------|\n" + for name, rtype, rval, dur in matches: + result += f"| {name} | {rtype} | {rval} | {dur} |\n" + return _text_result(result) + + +async def handle_export_role_stems(arguments: dict) -> Sequence[TextContent]: + project, tl = _require_timeline(arguments["filepath"]) + + stems: dict[str, list] = {} + for clip in tl.clips: + role = clip.audio_role or "unassigned" + stems.setdefault(role, []).append(clip) + + for cc in tl.connected_clips: + role = cc.role or "unassigned" + stems.setdefault(role, []).append(cc) + + result = f"# Audio Stem Plan for {tl.name}\n\n" + for role, clips in sorted(stems.items()): + total_dur = sum(c.duration_seconds for c in clips) + result += f"## {role.title()} ({len(clips)} clips, {format_duration(total_dur)})\n\n" + for c in clips: + result += f"- {c.name} ({format_duration(c.duration_seconds)})\n" + result += "\n" + + return _text_result(result) + + +HANDLERS = { + "assign_role": handle_assign_role, + "filter_by_role": handle_filter_by_role, + "export_role_stems": handle_export_role_stems, +} diff --git a/code/server_tools/subtitles.py b/code/server_tools/subtitles.py new file mode 100644 index 0000000..4f009a5 --- /dev/null +++ b/code/server_tools/subtitles.py @@ -0,0 +1,283 @@ +"""Legendas dinâmicas (geração → validação) — tool schemas and handlers. + +Extracted from server.py; see Engine/docs/03_SERVER_TOOLS.md for the tool catalog. +""" + +from __future__ import annotations + +import json +from pathlib import Path +from typing import Sequence + +from mcp.types import TextContent, Tool + +from fcpxml.media_intel import media_src_to_path +from fcpxml.model_manager import load_dynamic_subtitle_config +from fcpxml.models import DynamicSubtitleConfig, WordLook, WordStyle +from fcpxml.writer import FCPXMLModifier +from server_tools._shared import ( + _load_or_transcribe, + _markdown_table, + _setup_modifier, + _text_result, + _validate_filepath, +) + +TOOLS = [ + Tool( + name="validate_subtitle_layout", + description="Re-measure every title/subtitle in an FCPXML and report spatial collisions, frame and safe-area violations, and font fallbacks. Detects overlapping boxes only for titles on screen at the same time (half-open time intervals, so a title ending exactly as the next begins is never flagged). Returns a severity (none/render_tolerance/warning/probable/severe), the list of issues with suggested corrections, and summary counts.", + inputSchema={ + "type": "object", + "properties": { + "filepath": {"type": "string", "description": "Path to FCPXML file"}, + "safe_margin_x": {"type": "number", "default": 0.05, "description": "Fraction of frame width to inset from each side (0.05 = 5%)"}, + "safe_margin_y": {"type": "number", "default": 0.05, "description": "Fraction of frame height to inset from top/bottom"}, + "min_font_size": {"type": "number", "description": "Flag titles whose emitted fontSize is below this readable minimum"}, + "min_distance": {"type": "number", "description": "Flag same-block titles closer than this many pixels (insufficient_spacing)"}, + "max_distance": {"type": "number", "description": "Flag same-block titles farther than this many pixels (excessive_spacing)"}, + "output_format": {"type": "string", "enum": ["markdown", "json"], "default": "markdown", "description": "Report format"} + }, + "required": ["filepath"] + } + ), + Tool( + name="generate_dynamic_subtitles", + description="Generate progressive-composition subtitles as real, editable FCPXML title clips (the 'Text'/Basic Text template). Whisper's segments become sentences; each sentence is diagrammed as stacked blocks — supporting words grouped small in a grotesque, the sentence's key word alone and large in a display italic, body lines staggered to opposite edges. One <title> per block: each enters as its own words are spoken and stays on screen, so the sentence assembles itself, and every block clears at the same instant. Set granularity='word' for the older one-title-per-word rhythm. A sentence too tall for the band splits into successive compositions. These are TITLES, not captions: no subtitles role, so they render over the video without enabling caption display. Uses each media file's local Whisper word-level transcript (_transcript.json, auto-transcribes if missing). Style fields below (band_height through inactive_color) fall back to the style saved from the app's 'Legendas Dinâmicas' screen (~/.fcp-mcp-server/config.json via save_dynamic_subtitle_config) when omitted — pass a value here only to override that for one call. Non-destructive: writes a _dynamic_subtitles copy.", + inputSchema={ + "type": "object", + "properties": { + "filepath": {"type": "string", "description": "Path to FCPXML file"}, + "clip_name": {"type": "string", "description": "Only caption the clip with this name (default: all spine clips with matched source media)"}, + "model": {"type": "string", "default": "base", "description": "Whisper model size if transcription is needed"}, + "language": {"type": "string", "description": "ISO language code hint (e.g. 'en'); auto-detected if omitted"}, + "band_height": {"type": "number", "description": "Fraction of frame height the sentence block may fill before splitting into another block. Falls back to the saved style (default 0.22 — about three lines)"}, + "block_center_y": {"type": "number", "description": "Vertical centre of the block in canvas points; negative sits below frame centre. Falls back to the saved style (default -167, just under centre)"}, + "line_gap": {"type": "number", "description": "Air between stacked lines in canvas points. Lines are stacked on their real ink, so this is the whole distance beyond the glyphs themselves; negative values deliberately tuck each line into the one above. Falls back to the saved style (default 8)"}, + "granularity": {"type": "string", "enum": ["phrase", "word"], "default": "phrase", "description": "'phrase': one title per LINE of the composition, key word set large (the reference look). 'word': one title per word."}, + "emphasis_font": {"type": "string", "description": "Family for the key word (phrase mode). Must be installed on the editing Mac; unmeasured families fall back to estimated widths. Falls back to the saved style (default 'Playfair Display')"}, + "emphasis_face": {"type": "string", "description": "Face for the key word, e.g. 'Medium Italic' or a script/calligraphic face. Falls back to the saved style (default 'Medium Italic')"}, + "emphasis_size": {"type": "integer", "description": "Key-word size in canvas points, at the 2160x3840 reference frame. Falls back to the saved style (default 265)"}, + "emphasis_color": {"type": "string", "description": "RGBA (0-1, space-separated) for the key word (phrase mode). Defaults to active_color, so the block reads in a single colour unless the key word is deliberately set apart"}, + "text_scale": {"type": "number", "description": "Ratio between the title template's fontSize space and the canvas-point space it positions in. The \"Text\" template sizes type in frame pixels, so sizes are doubled on the way out. Falls back to the saved style (default 2.0). Lower it only if a template renders type larger than the chosen point size"}, + "font": {"type": "string", "description": "Title font family (supporting lines in phrase mode). Falls back to the saved style (default 'Helvetica Neue')"}, + "font_size": {"type": "integer", "description": "Supporting-line font size in canvas points, at the 2160x3840 reference frame. Falls back to the saved style (default 104)"}, + "active_color": {"type": "string", "description": "RGBA (0-1, space-separated) for even-indexed lines. Falls back to the saved style (default '1 1 1 1')"}, + "inactive_color": {"type": "string", "default": "0.7 0.7 0.7 1", "description": "RGBA (0-1, space-separated) for odd-indexed lines — alternates with active_color for visual variety between stacked lines"}, + "output_path": {"type": "string", "description": "Output path (default: adds _dynamic_subtitles suffix)"}, + }, + "required": ["filepath"] + } + ), +] + + +async def handle_validate_subtitle_layout(arguments: dict) -> Sequence[TextContent]: + """Validate title/subtitle layout for spatial collisions and safe-area + containment (collision.validate_titles over every <title> in the file).""" + filepath = _validate_filepath(arguments["filepath"], (".fcpxml", ".fcpxmld")) + modifier = FCPXMLModifier(filepath) + report = modifier.validate_subtitle_layout( + safe_margin_x=float(arguments.get("safe_margin_x", 0.05)), + safe_margin_y=float(arguments.get("safe_margin_y", 0.05)), + min_font_size=( + float(arguments["min_font_size"]) + if arguments.get("min_font_size") is not None else None + ), + min_distance=( + float(arguments["min_distance"]) + if arguments.get("min_distance") is not None else None + ), + max_distance=( + float(arguments["max_distance"]) + if arguments.get("max_distance") is not None else None + ), + ) + + if arguments.get("output_format") == "json": + return _text_result(json.dumps(report, indent=2)) + + summary = report["summary"] + lines = [ + "# Subtitle Layout Validation", + "", + f"## Summary (severity: {report['severity']})", + f"- **Titles**: {summary['title_count']}", + f"- **Issues**: {summary['issue_count']}", + f"- **Collisions**: {summary['spatial_collision']}", + f"- **Outside frame**: {summary['outside_frame']}", + f"- **Outside safe area**: {summary['outside_safe_area']}", + f"- **Font fallback**: {summary['font_missing']}", + f"- **Font too small**: {summary['font_too_small']}", + "", + ] + issues = report["issues"] + if issues: + lines.append(f"## Issues ({len(issues)})") + for issue in issues: + sev = issue["severity"].upper() + if issue["type"] == "spatial_collision": + corr = issue["suggested_correction"] + lines.append( + f"- [{sev}] collision: \"{issue['first_title']}\" x " + f"\"{issue['second_title']}\" " + f"(overlap {issue['overlap_width']:.0f}x" + f"{issue['overlap_height']:.0f} = " + f"{issue['overlap_area']:.0f}px, ratio " + f"{issue['overlap_ratio']:.2f}, move " + f"{corr['axis']} {corr['minimum_movement']:.0f}px)" + ) + else: + detail = issue.get("title", "") or issue.get("font", "") + lines.append(f"- [{sev}] {issue['type']}: {detail}".rstrip()) + else: + lines.append("_No issues found — no simultaneous titles overlap._") + + return _text_result("\n".join(lines)) + + +async def handle_generate_dynamic_subtitles(arguments: dict) -> Sequence[TextContent]: + """Generate per-word subtitle titles laid out as a block per sentence. + + Whisper's segments become sentences; each word becomes its own positioned + <title> connected clip, appearing as it is spoken and accumulating on + screen until the whole block clears at once. No compound clip. + + Reuses the same SOURCE-media -> TIMELINE mapping as ``transcript_markers`` + (``modifier.source_file_start`` per spine clip) so word timestamps land + at the correct position even across trimmed/multiple clips. + """ + model = arguments.get("model", "base") + language = arguments.get("language") + output_dir = arguments.get("output_dir") + clip_filter = arguments.get("clip_name") + + # Anything the caller didn't explicitly pass falls back to the style + # persisted from the "Legendas Dinâmicas" screen (~/.fcp-mcp-server/ + # config.json), not a hardcoded default — so the UI is the single place + # that configures the look, and every caller (app, MCP, this session) + # renders the same thing without threading 11 fields through every call. + saved = load_dynamic_subtitle_config() + body_color = arguments.get("active_color") or saved["active_color"] + config = DynamicSubtitleConfig( + style=WordStyle( + font=arguments.get("font") or saved["font"], + font_size=int(arguments.get("font_size", saved["font_size"])), + active_color=body_color, + inactive_color=arguments.get("inactive_color", "0.7 0.7 0.7 1"), + emphasis_look=WordLook( + int(arguments.get("emphasis_size", saved["emphasis_size"])), + arguments.get("emphasis_color") or saved["emphasis_color"] or body_color, + font=arguments.get("emphasis_font") or saved["emphasis_font"], + face=arguments.get("emphasis_face") or saved["emphasis_face"], + kerning=0.0, + ), + body_look=WordLook( + int(arguments.get("font_size", saved["font_size"])), + body_color, + font=arguments.get("font") or saved["font"], + face="Bold", + kerning=1.2, + ), + ), + band_height=float(arguments.get("band_height", saved["band_height"])), + block_center_y=float(arguments.get("block_center_y", saved["block_center_y"])), + granularity=arguments.get("granularity", "phrase"), + text_scale=float(arguments.get("text_scale", saved["text_scale"])), + line_gap=float(arguments.get("line_gap", saved["line_gap"])), + ) + + filepath, output_path, modifier = _setup_modifier(arguments, "_dynamic_subtitles") + + added: list[tuple[str, int, int]] = [] + skipped: list[tuple[str, str]] = [] + spine_clips = [el for _, el in modifier._iter_spine_clips()] + for el in spine_clips: + name = el.get("name", "") + if clip_filter and name != clip_filter: + continue + src = modifier.resources.get(el.get("ref", ""), {}).get("src", "") + media_path = media_src_to_path(src) + if not media_path or not Path(media_path).is_file(): + skipped.append((name, "media file missing")) + continue + data, reason = _load_or_transcribe(media_path, model, language, output_dir) + if data is None: + skipped.append((name, reason)) + continue + + clip_source_start = modifier.source_file_start(el).to_seconds() + clip_duration = modifier._parse_time(el.get("duration", "0s")).to_seconds() + window_end = clip_source_start + clip_duration + + clip_words = [ + { + "word": w.get("word", ""), + "start": float(w.get("start", 0.0)) - clip_source_start, + "end": float(w.get("end", 0.0)) - clip_source_start, + } + for w in data.get("words", []) + if clip_source_start <= float(w.get("start", 0.0)) < window_end + ] + if not clip_words: + skipped.append((name, "no words in clip's source range")) + continue + + # Sentence boundaries, rebased the same way, so each sentence becomes + # its own block of titles that builds up and then clears together. + # Overlap rather than containment: a segment straddling the clip's + # in-point still governs the words that made the cut. + clip_segments = [ + { + "start": float(s.get("start", 0.0)) - clip_source_start, + "end": float(s.get("end", 0.0)) - clip_source_start, + } + for s in data.get("segments", []) + if float(s.get("end", 0.0)) > clip_source_start + and float(s.get("start", 0.0)) < window_end + ] + + # Pass the element itself, not `name` — after ripple-cut/silence + # removal every fragment of an originally-named clip keeps the same + # `name`, so a name lookup here would resolve every clip in this + # loop to whichever one `self.clips` last indexed, stacking every + # clip's captions onto a single wrong spine element instead of each + # clip's own. See Engine/docs/05_EXPERIENCIAS.md, entry 2026-08-17. + lines = modifier.generate_dynamic_subtitles( + el, clip_words, config, segments=clip_segments + ) + added.append((name, len(lines), len(clip_words))) + + if not added: + text = "# Dynamic Subtitles\n\nNo captions generated — file unchanged (nothing saved)." + if skipped: + text += "\n\n## Skipped Clips\n" + _markdown_table( + ["Clip", "Reason"], [[n, r] for n, r in skipped] + ) + return _text_result(text) + + modifier.save(output_path) + total_lines = sum(lines for _, lines, _ in added) + total_words = sum(words for _, _, words in added) + result = "# Dynamic Subtitles Generated (local Whisper)\n\n## Summary\n" + result += ( + f"- **Clips Captioned**: {len(added)}\n" + f"- **Caption Lines (Title Clips)**: {total_lines}\n" + f"- **Total Words**: {total_words}\n\n" + ) + result += _markdown_table( + ["Clip", "Caption Lines", "Words"], + [[n, str(lines), str(words)] for n, lines, words in added], + ) + if skipped: + result += "\n## Skipped Clips\n" + _markdown_table( + ["Clip", "Reason"], [[n, r] for n, r in skipped] + ) + result += f"\n\nSaved to: `{output_path}`\n\n*Transcripts are cached as _transcript.json.*" + return _text_result(result) + + +HANDLERS = { + "validate_subtitle_layout": handle_validate_subtitle_layout, + "generate_dynamic_subtitles": handle_generate_dynamic_subtitles, +} diff --git a/code/server_tools/timeline.py b/code/server_tools/timeline.py new file mode 100644 index 0000000..844354b --- /dev/null +++ b/code/server_tools/timeline.py @@ -0,0 +1,400 @@ +"""Timeline & análise (Projeto) — tool schemas and handlers. + +Extracted from server.py; see Engine/docs/03_SERVER_TOOLS.md for the tool catalog. +""" + +from __future__ import annotations + +from typing import Sequence + +from mcp.types import TextContent, Tool + +from fcpxml.diff import compare_timelines +from fcpxml.models import MarkerType +from fcpxml.parser import FCPXMLParser +from fcpxml.writer import list_effects +from server_tools._shared import ( + _SANDBOX_ENABLED, + PROJECTS_DIR, + _require_timeline, + _text_result, + _validate_directory, + _validate_filepath, + find_fcpxml_files, + format_duration, + format_timecode, +) + +TOOLS = [ + Tool( + name="list_projects", + description="List all FCPXML projects in directory", + inputSchema={ + "type": "object", + "properties": { + "directory": {"type": "string", "description": "Directory to search (default: ~/Movies)"} + } + } + ), + Tool( + name="analyze_timeline", + description="Get comprehensive timeline statistics including duration, resolution, clip count, pacing metrics", + inputSchema={ + "type": "object", + "properties": {"filepath": {"type": "string", "description": "Path to FCPXML file"}}, + "required": ["filepath"] + } + ), + Tool( + name="list_clips", + description="List all clips with timecodes, durations, and metadata", + inputSchema={ + "type": "object", + "properties": { + "filepath": {"type": "string"}, + "limit": {"type": "integer", "description": "Max clips to return"} + }, + "required": ["filepath"] + } + ), + Tool( + name="list_markers", + description="Extract markers (chapter, todo, standard) with timestamps", + inputSchema={ + "type": "object", + "properties": { + "filepath": {"type": "string"}, + "marker_type": {"type": "string", "enum": ["all", "chapter", "todo", "standard", "completed"]}, + "format": {"type": "string", "enum": ["detailed", "youtube", "simple"]} + }, + "required": ["filepath"] + } + ), + Tool( + name="list_keywords", + description="Extract all keywords/tags from project", + inputSchema={ + "type": "object", + "properties": {"filepath": {"type": "string"}}, + "required": ["filepath"] + } + ), + Tool( + name="list_library_clips", + description="List all available clips in the library (source media, not yet on timeline)", + inputSchema={ + "type": "object", + "properties": { + "filepath": {"type": "string", "description": "Path to FCPXML file"}, + "keywords": {"type": "array", "items": {"type": "string"}, "description": "Filter by keywords"}, + "limit": {"type": "integer", "description": "Max clips to return"} + }, + "required": ["filepath"] + } + ), + Tool( + name="list_connected_clips", + description="List all connected clips (B-roll, titles, audio) with their lanes and parent clips", + inputSchema={ + "type": "object", + "properties": { + "filepath": {"type": "string", "description": "Path to FCPXML file"}, + "lane": {"type": "integer", "description": "Filter by lane number (positive=above, negative=below)"}, + }, + "required": ["filepath"] + } + ), + Tool( + name="list_compound_clips", + description="List compound clips (ref-clips) and their nested content", + inputSchema={ + "type": "object", + "properties": { + "filepath": {"type": "string", "description": "Path to FCPXML file"}, + }, + "required": ["filepath"] + } + ), + Tool( + name="list_roles", + description="List all audio/video roles used in the timeline with clip counts", + inputSchema={ + "type": "object", + "properties": { + "filepath": {"type": "string", "description": "Path to FCPXML file"}, + }, + "required": ["filepath"] + } + ), + Tool( + name="diff_timelines", + description="Compare two FCPXML files and report differences in clips, markers, transitions, and format", + inputSchema={ + "type": "object", + "properties": { + "filepath_a": {"type": "string", "description": "Path to first FCPXML file (baseline)"}, + "filepath_b": {"type": "string", "description": "Path to second FCPXML file (comparison)"}, + }, + "required": ["filepath_a", "filepath_b"] + } + ), + Tool( + name="list_effects", + description="List all available FCP transition effects with slugs and UUIDs", + inputSchema={ + "type": "object", + "properties": {}, + } + ), +] + + +async def handle_list_projects(arguments: dict) -> Sequence[TextContent]: + directory = arguments.get("directory", PROJECTS_DIR) + resolved_dir = _validate_directory( + directory, allowed_root=PROJECTS_DIR if _SANDBOX_ENABLED else None + ) + files = find_fcpxml_files(resolved_dir) + if not files: + return _text_result(f"No FCPXML files found in {directory}") + return _text_result(f"Found {len(files)} FCPXML file(s):\n" + "\n".join(f" - {f}" for f in files)) + + +async def handle_analyze_timeline(arguments: dict) -> Sequence[TextContent]: + project, tl = _require_timeline(arguments["filepath"]) + durs = [c.duration_seconds for c in tl.clips] + avg, med, mn, mx = (0, 0, 0, 0) if not durs else ( + sum(durs)/len(durs), sorted(durs)[len(durs)//2], min(durs), max(durs)) + return _text_result(f"""# Timeline Analysis: {tl.name} + +## Overview +- **Duration**: {format_duration(tl.duration.seconds)} +- **Resolution**: {tl.width}x{tl.height} @ {tl.frame_rate}fps + +## Clip Statistics +- **Total Clips**: {tl.total_clips} +- **Total Cuts**: {tl.total_cuts} +- **Transitions**: {len(tl.transitions)} + +## Pacing +- **Average**: {format_duration(avg)} +- **Median**: {format_duration(med)} +- **Shortest**: {format_duration(mn)} +- **Longest**: {format_duration(mx)} +- **Cuts/Minute**: {tl.cuts_per_minute:.1f} + +## Markers +- **Total**: {len(tl.markers)} +- **Chapters**: {len([m for m in tl.markers if m.marker_type == MarkerType.CHAPTER])} +""") + + +async def handle_list_clips(arguments: dict) -> Sequence[TextContent]: + project, tl = _require_timeline(arguments["filepath"]) + limit = arguments.get("limit") + clips = tl.clips[:limit] if limit else tl.clips + result = f"# Clips in {tl.name}\n\n| # | Name | Start | Duration | Keywords |\n|---|------|-------|----------|----------|\n" + for i, c in enumerate(clips, 1): + kws = ", ".join(k.value for k in c.keywords) if c.keywords else "-" + result += f"| {i} | {c.name} | {format_timecode(c.start)} | {format_duration(c.duration_seconds)} | {kws} |\n" + return _text_result(result) + + +async def handle_list_markers(arguments: dict) -> Sequence[TextContent]: + project, tl = _require_timeline(arguments["filepath"]) + markers = list(tl.markers) + for clip in tl.clips: + markers.extend(clip.markers) + marker_type = arguments.get("marker_type", "all") + if marker_type != "all": + markers = [m for m in markers if m.marker_type == MarkerType.from_string(marker_type)] + markers.sort(key=lambda m: m.start.frames) + fmt = arguments.get("format", "detailed") + if fmt == "youtube": + result = "# YouTube Chapters\n\n" + "\n".join(f"{m.to_youtube_timestamp()} {m.name}" for m in markers) + elif fmt == "simple": + result = "\n".join(f"{format_timecode(m.start)} - {m.name}" for m in markers) + else: + result = f"# Markers ({len(markers)})\n\n| TC | Name | Type |\n|---|------|------|\n" + result += "\n".join(f"| {format_timecode(m.start)} | {m.name} | {m.marker_type.value} |" for m in markers) + return _text_result(result) + + +async def handle_list_keywords(arguments: dict) -> Sequence[TextContent]: + project, tl = _require_timeline(arguments["filepath"]) + keywords = {} + for clip in tl.clips: + for kw in clip.keywords: + keywords.setdefault(kw.value, []).append(clip.name) + if not keywords: + return _text_result("No keywords found") + result = f"# Keywords ({len(keywords)})\n\n" + for kw, clips in sorted(keywords.items()): + result += f"**{kw}** ({len(clips)} clips)\n" + return _text_result(result) + + +async def handle_list_library_clips(arguments: dict) -> Sequence[TextContent]: + filepath = _validate_filepath(arguments["filepath"], ('.fcpxml', '.fcpxmld')) + parser = FCPXMLParser() + parser.parse_file(filepath) + keywords = arguments.get("keywords") + library_clips = parser.get_library_clips(keywords=keywords) + limit = arguments.get("limit") + if limit: + library_clips = library_clips[:limit] + if not library_clips: + return _text_result("No library clips found") + result = f"# Library Clips ({len(library_clips)} available)\n\n" + result += "| ID | Name | Duration | Has Video | Has Audio |\n" + result += "|----|------|----------|-----------|----------|\n" + for c in library_clips: + result += f"| {c['asset_id']} | {c['name']} | {format_duration(c['duration_seconds'])} | {'Y' if c['has_video'] else 'N'} | {'Y' if c['has_audio'] else 'N'} |\n" + result += "\n*Use `insert_clip` to add these to your timeline.*" + return _text_result(result) + + +async def handle_list_connected_clips(arguments: dict) -> Sequence[TextContent]: + project, tl = _require_timeline(arguments["filepath"]) + + lane_filter = arguments.get("lane") + clips = tl.connected_clips + if lane_filter is not None: + clips = [c for c in clips if c.lane == lane_filter] + + if not clips: + return _text_result("No connected clips found in timeline.") + + result = f"# Connected Clips in {tl.name}\n\n**Total**: {len(clips)}\n\n" + result += "| # | Name | Lane | Type | Duration | Parent | Role |\n" + result += "|---|------|------|------|----------|--------|------|\n" + for i, c in enumerate(clips, 1): + result += ( + f"| {i} | {c.name} | {c.lane} | {c.clip_type} | " + f"{format_duration(c.duration_seconds)} | {c.parent_clip_name} | " + f"{c.role or '-'} |\n" + ) + return _text_result(result) + + +async def handle_list_compound_clips(arguments: dict) -> Sequence[TextContent]: + project, tl = _require_timeline(arguments["filepath"]) + + if not tl.compound_clips: + return _text_result("No compound clips found in timeline.") + + result = f"# Compound Clips in {tl.name}\n\n" + for i, cc in enumerate(tl.compound_clips, 1): + result += f"### {i}. {cc.name}\n" + result += f"- **Ref ID**: {cc.ref_id}\n" + result += f"- **Duration**: {format_duration(cc.duration_seconds)}\n" + result += f"- **Clips inside**: {len(cc.clips)}\n\n" + return _text_result(result) + + +async def handle_list_roles(arguments: dict) -> Sequence[TextContent]: + project, tl = _require_timeline(arguments["filepath"]) + + audio_roles: dict[str, int] = {} + video_roles: dict[str, int] = {} + + for clip in tl.clips: + if clip.audio_role: + audio_roles[clip.audio_role] = audio_roles.get(clip.audio_role, 0) + 1 + if clip.video_role: + video_roles[clip.video_role] = video_roles.get(clip.video_role, 0) + 1 + + for cc in tl.connected_clips: + if cc.role: + # Determine type from clip_type + if cc.clip_type in ('audio', 'audio-clip'): + audio_roles[cc.role] = audio_roles.get(cc.role, 0) + 1 + else: + video_roles[cc.role] = video_roles.get(cc.role, 0) + 1 + + result = f"# Roles in {tl.name}\n\n" + if audio_roles: + result += "## Audio Roles\n\n| Role | Clips |\n|------|-------|\n" + for role, count in sorted(audio_roles.items()): + result += f"| {role} | {count} |\n" + else: + result += "## Audio Roles\n\nNo audio roles assigned.\n" + + result += "\n" + if video_roles: + result += "## Video Roles\n\n| Role | Clips |\n|------|-------|\n" + for role, count in sorted(video_roles.items()): + result += f"| {role} | {count} |\n" + else: + result += "## Video Roles\n\nNo video roles assigned.\n" + + return _text_result(result) + + +async def handle_diff_timelines(arguments: dict) -> Sequence[TextContent]: + filepath_a = _validate_filepath(arguments["filepath_a"], ('.fcpxml', '.fcpxmld')) + filepath_b = _validate_filepath(arguments["filepath_b"], ('.fcpxml', '.fcpxmld')) + + diff = compare_timelines(filepath_a, filepath_b) + + if not diff.has_changes: + return _text_result(( + f"# Timeline Diff: No Changes\n\n" + f"**{diff.timeline_a_name}** vs **{diff.timeline_b_name}** are identical." + )) + + result = ( + f"# Timeline Diff\n\n" + f"**Baseline**: {diff.timeline_a_name}\n" + f"**Comparison**: {diff.timeline_b_name}\n" + f"**Total changes**: {diff.total_changes}\n\n" + ) + + if diff.format_changes: + result += "## Format Changes\n\n" + for change in diff.format_changes: + result += f"- {change}\n" + result += "\n" + + clip_changes = [d for d in diff.clip_diffs if d.action != "unchanged"] + if clip_changes: + result += "## Clip Changes\n\n| Action | Clip | Details |\n|--------|------|--------|\n" + for d in clip_changes: + result += f"| {d.action.upper()} | {d.clip_name} | {d.details} |\n" + result += "\n" + + if diff.marker_diffs: + result += "## Marker Changes\n\n| Action | Marker | Details |\n|--------|--------|--------|\n" + for d in diff.marker_diffs: + result += f"| {d.action.upper()} | {d.marker_name} | {d.details} |\n" + result += "\n" + + if diff.transition_diffs: + result += "## Transition Changes\n\n" + for change in diff.transition_diffs: + result += f"- {change}\n" + + return _text_result(result) + + +async def handle_list_effects(arguments: dict) -> Sequence[TextContent]: + effects = list_effects() + lines = ["# Available FCP Transition Effects\n"] + for eff in effects: + lines.append(f"- **{eff['slug']}**: {eff['name']} (`{eff['uuid']}`)") + return _text_result("\n".join(lines)) + + +HANDLERS = { + "list_projects": handle_list_projects, + "analyze_timeline": handle_analyze_timeline, + "list_clips": handle_list_clips, + "list_markers": handle_list_markers, + "list_keywords": handle_list_keywords, + "list_library_clips": handle_list_library_clips, + "list_connected_clips": handle_list_connected_clips, + "list_compound_clips": handle_list_compound_clips, + "list_roles": handle_list_roles, + "diff_timelines": handle_diff_timelines, + "list_effects": handle_list_effects, +} diff --git a/code/server_tools/transcript.py b/code/server_tools/transcript.py new file mode 100644 index 0000000..7a20020 --- /dev/null +++ b/code/server_tools/transcript.py @@ -0,0 +1,235 @@ +"""Transcrição & edição por transcrição — tool schemas and handlers. + +Extracted from server.py; see Engine/docs/03_SERVER_TOOLS.md for the tool catalog. +""" + +from __future__ import annotations + +from pathlib import Path +from typing import Sequence + +from mcp.types import TextContent, Tool + +from fcpxml.media_intel import media_src_to_path +from fcpxml.transcribe import ( + DEFAULT_FILLERS, + find_filler_spans, + find_phrase_spans, + merge_ranges, + segments_to_srt, +) +from server_tools._shared import ( + _TRANSCRIBE_INSTALL_HINT, + TRANSCRIBE_MAX_MEDIA, + _cut_transcript_spans, + _load_or_transcribe, + _markdown_table, + _require_timeline, + _setup_modifier, + _text_result, + _transcript_cut_report, + _validate_output_path, + format_duration, +) + +TOOLS = [ + Tool( + name="transcribe_media", + description="Transcribe each clip's source media locally with word-level timestamps (faster-whisper). Writes a _transcript.json next to each media file (reused by edit_by_transcript / remove_filler_words so media is only transcribed once) and optionally an SRT for captions. Requires the optional [transcribe] extra; degrades to an install hint without it.", + inputSchema={ + "type": "object", + "properties": { + "filepath": {"type": "string", "description": "Path to FCPXML file"}, + "clip_name": {"type": "string", "description": "Only transcribe the clip with this name"}, + "model": {"type": "string", "default": "base", "description": "Whisper model size: tiny, base, small, medium, large-v3 (default base; larger = slower + more accurate)"}, + "language": {"type": "string", "description": "ISO language code hint (e.g. 'en'); auto-detected if omitted"}, + "write_srt": {"type": "boolean", "default": False, "description": "Also write a _transcript.srt next to each media file (plugs into import_srt_markers)"}, + }, + "required": ["filepath"] + } + ), + Tool( + name="edit_by_transcript", + description="Text-based editing: cut timeline content by what was SAID. mode=remove cuts every occurrence of the given phrases (with ripple); mode=keep_only keeps only the matched phrases and cuts everything else in each matched clip (clips with no matches are left untouched). Uses each media file's _transcript.json (auto-transcribes if missing). Non-destructive: writes a _transcript_edit copy.", + inputSchema={ + "type": "object", + "properties": { + "filepath": {"type": "string", "description": "Path to FCPXML file"}, + "phrases": {"type": "array", "items": {"type": "string"}, "description": "Spoken phrases to match (case/punctuation-insensitive)"}, + "mode": {"type": "string", "enum": ["remove", "keep_only"], "default": "remove", "description": "remove=cut matches out; keep_only=keep only matches"}, + "clip_name": {"type": "string", "description": "Only edit the clip with this name"}, + "model": {"type": "string", "default": "base", "description": "Whisper model size if transcription is needed"}, + "padding": {"type": "number", "default": 0.0, "description": "Seconds to widen each cut on both sides (0-2, default 0)"}, + "output_path": {"type": "string", "description": "Output path (default: adds _transcript_edit suffix)"}, + }, + "required": ["filepath", "phrases"] + } + ), + Tool( + name="remove_filler_words", + description="Cut filler words (um, uh, erm...) out of the timeline with ripple, using word-level transcripts of the real source audio. Conservative default filler list — words like 'like' and 'so' are only cut if you pass them explicitly. Uses each media file's _transcript.json (auto-transcribes if missing). Non-destructive: writes a _defillered copy.", + inputSchema={ + "type": "object", + "properties": { + "filepath": {"type": "string", "description": "Path to FCPXML file"}, + "fillers": {"type": "array", "items": {"type": "string"}, "description": "Filler words/phrases to cut (default: um, uh, uhh, umm, erm, ehm, mmm, hmm, mhm)"}, + "clip_name": {"type": "string", "description": "Only clean the clip with this name"}, + "model": {"type": "string", "default": "base", "description": "Whisper model size if transcription is needed"}, + "padding": {"type": "number", "default": 0.02, "description": "Seconds to widen each cut on both sides (0-2, default 0.02)"}, + "output_path": {"type": "string", "description": "Output path (default: adds _defillered suffix)"}, + }, + "required": ["filepath"] + } + ), +] + + +async def handle_transcribe_media(arguments: dict) -> Sequence[TextContent]: + model = arguments.get("model", "base") + language = arguments.get("language") + output_dir = arguments.get("output_dir") + write_srt = bool(arguments.get("write_srt", False)) + _, tl = _require_timeline(arguments["filepath"]) + clip_filter = arguments.get("clip_name") + + done: dict[str, dict | None] = {} + skipped: list[tuple[str, str]] = [] + rows: list[list[str]] = [] + srt_paths: list[str] = [] + for clip in tl.clips: + if clip_filter and clip.name != clip_filter: + continue + media_path = media_src_to_path(clip.media_path or "") + if not media_path or not Path(media_path).is_file(): + skipped.append((clip.name, "media file missing")) + continue + if media_path in done: + continue + if len(done) >= TRANSCRIBE_MAX_MEDIA: + skipped.append((clip.name, f"transcription cap reached ({TRANSCRIBE_MAX_MEDIA} media files)")) + continue + data, reason = _load_or_transcribe(media_path, model, language, output_dir) + done[media_path] = data + if data is None: + skipped.append((clip.name, reason)) + continue + if write_srt and data.get("segments"): + srt_name = Path(media_path).stem + "_transcript.srt" + srt_anchor = str(Path(output_dir).expanduser()) if output_dir else str(Path(media_path).parent) + srt_path = _validate_output_path( + str(Path(srt_anchor) / srt_name), + anchor_dir=srt_anchor, + ) + with open(srt_path, "w") as f: + f.write(segments_to_srt(data["segments"])) + srt_paths.append(srt_path) + preview = data.get("text", "")[:160] + rows.append([ + Path(media_path).name, + data.get("language", "?"), + str(len(data.get("words", []))), + format_duration(float(data.get("duration", 0.0))), + preview + ("…" if len(data.get("text", "")) > 160 else ""), + ]) + + result = f"""# Media Transcription (local Whisper) + +## Summary +- **Model**: {model} +- **Media Files Transcribed**: {len(rows)} +""" + if rows: + result += "\n## Transcripts (saved as _transcript.json next to each media file)\n" + result += _markdown_table( + ["Media", "Language", "Words", "Duration", "Preview"], rows + ) + "\n" + result += ( + "\n*Next: `edit_by_transcript` to cut by what was said, or " + "`remove_filler_words` to clean ums/uhs. Transcripts are cached — " + "media is only transcribed once.*" + ) + if srt_paths: + result += "\n\n## SRT Files\n" + "\n".join(f"- {p}" for p in srt_paths) + if skipped: + result += "\n## Skipped Clips\n" + _markdown_table( + ["Clip", "Reason"], [[name, reason] for name, reason in skipped] + ) + "\n" + if not rows and any("faster-whisper" in reason for _, reason in skipped): + result += _TRANSCRIBE_INSTALL_HINT + return _text_result(result) + + +async def handle_edit_by_transcript(arguments: dict) -> Sequence[TextContent]: + phrases = arguments.get("phrases") or [] + if not isinstance(phrases, list) or not all(isinstance(p, str) for p in phrases): + raise ValueError("phrases must be a list of strings") + phrases = [p for p in phrases if p.strip()] + if not phrases: + raise ValueError("phrases must contain at least one non-empty string") + mode = arguments.get("mode", "remove") + if mode not in ("remove", "keep_only"): + raise ValueError(f"mode must be 'remove' or 'keep_only', got {mode!r}") + padding = float(arguments.get("padding", 0.0)) + if not (0 <= padding <= 2): + raise ValueError(f"padding must be between 0 and 2 seconds, got {padding}") + model = arguments.get("model", "base") + language = arguments.get("language") + output_dir = arguments.get("output_dir") + + filepath, output_path, modifier = _setup_modifier(arguments, "_transcript_edit") + + def spans_fn(words): + return merge_ranges( + [span for phrase in phrases for span in find_phrase_spans(words, phrase)] + ) + + cuts_made, skipped = _cut_transcript_spans( + modifier, arguments.get("clip_name"), model, language, padding, + spans_fn, keep_only=(mode == "keep_only"), output_dir=output_dir, + ) + if cuts_made: + modifier.save(output_path) + verb = "kept only" if mode == "keep_only" else "removed" + return _transcript_cut_report( + "Transcript Edit", + [f"- **Mode**: {mode} ({verb} the matched phrases)", + f"- **Phrases**: {', '.join(repr(p) for p in phrases)}", + f"- **Padding**: {padding}s"], + cuts_made, skipped, output_path, + "*Transcripts are cached as _transcript.json. Original file untouched.*", + ) + + +async def handle_remove_filler_words(arguments: dict) -> Sequence[TextContent]: + fillers = arguments.get("fillers") or list(DEFAULT_FILLERS) + if not isinstance(fillers, list) or not all(isinstance(f, str) for f in fillers): + raise ValueError("fillers must be a list of strings") + padding = float(arguments.get("padding", 0.02)) + if not (0 <= padding <= 2): + raise ValueError(f"padding must be between 0 and 2 seconds, got {padding}") + model = arguments.get("model", "base") + language = arguments.get("language") + output_dir = arguments.get("output_dir") + + filepath, output_path, modifier = _setup_modifier(arguments, "_defillered") + + cuts_made, skipped = _cut_transcript_spans( + modifier, arguments.get("clip_name"), model, language, padding, + lambda words: merge_ranges(find_filler_spans(words, fillers)), + output_dir=output_dir, + ) + if cuts_made: + modifier.save(output_path) + return _transcript_cut_report( + "Filler Word Removal", + [f"- **Fillers**: {', '.join(fillers)}", f"- **Padding**: {padding}s"], + cuts_made, skipped, output_path, + "*Transcripts are cached as _transcript.json. Original file untouched.*", + ) + + +HANDLERS = { + "transcribe_media": handle_transcribe_media, + "edit_by_transcript": handle_edit_by_transcript, + "remove_filler_words": handle_remove_filler_words, +} diff --git a/code/server_tools/voice.py b/code/server_tools/voice.py new file mode 100644 index 0000000..8d63d1f --- /dev/null +++ b/code/server_tools/voice.py @@ -0,0 +1,751 @@ +"""Voz (análise → decisão → aplicação) — tool schemas and handlers. + +Extracted from server.py; see Engine/docs/03_SERVER_TOOLS.md for the tool catalog. +""" + +from __future__ import annotations + +import json +from pathlib import Path +from typing import List, Optional, Sequence, Tuple + +from mcp.types import TextContent, Tool + +from fcpxml.diarize import assign_speakers, build_speakers, diarization_capability, diarize +from fcpxml.emphasis import EmphasisWeights +from fcpxml.media_intel import media_src_to_path +from fcpxml.model_manager import ( + load_hf_token, + load_num_speakers, + load_voice_analysis_config, + save_voice_analysis_config, +) +from fcpxml.models import TimeValue +from fcpxml.voice_actions import parse_actions, resolve_actions, speaker_cut_actions +from fcpxml.voice_features import extract_energy, extract_pitch, features_capability +from fcpxml.voice_timeline import ( + build_voice_timeline, + enrich_words, + load_voice_timeline, + restrict_to_kept, + save_voice_timeline, + select_peaks, + suggest_zoom_windows, + voice_timeline_path, +) +from fcpxml.writer import FCPXMLModifier +from server_tools._shared import ( + _DIARIZATION_INSTALL_HINT, + _FEATURES_INSTALL_HINT, + _TRANSCRIBE_INSTALL_HINT, + AUDIO_MEDIA_EXTENSIONS, + MAX_MEDIA_FILE_SIZE, + _apply_placed_action, + _load_or_transcribe, + _markdown_table, + _setup_modifier, + _speaker_table, + _text_result, + _validate_filepath, + _validate_output_path, + _voice_analysis_config_text, + format_duration, +) + +TOOLS = [ + Tool( + name="diarize_media", + description="Identify WHO is speaking (speaker diarization) in an audio/video file using pyannote.audio, and assign SPEAKER_NN labels to each word/segment of its cached transcript. Writes a _diarization.json next to the media file. Requires the optional [diarization] extra (pyannote.audio) and a HuggingFace token with access to pyannote/speaker-diarization-3.1 (set once via save_hf_token or the HF_TOKEN argument); degrades to an install/token hint without them. Transcribes first if no _transcript.json is cached yet.", + inputSchema={ + "type": "object", + "properties": { + "media_path": {"type": "string", "description": "Path to audio/video file (.wav, .mp3, .m4a, .aac, .aif, .flac, .mov, .mp4)"}, + "hf_token": {"type": "string", "description": "HuggingFace token with pyannote/speaker-diarization-3.1 access (default: the persisted token from save_hf_token, if any)"}, + "num_speakers": {"type": "string", "description": "Known number of speakers, if you know it (speeds up and improves accuracy). Leave empty to auto-detect."}, + "model": {"type": "string", "default": "base", "description": "Whisper model size to use if transcription is needed (default base)"}, + "language": {"type": "string", "description": "ISO language code hint for transcription, if needed"}, + }, + "required": ["media_path"] + } + ), + Tool( + name="analyze_voice_features", + description="Analyze HOW a voice is speaking: pitch, energy, local speech rate, pauses, and a combined emphasis index (0-1) per transcribed word, using the persisted Voice Analysis settings (energy threshold, emphasis weights, emphasis cutoff — see save_voice_analysis_config). Writes a _voice_features.json next to the media file. Requires the optional [intelligence] extra (librosa); degrades to an install hint without it. Transcribes first if no _transcript.json is cached yet.", + inputSchema={ + "type": "object", + "properties": { + "media_path": {"type": "string", "description": "Path to audio/video file (.wav, .mp3, .m4a, .aac, .aif, .flac, .mov, .mp4)"}, + "model": {"type": "string", "default": "base", "description": "Whisper model size to use if transcription is needed (default base)"}, + "language": {"type": "string", "description": "ISO language code hint for transcription, if needed"}, + }, + "required": ["media_path"] + } + ), + Tool( + name="build_voice_timeline", + description="Build the consolidated voice timeline: WHAT was said (transcript), WHO said it (diarization), and HOW it was said (pitch/energy/rate/pauses -> emphasis index), merged into one AI-readable JSON written next to the media as _voice_timeline.json. This is the source of truth for automated editing — layered as summary -> segments -> words, with all acoustic values normalized 0-1 and documented inline, so a model can read the narrative shape and decide how to direct the edit. Every layer degrades independently: no librosa means acoustic values are 0, no HuggingFace token means a single default speaker; the document shape never changes.", + inputSchema={ + "type": "object", + "properties": { + "media_path": {"type": "string", "description": "Path to audio/video file (.wav, .mp3, .m4a, .aac, .aif, .flac, .mov, .mp4)"}, + "model": {"type": "string", "default": "base", "description": "Whisper model size to use if transcription is needed (default base)"}, + "language": {"type": "string", "description": "ISO language code hint for transcription, if needed"}, + "hf_token": {"type": "string", "description": "HuggingFace token for speaker diarization (default: the persisted token; omit to skip diarization)"}, + "num_speakers": {"type": "string", "description": "Known number of speakers, if any (default: the persisted setting, else auto-detect)"}, + "output_dir": {"type": "string", "description": "Folder to write _voice_timeline.json into (default: next to the media file)"}, + }, + "required": ["media_path"] + } + ), + Tool( + name="remove_speakers", + description="Cut everything one or more speakers say out of the timeline — the standard cleanup on an interview shoot, where the interviewer or a crew member talks during the take and only the subject should survive. Reads the media's _voice_timeline.json (build it first with build_voice_timeline, which must have run with diarization so speakers are separated). Call without speaker_ids to just LIST who was detected, with speaking share and sample lines, so you can tell who is who before cutting anything. Non-destructive: writes a _voice_edit copy.", + inputSchema={ + "type": "object", + "properties": { + "filepath": {"type": "string", "description": "Path to FCPXML file"}, + "media_path": {"type": "string", "description": "Media whose _voice_timeline.json holds the speakers (default: the timeline's first clip media)"}, + "speaker_ids": {"type": "array", "items": {"type": "string"}, "description": "Speakers to REMOVE (e.g. [\"SPEAKER_01\"]). Omit to only list the detected speakers without editing."}, + "padding": {"type": "number", "default": 0.15, "description": "Seconds trimmed inside each cut so the kept speaker's first syllable is never clipped (default 0.15)"}, + "output_path": {"type": "string", "description": "Output path (default: adds _voice_edit suffix)"}, + }, + "required": ["filepath"] + } + ), + Tool( + name="refine_voice_timeline", + description="Re-analyze a voice timeline over only the material that survives a set of cuts, then propose punch-in windows over it. Emphasis is RELATIVE — energy is scored against the loudest word of the recording — so once the loudest moment is cut (a laugh, an aside to the crew), every remaining score is measured against something the viewer will never see and the ranking points at the wrong words. Run this after deciding cuts and before deciding zooms. Cheap: it re-normalizes the already-measured numbers, never re-reads the audio. Times stay in ORIGINAL source seconds, so the result feeds straight back into apply_voice_actions.", + inputSchema={ + "type": "object", + "properties": { + "media_path": {"type": "string", "description": "Media whose _voice_timeline.json will be refined (build it first with build_voice_timeline)"}, + "cuts": { + "type": "array", + "description": "The ranges being REMOVED, in original source seconds. Pass the cut actions you already decided; anything overlapping them is excluded from the re-analysis.", + "items": { + "type": "object", + "properties": { + "start": {"type": "number", "description": "Start in original source seconds"}, + "end": {"type": "number", "description": "End in original source seconds"}, + }, + "required": ["start", "end"], + }, + }, + "min_gap": {"type": "number", "default": 8.0, "description": "Minimum seconds between two proposed zooms — effects stacked close together read as nervous editing (default 8.0)"}, + "max_zooms": {"type": "integer", "description": "Cap on how many zoom candidates to return (default: no cap — cut the list by rhythm yourself)"}, + "save": {"type": "boolean", "default": False, "description": "Also write the refined timeline as _voice_timeline_refined.json next to the media"}, + "output_dir": {"type": "string", "description": "Folder holding _voice_timeline.json (default: next to the media file)"}, + }, + "required": ["media_path", "cuts"] + } + ), + Tool( + name="apply_voice_actions", + description="Apply a list of editing decisions (from the rules engine, or from a model that read the _voice_timeline.json) to a timeline, producing FCPXML. Actions are validated first and reported per row, so one malformed decision never discards the edit. All action times are in ORIGINAL source seconds: cuts are resolved first and every other action is moved onto its post-cut position automatically, so decisions never land on the wrong frame. Actions pointing into removed material are dropped and reported, not silently slid. Non-destructive: writes a _voice_edit copy.", + inputSchema={ + "type": "object", + "properties": { + "filepath": {"type": "string", "description": "Path to FCPXML file"}, + "actions": { + "type": "array", + "description": "The decision list. Each item: {kind, start, end, params, reason, speaker}. kind is cut | zoom | text | marker. Times in original source seconds. zoom takes params.scale (1.0-3.0, default 1.3); text requires params.content.", + "items": { + "type": "object", + "properties": { + "kind": {"type": "string", "enum": ["cut", "zoom", "text", "marker"]}, + "start": {"type": "number", "description": "Start in original source seconds"}, + "end": {"type": "number", "description": "End in original source seconds"}, + "params": {"type": "object", "description": "kind-specific: {scale} for zoom, {content} for text"}, + "reason": {"type": "string", "description": "Why this decision was made — kept for review"}, + "speaker": {"type": "string", "description": "Speaker id this decision relates to, if any"}, + }, + "required": ["kind", "start", "end"], + }, + }, + "output_path": {"type": "string", "description": "Output path (default: adds _voice_edit suffix)"}, + }, + "required": ["filepath", "actions"] + } + ), + Tool( + name="get_voice_analysis_config", + description="Read the persisted Voice Analysis settings: energy threshold, emphasis-index weights (energy/pitch_variation/rate_variation/pause_before/duration), emphasis cutoff for punch-in candidates, and emotion detection toggle/sensitivity. Shared with the MacApp settings screen (~/.fcp-mcp-server/config.json).", + inputSchema={"type": "object", "properties": {}} + ), + Tool( + name="save_voice_analysis_config", + description="Persist Voice Analysis settings. Only the fields you pass are changed; omitted fields keep their current value. emphasis_weights don't need to sum to 1 (normalized internally). Shared with the MacApp settings screen (~/.fcp-mcp-server/config.json).", + inputSchema={ + "type": "object", + "properties": { + "energy_threshold": {"type": "number", "description": "0-1, how loud (normalized RMS) counts as 'high energy' (default 0.5)"}, + "emphasis_weights": { + "type": "object", + "description": "Any subset of {energy, pitch_variation, rate_variation, pause_before, duration} weights for the emphasis index", + "properties": { + "energy": {"type": "number"}, + "pitch_variation": {"type": "number"}, + "rate_variation": {"type": "number"}, + "pause_before": {"type": "number"}, + "duration": {"type": "number"}, + }, + }, + "peak_percentile": {"type": "number", "description": "Fraction of words selected as peaks, 0-1 (default 0.02 = top 2%). Selection is relative because the emphasis index's real range depends on the material — measured on a real interview it never passed 0.55."}, + "emphasis_floor": {"type": "number", "description": "0-1 minimum emphasis for a peak, guarding genuinely flat audio (default 0.25)"}, + "emotion_enabled": {"type": "boolean", "description": "Whether emotion detection runs as part of voice analysis (default false)"}, + "emotion_sensitivity": {"type": "number", "description": "0-1 confidence threshold to accept an emotion label (default 0.5)"}, + }, + } + ), +] + + +async def handle_diarize_media(arguments: dict) -> Sequence[TextContent]: + media_path = _validate_filepath( + arguments["media_path"], AUDIO_MEDIA_EXTENSIONS, max_size=MAX_MEDIA_FILE_SIZE + ) + token = str(arguments.get("hf_token") or "").strip() or load_hf_token() or None + num_speakers = str(arguments.get("num_speakers") or "").strip() + model = arguments.get("model", "base") + language = arguments.get("language") + + ok, message = diarization_capability(token) + if not ok: + return _text_result(f"# Speaker Diarization\n\n{message}{_DIARIZATION_INSTALL_HINT}") + + transcript, reason = _load_or_transcribe(media_path, model, language) + if transcript is None: + return _text_result( + f"# Speaker Diarization\n\nCould not obtain a transcript to diarize " + f"({reason}).{_TRANSCRIBE_INSTALL_HINT}" + ) + + tracks = diarize(media_path, token, num_speakers) + if tracks is None: + return _text_result( + "# Speaker Diarization\n\nDiarization failed — check the HuggingFace " + "token has accepted the pyannote/speaker-diarization-3.1 model terms, " + "and that the media file is readable." + ) + + segments, words = assign_speakers( + transcript.get("segments", []), transcript.get("words", []), tracks + ) + speakers = build_speakers(segments) + + diarization_data = { + "source": Path(media_path).name, + "speakers": speakers, + "segments": segments, + "words": words, + } + json_path = _validate_output_path( + str(Path(media_path).with_name(Path(media_path).stem + "_diarization.json")), + anchor_dir=str(Path(media_path).parent), + ) + with open(json_path, "w") as f: + json.dump(diarization_data, f, indent=2) + + result_text = f"""# Speaker Diarization + +## Summary +- **Source**: {Path(media_path).name} +- **Speakers Detected**: {len(speakers)} +- **Segments**: {len(segments)} +- **Diarization JSON**: {json_path} + +## Speakers +""" + result_text += _markdown_table( + ["ID", "Name"], [[s["id"], s["name"]] for s in speakers] + ) + "\n" + result_text += ( + "\n*Next: `build_voice_timeline` to cross this with acoustic features, " + "or use the segments/words directly for speaker-aware editing.*" + ) + return _text_result(result_text) + + +async def handle_analyze_voice_features(arguments: dict) -> Sequence[TextContent]: + media_path = _validate_filepath( + arguments["media_path"], AUDIO_MEDIA_EXTENSIONS, max_size=MAX_MEDIA_FILE_SIZE + ) + model = arguments.get("model", "base") + language = arguments.get("language") + + ok, message = features_capability() + if not ok: + return _text_result(f"# Voice Feature Analysis\n\n{message}{_FEATURES_INSTALL_HINT}") + + transcript, reason = _load_or_transcribe(media_path, model, language) + if transcript is None: + return _text_result( + f"# Voice Feature Analysis\n\nCould not obtain a transcript to analyze " + f"({reason}).{_TRANSCRIBE_INSTALL_HINT}" + ) + words = transcript.get("words", []) + if not words: + return _text_result("# Voice Feature Analysis\n\nNo words in transcript — nothing to analyze.") + + config = load_voice_analysis_config() + weights = EmphasisWeights.from_dict(config["emphasis_weights"]) + enriched = enrich_words( + words, extract_pitch(media_path), extract_energy(media_path), weights + ) + + energy_threshold = config["energy_threshold"] + high_energy_words = [w for w in enriched if w["energy_norm"] >= energy_threshold] + high_emphasis_words = select_peaks( + enriched, config["peak_percentile"], config["emphasis_floor"] + ) + + features_data = {"source": Path(media_path).name, "config": config, "words": enriched} + json_path = _validate_output_path( + str(Path(media_path).with_name(Path(media_path).stem + "_voice_features.json")), + anchor_dir=str(Path(media_path).parent), + ) + with open(json_path, "w") as f: + json.dump(features_data, f, indent=2) + + result_text = f"""# Voice Feature Analysis + +## Summary +- **Source**: {Path(media_path).name} +- **Words Analyzed**: {len(enriched)} +- **High-Energy Words** (>= {energy_threshold:.2f}): {len(high_energy_words)} +- **Peak Words** (top {config["peak_percentile"]:.1%}): {len(high_emphasis_words)} +- **Features JSON**: {json_path} + +## Top Emphasis Words +""" + top = sorted(enriched, key=lambda w: w["emphasis"], reverse=True)[:10] + result_text += _markdown_table( + ["Word", "Time", "Emphasis", "Energy", "Pitch Δ"], + [ + [ + w.get("word", ""), + f"{w.get('start', 0):.2f}s", + f"{w['emphasis']:.2f}", + f"{w['energy_norm']:.2f}", + f"{w['pitch_delta']:.2f}", + ] + for w in top + ], + ) + "\n" + result_text += ( + "\n*Thresholds and emphasis weights are configurable in Voice Analysis " + "settings (`save_voice_analysis_config`). Next: `diarize_media` to add " + "speaker labels.*" + ) + return _text_result(result_text) + + +async def handle_build_voice_timeline(arguments: dict) -> Sequence[TextContent]: + media_path = _validate_filepath( + arguments["media_path"], AUDIO_MEDIA_EXTENSIONS, max_size=MAX_MEDIA_FILE_SIZE + ) + model = arguments.get("model", "base") + language = arguments.get("language") + token = str(arguments.get("hf_token") or "").strip() or load_hf_token() or None + num_speakers = str(arguments.get("num_speakers") or "").strip() or load_num_speakers() + + transcript, reason = _load_or_transcribe(media_path, model, language) + if transcript is None: + return _text_result( + f"# Voice Timeline\n\nCould not obtain a transcript " + f"({reason}).{_TRANSCRIBE_INSTALL_HINT}" + ) + + config = load_voice_analysis_config() + timeline = build_voice_timeline( + media_path, + transcript, + hf_token=token, + num_speakers=num_speakers, + weights=EmphasisWeights.from_dict(config["emphasis_weights"]), + peak_percentile=config["peak_percentile"], + emphasis_floor=config["emphasis_floor"], + ) + + output_dir = arguments.get("output_dir") + json_path = Path(_validate_output_path( + str(voice_timeline_path(media_path, output_dir)), + anchor_dir=str(Path(output_dir) if output_dir else Path(media_path).parent), + )) + save_voice_timeline(timeline, json_path) + + summary = timeline["summary"] + layers = timeline["layers"] + result_text = f"""# Voice Timeline + +## Summary +- **Source**: {timeline["source"]} +- **Duration**: {format_duration(summary["duration"])} +- **Speakers**: {summary["speaker_count"]} +- **Segments**: {summary["segment_count"]} ({summary["word_count"]} words) +- **Average Emphasis**: {summary["avg_emphasis"]:.2f} +- **Peak Moments** ({summary["peak_selection"]}): {summary["peak_count"]} +- **Timeline JSON**: {json_path} + +## Analysis Layers +""" + result_text += _markdown_table( + ["Layer", "Status"], + [ + ["Transcript", "yes" if layers["transcript"] else "empty"], + [ + "Acoustics (pitch/energy)", + "yes" if layers["acoustics"] else "FAILED — every acoustic value is 0", + ], + ["Speakers", "yes" if layers["speakers"] else "not run — single default speaker"], + ], + ) + "\n" + + if summary["peak_moments"]: + result_text += "\n## Peak Moments\n" + result_text += _markdown_table( + ["Time", "Word", "Speaker", "Emphasis"], + [ + [ + f"{m['time']:.2f}s", + m["text"], + m["speaker"], + f"{m['emphasis']:.2f}", + ] + for m in summary["peak_moments"][:10] + ], + ) + "\n" + + result_text += ( + "\n*The JSON is layered summary -> segments -> words with normalized " + "0-1 values, ready to hand to a model for edit direction.*" + ) + return _text_result(result_text) + + +async def handle_remove_speakers(arguments: dict) -> Sequence[TextContent]: + filepath = _validate_filepath(arguments["filepath"], ('.fcpxml', '.fcpxmld')) + speaker_ids = arguments.get("speaker_ids") or [] + + media_path = arguments.get("media_path") + if not media_path: + modifier = FCPXMLModifier(filepath) + for _, clip_el in modifier._iter_spine_clips(): + src = modifier.resources.get(clip_el.get("ref", ""), {}).get("src", "") + candidate = media_src_to_path(src) + if candidate and Path(candidate).is_file(): + media_path = candidate + break + if not media_path: + return _text_result( + "# Speakers\n\nNo source media found for this timeline — pass `media_path` explicitly." + ) + + timeline = _read_voice_timeline(media_path, arguments.get("output_dir")) + if timeline is None: + return _text_result( + f"# Speakers\n\nNo voice timeline for `{Path(media_path).name}` yet.\n\n" + "Run `build_voice_timeline` on it first." + ) + + profiles = timeline.get("speakers", []) + header = f"# Speakers in {timeline.get('source', '')}\n\n" + _speaker_table(profiles) + "\n" + if len(profiles) < 2: + header += ( + "\n> Only one speaker is present. Either the recording really has one " + "voice, or diarization did not run — check the Models tab for the " + "HuggingFace token.\n" + ) + for p in profiles: + samples = [s for s in p.get("samples", []) if s] + if samples: + header += f"\n**{p['id']}** ({p.get('name', '')}) says things like:\n" + header += "".join(f"> {s}\n" for s in samples[:2]) + + if not speaker_ids: + return _text_result( + header + + "\n*Nothing was edited. Re-run with `speaker_ids` naming who to REMOVE — " + "typically the interviewer or crew, keeping the subject.*" + ) + + known = {p["id"] for p in profiles} + unknown = [s for s in speaker_ids if s not in known] + if unknown: + return _text_result( + header + f"\n**Unknown speaker(s): {', '.join(unknown)}** — nothing was edited." + ) + if set(speaker_ids) >= known: + return _text_result( + header + "\n**That would remove every speaker**, leaving nothing — nothing was edited." + ) + + actions = speaker_cut_actions( + timeline, speaker_ids, padding=float(arguments.get("padding", 0.15)) + ) + if not actions: + return _text_result(header + "\n No speech found for those speakers — nothing was edited.") + + result = await handle_apply_voice_actions({ + **arguments, + "actions": [a.as_dict() for a in actions], + }) + removed = sum(a.duration for a in actions) + return _text_result( + header + + f"\n## Removed\n- **Speakers cut**: {', '.join(speaker_ids)}\n" + + f"- **Speech removed**: {format_duration(removed)} across {len(actions)} segments\n\n" + + result[0].text + ) + + +def _read_voice_timeline(media_path: str, output_dir: Optional[str] = None) -> Optional[dict]: + """Load the cached voice timeline, project folder first. + + ``build_voice_timeline`` writes to the chosen project folder when one is + set and beside the media otherwise, so a reader that only checks one of + the two reports "no voice timeline yet" for a file that exists. Checking + both also keeps timelines built before the project folder existed + readable. + """ + if output_dir: + timeline = load_voice_timeline(voice_timeline_path(media_path, output_dir)) + if timeline is not None: + return timeline + return load_voice_timeline(voice_timeline_path(media_path)) + + +async def handle_refine_voice_timeline(arguments: dict) -> Sequence[TextContent]: + media_path = _validate_filepath( + arguments["media_path"], AUDIO_MEDIA_EXTENSIONS, max_size=MAX_MEDIA_FILE_SIZE + ) + timeline = _read_voice_timeline(media_path, arguments.get("output_dir")) + if timeline is None: + return _text_result( + f"# Refined Voice Timeline\n\nNo voice timeline for " + f"`{Path(media_path).name}` yet.\n\nRun `build_voice_timeline` on it first." + ) + + cut_ranges: List[Tuple[float, float]] = [] + rejected: List[str] = [] + for i, raw in enumerate(arguments.get("cuts") or []): + try: + start, end = float(raw["start"]), float(raw["end"]) + except (TypeError, ValueError, KeyError): + rejected.append(f"cut #{i}: start/end must be numbers") + continue + if end <= start: + rejected.append(f"cut #{i}: end ({end}) must be after start ({start})") + continue + cut_ranges.append((start, end)) + + config = load_voice_analysis_config() + refined = restrict_to_kept( + timeline, + cut_ranges, + weights=EmphasisWeights.from_dict(config["emphasis_weights"]), + peak_percentile=config["peak_percentile"], + emphasis_floor=config["emphasis_floor"], + ) + before, after = timeline["summary"], refined["summary"] + zooms = suggest_zoom_windows( + refined, + min_gap=float(arguments.get("min_gap", 8.0)), + max_zooms=arguments.get("max_zooms"), + ) + + result = f"""# Refined Voice Timeline + +Re-normalized over the material that survives {len(cut_ranges)} cut(s). + +""" + result += _markdown_table( + ["Measure", "Raw recording", "Survivors only"], + [ + ["Duration", format_duration(before["duration"]), format_duration(after["duration"])], + ["Segments", str(before["segment_count"]), str(after["segment_count"])], + ["Words", str(before["word_count"]), str(after["word_count"])], + ["Average emphasis", f"{before['avg_emphasis']:.3f}", f"{after['avg_emphasis']:.3f}"], + ["Peak moments", str(before["peak_count"]), str(after["peak_count"])], + ], + ) + "\n" + + if after["word_count"] == 0: + result += "\n> The cuts removed every word — nothing left to analyze.\n" + + if after["peak_moments"]: + result += "\n## Peak Moments (re-ranked)\n" + result += _markdown_table( + ["Time", "Word", "Speaker", "Emphasis"], + [ + [f"{m['time']:.2f}s", m["text"], m["speaker"], f"{m['emphasis']:.2f}"] + for m in after["peak_moments"][:10] + ], + ) + "\n" + + if zooms: + result += "\n## Zoom Candidates\n" + result += _markdown_table( + ["Start", "End", "Word", "Emphasis", "Line"], + [ + [f"{z['start']:.2f}s", f"{z['end']:.2f}s", z["word"], + f"{z['emphasis']:.2f}", z["line"][:60]] + for z in zooms + ], + ) + "\n" + else: + result += "\n## Zoom Candidates\n\nNone — no content words survived the cuts.\n" + + if arguments.get("save"): + p = Path(media_path) + json_path = Path(_validate_output_path( + str(p.with_name(p.stem + "_voice_timeline_refined.json")), + anchor_dir=str(p.parent), + )) + save_voice_timeline(refined, json_path) + result += f"\n**Refined JSON**: {json_path}\n" + + if rejected: + result += "\n## Rejected cuts\n" + "\n".join(f"- {r}" for r in rejected) + "\n" + + result += ( + "\n*Candidates, not obligations — cut the list by rhythm. Times are in " + "original source seconds, ready for `apply_voice_actions`.*" + ) + return _text_result(result) + + +async def handle_apply_voice_actions(arguments: dict) -> Sequence[TextContent]: + raw_actions = arguments.get("actions") + if raw_actions is None: + return _text_result("# Voice Actions\n\nNo `actions` provided — nothing to apply.") + + actions, errors = parse_actions(raw_actions) + if not actions: + text = "# Voice Actions\n\nNo valid actions to apply." + if errors: + text += "\n\n## Rejected\n" + "\n".join(f"- {e}" for e in errors) + return _text_result(text) + + cut_ranges, placed, dropped = resolve_actions(actions) + filepath, output_path, modifier = _setup_modifier(arguments, "_voice_edit") + + def source_window(clip_el) -> tuple[float, float]: + """The span of source media a spine clip actually uses.""" + start = modifier.source_file_start(clip_el).to_seconds() + duration = modifier._parse_time(clip_el.get("duration", "0s")).to_seconds() + return start, start + duration + + applied: list[list[str]] = [] + unplaced: list[str] = [] + + # Placements go on before cuts: they are anchored in source coordinates, + # and cutting afterwards ripples the spine around them. + # Cuts go FIRST. Cutting splits a clip into pieces and rewrites the + # spine around them, which would duplicate a zoom onto every piece and + # lose markers entirely. Cutting first means placements land on final, + # stable clips — and `resolve_actions` already moved their times onto + # the post-cut timeline, so they still point at the same moment. + cuts_made = 0 + for _, clip_el in modifier._iter_spine_clips(): + clip_start, clip_end = source_window(clip_el) + to_frame = modifier.snap_seconds_to_frame + ranges = [ + (to_frame(max(s, clip_start) - clip_start), to_frame(min(e, clip_end) - clip_start)) + for s, e in cut_ranges + if min(e, clip_end) > max(s, clip_start) + ] + ranges = [(a, b) for a, b in ranges if b > a] + if ranges and modifier.cut_clip_ranges(clip_el, ranges) > TimeValue.zero(): + cuts_made += len(ranges) + + def timeline_window(clip_el) -> tuple[float, float]: + """Where a spine clip sits on the timeline, in seconds.""" + offset = modifier._parse_time(clip_el.get("offset", "0s")).to_seconds() + duration = modifier._parse_time(clip_el.get("duration", "0s")).to_seconds() + return offset, offset + duration + + for action in placed: + host = next( + ( + clip_el + for _, clip_el in modifier._iter_spine_clips() + if timeline_window(clip_el)[0] <= action.start < timeline_window(clip_el)[1] + ), + None, + ) + if host is None: + unplaced.append( + f"{action.kind} @ {action.start:.2f}s — falls outside the edited timeline" + ) + continue + try: + what = _apply_placed_action(modifier, host, action, timeline_window(host)[0]) + applied.append([f"{action.start:.2f}s", what, host.get("name", ""), action.reason]) + except (ValueError, KeyError) as exc: + unplaced.append(f"{action.kind} @ {action.start:.2f}s — {exc}") + + # Rename the project so it does not land in the library indistinguishable + # from the original. FCP imports by the name in the XML, so an untouched + # name puts two same-named projects in the same event — and the edit looks + # like it did nothing, because the original is what gets opened. + project_name = "" + project_el = modifier.root.find(".//project") + if project_el is not None: + project_name = f"{project_el.get('name', 'Projeto')} — corte por voz" + project_el.set("name", project_name) + + modifier.save(output_path) + + result = f"""# Voice Actions Applied + +## Summary +- **Actions received**: {len(actions)} +- **Placed** (zoom/text/marker): {len(applied)} +- **Cuts applied**: {cuts_made} +- **Project name**: {project_name or '(unchanged)'} +- **Saved to**: `{output_path}` + +""" + if applied: + result += "## Applied\n" + _markdown_table( + ["Time", "Action", "Clip", "Reason"], applied[:40] + ) + "\n" + if dropped: + result += "\n## Dropped (pointed into removed material)\n" + "\n".join( + f"- {a.kind} @ {a.start:.2f}s — {a.reason or 'no reason given'}" for a in dropped + ) + "\n" + if unplaced: + result += "\n## Not placed\n" + "\n".join(f"- {u}" for u in unplaced) + "\n" + if errors: + result += "\n## Rejected\n" + "\n".join(f"- {e}" for e in errors) + "\n" + result += "\n*Non-destructive: the original file is untouched.*" + return _text_result(result) + + +async def handle_get_voice_analysis_config(arguments: dict) -> Sequence[TextContent]: + return _text_result(_voice_analysis_config_text(load_voice_analysis_config())) + + +async def handle_save_voice_analysis_config(arguments: dict) -> Sequence[TextContent]: + config = save_voice_analysis_config( + energy_threshold=arguments.get("energy_threshold"), + emphasis_weights=arguments.get("emphasis_weights"), + peak_percentile=arguments.get("peak_percentile"), + emphasis_floor=arguments.get("emphasis_floor"), + emotion_enabled=arguments.get("emotion_enabled"), + emotion_sensitivity=arguments.get("emotion_sensitivity"), + ) + return _text_result(_voice_analysis_config_text(config)) + + +HANDLERS = { + "diarize_media": handle_diarize_media, + "analyze_voice_features": handle_analyze_voice_features, + "build_voice_timeline": handle_build_voice_timeline, + "remove_speakers": handle_remove_speakers, + "refine_voice_timeline": handle_refine_voice_timeline, + "apply_voice_actions": handle_apply_voice_actions, + "get_voice_analysis_config": handle_get_voice_analysis_config, + "save_voice_analysis_config": handle_save_voice_analysis_config, +} diff --git a/code/tests/test_collision.py b/code/tests/test_collision.py new file mode 100644 index 0000000..7b45b00 --- /dev/null +++ b/code/tests/test_collision.py @@ -0,0 +1,266 @@ +"""Tests for collision detection and subtitle-layout validation (Fase 1). + +Covers the pure functions in ``fcpxml.collision`` (spatial/temporal overlap, +area/ratio classification, distance, separation suggestion, box measurement) +and the integration through ``FCPXMLModifier.validate_subtitle_layout`` over a +generated document — the post-generation guarantee the layout engine only +provides by construction. +""" + +import shutil +import tempfile +from pathlib import Path + +import pytest + +from fcpxml.collision import ( + FONT_MISSING, + FONT_TOO_SMALL, + OUTSIDE_FRAME, + OUTSIDE_SAFE_AREA, + OVERLAP_PROBABLE, + OVERLAP_RENDER_TOLERANCE, + OVERLAP_SEVERE, + SPATIAL_COLLISION, + Box, + blocking, + classify_overlap, + distance_between, + measure_title_box, + overlap_metrics, + separation_suggestion, + temporal_overlap, + validate_titles, +) +from fcpxml.models import DynamicSubtitleConfig +from fcpxml.writer import FCPXMLModifier + +SAMPLE = Path(__file__).parent.parent / "examples" / "sample.fcpxml" + +WORD_MODE = DynamicSubtitleConfig(granularity="word") + +WORDS = [ + {"word": "Hello", "start": 0.0, "end": 0.4}, + {"word": "there", "start": 0.4, "end": 0.8}, + {"word": "friend", "start": 0.8, "end": 1.3}, +] + + +@pytest.fixture +def temp_fcpxml(): + with tempfile.NamedTemporaryFile(suffix=".fcpxml", delete=False) as f: + shutil.copy(SAMPLE, f.name) + yield f.name + Path(f.name).unlink(missing_ok=True) + + +def _title(text, x=0.0, y=0.0, *, start=0.0, end=1.0, font="Helvetica Neue", + face=None, font_size=100.0, kerning=0.0, group=0): + return { + "text": text, + "x": x, + "y": y, + "start": start, + "end": end, + "font": font, + "face": face, + "font_size": font_size, + "kerning": kerning, + "group": group, + } + + +class TestBoxOverlap: + def test_touching_boxes_do_not_overlap(self): + a = Box(0, 100, 0, 50) + b = Box(100, 200, 0, 50) # shares the right edge + assert not a.overlaps(b) + assert overlap_metrics(a, b)["overlap_area"] == 0 + + def test_overlap_by_one_pixel_detected(self): + a = Box(0, 100, 0, 50) + b = Box(99, 200, 0, 50) # one-pixel horizontal overlap + assert a.overlaps(b) + metrics = overlap_metrics(a, b) + assert metrics["overlap_width"] == 1 + assert metrics["overlap_area"] == 50 + + def test_vertical_only_overlap(self): + a = Box(0, 100, 0, 50) + b = Box(0, 100, 49, 100) # one-pixel vertical overlap + assert a.overlaps(b) + assert overlap_metrics(a, b)["overlap_height"] == 1 + + +class TestTemporalOverlap: + def test_adjacent_intervals_are_not_simultaneous(self): + # [0, 1) and [1, 2) share no instant. + assert not temporal_overlap(0.0, 1.0, 1.0, 2.0) + assert not temporal_overlap(1.0, 2.0, 0.0, 1.0) + + def test_interleaved_intervals_overlap(self): + assert temporal_overlap(0.0, 2.0, 1.0, 3.0) + + def test_contained_interval_overlaps(self): + assert temporal_overlap(0.0, 5.0, 1.0, 2.0) + + def test_boundary_survives_float_noise_from_the_writer(self): + """Found on real footage: the writer sets one block's title duration + to make its end land EXACTLY on the next block's start (same exact + FCPXML fraction), but end here is re-derived as start + duration — + two independently-rounded floats — which isn't bit-identical to the + other title's start read as a single division of that same + fraction. 7 of 8 "severe" collisions from one real clip were this, + off by ~1e-13s, far below any frame boundary.""" + start_a, duration_a = 74883609 / 24000, 43043 / 24000 + end_a = start_a + duration_a # float addition, like validate_titles does + start_b = 74926652 / 24000 # the exact same instant, read directly + assert end_a != start_b # the float noise is real + assert abs(end_a - start_b) < 1e-6 # ...and far below one frame + assert not temporal_overlap(start_a, end_a, start_b, start_b + 1.0) + + +class TestClassifyOverlap: + def test_zero_area_is_no_conflict(self): + assert classify_overlap( + {"overlap_area": 0, "overlap_height": 0, "overlap_ratio": 0} + ) == "none" + + def test_ratio_dominates_small_height(self): + # A sliver of 5px that covers most of a tiny box is still severe. + metrics = {"overlap_area": 50, "overlap_height": 5, "overlap_ratio": 0.9} + assert classify_overlap(metrics) == OVERLAP_SEVERE + + def test_height_buckets(self): + def metrics(height): + return {"overlap_area": 1, "overlap_height": height, + "overlap_ratio": 0.0} + + assert classify_overlap(metrics(5)) == OVERLAP_RENDER_TOLERANCE + assert classify_overlap(metrics(6)) == "warning" + assert classify_overlap(metrics(21)) == OVERLAP_PROBABLE + assert classify_overlap(metrics(51)) == OVERLAP_SEVERE + + +class TestDistanceAndSeparation: + def test_horizontal_distance(self): + a = Box(0, 100, 0, 50) + b = Box(150, 200, 0, 50) + d = distance_between(a, b) + assert d["distance_x"] == 50 + assert d["distance_y"] == 0 + assert d["distance"] == 50 + + def test_vertical_distance(self): + a = Box(0, 100, 0, 50) + b = Box(0, 100, 80, 130) + d = distance_between(a, b) + assert d["distance_y"] == 30 + assert d["distance_x"] == 0 + + def test_separation_picks_smallest_axis(self): + a = Box(0, 100, 0, 100) + b = Box(90, 190, 95, 195) # 10px horizontal, 5px vertical penetration + suggestion = separation_suggestion(a, b) + assert suggestion["axis"] == "vertical" + assert suggestion["minimum_movement"] == 5 + + def test_blocking_severities(self): + assert blocking(OVERLAP_SEVERE) + assert blocking(OVERLAP_PROBABLE) + assert not blocking("warning") + assert not blocking("none") + + +class TestValidateTitles: + def test_clean_layout_has_no_issues(self): + report = validate_titles( + [_title("um", x=-200), _title("dois", x=200)], + 2160, 3840, + ) + assert report["severity"] == "none" + assert report["issues"] == [] + + def test_simultaneous_collision_reported(self): + titles = [_title("A", x=0), _title("B", x=0)] + report = validate_titles(titles, 2160, 3840) + collisions = [ + i for i in report["issues"] if i["type"] == SPATIAL_COLLISION + ] + assert len(collisions) == 1 + assert collisions[0]["severity"] == OVERLAP_SEVERE + assert collisions[0]["suggested_correction"]["axis"] in ( + "vertical", "horizontal", + ) + + def test_collision_across_times_ignored(self): + titles = [ + _title("A", x=0, start=0.0, end=1.0), + _title("B", x=0, start=1.0, end=2.0), + ] + report = validate_titles(titles, 2160, 3840) + assert not any( + i["type"] == SPATIAL_COLLISION for i in report["issues"] + ) + + def test_outside_frame_is_error(self): + report = validate_titles([_title("fora", x=5000)], 2160, 3840) + assert any(i["type"] == OUTSIDE_FRAME for i in report["issues"]) + + def test_outside_safe_area_is_warning_not_frame(self): + # Near the right edge: inside the frame, outside the 5% safe area. + report = validate_titles( + [_title("a", x=1000, font_size=40)], 2160, 3840, + ) + assert any(i["type"] == OUTSIDE_SAFE_AREA for i in report["issues"]) + assert not any(i["type"] == OUTSIDE_FRAME for i in report["issues"]) + + def test_resolution_changes_safe_area(self): + # Same x is fine on a wide frame but out of the safe area on a narrow one. + narrow = validate_titles([_title("a", x=1000, font_size=40)], 1920, 1080) + assert any(i["type"] == OUTSIDE_FRAME for i in narrow["issues"]) + + def test_font_missing_reported(self): + report = validate_titles( + [_title("oi", font="Comic Sans MS")], 2160, 3840, + ) + assert any(i["type"] == FONT_MISSING for i in report["issues"]) + + def test_font_too_small_reported(self): + report = validate_titles( + [_title("oi", font_size=10)], 2160, 3840, min_font_size=20, + ) + assert any(i["type"] == FONT_TOO_SMALL for i in report["issues"]) + + def test_measure_title_box_uses_real_width(self): + box = measure_title_box("ii", 100, x=0, y=0, font="Helvetica Neue") + # "ii" is the narrowest glyph; a single "W" is much wider. + wide = measure_title_box("W", 100, x=0, y=0, font="Helvetica Neue") + assert wide.width > box.width + + +class TestIntegration: + def test_generated_layout_is_clean(self, temp_fcpxml): + modifier = FCPXMLModifier(temp_fcpxml) + modifier.generate_dynamic_subtitles("Interview_A", WORDS, WORD_MODE) + report = modifier.validate_subtitle_layout() + assert report["summary"]["title_count"] >= 3 + assert report["summary"]["spatial_collision"] == 0 + assert not blocking(report["severity"]) + + def test_hand_edited_position_is_caught(self, temp_fcpxml): + modifier = FCPXMLModifier(temp_fcpxml) + titles = modifier.generate_dynamic_subtitles("Interview_A", WORDS, WORD_MODE) + + def position(el): + for p in el.findall("param"): + if p.get("name") == "Position": + return p + return None + + p0 = position(titles[0]) + position(titles[1]).set("value", p0.get("value")) + + report = modifier.validate_subtitle_layout() + assert report["summary"]["spatial_collision"] >= 1 + assert blocking(report["severity"]) diff --git a/code/tests/test_diarize_media_tool.py b/code/tests/test_diarize_media_tool.py new file mode 100644 index 0000000..cf5c584 --- /dev/null +++ b/code/tests/test_diarize_media_tool.py @@ -0,0 +1,108 @@ +"""Tests for the diarize_media MCP tool (server.handle_diarize_media). + +Diarization itself (pyannote.audio) is monkeypatched so these tests run +without the optional [diarization] extra or a HuggingFace token — matching +the existing TestDetectBeatsHandler pattern in test_media_intel.py. +""" + +import json + +import pytest + + +def _write_tiny_wav(path: str, seconds: float = 1.0) -> None: + import struct + import wave + + n_frames = int(44100 * seconds) + with wave.open(path, "w") as f: + f.setnchannels(1) + f.setsampwidth(2) + f.setframerate(44100) + f.writeframes(struct.pack("<%dh" % n_frames, *([0] * n_frames))) + + +class TestDiarizeMediaHandler: + async def test_reports_when_pyannote_unavailable(self, tmp_path, monkeypatch): + import server_tools.voice as server_mod + from server import handle_diarize_media + + wav = tmp_path / "clip.wav" + _write_tiny_wav(str(wav)) + monkeypatch.setattr( + server_mod, "diarization_capability", lambda token: (False, "Diarização indisponível: componente pyannote.audio ausente.") + ) + result = await handle_diarize_media({"media_path": str(wav)}) + text = result[0].text + assert "indisponível" in text.lower() or "unavailable" in text.lower() + assert "diarization" in text.lower() + + async def test_rejects_disallowed_extension(self, tmp_path): + from server import handle_diarize_media + + bad = tmp_path / "clip.txt" + bad.write_text("not audio") + with pytest.raises(ValueError): + await handle_diarize_media({"media_path": str(bad)}) + + async def test_writes_diarization_json_and_reports(self, tmp_path, monkeypatch): + import server_tools._shared as _shared_mod + import server_tools.voice as server_mod + from server import handle_diarize_media + + wav = tmp_path / "clip.wav" + _write_tiny_wav(str(wav), seconds=2.0) + + fake_transcript = { + "language": "en", + "duration": 2.0, + "text": "hello world", + "segments": [ + {"text": "hello", "start": 0.0, "end": 1.0}, + {"text": "world", "start": 1.0, "end": 2.0}, + ], + "words": [ + {"word": "hello", "start": 0.0, "end": 0.5, "confidence": 0.9}, + {"word": "world", "start": 1.0, "end": 1.5, "confidence": 0.9}, + ], + } + monkeypatch.setattr(_shared_mod, "transcribe", lambda *a, **k: fake_transcript) + monkeypatch.setattr(server_mod, "diarization_capability", lambda token: (True, "ok")) + monkeypatch.setattr( + server_mod, + "diarize", + lambda path, token, num_speakers="": [(0.0, 1.0, "A"), (1.0, 2.0, "B")], + ) + + result = await handle_diarize_media({"media_path": str(wav), "hf_token": "fake-token"}) + text = result[0].text + assert "Speaker" in text or "speaker" in text.lower() + + json_path = tmp_path / "clip_diarization.json" + assert str(json_path) in text + data = json.loads(json_path.read_text()) + assert len(data["speakers"]) == 2 + assert data["words"][0]["speaker_id"] == "SPEAKER_00" + assert data["words"][1]["speaker_id"] == "SPEAKER_01" + + async def test_reports_when_diarization_fails(self, tmp_path, monkeypatch): + import server_tools._shared as _shared_mod + import server_tools.voice as server_mod + from server import handle_diarize_media + + wav = tmp_path / "clip.wav" + _write_tiny_wav(str(wav)) + + fake_transcript = { + "language": "en", + "duration": 1.0, + "text": "hi", + "segments": [{"text": "hi", "start": 0.0, "end": 1.0}], + "words": [{"word": "hi", "start": 0.0, "end": 0.5, "confidence": 0.9}], + } + monkeypatch.setattr(_shared_mod, "transcribe", lambda *a, **k: fake_transcript) + monkeypatch.setattr(server_mod, "diarization_capability", lambda token: (True, "ok")) + monkeypatch.setattr(server_mod, "diarize", lambda *a, **k: None) + + result = await handle_diarize_media({"media_path": str(wav), "hf_token": "fake-token"}) + assert "failed" in result[0].text.lower() diff --git a/code/tests/test_dynamic_subtitles.py b/code/tests/test_dynamic_subtitles.py index 7a23503..5fd6c45 100644 --- a/code/tests/test_dynamic_subtitles.py +++ b/code/tests/test_dynamic_subtitles.py @@ -17,12 +17,27 @@ from pathlib import Path import pytest -from fcpxml.models import DynamicSubtitleConfig, TimeValue, WordStyle +from fcpxml.models import DynamicSubtitleConfig, TimeValue, WordLook, WordStyle from fcpxml.parser import parse_fcpxml -from fcpxml.text_layout import POINT_SCALE, REFERENCE_CANVAS_HEIGHT, ink_extent +from fcpxml.text_layout import ( + POINT_SCALE, + REFERENCE_BLOCK_LINE_GAP, + REFERENCE_CANVAS_HEIGHT, + TEXT_TEMPLATE_FONT_SCALE, + LayoutBox, + compose_sentence, + ink_extent, +) from fcpxml.writer import FCPXMLModifier SAMPLE = Path(__file__).parent.parent / "examples" / "sample.fcpxml" +def font_points(style) -> float: + """The style's size back in CANVAS POINTS. + + The emitted fontSize lives in the template's own space, which is + TEXT_TEMPLATE_FONT_SCALE times bigger than the space positions use, so any + check that mixes the two has to convert first.""" + return float(style.get("fontSize")) / TEXT_TEMPLATE_FONT_SCALE # The earlier rhythm: one title per WORD. The default is now the # progressive composition (one title per LINE), covered in @@ -106,7 +121,7 @@ class TestGenerateDynamicSubtitles: param_names = [p.get("name") for p in title.findall("param")] assert param_names == [ - "Position", "Layout Method", "Left Margin", "Right Margin", + "Position", "Build Out", "Layout Method", "Left Margin", "Right Margin", "Top Margin", "Bottom Margin", "Alignment", "Line Spacing", "Auto-Shrink", "Alignment", "Opacity", "Speed", "Custom Speed", "Apply Speed", @@ -132,7 +147,8 @@ class TestGenerateDynamicSubtitles: assert style_def.get("fontFace") == first.face scale = (1080 * POINT_SCALE) / REFERENCE_CANVAS_HEIGHT - assert style_def.get("fontSize") == str(round(first.font_size * scale)) + expected = round(first.font_size * scale) * TEXT_TEMPLATE_FONT_SCALE + assert float(style_def.get("fontSize")) == expected def test_font_size_scales_with_the_frame(self, temp_fcpxml): """A vertical 2160x3840 timeline must get the reference sizes back @@ -144,7 +160,9 @@ class TestGenerateDynamicSubtitles: title = modifier.generate_dynamic_subtitles("Interview_A", WORDS, WORD_MODE)[0] run = title.find("text/text-style") style_def = title.find(f"text-style-def[@id='{run.get('ref')}']/text-style") - assert style_def.get("fontSize") == str(WordStyle().rhythm[0].font_size) + assert float(style_def.get("fontSize")) == ( + WordStyle().rhythm[0].font_size * TEXT_TEMPLATE_FONT_SCALE + ) def test_position_is_keyframed_constant_hold(self, temp_fcpxml): """Regression (2026-08-17): position is a STATIC param in the "Text" @@ -685,7 +703,7 @@ class TestProgressiveComposition: style = self._style(t) top, bottom = ink_extent( t.find("text/text-style").text, - float(style.get("fontSize")), + font_points(style), font=style.get("font"), face=style.get("fontFace"), ) @@ -718,7 +736,7 @@ class TestProgressiveComposition: style = self._style(t) top, bottom = ink_extent( t.find("text/text-style").text, - float(style.get("fontSize")), + font_points(style), font=style.get("font"), face=style.get("fontFace"), ) @@ -762,3 +780,146 @@ class TestProgressiveComposition: spoken = " ".join(w["word"] for w in words) emitted = " ".join(t.find("text/text-style").text for t in titles) assert emitted == spoken, "no spoken word may be dropped or reordered" + + +class TestTemplateFontScale: + """The "Text" template sizes type in frame pixels but positions in canvas + points, so the emitted fontSize must be converted or the block renders in + the right place at half the chosen size.""" + + STYLE = WordStyle( + emphasis_look=WordLook(200, "1 1 1 1", font="Georgia", kerning=3.0), + body_look=WordLook(100, "1 1 1 1", font="Helvetica Neue", kerning=3.0), + ) + + def _styles(self, titles): + return [t.find(".//text-style-def/text-style") for t in titles] + + def _generate(self, path, **kwargs): + modifier = FCPXMLModifier(path) + titles = modifier.generate_dynamic_subtitles( + "Interview_A", WORDS, + DynamicSubtitleConfig(style=self.STYLE, **kwargs), + ) + return titles + + def test_emitted_size_is_the_layout_size_times_the_template_scale(self, temp_fcpxml): + unscaled = self._generate(temp_fcpxml, text_scale=1.0) + scaled = self._generate(temp_fcpxml, text_scale=TEXT_TEMPLATE_FONT_SCALE) + for plain, big in zip(self._styles(unscaled), self._styles(scaled)): + assert float(big.get("fontSize")) == ( + float(plain.get("fontSize")) * TEXT_TEMPLATE_FONT_SCALE + ) + + def test_default_config_applies_the_template_scale(self, temp_fcpxml): + assert DynamicSubtitleConfig().text_scale == TEXT_TEMPLATE_FONT_SCALE + default = self._styles(self._generate(temp_fcpxml)) + unscaled = self._styles(self._generate(temp_fcpxml, text_scale=1.0)) + assert [s.get("fontSize") for s in default] != [ + s.get("fontSize") for s in unscaled + ] + + def test_kerning_scales_with_the_font_size(self, temp_fcpxml): + """Kerning is in font units too — leaving it behind would tighten the + letter spacing to half as the type doubled.""" + unscaled = self._styles(self._generate(temp_fcpxml, text_scale=1.0)) + scaled = self._styles(self._generate(temp_fcpxml, text_scale=2.0)) + for plain, big in zip(unscaled, scaled): + if plain.get("kerning"): + assert float(big.get("kerning")) == float(plain.get("kerning")) * 2 + + def _positions(self, titles): + return [ + tuple(float(v) for v in t.find( + "param[@name='Position']").get("value").split()) + for t in titles + ] + + def test_position_is_converted_with_the_type(self, temp_fcpxml): + """The template reads fontSize and Position in the SAME space, so the + conversion has to reach both. Scaling only the type leaves the block at + the old spread with twice the type in it, and the lines collide.""" + plain = self._generate(temp_fcpxml, text_scale=1.0) + big = self._generate(temp_fcpxml, text_scale=2.0) + for (x, y), (x2, y2) in zip(self._positions(plain), self._positions(big)): + assert (x2, y2) == pytest.approx((x * 2, y * 2), rel=1e-4, abs=0.01) + + def test_type_and_spacing_keep_their_ratio_at_any_scale(self, temp_fcpxml): + """The invariant that broke in the field: the distance between two + lines, measured in font sizes, must not depend on the scale.""" + ratios = [] + for scale in (1.0, 2.0, 3.5): + titles = self._generate(temp_fcpxml, text_scale=scale) + ys = [y for _, y in self._positions(titles)] + sizes = [float(s.get("fontSize")) for s in self._styles(titles)] + ratios.append([ + (a - b) / size + for a, b, size in zip(ys, ys[1:], sizes) + ]) + for other in ratios[1:]: + assert other == pytest.approx(ratios[0], rel=1e-4, abs=1e-4) + + +class TestLineGap: + """The air between stacked lines is a design choice, negative included.""" + + STYLE = WordStyle( + emphasis_look=WordLook(200, "1 1 1 1", font="Georgia", kerning=0.0), + body_look=WordLook(100, "1 1 1 1", font="Helvetica Neue", kerning=0.0), + ) + PHRASE = "eu tinha muita dificuldade de encontrar roupa" + + def _blocks(self, gap): + """Compose in a band tall enough to hold every line, so the gap is the + only thing that changes — a short band would also change how many + lines fit, which is a different effect.""" + words = [ + {"word": w, "start": i * 0.3, "end": i * 0.3 + 0.3} + for i, w in enumerate(self.PHRASE.split()) + ] + box = LayoutBox(width=1080 * 0.92, height=100_000, center_y=0) + return compose_sentence(words, self.STYLE, box, line_gap=gap).blocks + + def test_default_matches_the_reference_gap(self): + assert DynamicSubtitleConfig().line_gap == REFERENCE_BLOCK_LINE_GAP + + def test_each_step_changes_by_exactly_the_gap(self): + """The stack places ink boxes edge to edge, so the gap is the whole + distance between two lines beyond their own ink.""" + zero = [b.y for b in self._blocks(0.0)] + loose = [b.y for b in self._blocks(50.0)] + assert len(zero) == len(loose) >= 2 + for plain, spaced in zip( + [a - b for a, b in zip(zero, zero[1:])], + [a - b for a, b in zip(loose, loose[1:])], + ): + assert spaced == pytest.approx(plain + 50.0) + + def test_a_negative_gap_overlaps_by_exactly_that_much(self): + """Negative is a supported look, not a failure: -40 tucks each line 40 + points into the one above rather than colliding by some amount the + caller cannot predict.""" + zero = [b.y for b in self._blocks(0.0)] + tucked = [b.y for b in self._blocks(-40.0)] + assert len(zero) == len(tucked) >= 2 + for plain, tight in zip( + [a - b for a, b in zip(zero, zero[1:])], + [a - b for a, b in zip(tucked, tucked[1:])], + ): + assert tight == pytest.approx(plain - 40.0) + + def test_the_gap_reaches_the_generated_titles(self, temp_fcpxml): + """The config field has to survive the trip to the XML.""" + def spread(gap): + modifier = FCPXMLModifier(temp_fcpxml) + titles = modifier.generate_dynamic_subtitles( + "Interview_A", WORDS, + DynamicSubtitleConfig(style=self.STYLE, line_gap=gap, text_scale=1.0), + ) + ys = [ + float(t.find("param[@name='Position']").get("value").split()[1]) + for t in titles + ] + return max(ys) - min(ys) + + assert spread(0.0) < spread(80.0) diff --git a/code/tests/test_emphasis.py b/code/tests/test_emphasis.py new file mode 100644 index 0000000..1aa235d --- /dev/null +++ b/code/tests/test_emphasis.py @@ -0,0 +1,112 @@ +"""Tests for fcpxml/emphasis.py — the emphasis index (pure, no audio needed).""" + +from fcpxml.emphasis import ( + EmphasisWeights, + annotate_emphasis, + compute_emphasis, + pause_weight, +) + + +def test_zero_input_gives_zero_score(): + assert compute_emphasis(0.0, 0.0, 0.0, 0.0, 0.0) == 0.0 + + +def test_max_input_gives_max_score(): + # pause sits inside the "dramatic beat" window, not a scene-change gap + score = compute_emphasis( + energy=1.0, pitch_delta=1.0, rate_delta=1.0, pause_before=1.5, word_duration=10.0 + ) + assert score == 1.0 + + +def test_negative_pitch_delta_uses_magnitude(): + a = compute_emphasis(0.0, pitch_delta=0.5, rate_delta=0.0, pause_before=0.0, word_duration=0.0) + b = compute_emphasis(0.0, pitch_delta=-0.5, rate_delta=0.0, pause_before=0.0, word_duration=0.0) + assert a == b > 0.0 + + +def test_duration_saturates_beyond_cap(): + at_cap = compute_emphasis(0.0, 0.0, 0.0, 0.0, word_duration=1.0, max_duration=1.0) + beyond = compute_emphasis(0.0, 0.0, 0.0, 0.0, word_duration=30.0, max_duration=1.0) + assert at_cap == beyond + + +class TestPauseWeight: + """A long gap is a scene change, not emphasis — it must not outrank a + word the speaker actually hit hard. Grounded in real footage where 6-9s + gaps were topping the emphasis ranking.""" + + def test_dramatic_beat_counts_fully(self): + assert pause_weight(1.5, max_pause=1.5) == 1.0 + + def test_short_beat_counts_proportionally(self): + assert pause_weight(0.75, max_pause=1.5) == 0.5 + + def test_long_gap_is_ignored(self): + assert pause_weight(8.7, max_pause=1.5, ignore_above=3.0) == 0.0 + + def test_no_pause_is_zero(self): + assert pause_weight(0.0) == 0.0 + + def test_cutoff_can_be_disabled(self): + assert pause_weight(30.0, max_pause=1.5, ignore_above=0.0) == 1.0 + + def test_scene_change_scores_below_a_loud_word(self): + gap = compute_emphasis(0.25, 0.04, 0.0, pause_before=6.2, word_duration=0.5) + loud = compute_emphasis(1.00, 0.28, 0.0, pause_before=1.9, word_duration=0.2) + assert loud > gap + + +def test_higher_energy_weight_increases_energy_contribution(): + low_weight = EmphasisWeights(energy=0.1, pitch_variation=0.0, rate_variation=0.0, pause_before=0.0, duration=0.0) + high_weight = EmphasisWeights(energy=1.0, pitch_variation=0.0, rate_variation=0.0, pause_before=0.0, duration=0.0) + # energy is the only nonzero factor for both weight sets, so normalized + # score should be identical regardless of the absolute weight value. + score_low = compute_emphasis(0.6, 0.0, 0.0, 0.0, 0.0, weights=low_weight) + score_high = compute_emphasis(0.6, 0.0, 0.0, 0.0, 0.0, weights=high_weight) + assert score_low == score_high + + +def test_all_zero_weights_returns_zero_not_error(): + zero_weights = EmphasisWeights(0.0, 0.0, 0.0, 0.0, 0.0) + assert compute_emphasis(1.0, 1.0, 1.0, 1.0, 1.0, weights=zero_weights) == 0.0 + + +def test_weights_round_trip_dict(): + w = EmphasisWeights(energy=0.4, pitch_variation=0.3, rate_variation=0.1, pause_before=0.1, duration=0.1) + restored = EmphasisWeights.from_dict(w.as_dict()) + assert restored == w + + +def test_from_dict_fills_missing_with_defaults(): + restored = EmphasisWeights.from_dict({"energy": 0.9}) + defaults = EmphasisWeights() + assert restored.energy == 0.9 + assert restored.pitch_variation == defaults.pitch_variation + + +def test_annotate_emphasis_adds_score_per_word(): + words = [ + {"word": "hi", "start": 0.0, "end": 0.3, "energy": 0.2, "pitch_delta": 0.1, "rate_delta": 0.1, "pause_before": 0.0}, + {"word": "WOW", "start": 1.0, "end": 1.5, "energy": 0.9, "pitch_delta": 0.8, "rate_delta": 0.7, "pause_before": 1.0}, + ] + annotated = annotate_emphasis(words) + assert len(annotated) == 2 + assert all("emphasis" in w for w in annotated) + assert annotated[1]["emphasis"] > annotated[0]["emphasis"] + + +def test_annotate_emphasis_does_not_mutate_input(): + words = [{"word": "hi", "start": 0.0, "end": 0.3, "energy": 0.5}] + annotate_emphasis(words) + assert "emphasis" not in words[0] + + +def test_annotate_emphasis_derives_duration_from_start_end(): + """A word with no explicit "duration" key gets it from end - start.""" + base = {"energy": 0.0, "pitch_delta": 0.0, "rate_delta": 0.0, "pause_before": 0.0} + derived = annotate_emphasis([{"word": "hi", "start": 1.0, "end": 1.5, **base}], max_duration=0.5) + explicit = annotate_emphasis([{"word": "hi", "start": 0.0, "end": 0.0, "duration": 0.5, **base}], max_duration=0.5) + # duration is the only nonzero factor in both, so scores must match + assert derived[0]["emphasis"] == explicit[0]["emphasis"] > 0.0 diff --git a/code/tests/test_media_intel.py b/code/tests/test_media_intel.py index 52af2d2..d3c16df 100755 --- a/code/tests/test_media_intel.py +++ b/code/tests/test_media_intel.py @@ -518,7 +518,7 @@ class TestDetectBeatsHandler: await handle_detect_beats({"media_path": str(bad)}) async def test_reports_when_librosa_unavailable(self, tmp_path, monkeypatch): - import server as server_mod + import server_tools.qc as server_mod from server import handle_detect_beats wav = tmp_path / "song.wav" diff --git a/code/tests/test_output_dir_routing.py b/code/tests/test_output_dir_routing.py new file mode 100644 index 0000000..6000ac9 --- /dev/null +++ b/code/tests/test_output_dir_routing.py @@ -0,0 +1,63 @@ +"""Tests for `output_dir` routing — the app's "Pasta do projeto" promise. + +The setting is documented in the UI as "everything generated is saved in +here". It used to be applied as a sandbox anchor only, while the filename +was still derived in the INPUT's directory — so any call whose output_dir +differed from the input's folder failed its own anchor check. +""" + +import pytest + +from server_tools._shared import _resolve_io_paths + + +@pytest.fixture +def project(tmp_path): + """An .fcpxml in one folder, with a separate chosen output folder.""" + source_dir = tmp_path / "media" + source_dir.mkdir() + fcpxml = source_dir / "Projeto.fcpxml" + fcpxml.write_text("<fcpxml version='1.13'/>") + chosen = tmp_path / "pasta do projeto" + chosen.mkdir() + return fcpxml, chosen + + +class TestOutputDirRouting: + def test_output_lands_in_the_chosen_folder(self, project): + fcpxml, chosen = project + _, output_path = _resolve_io_paths({"filepath": str(fcpxml), "output_dir": str(chosen)}, "_voice_edit") + assert output_path.startswith(str(chosen)) + + def test_filename_keeps_the_suffix_convention(self, project): + fcpxml, chosen = project + _, output_path = _resolve_io_paths({"filepath": str(fcpxml), "output_dir": str(chosen)}, "_voice_edit") + assert output_path.endswith("Projeto_voice_edit.fcpxml") + + def test_cross_directory_call_does_not_raise(self, project): + """The regression: output_dir different from the input folder used to + raise "output path escapes allowed directory" every single time.""" + fcpxml, chosen = project + _resolve_io_paths({"filepath": str(fcpxml), "output_dir": str(chosen)}, "_dynamic_subtitles") + + def test_without_output_dir_it_still_writes_beside_the_input(self, project): + fcpxml, _ = project + _, output_path = _resolve_io_paths({"filepath": str(fcpxml)}, "_modified") + assert output_path == str(fcpxml.parent / "Projeto_modified.fcpxml") + + def test_explicit_output_path_still_wins(self, project): + fcpxml, chosen = project + target = chosen / "nome escolhido.fcpxml" + _, output_path = _resolve_io_paths( + {"filepath": str(fcpxml), "output_dir": str(chosen), "output_path": str(target)}, "_voice_edit" + ) + assert output_path == str(target) + + def test_explicit_output_path_outside_the_anchor_is_rejected(self, project, tmp_path): + """The anchor must keep constraining explicit paths, not just names.""" + fcpxml, chosen = project + with pytest.raises(ValueError): + _resolve_io_paths( + {"filepath": str(fcpxml), "output_dir": str(chosen), + "output_path": str(tmp_path / "fora.fcpxml")}, "_voice_edit" + ) diff --git a/code/tests/test_project_config.py b/code/tests/test_project_config.py new file mode 100644 index 0000000..eeacd72 --- /dev/null +++ b/code/tests/test_project_config.py @@ -0,0 +1,83 @@ +"""Tests for the last-project settings — the folder/file the app reopens with. + +The config file (~/.fcp-mcp-server/config.json) is redirected to a tmp_path +so these never touch the developer's real settings. +""" + +import json + +import pytest + +from fcpxml import model_manager + + +@pytest.fixture(autouse=True) +def isolated_config(tmp_path, monkeypatch): + """Point model_manager's config file at a throwaway directory.""" + monkeypatch.setattr(model_manager, "_CONFIG_DIR", tmp_path) + monkeypatch.setattr(model_manager, "_CONFIG_FILE", tmp_path / "config.json") + return tmp_path / "config.json" + + +class TestLoadProjectConfig: + def test_defaults_when_nothing_stored(self): + assert model_manager.load_project_config() == model_manager.DEFAULT_PROJECT_CONFIG + + def test_non_dict_stored_value_falls_back(self, isolated_config): + isolated_config.write_text(json.dumps({"project": "nonsense"})) + assert model_manager.load_project_config() == model_manager.DEFAULT_PROJECT_CONFIG + + def test_reads_back_what_was_stored(self, tmp_path): + folder = tmp_path / "03 - Mastopexia" + folder.mkdir() + model_manager.save_project_config(folder=str(folder)) + assert model_manager.load_project_config()["folder"] == str(folder) + + def test_path_that_no_longer_exists_comes_back_empty(self, isolated_config, tmp_path): + """An unmounted volume must degrade to "nothing selected", not a dead path.""" + isolated_config.write_text(json.dumps({"project": {"folder": str(tmp_path / "gone")}})) + assert model_manager.load_project_config()["folder"] == "" + + def test_non_string_stored_value_is_ignored(self, isolated_config): + isolated_config.write_text(json.dumps({"project": {"folder": 42}})) + assert model_manager.load_project_config()["folder"] == "" + + +class TestSaveProjectConfig: + def test_omitted_field_keeps_its_current_value(self, tmp_path): + folder = tmp_path / "projeto" + folder.mkdir() + project = tmp_path / "projeto" / "Mastopexia.fcpxml" + project.write_text("<fcpxml/>") + model_manager.save_project_config(folder=str(folder), file=str(project)) + model_manager.save_project_config(file=str(project)) + assert model_manager.load_project_config()["folder"] == str(folder) + + def test_empty_string_clears_a_field(self, tmp_path): + folder = tmp_path / "projeto" + folder.mkdir() + model_manager.save_project_config(folder=str(folder)) + model_manager.save_project_config(folder="") + assert model_manager.load_project_config()["folder"] == "" + + def test_home_relative_path_is_expanded(self, isolated_config): + model_manager.save_project_config(folder="~/Movies") + stored = json.loads(isolated_config.read_text())["project"]["folder"] + assert not stored.startswith("~") + + def test_does_not_disturb_other_config_sections(self, isolated_config, tmp_path): + isolated_config.write_text(json.dumps({"language": "pt", "selected_model": "large-v3"})) + folder = tmp_path / "projeto" + folder.mkdir() + model_manager.save_project_config(folder=str(folder)) + data = json.loads(isolated_config.read_text()) + assert data["language"] == "pt" + assert data["selected_model"] == "large-v3" + + def test_returns_the_merged_config(self, tmp_path): + folder = tmp_path / "projeto" + folder.mkdir() + assert model_manager.save_project_config(folder=str(folder)) == { + "folder": str(folder), + "file": "", + } diff --git a/code/tests/test_refine_voice_timeline_tool.py b/code/tests/test_refine_voice_timeline_tool.py new file mode 100644 index 0000000..1330bbf --- /dev/null +++ b/code/tests/test_refine_voice_timeline_tool.py @@ -0,0 +1,167 @@ +"""Tests for the refine_voice_timeline MCP tool. + +The tool exists because emphasis is *relative*: cut the loudest moment of a +recording and every surviving score is still measured against something the +viewer never sees. These tests pin the re-normalization actually happening, +and the times staying in original source seconds so the result can be fed +straight back to apply_voice_actions. +""" + +import json + +import pytest + +from tests.test_voice_features_tool import _write_silent_wav +from tests.test_voice_timeline_tool import _TRANSCRIPT, patched, wav # noqa: F401 + +_ = _write_silent_wav, _TRANSCRIPT # re-exported fixtures need the imports + + +async def _build(wav_path): + from server import handle_build_voice_timeline + + await handle_build_voice_timeline({"media_path": str(wav_path)}) + + +@pytest.fixture +def two_candidates(monkeypatch): + """A transcript where BOTH lines carry a content word. + + The shared fixture's first line is "isso e" — two function words, which + ``suggest_zoom_windows`` skips by design, so it can never produce more + than one candidate to cap. + """ + import fcpxml.voice_timeline as vt + import server_tools._shared as _shared_mod + + transcript = { + "language": "pt", + "duration": 4.0, + "text": "cirurgia rapida seguranca total", + "segments": [ + {"text": "cirurgia rapida", "start": 0.0, "end": 1.0}, + {"text": "seguranca total", "start": 2.0, "end": 4.0}, + ], + "words": [ + {"word": "cirurgia", "start": 0.0, "end": 0.4, "confidence": 0.9}, + {"word": "rapida", "start": 0.5, "end": 0.7, "confidence": 0.9}, + {"word": "seguranca", "start": 2.0, "end": 2.9, "confidence": 0.9}, + {"word": "total", "start": 3.0, "end": 3.5, "confidence": 0.9}, + ], + } + monkeypatch.setattr(_shared_mod, "transcribe", lambda *a, **k: transcript) + monkeypatch.setattr(vt, "extract_pitch", lambda *a, **k: [(2.4, 260.0), (0.2, 120.0)]) + monkeypatch.setattr(vt, "extract_energy", lambda *a, **k: [(2.4, 0.95), (0.2, 0.10)]) + + +class TestRefineVoiceTimelineHandler: + async def test_requires_an_existing_timeline(self, wav): # noqa: F811 + from server import handle_refine_voice_timeline + + result = await handle_refine_voice_timeline({"media_path": str(wav), "cuts": []}) + assert "build_voice_timeline" in result[0].text + + async def test_rejects_disallowed_extension(self, tmp_path): + from server import handle_refine_voice_timeline + + bad = tmp_path / "clip.txt" + bad.write_text("not audio") + with pytest.raises(ValueError): + await handle_refine_voice_timeline({"media_path": str(bad), "cuts": []}) + + async def test_compares_raw_against_survivors(self, wav, patched): # noqa: F811 + from server import handle_refine_voice_timeline + + await _build(wav) + result = await handle_refine_voice_timeline( + {"media_path": str(wav), "cuts": [{"start": 0.0, "end": 1.0}]} + ) + text = result[0].text + assert "Survivors only" in text + assert "Average emphasis" in text + + async def test_cut_words_are_excluded(self, wav, patched): # noqa: F811 + from server import handle_refine_voice_timeline + + await _build(wav) + result = await handle_refine_voice_timeline( + {"media_path": str(wav), "cuts": [{"start": 0.0, "end": 1.0}], "save": True} + ) + assert "_voice_timeline_refined.json" in result[0].text + data = json.loads( + (wav.parent / "clip_voice_timeline_refined.json").read_text(encoding="utf-8") + ) + words = [w["text"] for s in data["segments"] for w in s["words"]] + assert "isso" not in words + assert "seguranca" in words + + async def test_times_stay_in_original_source_seconds(self, wav, patched): # noqa: F811 + """A cut at the head must NOT slide the survivors back to zero.""" + from server import handle_refine_voice_timeline + + await _build(wav) + await handle_refine_voice_timeline( + {"media_path": str(wav), "cuts": [{"start": 0.0, "end": 1.0}], "save": True} + ) + data = json.loads( + (wav.parent / "clip_voice_timeline_refined.json").read_text(encoding="utf-8") + ) + first = data["segments"][0]["words"][0] + assert first["start"] == pytest.approx(2.0) + + async def test_proposes_zoom_candidates(self, wav, patched): # noqa: F811 + from server import handle_refine_voice_timeline + + await _build(wav) + result = await handle_refine_voice_timeline({"media_path": str(wav), "cuts": []}) + assert "Zoom Candidates" in result[0].text + + async def test_max_zooms_caps_the_list(self, wav, two_candidates): # noqa: F811 + from server import handle_refine_voice_timeline + + await _build(wav) + + async def zoom_rows(**extra): + result = await handle_refine_voice_timeline( + {"media_path": str(wav), "cuts": [], "min_gap": 0.0, **extra} + ) + section = result[0].text.split("## Zoom Candidates", 1)[1] + return [ + ln for ln in section.splitlines() + if ln.startswith("| ") and ln.rstrip().endswith("|") and "Start" not in ln + ] + + assert len(await zoom_rows()) == 2 + assert len(await zoom_rows(max_zooms=1)) == 1 + + async def test_malformed_cut_is_reported_not_raised(self, wav, patched): # noqa: F811 + from server import handle_refine_voice_timeline + + await _build(wav) + result = await handle_refine_voice_timeline( + {"media_path": str(wav), "cuts": [{"start": 3.0, "end": 1.0}]} + ) + text = result[0].text + assert "Rejected cuts" in text + assert "must be after start" in text + + async def test_cutting_everything_says_so(self, wav, patched): # noqa: F811 + from server import handle_refine_voice_timeline + + await _build(wav) + result = await handle_refine_voice_timeline( + {"media_path": str(wav), "cuts": [{"start": 0.0, "end": 60.0}]} + ) + assert "removed every word" in result[0].text + + +class TestRefineVoiceTimelineRegistration: + async def test_tool_is_listed(self): + from server import list_tools + + assert "refine_voice_timeline" in {t.name for t in await list_tools()} + + async def test_tool_is_dispatched(self): + from server import TOOL_HANDLERS, handle_refine_voice_timeline + + assert TOOL_HANDLERS["refine_voice_timeline"] is handle_refine_voice_timeline diff --git a/code/tests/test_voice_actions.py b/code/tests/test_voice_actions.py new file mode 100644 index 0000000..7ffc797 --- /dev/null +++ b/code/tests/test_voice_actions.py @@ -0,0 +1,227 @@ +"""Tests for fcpxml/voice_actions.py — the decision contract. + +Pure functions over untrusted input (a model's decision list), so these +cover the rejection paths as carefully as the happy path. +""" + +import pytest + +from fcpxml.voice_actions import ( + MAX_TEXT_LENGTH, + VoiceAction, + merge_cut_ranges, + parse_actions, + resolve_actions, + shift_after_cuts, +) + + +class TestParseActions: + def test_accepts_bare_list(self): + actions, errors = parse_actions([{"kind": "cut", "start": 1.0, "end": 2.0}]) + assert len(actions) == 1 and errors == [] + + def test_accepts_actions_envelope(self): + actions, errors = parse_actions({"actions": [{"kind": "cut", "start": 1.0, "end": 2.0}]}) + assert len(actions) == 1 and errors == [] + + def test_rejects_non_list(self): + actions, errors = parse_actions("cortar tudo") + assert actions == [] and len(errors) == 1 + + def test_one_bad_row_does_not_discard_the_good_ones(self): + actions, errors = parse_actions([ + {"kind": "cut", "start": 1.0, "end": 2.0}, + {"kind": "teleport", "start": 3.0, "end": 4.0}, + {"kind": "zoom", "start": 5.0, "end": 6.0}, + ]) + assert len(actions) == 2 + assert len(errors) == 1 and "teleport" in errors[0] + + def test_rejects_unknown_kind(self): + _, errors = parse_actions([{"kind": "explode", "start": 0.0, "end": 1.0}]) + assert "explode" in errors[0] + + def test_rejects_non_numeric_times(self): + _, errors = parse_actions([{"kind": "cut", "start": "início", "end": 2.0}]) + assert "numbers" in errors[0] + + def test_rejects_negative_start(self): + _, errors = parse_actions([{"kind": "cut", "start": -1.0, "end": 2.0}]) + assert "negative" in errors[0] + + def test_rejects_end_before_start(self): + _, errors = parse_actions([{"kind": "cut", "start": 5.0, "end": 2.0}]) + assert "must be after" in errors[0] + + def test_rejects_zero_length(self): + _, errors = parse_actions([{"kind": "cut", "start": 2.0, "end": 2.0}]) + assert errors + + def test_rejects_row_that_is_not_an_object(self): + _, errors = parse_actions(["cortar aos 5s"]) + assert "expected an object" in errors[0] + + def test_kind_is_case_insensitive(self): + actions, _ = parse_actions([{"kind": "ZOOM", "start": 1.0, "end": 2.0}]) + assert actions[0].kind == "zoom" + + def test_preserves_reason_and_speaker(self): + actions, _ = parse_actions([ + {"kind": "zoom", "start": 1.0, "end": 2.0, + "reason": "argumento central", "speaker": "SPEAKER_01"} + ]) + assert actions[0].reason == "argumento central" + assert actions[0].speaker == "SPEAKER_01" + + +class TestZoomValidation: + def test_default_scale_when_absent(self): + actions, _ = parse_actions([{"kind": "zoom", "start": 1.0, "end": 2.0}]) + assert actions[0].params["scale"] == 1.3 + + def test_rejects_scale_below_one(self): + _, errors = parse_actions([ + {"kind": "zoom", "start": 1.0, "end": 2.0, "params": {"scale": 0.5}} + ]) + assert "outside" in errors[0] + + def test_rejects_absurd_scale(self): + _, errors = parse_actions([ + {"kind": "zoom", "start": 1.0, "end": 2.0, "params": {"scale": 50}} + ]) + assert "outside" in errors[0] + + def test_rejects_non_numeric_scale(self): + _, errors = parse_actions([ + {"kind": "zoom", "start": 1.0, "end": 2.0, "params": {"scale": "muito"}} + ]) + assert "must be a number" in errors[0] + + +class TestTextValidation: + def test_requires_content(self): + _, errors = parse_actions([{"kind": "text", "start": 1.0, "end": 2.0}]) + assert "params.content" in errors[0] + + def test_rejects_blank_content(self): + _, errors = parse_actions([ + {"kind": "text", "start": 1.0, "end": 2.0, "params": {"content": " "}} + ]) + assert "params.content" in errors[0] + + def test_truncates_overlong_content(self): + actions, _ = parse_actions([ + {"kind": "text", "start": 1.0, "end": 2.0, "params": {"content": "A" * 500}} + ]) + assert len(actions[0].params["content"]) == MAX_TEXT_LENGTH + + +class TestMergeCutRanges: + def test_sorts_and_merges_overlaps(self): + actions = [ + VoiceAction("cut", 5.0, 7.0), + VoiceAction("cut", 1.0, 3.0), + VoiceAction("cut", 2.0, 4.0), + ] + assert merge_cut_ranges(actions) == [(1.0, 4.0), (5.0, 7.0)] + + def test_merges_touching_ranges(self): + actions = [VoiceAction("cut", 1.0, 2.0), VoiceAction("cut", 2.0, 3.0)] + assert merge_cut_ranges(actions) == [(1.0, 3.0)] + + def test_ignores_non_cut_actions(self): + assert merge_cut_ranges([VoiceAction("zoom", 1.0, 2.0)]) == [] + + +class TestShiftAfterCuts: + def test_time_before_any_cut_is_unchanged(self): + assert shift_after_cuts(0.5, [(2.0, 4.0)]) == 0.5 + + def test_time_after_a_cut_moves_earlier(self): + assert shift_after_cuts(6.0, [(2.0, 4.0)]) == pytest.approx(4.0) + + def test_time_inside_a_cut_is_dropped(self): + assert shift_after_cuts(3.0, [(2.0, 4.0)]) is None + + def test_multiple_cuts_accumulate(self): + cuts = [(1.0, 2.0), (5.0, 7.0)] + assert shift_after_cuts(10.0, cuts) == pytest.approx(7.0) + + def test_no_cuts_is_identity(self): + assert shift_after_cuts(3.0, []) == 3.0 + + def test_boundary_start_of_cut_is_inside(self): + assert shift_after_cuts(2.0, [(2.0, 4.0)]) is None + + def test_boundary_end_of_cut_survives(self): + assert shift_after_cuts(4.0, [(2.0, 4.0)]) == pytest.approx(2.0) + + +class TestResolveActions: + def test_zoom_after_a_cut_is_moved_earlier(self): + actions = [VoiceAction("cut", 2.0, 4.0), VoiceAction("zoom", 6.0, 7.0)] + cuts, placed, dropped = resolve_actions(actions) + assert cuts == [(2.0, 4.0)] + assert dropped == [] + assert placed[0].start == pytest.approx(4.0) + assert placed[0].end == pytest.approx(5.0) + + def test_zoom_inside_a_cut_is_dropped_not_slid(self): + actions = [VoiceAction("cut", 2.0, 8.0), VoiceAction("zoom", 3.0, 4.0)] + _, placed, dropped = resolve_actions(actions) + assert placed == [] + assert len(dropped) == 1 + + def test_zoom_straddling_a_cut_edge_is_dropped(self): + actions = [VoiceAction("cut", 4.0, 8.0), VoiceAction("zoom", 3.0, 5.0)] + _, placed, dropped = resolve_actions(actions) + assert placed == [] and len(dropped) == 1 + + def test_cuts_are_not_returned_as_placed(self): + _, placed, _ = resolve_actions([VoiceAction("cut", 1.0, 2.0)]) + assert placed == [] + + def test_without_cuts_everything_keeps_its_time(self): + actions = [VoiceAction("zoom", 3.0, 4.0), VoiceAction("text", 5.0, 6.0)] + cuts, placed, dropped = resolve_actions(actions) + assert cuts == [] and dropped == [] + assert [(a.start, a.end) for a in placed] == [(3.0, 4.0), (5.0, 6.0)] + + def test_params_survive_the_shift(self): + actions = [ + VoiceAction("cut", 1.0, 2.0), + VoiceAction("text", 5.0, 6.0, params={"content": "SEGURANÇA"}), + ] + _, placed, _ = resolve_actions(actions) + assert placed[0].params["content"] == "SEGURANÇA" + + +class TestMarkersSurviveCutEdges: + """A marker is a point in time, not a span. The markers worth keeping are + precisely the ones flagging a join, which sit against a cut edge — so + requiring their nominal end to survive would drop exactly those.""" + + def test_marker_at_a_cut_edge_survives(self): + actions = [VoiceAction("cut", 21.9, 127.6), VoiceAction("marker", 21.85, 22.0)] + _, placed, dropped = resolve_actions(actions) + assert dropped == [] + assert placed[0].kind == "marker" + assert placed[0].start == pytest.approx(21.85) + + def test_marker_inside_removed_material_is_still_dropped(self): + actions = [VoiceAction("cut", 20.0, 100.0), VoiceAction("marker", 50.0, 50.2)] + _, placed, dropped = resolve_actions(actions) + assert placed == [] and len(dropped) == 1 + + def test_marker_keeps_its_length_after_shifting(self): + actions = [VoiceAction("cut", 0.0, 10.0), VoiceAction("marker", 20.0, 20.5)] + _, placed, _ = resolve_actions(actions) + assert placed[0].start == pytest.approx(10.0) + assert placed[0].duration == pytest.approx(0.5) + + def test_zoom_straddling_an_edge_is_still_dropped(self): + """Only markers get the point-action treatment — a span must fit.""" + actions = [VoiceAction("cut", 21.9, 127.6), VoiceAction("zoom", 21.0, 22.5)] + _, placed, dropped = resolve_actions(actions) + assert placed == [] and len(dropped) == 1 diff --git a/code/tests/test_voice_actions_tool.py b/code/tests/test_voice_actions_tool.py new file mode 100644 index 0000000..f8517e6 --- /dev/null +++ b/code/tests/test_voice_actions_tool.py @@ -0,0 +1,267 @@ +"""Tests for the apply_voice_actions MCP tool — decisions -> real FCPXML. + +Uses an inline fixture rather than examples/sample.fcpxml so the source +windows are explicit and the assertions can be exact. +""" + +import shutil + +import pytest + +from fcpxml.safe_xml import safe_parse + +_FIXTURE = """<?xml version="1.0" encoding="UTF-8"?> +<fcpxml version="1.13"> + <resources> + <format id="r1" name="FFVideoFormat1080p30" frameDuration="100/3000s" width="1920" height="1080"/> + <asset id="a1" name="entrevista" start="0s" duration="600/30s" hasVideo="1" hasAudio="1" format="r1"> + <media-rep kind="original-media" src="file:///media/entrevista.mov"/> + </asset> + </resources> + <library> + <event name="Ev"> + <project name="Proj"> + <sequence format="r1" duration="600/30s" tcStart="0s"> + <spine> + <asset-clip name="entrevista" ref="a1" offset="0s" start="0s" duration="600/30s"/> + </spine> + </sequence> + </project> + </event> + </library> +</fcpxml> +""" + + +@pytest.fixture +def project(tmp_path): + path = tmp_path / "proj.fcpxml" + path.write_text(_FIXTURE) + return path + + +def _out(project): + return project.with_name("proj_voice_edit.fcpxml") + + +class TestApplyVoiceActionsHandler: + async def test_no_actions_reports_instead_of_writing(self, project): + from server import handle_apply_voice_actions + + result = await handle_apply_voice_actions({"filepath": str(project)}) + assert "nothing to apply" in result[0].text.lower() + assert not _out(project).exists() + + async def test_all_invalid_actions_writes_nothing(self, project): + from server import handle_apply_voice_actions + + result = await handle_apply_voice_actions({ + "filepath": str(project), + "actions": [{"kind": "teleport", "start": 1.0, "end": 2.0}], + }) + assert "No valid actions" in result[0].text + assert "teleport" in result[0].text + assert not _out(project).exists() + + async def test_applies_zoom_into_the_hosting_clip(self, project): + from server import handle_apply_voice_actions + + result = await handle_apply_voice_actions({ + "filepath": str(project), + "actions": [{ + "kind": "zoom", "start": 5.0, "end": 6.0, + "params": {"scale": 1.4}, "reason": "argumento central", + }], + }) + assert "argumento central" in result[0].text + + tree = safe_parse(str(_out(project))) + transforms = tree.getroot().findall(".//adjust-transform") + assert len(transforms) == 1 + + async def test_applies_text_title(self, project): + from server import handle_apply_voice_actions + + await handle_apply_voice_actions({ + "filepath": str(project), + "actions": [{ + "kind": "text", "start": 3.0, "end": 4.0, + "params": {"content": "SEGURANÇA"}, + }], + }) + titles = safe_parse(str(_out(project))).getroot().findall(".//title") + assert len(titles) == 1 + texts = [t.text for t in titles[0].iter() if t.text] + assert any("SEGURANÇA" in t for t in texts) + + async def test_applies_marker(self, project): + from server import handle_apply_voice_actions + + await handle_apply_voice_actions({ + "filepath": str(project), + "actions": [{ + "kind": "marker", "start": 2.0, "end": 2.5, "reason": "virada", + }], + }) + markers = safe_parse(str(_out(project))).getroot().findall(".//marker") + assert len(markers) == 1 + + async def test_cut_shortens_the_timeline(self, project): + from server import handle_apply_voice_actions + + before = safe_parse(str(project)).getroot().find(".//asset-clip").get("duration") + result = await handle_apply_voice_actions({ + "filepath": str(project), + "actions": [{"kind": "cut", "start": 5.0, "end": 10.0, "reason": "digressão"}], + }) + assert "Cuts applied" in result[0].text + + clips = safe_parse(str(_out(project))).getroot().findall(".//asset-clip") + total = sum( + int(c.get("duration").split("/")[0]) / int(c.get("duration").split("/")[1].rstrip("s")) + for c in clips + ) + original = int(before.split("/")[0]) / int(before.split("/")[1].rstrip("s")) + assert total < original + + async def test_action_inside_a_cut_is_dropped_and_reported(self, project): + from server import handle_apply_voice_actions + + result = await handle_apply_voice_actions({ + "filepath": str(project), + "actions": [ + {"kind": "cut", "start": 4.0, "end": 12.0}, + {"kind": "zoom", "start": 6.0, "end": 7.0, "reason": "some no material cortado"}, + ], + }) + text = result[0].text + assert "Dropped" in text + assert "some no material cortado" in text + assert safe_parse(str(_out(project))).getroot().findall(".//adjust-transform") == [] + + async def test_action_beyond_the_media_is_reported_not_silent(self, project): + from server import handle_apply_voice_actions + + result = await handle_apply_voice_actions({ + "filepath": str(project), + "actions": [{"kind": "zoom", "start": 500.0, "end": 501.0}], + }) + assert "Not placed" in result[0].text + assert "outside the edited timeline" in result[0].text + + async def test_original_file_is_untouched(self, project): + from server import handle_apply_voice_actions + + original = project.read_text() + await handle_apply_voice_actions({ + "filepath": str(project), + "actions": [{"kind": "zoom", "start": 5.0, "end": 6.0}], + }) + assert project.read_text() == original + + async def test_mixed_valid_and_invalid_applies_the_valid_ones(self, project): + from server import handle_apply_voice_actions + + result = await handle_apply_voice_actions({ + "filepath": str(project), + "actions": [ + {"kind": "zoom", "start": 5.0, "end": 6.0}, + {"kind": "zoom", "start": 8.0, "end": 9.0, "params": {"scale": 99}}, + ], + }) + assert "Rejected" in result[0].text + assert len(safe_parse(str(_out(project))).getroot().findall(".//adjust-transform")) == 1 + + async def test_respects_explicit_output_path(self, project, tmp_path): + from server import handle_apply_voice_actions + + target = tmp_path / "custom.fcpxml" + await handle_apply_voice_actions({ + "filepath": str(project), + "actions": [{"kind": "marker", "start": 1.0, "end": 2.0}], + "output_path": str(target), + }) + assert target.exists() + + async def test_output_is_valid_parseable_fcpxml(self, project): + from server import handle_apply_voice_actions + + await handle_apply_voice_actions({ + "filepath": str(project), + "actions": [ + {"kind": "cut", "start": 2.0, "end": 4.0}, + {"kind": "zoom", "start": 10.0, "end": 11.0}, + {"kind": "text", "start": 12.0, "end": 13.0, "params": {"content": "OK"}}, + ], + }) + root = safe_parse(str(_out(project))).getroot() + assert root.tag == "fcpxml" + assert root.find(".//spine") is not None + + +class TestSampleFixtureStillParses: + """The applier must not corrupt a real-world document.""" + + async def test_real_sample_survives_a_zoom(self, tmp_path): + from server import handle_apply_voice_actions + + src = "examples/sample.fcpxml" + target = tmp_path / "sample.fcpxml" + shutil.copy(src, target) + result = await handle_apply_voice_actions({ + "filepath": str(target), + "actions": [{"kind": "marker", "start": 1.0, "end": 2.0, "reason": "teste"}], + }) + assert "Voice Actions Applied" in result[0].text + + +class TestPlacementsLandOnTheRightPieceAfterCuts: + """Cutting splits a clip into same-named pieces. Placing before cutting + duplicated the zoom onto every piece and lost markers outright; a + name-based lookup afterwards would always resolve to the first piece. + Both bugs shipped past the suite and only showed up on real footage.""" + + async def test_zoom_lands_on_exactly_one_piece(self, project): + from server import handle_apply_voice_actions + + await handle_apply_voice_actions({ + "filepath": str(project), + "actions": [ + {"kind": "cut", "start": 2.0, "end": 5.0}, + {"kind": "cut", "start": 8.0, "end": 12.0}, + {"kind": "zoom", "start": 15.0, "end": 16.0, "params": {"scale": 1.2}}, + ], + }) + root = safe_parse(str(_out(project))).getroot() + assert len(root.findall(".//spine/asset-clip")) == 3 + assert len(root.findall(".//adjust-transform")) == 1 + + async def test_zoom_lands_on_the_last_piece_not_the_first(self, project): + from server import handle_apply_voice_actions + + await handle_apply_voice_actions({ + "filepath": str(project), + "actions": [ + {"kind": "cut", "start": 2.0, "end": 5.0}, + {"kind": "zoom", "start": 15.0, "end": 16.0}, + ], + }) + clips = safe_parse(str(_out(project))).getroot().findall(".//spine/asset-clip") + # the zoom is at 15s source -> 12s after a 3s cut, i.e. the 2nd piece + assert clips[0].find("adjust-transform") is None + assert clips[1].find("adjust-transform") is not None + + async def test_markers_survive_the_cut(self, project): + from server import handle_apply_voice_actions + + result = await handle_apply_voice_actions({ + "filepath": str(project), + "actions": [ + {"kind": "cut", "start": 5.0, "end": 10.0}, + {"kind": "marker", "start": 4.9, "end": 5.05, "reason": "emenda"}, + {"kind": "marker", "start": 15.0, "end": 15.2, "reason": "depois"}, + ], + }) + assert "Dropped" not in result[0].text + markers = safe_parse(str(_out(project))).getroot().findall(".//marker") + assert len(markers) == 2 diff --git a/code/tests/test_voice_analysis_config.py b/code/tests/test_voice_analysis_config.py new file mode 100644 index 0000000..602f9a1 --- /dev/null +++ b/code/tests/test_voice_analysis_config.py @@ -0,0 +1,111 @@ +"""Tests for the Voice Analysis settings — persistence + MCP config tools. + +The config file (~/.fcp-mcp-server/config.json) is redirected to a tmp_path +so these never touch the developer's real settings. +""" + +import json + +import pytest + +from fcpxml import model_manager + + +@pytest.fixture(autouse=True) +def isolated_config(tmp_path, monkeypatch): + """Point model_manager's config file at a throwaway directory.""" + monkeypatch.setattr(model_manager, "_CONFIG_DIR", tmp_path) + monkeypatch.setattr(model_manager, "_CONFIG_FILE", tmp_path / "config.json") + return tmp_path / "config.json" + + +class TestLoadVoiceAnalysisConfig: + def test_defaults_when_nothing_stored(self): + cfg = model_manager.load_voice_analysis_config() + assert cfg == model_manager.DEFAULT_VOICE_ANALYSIS_CONFIG + + def test_defaults_are_not_shared_mutable_state(self): + cfg = model_manager.load_voice_analysis_config() + cfg["emphasis_weights"]["energy"] = 0.99 + fresh = model_manager.load_voice_analysis_config() + assert fresh["emphasis_weights"]["energy"] == 0.30 + + def test_malformed_stored_values_fall_back_to_defaults(self, isolated_config): + isolated_config.write_text(json.dumps({"voice_analysis": {"energy_threshold": "loud"}})) + cfg = model_manager.load_voice_analysis_config() + assert cfg["energy_threshold"] == 0.5 + + def test_non_dict_stored_value_falls_back(self, isolated_config): + isolated_config.write_text(json.dumps({"voice_analysis": "nonsense"})) + assert model_manager.load_voice_analysis_config() == model_manager.DEFAULT_VOICE_ANALYSIS_CONFIG + + def test_thresholds_are_clamped_to_unit_range(self, isolated_config): + isolated_config.write_text( + json.dumps({"voice_analysis": {"energy_threshold": 5.0, "emphasis_floor": -2.0}}) + ) + cfg = model_manager.load_voice_analysis_config() + assert cfg["energy_threshold"] == 1.0 + assert cfg["emphasis_floor"] == 0.0 + + +class TestSaveVoiceAnalysisConfig: + def test_saves_and_reloads(self): + model_manager.save_voice_analysis_config(energy_threshold=0.7, emotion_enabled=True) + cfg = model_manager.load_voice_analysis_config() + assert cfg["energy_threshold"] == 0.7 + assert cfg["emotion_enabled"] is True + + def test_omitted_fields_keep_current_value(self): + model_manager.save_voice_analysis_config(energy_threshold=0.7) + model_manager.save_voice_analysis_config(emphasis_floor=0.9) + cfg = model_manager.load_voice_analysis_config() + assert cfg["energy_threshold"] == 0.7 + assert cfg["emphasis_floor"] == 0.9 + + def test_partial_weight_update_keeps_other_weights(self): + model_manager.save_voice_analysis_config(emphasis_weights={"energy": 0.55}) + weights = model_manager.load_voice_analysis_config()["emphasis_weights"] + assert weights["energy"] == 0.55 + assert weights["pitch_variation"] == 0.25 + + def test_unknown_weight_key_is_ignored(self): + model_manager.save_voice_analysis_config(emphasis_weights={"loudness": 9.0}) + weights = model_manager.load_voice_analysis_config()["emphasis_weights"] + assert "loudness" not in weights + + def test_does_not_clobber_unrelated_config_keys(self): + model_manager.save_hf_token("tok123") + model_manager.save_voice_analysis_config(energy_threshold=0.7) + assert model_manager.load_hf_token() == "tok123" + + def test_returns_full_merged_config(self): + returned = model_manager.save_voice_analysis_config(energy_threshold=0.7) + assert returned == model_manager.load_voice_analysis_config() + + +class TestVoiceAnalysisConfigTools: + async def test_get_reports_current_settings(self): + from server import handle_get_voice_analysis_config + + result = await handle_get_voice_analysis_config({}) + text = result[0].text + assert "Voice Analysis Settings" in text + assert "Emphasis Weights" in text + + async def test_save_persists_and_echoes_back(self): + from server import handle_save_voice_analysis_config + + result = await handle_save_voice_analysis_config( + {"energy_threshold": 0.8, "emotion_enabled": True} + ) + assert "0.80" in result[0].text + cfg = model_manager.load_voice_analysis_config() + assert cfg["energy_threshold"] == 0.8 + assert cfg["emotion_enabled"] is True + + async def test_save_with_no_arguments_is_a_noop(self): + from server import handle_save_voice_analysis_config + + before = model_manager.load_voice_analysis_config() + await handle_save_voice_analysis_config({}) + assert model_manager.load_voice_analysis_config() == before diff --git a/code/tests/test_voice_features.py b/code/tests/test_voice_features.py new file mode 100644 index 0000000..2e0b336 --- /dev/null +++ b/code/tests/test_voice_features.py @@ -0,0 +1,137 @@ +"""Tests for fcpxml/voice_features.py — acoustic features. + +The pure helpers (speech rate, pauses, window averaging) need no audio. +The librosa-backed extractors are skipped when the optional [intelligence] +extra is absent, matching the pattern in test_media_intel.py. +""" + +import math +import struct +import wave + +import pytest + +from fcpxml.voice_features import ( + compute_pauses, + compute_speech_rate, + extract_energy, + extract_pitch, + features_capability, + word_pitch_energy, +) + +try: + import librosa # noqa: F401 + + LIBROSA = True +except ImportError: + LIBROSA = False + + +def _write_tone_wav(path: str, hz: float = 220.0, seconds: float = 2.0, rate: int = 22050) -> None: + n = int(rate * seconds) + frames = [int(20000 * math.sin(2 * math.pi * hz * i / rate)) for i in range(n)] + with wave.open(path, "w") as f: + f.setnchannels(1) + f.setsampwidth(2) + f.setframerate(rate) + f.writeframes(struct.pack("<%dh" % n, *frames)) + + +class TestComputePauses: + def test_first_word_pause_is_time_from_zero(self): + words = [{"start": 1.5, "end": 2.0}] + assert compute_pauses(words) == [1.5] + + def test_gap_between_words(self): + words = [{"start": 0.0, "end": 1.0}, {"start": 2.5, "end": 3.0}] + assert compute_pauses(words) == [0.0, 1.5] + + def test_overlapping_words_clamp_to_zero(self): + words = [{"start": 0.0, "end": 2.0}, {"start": 1.0, "end": 3.0}] + assert compute_pauses(words) == [0.0, 0.0] + + def test_empty_words(self): + assert compute_pauses([]) == [] + + +class TestComputeSpeechRate: + def test_rate_counts_words_in_trailing_window(self): + # 3 words within a 3s window -> 1.0 word/sec at the last one + words = [{"start": 0.0}, {"start": 1.0}, {"start": 2.0}] + rates = compute_speech_rate(words, window_seconds=3.0) + assert rates[-1] == pytest.approx(1.0) + + def test_old_words_fall_out_of_window(self): + words = [{"start": 0.0}, {"start": 100.0}] + rates = compute_speech_rate(words, window_seconds=3.0) + # only the word itself is in range at t=100 + assert rates[-1] == pytest.approx(1 / 3.0) + + def test_zero_window_is_not_a_division_error(self): + assert compute_speech_rate([{"start": 0.0}], window_seconds=0.0) == [0.0] + + def test_empty_words(self): + assert compute_speech_rate([]) == [] + + +class TestWordPitchEnergy: + def test_averages_track_values_within_word_span(self): + words = [{"word": "a", "start": 0.0, "end": 1.0}] + pitch = [(0.0, 100.0), (0.5, 200.0), (5.0, 999.0)] + energy = [(0.0, 0.2), (1.0, 0.4)] + out = word_pitch_energy(words, pitch, energy) + assert out[0]["pitch_hz"] == pytest.approx(150.0) + assert out[0]["energy"] == pytest.approx(0.3) + + def test_none_when_no_frames_in_span(self): + words = [{"word": "a", "start": 10.0, "end": 11.0}] + out = word_pitch_energy(words, [(0.0, 100.0)], [(0.0, 0.5)]) + assert out[0]["pitch_hz"] is None + assert out[0]["energy"] is None + + def test_none_tracks_degrade_gracefully(self): + out = word_pitch_energy([{"word": "a", "start": 0.0, "end": 1.0}], None, None) + assert out[0]["pitch_hz"] is None + assert out[0]["energy"] is None + + def test_does_not_mutate_input(self): + words = [{"word": "a", "start": 0.0, "end": 1.0}] + word_pitch_energy(words, [(0.0, 100.0)], None) + assert "pitch_hz" not in words[0] + + def test_word_shorter_than_hop_gets_none_not_crash(self): + """A word briefer than the frame spacing may contain no frame at all.""" + words = [{"word": "a", "start": 0.501, "end": 0.502}] + out = word_pitch_energy(words, [(0.0, 100.0), (1.0, 200.0)], None) + assert out[0]["pitch_hz"] is None + + +class TestExtractorsDegradeGracefully: + def test_missing_file_returns_none(self): + assert extract_pitch("/nonexistent/audio.wav") is None + assert extract_energy("/nonexistent/audio.wav") is None + + +@pytest.mark.skipif(not LIBROSA, reason="librosa not installed") +class TestExtractorsWithLibrosa: + def test_capability_is_available(self): + ok, _msg = features_capability() + assert ok is True + + def test_extracts_pitch_of_known_tone(self, tmp_path): + wav = tmp_path / "tone.wav" + _write_tone_wav(str(wav), hz=220.0, seconds=2.0) + track = extract_pitch(str(wav)) + assert track is not None and len(track) > 0 + hz_values = sorted(hz for _t, hz in track) + median = hz_values[len(hz_values) // 2] + assert median == pytest.approx(220.0, rel=0.1) + + def test_extracts_energy_track(self, tmp_path): + wav = tmp_path / "tone.wav" + _write_tone_wav(str(wav), seconds=1.0) + track = extract_energy(str(wav)) + assert track is not None and len(track) > 0 + assert all(rms >= 0 for _t, rms in track) + assert max(rms for _t, rms in track) > 0 diff --git a/code/tests/test_voice_features_tool.py b/code/tests/test_voice_features_tool.py new file mode 100644 index 0000000..ffb5b6f --- /dev/null +++ b/code/tests/test_voice_features_tool.py @@ -0,0 +1,147 @@ +"""Tests for the analyze_voice_features MCP tool. + +librosa and Whisper are monkeypatched so these run without the optional +extras, matching the pattern used by TestDetectBeatsHandler. +""" + +import json +import struct +import wave + +import pytest + + +def _write_silent_wav(path: str, seconds: float = 2.0) -> None: + n = int(44100 * seconds) + with wave.open(path, "w") as f: + f.setnchannels(1) + f.setsampwidth(2) + f.setframerate(44100) + f.writeframes(struct.pack("<%dh" % n, *([0] * n))) + + +_FAKE_TRANSCRIPT = { + "language": "pt", + "duration": 3.0, + "text": "isso e seguranca", + "segments": [{"text": "isso e seguranca", "start": 0.0, "end": 3.0}], + "words": [ + {"word": "isso", "start": 0.0, "end": 0.4, "confidence": 0.9}, + {"word": "e", "start": 0.5, "end": 0.7, "confidence": 0.9}, + {"word": "seguranca", "start": 2.0, "end": 2.9, "confidence": 0.9}, + ], +} + + +@pytest.fixture +def wav(tmp_path): + path = tmp_path / "clip.wav" + _write_silent_wav(str(path)) + return path + + +@pytest.fixture +def patched_analysis(monkeypatch): + """Make the tool's transcription + librosa extractors deterministic.""" + import server_tools._shared as _shared_mod + import server_tools.voice as server_mod + + monkeypatch.setattr(_shared_mod, "transcribe", lambda *a, **k: _FAKE_TRANSCRIPT) + monkeypatch.setattr(server_mod, "features_capability", lambda: (True, "ok")) + # "seguranca" (2.0-2.9s) is the loud, high-pitched, emphatic word + monkeypatch.setattr( + server_mod, + "extract_pitch", + lambda *a, **k: [(0.2, 120.0), (0.6, 118.0), (2.4, 260.0)], + ) + monkeypatch.setattr( + server_mod, + "extract_energy", + lambda *a, **k: [(0.2, 0.10), (0.6, 0.12), (2.4, 0.95)], + ) + + +class TestAnalyzeVoiceFeaturesHandler: + async def test_reports_when_librosa_unavailable(self, wav, monkeypatch): + import server_tools.voice as server_mod + from server import handle_analyze_voice_features + + monkeypatch.setattr( + server_mod, "features_capability", lambda: (False, "componente librosa ausente.") + ) + result = await handle_analyze_voice_features({"media_path": str(wav)}) + assert "librosa" in result[0].text.lower() + + async def test_rejects_disallowed_extension(self, tmp_path): + from server import handle_analyze_voice_features + + bad = tmp_path / "clip.txt" + bad.write_text("not audio") + with pytest.raises(ValueError): + await handle_analyze_voice_features({"media_path": str(bad)}) + + async def test_writes_features_json_with_emphasis_per_word(self, wav, patched_analysis): + from server import handle_analyze_voice_features + + result = await handle_analyze_voice_features({"media_path": str(wav)}) + text = result[0].text + + json_path = wav.parent / "clip_voice_features.json" + assert str(json_path) in text + data = json.loads(json_path.read_text()) + assert len(data["words"]) == 3 + assert all("emphasis" in w for w in data["words"]) + assert all(0.0 <= w["emphasis"] <= 1.0 for w in data["words"]) + + async def test_loudest_word_scores_highest_emphasis(self, wav, patched_analysis): + from server import handle_analyze_voice_features + + await handle_analyze_voice_features({"media_path": str(wav)}) + data = json.loads((wav.parent / "clip_voice_features.json").read_text()) + by_word = {w["word"]: w["emphasis"] for w in data["words"]} + assert by_word["seguranca"] > by_word["isso"] + assert by_word["seguranca"] > by_word["e"] + + async def test_persisted_config_is_embedded_in_output(self, wav, patched_analysis): + from server import handle_analyze_voice_features + + await handle_analyze_voice_features({"media_path": str(wav)}) + data = json.loads((wav.parent / "clip_voice_features.json").read_text()) + assert "energy_threshold" in data["config"] + assert "emphasis_weights" in data["config"] + + async def test_empty_transcript_reports_instead_of_crashing(self, wav, monkeypatch): + import server_tools._shared as _shared_mod + import server_tools.voice as server_mod + from server import handle_analyze_voice_features + + monkeypatch.setattr(server_mod, "features_capability", lambda: (True, "ok")) + monkeypatch.setattr( + _shared_mod, "transcribe", lambda *a, **k: {**_FAKE_TRANSCRIPT, "words": []} + ) + result = await handle_analyze_voice_features({"media_path": str(wav)}) + assert "no words" in result[0].text.lower() + + async def test_untranscribable_media_reports_install_hint(self, wav, monkeypatch): + import server_tools._shared as _shared_mod + import server_tools.voice as server_mod + from server import handle_analyze_voice_features + + monkeypatch.setattr(server_mod, "features_capability", lambda: (True, "ok")) + monkeypatch.setattr(_shared_mod, "transcribe", lambda *a, **k: None) + result = await handle_analyze_voice_features({"media_path": str(wav)}) + assert "faster-whisper" in result[0].text + + async def test_missing_pitch_track_degrades_without_crashing(self, wav, monkeypatch): + import server_tools._shared as _shared_mod + import server_tools.voice as server_mod + from server import handle_analyze_voice_features + + monkeypatch.setattr(_shared_mod, "transcribe", lambda *a, **k: _FAKE_TRANSCRIPT) + monkeypatch.setattr(server_mod, "features_capability", lambda: (True, "ok")) + monkeypatch.setattr(server_mod, "extract_pitch", lambda *a, **k: None) + monkeypatch.setattr(server_mod, "extract_energy", lambda *a, **k: None) + result = await handle_analyze_voice_features({"media_path": str(wav)}) + assert "Voice Feature Analysis" in result[0].text + data = json.loads((wav.parent / "clip_voice_features.json").read_text()) + assert all(w["emphasis"] >= 0.0 for w in data["words"]) diff --git a/code/tests/test_voice_timeline.py b/code/tests/test_voice_timeline.py new file mode 100644 index 0000000..d8f4dd8 --- /dev/null +++ b/code/tests/test_voice_timeline.py @@ -0,0 +1,338 @@ +"""Tests for fcpxml/voice_timeline.py — the consolidated AI-readable timeline. + +The document's shape is the contract downstream consumers (rules engine, a +model reading the JSON) rely on, so these tests pin the shape as much as +the values — including that it survives every analysis layer being absent. +""" + +import json + +import pytest + +from fcpxml.voice_timeline import ( + VOICE_TIMELINE_VERSION, + build_voice_timeline, + enrich_words, + load_voice_timeline, + save_voice_timeline, + voice_timeline_path, +) + +_TRANSCRIPT = { + "language": "pt", + "duration": 4.0, + "text": "isso e seguranca total", + "segments": [ + {"text": "isso e", "start": 0.0, "end": 1.0}, + {"text": "seguranca total", "start": 2.0, "end": 4.0}, + ], + "words": [ + {"word": "isso", "start": 0.0, "end": 0.4, "confidence": 0.9}, + {"word": "e", "start": 0.5, "end": 0.7, "confidence": 0.9}, + {"word": "seguranca", "start": 2.0, "end": 2.9, "confidence": 0.9}, + {"word": "total", "start": 3.0, "end": 3.5, "confidence": 0.9}, + ], +} + +# "seguranca" is the loud, high-pitched moment +_PITCH = [(0.2, 120.0), (0.6, 118.0), (2.4, 260.0), (3.2, 130.0)] +_ENERGY = [(0.2, 0.10), (0.6, 0.12), (2.4, 0.95), (3.2, 0.20)] + + +class TestEnrichWords: + def test_normalizes_energy_against_loudest_word(self): + enriched = enrich_words(_TRANSCRIPT["words"], _PITCH, _ENERGY) + loudest = max(enriched, key=lambda w: w["energy_norm"]) + assert loudest["word"] == "seguranca" + assert loudest["energy_norm"] == pytest.approx(1.0) + + def test_all_values_stay_within_unit_range(self): + enriched = enrich_words(_TRANSCRIPT["words"], _PITCH, _ENERGY) + for w in enriched: + for key in ("energy_norm", "pitch_delta", "rate_delta", "emphasis"): + assert 0.0 <= w[key] <= 1.0, f"{key} out of range on {w['word']}" + + def test_empty_words_returns_empty(self): + assert enrich_words([], _PITCH, _ENERGY) == [] + + def test_missing_tracks_give_zero_not_crash(self): + enriched = enrich_words(_TRANSCRIPT["words"], None, None) + assert all(w["energy_norm"] == 0.0 for w in enriched) + assert all(w["pitch_delta"] == 0.0 for w in enriched) + + +class TestBuildVoiceTimeline: + @pytest.fixture + def timeline(self, monkeypatch): + import fcpxml.voice_timeline as vt + + monkeypatch.setattr(vt, "extract_pitch", lambda *a, **k: _PITCH) + monkeypatch.setattr(vt, "extract_energy", lambda *a, **k: _ENERGY) + return build_voice_timeline("/tmp/clip.wav", _TRANSCRIPT) + + def test_document_has_all_top_level_layers(self, timeline): + for key in ("version", "source", "language", "scales", "summary", "speakers", "segments"): + assert key in timeline + assert timeline["version"] == VOICE_TIMELINE_VERSION + + def test_scales_document_every_word_metric(self, timeline): + word = timeline["segments"][0]["words"][0] + for metric in timeline["scales"]["word"]: + assert metric in word, f"{metric} documented in scales but absent from words" + + def test_scales_document_every_segment_metric(self, timeline): + segment = timeline["segments"][0] + for metric in timeline["scales"]["segment"]: + assert metric in segment, f"{metric} documented in scales but absent from segments" + + def test_take_boundary_flags_a_long_gap(self, timeline): + # the fixture has a 1s gap between its two segments -> not a boundary + assert timeline["segments"][1]["gap_before"] > 0 + assert timeline["segments"][1]["take_boundary"] is False + + def test_summary_counts_match_the_detail(self, timeline): + summary = timeline["summary"] + assert summary["segment_count"] == len(timeline["segments"]) + total_words = sum(len(s["words"]) for s in timeline["segments"]) + assert summary["word_count"] == total_words + + def test_words_are_grouped_under_their_segment(self, timeline): + first, second = timeline["segments"] + assert [w["text"] for w in first["words"]] == ["isso", "e"] + assert [w["text"] for w in second["words"]] == ["seguranca", "total"] + + def test_segment_aggregates_reflect_their_words(self, timeline): + loud_segment = timeline["segments"][1] + quiet_segment = timeline["segments"][0] + assert loud_segment["avg_energy"] > quiet_segment["avg_energy"] + assert loud_segment["peak_emphasis"] >= max(w["emphasis"] for w in loud_segment["words"]) + + def test_peak_moments_are_sorted_by_emphasis(self, timeline): + peaks = timeline["summary"]["peak_moments"] + assert peaks == sorted(peaks, key=lambda m: m["emphasis"], reverse=True) + + def test_defaults_to_single_speaker_without_token(self, timeline): + assert timeline["summary"]["speaker_count"] == 1 + assert all(w["speaker"] == "SPEAKER_00" for s in timeline["segments"] for w in s["words"]) + + def test_is_json_serializable(self, timeline): + # the whole point is handing this to a model / writing it to disk + assert json.loads(json.dumps(timeline, ensure_ascii=False))["version"] + + +class TestDegradesWithoutAnalysisLayers: + def test_shape_survives_with_no_acoustics(self, monkeypatch): + import fcpxml.voice_timeline as vt + + monkeypatch.setattr(vt, "extract_pitch", lambda *a, **k: None) + monkeypatch.setattr(vt, "extract_energy", lambda *a, **k: None) + timeline = build_voice_timeline("/tmp/clip.wav", _TRANSCRIPT) + assert timeline["summary"]["word_count"] == 4 + assert timeline["summary"]["avg_emphasis"] >= 0.0 + + def test_empty_transcript_still_yields_valid_document(self, monkeypatch): + import fcpxml.voice_timeline as vt + + monkeypatch.setattr(vt, "extract_pitch", lambda *a, **k: None) + monkeypatch.setattr(vt, "extract_energy", lambda *a, **k: None) + timeline = build_voice_timeline( + "/tmp/clip.wav", {"duration": 0.0, "segments": [], "words": []} + ) + assert timeline["segments"] == [] + assert timeline["summary"]["word_count"] == 0 + assert timeline["summary"]["avg_emphasis"] == 0.0 + + def test_progress_callback_is_reported(self, monkeypatch): + import fcpxml.voice_timeline as vt + + monkeypatch.setattr(vt, "extract_pitch", lambda *a, **k: None) + monkeypatch.setattr(vt, "extract_energy", lambda *a, **k: None) + seen: list[tuple[float, str]] = [] + build_voice_timeline("/tmp/clip.wav", _TRANSCRIPT, progress_cb=lambda f, s: seen.append((f, s))) + assert seen and all(0.0 <= f <= 1.0 for f, _ in seen) + + +class TestPersistence: + def test_round_trip(self, tmp_path): + path = tmp_path / "clip_voice_timeline.json" + timeline = {"version": "1.0", "segments": [], "summary": {}} + save_voice_timeline(timeline, path) + assert load_voice_timeline(path) == timeline + + def test_accented_text_stays_readable(self, tmp_path): + path = tmp_path / "t.json" + save_voice_timeline({"segments": [{"text": "segurança"}]}, path) + assert "segurança" in path.read_text(encoding="utf-8") + + def test_missing_file_returns_none(self, tmp_path): + assert load_voice_timeline(tmp_path / "absent.json") is None + + def test_malformed_json_returns_none(self, tmp_path): + path = tmp_path / "bad.json" + path.write_text("{not json") + assert load_voice_timeline(path) is None + + def test_wrong_shape_returns_none(self, tmp_path): + path = tmp_path / "other.json" + path.write_text('{"something": "else"}') + assert load_voice_timeline(path) is None + + def test_path_next_to_media_by_default(self): + assert voice_timeline_path("/media/clip.mov").name == "clip_voice_timeline.json" + + def test_path_honours_output_dir(self, tmp_path): + path = voice_timeline_path("/media/clip.mov", output_dir=str(tmp_path)) + assert path.parent == tmp_path + + +_RAW_WORDS = [ + # a loud outlier that will be cut, plus quieter material that survives + {"text": "GRITO", "start": 1.0, "end": 1.5, "speaker": "SPEAKER_00", + "energy": 0.5, "pitch_delta": 0.5, "rate_delta": 0.0, "pause_before": 0.0, + "emphasis": 0.5, "energy_raw": 1.0, "pitch_hz": 300.0}, + {"text": "mastopexia", "start": 10.0, "end": 10.8, "speaker": "SPEAKER_00", + "energy": 0.2, "pitch_delta": 0.1, "rate_delta": 0.0, "pause_before": 0.0, + "emphasis": 0.1, "energy_raw": 0.4, "pitch_hz": 190.0}, + {"text": "a", "start": 11.0, "end": 11.1, "speaker": "SPEAKER_00", + "energy": 0.15, "pitch_delta": 0.05, "rate_delta": 0.0, "pause_before": 0.0, + "emphasis": 0.08, "energy_raw": 0.3, "pitch_hz": 185.0}, +] + +_RESTRICT_TIMELINE = { + "version": "1.0", "source": "x.mp4", "speakers": [], + "segments": [ + {"start": 1.0, "end": 1.5, "speaker": "SPEAKER_00", "text": "GRITO", + "gap_before": 0.0, "take_boundary": False, "avg_energy": 0.5, + "peak_emphasis": 0.5, "words": [_RAW_WORDS[0]]}, + {"start": 10.0, "end": 11.1, "speaker": "SPEAKER_00", + "text": "mastopexia a", "gap_before": 8.5, "take_boundary": True, + "avg_energy": 0.17, "peak_emphasis": 0.1, "words": _RAW_WORDS[1:]}, + ], +} + + +class TestRestrictToKept: + """Emphasis is relative. Cut the loudest moment out and everything left + is still scored against something the viewer will never see, so the + surviving material has to be re-normalized on its own.""" + + def test_cut_words_are_dropped(self): + from fcpxml.voice_timeline import restrict_to_kept + + r = restrict_to_kept(_RESTRICT_TIMELINE, [(0.0, 5.0)]) + texts = [w["text"] for s in r["segments"] for w in s["words"]] + assert "GRITO" not in texts + assert "mastopexia" in texts + + def test_survivors_are_rescored_against_each_other(self): + from fcpxml.voice_timeline import restrict_to_kept + + r = restrict_to_kept(_RESTRICT_TIMELINE, [(0.0, 5.0)]) + word = next(w for s in r["segments"] for w in s["words"] if w["text"] == "mastopexia") + # was 0.2 against the shout's 1.0; alone it becomes the loudest + assert word["energy"] == pytest.approx(1.0) + + def test_empty_segments_are_removed(self): + from fcpxml.voice_timeline import restrict_to_kept + + r = restrict_to_kept(_RESTRICT_TIMELINE, [(0.0, 5.0)]) + assert len(r["segments"]) == 1 + + def test_no_cuts_keeps_everything(self): + from fcpxml.voice_timeline import restrict_to_kept + + r = restrict_to_kept(_RESTRICT_TIMELINE, []) + assert sum(len(s["words"]) for s in r["segments"]) == 3 + + def test_times_stay_in_original_source_seconds(self): + from fcpxml.voice_timeline import restrict_to_kept + + r = restrict_to_kept(_RESTRICT_TIMELINE, [(0.0, 5.0)]) + assert r["segments"][0]["start"] == 10.0 + + +class TestSuggestZoomWindows: + def test_skips_function_words(self): + from fcpxml.voice_timeline import restrict_to_kept, suggest_zoom_windows + + r = restrict_to_kept(_RESTRICT_TIMELINE, [(0.0, 5.0)]) + zooms = suggest_zoom_windows(r) + assert zooms and all(z["word"] != "a" for z in zooms) + + def test_window_runs_from_the_word_to_the_end_of_its_line(self): + from fcpxml.voice_timeline import restrict_to_kept, suggest_zoom_windows + + r = restrict_to_kept(_RESTRICT_TIMELINE, [(0.0, 5.0)]) + z = suggest_zoom_windows(r)[0] + assert z["start"] == 10.0 and z["end"] == 11.1 + + def test_min_gap_keeps_zooms_apart(self): + from fcpxml.voice_timeline import suggest_zoom_windows + + timeline = {"segments": [ + {"start": t, "end": t + 1.0, "text": "linha", + "words": [{"text": "palavra", "start": t, "end": t + 0.5, "emphasis": 0.5 - i * 0.01}]} + for i, t in enumerate([0.0, 1.0, 2.0, 30.0]) + ]} + zooms = suggest_zoom_windows(timeline, min_gap=8.0) + assert len(zooms) == 2 + + def test_max_zooms_caps_the_result(self): + from fcpxml.voice_timeline import suggest_zoom_windows + + timeline = {"segments": [ + {"start": t, "end": t + 1.0, "text": "linha", + "words": [{"text": "palavra", "start": t, "end": t + 0.5, "emphasis": 0.5}]} + for t in [0.0, 20.0, 40.0, 60.0] + ]} + assert len(suggest_zoom_windows(timeline, min_gap=8.0, max_zooms=2)) == 2 + + def test_results_are_in_chronological_order(self): + from fcpxml.voice_timeline import suggest_zoom_windows + + timeline = {"segments": [ + {"start": t, "end": t + 1.0, "text": "linha", + "words": [{"text": "palavra", "start": t, "end": t + 0.5, "emphasis": e}]} + for t, e in [(60.0, 0.9), (0.0, 0.5), (30.0, 0.7)] + ]} + zooms = suggest_zoom_windows(timeline, min_gap=8.0) + assert [z["start"] for z in zooms] == sorted(z["start"] for z in zooms) + + +class TestSentenceEnd: + """Transcription segments break on breath, not grammar — a sentence + routinely spans several. A zoom ending on a segment boundary releases + mid-thought, which is what makes a punch-in feel arbitrary.""" + + SEGS = [ + {"start": 0.0, "end": 5.0, "text": "Aquela mama com um formato, que dá aquele ar", + "take_boundary": False}, + {"start": 5.0, "end": 10.7, "text": "de elegância, isso é desejo de muitas mulheres, né?", + "take_boundary": False}, + {"start": 11.0, "end": 14.0, "text": "Com o tempo, o corpo muda.", "take_boundary": False}, + ] + + def test_extends_past_a_segment_that_does_not_end_a_sentence(self): + from fcpxml.voice_timeline import sentence_end + + assert sentence_end(self.SEGS, 0) == 10.7 + + def test_stops_at_terminal_punctuation(self): + from fcpxml.voice_timeline import sentence_end + + assert sentence_end(self.SEGS, 2) == 14.0 + + def test_never_runs_past_a_take_boundary(self): + from fcpxml.voice_timeline import sentence_end + + segs = [ + {"start": 0.0, "end": 5.0, "text": "frase sem fim", "take_boundary": False}, + {"start": 12.0, "end": 15.0, "text": "outra tomada", "take_boundary": True}, + ] + assert sentence_end(segs, 0) == 5.0 + + def test_last_segment_without_punctuation_ends_at_itself(self): + from fcpxml.voice_timeline import sentence_end + + segs = [{"start": 0.0, "end": 4.0, "text": "sem ponto final", "take_boundary": False}] + assert sentence_end(segs, 0) == 4.0 diff --git a/code/tests/test_voice_timeline_tool.py b/code/tests/test_voice_timeline_tool.py new file mode 100644 index 0000000..c89ad31 --- /dev/null +++ b/code/tests/test_voice_timeline_tool.py @@ -0,0 +1,107 @@ +"""Tests for the build_voice_timeline MCP tool.""" + +import json + +import pytest + +from tests.test_voice_features_tool import _write_silent_wav + +_TRANSCRIPT = { + "language": "pt", + "duration": 4.0, + "text": "isso e seguranca total", + "segments": [ + {"text": "isso e", "start": 0.0, "end": 1.0}, + {"text": "seguranca total", "start": 2.0, "end": 4.0}, + ], + "words": [ + {"word": "isso", "start": 0.0, "end": 0.4, "confidence": 0.9}, + {"word": "e", "start": 0.5, "end": 0.7, "confidence": 0.9}, + {"word": "seguranca", "start": 2.0, "end": 2.9, "confidence": 0.9}, + {"word": "total", "start": 3.0, "end": 3.5, "confidence": 0.9}, + ], +} + + +@pytest.fixture +def wav(tmp_path): + path = tmp_path / "clip.wav" + _write_silent_wav(str(path)) + return path + + +@pytest.fixture +def patched(monkeypatch): + """Deterministic transcription + acoustics, no optional extras needed.""" + import fcpxml.voice_timeline as vt + import server_tools._shared as _shared_mod + + monkeypatch.setattr(_shared_mod, "transcribe", lambda *a, **k: _TRANSCRIPT) + monkeypatch.setattr(vt, "extract_pitch", lambda *a, **k: [(2.4, 260.0), (0.2, 120.0)]) + monkeypatch.setattr(vt, "extract_energy", lambda *a, **k: [(2.4, 0.95), (0.2, 0.10)]) + + +class TestBuildVoiceTimelineHandler: + async def test_writes_timeline_json(self, wav, patched): + from server import handle_build_voice_timeline + + result = await handle_build_voice_timeline({"media_path": str(wav)}) + text = result[0].text + + json_path = wav.parent / "clip_voice_timeline.json" + assert str(json_path) in text + data = json.loads(json_path.read_text(encoding="utf-8")) + assert data["summary"]["word_count"] == 4 + assert len(data["segments"]) == 2 + + async def test_reports_which_layers_ran(self, wav, patched): + from server import handle_build_voice_timeline + + result = await handle_build_voice_timeline({"media_path": str(wav)}) + text = result[0].text + assert "Analysis Layers" in text + assert "Transcript" in text + + async def test_rejects_disallowed_extension(self, tmp_path): + from server import handle_build_voice_timeline + + bad = tmp_path / "clip.txt" + bad.write_text("not audio") + with pytest.raises(ValueError): + await handle_build_voice_timeline({"media_path": str(bad)}) + + async def test_untranscribable_media_reports_hint(self, wav, monkeypatch): + import server_tools._shared as _shared_mod + from server import handle_build_voice_timeline + + monkeypatch.setattr(_shared_mod, "transcribe", lambda *a, **k: None) + result = await handle_build_voice_timeline({"media_path": str(wav)}) + assert "faster-whisper" in result[0].text + + async def test_uses_persisted_peak_settings(self, wav, patched, monkeypatch): + """A wider percentile must surface more peak moments.""" + import server_tools.voice as server_mod + from server import handle_build_voice_timeline + + def config(percentile): + return { + "energy_threshold": 0.5, + "peak_percentile": percentile, + "emphasis_floor": 0.0, + "emphasis_weights": { + "energy": 0.30, "pitch_variation": 0.25, "rate_variation": 0.20, + "pause_before": 0.15, "duration": 0.10, + }, + "emotion_enabled": False, + "emotion_sensitivity": 0.5, + } + + monkeypatch.setattr(server_mod, "load_voice_analysis_config", lambda: config(0.01)) + await handle_build_voice_timeline({"media_path": str(wav)}) + strict = json.loads((wav.parent / "clip_voice_timeline.json").read_text()) + + monkeypatch.setattr(server_mod, "load_voice_analysis_config", lambda: config(1.0)) + await handle_build_voice_timeline({"media_path": str(wav)}) + loose = json.loads((wav.parent / "clip_voice_timeline.json").read_text()) + + assert loose["summary"]["peak_count"] > strict["summary"]["peak_count"] diff --git a/code/tests/test_writer.py b/code/tests/test_writer.py index 55380d4..f13255b 100755 --- a/code/tests/test_writer.py +++ b/code/tests/test_writer.py @@ -652,6 +652,7 @@ def test_add_zoom_creates_keyframed_transform(temp_fcpxml): """4 keyframes: 100% -> scale -> scale -> 100%, all within [start, end].""" modifier = FCPXMLModifier(temp_fcpxml) clip = modifier.add_zoom(clip_id='Broll_Studio', start=1.0, end=3.0, scale=1.3, ease=0.5) + frame = float(modifier.frame_duration_fraction()) transform = clip.find('adjust-transform') assert transform is not None @@ -660,9 +661,23 @@ def test_add_zoom_creates_keyframed_transform(temp_fcpxml): keyframes = param.find('keyframeAnimation').findall('keyframe') assert len(keyframes) == 4 assert [kf.get('value') for kf in keyframes] == ['1 1', '1.3 1.3', '1.3 1.3', '1 1'] - assert all(kf.get('interp') == 'ease' for kf in keyframes) + # Bare keyframes — only time and value — matching a zoom exported from + # FCP itself. It rejects 'interp' on this vector param (discarding the + # whole <param>), and its own export writes no 'curve' either. + assert not any(kf.get('interp') for kf in keyframes) + assert not any(kf.get('curve') for kf in keyframes) + assert all(set(kf.attrib) == {'time', 'value'} for kf in keyframes) + # Keyframe times are anchored in the clip's SOURCE timebase (its own + # `start`), not clip-relative. This fixture starts at 10s, so a zoom over + # clip seconds 1-3 must be written at 11-13s. Writing 1-3s here would put + # the animation outside the clip and FCP imports it as nothing. + origin = modifier._parse_time(clip.get('start', '0s')).to_seconds() times = [modifier._parse_time(kf.get('time')).to_seconds() for kf in keyframes] - assert times == pytest.approx([1.0, 1.5, 2.5, 3.0], abs=0.05) + assert origin > 0, "fixture must start off zero, or this asserts nothing" + # in over 0.5s, hold, then snap back on the very next frame + assert times == pytest.approx( + [origin + 1.0, origin + 1.5, origin + 3.0 - frame, origin + 3.0], abs=0.05 + ) assert times == sorted(times) @@ -699,8 +714,9 @@ def test_add_zoom_window_outside_clip_duration_raises(temp_fcpxml): def test_add_zoom_ease_too_long_for_window_raises(temp_fcpxml): modifier = FCPXMLModifier(temp_fcpxml) - with pytest.raises(ValueError, match="doesn't fit"): - modifier.add_zoom(clip_id='Broll_Studio', start=0.0, end=1.0, ease=1.0) + # start away from the clip head so the ramp-in is actually written + with pytest.raises(ValueError, match="don't fit"): + modifier.add_zoom(clip_id='Broll_Studio', start=1.0, end=2.0, ease=1.5) def test_change_speed_twice_no_duplicate_elements(temp_fcpxml): @@ -2067,3 +2083,258 @@ def test_remove_trailing_gaps_noop_without_gap(): assert len(children) == 1 assert children[0].tag == 'asset-clip' Path(f.name).unlink(missing_ok=True) + + +class TestNTSCFrameAlignment: + """23.976/29.97 timebases must not be reported as misaligned. + + Regression: the check used int(fps), so an exactly frame-aligned NTSC + duration (a whole multiple of 1001/24000s) was flagged as broken — + every NTSC project produced spurious warnings that buried real ones. + """ + + NTSC_DOC = """<?xml version="1.0" encoding="UTF-8"?> +<fcpxml version="1.13"> + <resources> + <format id="r1" name="FFVideoFormat1080p2398" frameDuration="1001/24000s" width="1920" height="1080"/> + <asset id="a1" name="v" start="0s" duration="24437413/24000s" hasVideo="1" format="r1"> + <media-rep kind="original-media" src="file:///v.mov"/> + </asset> + </resources> + <library><event name="E"><project name="P"> + <sequence format="r1" duration="24437413/24000s" tcStart="0s"> + <spine><asset-clip name="v" ref="a1" offset="0s" start="0s" duration="24437413/24000s"/></spine> + </sequence> + </project></event></library> +</fcpxml> +""" + + def _issues(self, xml): + from fcpxml.safe_xml import safe_fromstring + from fcpxml.writer import validate_fcpxml + + root = safe_fromstring(xml) + return [ + i for i in (validate_fcpxml(root) or []) + if "frame" in str(i.issue_type.value) + ] + + def test_aligned_ntsc_duration_is_not_flagged(self): + # 24437413/24000s is exactly 24413 frames of 1001/24000s + assert self._issues(self.NTSC_DOC) == [] + + def test_genuinely_misaligned_duration_is_still_flagged(self): + broken = self.NTSC_DOC.replace( + '<asset-clip name="v" ref="a1" offset="0s" start="0s" duration="24437413/24000s"/>', + '<asset-clip name="v" ref="a1" offset="0s" start="0s" duration="500/24000s"/>', + ) + assert len(self._issues(broken)) == 1 + + def test_message_names_the_real_rate_not_a_rounded_one(self): + broken = self.NTSC_DOC.replace('duration="24437413/24000s"/>', 'duration="500/24000s"/>') + issues = self._issues(broken) + assert issues and "23.976fps" in issues[0].message + + +class TestZoomPreservesExistingFraming: + """A clip may already carry the editor's reframe — rotation for footage + shot sideways, position, a base scale. add_zoom used to delete it, which + on real footage brought the zoomed section back rotated.""" + + FRAMED = """<?xml version="1.0" encoding="UTF-8"?> +<fcpxml version="1.13"> + <resources> + <format id="r1" frameDuration="100/3000s" width="1920" height="1080"/> + <asset id="a1" name="v" start="0s" duration="600/30s" hasVideo="1" format="r1"> + <media-rep kind="original-media" src="file:///v.mov"/> + </asset> + </resources> + <library><event name="E"><project name="P"> + <sequence format="r1" duration="600/30s" tcStart="0s"> + <spine> + <asset-clip name="v" ref="a1" offset="0s" start="0s" duration="600/30s"> + <adjust-transform position="0.16 0.66" rotation="90.1" scale="1.77311 1.77311"/> + </asset-clip> + </spine> + </sequence> + </project></event></library> +</fcpxml> +""" + + def _zoomed(self, tmp_path, scale=1.2): + from fcpxml.writer import FCPXMLModifier + + path = tmp_path / "framed.fcpxml" + path.write_text(self.FRAMED) + m = FCPXMLModifier(str(path)) + m.add_zoom(clip_id="v", start=1.0, end=5.0, scale=scale) + return m.root.find(".//adjust-transform") + + def test_rotation_and_position_survive(self, tmp_path): + t = self._zoomed(tmp_path) + assert t.get("rotation") == "90.1" + assert t.get("position") == "0.16 0.66" + + def test_animation_rests_at_the_existing_scale(self, tmp_path): + kfs = self._zoomed(tmp_path).findall(".//keyframe") + # first and last keyframe return to the clip's own framing, not to 1 + assert kfs[0].get("value").split()[0].startswith("1.77") + assert kfs[-1].get("value").split()[0].startswith("1.77") + + def test_peak_multiplies_the_existing_scale(self, tmp_path): + kfs = self._zoomed(tmp_path, scale=2.0).findall(".//keyframe") + peak = float(kfs[1].get("value").split()[0]) + assert peak == pytest.approx(1.77311 * 2.0, rel=1e-4) + + def test_only_one_transform_remains(self, tmp_path): + from fcpxml.writer import FCPXMLModifier + + path = tmp_path / "framed.fcpxml" + path.write_text(self.FRAMED) + m = FCPXMLModifier(str(path)) + m.add_zoom(clip_id="v", start=1.0, end=5.0, scale=1.2) + m.add_zoom(clip_id="v", start=6.0, end=9.0, scale=1.4) + assert len(m.root.findall(".//adjust-transform")) == 1 + + def test_second_zoom_on_same_clip_keeps_the_real_base_scale(self, tmp_path): + """Found on real footage: two zoom actions landing on disjoint + windows of the same post-cut clip. The second add_zoom call used to + see the already-animated <param name="scale"> from the first zoom + instead of a static attribute, read that as "no framing", and + default the base to 1.0 — silently shrinking the shot back to its + unframed size for the whole clip wherever no keyframe applied, and + discarding the first zoom's animation in the process.""" + from fcpxml.writer import FCPXMLModifier + + path = tmp_path / "framed.fcpxml" + path.write_text(self.FRAMED) + m = FCPXMLModifier(str(path)) + m.add_zoom(clip_id="v", start=1.0, end=5.0, scale=1.2) + m.add_zoom(clip_id="v", start=6.0, end=9.0, scale=1.4) + kfs = m.root.findall(".//keyframe") + + # Every rest keyframe returns to the clip's real base scale, never 1.0. + rest_values = {kf.get("value") for kf in (kfs[0], kfs[3], kfs[4], kfs[-1])} + assert rest_values == {"1.77311 1.77311"} + + # Both peaks survive — the second call didn't erase the first. + peaks = sorted(float(kf.get("value").split()[0]) for kf in (kfs[1], kfs[5])) + assert peaks[0] == pytest.approx(1.77311 * 1.2, rel=1e-4) + assert peaks[1] == pytest.approx(1.77311 * 1.4, rel=1e-4) + + def test_overlapping_zoom_on_same_clip_replaces_instead_of_stacking(self, tmp_path): + """Two windows that OVERLAP mean "redo this zoom", not "add another + one" — the old keyframes are stale and all of them go, matching + test_add_zoom_replaces_existing_zoom's contract.""" + from fcpxml.writer import FCPXMLModifier + + path = tmp_path / "framed.fcpxml" + path.write_text(self.FRAMED) + m = FCPXMLModifier(str(path)) + m.add_zoom(clip_id="v", start=1.0, end=5.0, scale=1.2) + m.add_zoom(clip_id="v", start=3.0, end=7.0, scale=1.5) + values = [kf.get("value") for kf in m.root.findall(".//keyframe")] + assert not any(v.startswith("2.1277") for v in values) # 1.77311*1.2 gone + assert any(v.startswith("2.6596") for v in values) # 1.77311*1.5 present + + def test_unframed_clip_still_rests_at_one(self, tmp_path): + from fcpxml.writer import FCPXMLModifier + + path = tmp_path / "plain.fcpxml" + path.write_text(self.FRAMED.replace( + '<adjust-transform position="0.16 0.66" rotation="90.1" scale="1.77311 1.77311"/>', '')) + m = FCPXMLModifier(str(path)) + m.add_zoom(clip_id="v", start=1.0, end=5.0, scale=1.3) + kfs = m.root.findall(".//keyframe") + assert kfs[0].get("value") == "1 1" + assert kfs[1].get("value") == "1.3 1.3" + + +class TestZoomShapeIsAsymmetric: + """The editorial shape: ramp in fast to land with the emphasised word, + hold through the impact phrase, then snap back in a single frame so the + video resumes its normal framing without a drift that draws the eye.""" + + def _times(self, temp_fcpxml, **kw): + from fcpxml.writer import FCPXMLModifier + + m = FCPXMLModifier(temp_fcpxml) + # Broll_Studio is 5s long; end=3.0 keeps clear of the hold-at-cut + # margin so these exercise the ordinary return-to-framing shape. + kw.setdefault('end', 3.0) + clip = m.add_zoom(clip_id='Broll_Studio', start=1.0, **kw) + origin = m._parse_time(clip.get('start', '0s')).to_seconds() + kfs = clip.find('adjust-transform').find('param').find('keyframeAnimation') + return m, origin, [m._parse_time(k.get('time')).to_seconds() - origin + for k in kfs.findall('keyframe')] + + def test_return_takes_a_single_frame(self, temp_fcpxml): + m, _, t = self._times(temp_fcpxml, scale=1.2) + frame = float(m.frame_duration_fraction()) + assert t[3] - t[2] == pytest.approx(frame, abs=0.005) + + def test_ramp_in_is_quick(self, temp_fcpxml): + """Fast enough to land with the emphasised word rather than drift.""" + _, _, t = self._times(temp_fcpxml, scale=1.2) + assert t[1] - t[0] == pytest.approx(0.25, abs=0.05) + + def test_peak_is_held_until_the_return(self, temp_fcpxml): + _, _, t = self._times(temp_fcpxml, scale=1.2) + # hold spans from the top of the ramp to one frame before the end + assert t[2] - t[1] > 1.4 + + def test_zoom_opening_at_a_cut_starts_already_zoomed(self, temp_fcpxml): + """The cut is the transition — ramping up from it reads as the shot + settling rather than as emphasis.""" + from fcpxml.writer import FCPXMLModifier + + m = FCPXMLModifier(temp_fcpxml) + clip = m.add_zoom(clip_id='Broll_Studio', start=0.1, end=3.0, scale=1.2) + kfs = clip.find('adjust-transform').find('param').find('keyframeAnimation') + values = [k.get('value') for k in kfs.findall('keyframe')] + assert values[0] != values[-1] # opens zoomed, returns to framing + assert values[0] == values[-2] # ...and was at the peak from frame one + + def test_opening_at_peak_can_be_forced_off(self, temp_fcpxml): + from fcpxml.writer import FCPXMLModifier + + m = FCPXMLModifier(temp_fcpxml) + clip = m.add_zoom( + clip_id='Broll_Studio', start=0.1, end=3.0, scale=1.2, start_at_peak=False + ) + kfs = clip.find('adjust-transform').find('param').find('keyframeAnimation') + assert len(kfs.findall('keyframe')) == 4 + + def test_ease_out_can_be_made_gradual(self, temp_fcpxml): + _, _, t = self._times(temp_fcpxml, scale=1.2, ease_out=1.0) + assert t[3] - t[2] == pytest.approx(1.0, abs=0.05) + + def test_zoom_reaching_the_cut_holds_instead_of_returning(self, temp_fcpxml): + """Returning right before a cut is wasted motion — the next clip + opens on its own framing, so the move back reads as a twitch.""" + _, _, t = self._times(temp_fcpxml, scale=1.2, end=5.0) + assert len(t) == 3 # rest, peak, still peak at the cut + + def test_hold_can_be_forced_off_at_a_cut(self, temp_fcpxml): + _, _, t = self._times(temp_fcpxml, scale=1.2, end=5.0, hold_at_end=False) + assert len(t) == 4 + + def test_hold_can_be_forced_on_mid_clip(self, temp_fcpxml): + _, _, t = self._times(temp_fcpxml, scale=1.2, end=3.0, hold_at_end=True) + assert len(t) == 3 + + def test_held_zoom_stays_at_the_peak(self, temp_fcpxml): + from fcpxml.writer import FCPXMLModifier + + m = FCPXMLModifier(temp_fcpxml) + clip = m.add_zoom(clip_id='Broll_Studio', start=1.0, end=5.0, scale=1.2) + kfs = clip.find('adjust-transform').find('param').find('keyframeAnimation') + values = [k.get('value') for k in kfs.findall('keyframe')] + assert values[-1] == values[-2] != values[0] + + def test_window_too_short_for_the_ramp_is_rejected(self, temp_fcpxml): + from fcpxml.writer import FCPXMLModifier + + m = FCPXMLModifier(temp_fcpxml) + with pytest.raises(ValueError, match="don't fit"): + m.add_zoom(clip_id='Broll_Studio', start=1.0, end=1.2, scale=1.2, ease=1.5)