Compare commits

9 Commits
Author SHA1 Message Date
João HenriqueandClaude Sonnet 5 7b5aed79ee feat(voz): legenda por ênfase, forced align, IA local e correções de zoom/revisão
Trabalho da branch feat/revisao-enfases: pipeline de edição por voz ganha
alinhamento forçado (whisperx), roteirização por LLM local (Ollama), e a
etapa 5 (revisão de frases) passa a refletir de verdade o que é aplicado.

- generate_subtitles_by_emphasis: legenda comum cobre o clipe inteiro,
  legenda dinâmica só nas frases de ênfase, e a comum é desativada
  (enabled="0") onde a dinâmica cobre, em vez de nunca ser gerada ali.
- validate_subtitle_layout ignora títulos com enabled="0" — corrige falso
  positivo de colisão contra o que está desativado no lugar dele.
- Corrige zoom/marcador sendo descartado quando a borda encosta exatamente
  no início de um corte.
- Etapa 5 do Assistente: recarrega quando as decisões da IA mudam (com
  fresh=true, ignorando a revisão salva antiga) — resolve a dessincronia
  entre "ativa" na tela e o que já foi cortado no FCPXML.
- Etapa "Processar" reaplica as decisões da revisão (_phrase_actions.json)
  antes da cadeia de remoção de silêncio/legendas — antes, desativar uma
  frase na etapa 5 não tinha efeito nenhum no vídeo final.
- Etapa "Concluído" fundida em "Processar" — abrir no Final Cut/Finder
  aparece assim que termina, sem slide extra.
- Palavra clicável na etapa 5 agora funciona como toggle (clique de novo
  desfaz) e mostra a própria ênfase (sublinhado colorido + peso da fonte).
- fcpxml/forced_align.py, fcpxml/llm_local.py, ai_edit.py: alinhamento
  fonético via whisperx e roteirização local via Ollama/Gemma.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-21 18:26:04 -04:00
João HenriqueandClaude Opus 5 711c397dfe fix: admin/api apontava para admin/code (inexistente) — crash no app
Ao dividir _shared.py em admin/api/*.py ontem, o cálculo
`Path(__file__).resolve().parent.parent / "code"` foi copiado sem ajustar
para o nível de diretório novo. No arquivo original (admin/models_api.py,
direto em admin/) dois `.parent` chegavam na raiz do repo. Em
admin/api/shared.py, um nível mais fundo, dois `.parent` param em admin/ —
e admin/code nunca existiu. sys.path nunca recebia code/, então toda ação
que passa por `server` (analisar voz, aplicar decisões) crashava o app com
ModuleNotFoundError: server_tools.

O bug sobreviveu a duas rodadas de validação da sessão anterior — lint
zero, 1454 testes verdes, comando testado manualmente pela ponte — porque
todos rodam num venv com install editável (__editable__.fcp_mcp_server.pth)
que já deixa fcpxml/server_tools importáveis por conta própria, mascarando
qualquer erro no cálculo manual de sys.path. Só o app real, no fallback sem
uv, expõe o bug.

Correção: o cálculo de sys.path sai de cada módulo de comando (estava
duplicado em nove arquivos) e passa a existir uma única vez em
admin/api/__init__.py, que roda antes de qualquer submódulo — nenhum
precisa mais da própria cópia.

O teste de regressão precisou de duas tentativas pelo mesmo motivo do bug:
a primeira versão também passava com o bug presente, por rodar no mesmo
venv "de sorte". Só ficou confiável isolando um subprocess que remove
site-packages do sys.path antes de importar — confirmado nos dois sentidos,
falha com o bug reintroduzido e passa com a correção
(TestCodeDirResolution).

Detalhe completo, incluindo por que o comando manual não pegou:
Engine/docs/05_EXPERIENCIAS.md #25.

Lint zerado, 1457 testes passando (3 novos), app compilado.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-20 09:50:04 -04:00
João HenriqueandClaude Opus 5 cbd9297751 docs(skill): alinhar editar-por-voz com a revisão humana da etapa 5
A skill decidia a edição sem saber que o JSON dela agora passa por uma tela
de revisão antes de virar FCPXML. Isso não é detalhe de fluxo: a etapa 5
traduz cada ação para o vocabulário dela, e sem conhecer essa tradução a
intenção da IA se perde no caminho — que é exatamente como uma decisão vira
"arbitrária" aos olhos de quem revisa.

Novo criterios/10-revisao-humana.md, com o que o app faz com cada ação:

- cut cobrindo >=60% da frase remove a linha; tocando só uma borda vira trim
  encaixado na fronteira de palavra. Corte de meia frase é ambíguo — passa
  do limiar e apaga a linha toda quando a intenção era aparar a hesitação.
- zoom ou text sobre uma frase marca ênfase, e ênfase significa DUAS coisas:
  zoom mais legenda dinâmica; as demais frases ficam com legenda comum. A
  escala vira o nível (1.15→leve, 1.3→média, 1.5→forte).
- sem ação, o nível é derivado do peak_emphasis; a decisão da IA sempre ganha.
- reason é exibido ao lado da frase na tela — é o que o editor lê antes de
  manter ou desfazer. Deixou de ser campo de log.

Consequência prática que faltava em 05-zoom.md: não espalhar zoom "por
segurança", porque cada um promove a frase em duas dimensões ao mesmo tempo.
Na dúvida, deixar sem — promover custa uma tecla, despromover custa mais.

Cada arquivo de critério ganhou cabeçalho de escopo (o que cobre, em que
fase), no mesmo padrão dos docs do Engine, para ler só o necessário.

Todas as afirmações numéricas do novo critério foram verificadas contra
fcpxml/phrase_review.py rodando, não assumidas.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-19 22:54:51 -04:00
João HenriqueandClaude Opus 5 dcdd73edb5 docs: varredura geral, documentação por função e regra de atualização
A documentação descrevia um sistema que não existe mais: 62/73 ferramentas
(são 74), writer.py e models.py como arquivos (viraram pacotes), 1032 testes
(são 1454), models_api.py descrito como "API FastAPI" (é ponte JSON) e o app
SwiftUI ausente por completo — 5.500 linhas que o usuário opera todo dia sem
uma linha de documentação.

Cada arquivo passa a ter uma função específica, com cabeçalho de escopo
dizendo o que cobre e o que NÃO cobre (com a seta para quem cobre). O objetivo
é ler só o necessário: doc fora do assunto custa tempo e processamento sem
entregar nada.

    01 arquitetura   camadas, duas portas de entrada, regras transversais
    02 módulos       mapa do engine, incluindo o pipeline de voz
    03 server/tools  as 74 tools, helpers e como criar uma nova
    08 app macOS     NOVO — build por swiftc, telas, ponte, etapa 5
    09 manutenção    NOVO — por onde começar, o que está aberto, sintoma→arquivo

CLAUDE.md ganha a seção "Documentação (MANDATORY)": tabela de roteamento
(qual arquivo abrir para cada tarefa) e a regra de que toda alteração de
código atualiza a doc no mesmo commit, com o mapa de o-que-mexeu → o-que-
atualizar. Doc velha engana mais que doc ausente.

O índice do 05_EXPERIENCIAS subiu para o topo: consultar "isso já quebrou
antes?" custava carregar 1.281 linhas antes de chegar na tabela.

Dívidas levantadas na varredura e registradas em 09 §2: etapa 6 ainda ignora
o phrase_review.json, offset de ~400ms do Whisper, MacApp sem teste, admin/
fora do lint, confirmações visuais pendentes no FCP, submódulo WHISPERX sujo.

Também corrigidos dois links quebrados no Engine/README que apontavam um
nível acima do certo.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-19 22:51:33 -04:00
João HenriqueandClaude Opus 5 ffaebb3f72 refactor: _shared.py vira subpacote, um módulo por papel
Eram 882 linhas de seis papéis sem relação, sob um nome que só dizia
"compartilhado" — o depósito onde tudo que servia a mais de um handler
acabava caindo.

    media       316   transcrição em cache, corte por fala, relatório
    paths       206   sandbox, limites, caminho de saída
    project     116   abrir projeto, preparar modifier/generator
    captions    112   SRT, VTT, listas com timestamp
    detection    99   flash frames, buracos, duplicados
    formatting   86   tabelas e relatórios dos handlers

O __init__ reexporta os 46 nomes, então os treze pontos que importam daqui
não mudaram.

_transcript_cut_report saiu de formatting para media: ele precisa do hint de
instalação e do _text_result, ou seja, é relatório de transcrição e não
formatação genérica — mover foi mais honesto que cruzar imports entre os
dois módulos.

Quatro testes patchavam `server_tools._shared.transcribe`; o nome agora é
ligado por _shared/media.py, então o patch passou a apontar para lá — mesmo
padrão da experiência #23.

Lint zerado, 1454 testes passando.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-19 22:35:02 -04:00
João HenriqueandClaude Opus 5 368bb62706 refactor: models.py vira pacote, um módulo por família
Eram 1.091 linhas com seis famílias de modelo sem relação entre si —
enumerações, tempo racional, timeline, geração, QC e legendas.

    timing     304   TimeValue e Timecode
    timeline   217   clipes, marcadores, lanes, projeto
    enums      183   tipos/cores de marcador, transições, ritmo
    subtitles  157   paleta e look das legendas dinâmicas
    qc         121   achados de QC e resultado de validação
    planning    93   rough cut, ritmo, montagem

O __init__ reexporta os 43 nomes, incluindo os com underscore que o writer
e a suíte já importavam, então nenhum ponto de uso mudou.

Lint zerado, 1454 testes passando.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-19 22:31:18 -04:00
João HenriqueandClaude Opus 5 6090e229e9 refactor: models_api.py vira ponto de entrada sobre admin/api/
A ponte JSON do app tinha 1.395 linhas e 37 comandos de oito assuntos
diferentes num arquivo só. Agora models_api.py guarda apenas a referência
dos comandos, a tabela de despacho e o main(); cada assunto virou um módulo
em admin/api/ (models, project, editing, zoom, subtitles, transcription,
voice, review), com a base comum em shared.py.

Nada muda para o app: ele continua chamando admin/models_api.py por caminho,
e os 37 comandos respondem igual — verificado rodando a ponte de verdade.

Duas coisas que a divisão obrigou a arrumar:

- A saída passa por `shared.emit` chamada pelo módulo, não pelo nome
  importado. Isso preserva a propriedade de que trocar `emit` num lugar só
  captura a saída de todos os comandos — que era acidental quando tudo
  morava no mesmo arquivo, e vira intencional agora.
- `_CANCEL` e o lock eram globais compartilhados. O registro de downloads
  foi para models.py, junto de quem o usa, com lock próprio: o antigo
  protegia ao mesmo tempo o dicionário e a escrita em stdout, duas coisas
  sem relação.

Também: admin/test_models_api.py estava fora de `testpaths` e nunca rodava.
Movido para code/tests/ e ligado ao gate — 1441 → 1454 testes
(ver Engine/docs/05_EXPERIENCIAS.md #24).

Lint zerado, 1454 testes passando.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-19 21:46:47 -04:00
João HenriqueandClaude Opus 5 4f5cf94443 refactor: writer.py vira pacote, um módulo por assunto
O writer tinha 4.199 linhas, das quais 3.300 numa única classe com dezoito
assuntos dentro. Achar o trecho de zoom exigia rolar por marcadores,
velocidade e legendas.

Agora é o pacote fcpxml/writer/, com um arquivo por assunto e o
FCPXMLModifier montado por composição de mixins. Mixins, e não objetos
separados, porque todas essas operações mexem no mesmo documento e nos
mesmos índices — separá-las em objetos independentes transformaria toda
chamada interna em travessia de fronteira sem nada em troca. A divisão que
importa aqui é de leitura, não de estado.

Nenhuma mudança de comportamento e nenhuma alteração nos ~50 pontos que
importam do writer: o __init__ re-exporta tudo, inclusive os nomes com
underscore que a suíte já usava.

    core      723   carga, índices, navegação na spine, save
    titles    600   títulos e legendas dinâmicas
    cut       333   dividir, cortar faixas, apagar
    speed     297   velocidade e zoom
    (+ 20 módulos menores)

Único ajuste de chamada: quatro testes faziam patch em
fcpxml.writer.subprocess, que agora mora em writer.document (ver
Engine/docs/05_EXPERIENCIAS.md #23).

Lint zerado, 1441 testes passando.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-19 21:38:49 -04:00
João HenriqueandClaude Opus 5 1bebee4359 feat: etapa 5 do assistente — revisão de ênfases com timeline
Transforma a etapa "colar decisões" numa tela de lapidação: a sugestão da
IA chega carregada e o editor afina frase a frase o que é ênfase e o que
fica fora. Essa marcação é o norte da etapa 6 — só as frases com ênfase
recebem zoom e legenda dinâmica; as demais ficam com legenda comum.

O campo de colar o JSON sobe para a etapa 4, então a numeração das etapas
não muda e a etapa 6 segue intacta.

Backend (fcpxml/phrase_review.py):
- build_phrase_review funde o _voice_timeline.json com as actions da IA
- trim por frase que anda em fronteira de palavra; corte parcial da IA
  chega como trim em vez de ser arredondado fora
- phrase_review_to_actions volta a cuts/zooms + emphasis_spans
- merge_saved_decisions reaplica só as decisões salvas sobre uma revisão
  remontada da análise atual, para reprocessar a voz não ficar mascarado
- resolve_source acha a mídia: o voice timeline guarda só o nome do arquivo

App (SwiftUI):
- layout de sala de edição: preview em cima, inspector à direita, timeline
  atravessando embaixo com seis trilhas rotuladas
- preview enquadra no formato de entrega lido do .fcpxml (fonte horizontal,
  projeto vertical), com alternância para a mídia original
- reprodução pula os trechos removidos e para no fim do trecho
- zoom manual por trecho marcado, sem guardar escala: a forma vem das
  configurações de Análise de Voz no render
- emoção da fala exposta por frase

Correções encontradas no caminho:
- VideoPlayer (AVKit) aborta em runtime no app compilado por swiftc;
  trocado por AVPlayerLayer (ver Engine/docs/05_EXPERIENCIAS.md #22)
- teste que ainda afirmava o default zoom scale=1.3 removido do parser (#21)

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-19 21:29:27 -04:00
112 changed files with 15549 additions and 7964 deletions
+25 -2
View File
@@ -28,6 +28,22 @@ minutos.
`apply_voice_actions` aplica direto, mas é para **teste**. O produto do seu `apply_voice_actions` aplica direto, mas é para **teste**. O produto do seu
trabalho é a lista de decisões. trabalho é a lista de decisões.
**Caminho automatizado (sem wizard, sem copiar-e-colar):** a tool
`generate_voice_script` (MCP) / comando `generate_voice_script` (ponte do app)
corre o fluxo fechado: transcreve → `build_voice_timeline` → entrega a timeline
a um **modelo local Ollama (Gemma 3 / Llama)** que age exatamente como este
skill descreve (separa roteiro de bastidor, escolhe tomadas, decide zoom/corte)
→ devolve o roteiro legível **e** o JSON de ações, e opcionalmente aplica no
FCPXML. O cliente fica em `code/fcpxml/llm_local.py`; o prompt que embute este
contrato está em `_SYSTEM_PROMPT`. Use essa tool quando o usuário pedir para
"rodar tudo internamente" ou "gerar o roteiro por IA local".
**Para onde ela vai (modo manual):** o usuário cola o seu JSON no app, e ele
abre na etapa 5 do Assistente — uma tela onde cada frase do roteiro aparece com
a sua decisão já marcada, para ser revisada antes de gerar. Você é o **ponto de
partida** da edição, não a palavra final; escreva decisões defensáveis e motivos
legíveis. Como o app traduz cada ação sua: `criterios/10-revisao-humana.md`.
## Ordem de trabalho ## Ordem de trabalho
Siga nesta ordem. Pular a Fase 2 ou a 3 leva a decisões erradas. Siga nesta ordem. Pular a Fase 2 ou a 3 leva a decisões erradas.
@@ -44,6 +60,9 @@ Siga nesta ordem. Pular a Fase 2 ou a 3 leva a decisões erradas.
| **7** | Cortar a lista pelo ritmo | `criterios/07-ritmo.md` | | **7** | Cortar a lista pelo ritmo | `criterios/07-ritmo.md` |
| **8** | Montar o JSON de saída | `criterios/08-formato-de-saida.md` | | **8** | Montar o JSON de saída | `criterios/08-formato-de-saida.md` |
**Antes da Fase 5, leia `criterios/10-revisao-humana.md`.** Ele descreve o que
o app faz com o seu JSON — e muda *como* escrever cortes e zooms, não só quais.
## As três armadilhas ## As três armadilhas
Cada uma já causou erro silencioso em material real: Cada uma já causou erro silencioso em material real:
@@ -69,10 +88,14 @@ disso e o efeito cai no frame errado — sem erro visível.
``` ```
build_voice_timeline → [você decide] → refine_voice_timeline → [você corta build_voice_timeline → [você decide] → refine_voice_timeline → [você corta
pelo ritmo] → apply_voice_actions → remove_media_silence → pelo ritmo] → [revisão humana na etapa 5 do app] → apply_voice_actions →
generate_dynamic_subtitles remove_media_silence → generate_dynamic_subtitles
``` ```
A revisão humana entra entre a sua decisão e a aplicação. É por isso que o
`reason` importa tanto: ele é lido ali, na hora de decidir se a sua escolha
fica.
Vícios de linguagem e lacunas longas entram na **sua** lista, num ripple só Vícios de linguagem e lacunas longas entram na **sua** lista, num ripple só
(`06-texto-corte-marcador.md`). Silêncio fino e legendas vêm depois. (`06-texto-corte-marcador.md`). Silêncio fino e legendas vêm depois.
@@ -1,5 +1,8 @@
# 01 — Leitura do JSON # 01 — Leitura do JSON
> **Escopo:** Como ler o voice_timeline em camadas, sem recalcular o que já foi medido.
> **Quando:** Fase 1 — ver a ordem de trabalho em `../SKILL.md`.
O arquivo `<mídia>_voice_timeline.json` é a entrada de todo o trabalho. O arquivo `<mídia>_voice_timeline.json` é a entrada de todo o trabalho.
Leia em camadas, de cima para baixo, e só desça quando precisar. Leia em camadas, de cima para baixo, e só desça quando precisar.
@@ -50,14 +53,19 @@ use para decidir; existem para permitir a reanálise da Fase 2.
## O timestamp por palavra tem um viés conhecido ## O timestamp por palavra tem um viés conhecido
O início de cada palavra vem sistematicamente **adiantado em ~0,3-0,5s** em O início de cada palavra vem sistematicamente **adiantado em ~0,3-0,5s** em
relação ao ataque real da fala — medido em material real com ffmpeg relação ao ataque real da fala — medido em material real com ffmpeg (`astats`),
(`astats`), consistente em 6 pontos do mesmo vídeo. O fim da palavra não consistente em 6 pontos do mesmo vídeo. O fim da palavra não tem esse problema
tem esse problema (erro de poucos centésimos). Causa: `word_timestamps` do (erro de poucos centésimos). Causa: `word_timestamps` do faster-whisper deriva
faster-whisper deriva por atenção cruzada, sem alinhamento forçado — ver por atenção cruzada, sem alinhamento forçado — ver `05_EXPERIENCIAS.md`, entrada
`05_EXPERIENCIAS.md`, entrada de 2026-08-19. de 2026-08-19.
**Quando o pipeline já corrigiu isso:** se `layers.alignment` for `true`
(transcript gerado com alinhamento forçado fonético via whisperx, implementado
depois desse aviso), o viés foi removido na origem — **não aplique o offset
manual** abaixo. O aviso vale só para transcripts antigos sem `layers.alignment`.
Isso não é "reestimar no olho" — é um bug de medição na fonte, não um Isso não é "reestimar no olho" — é um bug de medição na fonte, não um
julgamento seu. Na prática: julgamento seu. Na prática (somente sem `layers.alignment`):
- Ao posicionar um `zoom` cujo `start` precisa cair exatamente na palavra - Ao posicionar um `zoom` cujo `start` precisa cair exatamente na palavra
(não uma frase inteira), some **+0,3 a +0,4s** ao timestamp do JSON antes (não uma frase inteira), some **+0,3 a +0,4s** ao timestamp do JSON antes
@@ -66,6 +74,3 @@ julgamento seu. Na prática:
- **Não aplique essa correção a `gap_before` para decidir corte** — a régua - **Não aplique essa correção a `gap_before` para decidir corte** — a régua
de silêncio (`06-texto-corte-marcador.md`) já é conservadora o bastante de silêncio (`06-texto-corte-marcador.md`) já é conservadora o bastante
para absorver esse erro; corrigir os dois ao mesmo tempo é redundante. para absorver esse erro; corrigir os dois ao mesmo tempo é redundante.
- Se um dia o pipeline ganhar alinhamento forçado (WhisperX), este aviso
perde a razão de existir — confira se `layers` ou a versão do documento
já indicam isso antes de aplicar o offset manualmente.
@@ -1,5 +1,8 @@
# 02 — Triagem: roteiro vs. conversa de bastidor # 02 — Triagem: roteiro vs. conversa de bastidor
> **Escopo:** Separar o texto do roteiro da conversa de bastidor — tarefa de texto, nunca de limiar.
> **Quando:** Fase 2 — ver a ordem de trabalho em `../SKILL.md`.
**Primeira coisa a fazer, antes de qualquer decisão de efeito.** **Primeira coisa a fazer, antes de qualquer decisão de efeito.**
Material bruto de gravação quase nunca é uma tomada só. A pessoa lê o Material bruto de gravação quase nunca é uma tomada só. A pessoa lê o
@@ -1,5 +1,8 @@
# 03 — Escolha da melhor tomada # 03 — Escolha da melhor tomada
> **Escopo:** Qual tomada de cada frase sobrevive, e o que fazer em caso de empate.
> **Quando:** Fase 3 — ver a ordem de trabalho em `../SKILL.md`.
A mesma frase costuma aparecer 2, 3, 4 vezes. Seu trabalho é ficar com A mesma frase costuma aparecer 2, 3, 4 vezes. Seu trabalho é ficar com
**uma**. **uma**.
@@ -1,5 +1,8 @@
# 04 — Reanálise do material que sobrou # 04 — Reanálise do material que sobrou
> **Escopo:** Renormalizar a ênfase sobre o que sobrou, antes de escolher zooms.
> **Quando:** Fase 4 — ver a ordem de trabalho em `../SKILL.md`.
**Não escolha zooms com os números da análise bruta.** **Não escolha zooms com os números da análise bruta.**
## O problema ## O problema
@@ -1,5 +1,8 @@
# 05 — Zoom (punch-in) # 05 — Zoom (punch-in)
> **Escopo:** Onde dar punch-in, qual janela e qual escala — e o que a escala significa além do zoom.
> **Quando:** Fase 5 — ver a ordem de trabalho em `../SKILL.md`.
## Quando usar ## Quando usar
No momento em que o argumento vira. Um pico acústico só merece zoom se for No momento em que o argumento vira. Um pico acústico só merece zoom se for
@@ -46,14 +49,24 @@ automático acerta na quase totalidade dos casos.
## Escala ## Escala
| Valor | Uso | | Valor | Uso | Vira, na tela de revisão |
|---|---| |---|---|---|
| 1,15 | sutil | | 1,15 | sutil | ênfase **1 — Leve** |
| 1,18 – 1,3 | padrão | | 1,18 – 1,3 | padrão | ênfase **2 — Média** |
| 1,5 | forte | | 1,5 | forte | ênfase **3 — Forte** |
Em vídeo institucional, fique na faixa baixa. Acima de 3,0 é rejeitado. Em vídeo institucional, fique na faixa baixa. Acima de 3,0 é rejeitado.
**A escala tem um segundo efeito, e ele é maior que o zoom.** A frase que
recebe um zoom é marcada como **ênfase** na etapa 5, e frase de ênfase recebe
**legenda dinâmica**; as demais ficam com legenda comum. Ou seja: escolher onde
dar zoom é também escolher onde o texto ganha tratamento tipográfico.
Consequência prática: **não espalhe zoom "por segurança"**. Cada um promove uma
frase a destaque em duas dimensões ao mesmo tempo. Na dúvida, deixe sem — o
editor promove numa tecla, e despromover custa mais que promover.
Detalhe: `10-revisao-humana.md`.
O zoom é **relativo ao enquadramento existente**: se o clipe já tem escala O zoom é **relativo ao enquadramento existente**: se o clipe já tem escala
1,77 (material gravado de lado e reenquadrado), um zoom 1,18 anima de 1,77 1,77 (material gravado de lado e reenquadrado), um zoom 1,18 anima de 1,77
para 2,09 e preserva rotação e posição. para 2,09 e preserva rotação e posição.
@@ -1,5 +1,8 @@
# 06 — Texto, corte e marcador # 06 — Texto, corte e marcador
> **Escopo:** Texto na tela, o que cortar (inclui muletas e lacunas) e quando marcar.
> **Quando:** Fase 6 — ver a ordem de trabalho em `../SKILL.md`.
## Texto ## Texto
Para fixar um **conceito, número ou nome** que o espectador precisa reter. Para fixar um **conceito, número ou nome** que o espectador precisa reter.
@@ -1,5 +1,8 @@
# 07 — Ritmo # 07 — Ritmo
> **Escopo:** Quantos efeitos cabem: os tetos e como escolher o que fica.
> **Quando:** Fase 7 — ver a ordem de trabalho em `../SKILL.md`.
**O erro mais comum é efeito demais.** Cansa mais que efeito de menos, e **O erro mais comum é efeito demais.** Cansa mais que efeito de menos, e
denuncia edição automática. denuncia edição automática.
@@ -1,5 +1,8 @@
# 08 — Formato de saída # 08 — Formato de saída
> **Escopo:** O JSON de entrega: estrutura, regras e como o programa trata erros.
> **Quando:** Fase 8 — ver a ordem de trabalho em `../SKILL.md`.
O produto do seu trabalho é **este JSON**. É ele que vai para o programa O produto do seu trabalho é **este JSON**. É ele que vai para o programa
gerar o FCPXML. Você nunca escreve XML. gerar o FCPXML. Você nunca escreve XML.
@@ -53,6 +56,19 @@ uma. Um `reason` vazio é sinal de decisão sem critério.
Inclua o dado que embasou: *"abertura: 'Aquela mama' (ênfase 0.42)"* é útil; Inclua o dado que embasou: *"abertura: 'Aquela mama' (ênfase 0.42)"* é útil;
*"zoom"* não é. *"zoom"* não é.
Não é campo de log: o texto é **exibido na tela de revisão**, ao lado da frase,
e é o que o editor lê antes de manter ou desfazer o que você decidiu.
### 6. Corte: alinhe à intenção
A tela lê cada `cut` contra as frases da transcrição:
- cobre **≥ 60%** de uma frase → aquela frase é **removida**;
- toca só o **começo** ou só o **fim** → vira **trim** (a frase fica, aparada).
Então corte a frase **inteira** quando quiser removê-la, e corte **só da borda
até a palavra** quando quiser aparar uma hesitação. Um corte de meia frase é
ambíguo — passa de 60% e apaga a linha toda. Detalhe: `10-revisao-humana.md`.
## Como o programa trata erros ## Como o programa trata erros
- **Ação inválida** → rejeitada e reportada **individualmente**. Uma linha - **Ação inválida** → rejeitada e reportada **individualmente**. Uma linha
@@ -1,5 +1,8 @@
# 09 — Quando a análise veio incompleta # 09 — Quando a análise veio incompleta
> **Escopo:** O que fazer quando uma camada da análise não rodou.
> **Quando:** Fase 0 — ver a ordem de trabalho em `../SKILL.md`.
O bloco `layers` no topo do JSON diz **o que de fato rodou**. Leia antes de O bloco `layers` no topo do JSON diz **o que de fato rodou**. Leia antes de
qualquer outra coisa. qualquer outra coisa.
@@ -0,0 +1,121 @@
# 10 — A revisão humana: o que acontece com o seu JSON
> **Escopo:** O que o app faz com o seu JSON na etapa 5 — muda como escrever as ações.
> **Quando:** ler antes da Fase 5 — ver a ordem de trabalho em `../SKILL.md`.
> Leia antes de decidir cortes e zooms. Muda **como** escrever as ações, não
> apenas quais.
Seu JSON não vai direto para o FCPXML. Ele é colado no app e abre na **etapa 5
do Assistente**, uma tela onde o editor vê cada frase do roteiro com a sua
decisão já aplicada e lapida antes de gerar.
Isso tem duas consequências práticas:
1. **Suas decisões são lidas por uma pessoa, frase a frase.** Uma decisão sem
motivo explícito parece arbitrária — e será desfeita.
2. **A tela traduz suas ações para o vocabulário dela.** Se você não escrever
as ações do jeito que essa tradução espera, a intenção se perde no caminho.
---
## Como cada ação sua é lida
O app quebra a gravação em **frases** (os segmentos do voice timeline) e
projeta suas ações sobre elas.
### `cut`
| O corte cobre… | Vira | Na tela |
|---|---|---|
| **≥ 60%** da frase | frase **desativada** | apagada, riscada, reativável num clique |
| só o **começo** ou só o **fim** | **trim** da frase | a frase fica, aparada nas pontas |
| um pedaço no **meio** | nada em si | só conta para a regra dos 60% |
O trim é **encaixado na fronteira de palavra** mais próxima. Você não precisa
acertar o frame: mire na palavra onde a frase deve começar ou terminar.
**O que isso pede de você:** decida se está removendo *a linha* ou *aparando*
uma ponta, e escreva o corte de acordo.
- Removendo a linha → corte a frase inteira, de ponta a ponta.
- Aparando um falso começo → corte só da borda até a palavra onde a fala
engata. Um corte que cobre meia frase é ambíguo: passa de 60% e apaga a linha
toda, quando você só queria tirar a hesitação.
### `zoom` e `text`
Qualquer `zoom` ou `text` que toque uma frase marca aquela frase como
**ênfase** — e ênfase, nesta tela, significa **duas coisas**:
> **A frase de ênfase recebe zoom E legenda dinâmica. As demais recebem
> legenda comum.**
O nível vem da sua `scale`:
| `scale` | Nível na tela | |
|---|---|---|
| 1,15 | 1 — Leve | |
| 1,3 | 2 — Média | |
| 1,5 | 3 — Forte | |
| omitida, ou uma ação `text` | 2 — Média | padrão |
Sem nenhuma ação sua, a tela deriva o nível do `peak_emphasis` da frase
(< 0,25 → sem ênfase; < 0,45 → leve; < 0,65 → média; acima → forte). **A sua
decisão sempre ganha da derivação automática.**
**O que isso pede de você:** escolher a escala com intenção. Ela não é só
"quanto amplia" — é o peso que aquela frase terá no vídeo inteiro, incluindo o
tratamento da legenda. Um zoom leve numa frase de apoio não é neutro: promove
aquela frase a destaque tipográfico também.
### `marker`
Não altera a frase. Continua sendo o seu recado para o editor conferir uma
emenda — e é a ferramenta certa quando você está em dúvida (ver
`03-escolha-da-melhor-tomada.md`).
---
## `reason` aparece na tela
Não é campo de log. O texto que você escreve em `reason` é exibido para o
editor ao lado da frase selecionada, e é o que ele lê antes de manter ou
desfazer a sua decisão.
Escreva para quem está com pressa e vai decidir na hora:
- **Bom:** `"fecho, pico em 'devolver' (ênfase 0.34) — escala mais forte por ser o fechamento da peça"`
- **Ruim:** `"zoom"` · `"corte necessário"` · `"melhor tomada"`
A regra prática: se o `reason` não contém **o dado** que embasou (a palavra, o
número, a comparação entre tomadas), você provavelmente não tinha critério —
tinha impressão.
---
## O que a tela NÃO desfaz por você
- **Tempo errado continua errado.** A tela mostra suas ações no eixo da mídia
original; se você compensou para pós-corte, tudo aparece no lugar errado e o
editor não tem como adivinhar o que você quis dizer.
- **Excesso de zoom continua excesso.** A tela não impõe o teto de 2–4 por
minuto (`07-ritmo.md`) — ela mostra o que você mandou. Efeito demais chega
ao editor como trabalho de limpeza.
- **Frase promovida a ênfase sem querer.** Como zoom e legenda dinâmica andam
juntos, espalhar zooms "de segurança" enche o vídeo de legenda dinâmica. Na
dúvida, deixe sem — o editor promove; é mais barato que despromover.
---
## Depois da revisão
O editor pode, na tela: mudar o nível de ênfase (0–3), desativar ou reativar
frases, corrigir o texto, aparar as pontas por palavra, reclassificar entre
roteiro e bastidor e acrescentar zooms manuais em trechos arbitrários.
O resultado vira um `_phrase_review.json` e o `_phrase_actions.json` derivado —
e é esse que a geração usa. **Seu JSON é o ponto de partida da conversa, não a
palavra final.** Trabalhe para ser um bom ponto de partida: decisões
defensáveis, motivos legíveis e nenhuma escolha que o editor precise desfazer
antes de começar.
+88 -15
View File
@@ -9,25 +9,91 @@ normalmente; a regra é sobre a comunicação com o usuário.
## What This Is ## What This Is
MCP server that reads/writes Final Cut Pro XML (FCPXML) files. 73 tools for timeline analysis, batch editing, QC, generation, multi-track support, media relink, NLE export, transcript-based editing (local Whisper), and LIVE FCP control (push_to_fcp / list_fcp_libraries via Apple events). Reads FCPXML 1.8–1.14 (incl. `.fcpxmld` bundles with sidecar preservation), writes 1.13 by default. Dual-mode (XML + Live) direction: `code/docs/CAPABILITY-AUDIT-2026-06.md`. MCP server that reads/writes Final Cut Pro XML (FCPXML) files. 77 tools for timeline analysis, batch editing, QC, generation, multi-track support, media relink, NLE export, transcript-based editing (local Whisper), LIVE FCP control (push_to_fcp / list_fcp_libraries via Apple events), and local-LLM voice scripting (editar-por-voz against Ollama/Gemma 3). Reads FCPXML 1.8–1.14 (incl. `.fcpxmld` bundles with sidecar preservation), writes 1.13 by default. Dual-mode (XML + Live) direction: `code/docs/CAPABILITY-AUDIT-2026-06.md`.
## Architecture ## Architecture
Toda a estrutura do projeto fica em `code/`. A pasta `admin/` fica na raiz (fora de `code/`). Toda a estrutura do projeto fica em `code/`. A pasta `admin/` fica na raiz (fora de `code/`).
``` Há **duas portas de entrada** para o mesmo engine: o MCP (Claude decide a
code/server.py — MCP server entry point. All 62 tool definitions, handlers, resources, prompts. edição) e a ponte JSON (o app macOS opera). Nenhuma das duas tem lógica de
Dispatch dict pattern: TOOL_HANDLERS maps tool names → async handler functions. timeline — as duas delegam a `fcpxml/`.
code/fcpxml/parser.py — Reads FCPXML → Python objects (Timeline, Clip, ConnectedClip, Marker, etc.)
code/fcpxml/writer.py — Writes modifications back to FCPXML. Handles markers, trimming, gaps, transitions.
code/fcpxml/rough_cut.py — Generates new timelines from source clips (rough cuts, montages, A/B rolls).
code/fcpxml/diff.py — Timeline comparison engine. Detects added/removed/moved/trimmed clips & markers.
code/fcpxml/export.py — DaVinci Resolve FCPXML v1.9 export + FCP7 XMEML v5 export for cross-NLE workflows.
code/fcpxml/models.py — Data classes: TimeValue, Timecode, Clip, ConnectedClip, CompoundClip, Timeline, etc.
code/fcpxml/media_intel.py — Real media analysis. Audio silence detection + beat detection.
code/fcpxml/dtd.py — Validates output against Apple's official DTDs.
``` ```
code/server.py — MCP entry point (592 linhas). Só dispatch: TOOL_HANDLERS.
code/server_tools/ — Os handlers das 77 tools, um módulo por categoria.
code/server_tools/_shared/ — Helpers compartilhados (paths, project, formatting,
captions, detection, media).
code/fcpxml/parser.py — Reads FCPXML → Python objects (Timeline, Clip, Marker…)
code/fcpxml/writer/ — PACOTE. Edição/escrita de FCPXML. FCPXMLModifier é
montado por mixins, um por assunto (markers, trim,
speed, titles, cut, silence…). Ver writer/modifier.py.
code/fcpxml/models/ — PACOTE. Data classes por família: timing, timeline,
enums, subtitles, qc, planning.
code/fcpxml/rough_cut.py — Generates new timelines (rough cuts, montages, A/B rolls).
code/fcpxml/diff.py — Timeline comparison engine.
code/fcpxml/export.py — DaVinci Resolve v1.9 + FCP7 XMEML v5 export.
code/fcpxml/media_intel.py — Silence detection + beat detection.
code/fcpxml/dtd.py — Validates output against Apple's official DTDs.
code/fcpxml/voice_*.py — Pipeline de voz: features → emphasis → voice_timeline
→ voice_actions → phrase_review. Ver Engine/docs/02.
admin/models_api.py — Ponte JSON com o app: docstring de comandos + dispatch.
admin/api/ — Os 37 comandos, um módulo por assunto.
code/MacApp/Sources/ — App SwiftUI. Compilado por swiftc (sem Xcode/SPM).
```
Os dois `__init__.py` de pacote (`writer/`, `models/`) reexportam tudo, então
`from .writer import FCPXMLModifier` e `from .models import TimeValue` seguem
valendo em todo o projeto.
## Documentação (MANDATORY)
A documentação viva fica em `code/Engine/docs/`. Cada arquivo tem **uma função
específica** — leia só o que a tarefa exige, não o conjunto. Carregar
documentação que não é do assunto custa tempo e processamento sem entregar nada.
### Qual arquivo abrir
| Sua tarefa | Abra | Não precisa de |
|-----------|------|----------------|
| Entender como o sistema é dividido | `01_ARCHITECTURE.md` | o resto |
| Achar onde mora uma função do engine | `02_MODULES.md` | 01, 03 |
| Criar/alterar uma ferramenta MCP | `03_SERVER_TOOLS.md` | 08 |
| Entender ou rodar os testes | `04_TESTS_AND_WORKFLOW.md` | — |
| "Isso já quebrou antes?" | `05_EXPERIENCIAS.md` — **só o índice no topo** | as entradas que não são a sua |
| Checklist antes de fechar | `06_BOAS_PRATICAS.md` | — |
| Mexer no app / no Assistente | `08_APP_MACOS.md` | 02, 03 |
| Escolher o que fazer, ver o que está aberto | `09_MANUTENCAO.md` | — |
Quando não souber por onde começar: `09_MANUTENCAO.md`. Ele roteia para o resto.
### Regra de atualização (obrigatória)
**Toda alteração de código atualiza a documentação no mesmo commit.** Doc velha
engana mais do que doc ausente — quem lê confia nela e erra com confiança.
| Você alterou | Atualize |
|--------------|----------|
| Estrutura de pastas, camadas ou dependências | `01_ARCHITECTURE.md` |
| Criou/moveu/dividiu módulo em `fcpxml/` | `02_MODULES.md` (tabela + linhas) |
| Criou/removeu ferramenta MCP | `03_SERVER_TOOLS.md` + contagem no `CLAUDE.md` |
| Comando da ponte | docstring de `admin/models_api.py` + `08_APP_MACOS.md` |
| Tela ou fluxo do app | `08_APP_MACOS.md` |
| Resolveu ou abriu uma dívida | `09_MANUTENCAO.md` §2 |
| Bateu num problema estrutural ou erro recorrente | `05_EXPERIENCIAS.md` + **índice no topo** |
Se um número (tools, testes, linhas) mudou, corrija onde ele aparece. Se um
documento divergir do código, **o código está certo** — conserte o documento.
### Ao escrever documentação
- **Um assunto por arquivo.** Se um doc começar a cobrir dois, divida.
- **Diga o que não está ali** e para onde ir — economiza a leitura seguinte.
- **Fatos verificados**, não suposições: rode o comando e use o número real.
- **Registre o porquê**, não só o quê. O "o quê" está no código; o "por quê"
se perde, e é o que evita alguém desfazer uma decisão por engano.
## Key Patterns ## Key Patterns
@@ -69,12 +135,19 @@ não passar. Equivalente a rodar manualmente os dois comandos abaixo.
Sempre que uma alteração for feita no app (MacApp/) durante o período de Sempre que uma alteração for feita no app (MacApp/) durante o período de
implementação, **compile e rode o programa localmente no computador** para implementação, **compile e rode o programa localmente no computador** para
validar visualmente a alteração, além de rodar os testes: validar visualmente a alteração, além de rodar os testes. O comando padrão
para isso — que fecha a instância anterior, recompila e abre o app para
conferência — é:
```bash ```bash
cd code && ./MacApp/build_app.sh --run # compila e abre o app localmente admin/run_app.command # compila e abre o app localmente (padrão de revisão)
``` ```
Equivalente a `cd code && ./MacApp/build_app.sh --run`, mas desacoplado do
Terminal. **Toda vez que uma alteração for concluída, rode este arquivo
automaticamente** para já conseguirmos revisar o que foi feito antes de
fechar a tarefa.
Regra geral: após qualquer alteração, o app deve ser executado localmente Regra geral: após qualquer alteração, o app deve ser executado localmente
antes de concluir a tarefa. Se houver erro de compilação, corrija antes de antes de concluir a tarefa. Se houver erro de compilação, corrija antes de
seguir. seguir.
@@ -88,7 +161,7 @@ CI runs both on every push to main. If either fails, the commit gets an X on Git
## Testing ## Testing
1342 tests across 34 files. `test_models.py` covers TimeValue arithmetic, Timecode parsing/formatting, Clip properties, validation models, and Timeline helpers. `test_writer.py` covers insert_clip, add_marker (all types), trim_clip, delete_clip, split_clip, and change_speed operations. `test_server.py` covers MCP tool handlers, parsers, and dispatch. `test_rough_cut.py` covers RoughCutGenerator. `test_features_v05.py` covers connected clips, roles, timeline diff, reformat, silence detection, export, and backward compatibility. `test_marker_pipeline.py` covers build_marker_element shared builder, batch auto-modes, clip index duplicate-name behavior, and write_fcpxml output format. `test_refactored_helpers.py` covers _index_elements, _iter_spine_clips, _find_spine_clip_at_seconds, _resolve_clip_duration, _make_asset_clip, _format_batch_result, and serialize_xml edge cases. `test_transcribe.py` covers phrase/filler span matching, range merge/invert algebra, whisper graceful degradation, and transcript-driven handler cuts against cached transcripts. `test_media_intel.py` covers silencedetect stderr parsing, source-to-timeline mapping, parameter bounds, and real-WAV ffmpeg integration (skips without ffmpeg; CI installs it). Tests use `examples/sample.fcpxml` as fixture data and inline XML fixtures. Tests create temp files and clean up after. 1498 tests across 43 files, all under `code/tests/`. Um teste fora dessa pasta não roda (`testpaths = ["tests"]`) — se você criar um em outro lugar, confirme que a contagem total subiu. Cobertura por área: `test_models.py` (TimeValue/Timecode/Clip/Timeline), `test_writer.py` (insert/marker/trim/delete/split/speed), `test_server.py` (handlers e dispatch), `test_rough_cut.py`, `test_features_v05.py` (connected clips, roles, diff, reformat, silêncio, export), `test_marker_pipeline.py`, `test_refactored_helpers.py`, `test_transcribe.py`, `test_media_intel.py` (pula sem ffmpeg; o CI instala), `test_phrase_review.py` (revisão de frases da etapa 5) e `test_models_api.py` (comandos da ponte). Fixtures: `examples/sample.fcpxml` e XML inline. Os testes criam temporários e limpam depois.
## FCPXML Gotchas ## FCPXML Gotchas
+20
View File
@@ -0,0 +1,20 @@
"""Comandos da ponte JSON usada pelo app, agrupados por assunto.
O setup de sys.path mora aqui, e só aqui, porque o pacote é importado antes de
qualquer um dos seus módulos (`from admin.api import models, voice, ...`
dispara este arquivo primeiro). Cada módulo de comando importa `fcpxml.*`
antes de importar `.shared` — sem o path já pronto neste ponto, o primeiro
desses imports falha com `ModuleNotFoundError`. Repetir o cálculo em cada
módulo (como era antes) é frágil por ordem: o app roda `admin/models_api.py`
por caminho absoluto, então `__file__` está sempre correto, mas cada arquivo
que refizesse essa conta um nível de diretório errado — como aconteceu quando
`_shared.py` virou este pacote e `admin/code` (inexistente) saiu no lugar de
`code/` — quebrava em silêncio até alguém rodar o comando de verdade.
"""
import sys
from pathlib import Path
_CODE_DIR = str(Path(__file__).resolve().parent.parent.parent / "code")
if _CODE_DIR not in sys.path:
sys.path.insert(0, _CODE_DIR)
+122
View File
@@ -0,0 +1,122 @@
"""Edições no projeto: silêncio, corte por texto, preenchimento, marcadores.
Extraído de models_api.py — a tabela de comandos segue lá.
"""
from __future__ import annotations
import asyncio
from pathlib import Path
from fcpxml.model_manager import (
load_silence_config,
save_silence_config,
)
from . import shared
from .shared import (
_derived_output,
_emit_no_change_or_error,
)
def cmd_remove_silences(args: dict) -> int:
"""Run the canonical server silence remover into a suffixed copy."""
path = str(args.get("path", ""))
if not path or not Path(path).exists():
shared.emit({"ok": False, "error": "Arquivo de projeto não encontrado."})
return 1
try:
from server import handle_remove_media_silence
output = _derived_output(path, "_silence_removed", args)
contents = asyncio.run(handle_remove_media_silence({**args, "filepath": path, "output_path": output}))
message = "\n".join(getattr(content, "text", str(content)) for content in contents)
if not Path(output).exists():
return _emit_no_change_or_error(path, message)
shared.emit({"ok": True, "path": output, "message": message})
return 0
except Exception as exc:
shared.emit({"ok": False, "error": str(exc)})
return 1
def cmd_edit_by_transcript(args: dict) -> int:
"""Cut (or keep only) spoken phrases, using each media's cached transcript."""
path = str(args.get("path", ""))
phrases = args.get("phrases") or []
if not path or not Path(path).exists():
shared.emit({"ok": False, "error": "Arquivo de projeto não encontrado."})
return 1
if not isinstance(phrases, list) or not [p for p in phrases if str(p).strip()]:
shared.emit({"ok": False, "error": "Informe ao menos uma frase para cortar."})
return 1
try:
from server import handle_edit_by_transcript
output = _derived_output(path, "_transcript_edit", args)
contents = asyncio.run(handle_edit_by_transcript({**args, "filepath": path, "output_path": output}))
message = "\n".join(getattr(content, "text", str(content)) for content in contents)
if not Path(output).exists():
shared.emit({"ok": False, "error": message})
return 1
shared.emit({"ok": True, "path": output, "message": message})
return 0
except Exception as exc:
shared.emit({"ok": False, "error": str(exc)})
return 1
def cmd_remove_filler_words(args: dict) -> int:
"""Cut filler words (um, uh, ...) out, using each media's cached transcript."""
path = str(args.get("path", ""))
if not path or not Path(path).exists():
shared.emit({"ok": False, "error": "Arquivo de projeto não encontrado."})
return 1
try:
from server import handle_remove_filler_words
output = _derived_output(path, "_defillered", args)
contents = asyncio.run(handle_remove_filler_words({**args, "filepath": path, "output_path": output}))
message = "\n".join(getattr(content, "text", str(content)) for content in contents)
if not Path(output).exists():
return _emit_no_change_or_error(path, message)
shared.emit({"ok": True, "path": output, "message": message})
return 0
except Exception as exc:
shared.emit({"ok": False, "error": str(exc)})
return 1
def cmd_transcript_markers(args: dict) -> int:
"""Add a marker per transcribed segment, using each media's cached transcript."""
path = str(args.get("path", ""))
if not path or not Path(path).exists():
shared.emit({"ok": False, "error": "Arquivo de projeto não encontrado."})
return 1
try:
from server import handle_transcript_markers
output = _derived_output(path, "_transcript_markers", args)
contents = asyncio.run(handle_transcript_markers({**args, "filepath": path, "output_path": output}))
message = "\n".join(getattr(content, "text", str(content)) for content in contents)
if not Path(output).exists():
shared.emit({"ok": False, "error": message})
return 1
shared.emit({"ok": True, "path": output, "message": message})
return 0
except Exception as exc:
shared.emit({"ok": False, "error": str(exc)})
return 1
def cmd_silence_config(args: dict) -> int:
"""Read the persisted silence thresholds (noise floor, duration, padding)."""
shared.emit({"ok": True, **load_silence_config()})
return 0
def cmd_set_silence_config(args: dict) -> int:
"""Persist silence thresholds. Only the given fields change."""
config = save_silence_config(
noise_db=args.get("noise_db"),
min_silence=args.get("min_silence"),
padding=args.get("padding"),
)
shared.emit({"ok": True, **config})
return 0
+137
View File
@@ -0,0 +1,137 @@
"""Catálogo de modelos: listar, baixar, escolher, apagar.
Extraído de models_api.py — a tabela de comandos segue lá.
"""
from __future__ import annotations
import shutil
import subprocess
import threading
from fcpxml.diarize import (
diarization_capability,
)
from fcpxml.model_manager import (
download_model,
get_models_dir,
is_model_downloaded,
list_installed_models,
load_catalog,
load_hf_token,
load_num_speakers,
load_selected_model,
load_transcript_language,
model_cache_dir,
save_models_dir,
save_selected_model,
save_transcript_language,
)
from . import shared
from .shared import (
RECOMMENDED,
)
# Downloads em andamento, para o comando `cancel` conseguir interrompê-los.
# Mora aqui, e não no shared, porque só `download` e `cancel` o tocam — e o
# lock é próprio: ele protege este dicionário, não a saída em stdout.
_CANCEL: dict[str, threading.Event] = {}
_CANCEL_LOCK = threading.Lock()
def cmd_catalog() -> None:
catalog = load_catalog()
installed = list_installed_models()
diar_ok, diar_msg = diarization_capability(load_hf_token())
shared.emit(
{
"models": catalog,
"installed": installed,
"selected": load_selected_model(),
"language": load_transcript_language(),
"models_dir": str(get_models_dir()),
"installed_count": len(installed),
"recommended": list(RECOMMENDED),
"diarization": diar_ok,
"diarization_message": diar_msg,
"hf_token_set": bool(load_hf_token()),
"num_speakers": load_num_speakers(),
}
)
def cmd_download(args: dict) -> int:
model = str(args.get("model", ""))
if model not in _model_names():
shared.emit({"type": "error", "message": f"Modelo desconhecido: {model}"})
return 1
ev = threading.Event()
with _CANCEL_LOCK:
_CANCEL[model] = ev
try:
download_model(model, progress_cb=lambda f: shared.emit({"type": "progress", "fraction": f}), cancel_event=ev)
installed = is_model_downloaded(model)
shared.emit({"type": "done", "installed": installed})
if installed:
save_selected_model(model)
return 0 if installed else 1
except Exception as exc:
shared.emit({"type": "error", "message": str(exc)})
return 1
finally:
with _CANCEL_LOCK:
_CANCEL.pop(model, None)
def cmd_cancel(args: dict) -> None:
model = str(args.get("model", ""))
ev = _CANCEL.get(model)
if ev is not None:
ev.set()
shared.emit({"ok": True})
def cmd_select(args: dict) -> None:
model = str(args.get("model", ""))
if not is_model_downloaded(model):
shared.emit({"ok": False, "error": "Modelo não está instalado."})
return
save_selected_model(model)
shared.emit({"ok": True, "selected": load_selected_model()})
def cmd_set_language(args: dict) -> int:
"""Persist the transcription language (the default for every transcription)."""
lang = str(args.get("language", "auto"))
try:
saved = save_transcript_language(lang)
except ValueError as exc:
shared.emit({"ok": False, "error": str(exc)})
return 1
shared.emit({"ok": True, "language": saved})
return 0
def cmd_delete(args: dict) -> None:
model = str(args.get("model", ""))
try:
shutil.rmtree(model_cache_dir(model), ignore_errors=True)
except Exception:
pass
shared.emit({"ok": True})
def cmd_open_finder(args: dict) -> None:
target = str(args.get("path") or model_cache_dir(str(args.get("model", ""))))
try:
subprocess.Popen(["open", target])
except OSError:
pass
shared.emit({"ok": True})
def cmd_set_models_dir(args: dict) -> int:
try:
d = save_models_dir(str(args.get("dir", "")))
shared.emit({"ok": True, "models_dir": d})
return 0
except ValueError as exc:
shared.emit({"ok": False, "error": str(exc)})
return 1
def _model_names() -> list[str]:
return [m["internal_name"] for m in load_catalog()]
+69
View File
@@ -0,0 +1,69 @@
"""Projeto: inspecionar o .fcpxml e lembrar a pasta/arquivo em uso.
Extraído de models_api.py — a tabela de comandos segue lá.
"""
from __future__ import annotations
from pathlib import Path
from fcpxml.model_manager import (
load_project_config,
save_project_config,
)
from fcpxml.parser import parse_fcpxml
from . import shared
def cmd_inspect(args: dict) -> int:
"""Validate an FCPXML file and return a summary of its projects/timelines."""
path = str(args.get("path", ""))
if not path:
shared.emit({"ok": False, "error": "Nenhum arquivo informado."})
return 1
if not Path(path).exists():
shared.emit({"ok": False, "error": "Arquivo não encontrado."})
return 1
try:
proj = parse_fcpxml(path)
except Exception as exc:
shared.emit({"ok": False, "error": f"Erro ao ler o projeto: {exc}"})
return 1
timelines = []
for tl in proj.timelines:
timelines.append(
{
"name": tl.name,
"duration_seconds": round(tl.duration.seconds, 3),
"frame_rate": round(tl.frame_rate, 3),
"width": tl.width,
"height": tl.height,
"clips": tl.total_clips,
"cuts": tl.total_cuts,
"connected": len(tl.connected_clips),
"markers": len(tl.markers),
}
)
shared.emit(
{
"ok": True,
"path": path,
"name": proj.name,
"fcpxml_version": proj.fcpxml_version,
"timelines": timelines,
}
)
return 0
def cmd_project_config(args: dict) -> int:
"""Read the last project folder/file the app was working on."""
shared.emit({"ok": True, **load_project_config()})
return 0
def cmd_set_project_config(args: dict) -> int:
"""Persist the last project folder/file. Only the given fields change."""
config = save_project_config(folder=args.get("folder"), file=args.get("file"))
shared.emit({"ok": True, **config})
return 0
+90
View File
@@ -0,0 +1,90 @@
"""Revisão de frases: montar a tela de ênfases e salvar o que foi decidido.
Extraído de models_api.py — a tabela de comandos segue lá.
"""
from __future__ import annotations
import json
from pathlib import Path
from . import shared
def cmd_build_phrase_review(args: dict) -> int:
"""Build the reviewable script (phrases + the AI's decisions) for the wizard.
`voice_timeline` points at the _voice_timeline.json; `actions` carries the
decision list the model returned (inline, in any of the shapes the skill
emits). The review is always rebuilt from the current analysis, then the
decisions saved on a previous visit are laid back over it — reopening the
step must show the edits the user left there without freezing the acoustics
as they were when they left.
"""
from fcpxml.phrase_review import (
build_phrase_review,
load_phrase_review,
merge_saved_decisions,
)
timeline_path = str(args.get("voice_timeline", ""))
if not timeline_path or not Path(timeline_path).exists():
shared.emit({"ok": False, "error": "Análise de voz (voice_timeline.json) não encontrada."})
return 1
try:
with open(timeline_path, encoding="utf-8") as fh:
timeline = json.load(fh)
except (OSError, ValueError) as exc:
shared.emit({"ok": False, "error": f"Erro ao ler a análise de voz: {exc}"})
return 1
extra = [d for d in (args.get("output_dir"), args.get("media_dir")) if d]
review = build_phrase_review(
timeline,
args.get("actions"),
voice_timeline_path=timeline_path,
extra_dirs=extra,
)
saved = None if args.get("fresh") else load_phrase_review(timeline_path)
review = merge_saved_decisions(review, saved)
shared.emit({"ok": True, "reused": saved is not None, **review})
return 0
def cmd_save_phrase_review(args: dict) -> int:
"""Persist the edited review and the actions derived from it."""
from fcpxml.phrase_review import save_phrase_review
timeline_path = str(args.get("voice_timeline", ""))
if not timeline_path:
shared.emit({"ok": False, "error": "Caminho da análise de voz não informado."})
return 1
phrases = args.get("phrases")
if not isinstance(phrases, list):
shared.emit({"ok": False, "error": "Nenhuma frase para salvar."})
return 1
review = {
"version": args.get("version", "1.0"),
"source": args.get("source", ""),
"duration": args.get("duration", 0.0),
"speakers": args.get("speakers", []),
"phrases": phrases,
"zooms": args.get("zooms", []),
}
try:
review_path, actions_path = save_phrase_review(timeline_path, review)
except OSError as exc:
shared.emit({"ok": False, "error": f"Erro ao salvar a revisão: {exc}"})
return 1
shared.emit({
"ok": True,
"review_path": str(review_path),
"actions_path": str(actions_path),
"emphasis_count": sum(1 for p in phrases if int(p.get("emphasis", 0) or 0) >= 1),
"removed_count": sum(1 for p in phrases if not p.get("active", True)),
})
return 0
+156
View File
@@ -0,0 +1,156 @@
"""Base comum dos comandos da ponte: saída JSON, caminhos derivados e cache.
A saída passa toda por `emit`. Os módulos de comando chamam `shared.emit(...)`
pelo módulo, e não pelo nome importado, de propósito: assim trocar `emit` num
lugar só — como a suíte faz para capturar a saída — continua alcançando todos
os comandos, o que deixaria de valer se cada um tivesse ligado o nome no seu
próprio import.
"""
from __future__ import annotations
import json
import os
import sys
import threading
from pathlib import Path
from typing import Any
from fcpxml.diarize import build_speakers
from fcpxml.media_intel import media_src_to_path
from fcpxml.parser import parse_fcpxml
RECOMMENDED = ("large-v3", "distil-large-v3", "small", "base")
def _derived_output(path: str, suffix: str, args: dict) -> str:
"""Resolve a derived XML path, optionally inside the chosen output folder."""
output_dir = str(args.get("output_dir", "")).strip()
if output_dir:
directory = Path(output_dir).expanduser()
directory.mkdir(parents=True, exist_ok=True)
source = Path(path)
extension = ".fcpxmld" if source.is_dir() else source.suffix
return str(directory / f"{source.stem}{suffix}{extension}")
from server import generate_output_path
return generate_output_path(path, suffix)
def _is_no_change_message(message: str) -> bool:
"""Whether a tool completed cleanly without needing to save a new file."""
text = message.lower()
return any(
token in text
for token in (
"no cuts to make",
"no silence",
"file unchanged",
"nothing saved",
)
)
def _emit_no_change_or_error(path: str, message: str) -> int:
if _is_no_change_message(message):
emit({"ok": True, "path": path, "unchanged": True, "message": message})
return 0
emit({"ok": False, "error": message})
return 1
# Serializa a escrita em stdout. A ponte é JSON-lines: dois comandos
# escrevendo ao mesmo tempo entrelaçariam documentos e o app leria lixo.
_OUT_LOCK = threading.Lock()
def emit(obj: Any) -> None:
sys.stdout.write(json.dumps(obj, ensure_ascii=False) + "\n")
sys.stdout.flush()
def _transcript_json_path(media_path: str, output_dir: str = "") -> Path:
"""Where the ``_transcript.json`` for ``media_path`` lives.
When ``output_dir`` (the user-selected project folder) is set, the
transcript is saved/read there — never next to the source media, which
may sit on a read-only volume or a Final Cut Library the user never
browses. Falls back to the media's own folder only when no project
folder has been chosen (legacy/MCP callers).
"""
p = Path(media_path)
if output_dir:
directory = Path(output_dir).expanduser()
directory.mkdir(parents=True, exist_ok=True)
return directory / f"{p.stem}_transcript.json"
return p.with_name(p.stem + "_transcript.json")
def _save_json_atomic(path: Path, data: Any) -> None:
"""Write ``data`` to ``path`` atomically and validate the result on disk.
Mirrors the reference WHISPERX save path: write a ``.tmp``, ``os.replace``
into place, then confirm the file exists, is non-empty, and parses as JSON.
"""
tmp_path = str(path) + ".tmp"
with open(tmp_path, "w", encoding="utf-8") as fh:
json.dump(data, fh, ensure_ascii=False, indent=2)
os.replace(tmp_path, path)
if not path.exists() or os.path.getsize(path) == 0:
raise RuntimeError("O arquivo salvo está vazio ou não foi encontrado.")
with open(path, encoding="utf-8") as fh:
json.load(fh)
def _project_media_paths(path: str) -> list[str]:
proj = parse_fcpxml(path)
tl = proj.primary_timeline or (proj.timelines[0] if proj.timelines else None)
media_paths: list[str] = []
if tl is not None:
for clip in getattr(tl, "clips", []):
mp = media_src_to_path(clip.media_path or "")
if mp and Path(mp).is_file() and mp not in media_paths:
media_paths.append(mp)
return media_paths
def _project_media_rotations(path: str) -> dict[str, float]:
"""Degrees each source media was rotated by via a Transform filter on its
clip in the FCPXML — keyed by the same resolved media path
``_project_media_paths`` returns, so the two can be joined by media_path."""
proj = parse_fcpxml(path)
tl = proj.primary_timeline or (proj.timelines[0] if proj.timelines else None)
rotations: dict[str, float] = {}
if tl is not None:
for clip in getattr(tl, "clips", []):
mp = media_src_to_path(clip.media_path or "")
if mp and clip.rotation:
rotations[mp] = clip.rotation
return rotations
def _voice_timeline_json_path(media_path: str, output_dir: str = "") -> Path:
p = Path(media_path)
if output_dir:
directory = Path(output_dir).expanduser()
directory.mkdir(parents=True, exist_ok=True)
return directory / f"{p.stem}_voice_timeline.json"
return p.with_name(p.stem + "_voice_timeline.json")
def _load_cached_voice_timeline(json_path: Path, media_path: str) -> dict | None:
try:
with open(json_path, encoding="utf-8") as fh:
data = json.load(fh)
except (OSError, json.JSONDecodeError, UnicodeDecodeError):
return None
if not isinstance(data, dict):
return None
if data.get("source") != Path(media_path).name:
return None
if not isinstance(data.get("segments"), list):
return None
return data
def _load_cached_transcript(json_path: Path) -> dict | None:
"""Return a valid cached transcript dict, or ``None`` if absent/unreadable."""
if not json_path.is_file():
return None
try:
data = json.loads(json_path.read_text(encoding="utf-8"))
except (OSError, ValueError):
return None
if isinstance(data, dict) and isinstance(data.get("words"), list):
if "speakers" not in data:
data["speakers"] = build_speakers(data.get("segments", []))
return data
return None
+234
View File
@@ -0,0 +1,234 @@
"""Legendas: dinâmicas, comuns, SRT e as configurações de estilo.
Extraído de models_api.py — a tabela de comandos segue lá.
"""
from __future__ import annotations
import asyncio
from pathlib import Path
from fcpxml.media_intel import media_src_to_path
from fcpxml.model_manager import (
load_dynamic_subtitle_config,
load_plain_subtitle_config,
save_dynamic_subtitle_config,
save_plain_subtitle_config,
)
from fcpxml.writer import FCPXMLModifier
from . import shared
from .shared import (
_derived_output,
_emit_no_change_or_error,
_load_cached_transcript,
_transcript_json_path,
)
def cmd_generate_dynamic_subtitles(args: dict) -> int:
"""Generate word-by-word ("karaoke") caption compound clips, one per line,
using each media's cached transcript."""
path = str(args.get("path", ""))
if not path or not Path(path).exists():
shared.emit({"ok": False, "error": "Arquivo de projeto não encontrado."})
return 1
try:
from server import handle_generate_dynamic_subtitles
output = _derived_output(path, "_dynamic_subtitles", args)
contents = asyncio.run(
handle_generate_dynamic_subtitles({**args, "filepath": path, "output_path": output})
)
message = "\n".join(getattr(content, "text", str(content)) for content in contents)
if not Path(output).exists():
shared.emit({"ok": False, "error": message})
return 1
shared.emit({"ok": True, "path": output, "message": message})
return 0
except Exception as exc:
shared.emit({"ok": False, "error": str(exc)})
return 1
def cmd_generate_plain_subtitles(args: dict) -> int:
"""Generate simple static editable subtitle title clips."""
path = str(args.get("path", ""))
if not path or not Path(path).exists():
shared.emit({"ok": False, "error": "Arquivo de projeto não encontrado."})
return 1
try:
from server import handle_generate_plain_subtitles
output = _derived_output(path, "_plain_subtitles", args)
contents = asyncio.run(
handle_generate_plain_subtitles({**args, "filepath": path, "output_path": output})
)
message = "\n".join(getattr(content, "text", str(content)) for content in contents)
if not Path(output).exists():
return _emit_no_change_or_error(path, message)
shared.emit({"ok": True, "path": output, "message": message})
return 0
except Exception as exc:
shared.emit({"ok": False, "error": str(exc)})
return 1
def cmd_export_srt(args: dict) -> int:
"""Write a captions .srt synced to the edited timeline.
Each transcribed segment is mapped from its SOURCE-media timestamp to its
real TIMELINE position (``clip_offset + (seg_start - clip_source_start)``),
so captions only cover the frames that remain after cuts/silence removal —
not the whole source file. One .srt is produced per media, in timeline order.
"""
path = str(args.get("path", ""))
output_dir = str(args.get("output_dir", "")).strip()
if not path or not Path(path).exists():
shared.emit({"ok": False, "error": "Arquivo de projeto não encontrado."})
return 1
try:
modifier = FCPXMLModifier(path)
except Exception as exc:
shared.emit({"ok": False, "error": f"Erro ao ler o projeto: {exc}"})
return 1
# Group spine clips by media so each transcript is loaded once.
by_media: dict[str, list] = {}
for _, el in modifier._iter_spine_clips():
src = modifier.resources.get(el.get("ref", ""), {}).get("src", "")
mp = media_src_to_path(src)
if not mp or not Path(mp).is_file():
continue
by_media.setdefault(mp, []).append(el)
# Never emit a caption past the end of the project — Final Cut rejects an
# SRT whose last cue overruns the timeline ("subtitle extends beyond project
# duration"). Clamp every mapped cue end to this ceiling.
timeline_total = modifier._timeline_duration().to_seconds()
srt_paths: list[str] = []
for mp, clips in by_media.items():
cached = _load_cached_transcript(_transcript_json_path(mp, output_dir))
if cached is None:
continue
segments = cached.get("segments") or []
if not segments:
continue
rows: list[tuple[float, float, str, int]] = []
for el in clips:
clip_source_start = modifier.source_file_start(el).to_seconds()
clip_duration = modifier._parse_time(el.get("duration", "0s")).to_seconds()
clip_offset = modifier._parse_time(el.get("offset", "0s")).to_seconds()
window_end = clip_source_start + clip_duration
for seg_index, seg in enumerate(segments):
seg_start = float(seg.get("start", 0.0))
seg_end = float(seg.get("end", seg_start))
text = seg.get("text", "").strip()
if not text or seg_end <= seg_start:
continue
# Intersect the complete source segment with this kept clip.
# Testing only seg_start loses speech whose first words fall in
# a removed range; interval intersection preserves the part
# that remains and avoids duplicating a segment wholesale.
source_start = max(seg_start, clip_source_start)
source_end = min(seg_end, window_end)
if source_end <= source_start:
continue
tl_start = clip_offset + (source_start - clip_source_start)
tl_end = clip_offset + (source_end - clip_source_start)
tl_start = max(0.0, min(tl_start, timeline_total))
tl_end = max(0.0, min(tl_end, timeline_total))
if tl_end > tl_start:
rows.append((tl_start, tl_end, text, seg_index))
if not rows:
continue
rows.sort(key=lambda r: (r[0], r[1], r[3]))
# Merge only pieces from the same original Whisper segment when their
# mapped intervals touch. Never merge unrelated speech or invent time.
merged: list[tuple[float, float, str, int]] = []
for row in rows:
if merged and row[3] == merged[-1][3] and row[0] <= merged[-1][1] + 0.001:
prev = merged[-1]
merged[-1] = (prev[0], max(prev[1], row[1]), prev[2], prev[3])
else:
merged.append(row)
blocks = []
for index, (s, e, text, _) in enumerate(merged, 1):
start_stamp = srt_stamp(s)
end_stamp = srt_stamp(e)
# Millisecond SRT precision can collapse a sub-millisecond span;
# omit it rather than emit an invalid zero-duration cue.
if start_stamp == end_stamp:
continue
blocks.append(f"{index}\n{start_stamp} --> {end_stamp}\n{text}\n")
if not blocks:
continue
out = (
Path(output_dir).expanduser() / f"{Path(mp).stem}_captions.srt"
if output_dir
else Path(mp).with_name(Path(mp).stem + "_captions.srt")
)
if output_dir:
out.parent.mkdir(parents=True, exist_ok=True)
try:
out.write_text("\n".join(blocks), encoding="utf-8")
except OSError as exc:
shared.emit({"ok": False, "error": f"Não foi possível salvar a legenda: {exc}"})
return 1
srt_paths.append(str(out))
if not srt_paths:
shared.emit({"ok": False, "error": "Nenhuma transcrição encontrada. Transcreva o projeto primeiro."})
return 1
shared.emit({"ok": True, "paths": srt_paths, "message": f"{len(srt_paths)} legenda(s) .srt sincronizada(s) com o corte."})
return 0
def srt_stamp(seconds: float) -> str:
"""Format float seconds as ``HH:MM:SS,mmm`` (SRT uses a comma).
Uses ``floor`` (not ``round``) so a timestamp never rounds up past a frame
boundary — an SRT cue ending on the last frame must not overrun the
project duration, or Final Cut flags it as extending beyond the project.
"""
ms = int((seconds if seconds > 0 else 0.0) * 1000)
h, rem = divmod(ms, 3600000)
m, rem = divmod(rem, 60000)
s, ms = divmod(rem, 1000)
return f"{h:02d}:{m:02d}:{s:02d},{ms:03d}"
def cmd_dynamic_subtitle_config(args: dict) -> int:
"""Read the persisted dynamic-subtitle style (font, size, color, layout)."""
shared.emit({"ok": True, **load_dynamic_subtitle_config()})
return 0
def cmd_set_dynamic_subtitle_config(args: dict) -> int:
"""Persist dynamic-subtitle style fields. Only the given fields change."""
config = save_dynamic_subtitle_config(**{
k: args.get(k) for k in (
"band_height", "block_center_y", "line_gap", "font", "font_size",
"emphasis_font", "emphasis_face", "emphasis_size",
"active_color", "emphasis_color", "text_scale",
)
})
shared.emit({"ok": True, **config})
return 0
def cmd_plain_subtitle_config(args: dict) -> int:
"""Read the persisted simple subtitle style."""
shared.emit({"ok": True, **load_plain_subtitle_config()})
return 0
def cmd_set_plain_subtitle_config(args: dict) -> int:
"""Persist simple subtitle style fields. Only the given fields change."""
config = save_plain_subtitle_config(**{
k: args.get(k) for k in (
"font", "font_size", "font_color", "max_words",
"position_y", "uppercase", "keep_punctuation", "text_scale",
)
})
shared.emit({"ok": True, **config})
return 0
+185
View File
@@ -0,0 +1,185 @@
"""Transcrição e locutores.
Extraído de models_api.py — a tabela de comandos segue lá.
"""
from __future__ import annotations
import json
from pathlib import Path
from fcpxml.diarize import (
assign_speakers,
build_speakers,
diarization_capability,
diarize,
)
from fcpxml.media_intel import media_src_to_path
from fcpxml.model_manager import (
is_model_downloaded,
load_hf_token,
load_num_speakers,
load_selected_model,
load_transcript_language,
save_hf_token,
save_num_speakers,
)
from fcpxml.parser import parse_fcpxml
from fcpxml.transcribe import transcribe
from . import shared
from .shared import (
_load_cached_transcript,
_save_json_atomic,
_transcript_json_path,
)
def cmd_transcribe(args: dict) -> int:
proj_path = str(args.get("path", ""))
output_dir = str(args.get("output_dir", "")).strip()
# Honra o modelo selecionado no programa quando nenhum é passado.
model = str(args.get("model", "") or load_selected_model() or "")
language = args.get("language")
if language is None:
language = load_transcript_language()
if language == "auto":
language = None
if not proj_path:
shared.emit({"type": "error", "message": "Nenhum projeto selecionado."})
return 1
if not output_dir:
shared.emit({"type": "error", "message": "Selecione a pasta do projeto antes de transcrever."})
return 1
if not model or not is_model_downloaded(model):
shared.emit(
{
"type": "error",
"message": "Nenhum modelo de transcrição instalado. Baixe e selecione um modelo na aba Modelos.",
}
)
return 1
token = str(args.get("hf_token") or load_hf_token() or "")
if args.get("num_speakers") is not None:
num_speakers = str(args.get("num_speakers"))
else:
num_speakers = load_num_speakers()
# Load project.
try:
proj = parse_fcpxml(proj_path)
except Exception as exc:
shared.emit({"type": "error", "message": f"Erro ao ler o projeto: {exc}"})
return 1
tl = proj.primary_timeline or (proj.timelines[0] if proj.timelines else None)
media_paths: list[str] = []
if tl is not None:
for clip in getattr(tl, "clips", []):
mp = media_src_to_path(clip.media_path or "")
if mp and Path(mp).is_file() and mp not in media_paths:
media_paths.append(mp)
if not media_paths:
shared.emit({"type": "error", "message": "Nenhum arquivo de mídia acessível encontrado."})
return 1
total = len(media_paths)
results: list[dict] = []
for i, mp in enumerate(media_paths, 1):
stage = f"Transcrevendo {Path(mp).name} ({i}/{total})…"
shared.emit({"type": "progress", "fraction": (i - 1) / total, "stage": stage})
json_path = _transcript_json_path(mp, output_dir)
cached = _load_cached_transcript(json_path)
if cached is not None:
shared.emit({"type": "progress", "fraction": i / total, "stage": stage})
results.append(_result_row(mp, cached))
continue
def _on_progress(file_fraction: float, _i: int = i, _stage: str = stage) -> None:
# Blend this file's own progress into the overall fraction so a
# single-media project doesn't jump straight to 100% before the
# actual (slow) decoding work has even started.
overall = (_i - 1 + file_fraction) / total
shared.emit({"type": "progress", "fraction": overall, "stage": _stage})
data = transcribe(mp, model_size=model, language=language, progress_cb=_on_progress)
if data is None:
shared.emit({"type": "error", "message": f"Não foi possível transcrever: {Path(mp).name}"})
return 1
# Diarização opcional (necessita token HF): assina speaker por segmento/palavra.
if token:
tracks = diarize(mp, token, num_speakers)
segments, words = assign_speakers(
data.get("segments", []), data.get("words", []), tracks
)
data = {**data, "segments": segments, "words": words}
data["speakers"] = build_speakers(data.get("segments", []))
payload = {
"schema_version": "1.0",
"source": Path(mp).name,
"model": model,
**data,
}
try:
_save_json_atomic(json_path, payload)
except (OSError, RuntimeError, ValueError) as exc:
shared.emit({"type": "error", "message": f"Não foi possível salvar o JSON: {exc}"})
return 1
results.append(_result_row(mp, data))
shared.emit({"type": "result", "transcripts": results})
return 0
def cmd_rename_speakers(args: dict) -> int:
"""Apply real names to speakers already saved in a transcript JSON."""
json_path = Path(str(args.get("path", "")))
names = args.get("speakers") or {}
if not json_path.is_file():
shared.emit({"type": "error", "message": "Transcrição não encontrada."})
return 1
try:
data = json.loads(json_path.read_text(encoding="utf-8"))
except (OSError, ValueError) as exc:
shared.emit({"type": "error", "message": f"Não foi possível ler o JSON: {exc}"})
return 1
mapping = {str(sid): str(name).strip() for sid, name in (names or {}).items()}
for sp in data.get("speakers", []):
sid = str(sp.get("id", ""))
if mapping.get(sid):
sp["name"] = mapping[sid]
try:
_save_json_atomic(json_path, data)
except (OSError, RuntimeError, ValueError) as exc:
shared.emit({"type": "error", "message": f"Não foi possível salvar: {exc}"})
return 1
shared.emit({"ok": True, "speakers": data.get("speakers", [])})
return 0
def cmd_set_diarization(args: dict) -> int:
"""Persist the HuggingFace token and expected speaker count for diarization."""
token = args.get("token")
num = args.get("num_speakers")
if token is not None:
save_hf_token(str(token))
if num is not None:
save_num_speakers(str(num))
ok, msg = diarization_capability(load_hf_token())
shared.emit({"ok": True, "diarization": ok, "diarization_message": msg, "num_speakers": load_num_speakers()})
return 0
def _result_row(mp: str, data: dict) -> dict:
words = data.get("words", [])
preview = (data.get("text", "") or "")[:160]
speakers = data.get("speakers") or []
return {
"media": Path(mp).name,
"language": data.get("language", "?"),
"words": len(words),
"duration": float(data.get("duration", 0.0)),
"preview": preview,
"saved": str(_transcript_json_path(mp)),
"speakers": [s.get("name", s.get("id", "")) for s in speakers],
}
+300
View File
@@ -0,0 +1,300 @@
"""Análise de voz e aplicação das decisões de edição.
Extraído de models_api.py — a tabela de comandos segue lá.
"""
from __future__ import annotations
import asyncio
import json
import re
from pathlib import Path
from fcpxml.model_manager import (
load_hf_token,
load_num_speakers,
load_selected_model,
load_transcript_language,
load_voice_analysis_config,
save_voice_analysis_config,
)
from . import shared
from .shared import (
_load_cached_transcript,
_load_cached_voice_timeline,
_project_media_paths,
_project_media_rotations,
_transcript_json_path,
_voice_timeline_json_path,
)
def cmd_analyze_voice(args: dict) -> int:
"""Build the voice timeline (transcript+diarization+acoustics -> emphasis)
for every unique source media in the project, so `refine_voice_timeline`
and friends have something to read without ever reopening the audio.
Analysis only — writes _voice_timeline.json next to each media, doesn't
touch the project XML. `path` passes through unchanged so it composes
with the other batch steps (silence removal, captions) regardless of
where in the list it runs.
"""
path = str(args.get("path", ""))
if not path or not Path(path).exists():
shared.emit({"ok": False, "error": "Arquivo de projeto não encontrado."})
return 1
model = str(args.get("model", "") or load_selected_model() or "")
language = args.get("language")
if language is None:
language = load_transcript_language()
if language == "auto":
language = None
token = str(args.get("hf_token") or load_hf_token() or "")
num_speakers = str(args.get("num_speakers") or load_num_speakers() or "")
try:
media_paths = _project_media_paths(path)
rotations = _project_media_rotations(path)
except Exception as exc:
shared.emit({"ok": False, "error": f"Erro ao ler o projeto: {exc}"})
return 1
if not media_paths:
shared.emit({"ok": False, "error": "Nenhum arquivo de mídia acessível encontrado."})
return 1
from server import handle_build_voice_timeline
messages: list[str] = []
output_dir = str(args.get("output_dir") or "").strip()
existing: list[Path] = []
for mp in media_paths:
timeline_path = _voice_timeline_json_path(mp, output_dir)
if _load_cached_voice_timeline(timeline_path, mp) is not None:
existing.append(timeline_path)
if existing and len(existing) == len(media_paths) and not bool(args.get("force_reprocess", False)):
message = "# Voice Timeline Cache\n\n"
message += "Reaproveitando análise de voz existente. Nada foi reprocessado.\n\n"
for timeline_path in existing:
message += f"- **Timeline JSON**: {timeline_path}\n"
shared.emit({
"ok": True,
"path": path,
"reused": True,
"timelines": [str(p) for p in existing],
"message": message,
})
return 0
for mp in media_paths:
transcript_path = _transcript_json_path(mp, output_dir)
reused_prefix = ""
if _load_cached_transcript(transcript_path) is not None:
reused_prefix = f"# Cache\n\nReaproveitando transcrição existente: `{transcript_path}`\n\n"
try:
contents = asyncio.run(handle_build_voice_timeline({
"media_path": mp, "model": model, "language": language,
"hf_token": token, "num_speakers": num_speakers,
"output_dir": output_dir, "rotation": rotations.get(mp, 0.0),
}))
except Exception as exc:
shared.emit({"ok": False, "error": f"Falha analisando {Path(mp).name}: {exc}"})
return 1
messages.append(reused_prefix + "\n".join(getattr(c, "text", str(c)) for c in contents))
shared.emit({"ok": True, "path": path, "message": "\n\n---\n\n".join(messages)})
return 0
def cmd_acoustics_capability(args: dict) -> int:
"""Whether librosa (pitch/energy extraction) is installed in this venv.
Surfaces `features_capability()` — previously computed but never
exposed to the app, so `layers.acoustics: false` in a voice timeline
had no explanation the user could act on.
"""
from fcpxml.voice_features import features_capability
ok, msg = features_capability()
shared.emit({"ok": True, "available": ok, "message": msg})
return 0
def cmd_voice_analysis(args: dict) -> int:
"""Read the persisted voice-analysis settings (energy/emphasis/emotion)."""
config = load_voice_analysis_config()
shared.emit({"ok": True, **config, "emphasis_threshold": config["emphasis_floor"]})
return 0
def cmd_set_voice_analysis(args: dict) -> int:
"""Persist voice-analysis settings. Only the given fields change."""
weights = args.get("emphasis_weights")
config = save_voice_analysis_config(
energy_threshold=args.get("energy_threshold"),
emphasis_weights=weights if isinstance(weights, dict) else None,
emphasis_floor=args.get("emphasis_threshold"),
emotion_enabled=args.get("emotion_enabled"),
emotion_sensitivity=args.get("emotion_sensitivity"),
zoom_scale=args.get("zoom_scale"),
zoom_mode=args.get("zoom_mode"),
zoom_ease_in=args.get("zoom_ease_in"),
zoom_ease_out=args.get("zoom_ease_out"),
)
shared.emit({"ok": True, **config})
return 0
def cmd_apply_voice_actions(args: dict) -> int:
"""Apply a decision list (cuts/zooms/texts/markers) to the project XML.
The list is produced by a model reading the _voice_timeline.json — this
is the step that turns those decisions into an edit, and the one the
batch chain was missing: without it the app could measure the voice and
caption the result, but never cut by it.
`actions_path` points at the JSON; either a bare list or the
``{"actions": [...]}`` wrapper the skill emits is accepted. Times stay in
ORIGINAL source seconds — the handler resolves cuts first and shifts
everything else itself.
"""
path = str(args.get("path", ""))
if not path or not Path(path).exists():
shared.emit({"ok": False, "error": "Arquivo de projeto não encontrado."})
return 1
actions = args.get("actions")
if actions is None:
actions_path = str(args.get("actions_path", ""))
if not actions_path or not Path(actions_path).exists():
shared.emit({"ok": False, "error": "Arquivo de decisões (JSON) não encontrado."})
return 1
try:
with open(actions_path, encoding="utf-8") as fh:
loaded = json.load(fh)
except (OSError, ValueError) as exc:
shared.emit({"ok": False, "error": f"Erro ao ler as decisões: {exc}"})
return 1
actions = loaded.get("actions") if isinstance(loaded, dict) else loaded
# The documented output format is {"source": ..., "actions": [...]} —
# callers passing that whole object inline (e.g. the wizard pasting the
# skill's JSON verbatim) need the same unwrap the actions_path branch
# above already does, or a well-formed payload gets rejected as
# "malformed" for having one extra layer of nesting.
if isinstance(actions, dict):
actions = actions.get("actions")
if not isinstance(actions, list) or not actions:
shared.emit({"ok": False, "error": "A lista de decisões está vazia ou malformada."})
return 1
from server import handle_apply_voice_actions
try:
contents = asyncio.run(handle_apply_voice_actions({
"filepath": path,
"actions": actions,
"output_dir": args.get("output_dir"),
}))
except Exception as exc:
shared.emit({"ok": False, "error": f"Falha ao aplicar as decisões: {exc}"})
return 1
message = "\n".join(getattr(c, "text", str(c)) for c in contents)
# The handler reports dropped/rejected actions individually; hand the
# whole report back so the app can surface them instead of only the count.
out_path = path
for line in message.splitlines():
if line.startswith("- **Saved to**:"):
out_path = line.split("`")[1] if "`" in line else path
break
shared.emit({"ok": True, "path": out_path, "message": message})
return 0
def cmd_generate_voice_script(args: dict) -> int:
"""Run the ENTIRE voice-edit pass against a LOCAL model, inside the engine.
Transcribe (cached) -> build the voice timeline -> hand it to a local
Ollama model (Gemma 3 / Llama) that directs the edit -> return the readable
script (roteiro) and the action JSON, and optionally apply to a FCPXML. No
wizard, no copy-paste: the model's decisions are validated and applied by
the same pipeline the rules engine uses.
Args (all optional except one of ``media_path`` / ``voice_timeline``):
media_path audio/video to analyze and direct (required when there is
no voice_timeline yet)
voice_timeline path to an existing _voice_timeline.json; when given the
analysis is reused and media_path is not required
filepath optional FCPXML to apply the decisions to (non-destructive)
model local model Ollama serves (default gemma3:12b)
base_url Ollama base URL (default http://localhost:11434)
model_size whisper size if transcription is needed
language ISO language hint for transcription
hf_token HuggingFace token for diarization
num_speakers known speaker count, if any
output_dir folder for the timeline/review/actions JSON
apply_to_fcpxml apply to filepath when given (default true)
-> {"ok": true, "message": "...", "roteiro_path", "actions_path",
"applied_path"} or {"ok": false, "error": "..."}
"""
media_path = str(args.get("media_path", ""))
voice_timeline = str(args.get("voice_timeline", ""))
if not voice_timeline and (not media_path or not Path(media_path).exists()):
shared.emit({"ok": False, "error": "Arquivo de mídia não encontrado (informe media_path ou voice_timeline)."})
return 1
from server import handle_generate_voice_script
try:
contents = asyncio.run(handle_generate_voice_script({
"media_path": media_path,
"voice_timeline": args.get("voice_timeline"),
"filepath": args.get("filepath"),
"model": args.get("model"),
"base_url": args.get("base_url"),
"model_size": args.get("model_size"),
"language": args.get("language"),
"hf_token": args.get("hf_token"),
"num_speakers": args.get("num_speakers"),
"output_dir": args.get("output_dir"),
"apply_to_fcpxml": args.get("apply_to_fcpxml", True),
}))
except Exception as exc:
shared.emit({"ok": False, "error": f"Falha ao gerar roteiro por IA local: {exc}"})
return 1
message = "\n".join(getattr(c, "text", str(c)) for c in contents)
def _path_after(label: str) -> str:
m = re.search(rf"\*\*{label}\*\*: (.+)", message)
return m.group(1).strip() if m else ""
roteiro_path = _path_after(r"Roteiro \(legível\)")
actions_path = _path_after("Ações JSON")
applied_path = ""
for line in message.splitlines():
if line.startswith("- **Saved to**:"):
applied_path = line.split("`")[1] if "`" in line else ""
break
shared.emit({
"ok": True,
"message": message,
"roteiro_path": roteiro_path,
"actions_path": actions_path,
"applied_path": applied_path,
})
return 0
def cmd_list_ollama_models(args: dict) -> int:
"""List the models Ollama currently serves, for the app's model picker.
Args:
base_url Ollama base URL (default http://localhost:11434)
-> {"ok": true, "models": ["gemma3:12b", ...]} (empty list if Ollama
is unreachable, so the UI can fall back to a text field)
"""
from fcpxml.llm_local import list_ollama_models
base_url = str(args.get("base_url") or "http://localhost:11434")
models = list_ollama_models(base_url=base_url)
shared.emit({"ok": True, "models": models})
return 0
+107
View File
@@ -0,0 +1,107 @@
"""Zoom (punch-in): por janela, por clipe e por trecho da transcrição.
Extraído de models_api.py — a tabela de comandos segue lá.
"""
from __future__ import annotations
import asyncio
from pathlib import Path
from . import shared
from .shared import (
_derived_output,
_load_cached_transcript,
_transcript_json_path,
)
def cmd_add_zoom(args: dict) -> int:
"""Add an ease-in/ease-out punch-in zoom to one clip."""
path = str(args.get("path", ""))
clip_id = str(args.get("clip_id", "")).strip()
if not path or not Path(path).exists():
shared.emit({"ok": False, "error": "Arquivo de projeto não encontrado."})
return 1
if not clip_id:
shared.emit({"ok": False, "error": "Informe o nome do clipe."})
return 1
try:
from server import handle_add_zoom
output = _derived_output(path, "_zoom", args)
contents = asyncio.run(handle_add_zoom({**args, "filepath": path, "output_path": output}))
message = "\n".join(getattr(content, "text", str(content)) for content in contents)
if not Path(output).exists():
shared.emit({"ok": False, "error": message})
return 1
shared.emit({"ok": True, "path": output, "message": message})
return 0
except Exception as exc:
shared.emit({"ok": False, "error": str(exc)})
return 1
def cmd_zoom_clips(args: dict) -> int:
"""Return timeline clips with enough identity for the zoom picker."""
path = Path(str(args.get("path", "")))
output_dir = str(args.get("output_dir", "")).strip()
if not path.exists():
shared.emit({"ok": False, "error": "Arquivo de projeto não encontrado."})
return 1
try:
from server import _require_timeline
_, timeline = _require_timeline(str(path))
clips = []
for index, clip in enumerate(timeline.clips):
media = clip.media_path or ""
cached = _load_cached_transcript(_transcript_json_path(media, output_dir)) if media else None
clips.append({
"id": f"{index}:{clip.start.seconds:.6f}",
"index": index,
"name": clip.name,
"start": clip.start.seconds,
"duration": clip.duration_seconds,
"media": Path(media).name if media else "",
"preview": ((cached or {}).get("text", "") or "")[:180],
"has_transcript": cached is not None,
})
shared.emit({"ok": True, "clips": clips})
return 0
except Exception as exc:
shared.emit({"ok": False, "error": str(exc)})
return 1
def cmd_zoom_segments(args: dict) -> int:
"""Return sentence/word ranges for one timeline clip."""
path = Path(str(args.get("path", "")))
output_dir = str(args.get("output_dir", "")).strip()
try:
from server import _require_timeline
_, timeline = _require_timeline(str(path))
index = int(args.get("index", -1))
if index < 0 or index >= len(timeline.clips):
raise ValueError("Clipe selecionado não existe.")
clip = timeline.clips[index]
if not clip.media_path:
raise ValueError("Este clipe não possui mídia associada.")
data = _load_cached_transcript(_transcript_json_path(clip.media_path, output_dir))
if data is None:
shared.emit({"ok": True, "segments": [], "message": "Transcreva este clipe primeiro."})
return 0
segments = []
for number, segment in enumerate(data.get("segments", [])):
text = str(segment.get("text", "")).strip()
if text:
segments.append({
"id": number,
"start": float(segment.get("start", 0)),
"end": float(segment.get("end", 0)),
"text": text,
})
shared.emit({"ok": True, "segments": segments})
return 0
except Exception as exc:
shared.emit({"ok": False, "error": str(exc)})
return 1
+96 -977
View File
File diff suppressed because it is too large Load Diff
+7 -8
View File
@@ -20,7 +20,6 @@ import subprocess
import sys import sys
import threading import threading
from pathlib import Path from pathlib import Path
from typing import Optional
import flet as ft import flet as ft
@@ -29,8 +28,8 @@ _CODE_DIR = str(Path(__file__).resolve().parent.parent / "code")
if _CODE_DIR not in sys.path: if _CODE_DIR not in sys.path:
sys.path.insert(0, _CODE_DIR) sys.path.insert(0, _CODE_DIR)
from fcpxml.media_intel import media_src_to_path # noqa: E402 from fcpxml.media_intel import media_src_to_path
from fcpxml.model_manager import ( # noqa: E402 from fcpxml.model_manager import (
download_model, download_model,
get_models_dir, get_models_dir,
is_model_downloaded, is_model_downloaded,
@@ -41,8 +40,8 @@ from fcpxml.model_manager import ( # noqa: E402
save_models_dir, save_models_dir,
save_selected_model, save_selected_model,
) )
from fcpxml.parser import parse_fcpxml # noqa: E402 from fcpxml.parser import parse_fcpxml
from fcpxml.transcribe import transcribe # noqa: E402 from fcpxml.transcribe import transcribe
logger = logging.getLogger(__name__) logger = logging.getLogger(__name__)
@@ -101,9 +100,9 @@ class ModelManagerApp:
def __init__(self, page: ft.Page) -> None: def __init__(self, page: ft.Page) -> None:
self.page = page self.page = page
self.selected = load_selected_model() self.selected = load_selected_model()
self.downloading: Optional[str] = None self.downloading: str | None = None
self._cancel_events: dict[str, threading.Event] = {} self._cancel_events: dict[str, threading.Event] = {}
self._picker: Optional[ft.FilePicker] = None self._picker: ft.FilePicker | None = None
# ── helpers ──────────────────────────────────────────────────────────── # ── helpers ────────────────────────────────────────────────────────────
@@ -121,7 +120,7 @@ class ModelManagerApp:
self._picker = ft.FilePicker() self._picker = ft.FilePicker()
self._picker.on_result = self._on_file_picked self._picker.on_result = self._on_file_picked
self.page.overlay.append(self._picker) self.page.overlay.append(self._picker)
self._pending_target: Optional[dict] = None self._pending_target: dict | None = None
def _on_file_picked(self, e) -> None: def _on_file_picked(self, e) -> None:
if self._pending_target == "project": if self._pending_target == "project":
-243
View File
@@ -1,243 +0,0 @@
"""Tests for admin/models_api.py — the SwiftUI JSON bridge commands.
Focused on the transcription-flow changes: atomic save, speaker renaming, and
the "use the selected model" default plus model-availability guard.
"""
import json
import admin.models_api as api
def _capture(monkeypatch):
captured: list[dict] = []
def _emit(obj):
captured.append(obj)
monkeypatch.setattr(api, "_emit", _emit)
return captured
def test_save_json_atomic(tmp_path):
p = tmp_path / "t.json"
api._save_json_atomic(p, {"a": [1, 2], "text": "olá"})
assert p.exists()
assert not (tmp_path / "t.json.tmp").exists()
assert json.loads(p.read_text(encoding="utf-8"))["text"] == "olá"
def test_rename_speakers(tmp_path, monkeypatch):
captured = _capture(monkeypatch)
p = tmp_path / "t.json"
p.write_text(
json.dumps(
{
"speakers": [
{"id": "SPEAKER_00", "name": "Speaker 1"},
{"id": "SPEAKER_01", "name": "Speaker 2"},
]
}
),
encoding="utf-8",
)
assert api.cmd_rename_speakers({"path": str(p), "speakers": {"SPEAKER_01": "Erika"}}) == 0
assert captured[0]["ok"] is True
saved = json.loads(p.read_text(encoding="utf-8"))
assert saved["speakers"][0]["name"] == "Speaker 1"
assert saved["speakers"][1]["name"] == "Erika"
def test_rename_speakers_missing_file(monkeypatch):
captured = _capture(monkeypatch)
assert api.cmd_rename_speakers({"path": "/nonexistent/x.json"}) == 1
assert captured[0]["type"] == "error"
def test_transcribe_requires_output_dir(monkeypatch):
captured = _capture(monkeypatch)
monkeypatch.setattr(api, "load_selected_model", lambda: "small")
monkeypatch.setattr(api, "is_model_downloaded", lambda m: True)
assert api.cmd_transcribe({"path": "/some/project.fcpxml"}) == 1
assert captured[0]["type"] == "error"
assert "pasta do projeto" in captured[0]["message"]
def test_transcribe_requires_installed_model(monkeypatch, tmp_path):
captured = _capture(monkeypatch)
monkeypatch.setattr(api, "load_selected_model", lambda: "")
monkeypatch.setattr(api, "is_model_downloaded", lambda m: False)
assert api.cmd_transcribe({"path": "/some/project.fcpxml", "output_dir": str(tmp_path)}) == 1
assert captured[0]["type"] == "error"
assert "instalado" in captured[0]["message"]
def test_transcribe_defaults_to_selected_model(monkeypatch, tmp_path):
captured = _capture(monkeypatch)
monkeypatch.setattr(api, "load_selected_model", lambda: "small")
monkeypatch.setattr(api, "is_model_downloaded", lambda m: m == "small")
class FakeTL:
clips = []
class FakeProject:
primary_timeline = None
timelines = [FakeTL()]
monkeypatch.setattr(api, "parse_fcpxml", lambda p: FakeProject())
# No media accessible -> reaches the media-path check (past model validation).
assert api.cmd_transcribe({"path": "/some/project.fcpxml", "output_dir": str(tmp_path)}) == 1
assert captured[0]["type"] == "error"
assert "mídia" in captured[0]["message"]
def test_set_language_persists(monkeypatch):
captured = _capture(monkeypatch)
assert api.cmd_set_language({"language": "pt"}) == 0
assert captured[0]["ok"] is True
assert captured[0]["language"] == "pt"
assert api.load_transcript_language() == "pt"
def test_set_language_rejects_unknown(monkeypatch):
captured = _capture(monkeypatch)
assert api.cmd_set_language({"language": "xx"}) == 1
assert captured[0]["ok"] is False
assert "language" in captured[0]["error"]
def test_transcribe_defaults_language_to_persisted(monkeypatch, tmp_path):
monkeypatch.setattr(api, "load_selected_model", lambda: "small")
monkeypatch.setattr(api, "is_model_downloaded", lambda m: m == "small")
monkeypatch.setattr(api, "load_transcript_language", lambda: "pt")
media = tmp_path / "clip.mov"
media.write_bytes(b"fake")
class FakeClip:
media_path = ""
class FakeTL:
clips = [FakeClip()]
class FakeProject:
primary_timeline = None
timelines = [FakeTL()]
monkeypatch.setattr(api, "parse_fcpxml", lambda p: FakeProject())
monkeypatch.setattr(api, "media_src_to_path", lambda mp: str(media))
called = {}
monkeypatch.setattr(
api, "transcribe", lambda mp, model_size, language, **kw: called.update(lang=language)
)
assert api.cmd_transcribe({"path": "/some/project.fcpxml", "output_dir": str(tmp_path / "out")}) == 1
assert called["lang"] == "pt"
def test_srt_stamp_format():
assert api.srt_stamp(0.0) == "00:00:00,000"
assert api.srt_stamp(1.5) == "00:00:01,500"
assert api.srt_stamp(3661.234) == "01:01:01,234"
_FCPXML_SAMPLE = """<?xml version="1.0" encoding="UTF-8"?>
<fcpxml version="1.13">
<resources>
<asset id="r1" name="clip" uid="u1" start="0s" duration="100s"
hasVideo="1" format="f1" hasAudio="1">
<media-rep kind="original-media" src="file:///tmp/clip.mp4"/>
</asset>
<format id="f1" name="FFVideoFormat1080p25" frameDuration="1/25s" width="1920" height="1080"/>
</resources>
<library>
<event name="Event">
<project name="P">
<sequence format="f1">
<spine>
<asset-clip ref="r1" offset="0s" start="10s" duration="10s" name="clip"/>
<gap name="Espaço" offset="10s" duration="90s" start="10s"/>
</spine>
</sequence>
</project>
</event>
</library>
</fcpxml>
"""
def test_cmd_export_srt_maps_to_edited_timeline(tmp_path, monkeypatch):
"""Captions must reflect the EDITED timeline, not the whole source file."""
captured = _capture(monkeypatch)
project = tmp_path / "proj.fcpxml"
project.write_text(_FCPXML_SAMPLE, encoding="utf-8")
media = tmp_path / "clip.mp4"
media.write_bytes(b"fake")
# Transcript covers 0..100s; the clip only USES source 10..20s -> timeline 0..10s.
transcript = {
"words": [],
"segments": [
{"start": 5.0, "end": 6.0, "text": "antes do corte"},
{"start": 12.0, "end": 14.0, "text": "dentro do corte"},
{"start": 50.0, "end": 51.0, "text": "depois do corte"},
]
}
tj = api._transcript_json_path(media)
tj.parent.mkdir(parents=True, exist_ok=True)
api._save_json_atomic(tj, transcript)
monkeypatch.setattr(api, "media_src_to_path", lambda src: str(media))
assert api.cmd_export_srt({"path": str(project)}) == 0
assert captured[0]["ok"] is True
srt = tmp_path / "clip_captions.srt"
assert srt.exists()
text = srt.read_text(encoding="utf-8")
# Only the segment inside the used source window (12s) survives.
assert "dentro do corte" in text
assert "antes do corte" not in text
assert "depois do corte" not in text
# Mapped to timeline 0..10s -> the 12s source segment lands at 2s.
assert "00:00:02,000 --> 00:00:04,000" in text
def test_cmd_export_srt_no_transcript(tmp_path, monkeypatch):
captured = _capture(monkeypatch)
project = tmp_path / "proj.fcpxml"
project.write_text(_FCPXML_SAMPLE, encoding="utf-8")
media = tmp_path / "clip.mp4"
media.write_bytes(b"fake")
monkeypatch.setattr(api, "media_src_to_path", lambda src: str(media))
assert api.cmd_export_srt({"path": str(project)}) == 1
assert captured[0]["ok"] is False
def test_cmd_export_srt_clamps_past_project_duration(tmp_path, monkeypatch):
"""A segment ending after the last clip must be clamped to the project end.
Final Cut rejects an SRT whose final cue overruns the timeline
("subtitle extends beyond project duration").
"""
captured = _capture(monkeypatch)
project = tmp_path / "proj.fcpxml"
project.write_text(_FCPXML_SAMPLE, encoding="utf-8")
media = tmp_path / "clip.mp4"
media.write_bytes(b"fake")
# Clip uses source 10..20s -> timeline 0..10s. A segment 12..30s maps to
# timeline 2..20s, but the project only lasts 10s: must clamp end to 10s.
transcript = {
"words": [],
"segments": [
{"start": 12.0, "end": 30.0, "text": "longa fala"},
]
}
tj = api._transcript_json_path(media)
tj.parent.mkdir(parents=True, exist_ok=True)
api._save_json_atomic(tj, transcript)
monkeypatch.setattr(api, "media_src_to_path", lambda src: str(media))
assert api.cmd_export_srt({"path": str(project)}) == 0
assert captured[0]["ok"] is True
srt = tmp_path / "clip_captions.srt"
text = srt.read_text(encoding="utf-8")
# Timeline is 10s; the cue must not end past it.
assert "00:00:02,000 --> 00:00:10,000" in text
assert "00:00:20,000" not in text
+1
View File
@@ -0,0 +1 @@
analysis/
+43 -32
View File
@@ -10,12 +10,20 @@ opera **fora** do Final Cut Pro: você exporta o XML, o servidor processa o
documento como dados estruturados e devolve um XML modificado para importação. documento como dados estruturados e devolve um XML modificado para importação.
Nada é patcheado, nenhuma API privada é usada. Nada é patcheado, nenhuma API privada é usada.
Toda a análise foi feita a partir do código-fonte em `server.py` e `fcpxml/`. Toda a análise foi feita a partir do código-fonte. Este README é a visão
geral; o detalhe módulo a módulo mora em `docs/02_MODULES.md`, que é o
documento a manter atualizado quando a estrutura mudar.
> **Guia rápido:** [01 Arquitetura](docs/01_ARCHITECTURE.md) · > **Começando agora?** Leia [01 Arquitetura](docs/01_ARCHITECTURE.md) e depois
> [09 Manutenção](docs/09_MANUTENCAO.md) — o primeiro diz como o sistema é
> dividido, o segundo diz por onde começar a mexer e o que está em aberto.
>
> **Guia completo:** [01 Arquitetura](docs/01_ARCHITECTURE.md) ·
> [02 Módulos](docs/02_MODULES.md) · [03 Server/Tools](docs/03_SERVER_TOOLS.md) · > [02 Módulos](docs/02_MODULES.md) · [03 Server/Tools](docs/03_SERVER_TOOLS.md) ·
> [04 Testes & Workflow](docs/04_TESTS_AND_WORKFLOW.md) · > [04 Testes & Workflow](docs/04_TESTS_AND_WORKFLOW.md) ·
> [05 Experiências](docs/05_EXPERIENCIAS.md) · [06 Boas Práticas](docs/06_BOAS_PRATICAS.md) > [05 Experiências](docs/05_EXPERIENCIAS.md) · [06 Boas Práticas](docs/06_BOAS_PRATICAS.md) ·
> [07 Projeto Ativo no FCP](docs/07_ESTUDO_PROJETO_ATIVO_FCP.md) ·
> [08 App macOS](docs/08_APP_MACOS.md) · [09 Manutenção](docs/09_MANUTENCAO.md)
--- ---
@@ -26,7 +34,7 @@ Toda a análise foi feita a partir do código-fonte em `server.py` e `fcpxml/`.
Python, e reescreve de volta sem perda de sidecars (object tracking, Python, e reescreve de volta sem perda de sidecars (object tracking,
Cinematic). Cinematic).
2. **Uma camada MCP de 62 ferramentas** — expõe análise, edição em lote, QC, 2. **Uma camada MCP de 74 ferramentas** — expõe análise, edição em lote, QC,
geração, exportação cross-NLE, inteligência de mídia (silêncio/beats) e geração, exportação cross-NLE, inteligência de mídia (silêncio/beats) e
edição baseada em transcrição, tudo acessível por um cliente MCP (Claude). edição baseada em transcrição, tudo acessível por um cliente MCP (Claude).
@@ -40,7 +48,7 @@ Toda a análise foi feita a partir do código-fonte em `server.py` e `fcpxml/`.
| Camada | Tecnologia | | Camada | Tecnologia |
|--------|-----------| |--------|-----------|
| Linguagem | **Python 3.10+** (~7.1k linhas em `server.py` + `fcpxml/`) | | Linguagem | **Python 3.10+** (~13k linhas em `server.py`, `server_tools/` e `fcpxml/`) |
| Protocolo MCP | **mcp** (`mcp` SDK), servidor por stdio | | Protocolo MCP | **mcp** (`mcp` SDK), servidor por stdio |
| Parsing XML | **defusedxml** em todos os 4 entry points + `lxml`/`ElementTree` | | Parsing XML | **defusedxml** em todos os 4 entry points + `lxml`/`ElementTree` |
| Tempo racional | frações `numerador/denominador` no formato `"600/2400s"` | | Tempo racional | frações `numerador/denominador` no formato `"600/2400s"` |
@@ -56,26 +64,29 @@ Toda a análise foi feita a partir do código-fonte em `server.py` e `fcpxml/`.
``` ```
G-ART/ G-ART/
├── server.py # MCP server — 62 tools, prompts, resources, dispatch ├── CLAUDE.md # Regras do projeto para o agente
├── fcpxml/ # "Engine" — biblioteca Python de núcleo ├── admin/ # Ponte com o app (fora de code/)
│ ├── models.py # TimeValue, Timecode, Clip, Timeline, enums, QC models │ ├── models_api.py # Entry point: docstring dos comandos + dispatch
│ ├── parser.py # FCPXML → objetos Python (spine, connected clips, roles) │ └── api/ # Os 37 comandos, um módulo por assunto
│ ├── writer.py # Modifica e grava FCPXML (markers, trim, gaps, speed) └── code/
│ ├── rough_cut.py # Gera timelines novas (rough cuts, montages, A/B) ├── server.py # MCP entry point — só dispatch
│ ├── diff.py # Motor de comparação de timelines ├── server_tools/ # Handlers das 74 tools + _shared/
│ ├── export.py # Export DaVinci Resolve v1.9 + FCP7 XMEML v5 ├── fcpxml/ # "Engine" — biblioteca Python de núcleo
│ ├── media_intel.py # Detecção real de silêncio (ffmpeg) e beats (librosa) │ ├── writer/ # PACOTE: edição/escrita (mixins por assunto)
│ ├── transcribe.py # Transcrição Whisper local + edição por transcrição │ ├── models/ # PACOTE: dados por família (timing, timeline…)
│ ├── templates.py # Templates de timeline (intro/outro, lower thirds) │ ├── parser.py # FCPXML → objetos Python
│ ├── live.py # Modo Live — push_to_fcp / list_fcp_libraries │ ├── rough_cut.py # Gera timelines novas
│ ├── safe_xml.py # Wrappers defusedxml + serialize_xml() │ ├── voice_*.py # Pipeline de voz (features → timeline → actions)
│ └── dtd.py # Validação contra DTDs oficiais da Apple │ ├── phrase_review.py # Revisão de frases da etapa 5
├── Engine/ # Esta documentação da arquitetura │ ├── text_layout.py # Diagramação das legendas
├── admin/ # Scripts de manutenção (graphify.sh, graphify.md) │ ├── live.py # Modo Live — push_to_fcp
├── docs/ # WORKFLOWS, CAPABILITY-AUDIT, specs │ ├── safe_xml.py # defusedxml + serialize_xml()
├── examples/ # Fixture de teste (sample.fcpxml) │ └── dtd.py # Validação contra DTDs da Apple
├── tests/ # 1032 testes / 24 suítes ├── MacApp/Sources/ # App SwiftUI (compilado por swiftc)
└── tools/ # Pacote Python (__init__) ├── Engine/ # Esta documentação
├── docs/ # WORKFLOWS, CAPABILITY-AUDIT, specs
├── examples/ # Fixture de teste (sample.fcpxml)
└── tests/ # 1.466 testes / 42 suítes
``` ```
--- ---
@@ -102,7 +113,7 @@ TimeValue(600, 2400) # "600/2400s" == 0.25s
- Soma/subtração compartilham um único caminho `_binop()` (fast-path de mesmo - Soma/subtração compartilham um único caminho `_binop()` (fast-path de mesmo
denominador + alinhamento por LCM). denominador + alinhamento por LCM).
### 4.2 Modelos principais — `models.py` ### 4.2 Modelos principais — `models/`
| Classe | Função | | Classe | Função |
|--------|--------| |--------|--------|
@@ -126,8 +137,8 @@ escrita. `from_xml_element` faz match estrito do atributo `completed`
| Subsistema | Módulo | Função | | Subsistema | Módulo | Função |
|-----------|--------|--------| |-----------|--------|--------|
| Parser | `parser.py` | FCPXML → objetos Python: espinha, connected clips, secondary storylines, roles | | Parser | `parser.py` | FCPXML → objetos Python: espinha, connected clips, secondary storylines, roles |
| Modifier | `writer.FCPXMLModifier` | Edição index-based (clips/resources/formats dicts) do documento existente | | Modifier | `writer/` (`FCPXMLModifier`) | Edição index-based (clips/resources/formats dicts) do documento existente |
| Writer | `writer.FCPXMLWriter` | Gera FCPXML novo a partir de objetos Python | | Writer | `writer/generator.py` | Gera FCPXML novo a partir de objetos Python |
| Rough cut | `rough_cut.py` | Gera timelines (rough cuts, montages, A/B roll) | | Rough cut | `rough_cut.py` | Gera timelines (rough cuts, montages, A/B roll) |
| Diff | `diff.py` | Compara timelines — detecta added/removed/moved/trimmed | | Diff | `diff.py` | Compara timelines — detecta added/removed/moved/trimmed |
| Export | `export.py` | DaVinci Resolve v1.9 + FCP7 XMEML v5 | | Export | `export.py` | DaVinci Resolve v1.9 + FCP7 XMEML v5 |
@@ -151,7 +162,7 @@ assíncrono:
TOOL_HANDLERS = { TOOL_HANDLERS = {
"analyze_timeline": handle_analyze_timeline, "analyze_timeline": handle_analyze_timeline,
"list_clips": handle_list_clips, "list_clips": handle_list_clips,
# ... 62 tools # ... 74 tools, todos em server_tools/
} }
``` ```
@@ -249,7 +260,7 @@ correção, sem depender de lembrar dos dois comandos no pre-commit.
- [docs/02_MODULES.md](docs/02_MODULES.md) — guia módulo a módulo do `fcpxml/` - [docs/02_MODULES.md](docs/02_MODULES.md) — guia módulo a módulo do `fcpxml/`
(responsabilidade, tamanho, APIs públicas). (responsabilidade, tamanho, APIs públicas).
- [docs/03_SERVER_TOOLS.md](docs/03_SERVER_TOOLS.md) — a camada MCP `server.py`, - [docs/03_SERVER_TOOLS.md](docs/03_SERVER_TOOLS.md) — a camada MCP `server.py`,
62 ferramentas, helpers e o padrão de handler. 74 ferramentas, helpers e o padrão de handler.
- [docs/04_TESTS_AND_WORKFLOW.md](docs/04_TESTS_AND_WORKFLOW.md) — suíte de testes, - [docs/04_TESTS_AND_WORKFLOW.md](docs/04_TESTS_AND_WORKFLOW.md) — suíte de testes,
fluxo de trabalho (lint + pytest), execução e estado atual do sistema. fluxo de trabalho (lint + pytest), execução e estado atual do sistema.
- [docs/05_EXPERIENCIAS.md](docs/05_EXPERIENCIAS.md) — **memória de projeto**: - [docs/05_EXPERIENCIAS.md](docs/05_EXPERIENCIAS.md) — **memória de projeto**:
@@ -259,10 +270,10 @@ correção, sem depender de lembrar dos dois comandos no pre-commit.
programação** a aplicar em toda alteração/correção; inclui checklist final. programação** a aplicar em toda alteração/correção; inclui checklist final.
### Outros documentos ### Outros documentos
- [../CLAUDE.md](../CLAUDE.md) — visão geral, key patterns, execução e pre-commit. - [../CLAUDE.md](../../CLAUDE.md) — visão geral, key patterns, execução e pre-commit.
- [../docs/CAPABILITY-AUDIT-2026-06.md](../docs/CAPABILITY-AUDIT-2026-06.md) — - [../docs/CAPABILITY-AUDIT-2026-06.md](../docs/CAPABILITY-AUDIT-2026-06.md) —
auditoria do ecossistema e roadmap dual-mode (XML + Live). auditoria do ecossistema e roadmap dual-mode (XML + Live).
- [../docs/WORKFLOWS.md](../docs/WORKFLOWS.md) — 8 receitas de workflow de produção. - [../docs/WORKFLOWS.md](../docs/WORKFLOWS.md) — 8 receitas de workflow de produção.
- [../docs/specs/](../docs/specs/) — schemas de tools, estrutura FCPXML, pseudocódigo - [../docs/specs/](../docs/specs/) — schemas de tools, estrutura FCPXML, pseudocódigo
do writer, algoritmo de rough cut, implementação do server, roadmap, modelos. do writer, algoritmo de rough cut, implementação do server, roadmap, modelos.
- [../admin/graphify.md](../admin/graphify.md) — pipeline de graphify do código. - [../admin/graphify.md](../../admin/graphify.md) — pipeline de graphify do código.
+132 -64
View File
@@ -1,109 +1,177 @@
# 01 — Arquitetura do Sistema (G-ART / fcp-mcp-server) # 01 — Arquitetura do Sistema (G-ART / fcp-mcp-server)
> Referência canônica de como o sistema está dividido e implementado. Leia este > **Escopo:** Como o sistema é dividido em camadas e onde cada responsabilidade mora.
> documento antes de qualquer mudança de código. > **Não cobre:** Detalhe módulo a módulo (→ 02) · ferramentas MCP (→ 03) · app (→ 08)
> Referência canônica de como o sistema está dividido. Leia antes de qualquer
> mudança de código. Se algo aqui divergir do código, **o código está certo e
> este documento está velho** — corrija-o no mesmo commit.
Última varredura: 2026-08-19 · 77 ferramentas MCP · 1.498 testes · versão `0.6.35`
---
## 1. Visão de cima (camadas) ## 1. Visão de cima (camadas)
O sistema é um **servidor MCP em Python** que lê/analisa/reescreve arquivos O sistema lê, analisa e reescreve **FCPXML** do Final Cut Pro. Ele opera *fora*
**FCPXML** do Final Cut Pro. Há **três camadas** bem separadas: do FCP: você exporta o XML, o programa processa como dados estruturados e
devolve um XML para importar. Nada é patcheado, nenhuma API privada é usada.
São **quatro camadas**, e o ponto importante é que existem **duas portas de
entrada diferentes** para o mesmo motor:
``` ```
┌─────────────────────────────────────────────────────────────┐ ┌──────────────────────────┐ ┌──────────────────────────────┐
│ admin/ — Aplicações complementares (fora do MCP) │ │ MacApp/ (SwiftUI) │ │ Cliente MCP (Claude) │
│ models_api.py API (FastAPI) p/ gerenciar modelos │ │ O app que o usuário usa │ │ Conversa, decide a edição │
│ models_gui.py UI desktop (Flet) p/ gerenciar modelos │ └───────────┬──────────────┘ └───────────────┬──────────────┘
│ graphify.sh/.md Pipeline de graphify do código │ │ subprocesso + JSON-lines │ JSON-RPC (stdio)
├─────────────────────────────────────────────────────────────┤ ▼ ▼
│ server.py — CAMADA MCP / TRANSPORTE (NÃO tem lógica) │ ┌──────────────────────────┐ ┌──────────────────────────────┐
│ 73 tools, handlers, prompts, resources, dispatch │ │ admin/models_api.py │ │ server.py + server_tools/ │
│ Só valida entrada/saída e traduz JSON-RPC → chamadas │ │ + admin/api/ │ │ 77 tools, dispatch, schemas │
├─────────────────────────────────────────────────────────────┤ │ 37 comandos da ponte │ │ NÃO tem lógica de timeline │
│ fcpxml/ — "ENGINE" = NÚCLEO PURO Python (desacoplado) │ └───────────┬──────────────┘ └───────────────┬──────────────┘
│ Não conhece MCP nem argumentos de tool. │ └───────────────┬────────────────────┘
│ Trabalha com objetos Python e XML. │ ▼
│ É o foco / onde quase tudo mora. │ ┌───────────────────────────────────┐
└─────────────────────────────────────────────────────────────┘ │ fcpxml/ — O ENGINE │
│ Núcleo puro Python, desacoplado. │
│ Não conhece MCP nem o app. │
│ É onde quase tudo mora. │
└───────────────────────────────────┘
``` ```
**Regra de arquitetura:** `server.py` NUNCA implementa lógica de timeline — **A regra que sustenta tudo:** nem `server.py` nem `admin/api/` implementam
ele delega ao `fcpxml/`. Tudo em `fcpxml/` é testável isoladamente (1032 testes). lógica de timeline. Os dois validam entrada, chamam o engine e formatam a
saída. Toda regra de negócio é testável sem MCP e sem app.
## 2. Regras transversais (convenções em todo o código) **Por que duas portas.** O MCP existe para o julgamento editorial — qual tomada
usar, onde dar zoom — que é conversa com uma IA. A ponte existe para o que o
usuário faz sozinho no app — transcrever, configurar, processar. As duas caem
no mesmo engine, então uma correção ali vale para as duas.
---
## 2. Regras transversais (valem em todo o código)
| Conceito | Regra | | Conceito | Regra |
|----------|-------| |----------|-------|
| **Tempo** | `TimeValue` fração racional `"600/2400s"`. Nunca use float p/ tempo. | | **Tempo** | `TimeValue`, fração racional `"600/2400s"`. **Nunca float para tempo.** |
| **I/O paths** | Sempre via helpers `_validate_filepath` / `_validate_output_path` (sandbox). | | **Tempo de decisão** | Ações de voz usam sempre segundos da **mídia original**, nunca pós-corte. |
| **Nome de saída** | Nunca sobrescrever original: `output_<suffix>.fcpxml`. | | **I/O paths** | Sempre via `_validate_filepath` / `_validate_output_path` (sandbox). |
| **Segurança XML** | Sempre `defusedxml` (via `safe_xml.py`). Nunca `xml.etree` direto. | | **Nome de saída** | Nunca sobrescrever o original: `generate_output_path()` gera `_suffix`. |
| **Segurança XML** | Sempre `defusedxml` via `safe_xml.py`. Nunca `xml.etree` direto para ler. |
| **Deps opcionais** | `librosa`/`ffmpeg`/`huggingface_hub` importados **lazy**, degradam com `None`. | | **Deps opcionais** | `librosa`/`ffmpeg`/`huggingface_hub` importados **lazy**, degradam com `None`. |
| **Lint** | `ruff check . --exclude docs/` — zero erros. | | **Idioma** | Comunicação com o usuário em português. Código e comentários em inglês. |
| **Validação pós-correção** | `./Engine/run_after_fix.sh` SEMPRE após cada correção. | | **Validação** | `./Engine/run_after_fix.sh` **sempre** após cada correção. |
| **App** | Alterou `MacApp/`? Compile e rode: `admin/run_app.command` (padrão de revisão; equivale a `./MacApp/build_app.sh --run`). |
## 3. Fluxo de um request (round-trip) ---
## 3. Fluxo de um request
### Pela porta MCP (Claude decidindo a edição)
``` ```
Cliente MCP (Claude) Cliente MCP ──JSON-RPC──► server.py
│ JSON-RPC (stdio) │ TOOL_HANDLERS[nome]
▼ ▼
server.py ── dispatcher (TOOL_HANDLERS) server_tools/<categoria>.py
│ valida path, parseia projeto, chama engine │ _shared/: valida path, parseia projeto
▼ ▼
fcpxml/parser.py XML → objetos fcpxml/ (parser → writer → safe_xml)
fcpxml/writer.py edita / grava
fcpxml/rough_cut.py gera novas timelines
fcpxml/export.py cross-NLE
▼ ▼
output_<suffix>.fcpxml (original intocado) projeto_<suffix>.fcpxml (original intocado)
▼
Final Cut Pro: File → Import → XML (ou push_to_fcp, sem cliques)
``` ```
### Pela porta do app (usuário operando)
```
MacApp ──Process + argv JSON──► admin/models_api.py
│ handlers[comando]
▼
admin/api/<assunto>.py
│ shared.emit() devolve JSON-lines
▼
fcpxml/ (ou chama um handler do server)
▼
arquivo gerado + caminho de volta ao app
```
A saída da ponte é **JSON-lines**: um documento JSON por linha, para que
comandos longos transmitam progresso enquanto rodam. Toda escrita passa por
`admin/api/shared.py::emit`, que serializa o acesso a stdout — dois comandos
escrevendo ao mesmo tempo entrelaçariam documentos.
---
## 4. Dual-mode: XML + Live ## 4. Dual-mode: XML + Live
O sistema opera em **dois modos complementares**:
- **Modo XML (principal):** exporta FCPXML, processa como dados, reimporta. - **Modo XML (principal):** exporta FCPXML, processa como dados, reimporta.
Roda fora do FCP. Nenhuma API privada. - **Modo Live (`fcpxml/live.py`):** *push* do FCPXML direto para o FCP em
- **Modo Live (`fcpxml/live.py`):** *push* do FCPXML direto p/ o FCP em execução via Apple events oficiais (`Open Document`). Leitura de bibliotecas
execução via Apple events oficiais (`Open Document`), com `import-options`. via AppleScript read-only.
Leitura de bibliotecas via AppleScript read-only.
**Assimetria estrutural:** import é scriptable, mas a Apple não oferece export **Assimetria estrutural:** import é scriptable, mas a Apple não oferece export
programático — round-trips voltam pelas ferramentas XML. programático. Round-trips sempre voltam pelas ferramentas XML.
---
## 5. Onde está cada responsabilidade ## 5. Onde está cada responsabilidade
| Responsabilidade | Fica em | | Responsabilidade | Fica em |
|------------------|---------| |------------------|---------|
| Modelos de dados (tempo, clips, markers) | `fcpxml/models.py` | | Modelos de dados (tempo, clips, markers, QC, legendas) | `fcpxml/models/` |
| Parse FCPXML → objetos | `fcpxml/parser.py` | | Parse FCPXML → objetos | `fcpxml/parser.py` |
| Editing/escrita (modifier + writer) | `fcpxml/writer.py` | | Edição e escrita de FCPXML | `fcpxml/writer/` |
| Geração de timeline nova | `fcpxml/rough_cut.py` | | Geração de timeline nova | `fcpxml/rough_cut.py` |
| Comparação de timelines | `fcpxml/diff.py` | | Comparação de timelines | `fcpxml/diff.py` |
| Export cross-NLE (Resolve, FCP7) | `fcpxml/export.py` | | Export cross-NLE (Resolve, FCP7) | `fcpxml/export.py` |
| Inteligência de mídia (silêncio/beats) | `fcpxml/media_intel.py` | | Silêncio e beats | `fcpxml/media_intel.py` |
| Transcrição Whisper local | `fcpxml/transcribe.py` | | Transcrição Whisper | `fcpxml/transcribe.py` |
| Diarização (quem falou) | `fcpxml/diarize.py` |
| Ênfase acústica | `fcpxml/emphasis.py`, `fcpxml/voice_features.py` |
| Timeline de voz (o JSON que a IA lê) | `fcpxml/voice_timeline.py` |
| Decisões de edição (cut/zoom/text/marker) | `fcpxml/voice_actions.py` |
| Revisão de frases da etapa 5 | `fcpxml/phrase_review.py` |
| Layout de legendas e métricas de fonte | `fcpxml/text_layout.py`, `font_metrics.py`, `collision.py` |
| Gestão de modelos Whisper | `fcpxml/model_manager.py` | | Gestão de modelos Whisper | `fcpxml/model_manager.py` |
| Templates de timeline | `fcpxml/templates.py` |
| Controle Live do FCP | `fcpxml/live.py` | | Controle Live do FCP | `fcpxml/live.py` |
| Segurança XML (`defusedxml`, `serialize_xml`) | `fcpxml/safe_xml.py` | | Segurança XML | `fcpxml/safe_xml.py` |
| Validação contra DTDs da Apple | `fcpxml/dtd.py` | | Validação contra DTDs da Apple | `fcpxml/dtd.py` |
| Transporte MCP (73 tools) | `server.py` | | Transporte MCP (77 tools) | `server.py` + `server_tools/` |
| Ponte com o app (37 comandos) | `admin/models_api.py` + `admin/api/` |
| Interface do usuário | `MacApp/Sources/` |
## 6. Mapa de dependências (você está aqui se for mexer no X → quem tocar) ---
## 6. Mapa de dependências
``` ```
server.py ──► fcpxml/parser, writer, rough_cut, export, diff, MacApp/ ──► admin/models_api.py (subprocesso, por caminho)
media_intel, transcribe, templates, live, dtd admin/api/ ──► fcpxml/* e, para algumas operações, server.py
admin/models_gui.py ──► fcpxml/media_intel, model_manager, server.py ──► server_tools/*
parser, transcribe server_tools/* ──► server_tools/_shared/ ──► fcpxml/*
admin/models_api.py ──► fcpxml/model_manager fcpxml/writer/ ──► fcpxml/models/, safe_xml, dtd, text_layout, collision
fcpxml/writer.py ──► fcpxml/models, safe_xml, dtd fcpxml/models/ ──► fcpxml/text_layout (só o pacote subtitles)
fcpxml/__init__.py ──► reexporta a API pública fcpxml/__init__.py ──► reexporta a API pública
``` ```
> Se você cria uma **nova ferramenta MCP**, o trabalho principal é em `fcpxml/` **A seta que não existe, e não deve existir:** `fcpxml/` nunca importa de
> (função pura + testes). O handler em `server.py` fica fino: validação de `server_tools/`, de `admin/` ou de qualquer coisa que saiba o que é uma tool.
> caminho → `_parse_project` → chama a função → `_text_result`. Se você precisar disso, a lógica está no lugar errado.
---
## 7. Criando algo novo — por onde começar
| Você quer… | Comece por |
|-----------|-----------|
| Uma **ferramenta MCP** nova | Função pura em `fcpxml/` + teste. O handler em `server_tools/` fica fino. |
| Um **comando do app** novo | Mesmo caminho, e exponha em `admin/api/<assunto>.py` + tabela em `models_api.py`. |
| Uma **tela** nova | `MacApp/Sources/`, consumindo comandos que já existem na ponte. |
| Uma **regra de edição** nova | `fcpxml/` sempre. Se você está escrevendo `if` sobre timeline fora de `fcpxml/`, pare. |
O trabalho principal é **sempre** no engine. As camadas de cima são finas de
propósito: é o que permite testar 1.498 casos sem abrir o app nem subir o MCP.
+174 -96
View File
@@ -1,114 +1,192 @@
# 02 — Módulos do Engine (`fcpxml/`) # 02 — Módulos do Engine (`fcpxml/`)
Guia módulo a módulo do núcleo Python. Tamanho em linhas, responsabilidade e as > **Escopo:** Mapa do engine `fcpxml/`: qual módulo faz o quê e onde mexer.
funções/classes públicas de cada um. APIs públicas são reexportadas em > **Não cobre:** Camadas e regras gerais (→ 01) · handlers MCP (→ 03) · o que está aberto (→ 09)
`fcpxml/__init__.py` (fonte da verdade para o `__all__`).
## Versão atual Mapa módulo a módulo do núcleo Python: onde cada coisa mora e o que ela faz.
`__version__ = "0.6.35"` — ver `fcpxml/__init__.py`. A API pública é reexportada em `fcpxml/__init__.py` — essa é a fonte da verdade
do `__all__`.
Versão: `0.6.35` · Última varredura: 2026-08-19
> **Por que existem pacotes aqui.** `writer.py` tinha 4.199 linhas e `models.py`
> 1.091, cada um com muitos assuntos dentro. Viraram pacotes com um módulo por
> assunto. Do lado de fora **nada mudou**: `from .writer import FCPXMLModifier`
> e `from .models import TimeValue` seguem valendo, porque os `__init__.py`
> reexportam tudo — inclusive os nomes com underscore que a suíte usa.
--- ---
| Módulo | Linhas | Papel | ## Visão geral
|--------|-------:|-------|
| `models.py` | 930 | Data classes e enums (tempo, clips, markers, QC) | | Módulo / pacote | Linhas | Papel |
| `parser.py` | 367 | FCPXML → objetos Python | |-----------------|-------:|-------|
| `writer.py` | 3154 | Edição e escrita de FCPXML (o maior) | | `writer/` | 4.687 | **Edição e escrita de FCPXML** — o coração |
| `models/` | 1.195 | Data classes e enums |
| `text_layout.py` | 901 | Diagramação das legendas dinâmicas |
| `rough_cut.py` | 798 | Geração de timelines novas | | `rough_cut.py` | 798 | Geração de timelines novas |
| `dtd.py` | 112 | Validação contra DTDs oficiais | | `model_manager.py` | 748 | Modelos Whisper: catálogo, download, config |
| `safe_xml.py` | 113 | Wrappers `defusedxml` + `serialize_xml()` | | `voice_timeline.py` | 600 | O JSON de voz que a IA lê |
| `media_intel.py` | 173 | Silêncio (ffmpeg) e beats (librosa) | | `phrase_review.py` | 547 | Revisão de frases (etapa 5 do assistente) |
| `transcribe.py` | 184 | Transcrição Whisper + edição por transcrição | | `collision.py` | 472 | Colisão entre títulos na tela |
| `model_manager.py` | 298 | Gestão de modelos Whisper (cache/catálogo) | | `font_metrics.py` | 445 | Largura real de glifos por fonte |
| `export.py` | 226 | Export DaVinci Resolve v1.9 + FCP7 XMEML v5 |
| `diff.py` | 269 | Comparação de timelines |
| `live.py` | 273 | Modo Live — push_to_fcp / list_fcp_libraries |
| `templates.py` | 387 | Templates de timeline | | `templates.py` | 387 | Templates de timeline |
| `__init__.py` | 139 | Reexporta API pública | | `parser.py` | 367 | FCPXML → objetos Python |
| `transcribe.py` | 332 | Transcrição Whisper e corte por texto |
| `forced_align.py` | 181 | Alinhamento forçado opcional (whisperx/wav2vec2) que corrige o viés de ~0,4s no início das palavras |
| `live.py` | 273 | Modo Live (push_to_fcp) |
| `diff.py` | 269 | Comparação de timelines |
| `voice_actions.py` | 263 | Decisões de edição (cut/zoom/text/marker) |
| `export.py` | 226 | Export Resolve v1.9 + FCP7 XMEML v5 |
| `voice_features.py` | 220 | Pitch, energia, ritmo, pausas |
| `diarize.py` | 180 | Quem falou (pyannote) |
| `media_intel.py` | 177 | Silêncio (ffmpeg) e beats (librosa) |
| `emphasis.py` | 133 | Índice de ênfase por palavra |
| `safe_xml.py` | 113 | `defusedxml` + `serialize_xml()` |
| `dtd.py` | 112 | Validação contra os DTDs da Apple |
--- ---
## `models.py` — modelos e enums ## `writer/` — edição e escrita
Single source of truth para estrutura de dados. NUNCA mexa aqui sem rodar
`test_models.py`.
- **Tempo:** `TimeValue` (fração racional), `Timecode`. O `FCPXMLModifier` é montado por **composição de mixins**: um mixin por assunto
- **Clips:** `Clip`, `VideoClip`, `AudioClip`, `ConnectedClip` (lane), editorial, todos operando sobre o mesmo documento e os mesmos índices.
`CompoundClip`, `Transition`.
- **Contêineres:** `Timeline`, `Project`, `Keyword`.
- **Markers:** `Marker`, `MarkerType`, `MarkerColor`, `MARKER_XML_TAGS`.
`MarkerType` é o dono da serialização (`from_string`/`from_xml_element`/`xml_attrs`).
Match estrito do atributo `completed` (`'0'`/`'1'`, sem padding).
- **QC:** `SilenceCandidate`, `FlashFrame`, `GapInfo`, `DuplicateGroup`,
`ValidationIssue`, `ValidationResult`.
- **Geração:** `SegmentSpec`, `PacingConfig`, `PacingStyle`, `RoughCutResult`.
## `parser.py` — leitura | Módulo | Linhas | Conteúdo |
- `parse_fcpxml(path)` → `Project`. |--------|-------:|----------|
- `FCPXMLParser` — lê spine, connected clips (lanes), secondary storylines, roles. | `core.py` | 723 | `ModifierCore`: carga, índices, navegação na spine, `save` |
| `titles.py` | 600 | Títulos de texto e legendas dinâmicas |
| `cut.py` | 333 | Dividir, cortar faixas, apagar |
| `speed.py` | 297 | Velocidade e zoom (punch-in) |
| `helpers.py` | 279 | Sanitização, escalas, construtores de elemento |
| `rapid.py` | 240 | Flash frames, rapid trim, preencher buracos |
| `validation.py` | 232 | Verificações estruturais antes de salvar |
| `compound.py` | 196 | Compound clips: criar e achatar |
| `silence.py` | 185 | Detectar e remover silêncio |
| `document.py` | 170 | Assets de vídeo, timebases, `write_fcpxml` |
| `markers.py` | 165 | Marcadores: um, por timecode, em lote |
| `audio.py` | 162 | Clipes de áudio e cama musical |
| `generator.py` | 147 | `FCPXMLWriter` — cria documento do zero |
| `reorder.py` | 126 | Reordenar e recalcular offsets |
| `trim.py` | 125 | Aparar e propagar o ripple |
| `transitions.py` | 94 | Transições entre vizinhos |
| `relink.py` | 94 | Repontar mídia |
| `insert.py` | 78 | Inserir clipes na spine |
| `modifier.py` | 64 | Monta a classe a partir dos mixins |
| `selection.py` | 57 | Selecionar por palavra-chave |
| `api.py` | 55 | Atalhos de uma linha |
| `connected.py` | 49 | Clipes conectados (lanes) |
| `roles.py` | 43 | Atribuir roles |
| `reformat.py` | 43 | Reenquadrar resolução |
## `writer.py` — o coração (3154 linhas) **Onde mexer:** ache o assunto na tabela e abra só aquele arquivo. Se a sua
Duas classes principais: mudança precisa de dois mixins ao mesmo tempo, provavelmente o que você quer
é um método novo no `core.py` que os dois chamem.
- **`FCPXMLModifier`** — edita documento existente de forma index-based **Cuidado:** os mixins compartilham `self`. Um método novo que colida de nome
(dicts de `clips`/`resources`/`formats`), imune a ambiguidade de nomes duplicados. com outro mixin sobrescreve em silêncio — a ordem em `modifier.py` decide quem
Métodos: `insert_clip`, `add_marker`, `trim_clip`, `delete_clip`, `split_clip`, ganha. Hoje nenhum colide; mantenha assim.
`change_speed`, `cut_clip_ranges` (usado pela remoção de silêncio), etc.
- **`FCPXMLWriter`** — gera FCPXML novo a partir de objetos Python.
Helpers de nível de arquivo: `modify_fcpxml`, `add_marker_to_file`,
`trim_clip_in_file`, `build_marker_element`, `write_fcpxml`, `validate_fcpxml`,
`list_effects`, `FCP_EFFECTS`.
## `rough_cut.py` — geração
- `RoughCutGenerator`, `generate_rough_cut`, `generate_segmented_rough_cut`.
## `media_intel.py` — inteligência de mídia (v0.10)
- Silêncio via `ffmpeg silencedetect` (subprocess limitado), `remove_silence_candidates`,
mapeamento source→timeline.
- Beats via `librosa` (import lazy, extra `[intelligence]`).
- Degrada para `None` quando `ffmpeg` ausente.
## `transcribe.py` — Whisper local
- `transcribe(media_path, model_size, language)` → dict com `words` (spans).
- `ALLOWED_MODELS` — allowlist de nomes de modelo (também usado por `model_manager`).
- Edição por transcrição: remove filler words, aparar por transcrição.
## `model_manager.py` — gestão de modelos
Catálogo `models.json` + cache no HF hub. Config em `~/.fcp-mcp-server/config.json`.
Funções: `get/save_models_dir`, `list_installed_models`, `download_model`,
`delete_model`, `get/load_selected_model`, `save_selected_model`, `load_catalog`.
Permite cancelamento de download via `threading.Event`. Segue convenções:
allowlist, lazy imports, degradação graciosa.
## `export.py` — cross-NLE
- `DaVinciExporter` — FCPXML v1.9 p/ DaVinci Resolve.
- Export FCP7 XMEML v5.
## `diff.py` — comparação
- `compare_timelines`, `TimelineDiff`, `ClipDiff`, `MarkerDiff`.
- Detecta added/removed/moved/trimmed clips & markers.
## `live.py` — FCP ao vivo (macOS)
- `push_to_fcp(path, library, options)` — Apple event *Open Document* + `<import-options>`.
Requer `.fcpbundle` p/ zero-click real.
- `list_fcp_libraries()` — AppleScript read-only.
## `templates.py`
- `Template`, `TemplateSlot`, `ClipSpec`, `BUILTIN_TEMPLATES`, `apply_template`,
`list_templates`. Estruturas prontas: intro/outro, lower thirds, music video.
## `safe_xml.py`
Wrappers `defusedxml` centralizados + `serialize_xml()`. Todo parse/escrita passa aqui.
## `dtd.py`
Valida output contra DTDs oficiais no bundle do FCP (via `xmllint`; exige o caminho
do DTD percent-encoded por causa dos espaços em "Final Cut Pro.app").
--- ---
## Como adicionar um módulo novo ## `models/` — dados e enums
1. Criar `fcpxml/<seu_modulo>.py` — função pura, sem conhecer MCP.
2. Reexportar classes/funções em `fcpxml/__init__.py` (`__all__`). Fonte única da estrutura de dados. **Nunca mexa aqui sem rodar `test_models.py`.**
3. Cobrir em `tests/test_<seu_modulo>.py`.
4. Rodar `./Engine/run_after_fix.sh`. | Módulo | Linhas | Conteúdo |
|--------|-------:|----------|
| `timing.py` | 304 | `TimeValue` (fração racional), `Timecode` |
| `timeline.py` | 217 | `Clip`, `ConnectedClip`, `CompoundClip`, `Timeline`, `Project`, `Marker` |
| `enums.py` | 183 | `MarkerType`, `MarkerColor`, `TransitionType`, `PacingStyle`… |
| `subtitles.py` | 157 | `WordLook`, `WordStyle`, `DynamicSubtitleConfig`, paleta |
| `qc.py` | 121 | `FlashFrame`, `GapInfo`, `DuplicateGroup`, `ValidationIssue` |
| `planning.py` | 93 | `SegmentSpec`, `PacingConfig`, `RoughCutResult`, `MontageConfig` |
`MarkerType` é o dono da serialização de marcador (`from_string`,
`from_xml_element`, `xml_attrs`) — não reimplemente isso em outro lugar.
---
## O caminho da voz (do áudio à decisão)
Estes seis módulos formam um pipeline. É o fluxo mais novo e o menos óbvio do
projeto, então vale ler nesta ordem:
```
transcribe.py áudio → palavras com tempo
+
diarize.py quem falou cada trecho
+
voice_features.py pitch, energia, ritmo, pausas
▼
emphasis.py combina tudo num índice 0–1 por palavra
▼
voice_timeline.py monta o _voice_timeline.json ◄── é isto que a IA lê
▼
[decisão: skill "editar-por-voz", ou a mão do usuário]
▼
voice_actions.py valida a lista de cut/zoom/text/marker
▼
phrase_review.py funde tudo em frases revisáveis (etapa 5 do app)
▼
writer/ aplica no FCPXML
```
**Regra de ouro do pipeline:** toda ação carrega tempo da **mídia original**,
nunca pós-corte. Cortes deslocam tudo depois deles; resolver o deslocamento só
na hora de aplicar (`shift_after_cuts`) elimina uma classe inteira de bug.
### `voice_timeline.py` — o contrato com a IA
Saída em camadas, para um modelo raciocinar do topo e descer só onde importa:
```
{version, source, language,
layers: {transcript, acoustics, speakers, emotion} ← o que rodou de verdade
scales: {…} ← como ler cada número
summary: {…}
speakers: [...]
segments: [{start, end, speaker, text, gap_before, take_boundary,
avg_energy, peak_emphasis, emotion, emotion_confidence,
words: [{text, start, end, energy, pitch_delta, rate_delta,
pause_before, emphasis}]}]}
```
`layers` existe para separar *"a fala é monótona"* de *"a análise acústica nunca
carregou"* — os dois deixam os mesmos zeros nos dados.
### `phrase_review.py` — a revisão humana
Junta o timeline de voz com as ações da IA numa lista de frases editáveis, e
converte de volta. Frase inativa vira `cut`; ênfase ≥ 1 vira `zoom` mais um
`emphasis_spans` que a etapa de legendas usa. O trim de cada frase anda em
**fronteira de palavra** — cortar é apontar para uma palavra, nunca caçar frame.
---
## Legendas dinâmicas (três módulos que andam juntos)
| Módulo | Papel |
|--------|-------|
| `text_layout.py` | Quebra a frase em linhas e posiciona cada palavra |
| `font_metrics.py` | Largura real de cada glifo na fonte escolhida |
| `collision.py` | Detecta título saindo do quadro ou colidindo com outro |
Estes três não estão divididos porque **cada um já é um assunto só**. O
`text_layout.py` tem 901 linhas de um problema coeso: diagramação.
---
## Armadilhas do FCPXML (custaram sessões de depuração)
- Tempo é fração: `"3600/2400s"` = 1,5 s.
- `offset` é posição na timeline; `start` é o in-point da mídia.
- `<asset-clip>` (biblioteca) é diferente de `<clip>` (timeline).
- Marcadores são **filhos** do clipe, não irmãos.
- `.fcpxmld` é um **diretório** — sidecars precisam ser copiados no save, ou
dados de object tracking e Cinematic são destruídos.
- Negrito no FCP é `bold="1"` (atributo); itálico é `fontFace` + `italic="1"`.
- `id` de `<text-style-def>` precisa ser XML Name válido — acento, espaço ou
dígito inicial fazem o FCP recusar o arquivo inteiro.
- `code/examples/sample.fcpxml` **não** é DTD-conformante. Não use como fixture
de validade.
+89 -25
View File
@@ -1,26 +1,49 @@
# 03 — Camada MCP (`server.py`) — 73 ferramentas # 03 — Camada MCP (`server.py` + `server_tools/`) — 77 ferramentas
`server.py` (3824 linhas) é a camada de transporte. Não tem lógica de timeline — > **Escopo:** As 77 ferramentas MCP: helpers, categorias e como criar uma nova.
mapeia nome → handler e delega ao Engine. O dispatch é um dicionário > **Não cobre:** Lógica de edição, que mora no engine (→ 02) · comandos do app (→ 08)
`TOOL_HANDLERS` (padrão de despacho, sem cadeias gigantes de if/elif).
`server.py` (592 linhas) é só o transporte: dispatch por dicionário
`TOOL_HANDLERS`, sem cadeia de if/elif e **sem lógica de timeline**. Os handlers
moram em `server_tools/`, um módulo por categoria, e os helpers que todos usam
em `server_tools/_shared/`.
```
server_tools/
editing.py (649) qc.py (696) voice.py (754) timeline.py (400)
subtitles.py markers_import generation.py transcript.py
export.py roles.py live.py
_shared/ ← helpers compartilhados, ver abaixo
```
## Helpers centrais (use-os, não reinvente) ## Helpers centrais (use-os, não reinvente)
| Helper | Linha | Função | Todos reexportados por `server_tools/_shared`, então `from ._shared import X`
|--------|------:|--------| continua funcionando. A coluna diz o módulo real, para quando você precisar
| `_check_json_depth()` | 83 | Rejeita payloads além de 50 níveis | **editar** o helper — ou apontar um `monkeypatch` para ele.
| `_validate_filepath()` | 103 | Sandbox de entrada |
| `_validate_output_path()` | 149 | Sandbox de saída |
| `_format_clip_table()` | 245 | Renderização de tabela |
| `_markdown_table()` | 259 | Renderização de tabela markdown |
| `_parse_project()` | 319 | Parseia FCPXML → `(tree, timeline, project)`; quase todos os handlers começam aqui |
| `_resolve_io_paths()` | 357 | Validação de caminho de entrada/saída |
| `_setup_modifier()` | 390 | Prepara modifier com validação |
| `_setup_generator()` | 414 | Prepara generator com validação |
| `_parse_timestamp_parts()` | 433 | Parse de timestamps (min:seg, H:MM:SS, SMPTE) |
| `_detect_flash_frames/gaps/duplicate_groups()` | 1667+ | Detectores de QC |
## As 73 ferramentas por categoria | Helper | Mora em | Função |
|--------|---------|--------|
| `_validate_filepath()` | `_shared/paths.py` | Sandbox de entrada |
| `_validate_output_path()` | `_shared/paths.py` | Sandbox de saída |
| `_check_json_depth()` | `_shared/paths.py` | Rejeita payloads além de 50 níveis |
| `generate_output_path()` | `_shared/paths.py` | Nome derivado, sem tocar no original |
| `_resolve_io_paths()` | `_shared/paths.py` | Entrada + saída de uma vez |
| `_parse_project()` | `_shared/project.py` | FCPXML → `(tree, timeline, project)`; quase todo handler começa aqui |
| `_setup_modifier()` | `_shared/project.py` | Prepara modifier já validado |
| `_setup_generator()` | `_shared/project.py` | Prepara generator já validado |
| `_text_result()` | `_shared/project.py` | Envolve o texto em `TextContent` MCP |
| `_markdown_table()` | `_shared/formatting.py` | Tabela markdown |
| `_format_clip_table()` | `_shared/formatting.py` | Tabela de clipes |
| `_format_batch_result()` | `_shared/formatting.py` | Relatório de operação em lote |
| `_parse_timestamp_parts()` | `_shared/captions.py` | min:seg, H:MM:SS, SMPTE |
| `parse_srt()` / `parse_vtt()` | `_shared/captions.py` | Legendas coladas |
| `_detect_flash_frames/gaps/duplicate_groups()` | `_shared/detection.py` | Detectores de QC |
| `_load_or_transcribe()` | `_shared/media.py` | Transcrição com cache em disco |
| `_cut_transcript_spans()` | `_shared/media.py` | Corte por trecho falado |
| `_apply_placed_action()` | `_shared/media.py` | Aplica zoom/text/marker já posicionado |
## As 77 ferramentas por categoria
### Timeline & análise (Projeto) ### Timeline & análise (Projeto)
`list_projects`, `analyze_timeline`, `list_clips`, `list_markers`, `list_connected_clips`, `list_projects`, `analyze_timeline`, `list_clips`, `list_markers`, `list_connected_clips`,
@@ -59,18 +82,31 @@ mapeia nome → handler e delega ao Engine. O dispatch é um dicionário
### Voz (análise → decisão → aplicação) ### Voz (análise → decisão → aplicação)
`analyze_voice_features`, `build_voice_timeline`, `refine_voice_timeline`, `analyze_voice_features`, `build_voice_timeline`, `refine_voice_timeline`,
`remove_speakers`, `apply_voice_actions`, `get_voice_analysis_config`, `remove_speakers`, `apply_voice_actions`, `generate_voice_script`,
`save_voice_analysis_config`. `get_voice_analysis_config`, `save_voice_analysis_config`.
O fluxo é sempre o mesmo: `build_voice_timeline` mede (caro, roda uma vez) → O fluxo manual é: `build_voice_timeline` mede (caro, roda uma vez) →
o modelo decide os cortes → **`refine_voice_timeline` renormaliza sobre o que o modelo decide os cortes → **`refine_voice_timeline` renormaliza sobre o que
sobrou** (barato, sem reabrir áudio) e propõe as janelas de zoom → o modelo sobrou** (barato, sem reabrir áudio) e propõe as janelas de zoom → o modelo
corta a lista pelo ritmo → `apply_voice_actions` aplica. Pular a renormalização corta a lista pelo ritmo → `apply_voice_actions` aplica. Pular a renormalização
faz o ranking de ênfase apontar para as palavras erradas (ver faz o ranking de ênfase apontar para as palavras erradas (ver
`05_EXPERIENCIAS.md`). `05_EXPERIENCIAS.md`).
`generate_voice_script` é o fluxo **automático e fechado** (sem wizard, sem
copiar-e-colar): transcreve (cache) → `build_voice_timeline` → entrega a
timeline a um **modelo local Ollama** que dirige a edição → devolve o roteiro
legível (markdown) **e** o JSON de ações, e opcionalmente aplica num FCPXML.
O cliente fica em `fcpxml/llm_local.py`; o modelo é tratado como entrada não
confiável e cada ação é validada por `parse_actions`. Padrão:
`qwen2.5:7b-instruct-q4_K_M` (troca de `gemma3:12b` — não cabia em máquina de
8GB de RAM; Gemma 3 4B foi testado antes e falhou por apagar o roteiro
principal em vez de só cortar bastidor). Passe `model=` para usar outro
servido pelo Ollama.
### Legendas dinâmicas (geração → validação → aplicação) ### Legendas dinâmicas (geração → validação → aplicação)
`generate_dynamic_subtitles`, `validate_subtitle_layout`, `transcript_markers`. `generate_dynamic_subtitles`, `generate_plain_subtitles`,
`generate_subtitles_by_emphasis`, `validate_subtitle_layout`,
`transcript_markers`.
**Sempre gere e depois valide — nunca dê a geração como pronta sem **Sempre gere e depois valide — nunca dê a geração como pronta sem
`validate_subtitle_layout`.** A composição garante "sem sobreposição" só `validate_subtitle_layout`.** A composição garante "sem sobreposição" só
@@ -88,6 +124,23 @@ severidade probable/severe → investigar CADA colisão pela fração exata do
XML antes de mudar código (ver checklist abaixo) XML antes de mudar código (ver checklist abaixo)
``` ```
**`generate_subtitles_by_emphasis`** gera as duas legendas numa passada só —
mas não divide as palavras entre elas. A comum é gerada **completa, do início
ao fim do clipe**, sempre; a dinâmica é gerada só sobre as frases marcadas
como ênfase na etapa 5 (zoom aplicado, nível ≥ 1); e onde a dinâmica cobre um
trecho, os títulos comuns daquele trecho recebem `enabled="0"` — continuam no
XML (editáveis/reativáveis no Final Cut), só não são desenhados. É a tradução
literal de `10-revisao-humana.md` (skill `editar-por-voz`): "a frase de
ênfase recebe zoom E legenda dinâmica; as demais recebem legenda comum" —
sem nunca deixar um vão sem legenda nenhuma se a ênfase for desativada depois
(a comum já estava lá, só desligada). A decisão vem de
`<mídia>_phrase_actions.json["emphasis_spans"]`, escrito por
`save_phrase_review` quando o editor termina a etapa 5 — sem esse arquivo (ou
sem `zoom`/`text` marcados na revisão), a tool gera só a comum, tudo ligado,
e avisa no relatório ("Sem revisão de ênfase"). Não expõe overrides de estilo
por chamada — usa a config salva ("Legendas Dinâmicas"/plain); para estilo
pontual, use `generate_dynamic_subtitles`/`generate_plain_subtitles` direto.
**Antes de atribuir uma colisão ao gerador, confirme que é o gerador.** **Antes de atribuir uma colisão ao gerador, confirme que é o gerador.**
Um `<title>` de nome estranho (`ref` diferente, params tipo `Auto-Shrink`/ Um `<title>` de nome estranho (`ref` diferente, params tipo `Auto-Shrink`/
`Left Margin` que `_make_text_title_clip` nunca escreve) é conteúdo humano `Left Margin` que `_make_text_title_clip` nunca escreve) é conteúdo humano
@@ -144,7 +197,18 @@ Regras:
- Sempre retornam via `_text_result(text)` (envolve o texto em `TextContent` MCP). - Sempre retornam via `_text_result(text)` (envolve o texto em `TextContent` MCP).
## Para adicionar uma ferramenta nova ## Para adicionar uma ferramenta nova
1. Escrever a função no módulo do Engine (`fcpxml/…`) + testes.
2. Criar `handle_<nome>` em `server.py` seguindo o padrão acima. 1. **Escrever a função no Engine** (`fcpxml/…`) com testes. É aqui que mora o
3. Registrar no dicionário `TOOL_HANDLERS`. trabalho de verdade; o resto é encanamento.
2. **Criar `handle_<nome>`** em `server_tools/<categoria>.py`, seguindo o padrão
acima. Escolha a categoria pelo assunto, não pelo tamanho do arquivo.
3. **Declarar o schema** (`Tool(...)`) no mesmo módulo.
4. **Registrar** no `TOOL_HANDLERS` de `server.py`.
5. Rodar `./Engine/run_after_fix.sh`.
Se a ferramenta também deve aparecer no app, exponha um comando equivalente em
`admin/api/<assunto>.py` e registre na tabela de `admin/models_api.py` — ver
`08_APP_MACOS.md`. Uma capacidade que só existe como tool MCP **não existe para
quem usa o app** (foi exatamente o que aconteceu com `apply_voice_actions`,
`05_EXPERIENCIAS.md` #20).
4. Rodar `./Engine/run_after_fix.sh`. 4. Rodar `./Engine/run_after_fix.sh`.
@@ -1,5 +1,8 @@
# 04 — Testes, Fluxo de Trabalho e Estado Atual # 04 — Testes, Fluxo de Trabalho e Estado Atual
> **Escopo:** Como rodar e escrever testes, e o gate antes de commitar.
> **Não cobre:** O que testar em cada módulo (→ 02) · checklist de qualidade (→ 06)
## 1. Suíte de testes ## 1. Suíte de testes
**1032 testes em 24 arquivos** em `tests/`. Rode com `uv run pytest tests/ -v`. **1032 testes em 24 arquivos** em `tests/`. Rode com `uv run pytest tests/ -v`.
+231 -24
View File
@@ -11,6 +11,46 @@ houver uma correção ou trabalho em torno dele, **adicione um registro aqui**
antes de prosseguir. Um problema que se repete em várias tentativas é sinal de antes de prosseguir. Um problema que se repete em várias tentativas é sinal de
que merece entrada. que merece entrada.
> **Como usar:** o índice abaixo é o ponto de entrada. Procure o sintoma
> aqui primeiro; só abra a entrada completa (mais abaixo) se ela for a sua.
> As entradas ficam em ordem cronológica depois do índice.
## Resumo rápido (índice)
| # | Data | Problema | Estado |
|---|------|----------|--------|
| 1 | 2026-08-14 | Início do registro de experiências | `resolvido` |
| 4 | 2026-08-14 | XML fora da grade de frame em NTSC (23.976/29.97fps), confirmado no FCP | `resolvido` |
| 5 | 2026-08-14 | `TimeValue.from_timecode` corrompia segundos decimais em NTSC (3º ponto do bug) | `resolvido` |
| 6 | 2026-08-17 | Clipe-fantasma de 1 frame no início/fim após remoção de silêncio (padding sem vizinho na borda) | `resolvido` |
| 7 | 2026-08-17 | Legendas dinâmicas sobrepondo entre clipes (título conectado não é aparado pelo out-point do pai) | `resolvido` |
| 8 | 2026-08-17 | Importação recusada: `id` de `<text-style-def>` derivado do texto (acentos/espaços/dígito inicial) não é XML Name válido | `resolvido` |
| 9 | 2026-08-17 | Legendas palavra a palavra centradas em vez da composição progressiva diagramada (bloco por trecho, palavra-chave em display italic) | `resolvido` |
| 10 | 2026-08-17 | Cedilha/acentos da display italic invadindo a linha vizinha: empilhamento passou a usar a tinta real por classe de glifo | `resolvido` |
| 11 | 2026-08-18 | Preview das legendas dinâmicas desproporcional ao render do FCP (stagger/gap/canvas divergentes) e `inactive_color` exposto sem efeito | `resolvido` |
| 12 | 2026-08-18 | Espaço de coordenadas do modelo "Text": `fontSize`, `kerning` e `Position` no espaço do quadro — converter só o tamanho descolou o espaçamento | `resolvido` |
| 13 | 2026-08-19 | Reanálise de ênfase implementada no Engine mas sem ferramenta MCP — Fase 4 da skill era inexecutável | `resolvido` |
| 14 | 2026-08-19 | Offset sistemático de ~0,4s no timing por palavra (faster-whisper sem alinhamento forçado) — agora corrigido em pipeline por alinhamento forçado opcional | `resolvido` |
| 15 | 2026-08-19 | `add_zoom` perdia o enquadramento real (voltava a 100%) quando dois zooms caiam no mesmo clipe pós-corte; agora empilha ou substitui conforme as janelas se sobrepõem | `resolvido` |
| 16 | 2026-08-19 | `validate_subtitle_layout` acusava colisão severa em títulos que só se tocam na borda, por não-associatividade de float; 7 de 8 colisões reportadas no teste real eram falso positivo | `resolvido` |
| 17 | 2026-08-19 | Linha de ênfase das legendas dinâmicas sem limite de largura — palavra longa/maiúscula estourava o frame inteiro; auto-fit encolhe até caber, nunca abaixo do corpo | `resolvido` |
| 18 | 2026-08-19 | Legendas dinâmicas geradas com `bold="0" fontFace="Bold"` não renderizam no FCP — negrito deve ser `bold="1"` (atributo) e itálico `fontFace`+`italic="1"` | `resolvido` |
| 19 | 2026-08-19 | `output_dir` usado só como cerca de validação e nunca como destino — toda chamada entre pastas falhava acusando o caminho que ela mesma gerou | `resolvido` |
| 20 | 2026-08-19 | `apply_voice_actions` ausente da ponte e do encadeamento do app — dava para analisar e legendar, não para cortar | `resolvido` |
| 21 | 2026-08-19 | Teste ainda afirmava o default `zoom scale=1.3` removido do parser (agora vem do `zoom_scale` do usuário) | `resolvido` |
| 22 | 2026-08-19 | `VideoPlayer` (AVKit) aborta em runtime no app compilado por `swiftc` — etapa 5 fechava o app; trocado por `AVPlayerLayer` | `resolvido` |
| 23 | 2026-08-19 | Dividir `writer.py` em pacote quebrou `@patch('fcpxml.writer.subprocess')` — a suíte protege comportamento, não localização | `resolvido` |
| 24 | 2026-08-19 | `admin/test_models_api.py` existia mas estava fora de `testpaths` — 13 testes que nunca rodaram | `resolvido` |
| 25 | 2026-08-20 | `admin/api/shared.py` apontava para `admin/code` (inexistente) após a divisão — install editável mascarou o bug em toda validação anterior | `resolvido` |
| 26 | 2026-08-21 | `generate_voice_script` (IA local/Ollama) caía com "Falha ao gerar roteiro por IA local" — prompt embutia a timeline inteira (47k tokens) e estourava `num_ctx`; e `response.json()` de conexão caída escapava como `JSONDecodeError` | `resolvido` |
> Mantenha o índice acima sempre sincronizado com as entradas mais recentes.
---
## Entradas (ordem cronológica)
--- ---
## Como registrar (template de entrada) ## Como registrar (template de entrada)
@@ -79,9 +119,9 @@ Use o bloco abaixo como modelo. Uma entrada = um problema resolvido/reconhecido.
- **Por que isso importa mais do que parece:** o erro contamina toda decisão temporal a jusante — zoom disparava ~0,4s antes da palavra-alvo, `gap_before` subestimava pausas reais na mesma medida (o que afeta diretamente a régua de silêncio recém-adotada), e as folgas de corte saíam erradas nas emendas. - **Por que isso importa mais do que parece:** o erro contamina toda decisão temporal a jusante — zoom disparava ~0,4s antes da palavra-alvo, `gap_before` subestimava pausas reais na mesma medida (o que afeta diretamente a régua de silêncio recém-adotada), e as folgas de corte saíam erradas nas emendas.
- **Decisão tomada:** não rodei `remove_media_silence` bruto sobre o corte. A detecção (ffmpeg, limiar -30dB/0,5s) não distingue "batida entre frases dentro da régua de 1,5s" de "ar morto de emenda" — cortar ambos teria apertado frases fluidas. Corrigi os tempos manualmente medindo o ataque real nos pontos críticos (cabeça, 2 emendas, cauda, 3 zooms) e refiz o corte numa passada só. - **Decisão tomada:** não rodei `remove_media_silence` bruto sobre o corte. A detecção (ffmpeg, limiar -30dB/0,5s) não distingue "batida entre frases dentro da régua de 1,5s" de "ar morto de emenda" — cortar ambos teria apertado frases fluidas. Corrigi os tempos manualmente medindo o ataque real nos pontos críticos (cabeça, 2 emendas, cauda, 3 zooms) e refiz o corte numa passada só.
- **Solução adotada (paliativa, aplicada manualmente neste teste):** medir o RMS real com `ffmpeg -af astats=metadata=1:reset=1:length=0.05,ametadata=print` em janelas curtas ao redor de cada ponto crítico antes de fixar um corte ou zoom que dependa de precisão de frame. Não é o padrão do sistema — é o que cobre a lacuna até o alinhamento forçado existir. - **Solução adotada (paliativa, aplicada manualmente neste teste):** medir o RMS real com `ffmpeg -af astats=metadata=1:reset=1:length=0.05,ametadata=print` em janelas curtas ao redor de cada ponto crítico antes de fixar um corte ou zoom que dependa de precisão de frame. Não é o padrão do sistema — é o que cobre a lacuna até o alinhamento forçado existir.
- **Solução estrutural ainda pendente:** ligar o WhisperX (ou alinhamento forçado equivalente) em `transcribe.py`, o que levaria o erro de ~400ms para ~30ms e corrigiria zoom, corte e `gap_before` de uma vez, sem paliativo por projeto. Não implementado ainda — é mudança de pipeline, exige regerar todos os `_transcript.json`/`_voice_timeline.json` existentes. - **Solução estrutural implementada:** `transcribe.py` agora roda alinhamento forçado fonético (wav2vec2 via whisperx) como passo opcional pós-transcrição, em `fcpxml/forced_align.py` (classe `ForcedAligner`). O erro cai de ~400ms para ~30ms e corrige zoom, corte e `gap_before` de uma vez. É **dependência opcional** (`[align]` extra / pacote `whisperx` do PyPI) — quando ausente ou em qualquer falha, degrada e devolve os tempos brutos sem quebrar a transcrição. O `transcript` traz `"alignment": true/false` e o `voice_timeline` expõe `layers.alignment`, para quem lê o JSON saber se o offset manual ainda é necessário. Não reaproveitamos código da pasta `WHISPERX/` local (problemas conhecidos) — só a ideia documentada aqui. Exige regerar os `_transcript.json`/`_voice_timeline.json` existentes para aplicar nos caches antigos.
- **Aprendizado:** "não reestime tempos no olho" (critério 01) continua certo para decisão *editorial* — mas não cobre erro sistemático de *medição* na fonte dos tempos. Um offset constante e na mesma direção, em vários pontos do material, é sinal de bug no pipeline de transcrição, não de julgamento errado sobre o material. Vale conferir com uma amostra de áudio real antes de confiar cegamente em timestamp de word-level de qualquer fonte nova. - **Aprendizado:** "não reestime tempos no olho" (critério 01) continua certo para decisão *editorial* — mas não cobre erro sistemático de *medição* na fonte dos tempos. Um offset constante e na mesma direção, em vários pontos do material, é sinal de bug no pipeline de transcrição, não de julgamento errado sobre o material. Vale conferir com uma amostra de áudio real antes de confiar cegamente em timestamp de word-level de qualquer fonte nova.
- **Estado:** `parcialmente resolvido` — paliativo documentado e aplicado neste teste; correção estrutural (WhisperX) pendente de implementação. - **Estado:** `resolvido` — alinhamento forçado implementado em `transcribe.py`/`fcpxml/forced_align.py`; paliativo de medição manual mantido apenas para transcripts antigos sem `layers.alignment=true`.
--- ---
@@ -1183,27 +1223,194 @@ o outro; percentil entrega um punhado útil nos dois casos.
--- ---
## Resumo rápido (índice) ## 21 — 2026-08-19 — Teste travado no default antigo de `zoom scale`
| # | Data | Problema | Estado | - **Sintoma:** `tests/test_voice_actions.py::test_default_scale_when_absent`
|---|------|----------|--------| quebrando com `KeyError: 'scale'`, sem relação com a alteração em curso.
| 1 | 2026-08-14 | Início do registro de experiências | `resolvido` | - **Causa raiz:** `parse_actions` deixou de carimbar `scale=1.3` quando o
| 4 | 2026-08-14 | XML fora da grade de frame em NTSC (23.976/29.97fps), confirmado no FCP | `resolvido` | parâmetro vem ausente, justamente para que
| 5 | 2026-08-14 | `TimeValue.from_timecode` corrompia segundos decimais em NTSC (3º ponto do bug) | `resolvido` | `server_tools/_shared.py` use o `zoom_scale` configurado pelo usuário. O
| 6 | 2026-08-17 | Clipe-fantasma de 1 frame no início/fim após remoção de silêncio (padding sem vizinho na borda) | `resolvido` | teste continuou afirmando o default antigo, então passou a acusar como erro
| 7 | 2026-08-17 | Legendas dinâmicas sobrepondo entre clipes (título conectado não é aparado pelo out-point do pai) | `resolvido` | exatamente o comportamento desejado.
| 8 | 2026-08-17 | Importação recusada: `id` de `<text-style-def>` derivado do texto (acentos/espaços/dígito inicial) não é XML Name válido | `resolvido` | - **Solução adotada:** teste reescrito para o contrato novo — um `scale`
| 9 | 2026-08-17 | Legendas palavra a palavra centradas em vez da composição progressiva diagramada (bloco por trecho, palavra-chave em display italic) | `resolvido` | omitido tem que chegar ausente ao aplicador (`test_absent_scale_is_left_absent`).
| 10 | 2026-08-17 | Cedilha/acentos da display italic invadindo a linha vizinha: empilhamento passou a usar a tinta real por classe de glifo | `resolvido` | - **Aprendizado:** quando um default sai do parser e vira configuração, o teste
| 11 | 2026-08-18 | Preview das legendas dinâmicas desproporcional ao render do FCP (stagger/gap/canvas divergentes) e `inactive_color` exposto sem efeito | `resolvido` | que afirmava o valor antigo passa a defender o bug. Ao remover um default,
| 12 | 2026-08-18 | Espaço de coordenadas do modelo "Text": `fontSize`, `kerning` e `Position` no espaço do quadro — converter só o tamanho descolou o espaçamento | `resolvido` | procure o teste que o fixava no mesmo commit — senão ele fica dizendo o
| 13 | 2026-08-19 | Reanálise de ênfase implementada no Engine mas sem ferramenta MCP — Fase 4 da skill era inexecutável | `resolvido` | contrário do código, e a próxima pessoa perde tempo achando que quebrou algo.
| 14 | 2026-08-19 | Offset sistemático de ~0,4s no timing por palavra (faster-whisper sem alinhamento forçado) — corrigido manualmente no teste, WhisperX pendente | `parcialmente resolvido` | - **Estado:** `resolvido`
| 15 | 2026-08-19 | `add_zoom` perdia o enquadramento real (voltava a 100%) quando dois zooms caiam no mesmo clipe pós-corte; agora empilha ou substitui conforme as janelas se sobrepõem | `resolvido` |
| 16 | 2026-08-19 | `validate_subtitle_layout` acusava colisão severa em títulos que só se tocam na borda, por não-associatividade de float; 7 de 8 colisões reportadas no teste real eram falso positivo | `resolvido` |
| 17 | 2026-08-19 | Linha de ênfase das legendas dinâmicas sem limite de largura — palavra longa/maiúscula estourava o frame inteiro; auto-fit encolhe até caber, nunca abaixo do corpo | `resolvido` |
| 18 | 2026-08-19 | Legendas dinâmicas geradas com `bold="0" fontFace="Bold"` não renderizam no FCP — negrito deve ser `bold="1"` (atributo) e itálico `fontFace`+`italic="1"` | `resolvido` |
| 19 | 2026-08-19 | `output_dir` usado só como cerca de validação e nunca como destino — toda chamada entre pastas falhava acusando o caminho que ela mesma gerou | `resolvido` |
| 20 | 2026-08-19 | `apply_voice_actions` ausente da ponte e do encadeamento do app — dava para analisar e legendar, não para cortar | `resolvido` |
> Mantenha o índice acima sempre sincronizado com as entradas mais recentes. ---
## 22 — 2026-08-19 — `VideoPlayer` (AVKit) derruba o app compilado por `swiftc`
- **Sintoma:** "G-ART encerrou inesperadamente" (SIGABRT) toda vez que o
assistente entrava na etapa 5. Nada aparecia na tela antes do crash.
- **Causa raiz:** o app é montado invocando `swiftc` direto
(`MacApp/build_app.sh`), não pelo Xcode. Nesse modo o runtime não consegue
resolver a superclasse Objective-C de `VideoPlayer`:
`failed to demangle superclass of VideoPlayerView from mangled name
'So12AVPlayerViewC'` → `getSuperclassMetadata` chama `fatalError`. É erro de
runtime, então a compilação passa limpa e o problema só aparece ao abrir a
view.
- **Solução adotada:** trocar `VideoPlayer` por um `AVPlayerLayer` dentro de um
`NSViewRepresentable` (`PlayerSurface`/`PlayerLayerView` em
`PhraseReviewView.swift`). Só depende de AVFoundation, que linka normalmente.
Os controles de transporte já viviam na barra da timeline, então não se perde
nada com a chrome do AVKit.
- **Aprendizado:** compilar limpo não prova que um componente de framework
existe em runtime neste build. Ao usar uma view SwiftUI que embrulha uma
classe AppKit/ObjC (AVKit, WebKit, MapKit), abra a tela de fato antes de
concluir. Um harness pequeno (`swiftc` com os mesmos fontes + um `@main` que
monta só aquela view e sai) reproduz o crash em segundos, sem precisar
navegar o app inteiro até lá.
- **Estado:** `resolvido`
---
## 23 — 2026-08-19 — Dividir um módulo em pacote quebra quem faz `patch` nele
- **Sintoma:** ao transformar `fcpxml/writer.py` (4.199 linhas) no pacote
`fcpxml/writer/`, quatro testes passaram a falhar com
`AttributeError: module 'fcpxml.writer' has no attribute 'subprocess'` —
embora nenhuma linha de lógica tivesse mudado.
- **Causa raiz:** os testes usavam `@patch('fcpxml.writer.subprocess.run')`.
Isso não depende da API pública, e sim de *onde o import mora*: com o
módulo dividido, `subprocess` passou a ser importado por
`fcpxml/writer/document.py`, então o alvo do patch deixou de existir.
Re-exportar no `__init__` não resolveria — substituir
`fcpxml.writer.subprocess` não afeta a referência que `document` já tem.
- **Solução adotada:** apontar o patch para o módulo real
(`fcpxml.writer.document.subprocess.run`). Duas armadilhas do tipo foram
evitadas antes: imports relativos precisam de um ponto a mais ao descer um
nível (`from .models` → `from ..models`), inclusive os que ficam *dentro*
de funções, e o `__all__` precisa listar os nomes com underscore que o
resto do projeto já importava, senão a divisão vira quebra de API.
- **Aprendizado:** a suíte protege comportamento, não localização. Antes de
dividir um módulo, procure por `patch('<modulo>.` e por imports relativos
escondidos dentro de funções — são as duas coisas que uma refatoração
puramente mecânica quebra em silêncio, e as únicas que os testes pegam
tarde.
- **Estado:** `resolvido`
---
## 24 — 2026-08-19 — Teste existia, mas estava fora da suíte
- **Sintoma:** `admin/test_models_api.py` (13 testes) nunca rodava. Não
falhava — simplesmente não era coletado, então `models_api.py` figurava
como "coberto" sem que uma única asserção fosse executada em nenhum
commit.
- **Causa raiz:** `testpaths = ["tests"]` no `pyproject.toml`, com o pytest
rodando de `code/`. O arquivo morava em `admin/`, fora do alcance. Rodá-lo
à mão também falhava (`ModuleNotFoundError: admin`), porque a raiz do
repositório não entra no `sys.path` — ou seja, o único jeito de executá-lo
exigia saber de antemão que ele existia e como.
- **Solução adotada:** movido para `code/tests/test_models_api.py`, com o
insert da raiz do repositório no `sys.path` ao lado do import que precisa
dele. Passou a rodar no gate: 1441 → 1454 testes.
- **Aprendizado:** um teste fora de `testpaths` é pior que teste nenhum — ele
dá a sensação de rede sem ser rede. Ao mover ou criar teste fora da pasta
padrão, confirme que a contagem total subiu; se não subiu, ele não está
rodando. Vale também para o lint: `admin/` ainda não é coberto pelo
`run_after_fix.sh`, que roda só dentro de `code/`.
- **Estado:** `resolvido`
---
## 25 — 2026-08-20 — `admin/api/shared.py` apontava para `admin/code` (inexistente)
- **Sintoma:** app do usuário crashava em toda ação que passa por `server`
(ex: "Analisar voz"), com `ModuleNotFoundError: No module named
'server_tools'`. Sobreviveu a **duas rodadas de validação minha** na sessão
anterior — lint zero, 1454 testes verdes, comando testado manualmente pela
ponte — sem nenhuma delas pegar o bug.
- **Causa raiz:** ao dividir `admin/_shared.py` (#25 da sessão de refatoração,
commit `ffaebb3`) em `admin/api/*.py`, o cálculo
`Path(__file__).resolve().parent.parent / "code"` foi copiado sem ajuste.
No arquivo original (`admin/models_api.py`, direto em `admin/`), dois
`.parent` chegam na raiz do repo. Em `admin/api/shared.py`, um nível mais
fundo, dois `.parent` param em `admin/` — e `admin/code` nunca existiu.
`sys.path` nunca recebia `code/`, então `import server_tools` (que só
funciona com `code/` no path) falhava assim que qualquer handler tentava
`from server import ...`.
- **Por que passou pela validação anterior:** todo teste que exercitava esse
caminho importava `admin.api.*` **dentro do processo do pytest**, que já
roda com `cwd=code/` sob um venv com **install editável**
(`__editable__.fcp_mcp_server*.pth`) — isso já deixa `fcpxml`/`server_tools`
importáveis por conta própria, mascarando qualquer erro no cálculo manual
de `sys.path`. O teste manual pela ponte (`uv run python
admin/models_api.py analyze_voice ...`) tem o mesmo problema: `uv run`
ativa o mesmo venv com o mesmo install editável. **Só o app real, chamando
o fallback `python3` sem `uv` ou um venv sem o install editável, expõe o
bug** — que é exatamente a diferença entre o ambiente de teste e o do
usuário.
- **Solução adotada:** o cálculo de `sys.path` saiu de cada módulo de
comando e passou a existir **uma única vez**, em `admin/api/__init__.py`
— que roda antes de qualquer submódulo do pacote, então nenhum deles
precisa da própria cópia. `.parent.parent.parent` (três níveis: `api/` →
`admin/` → raiz → `code/`).
- **Como o teste de regressão foi validado (e por que precisou de duas
tentativas):** a primeira versão do teste também passava com o bug
presente, pelo mesmo motivo do parágrafo acima — rodava em processo com o
install editável ativo. Só ficou confiável rodando um `subprocess` limpo
que remove manualmente qualquer entrada `site-packages` de `sys.path`
antes de importar, isolando o mecanismo real que o `__init__.py` precisa
fornecer. Confirmado nos dois sentidos: falha com o bug reintroduzido,
passa com a correção (`tests/test_models_api.py::TestCodeDirResolution`).
- **Aprendizado:** um install editável no venv de teste é uma segunda fonte
de verdade que mascara bugs de `sys.path` — o mesmo defeito de "a suíte
passa mas o comportamento real não bate" da entrada #23, só que desta vez
nem *rodar o comando manualmente* pegou, porque o `uv run` usado para
testar caía no mesmo venv "de sorte" que o app não usa. Ao validar correção
de caminho/import, rodar num ambiente que não tenha as dependências
instaladas por fora do mecanismo sendo testado — ou o teste prova que o
ambiente de teste está bem configurado, não que o código está certo.
- **Estado:** `resolvido`
---
## Entrada #26 — Prompt da IA local estoura o contexto do Ollama (e erro de parse escapa)
- **Sintoma:** botão "Gerar roteiro por IA local" (etapa 4 do assistente)
devolvia "Falha ao gerar roteiro por IA local". Rodando a ponte direto, o
erro real aparecia como *"Server disconnected without sending a response"*
ou *"Connection refused"* do Ollama, e 0 decisões ("Decisões do modelo: 0").
- **Causa raiz (dupla):**
1. `build_edit_messages` embutia o JSON da voice timeline **inteiro** no
prompt. Uma gravação de 3min vira ~188KB / **~47k tokens** (cada palavra
carrega energia, pitch, arousal, valence, `samples`…). Como `num_ctx`
estava em 32768, o prompt estourava a janela e o Ollama **dropava a
conexão** sem resposta.
2. Quando a conexão cai sem resposta, `httpx` entrega um body vazio e
`response.json()` lançava `JSONDecodeError` — que **não** é
`httpx.HTTPError`, então escapava do `try/except` de `ollama_chat` e
virava a exceção genérica que o `cmd_generate_voice_script` transforma
em `ok:false` com a mensagem "Falha ao gerar roteiro por IA local: …".
- **Correção (em `fcpxml/llm_local.py` + `server_tools/voice.py`):**
- `build_edit_messages` agora projeta a timeline (**`_project_timeline`**):
mantém só `text`/`start`/`end`/`speaker`/`emphasis`/`pause_before` das
palavras e `id`/`name` dos locutores; descarta `layers`, `scales`,
`samples` e os floats de áudio. Caiu de ~47k para **~17k tokens** (69KB).
- Salvaguarda `_shrink_to_fit`: se ainda passar de `max_chars` (110k),
remove os `words` dos segmentos de menor `peak_emphasis` até caber.
- `ollama_chat` envolve `post`+`raise_for_status`+`json()` num único
`except Exception` que relança como `RuntimeError` claro — fim do
`JSONDecodeError` escapando.
- `_extract_json` agora desembrulha a lista de 1 elemento `[{source,
actions}]` que alguns modelos devolvem, senão o `parse_actions` tratava o
objeto-wrapper como uma ação sem `kind` e rejeitava tudo (0 decisões).
- `handle_generate_voice_script` levanta `RuntimeError` com a causa quando o
modelo não devolve nenhuma decisão utilizável, então o app mostra a
mensagem real ("O modelo local não devolveu decisões utilizáveis: …")
em vez do genérico.
- **Validação:** `tests/test_llm_local.py` ganhou `test_build_edit_messages_is_compact`
(prompt < raw, sem `samples`/`energy_raw`/`pitch_hz`) e
`test_ollama_chat_wraps_empty_response`. Ponte testada com Ollama mockado
nos dois sentidos (sucesso aplica; falha → `ok:false` com msg clara).
- **Estado:** `resolvido`
> **Aprendizado:** modelo local tem contexto finito — nunca embutir o objeto
> de análise cru no prompt; projetar só o que a decisão usa. E qualquer parse
> de resposta de servidor local deve tratar body vazio/quebrado como erro de
> transporte, não como sucesso mudo.
+3
View File
@@ -1,5 +1,8 @@
# 06 — Boas Práticas de Programação (G-ART) # 06 — Boas Práticas de Programação (G-ART)
> **Escopo:** Checklist de qualidade a aplicar antes de dar algo por pronto.
> **Não cobre:** Por onde começar uma tarefa (→ 09) · o que já quebrou (→ 05)
> **Propósito:** registrar as melhores práticas de programação a serem aplicadas > **Propósito:** registrar as melhores práticas de programação a serem aplicadas
> **sempre** que qualquer alteração ou correção for feita neste programa. > **sempre** que qualquer alteração ou correção for feita neste programa.
> Servem de checklist obrigatório antes de concluir qualquer mudança. > Servem de checklist obrigatório antes de concluir qualquer mudança.
+204
View File
@@ -0,0 +1,204 @@
# 08 — O app macOS (`MacApp/`) e o Assistente
> **Escopo:** O app SwiftUI e o Assistente: build, telas, ponte e a etapa 5.
> **Não cobre:** Engine Python (→ 02) · ferramentas MCP (→ 03)
O app SwiftUI é como o usuário opera o sistema sem abrir terminal nem conversar
com uma IA. São ~5.500 linhas em `MacApp/Sources/`, e ele **não tem lógica de
edição**: tudo que ele faz é montar argumentos, chamar a ponte Python e mostrar
o resultado.
Última varredura: 2026-08-19
---
## 1. Como o app é construído — leia antes de mexer
**Não existe `.xcodeproj` nem `Package.swift`.** O app é compilado invocando o
`swiftc` direto sobre `MacApp/Sources/*.swift`:
```bash
cd code && ./MacApp/build_app.sh # compila e monta o .app
admin/run_app.command # compila, fecha a instância antiga e abre (padrão de revisão)
```
Consequências práticas, todas já sentidas:
- **Arquivo novo em `Sources/` entra sozinho** no build. Não há lista de alvos.
- **Não dá para adicionar dependência SPM** sem antes migrar o build inteiro.
- **Compilar não prova que roda.** Componentes SwiftUI que embrulham classes
Objective-C podem falhar só em tempo de execução, ao abrir a tela. Foi o que
aconteceu com `VideoPlayer` (AVKit): compilava limpo e abortava ao abrir a
etapa 5 (`05_EXPERIENCIAS.md` #22). Por isso a regra: **alterou a interface,
abra a tela de fato.**
### Testando uma tela sem navegar o app inteiro
Um harness de vinte linhas compila os mesmos fontes com um `@main` próprio que
monta só a tela em questão. Reproduz crash de runtime em segundos:
```bash
swiftc -parse-as-library -sdk "$(xcrun --sdk macosx --show-sdk-path)" \
-target arm64-apple-macosx26.0 \
MacApp/Sources/PhraseReviewView.swift MacApp/Sources/PhraseReviewModel.swift \
MacApp/Sources/TimelineTracksView.swift MacApp/Sources/Models.swift \
MacApp/Sources/PythonBridge.swift /tmp/HarnessMain.swift -o /tmp/harness
```
O `@main` do harness carrega a tela, imprime o que interessa e chama
`NSApplication.shared.terminate` — dá para afirmar "abriu e funcionou" sem
depender de screenshot.
---
## 2. Estrutura das telas
| Arquivo | Linhas | Papel |
|---------|-------:|-------|
| `WizardView.swift` | 808 | **O Assistente** — fluxo guiado de 7 etapas |
| `TranscriptionView.swift` | 843 | Transcrição avulsa e processamento em lote |
| `ModelDownloadView.swift` | 545 | Catálogo e download de modelos Whisper |
| `CaptionsView.swift` | 545 | Legendas dinâmicas: estilo + preview ao vivo |
| `TimelineTracksView.swift` | 506 | Timeline com trilhas, zoom e playhead |
| `PhraseReviewModel.swift` | 429 | Estado da etapa 5: frases, player, zooms |
| `PhraseReviewView.swift` | 413 | Etapa 5: preview + inspector de frases |
| `VoiceAnalysisView.swift` | 322 | Parâmetros do motor de ênfase |
| `ProjectView.swift` | 293 | Inspeção do `.fcpxml` |
| `Models.swift` | 274 | Espelhos Swift do JSON da ponte |
| `PythonBridge.swift` | 230 | **A ponte** — ver seção 3 |
| `SubtitlePreviewView.swift` | 218 | Preview 9:16 das legendas |
| `App.swift` | 69 | `NavigationSplitView` e as abas |
Abas (`ActiveTab` em `App.swift`): Assistente · Projeto · Legendas · Análise de
Voz · Modelos · Sobre. As cinco últimas são "Avançado" — atalhos para operações
soltas. O Assistente é o caminho principal.
---
## 3. `PythonBridge.swift` — como o app fala com o Python
O app lança `admin/models_api.py` como **subprocesso**, passando o comando e um
JSON como `argv`, e lê **JSON-lines** no stdout.
```swift
PythonBridge.call(command: "build_phrase_review",
arguments: ["voice_timeline": path]) { result, error in … }
```
Dois pontos que já causaram problema e estão resolvidos no código — não os
desfaça sem entender:
- **`uv run` precisa rodar com cwd em `code/`.** O `uv` escolhe o ambiente pelo
diretório do processo, não pelo caminho do script. Rodar da raiz fazia o `uv`
criar um segundo `.venv` vazio e ignorar tudo que estava instalado em
`code/.venv` — librosa e pyannote instalavam com sucesso e o app insistia que
faltavam.
- **`scriptURL` procura `admin/models_api.py`** subindo diretórios a partir do
cwd, do bundle e do home. É o que faz o app funcionar tanto rodando do Xcode
quanto do `.app` montado.
- **O `sys.path` que torna `fcpxml`/`server_tools` importáveis dentro de
`admin/api/` mora só em `admin/api/__init__.py`.** Não copie esse cálculo
para um módulo de comando individual — foi exatamente essa cópia,
desatualizada em um nível de diretório, que quebrou toda ação que passa por
`server` (`05_EXPERIENCIAS.md` #25). E não confie em "testei com `uv run` e
funcionou": esse comando roda no mesmo venv com install editável que
mascara esse tipo de erro. O teste que pega de verdade é
`tests/test_models_api.py::TestCodeDirResolution`.
Para adicionar um comando: função em `admin/api/<assunto>.py`, registro na
tabela de `admin/models_api.py`, e `PythonBridge.call` do lado Swift. Os 37
comandos e seus formatos estão documentados no docstring de `models_api.py`.
---
## 4. O Assistente — as 7 etapas
`WizardStep` (`WizardView.swift`) é um enum sequencial; `canAdvance` decide
quando o botão "Continuar" libera.
| # | Etapa | O que acontece | Comando da ponte |
|---|-------|----------------|------------------|
| 1 | Projeto | Escolhe a pasta de saída e o `.fcpxml` | `project_config` |
| 2 | Transcrever | Transcreve toda a mídia do projeto | `transcribe` |
| 3 | Analisar voz | Mede ênfase, locutores, emoção | `analyze_voice` |
| 4 | Decisões da IA | Copia para o chat **ou** gera por IA local (Ollama/Gemma 3), aplica | `apply_voice_actions` / `generate_voice_script` |
| 5 | **Revisar ênfases** | Lapida frase a frase — ver seção 5 | `build_phrase_review` / `save_phrase_review` |
| 6 | Processar | Silêncios, preenchimento, legendas | vários, em cadeia |
| 7 | Concluído | Abre no FCP ou mostra no Finder | — |
**A etapa 4 tem duas saídas:**
- **Manual (chat):** o app monta o pedido pronto no clipboard (skill `editar-por-voz`) e recebe o JSON de volta — o julgamento de qual tomada usar e onde dar zoom fica com a IA numa conversa.
- **Automática (IA local):** botão "Gerar roteiro por IA local (Ollama/Gemma 3)". Ele manda a *voice timeline inteira* (o arquivo) junto com o brief para um modelo local (Ollama), que decide cortes/zooms/textos de uma vez, devolve o roteiro legível + o JSON de ações e já aplica no FCPXML (non-destructive). Não precisa sair do app nem colar nada. O modelo é escolhido num **picker que lista os modelos instalados no Ollama** (populado via `list_ollama_models` quando a etapa abre); se o Ollama estiver fora do ar, cai para um campo de texto livre. Troque para `llama3` etc. se tiver outro modelo. Requer o Ollama rodando em `localhost:11434`.
**Etapa 1 — armadilha registrada:** não escolha como "o projeto" um arquivo já
gerado pelo fluxo (`_voice_edit`, `_silence_removed`, …). Os cortes de voz
assumem timestamps da mídia **original**; reaplicá-los sobre um arquivo já
cortado desloca tudo em silêncio. O wizard avisa (`looksLikeGeneratedFile`).
---
## 5. Etapa 5 — a sala de edição
Única tela que ocupa a janela toda: o corpo do wizard é uma coluna de 640pt, e
essa etapa escapa dela porque precisa da largura (`step == .revisar` em
`WizardView.body`).
```
┌────────────────────────────┬──────────────┐
│ Preview (AVPlayerLayer) │ Inspector │
│ enquadrado no formato │ de frases │
│ de entrega do projeto │ │
├────────────────────────────┴──────────────┤
│ Timeline: 6 trilhas, zoom, playhead │
└───────────────────────────────────────────┘
```
**Trilhas:** zooms · frases · energia por palavra · emoção · locutor ·
roteiro/bastidor. Todas desenhadas sobre o mesmo eixo de tempo, com uma coluna
fixa à esquerda nomeando cada uma.
**O que o usuário decide por frase:** nível de ênfase (0–3), ativo/inativo,
texto, roteiro/bastidor e o trim das pontas. O trim anda em **fronteira de
palavra** — cortar é apontar para uma palavra, arrastando a borda do bloco ou
clicando na palavra no inspector.
**Zoom manual:** arrastar na timeline marca um trecho; botão direito cria um
zoom nele. O zoom guarda **só o quando** — escala e ramp vêm das configurações
de Análise de Voz no momento do render, então mudar lá restiliza todos.
**Decisões de implementação que parecem detalhe e não são:**
- **O preview não renderiza nada.** Ele toca a mídia original e *pula* os
trechos removidos. Renderizar para conferir um toggle poria minutos entre a
decisão e o resultado. O observador roda a 60 Hz porque o período dele é
exatamente quanto de material cortado dá para ouvir antes do pulo.
- **O enquadramento é o do projeto, não o da mídia.** As gravações são
horizontais e a entrega é vertical; o app lê o formato do `.fcpxml`
(`inspect`) e mostra o corte central aproximado, com um selo para alternar
para a mídia original. O enquadramento real de cada clipe vem do FCP — o
preview é aproximação, e o selo diz isso.
- **Nada é processado aqui.** "Continuar" grava o `_phrase_review.json` e o
`_phrase_actions.json` derivado dele. A geração é da etapa 6.
- **A revisão é sempre remontada da análise atual**, com as decisões salvas
reaplicadas por cima (`merge_saved_decisions`). Assim refazer a análise de voz
melhora a tela em vez de ficar mascarado por uma cópia velha; uma decisão cuja
frase se moveu mais de 0,25 s é descartada em vez de colar na frase errada.
---
## 6. Estado atual e o que falta
**Funciona e foi verificado:** carga das frases com decisões da IA, as 6
trilhas, seleção sincronizada nos três painéis, trim por palavra, zoom manual,
reprodução parando no ponto exato (erro de 0 ms medido), pulo dos trechos
removidos, enquadramento vertical, gravação ao avançar.
**Ainda em aberto:**
- **A etapa 6 não consome o `_phrase_review.json`.** A ligação — zoom e legenda
dinâmica só nas frases de ênfase, legenda comum no resto — é a próxima tarefa.
- **`MacApp/` não tem teste automatizado.** A rede é o harness da seção 1 e o
olho do usuário. Toda mudança de interface precisa ser aberta de fato.
- **O preview aproxima o reenquadramento vertical** pelo corte central; se os
clipes forem reposicionados no FCP, diverge.
+157
View File
@@ -0,0 +1,157 @@
# 09 — Manutenção: onde mexer, o que está aberto, o que dói
> **Escopo:** Por onde começar cada tipo de tarefa, o que está aberto e onde dói.
> **Não cobre:** Como as coisas funcionam — este doc roteia para quem explica
Este é o documento de rota. Os outros descrevem o que **é**; este diz o que
**fazer** e por onde começar quando chega uma implementação, uma melhoria ou
uma correção.
Última varredura: 2026-08-19 · 1.466 testes · lint zerado
---
## 1. Chegou uma tarefa — por onde começo?
| A tarefa é… | Comece em | Não esqueça |
|-------------|-----------|-------------|
| Regra nova de edição (corte, zoom, legenda) | `fcpxml/<módulo>` + teste | Expor na tool **e** na ponte, senão só metade dos usuários alcança |
| Corrigir XML que o FCP recusa | `fcpxml/writer/` + `dtd.py` | Validar contra o DTD real, não só o teste |
| Mudança visível na interface | `MacApp/Sources/` | **Abrir a tela** — compilar não prova nada (§4) |
| Comando novo para o app | `admin/api/<assunto>.py` | Registrar na tabela de `models_api.py` |
| Ferramenta MCP nova | `server_tools/<categoria>.py` | Schema `Tool(...)` + `TOOL_HANDLERS` |
| Ajuste de análise de voz | `fcpxml/voice_*`, `emphasis.py` | Regerar os `_voice_timeline.json` de teste |
| "Está lento" / "está errado" e não sei onde | §5 (mapa de sintomas) | — |
**A pergunta que resolve 90% das dúvidas de lugar:** essa lógica precisa saber
o que é uma tool MCP ou uma tela? Se não precisa — e quase nunca precisa — ela
vai para `fcpxml/`.
---
## 2. O que está aberto agora
Ordenado por quanto atrapalha, não por esforço.
### 2.1 A etapa 6 ignora a revisão de ênfases
O usuário lapida as frases na etapa 5, o `_phrase_review.json` é gravado — e a
etapa 6 ainda processa como antes. Falta ligar: **zoom e legenda dinâmica só
nas frases de ênfase, legenda comum no resto**. É a continuação natural do
trabalho da etapa 5 e o item mais valioso da lista.
→ `MacApp/Sources/WizardView.swift` (`finalizeProcessing`), `admin/api/subtitles.py`,
`fcpxml/phrase_review.py` (`emphasis_spans` já é produzido e ninguém consome).
### 2.2 Offset de ~400 ms no timing por palavra
O faster-whisper sem alinhamento forçado erra o início de cada palavra em
~0,4 s. Isso desloca zoom, corte e `gap_before` de uma vez. Há paliativo
aplicado por projeto; a correção estrutural é ligar o **WhisperX** (ou
alinhamento equivalente) em `transcribe.py`, o que levaria o erro para ~30 ms.
Custo real: regerar todos os `_transcript.json` e `_voice_timeline.json`
existentes. → `05_EXPERIENCIAS.md` #14, estado `parcialmente resolvido`.
### 2.3 `MacApp/` não tem teste automatizado
5.500 linhas de Swift sem uma asserção. A rede hoje é o harness manual (§4) e
o olho do usuário. Não é para sair criando suíte de UI — mas lógica pura que
foi parar na camada de tela (cálculo de trim, mapeamento de tempo) deveria
descer para o Python, onde já existe rede.
### 2.4 `admin/` fica fora do lint
`run_after_fix.sh` roda o ruff de dentro de `code/`, então `admin/` — 1.751
linhas de código que o app depende para funcionar — nunca é verificado.
Incluir mexe no gate, então é decisão consciente, não esquecimento.
### 2.5 Confirmações visuais pendentes no FCP
Várias entradas do `05_EXPERIENCIAS.md` estão marcadas como resolvidas *no XML*
— testes verdes, DTD válido — mas **pendentes de importação real no Final Cut**.
XML válido não é o mesmo que XML que renderiza como o esperado. Ao mexer em
legenda, zoom ou keyframe, a confirmação final é abrir no FCP.
### 2.6 Submódulo `WHISPERX` com conteúdo modificado e não commitado
Está fora dos commits de propósito, porque ninguém verificou o que mudou lá
dentro. Precisa ser olhado e resolvido — ou commitado, ou revertido.
---
## 3. Onde o código ainda é grande (e onde isso não é problema)
Quatro arquivos foram divididos (`writer.py`, `models.py`, `models_api.py`,
`_shared.py`): 6.685 linhas concentradas viraram 43 módulos.
O que sobrou grande, e o diagnóstico honesto de cada um:
| Arquivo | Linhas | Vale dividir? |
|---------|-------:|---------------|
| `fcpxml/text_layout.py` | 901 | **Não.** É diagramação — um assunto coeso. |
| `fcpxml/rough_cut.py` | 798 | **Não.** É geração de timeline, um assunto. |
| `fcpxml/model_manager.py` | 748 | Talvez: mistura catálogo, download e config. |
| `server_tools/voice.py` | 754 | Talvez, se crescer mais. |
| `MacApp/TranscriptionView.swift` | 843 | Sim, quando for mexer nela. |
| `MacApp/WizardView.swift` | 808 | Sim: sete etapas num `switch` só. |
**Critério, não número:** divida quando o arquivo tiver **assuntos** que não se
falam. Um arquivo grande de um assunto só é mais fácil de ler que seis arquivos
pequenos que você precisa abrir juntos. Código picado sem motivo atrapalha tanto
quanto arquivo gigante.
---
## 4. Checklist antes de dar algo por pronto
```bash
cd code && ./Engine/run_after_fix.sh # lint zerado + 1.466 testes
admin/run_app.command # se mexeu no app (padrão de revisão)
```
E, além do script:
- [ ] **Mexeu na interface? Abriu a tela?** Compilar não prova que roda —
`VideoPlayer` compilava e abortava (`05_EXPERIENCIAS.md` #22).
- [ ] **Mexeu em XML? Importou no FCP?** DTD válido ≠ renderiza certo.
- [ ] **Dividiu ou moveu módulo?** Procure `patch('<módulo>.` e imports
relativos dentro de funções — é o que quebra em silêncio (#23).
- [ ] **Criou teste fora de `code/tests/`?** Confirme que a contagem total
subiu. Teste fora de `testpaths` não roda e dá falsa sensação de rede (#24).
- [ ] **Problema estrutural ou erro recorrente?** Registre em
`05_EXPERIENCIAS.md` com o índice atualizado.
- [ ] **Documentação divergiu?** Corrija no mesmo commit. Doc velha engana mais
que doc ausente.
---
## 5. Mapa de sintomas → onde olhar
| Sintoma | Suspeite de | Arquivo |
|---------|-------------|---------|
| FCP recusa o arquivo ao importar | `id` inválido, ordem de filhos, timebase | `writer/validation.py`, `dtd.py` |
| Título importa mas não aparece | Template/uid Motion que não resolve | `writer/titles.py` |
| Corte no lugar errado | Tempo pós-corte usado como se fosse original | `voice_actions.py` (`shift_after_cuts`) |
| Zoom no lugar errado | Idem, ou offset de timing do Whisper | §2.2 |
| Legenda sobrepondo | Layout ou conteúdo antigo no arquivo | `collision.py`, `text_layout.py` |
| "Ênfase" apontando para palavra à toa | Falta renormalizar após o corte | `refine_voice_timeline` |
| App diz que falta librosa/pyannote | `uv run` com cwd errado | `PythonBridge.swift` (§3 do doc 08) |
| App crasha com `ModuleNotFoundError: server_tools` | `sys.path` de `admin/api/` mal calculado | `05_EXPERIENCIAS.md` #25 |
| Tela do app fecha o programa | Componente de framework que só falha em runtime | `05_EXPERIENCIAS.md` #22 |
| Comando existe no MCP mas não no app | Falta expor na ponte | `admin/api/`, #20 |
---
## 6. Convenções que não são negociáveis
Estão em `01_ARCHITECTURE.md` §2 e valem repetir as três que mais custaram:
1. **Tempo é fração racional.** Float para tempo produz drift que só aparece
depois de dez operações encadeadas.
2. **Ação de voz é sempre em tempo da mídia original.** Nunca pós-corte.
3. **Original nunca é sobrescrito.** Toda saída ganha sufixo.
---
## Documentos relacionados
- [01 Arquitetura](01_ARCHITECTURE.md) — camadas e onde cada coisa mora
- [02 Módulos](02_MODULES.md) — mapa do engine, módulo a módulo
- [03 Server/Tools](03_SERVER_TOOLS.md) — as 77 ferramentas MCP
- [04 Testes & Workflow](04_TESTS_AND_WORKFLOW.md)
- [05 Experiências](05_EXPERIENCIAS.md) — o que já quebrou e por quê
- [06 Boas Práticas](06_BOAS_PRATICAS.md)
- [08 App macOS](08_APP_MACOS.md) — o app e o Assistente
+12 -4
View File
@@ -12,6 +12,7 @@ struct GArtApp: App {
} }
enum ActiveTab: Hashable { enum ActiveTab: Hashable {
case wizard
case project case project
case captions case captions
case voiceAnalysis case voiceAnalysis
@@ -20,19 +21,23 @@ enum ActiveTab: Hashable {
} }
struct ContentView: View { struct ContentView: View {
@State private var activeTab: ActiveTab? = .project @State private var activeTab: ActiveTab? = .wizard
var body: some View { var body: some View {
NavigationSplitView { NavigationSplitView {
List(selection: $activeTab) { List(selection: $activeTab) {
Label("Assistente", systemImage: "wand.and.stars")
.tag(ActiveTab.wizard)
Section("Avançado") {
Label("Projeto", systemImage: "film") Label("Projeto", systemImage: "film")
.tag(ActiveTab.project) .tag(ActiveTab.project)
Label("Legendas Dinâmicas", systemImage: "captions.bubble") Label("Legendas", systemImage: "captions.bubble")
.tag(ActiveTab.captions) .tag(ActiveTab.captions)
Label("Análise de Voz", systemImage: "waveform") Label("Análise de Voz", systemImage: "waveform")
.tag(ActiveTab.voiceAnalysis) .tag(ActiveTab.voiceAnalysis)
Label("Modelos", systemImage: "tray.and.arrow.down") Label("Modelos", systemImage: "tray.and.arrow.down")
.tag(ActiveTab.models) .tag(ActiveTab.models)
}
Label("Sobre", systemImage: "info.circle") Label("Sobre", systemImage: "info.circle")
.tag(ActiveTab.about) .tag(ActiveTab.about)
} }
@@ -40,19 +45,22 @@ struct ContentView: View {
.navigationSplitViewColumnWidth(min: 180, ideal: 200) .navigationSplitViewColumnWidth(min: 180, ideal: 200)
} detail: { } detail: {
switch activeTab { switch activeTab {
case .wizard, nil:
WizardView().id(UUID())
.navigationTitle("Assistente")
case .project: case .project:
ProjectView().id(UUID()) ProjectView().id(UUID())
.navigationTitle("Projeto") .navigationTitle("Projeto")
case .captions: case .captions:
CaptionsView().id(UUID()) CaptionsView().id(UUID())
.navigationTitle("Legendas Dinâmicas") .navigationTitle("Legendas")
case .voiceAnalysis: case .voiceAnalysis:
VoiceAnalysisView().id(UUID()) VoiceAnalysisView().id(UUID())
.navigationTitle("Análise de Voz") .navigationTitle("Análise de Voz")
case .models: case .models:
ModelDownloadView().id(UUID()) ModelDownloadView().id(UUID())
.navigationTitle("Modelos") .navigationTitle("Modelos")
case .about, nil: case .about:
AboutView() AboutView()
.navigationTitle("Sobre") .navigationTitle("Sobre")
} }
+149 -10
View File
@@ -19,6 +19,7 @@ import UniformTypeIdentifiers
/// assunto. /// assunto.
struct CaptionsView: View { struct CaptionsView: View {
@State private var config = CaptionStyleConfig.defaults @State private var config = CaptionStyleConfig.defaults
@State private var plainConfig = PlainSubtitleConfig.defaults
@State private var isLoading = true @State private var isLoading = true
@State private var errorMessage: String? @State private var errorMessage: String?
@@ -30,15 +31,31 @@ struct CaptionsView: View {
@AppStorage("capSampleAfter") private var sampleAfter = "sua legenda" @AppStorage("capSampleAfter") private var sampleAfter = "sua legenda"
@AppStorage("capShowGuides") private var showsGuides = true @AppStorage("capShowGuides") private var showsGuides = true
private let fontChoices = [ /// Todas as famílias de fonte instaladas no macOS (sistema + usuário), as
"Helvetica Neue", "Helvetica", "Arial", "Avenir Next", /// usadas por padrão primeiro, para o seletor listar tudo sem hardcode.
"Futura", "SF Pro Display", "Georgia", "Impact", private static let installedFontFamilies: [String] = {
] var families = NSFontManager.shared.availableFontFamilies
.sorted { $0.localizedCaseInsensitiveCompare($1) == .orderedAscending }
let preferred = ["Helvetica Neue", "Playfair Display", "Georgia", "Didot"]
for family in preferred.reversed() {
if let idx = families.firstIndex(of: family) {
families.remove(at: idx)
families.insert(family, at: 0)
}
}
return families
}()
private let emphasisFontChoices = [ /// Lista para um picker: todas as famílias instaladas e, se o valor salvo
"Playfair Display", "Georgia", "Didot", "Futura", /// não estiver entre elas (ex.: fonte de outro Mac), ele entra no topo
"Avenir Next", "Times New Roman", "Helvetica Neue", "Impact", /// para o seletor continuar exibindo a escolha atual.
] private func fontChoices(for current: String) -> [String] {
var list = Self.installedFontFamilies
if !list.contains(current) {
list.insert(current, at: 0)
}
return list
}
private let emphasisFaceChoices = [ private let emphasisFaceChoices = [
"Medium Italic", "Italic", "Bold Italic", "Bold", "Regular", "Light Italic", "Medium Italic", "Italic", "Bold Italic", "Bold", "Regular", "Light Italic",
@@ -53,6 +70,13 @@ struct CaptionsView: View {
) )
} }
private func plainBound<T>(_ keyPath: WritableKeyPath<PlainSubtitleConfig, T>) -> Binding<T> {
Binding(
get: { plainConfig[keyPath: keyPath] },
set: { plainConfig[keyPath: keyPath] = $0; savePlain() }
)
}
private func colorBound(_ keyPath: WritableKeyPath<CaptionStyleConfig, String>) -> Binding<Color> { private func colorBound(_ keyPath: WritableKeyPath<CaptionStyleConfig, String>) -> Binding<Color> {
Binding( Binding(
get: { Color(rgbaString: config[keyPath: keyPath]) }, get: { Color(rgbaString: config[keyPath: keyPath]) },
@@ -60,6 +84,13 @@ struct CaptionsView: View {
) )
} }
private func plainColorBound(_ keyPath: WritableKeyPath<PlainSubtitleConfig, String>) -> Binding<Color> {
Binding(
get: { Color(rgbaString: plainConfig[keyPath: keyPath]) },
set: { plainConfig[keyPath: keyPath] = $0.fcpxmlColorString; savePlain() }
)
}
var body: some View { var body: some View {
HSplitView { HSplitView {
previewColumn previewColumn
@@ -152,6 +183,7 @@ struct CaptionsView: View {
positionSection positionSection
bodySection bodySection
emphasisSection emphasisSection
plainSubtitleSection
calibrationSection calibrationSection
} }
if let errorMessage { if let errorMessage {
@@ -195,7 +227,7 @@ struct CaptionsView: View {
private var bodySection: some View { private var bodySection: some View {
Section("Linhas de apoio") { Section("Linhas de apoio") {
Picker("Fonte", selection: bound(\.font)) { Picker("Fonte", selection: bound(\.font)) {
ForEach(fontChoices, id: \.self) { Text($0).tag($0) } ForEach(fontChoices(for: config.font), id: \.self) { Text($0).tag($0) }
} }
slider( slider(
"Tamanho", "Tamanho",
@@ -210,7 +242,7 @@ struct CaptionsView: View {
private var emphasisSection: some View { private var emphasisSection: some View {
Section("Palavra de ênfase") { Section("Palavra de ênfase") {
Picker("Fonte", selection: bound(\.emphasisFont)) { Picker("Fonte", selection: bound(\.emphasisFont)) {
ForEach(emphasisFontChoices, id: \.self) { Text($0).tag($0) } ForEach(fontChoices(for: config.emphasisFont), id: \.self) { Text($0).tag($0) }
} }
Picker("Estilo", selection: bound(\.emphasisFace)) { Picker("Estilo", selection: bound(\.emphasisFace)) {
ForEach(emphasisFaceChoices, id: \.self) { Text($0).tag($0) } ForEach(emphasisFaceChoices, id: \.self) { Text($0).tag($0) }
@@ -225,6 +257,35 @@ struct CaptionsView: View {
} }
} }
private var plainSubtitleSection: some View {
Section("Legenda comum") {
Picker("Fonte", selection: plainBound(\.font)) {
ForEach(fontChoices(for: plainConfig.font), id: \.self) { Text($0).tag($0) }
}
slider(
"Tamanho",
value: plainBound(\.fontSize), in: 28...300, step: 1,
readout: "\(Int(plainConfig.fontSize))pt",
help: "Tamanho da legenda comum editável no Final Cut."
)
slider(
"Máximo de palavras",
value: plainBound(\.maxWords), in: 1...14, step: 1,
readout: "\(Int(plainConfig.maxWords))",
help: "Quantidade máxima de palavras por bloco de legenda."
)
slider(
"Altura",
value: plainBound(\.positionY), in: -1200...300, step: 1,
readout: "\(Int(plainConfig.positionY))",
help: "Posição vertical da legenda comum no quadro; valores mais negativos descem."
)
ColorPicker("Cor", selection: plainColorBound(\.fontColor), supportsOpacity: true)
Toggle("Usar letra maiúscula", isOn: plainBound(\.uppercase))
Toggle("Manter vírgula e ponto", isOn: plainBound(\.keepPunctuation))
}
}
private var calibrationSection: some View { private var calibrationSection: some View {
Section { Section {
slider( slider(
@@ -305,18 +366,33 @@ struct CaptionsView: View {
} else if let error { } else if let error {
errorMessage = error errorMessage = error
} }
PythonBridge.call(command: "plain_subtitle_config") { plainResult, plainError in
DispatchQueue.main.async {
if let plainResult {
plainConfig = PlainSubtitleConfig(from: plainResult)
} else if let plainError {
errorMessage = plainError
}
isLoading = false isLoading = false
continuation.resume() continuation.resume()
} }
} }
} }
} }
}
}
private func save() { private func save() {
PythonBridge.call(command: "set_dynamic_subtitle_config", arguments: config.arguments()) { _, error in PythonBridge.call(command: "set_dynamic_subtitle_config", arguments: config.arguments()) { _, error in
DispatchQueue.main.async { errorMessage = error } DispatchQueue.main.async { errorMessage = error }
} }
} }
private func savePlain() {
PythonBridge.call(command: "set_plain_subtitle_config", arguments: plainConfig.arguments()) { _, error in
DispatchQueue.main.async { errorMessage = error }
}
}
} }
/// O estilo das legendas dinâmicas, no formato que a tela edita e o bridge /// O estilo das legendas dinâmicas, no formato que a tela edita e o bridge
@@ -402,6 +478,69 @@ struct CaptionStyleConfig {
} }
} }
struct PlainSubtitleConfig {
var font: String
var fontSize: Double
var fontColor: String
var maxWords: Double
var positionY: Double
var uppercase: Bool
var keepPunctuation: Bool
var textScale: Double
static let defaults = PlainSubtitleConfig(
font: "Helvetica Neue",
fontSize: 82,
fontColor: "1 1 1 1",
maxWords: 7,
positionY: -820,
uppercase: false,
keepPunctuation: true,
textScale: 2.0
)
init(from json: [String: Any]) {
let d = PlainSubtitleConfig.defaults
self.init(
font: json["font"] as? String ?? d.font,
fontSize: (json["font_size"] as? NSNumber)?.doubleValue ?? d.fontSize,
fontColor: json["font_color"] as? String ?? d.fontColor,
maxWords: (json["max_words"] as? NSNumber)?.doubleValue ?? d.maxWords,
positionY: (json["position_y"] as? NSNumber)?.doubleValue ?? d.positionY,
uppercase: json["uppercase"] as? Bool ?? d.uppercase,
keepPunctuation: json["keep_punctuation"] as? Bool ?? d.keepPunctuation,
textScale: (json["text_scale"] as? NSNumber)?.doubleValue ?? d.textScale
)
}
init(
font: String, fontSize: Double, fontColor: String, maxWords: Double,
positionY: Double, uppercase: Bool, keepPunctuation: Bool, textScale: Double
) {
self.font = font
self.fontSize = fontSize
self.fontColor = fontColor
self.maxWords = maxWords
self.positionY = positionY
self.uppercase = uppercase
self.keepPunctuation = keepPunctuation
self.textScale = textScale
}
func arguments() -> [String: Any] {
[
"font": font,
"font_size": Int(fontSize),
"font_color": fontColor,
"max_words": Int(maxWords),
"position_y": positionY,
"uppercase": uppercase,
"keep_punctuation": keepPunctuation,
"text_scale": textScale,
]
}
}
extension Color { extension Color {
/// Parses an FCPXML "R G B A" space-separated 0-1 string into a Color. /// Parses an FCPXML "R G B A" space-separated 0-1 string into a Color.
init(rgbaString: String) { init(rgbaString: String) {
+99 -1
View File
@@ -15,6 +15,11 @@ struct ModelDownloadView: View {
@State private var hfTokenText: String = "" @State private var hfTokenText: String = ""
@State private var numSpeakersText: String = "" @State private var numSpeakersText: String = ""
@State private var language: String = "auto" @State private var language: String = "auto"
@State private var acousticsAvailable: Bool?
@State private var acousticsMessage: String = ""
@State private var isInstallingAcoustics = false
@State private var acousticsInstallLog: String = ""
@State private var acousticsInstallError: String?
private let languages: [(String, String)] = [ private let languages: [(String, String)] = [
("auto", "Detectar automaticamente"), ("auto", "Detectar automaticamente"),
@@ -33,6 +38,7 @@ struct ModelDownloadView: View {
var body: some View { var body: some View {
Form { Form {
storageSection storageSection
acousticsSection
diarizationSection diarizationSection
if let errorMessage { if let errorMessage {
Section { Section {
@@ -65,7 +71,7 @@ struct ModelDownloadView: View {
} }
} }
.formStyle(.grouped) .formStyle(.grouped)
.task { await refresh() } .task { await refresh(); checkAcoustics() }
} }
// MARK: - Transcription language // MARK: - Transcription language
@@ -95,6 +101,98 @@ struct ModelDownloadView: View {
PythonBridge.call(command: "set_language", arguments: ["language": code]) { _, _ in } PythonBridge.call(command: "set_language", arguments: ["language": code]) { _, _ in }
} }
// MARK: - Acoustic analysis (librosa)
/// A ênfase de voz (pitch/energia) precisa do `librosa`, que é uma
/// dependência opcional — sem ela `layers.acoustics` vem `false` na
/// análise e a decisão de zoom fica sem base real. Antes disso só dava
/// pra descobrir lendo o JSON exportado; agora o app já diz e resolve.
private var acousticsSection: some View {
Section {
VStack(alignment: .leading, spacing: 10) {
if let acousticsAvailable {
Label(
acousticsMessage.isEmpty
? (acousticsAvailable ? "Disponível" : "Indisponível")
: acousticsMessage,
systemImage: acousticsAvailable ? "checkmark.circle.fill" : "exclamationmark.triangle.fill"
)
.font(.caption)
.foregroundStyle(acousticsAvailable ? Color.green : Color.orange)
} else {
Label("Verificando…", systemImage: "hourglass")
.font(.caption).foregroundStyle(.secondary)
}
if acousticsAvailable == false {
Button {
installAcoustics()
} label: {
if isInstallingAcoustics {
HStack { ProgressView().controlSize(.small); Text("Instalando…") }
} else {
Label("Instalar (uv sync --all-extras)", systemImage: "arrow.down.circle")
}
}
.disabled(isInstallingAcoustics)
if !acousticsInstallLog.isEmpty {
ScrollView {
Text(acousticsInstallLog)
.font(.system(.caption2, design: .monospaced))
.foregroundStyle(.secondary)
.frame(maxWidth: .infinity, alignment: .leading)
}
.frame(height: 90)
.background(RoundedRectangle(cornerRadius: 6).fill(Color.secondary.opacity(0.06)))
}
if let acousticsInstallError {
Label(acousticsInstallError, systemImage: "xmark.circle.fill")
.font(.caption).foregroundStyle(.red)
}
}
}
} header: {
Text("Análise Acústica (zoom por voz)")
} footer: {
Text("Mede a energia e o tom de voz de verdade, para os candidatos a zoom da edição por voz. Sem isso, a análise ainda transcreve e decide cortes pelo texto — só o zoom fica sem base acústica.")
.font(.caption)
.foregroundStyle(.secondary)
}
}
private func checkAcoustics() {
PythonBridge.call(command: "acoustics_capability") { result, err in
DispatchQueue.main.async {
guard let result, result["ok"] as? Bool == true else { return }
acousticsAvailable = result["available"] as? Bool
acousticsMessage = result["message"] as? String ?? ""
}
}
}
private func installAcoustics() {
isInstallingAcoustics = true
acousticsInstallLog = ""
acousticsInstallError = nil
// --all-extras, não só "intelligence": `uv sync` substitui o
// ambiente pelos extras pedidos em vez de somar, então um sync
// parcial aqui derrubaria dev/transcribe/diarização já instalados.
PythonBridge.runUV(arguments: ["sync", "--all-extras"]) { line in
DispatchQueue.main.async {
acousticsInstallLog += (acousticsInstallLog.isEmpty ? "" : "\n") + line
}
} completion: { code, err in
DispatchQueue.main.async {
isInstallingAcoustics = false
if code != 0 {
acousticsInstallError = err ?? "Falha ao instalar."
}
checkAcoustics()
}
}
}
// MARK: - Diarization // MARK: - Diarization
private var diarizationSection: some View { private var diarizationSection: some View {
+139
View File
@@ -116,6 +116,145 @@ struct ZoomClip: Identifiable {
} }
} }
/// One word inside a phrase, with the acoustics that justify an emphasis.
struct ReviewWord: Identifiable {
let id: Int
let text: String
let start: Double
let end: Double
let energy: Double
let emphasis: Double
init(id: Int, json: [String: Any]) {
self.id = id
text = json["text"] as? String ?? ""
start = json["start"] as? Double ?? 0
end = json["end"] as? Double ?? 0
energy = json["energy"] as? Double ?? 0
emphasis = json["emphasis"] as? Double ?? 0
}
}
/// A phrase in the review step — one spoken line plus the decision made about
/// it. Mirrors `fcpxml/phrase_review.py`; `emphasis` is 0–3 and everything
/// mutable here is what the editor is allowed to change.
struct ReviewPhrase: Identifiable {
let id: Int
let start: Double
let end: Double
var trimStart: Double
var trimEnd: Double
var text: String
let speaker: String
var active: Bool
var emphasis: Int
var track: String
let peakEmphasis: Double
let emotion: String
let emotionConfidence: Double
let takeBoundary: Bool
let gapBefore: Double
let reason: String
let words: [ReviewWord]
static let trackScript = "roteiro"
static let trackBackstage = "bastidor"
/// Delivery emotion as the analysis names it, in the user's language plus a
/// glyph — the label alone is too easy to skim past in a dense list.
static func emotionLabel(_ emotion: String) -> (String, String) {
switch emotion {
case "excited": return ("Empolgado", "flame")
case "tense": return ("Tenso", "bolt")
case "calm": return ("Calmo", "leaf")
case "reflective": return ("Reflexivo", "moon")
default: return ("Neutro", "circle")
}
}
init(json: [String: Any]) {
id = json["index"] as? Int ?? 0
start = json["start"] as? Double ?? 0
end = json["end"] as? Double ?? 0
trimStart = json["trim_start"] as? Double ?? (json["start"] as? Double ?? 0)
trimEnd = json["trim_end"] as? Double ?? (json["end"] as? Double ?? 0)
text = json["text"] as? String ?? ""
speaker = json["speaker"] as? String ?? ""
active = json["active"] as? Bool ?? true
emphasis = json["emphasis"] as? Int ?? 0
track = json["track"] as? String ?? ReviewPhrase.trackScript
peakEmphasis = json["peak_emphasis"] as? Double ?? 0
emotion = json["emotion"] as? String ?? "neutral"
emotionConfidence = json["emotion_confidence"] as? Double ?? 0
takeBoundary = json["take_boundary"] as? Bool ?? false
gapBefore = json["gap_before"] as? Double ?? 0
reason = json["reason"] as? String ?? ""
words = (json["words"] as? [[String: Any]] ?? [])
.enumerated().map { ReviewWord(id: $0.offset, json: $0.element) }
}
var asJSON: [String: Any] {
[
"index": id,
"start": start,
"end": end,
"trim_start": trimStart,
"trim_end": trimEnd,
"text": text,
"speaker": speaker,
"active": active,
"emphasis": emphasis,
"track": track,
"reason": reason,
]
}
var isBackstage: Bool { track == ReviewPhrase.trackBackstage }
var isTrimmed: Bool { trimStart > start + 0.001 || trimEnd < end - 0.001 }
var timecode: String {
String(format: "%02d:%02d", Int(start) / 60, Int(start) % 60)
}
/// The word boundaries a trim handle is allowed to land on.
func snap(_ time: Double, edge: TrimEdge) -> Double {
let boundaries = words.map { edge == .start ? $0.start : $0.end }.filter { $0 > 0 }
guard let nearest = boundaries.min(by: { abs($0 - time) < abs($1 - time) }) else {
return time
}
return nearest
}
}
enum TrimEdge { case start, end }
/// A punch-in the editor placed by hand over an arbitrary range, next to the
/// whole-phrase zoom that an emphasis level produces. It stores only *when* —
/// the scale and the ramp come from the Voice Analysis settings at render time.
struct ManualZoom: Identifiable {
let id = UUID()
var start: Double
var end: Double
/// Below this a punch-in has no room to ramp in and back out; the writer
/// rejects the window, so offering it would place nothing.
static let minimumDuration: Double = 0.4
init(start: Double, end: Double) {
self.start = start
self.end = end
}
init?(json: [String: Any]) {
guard let start = json["start"] as? Double, let end = json["end"] as? Double,
end - start >= ManualZoom.minimumDuration
else { return nil }
self.start = start
self.end = end
}
var asJSON: [String: Any] { ["start": start, "end": end] }
}
struct ZoomSegment: Identifiable { struct ZoomSegment: Identifiable {
let id: Int let id: Int
let start: Double let start: Double
+442
View File
@@ -0,0 +1,442 @@
import AVFoundation
import Combine
import Foundation
/// State behind the wizard's emphasis-review step.
///
/// Holds the phrases, the selection, and the player — together, because they
/// are one thing to the user: clicking a phrase moves the playhead, playing
/// moves the selection, and skipping a removed line only works if whoever owns
/// playback also knows which lines are removed.
///
/// The preview deliberately plays the *original* media and jumps over whatever
/// the edit removes, instead of rendering a cut first. Rendering to check a
/// toggle would put minutes between a decision and its result; jumping gives
/// the same reading instantly, and the real cut is generated later from the
/// exact same phrase list.
@MainActor
final class PhraseReviewModel: ObservableObject {
@Published var phrases: [ReviewPhrase] = []
@Published var selection: Int?
@Published var isLoading = false
@Published var errorMessage: String?
@Published var currentTime: Double = 0
@Published var isPlaying = false
@Published var pixelsPerSecond: Double = 40
@Published var skipRemoved = true
@Published var zooms: [ManualZoom] = []
/// In/out the editor dragged on the timeline, in source seconds.
@Published var rangeStart: Double?
@Published var rangeEnd: Double?
private(set) var source = ""
private(set) var sourcePath = ""
private(set) var duration: Double = 0
private(set) var speakers: [String] = []
private(set) var emotionAvailable = false
private(set) var player: AVPlayer?
private var voiceTimelinePath = ""
private var timeObserver: Any?
private var playbackLimit: Double?
let minPixelsPerSecond: Double = 8
let maxPixelsPerSecond: Double = 400
deinit {
if let timeObserver, let player {
player.removeTimeObserver(timeObserver)
}
}
// MARK: - Carregar
/// Builds the review from the voice timeline plus whatever the AI decided.
/// A review saved on a previous visit wins — see `cmd_build_phrase_review` —
/// UNLESS `fresh` is true, in which case that saved review is ignored and
/// `active`/`emphasis`/etc. come straight from this call's `decisionsJSON`.
/// Pass `fresh: true` when the decisions themselves changed since the
/// review was last built (the caller re-pasted/regenerated the AI's JSON
/// and re-ran `apply_voice_actions`) — otherwise the saved review from the
/// PREVIOUS decisions silently wins over the fresh cut it should reflect,
/// which is exactly the desync the wizard's "active" toggle showed against
/// the just-reapplied FCPXML.
func load(voiceTimelinePath: String, decisionsJSON: String,
outputFolder: String? = nil, mediaFolder: String? = nil, fresh: Bool = false) {
self.voiceTimelinePath = voiceTimelinePath
isLoading = true
errorMessage = nil
var arguments: [String: Any] = ["voice_timeline": voiceTimelinePath]
if let outputFolder { arguments["output_dir"] = outputFolder }
if let mediaFolder { arguments["media_dir"] = mediaFolder }
if fresh { arguments["fresh"] = true }
if let data = decisionsJSON.data(using: .utf8),
let parsed = try? JSONSerialization.jsonObject(with: data) {
arguments["actions"] = parsed
}
PythonBridge.call(command: "build_phrase_review", arguments: arguments) { [weak self] result, error in
Task { @MainActor in
guard let self else { return }
self.isLoading = false
if let error {
self.errorMessage = error
return
}
guard let result, result["ok"] as? Bool == true else {
self.errorMessage = result?["error"] as? String ?? "Não foi possível montar a revisão."
return
}
self.apply(result)
}
}
}
private func apply(_ result: [String: Any]) {
source = result["source"] as? String ?? ""
// The timeline JSON stores only the media's file name; the bridge
// resolves it to something openable (see phrase_review.resolve_source).
sourcePath = result["source_path"] as? String ?? ""
duration = result["duration"] as? Double ?? 0
speakers = result["speakers"] as? [String] ?? []
emotionAvailable = result["emotion_available"] as? Bool ?? false
phrases = (result["phrases"] as? [[String: Any]] ?? []).map { ReviewPhrase(json: $0) }
zooms = (result["zooms"] as? [[String: Any]] ?? []).compactMap { ManualZoom(json: $0) }
selection = phrases.first?.id
if let errors = result["errors"] as? [String], !errors.isEmpty {
errorMessage = "A IA mandou \(errors.count) decisão(ões) que não deu para ler — o resto foi aplicado."
}
preparePlayer()
}
/// Point the preview at a media file the user chose by hand — the way out
/// when the footage moved somewhere the automatic lookup can't reach.
func useMedia(at path: String) {
sourcePath = path
preparePlayer()
}
private func preparePlayer() {
guard !sourcePath.isEmpty, FileManager.default.fileExists(atPath: sourcePath) else {
player = nil
return
}
if let timeObserver, let player {
player.removeTimeObserver(timeObserver)
self.timeObserver = nil
}
let asset = AVURLAsset(url: URL(fileURLWithPath: sourcePath))
let player = AVPlayer(playerItem: AVPlayerItem(asset: asset))
self.player = player
// 60 Hz: the same observer drives the playhead *and* decides when to
// jump a removed stretch, so its period is the worst-case amount of cut
// material that can be heard before the skip lands. At 20 Hz that was an
// audible blip on every join.
let interval = CMTime(seconds: 1.0 / 60.0, preferredTimescale: 600)
timeObserver = player.addPeriodicTimeObserver(forInterval: interval, queue: .main) { [weak self] time in
Task { @MainActor in
self?.tick(time.seconds)
}
}
}
// MARK: - Reprodução
private func tick(_ time: Double) {
currentTime = time
guard isPlaying else { return }
// Playing a single phrase or a marked range stops at its out point
// instead of running on into the rest of the take.
if let limit = playbackLimit, time >= limit {
pause()
seek(to: limit)
return
}
if skipRemoved, let jump = nextKeptTime(after: time), jump > time {
seek(to: jump)
}
if let phrase = phrase(at: time), selection != phrase.id {
selection = phrase.id
}
}
/// Where playback should resume when `time` lands on removed material.
/// Returns nil when the time is on material that survives.
func nextKeptTime(after time: Double) -> Double? {
for phrase in phrases where time >= phrase.start - 0.001 && time < phrase.end {
if !phrase.active { return phrase.end }
if time < phrase.trimStart { return phrase.trimStart }
if time >= phrase.trimEnd { return phrase.end }
return nil
}
return nil
}
func togglePlay() {
if isPlaying {
pause()
} else {
playbackLimit = nil
play()
}
}
private func play() {
guard let player else { return }
if skipRemoved, let jump = nextKeptTime(after: currentTime) { seek(to: jump) }
player.play()
isPlaying = true
}
func pause() {
player?.pause()
isPlaying = false
playbackLimit = nil
}
/// Play exactly one span and stop — how a cut is judged: in context, at
/// speed, without hunting for the out point by hand.
func playRange(from start: Double, to end: Double) {
guard end > start else { return }
seek(to: start)
playbackLimit = end
player?.play()
isPlaying = true
}
func playSelectedPhrase() {
guard let selection, let phrase = phrases.first(where: { $0.id == selection })
else { return }
playRange(from: phrase.active ? phrase.trimStart : phrase.start,
to: phrase.active ? phrase.trimEnd : phrase.end)
}
func seek(to time: Double) {
currentTime = max(0, time)
player?.seek(to: CMTime(seconds: max(0, time), preferredTimescale: 600),
toleranceBefore: .zero, toleranceAfter: .zero)
}
/// Move the playhead to a phrase and select it.
func goTo(phraseID: Int) {
guard let phrase = phrases.first(where: { $0.id == phraseID }) else { return }
selection = phraseID
seek(to: phrase.active ? phrase.trimStart : phrase.start)
}
func phrase(at time: Double) -> ReviewPhrase? {
phrases.first { time >= $0.start && time < $0.end }
}
func selectNeighbour(_ delta: Int) {
guard let selection, let index = phrases.firstIndex(where: { $0.id == selection }) else {
if let first = phrases.first { goTo(phraseID: first.id) }
return
}
let next = min(max(0, index + delta), phrases.count - 1)
goTo(phraseID: phrases[next].id)
}
// MARK: - Edições
private func update(_ id: Int, _ change: (inout ReviewPhrase) -> Void) {
guard let index = phrases.firstIndex(where: { $0.id == id }) else { return }
change(&phrases[index])
}
func setEmphasis(_ level: Int, for id: Int) {
update(id) { $0.emphasis = min(3, max(0, level)) }
}
func toggleActive(_ id: Int) {
update(id) { $0.active.toggle() }
}
func setTrack(_ track: String, for id: Int) {
update(id) { $0.track = track }
}
func setText(_ text: String, for id: Int) {
update(id) { $0.text = text }
}
/// Trim a phrase's head or tail, landing on a word boundary.
/// A trim that would swallow the whole line is refused — deactivating the
/// phrase is the way to remove it, and doing it by accident with a drag
/// would lose the emphasis decision along with the line.
func trim(_ id: Int, edge: TrimEdge, to time: Double) {
update(id) { phrase in
let snapped = phrase.snap(time, edge: edge)
switch edge {
case .start:
let value = min(max(phrase.start, snapped), phrase.trimEnd - 0.1)
if value < phrase.trimEnd { phrase.trimStart = value }
case .end:
let value = max(min(phrase.end, snapped), phrase.trimStart + 0.1)
if value > phrase.trimStart { phrase.trimEnd = value }
}
}
}
func resetTrim(_ id: Int) {
update(id) { $0.trimStart = $0.start; $0.trimEnd = $0.end }
}
/// Trim everything before/after a given word — the text-first way to cut,
/// since the editor reads the line and points at where it should begin.
/// Clicking the word that is ALREADY that edge toggles it back off —
/// the trim on that side resets to the phrase's own start/end — so the
/// same click that sets a boundary also clears it, instead of needing
/// the separate "Inteira" button for a one-sided undo.
func trimToWord(_ word: ReviewWord, edge: TrimEdge, in id: Int) {
guard let phrase = phrases.first(where: { $0.id == id }) else { return }
let epsilon = 0.001
switch edge {
case .start where abs(word.start - phrase.trimStart) < epsilon:
update(id) { $0.trimStart = $0.start }
case .end where abs(word.end - phrase.trimEnd) < epsilon:
update(id) { $0.trimEnd = $0.end }
default:
trim(id, edge: edge, to: edge == .start ? word.start : word.end)
}
}
// MARK: - Trecho marcado e zooms
var hasRange: Bool {
guard let rangeStart, let rangeEnd else { return false }
return rangeEnd - rangeStart >= ManualZoom.minimumDuration
}
var rangeSpan: (start: Double, end: Double)? {
guard let rangeStart, let rangeEnd, rangeEnd > rangeStart else { return nil }
return (rangeStart, rangeEnd)
}
func setRange(from start: Double, to end: Double) {
rangeStart = min(start, end)
rangeEnd = max(start, end)
}
func clearRange() {
rangeStart = nil
rangeEnd = nil
}
/// Add a punch-in over the marked range. Scale and ramp are not stored:
/// they come from the "Análise de Voz" settings when the edit is rendered,
/// so changing the look there restyles every zoom at once.
func addZoomForRange() {
guard let span = rangeSpan, span.end - span.start >= ManualZoom.minimumDuration
else { return }
zooms.append(ManualZoom(start: span.start, end: span.end))
zooms.sort { $0.start < $1.start }
clearRange()
}
func addZoomForPhrase(_ id: Int) {
guard let phrase = phrases.first(where: { $0.id == id }) else { return }
zooms.append(ManualZoom(start: phrase.trimStart, end: phrase.trimEnd))
zooms.sort { $0.start < $1.start }
}
func removeZoom(_ id: UUID) {
zooms.removeAll { $0.id == id }
}
func zoom(at time: Double) -> ManualZoom? {
zooms.first { time >= $0.start && time <= $0.end }
}
func setEmphasisForAll(_ level: Int) {
for index in phrases.indices where phrases[index].active {
phrases[index].emphasis = level
}
}
// MARK: - Resumo e gravação
var emphasisCount: Int { phrases.filter { $0.active && $0.emphasis >= 1 }.count }
var removedCount: Int { phrases.filter { !$0.active }.count }
var keptDuration: Double {
phrases.filter { $0.active }.reduce(0) { $0 + ($1.trimEnd - $1.trimStart) }
}
// MARK: - Tempo compactado (sem os vãos do que foi cortado)
/// Kept spans of source media, in order, each carrying the position it
/// lands at once every removed stretch between phrases is squeezed out.
/// The timeline draws and scrubs in this space so it reads like the cut
/// itself instead of the raw take with holes in it.
private var keptSegments: [(rawStart: Double, rawEnd: Double, compactStart: Double)] {
var offset = 0.0
var segments: [(Double, Double, Double)] = []
for phrase in phrases.sorted(by: { $0.start < $1.start }) where phrase.active {
guard phrase.trimEnd > phrase.trimStart else { continue }
segments.append((phrase.trimStart, phrase.trimEnd, offset))
offset += phrase.trimEnd - phrase.trimStart
}
return segments
}
/// Maps a raw source-media time to its position on the compacted timeline.
/// Time inside removed material collapses to the boundary of the nearest
/// kept segment, so cut stretches take up no space at all.
func compactTime(_ raw: Double) -> Double {
let segments = keptSegments
for segment in segments {
if raw < segment.rawStart { return segment.compactStart }
if raw <= segment.rawEnd { return segment.compactStart + (raw - segment.rawStart) }
}
guard let last = segments.last else { return 0 }
return raw >= last.rawEnd ? last.compactStart + (last.rawEnd - last.rawStart) : 0
}
/// The inverse of `compactTime`: where a click on the compacted timeline
/// lands in the raw source media, for seeking and scrubbing.
func rawTime(fromCompact compact: Double) -> Double {
let segments = keptSegments
for segment in segments {
let compactEnd = segment.compactStart + (segment.rawEnd - segment.rawStart)
if compact <= compactEnd {
return segment.rawStart + max(0, compact - segment.compactStart)
}
}
return segments.last?.rawEnd ?? 0
}
/// Persists the edited review plus the actions derived from it. Called when
/// the wizard advances — the render itself happens in the next step.
/// Persists the edited review and hands back BOTH paths it wrote:
/// `review_path` (the human-readable `_phrase_review.json`) and
/// `actions_path` (`_phrase_actions.json`, the cut/zoom list derived from
/// it — what `finalizeProcessing` needs to actually apply the review's
/// active/inactive decisions instead of just filing them away).
func save(completion: @escaping (_ reviewPath: String?, _ actionsPath: String?) -> Void) {
guard !voiceTimelinePath.isEmpty, !phrases.isEmpty else {
completion(nil, nil)
return
}
let arguments: [String: Any] = [
"voice_timeline": voiceTimelinePath,
"source": source,
"duration": duration,
"speakers": speakers,
"phrases": phrases.map { $0.asJSON },
"zooms": zooms.map { $0.asJSON },
]
PythonBridge.call(command: "save_phrase_review", arguments: arguments) { result, error in
Task { @MainActor in
if let error {
completion(nil, nil)
_ = error
return
}
completion(result?["review_path"] as? String, result?["actions_path"] as? String)
}
}
}
}
+313
View File
@@ -0,0 +1,313 @@
import SwiftUI
/// The wizard's emphasis-review step.
///
/// Every decision here is about a *sentence* read from the original
/// transcription, so the phrases are listed in full — each line shows the text
/// as it will be said, a switch to keep or drop it from the cut, and the
/// emphasis level. Selecting a line in the list also selects its block on the
/// timeline below, and vice-versa.
struct PhraseReviewView: View {
@ObservedObject var model: PhraseReviewModel
var body: some View {
VSplitView {
VStack(spacing: 0) {
inspectorHeader
Divider()
List(selection: $model.selection) {
ForEach($model.phrases) { $phrase in
PhraseRow(phrase: $phrase, model: model)
.tag(phrase.id)
}
}
.listStyle(.inset)
.onChange(of: model.selection) { _, newValue in
if let newValue { model.goTo(phraseID: newValue) }
}
Divider()
summaryBar
}
.frame(minHeight: 240)
TimelineTracksView(model: model)
.frame(minHeight: 190, idealHeight: 210)
}
.overlay { if model.isLoading { loadingOverlay } }
.focusable()
.onKeyPress(.space) { model.togglePlay(); return .handled }
.onKeyPress(.return) { model.playSelectedPhrase(); return .handled }
.onKeyPress(.leftArrow) { model.selectNeighbour(-1); return .handled }
.onKeyPress(.rightArrow) { model.selectNeighbour(1); return .handled }
.onKeyPress(characters: .decimalDigits) { press in
guard let level = Int(press.characters), (0...3).contains(level),
let selection = model.selection else { return .ignored }
model.setEmphasis(level, for: selection)
return .handled
}
}
private var loadingOverlay: some View {
ZStack {
Color(nsColor: .windowBackgroundColor).opacity(0.85)
VStack(spacing: 10) {
ProgressView()
Text("Montando a revisão…").font(.callout).foregroundStyle(.secondary)
}
}
}
private var summaryBar: some View {
HStack(spacing: 16) {
summaryItem("text.quote", "\(model.phrases.count) frases")
summaryItem("sparkles", "\(model.emphasisCount) com ênfase")
summaryItem("scissors", "\(model.removedCount) fora do corte")
summaryItem("clock", durationLabel(model.keptDuration))
if !model.zooms.isEmpty {
summaryItem("plus.magnifyingglass", "\(model.zooms.count) zooms")
}
Spacer()
if let phrase = selectedPhrase, !phrase.reason.isEmpty {
Label(phrase.reason, systemImage: "brain")
.font(.caption).foregroundStyle(.secondary)
.lineLimit(1).truncationMode(.tail)
}
}
.padding(.horizontal, 14)
.padding(.vertical, 8)
}
private func summaryItem(_ icon: String, _ text: String) -> some View {
Label(text, systemImage: icon).font(.caption).foregroundStyle(.secondary)
}
private func durationLabel(_ seconds: Double) -> String {
String(format: "%02d:%02d finais", Int(seconds) / 60, Int(seconds) % 60)
}
private func pickMedia() {
let panel = NSOpenPanel()
panel.canChooseFiles = true
panel.canChooseDirectories = false
panel.allowsMultipleSelection = false
panel.prompt = "Usar esta mídia"
panel.message = model.source.isEmpty
? "Escolha o arquivo de vídeo desta gravação."
: "Escolha onde está \(model.source)."
if panel.runModal() == .OK, let url = panel.url {
model.useMedia(at: url.path)
}
}
private var selectedPhrase: ReviewPhrase? {
guard let selection = model.selection else { return nil }
return model.phrases.first { $0.id == selection }
}
// MARK: - Inspector de frases
private var inspectorPane: some View {
VStack(spacing: 0) {
inspectorHeader
Divider()
List(selection: $model.selection) {
ForEach($model.phrases) { $phrase in
PhraseRow(phrase: $phrase, model: model)
.tag(phrase.id)
}
}
.listStyle(.inset)
.onChange(of: model.selection) { _, newValue in
if let newValue { model.goTo(phraseID: newValue) }
}
}
}
private var inspectorHeader: some View {
VStack(alignment: .leading, spacing: 6) {
Text("Frases").font(.headline)
Text("Só as frases com ênfase recebem zoom e legenda dinâmica. O resto fica com legenda comum.")
.font(.caption).foregroundStyle(.secondary)
if !model.emotionAvailable {
Label("Emoção da fala não foi detectada nesta análise — ligue em Avançado → Análise de Voz e refaça o passo 3.",
systemImage: "waveform.path.ecg")
.font(.caption2).foregroundStyle(.secondary)
}
HStack(spacing: 8) {
Button("Limpar ênfases") { model.setEmphasisForAll(0) }
.buttonStyle(.link).font(.caption)
Spacer()
Text("0–3 no teclado · ← → navega")
.font(.caption2).foregroundStyle(.secondary)
}
}
.padding(12)
}
}
/// One phrase in the inspector: the line as it will be said, plus every
/// decision attached to it. Kept in one row on purpose — jumping to a separate
/// detail pane to set a toggle would double the clicks on the most repeated
/// action in the screen.
private struct PhraseRow: View {
@Binding var phrase: ReviewPhrase
@ObservedObject var model: PhraseReviewModel
@State private var isEditing = false
var body: some View {
VStack(alignment: .leading, spacing: 6) {
HStack(spacing: 6) {
Text(phrase.timecode)
.font(.system(.caption2, design: .monospaced))
.foregroundStyle(.secondary)
if phrase.takeBoundary {
Image(systemName: "scissors.badge.ellipsis")
.font(.caption2).foregroundStyle(.orange)
.help("Nova tomada começa aqui")
}
if phrase.isTrimmed {
Image(systemName: "arrow.left.and.right.square")
.font(.caption2).foregroundStyle(.blue)
.help("Frase cortada nas pontas")
}
if model.emotionAvailable {
emotionChip
}
Spacer()
Toggle("", isOn: $phrase.active)
.toggleStyle(.switch)
.controlSize(.mini)
.labelsHidden()
.help(phrase.active ? "No corte" : "Fora do corte")
}
if isEditing {
TextField("Texto da frase", text: $phrase.text, axis: .vertical)
.textFieldStyle(.roundedBorder)
.font(.callout)
.onSubmit { isEditing = false }
} else {
Text(phrase.text.isEmpty ? "(sem texto)" : phrase.text)
.font(.callout)
.foregroundStyle(phrase.active ? .primary : .secondary)
.strikethrough(!phrase.active)
.onTapGesture(count: 2) { isEditing = true }
}
HStack(spacing: 8) {
Picker("", selection: $phrase.emphasis) {
ForEach(0..<4, id: \.self) { level in
Text(EmphasisPalette.label(level)).tag(level)
}
}
.pickerStyle(.segmented)
.controlSize(.mini)
.labelsHidden()
.disabled(!phrase.active)
Picker("", selection: $phrase.track) {
Text("Roteiro").tag(ReviewPhrase.trackScript)
Text("Bastidor").tag(ReviewPhrase.trackBackstage)
}
.pickerStyle(.menu)
.controlSize(.mini)
.labelsHidden()
.frame(width: 92)
}
if model.selection == phrase.id && !phrase.words.isEmpty {
wordTrimmer
}
}
.padding(.vertical, 4)
.opacity(phrase.active ? 1 : 0.55)
}
/// The delivery emotion the acoustics suggest. Shown faded below its own
/// confidence: a guess the analysis is unsure about should not compete for
/// attention with the emphasis decision, which is the point of the row.
private var emotionChip: some View {
let (label, icon) = ReviewPhrase.emotionLabel(phrase.emotion)
return Label(label, systemImage: icon)
.font(.caption2)
.padding(.horizontal, 5)
.padding(.vertical, 1)
.background(
Capsule().fill(Color.secondary.opacity(0.12))
)
.foregroundStyle(phrase.emotionConfidence >= 0.5 ? .secondary : .tertiary)
.help("Emoção da entrega: \(label) — confiança \(Int(phrase.emotionConfidence * 100))%")
}
/// Trimming by pointing at the transcript: click a word to start the phrase
/// there, option-click to end it there. Same edit as dragging the block's
/// edge on the timeline, but reachable while reading the line.
private var wordTrimmer: some View {
VStack(alignment: .leading, spacing: 4) {
HStack(spacing: 4) {
Text("Cortar pelas palavras").font(.caption2).foregroundStyle(.secondary)
Spacer()
if phrase.isTrimmed {
Button("Inteira") { model.resetTrim(phrase.id) }
.buttonStyle(.link).font(.caption2)
}
}
FlowWords(words: phrase.words, phrase: phrase) { word, edge in
model.trimToWord(word, edge: edge, in: phrase.id)
}
Text("Clique = começa/desfaz aqui · ⌥clique = termina/desfaz aqui · sublinhado = ênfase da palavra")
.font(.caption2).foregroundStyle(.tertiary)
}
.padding(.top, 2)
}
}
/// The phrase's words as wrapping chips, dimmed where they fall outside the
/// trim and underlined where the acoustics mark them as an emphasis peak —
/// the same word-level signal `05-zoom.md` picks a punch-in's `start` from,
/// made visible instead of buried in the JSON.
private struct FlowWords: View {
let words: [ReviewWord]
let phrase: ReviewPhrase
let onTrim: (ReviewWord, TrimEdge) -> Void
var body: some View {
// A LazyVGrid with adaptive columns wraps chips without a custom layout;
// phrases are short enough that the slight raggedness beats the cost of
// hand-rolling a flow layout here.
LazyVGrid(columns: [GridItem(.adaptive(minimum: 44), spacing: 3)],
alignment: .leading, spacing: 3) {
ForEach(words) { word in
let kept = word.start >= phrase.trimStart - 0.001 && word.end <= phrase.trimEnd + 0.001
let level = EmphasisPalette.levelFromScore(word.emphasis)
Text(word.text)
.font(.caption2)
.fontWeight(level >= 2 ? .semibold : .regular)
.padding(.horizontal, 4)
.padding(.vertical, 2)
.background(
RoundedRectangle(cornerRadius: 3)
.fill(kept ? Color.accentColor.opacity(0.12) : Color.secondary.opacity(0.08))
)
.overlay(alignment: .bottom) {
if level >= 1 {
Rectangle()
.fill(EmphasisPalette.color(level))
.frame(height: 2)
.padding(.horizontal, 3)
}
}
.foregroundStyle(kept ? .primary : .secondary)
.strikethrough(!kept)
.help(
level >= 1
? "Ênfase \(EmphasisPalette.label(level).lowercased()) (\(Int(word.emphasis * 100))%)"
: "Sem ênfase"
)
.onTapGesture {
onTrim(word, NSEvent.modifierFlags.contains(.option) ? .end : .start)
}
}
}
}
}
+69 -2
View File
@@ -44,12 +44,26 @@ enum PythonBridge {
return ["python3", scriptURL.path] return ["python3", scriptURL.path]
} }
/// `admin/models_api.py` lives outside `code/`, but its dependencies
/// (`pyproject.toml`, `.venv`) live inside it. `uv run` picks the
/// environment from the process's cwd, not from the script path — so
/// running with cwd at the repo root made `uv` create/use a second,
/// empty `.venv` there, silently ignoring everything installed into
/// `code/.venv` (this cost a real debugging session: librosa/pyannote
/// installed successfully but the app kept reporting them missing).
/// Every `uv run` must share the same cwd as `uv sync` to see the same
/// environment.
static var workingDirectory: URL { static var workingDirectory: URL {
projectRoot codeDirectory
}
/// Directory containing `pyproject.toml` — where `uv sync` must run from.
static var codeDirectory: URL {
projectRoot.appendingPathComponent("code")
} }
/// Locate `uv` on PATH or in common install locations. /// Locate `uv` on PATH or in common install locations.
private static func findUV() -> String? { static func findUV() -> String? {
if let onPath = which("uv") { return onPath } if let onPath = which("uv") { return onPath }
let candidates = [ let candidates = [
"/usr/local/bin/uv", "/usr/local/bin/uv",
@@ -148,6 +162,59 @@ enum PythonBridge {
} }
} }
// MARK: - uv sync (installing optional extras, e.g. acoustic analysis)
/// Runs `uv <arguments>` from `codeDirectory` (where `pyproject.toml`
/// lives), streaming each output line as plain text — used for
/// `sync --extra intelligence` so "Modelos" can install the librosa
/// extra without the user opening a terminal.
static func runUV(arguments: [String],
onLine: @escaping (String) -> Void,
completion: @escaping (Int, String?) -> Void) {
guard let uv = findUV() else {
completion(1, "uv não encontrado. Instale com: curl -LsSf https://astral.sh/uv/install.sh | sh")
return
}
let process = Process()
process.executableURL = URL(fileURLWithPath: "/usr/bin/env")
process.arguments = [uv] + arguments
process.currentDirectoryURL = codeDirectory
let pipe = Pipe()
process.standardOutput = pipe
process.standardError = pipe
var buffer = ""
let lock = NSLock()
pipe.fileHandleForReading.readabilityHandler = { handle in
let data = handle.availableData
guard !data.isEmpty, let s = String(data: data, encoding: .utf8) else { return }
lock.lock()
buffer += s
let parts = buffer.split(separator: "\n", omittingEmptySubsequences: false)
buffer = String(parts.last ?? "")
let lines = parts.dropLast()
lock.unlock()
for line in lines where !line.isEmpty { onLine(String(line)) }
}
process.terminationHandler = { p in
pipe.fileHandleForReading.readabilityHandler = nil
lock.lock()
let last = buffer.trimmingCharacters(in: .whitespacesAndNewlines)
buffer = ""
lock.unlock()
if !last.isEmpty { onLine(last) }
completion(Int(p.terminationStatus), p.terminationStatus == 0 ? nil : "uv sync terminou com erro (código \(p.terminationStatus)).")
}
do {
try process.run()
} catch {
completion(1, error.localizedDescription)
}
}
// MARK: - Convenience: single JSON result // MARK: - Convenience: single JSON result
/// Runs a command and delivers the first parsed JSON document as the result. /// Runs a command and delivers the first parsed JSON document as the result.
@@ -0,0 +1,529 @@
import SwiftUI
/// Colors shared by the timeline and the inspector, so a block and its row in
/// the list always read as the same thing.
enum EmphasisPalette {
static func color(_ level: Int) -> Color {
switch level {
case 1: return Color.blue
case 2: return Color.orange
case 3: return Color.pink
default: return Color.secondary
}
}
static func label(_ level: Int) -> String {
switch level {
case 1: return "Leve"
case 2: return "Média"
case 3: return "Forte"
default: return "Sem"
}
}
/// The same 0–3 tiers a phrase's `emphasis` uses, derived from a raw 0–1
/// acoustic score — the thresholds `10-revisao-humana.md` documents for
/// deriving a phrase's level from `peak_emphasis` when no explicit zoom
/// was set, reused here per WORD so a word chip and a phrase row read as
/// the same scale.
static func levelFromScore(_ score: Double) -> Int {
switch score {
case ..<0.25: return 0
case ..<0.45: return 1
case ..<0.65: return 2
default: return 3
}
}
static func speakerColor(_ speaker: String, among speakers: [String]) -> Color {
let palette: [Color] = [.teal, .purple, .green, .indigo, .brown, .cyan]
guard let index = speakers.firstIndex(of: speaker) else { return .gray }
return palette[index % palette.count]
}
}
/// The timeline strip: four stacked tracks over one shared time axis.
///
/// Phrases are laid out as real views rather than drawn into a Canvas, because
/// every one of them is a target — click to select, drag its edge to trim,
/// right-click to change emphasis. The dense per-word energy track *is* a
/// Canvas: it has thousands of bars and nothing to hit.
struct TimelineTracksView: View {
@ObservedObject var model: PhraseReviewModel
private let rulerHeight: CGFloat = 18
private let phraseHeight: CGFloat = 46
private let energyHeight: CGFloat = 34
private let stripHeight: CGFloat = 12
private let handleWidth: CGFloat = 8
private let gutterWidth: CGFloat = 92
private let trackSpacing: CGFloat = 4
private var pps: CGFloat { CGFloat(model.pixelsPerSecond) }
/// Width follows the *kept* duration, not the raw take's — the timeline
/// draws the cut, so removed stretches take no horizontal space.
private var contentWidth: CGFloat { max(320, CGFloat(model.keptDuration) * pps) }
/// Name, icon and height of each lane, in the order they stack. The gutter
/// and the tracks are built from this one list so a label can never drift
/// off the lane it names.
private var lanes: [(label: String, icon: String, height: CGFloat)] {
[
("", "", rulerHeight),
("Zooms", "plus.magnifyingglass", stripHeight + 6),
("Frases", "text.quote", phraseHeight),
("Energia", "waveform", energyHeight),
("Emoção", "face.smiling", stripHeight),
("Locutor", "person.wave.2", stripHeight),
("Roteiro", "list.bullet.rectangle", stripHeight),
]
}
var body: some View {
VStack(spacing: 0) {
toolbar
Divider()
HStack(alignment: .top, spacing: 0) {
gutter
Divider()
timelineScroller
}
}
.background(Color(nsColor: .underPageBackgroundColor))
}
/// Fixed column naming each lane. Without it the stripes are six colours
/// with no way to tell which one is emotion and which one is the speaker.
private var gutter: some View {
VStack(alignment: .leading, spacing: trackSpacing) {
ForEach(lanes.indices, id: \.self) { index in
let lane = lanes[index]
HStack(spacing: 4) {
if !lane.icon.isEmpty {
Image(systemName: lane.icon).font(.system(size: 9))
}
Text(lane.label).font(.system(size: 10))
Spacer(minLength: 0)
}
.foregroundStyle(.secondary)
.frame(height: lane.height, alignment: .center)
}
}
.padding(.horizontal, 8)
.padding(.vertical, 8)
.frame(width: gutterWidth, alignment: .leading)
}
private var timelineScroller: some View {
ScrollViewReader { proxy in
ScrollView([.horizontal]) {
ZStack(alignment: .topLeading) {
VStack(alignment: .leading, spacing: trackSpacing) {
ruler
zoomTrack
phraseTrack
energyTrack
emotionTrack
speakerTrack
scriptTrack
}
.frame(width: contentWidth, alignment: .leading)
rangeOverlay
playhead
// Anchors the auto-scroll: one invisible marker per
// phrase, so selecting a line off-screen brings it in.
ForEach(model.phrases) { phrase in
Color.clear
.frame(width: 1, height: 1)
.offset(x: x(phrase.start))
.id(phrase.id)
}
}
.padding(.vertical, 8)
.contentShape(Rectangle())
.gesture(scrubGesture)
.contextMenu { timelineMenu }
}
.onChange(of: model.selection) { _, newValue in
guard let newValue else { return }
withAnimation(.easeOut(duration: 0.2)) {
proxy.scrollTo(newValue, anchor: .center)
}
}
}
}
// MARK: - Barra de controles
private var toolbar: some View {
HStack(spacing: 12) {
Button {
model.togglePlay()
} label: {
Image(systemName: model.isPlaying ? "pause.fill" : "play.fill")
}
.buttonStyle(.borderless)
.help("Reproduzir (espaço)")
.disabled(model.player == nil)
Text(timecode(model.currentTime))
.font(.system(.caption, design: .monospaced))
.foregroundStyle(.secondary)
Button {
model.playSelectedPhrase()
} label: {
Image(systemName: "play.rectangle")
}
.buttonStyle(.borderless)
.help("Tocar só a frase selecionada (⏎)")
.disabled(model.player == nil || model.selection == nil)
Toggle("Pular removidos", isOn: $model.skipRemoved)
.toggleStyle(.checkbox)
.font(.caption)
.help("Durante a reprodução, salta os trechos desativados — mostra como o corte ficou.")
Button {
model.addZoomForRange()
} label: {
Label("Zoom no trecho", systemImage: "plus.magnifyingglass")
}
.buttonStyle(.borderless)
.font(.caption)
.disabled(!model.hasRange)
.help("Arraste na timeline para marcar um trecho e crie um zoom nele. A escala vem de Análise de Voz.")
Spacer()
legend
Spacer()
Image(systemName: "minus.magnifyingglass").foregroundStyle(.secondary)
Slider(value: $model.pixelsPerSecond,
in: model.minPixelsPerSecond...model.maxPixelsPerSecond)
.frame(width: 130)
Image(systemName: "plus.magnifyingglass").foregroundStyle(.secondary)
}
.padding(.horizontal, 12)
.padding(.vertical, 8)
}
private var legend: some View {
HStack(spacing: 10) {
ForEach(0..<4, id: \.self) { level in
HStack(spacing: 4) {
RoundedRectangle(cornerRadius: 2)
.fill(EmphasisPalette.color(level))
.frame(width: 10, height: 10)
Text(EmphasisPalette.label(level)).font(.caption2)
}
}
}
.foregroundStyle(.secondary)
}
// MARK: - Trilhas
private var ruler: some View {
Canvas { context, size in
let step = tickStep()
var time = 0.0
while time <= model.keptDuration {
let position = compactX(time)
context.stroke(
Path { $0.move(to: CGPoint(x: position, y: size.height - 6))
$0.addLine(to: CGPoint(x: position, y: size.height)) },
with: .color(.secondary.opacity(0.5))
)
context.draw(
Text(timecode(time)).font(.system(size: 9, design: .monospaced))
.foregroundColor(.secondary),
at: CGPoint(x: position + 18, y: 6)
)
time += step
}
}
.frame(width: contentWidth, height: rulerHeight)
}
private var phraseTrack: some View {
ZStack(alignment: .topLeading) {
RoundedRectangle(cornerRadius: 4)
.fill(Color.secondary.opacity(0.06))
.frame(width: contentWidth, height: phraseHeight)
ForEach(model.phrases) { phrase in
phraseBlock(phrase)
}
}
.frame(width: contentWidth, height: phraseHeight, alignment: .topLeading)
}
@ViewBuilder
private func phraseBlock(_ phrase: ReviewPhrase) -> some View {
let isSelected = model.selection == phrase.id
let color = EmphasisPalette.color(phrase.emphasis)
let fullWidth = max(2, width(from: phrase.start, to: phrase.end))
let keptWidth = max(1, width(from: phrase.trimStart, to: phrase.trimEnd))
ZStack(alignment: .topLeading) {
// The whole line, dim — what is there before the edit.
RoundedRectangle(cornerRadius: 4)
.fill(color.opacity(phrase.active ? 0.15 : 0.10))
.frame(width: fullWidth, height: phraseHeight)
// What survives: the kept span, drawn solid over it.
RoundedRectangle(cornerRadius: 4)
.fill(color.opacity(phrase.active ? 0.55 : 0.12))
.frame(width: keptWidth, height: phraseHeight)
.offset(x: width(from: phrase.start, to: phrase.trimStart))
Text(phrase.text)
.font(.system(size: 10))
.lineLimit(2)
.padding(.horizontal, 4)
.frame(width: fullWidth, height: phraseHeight, alignment: .topLeading)
.foregroundStyle(phrase.active ? .primary : .secondary)
.strikethrough(!phrase.active)
RoundedRectangle(cornerRadius: 4)
.stroke(isSelected ? Color.accentColor : color.opacity(0.4),
lineWidth: isSelected ? 2 : 1)
.frame(width: fullWidth, height: phraseHeight)
if isSelected && phrase.active {
trimHandle(phrase, edge: .start)
trimHandle(phrase, edge: .end)
}
}
.frame(width: fullWidth, height: phraseHeight, alignment: .topLeading)
.offset(x: x(phrase.start))
.contentShape(Rectangle())
.onTapGesture { model.goTo(phraseID: phrase.id) }
.contextMenu { phraseMenu(phrase) }
.help(phrase.reason.isEmpty ? phrase.text : "\(phrase.text)\n— \(phrase.reason)")
}
private func trimHandle(_ phrase: ReviewPhrase, edge: TrimEdge) -> some View {
let offset = edge == .start
? width(from: phrase.start, to: phrase.trimStart)
: width(from: phrase.start, to: phrase.trimEnd) - handleWidth
return RoundedRectangle(cornerRadius: 2)
.fill(Color.accentColor)
.frame(width: handleWidth, height: phraseHeight)
.offset(x: offset)
.gesture(
DragGesture(minimumDistance: 1)
.onChanged { value in
let compactOrigin = model.compactTime(phrase.start)
let time = model.rawTime(fromCompact: compactOrigin + Double(value.location.x / pps))
model.trim(phrase.id, edge: edge, to: time)
}
)
.help(edge == .start ? "Arraste para cortar o começo (pula de palavra em palavra)"
: "Arraste para cortar o fim (pula de palavra em palavra)")
}
@ViewBuilder
private func phraseMenu(_ phrase: ReviewPhrase) -> some View {
Button("Tocar esta frase") {
model.goTo(phraseID: phrase.id)
model.playSelectedPhrase()
}
Button(phrase.active ? "Remover do corte" : "Trazer de volta") {
model.toggleActive(phrase.id)
}
Button("Adicionar zoom nesta frase") { model.addZoomForPhrase(phrase.id) }
Divider()
ForEach(0..<4, id: \.self) { level in
Button("Ênfase: \(EmphasisPalette.label(level))") {
model.setEmphasis(level, for: phrase.id)
}
}
Divider()
Button(phrase.isBackstage ? "Marcar como roteiro" : "Marcar como bastidor") {
model.setTrack(phrase.isBackstage ? ReviewPhrase.trackScript : ReviewPhrase.trackBackstage,
for: phrase.id)
}
if phrase.isTrimmed {
Divider()
Button("Desfazer corte da frase") { model.resetTrim(phrase.id) }
}
}
/// Per-word energy/emphasis, straight from the voice timeline — the closest
/// thing to a waveform without opening the audio again.
private var energyTrack: some View {
Canvas { context, size in
for phrase in model.phrases {
for word in phrase.words {
let start = x(word.start)
let barWidth = max(1, width(from: word.start, to: word.end) - 1)
let height = size.height * CGFloat(max(0.04, word.energy))
let rect = CGRect(x: start, y: size.height - height,
width: barWidth, height: height)
let color = word.emphasis >= 0.65 ? Color.pink
: word.emphasis >= 0.45 ? Color.orange
: Color.secondary
context.fill(Path(rect),
with: .color(color.opacity(phrase.active ? 0.6 : 0.2)))
}
}
}
.frame(width: contentWidth, height: energyHeight)
.background(RoundedRectangle(cornerRadius: 4).fill(Color.secondary.opacity(0.06)))
}
private var speakerTrack: some View {
stripTrack { phrase in
EmphasisPalette.speakerColor(phrase.speaker, among: model.speakers)
}
}
private var scriptTrack: some View {
stripTrack { phrase in phrase.isBackstage ? Color.gray : Color.mint }
}
private func stripTrack(_ color: @escaping (ReviewPhrase) -> Color) -> some View {
Canvas { context, size in
for phrase in model.phrases {
let rect = CGRect(x: x(phrase.start), y: 0,
width: max(1, width(from: phrase.start, to: phrase.end)),
height: size.height)
context.fill(Path(roundedRect: rect, cornerRadius: 2),
with: .color(color(phrase).opacity(phrase.active ? 0.7 : 0.2)))
}
}
.frame(width: contentWidth, height: stripHeight)
}
private var playhead: some View {
Rectangle()
.fill(Color.red)
.frame(width: 1.5)
.offset(x: x(model.currentTime))
.allowsHitTesting(false)
}
/// One gesture, two meanings, decided by whether the mouse moved: a click
/// parks the playhead, a drag marks in/out. Splitting them across separate
/// controls would mean choosing a tool before every action, which is
/// exactly the ceremony this screen is meant to avoid.
private var scrubGesture: some Gesture {
DragGesture(minimumDistance: 0)
.onChanged { value in
let from = model.rawTime(fromCompact: Double(value.startLocation.x / pps))
let to = model.rawTime(fromCompact: Double(value.location.x / pps))
if abs(value.translation.width) > 3 {
model.setRange(from: from, to: to)
model.seek(to: min(from, to))
} else {
model.clearRange()
model.seek(to: to)
}
}
}
/// The marked in/out, drawn over every track so the span reads against the
/// phrases and the energy at once.
private var rangeOverlay: some View {
Group {
if let span = model.rangeSpan {
Rectangle()
.fill(Color.accentColor.opacity(0.18))
.overlay(Rectangle().stroke(Color.accentColor.opacity(0.6), lineWidth: 1))
.frame(width: max(1, width(from: span.start, to: span.end)))
.offset(x: x(span.start))
.allowsHitTesting(false)
}
}
}
@ViewBuilder
private var timelineMenu: some View {
if model.hasRange, let span = model.rangeSpan {
Button("Adicionar zoom no trecho (\(secondsLabel(span.end - span.start)))") {
model.addZoomForRange()
}
Button("Tocar o trecho") { model.playRange(from: span.start, to: span.end) }
Button("Limpar seleção") { model.clearRange() }
} else {
Text("Arraste na timeline para marcar um trecho")
}
if let zoom = model.zoom(at: model.currentTime) {
Divider()
Button("Remover o zoom daqui") { model.removeZoom(zoom.id) }
}
}
private func secondsLabel(_ seconds: Double) -> String {
String(format: "%.1fs", seconds)
}
/// Punch-ins, on their own lane above the script: they are a second layer
/// over the same time, not a property of a phrase.
private var zoomTrack: some View {
ZStack(alignment: .topLeading) {
RoundedRectangle(cornerRadius: 3)
.fill(Color.secondary.opacity(0.06))
.frame(width: contentWidth, height: stripHeight + 6)
ForEach(model.zooms) { zoom in
RoundedRectangle(cornerRadius: 3)
.fill(Color.yellow.opacity(0.55))
.overlay(
Image(systemName: "plus.magnifyingglass")
.font(.system(size: 8)).foregroundStyle(.black.opacity(0.6))
)
.frame(width: max(6, width(from: zoom.start, to: zoom.end)),
height: stripHeight + 6)
.offset(x: x(zoom.start))
.help("Zoom marcado — \(secondsLabel(zoom.end - zoom.start)). A escala vem de Análise de Voz.")
.contextMenu {
Button("Remover este zoom") { model.removeZoom(zoom.id) }
}
}
}
.frame(width: contentWidth, height: stripHeight + 6, alignment: .topLeading)
}
/// Delivery emotion per phrase — the fourth signal to read against the text.
private var emotionTrack: some View {
stripTrack { phrase in
switch phrase.emotion {
case "excited": return .orange
case "tense": return .red
case "calm": return .blue
case "reflective": return .purple
default: return .secondary
}
}
}
// MARK: - Escala
/// Pixel position of a raw source-media time, after collapsing whatever
/// lies between it and the previous kept phrase.
private func x(_ time: Double) -> CGFloat { compactX(model.compactTime(time)) }
/// Pixel position of a time already in the compacted (edited) timeline —
/// used for the ruler and playhead, which think in that space directly.
private func compactX(_ compactTime: Double) -> CGFloat { CGFloat(compactTime) * pps }
private func width(from: Double, to: Double) -> CGFloat {
max(0, CGFloat(model.compactTime(to) - model.compactTime(from)) * pps)
}
/// Ruler spacing that keeps labels ~80pt apart at any zoom.
private func tickStep() -> Double {
let candidates: [Double] = [1, 2, 5, 10, 15, 30, 60, 120, 300, 600]
let wanted = 80 / Double(pps)
return candidates.first { $0 >= wanted } ?? 600
}
private func timecode(_ seconds: Double) -> String {
let total = Int(seconds.rounded(.down))
return String(format: "%02d:%02d", total / 60, total % 60)
}
}
+3 -3
View File
@@ -271,7 +271,7 @@ struct TranscriptionView: View {
Toggle("Marcar o que foi dito na timeline", isOn: $batchMarkers) Toggle("Marcar o que foi dito na timeline", isOn: $batchMarkers)
Divider() Divider()
Toggle("Exportar legendas SRT", isOn: $batchSubtitles) Toggle("Gerar legenda comum (texto editável no FCP)", isOn: $batchSubtitles)
Divider() Divider()
batchOptionRow( batchOptionRow(
@@ -793,7 +793,7 @@ struct TranscriptionView: View {
if batchFillers { operations.append("remove_filler_words") } if batchFillers { operations.append("remove_filler_words") }
if batchPhrases { operations.append("edit_by_transcript") } if batchPhrases { operations.append("edit_by_transcript") }
if batchMarkers { operations.append("transcript_markers") } if batchMarkers { operations.append("transcript_markers") }
if batchSubtitles { operations.append("export_srt") } if batchSubtitles { operations.append("generate_plain_subtitles") }
// Runs last, on the timing already cut by any earlier steps (see the // Runs last, on the timing already cut by any earlier steps (see the
// "Abrir no Final Cut Pro" fallback chain and exportSubtitles()'s own // "Abrir no Final Cut Pro" fallback chain and exportSubtitles()'s own
// preference for `processedPath` — same reasoning). // preference for `processedPath` — same reasoning).
@@ -834,7 +834,7 @@ struct TranscriptionView: View {
} }
let nextPath = result?["path"] as? String ?? currentPath let nextPath = result?["path"] as? String ?? currentPath
if operation == "remove_silences" { processedPath = nextPath } if operation == "remove_silences" { processedPath = nextPath }
if operation == "export_srt" { subtitlePaths = result?["paths"] as? [String] ?? [] } if operation == "generate_plain_subtitles" { subtitlePaths = [nextPath] }
if operation == "generate_dynamic_subtitles" { dynamicSubtitlesPath = nextPath } if operation == "generate_dynamic_subtitles" { dynamicSubtitlesPath = nextPath }
processBatchStep(operations, index: index + 1, currentPath: nextPath, outputFolder: outputFolder) processBatchStep(operations, index: index + 1, currentPath: nextPath, outputFolder: outputFolder)
} }
+68 -4
View File
@@ -23,6 +23,7 @@ struct VoiceAnalysisView: View {
} else { } else {
energySection energySection
emphasisSection emphasisSection
zoomSection
weightsSection weightsSection
emotionSection emotionSection
resetSection resetSection
@@ -90,6 +91,44 @@ struct VoiceAnalysisView: View {
} }
} }
private var zoomSection: some View {
Section {
sliderRow(
title: "Zoom na ênfase",
value: $config.zoomScale,
range: 1.0...3.0,
readout: "\(Int(config.zoomScale * 100))%",
help: "Fator aplicado nos punch-ins de ênfase. 130% equivale a escala 1,30 no Final Cut."
)
Picker("Movimento", selection: $config.zoomMode) {
Text("Zoom in e out").tag("in_out")
Text("Só zoom in").tag("in")
Text("Só zoom out").tag("out")
}
.onChange(of: config.zoomMode) { _, _ in save() }
sliderRow(
title: "Velocidade do zoom in",
value: $config.zoomEaseIn,
range: 0.05...2.0,
readout: String(format: "%.2fs", config.zoomEaseIn),
help: "Duração da entrada do zoom. Menor é mais rápido."
)
sliderRow(
title: "Velocidade do zoom out",
value: $config.zoomEaseOut,
range: 0.01...2.0,
readout: String(format: "%.2fs", config.zoomEaseOut),
help: "Duração da saída do zoom. Menor é mais seco."
)
} header: {
Text("Zoom de Ênfase")
} footer: {
Text("Esses valores viram o padrão para ações de zoom que não trouxerem scale/ease/ease_out no JSON da edição por voz.")
.font(.caption)
.foregroundStyle(.secondary)
}
}
// MARK: - Emoção // MARK: - Emoção
private var emotionSection: some View { private var emotionSection: some View {
@@ -129,13 +168,14 @@ struct VoiceAnalysisView: View {
title: String, title: String,
value: Binding<Double>, value: Binding<Double>,
range: ClosedRange<Double> = 0...1, range: ClosedRange<Double> = 0...1,
readout: String? = nil,
help: String? = nil help: String? = nil
) -> some View { ) -> some View {
VStack(alignment: .leading, spacing: 2) { VStack(alignment: .leading, spacing: 2) {
HStack { HStack {
Text(title) Text(title)
Spacer() Spacer()
Text(String(format: "%.2f", value.wrappedValue)) Text(readout ?? String(format: "%.2f", value.wrappedValue))
.monospacedDigit() .monospacedDigit()
.foregroundStyle(.secondary) .foregroundStyle(.secondary)
} }
@@ -188,6 +228,10 @@ struct VoiceAnalysisConfig {
var weightDuration: Double var weightDuration: Double
var emotionEnabled: Bool var emotionEnabled: Bool
var emotionSensitivity: Double var emotionSensitivity: Double
var zoomScale: Double
var zoomMode: String
var zoomEaseIn: Double
var zoomEaseOut: Double
static let defaults = VoiceAnalysisConfig( static let defaults = VoiceAnalysisConfig(
energyThreshold: 0.5, energyThreshold: 0.5,
@@ -198,7 +242,11 @@ struct VoiceAnalysisConfig {
weightPause: 0.15, weightPause: 0.15,
weightDuration: 0.10, weightDuration: 0.10,
emotionEnabled: false, emotionEnabled: false,
emotionSensitivity: 0.5 emotionSensitivity: 0.5,
zoomScale: 1.30,
zoomMode: "in_out",
zoomEaseIn: 0.25,
zoomEaseOut: 0.04
) )
init( init(
@@ -210,7 +258,11 @@ struct VoiceAnalysisConfig {
weightPause: Double, weightPause: Double,
weightDuration: Double, weightDuration: Double,
emotionEnabled: Bool, emotionEnabled: Bool,
emotionSensitivity: Double emotionSensitivity: Double,
zoomScale: Double,
zoomMode: String,
zoomEaseIn: Double,
zoomEaseOut: Double
) { ) {
self.energyThreshold = energyThreshold self.energyThreshold = energyThreshold
self.emphasisThreshold = emphasisThreshold self.emphasisThreshold = emphasisThreshold
@@ -221,6 +273,10 @@ struct VoiceAnalysisConfig {
self.weightDuration = weightDuration self.weightDuration = weightDuration
self.emotionEnabled = emotionEnabled self.emotionEnabled = emotionEnabled
self.emotionSensitivity = emotionSensitivity self.emotionSensitivity = emotionSensitivity
self.zoomScale = zoomScale
self.zoomMode = zoomMode
self.zoomEaseIn = zoomEaseIn
self.zoomEaseOut = zoomEaseOut
} }
/// Lê a resposta do bridge, caindo no padrão para qualquer campo ausente. /// Lê a resposta do bridge, caindo no padrão para qualquer campo ausente.
@@ -236,7 +292,11 @@ struct VoiceAnalysisConfig {
weightPause: weights["pause_before"] as? Double ?? defaults.weightPause, weightPause: weights["pause_before"] as? Double ?? defaults.weightPause,
weightDuration: weights["duration"] as? Double ?? defaults.weightDuration, weightDuration: weights["duration"] as? Double ?? defaults.weightDuration,
emotionEnabled: json["emotion_enabled"] as? Bool ?? defaults.emotionEnabled, emotionEnabled: json["emotion_enabled"] as? Bool ?? defaults.emotionEnabled,
emotionSensitivity: json["emotion_sensitivity"] as? Double ?? defaults.emotionSensitivity emotionSensitivity: json["emotion_sensitivity"] as? Double ?? defaults.emotionSensitivity,
zoomScale: json["zoom_scale"] as? Double ?? defaults.zoomScale,
zoomMode: json["zoom_mode"] as? String ?? defaults.zoomMode,
zoomEaseIn: json["zoom_ease_in"] as? Double ?? defaults.zoomEaseIn,
zoomEaseOut: json["zoom_ease_out"] as? Double ?? defaults.zoomEaseOut
) )
} }
@@ -253,6 +313,10 @@ struct VoiceAnalysisConfig {
], ],
"emotion_enabled": emotionEnabled, "emotion_enabled": emotionEnabled,
"emotion_sensitivity": emotionSensitivity, "emotion_sensitivity": emotionSensitivity,
"zoom_scale": zoomScale,
"zoom_mode": zoomMode,
"zoom_ease_in": zoomEaseIn,
"zoom_ease_out": zoomEaseOut,
] ]
} }
} }
+967
View File
@@ -0,0 +1,967 @@
import SwiftUI
import AppKit
/// Guia passo a passo do fluxo completo: projeto → transcrição → análise de
/// voz → copiar para o chat e trazer as decisões → revisar as ênfases →
/// processamento final. Existe para que o usuário não precise entender a ordem
/// certa de botões espalhados em várias abas — cada etapa só libera a próxima
/// quando o passo anterior terminou, e a "ponte" com o chat (que hoje exigia
/// sair do app e escolher um arquivo na mão) vira copiar/colar assistido
/// dentro da própria tela.
enum WizardStep: Int, CaseIterable, Identifiable {
case projeto, transcricao, analise, exportarChat, revisar, finalizar
var id: Int { rawValue }
var titulo: String {
switch self {
case .projeto: return "Projeto"
case .transcricao: return "Transcrever"
case .analise: return "Analisar voz"
case .exportarChat: return "Decisões da IA"
case .revisar: return "Revisar ênfases"
case .finalizar: return "Processar"
}
}
}
struct WizardView: View {
@State private var step: WizardStep = .projeto
// Passo 1 — projeto
@State private var outputFolder: String?
@State private var projectPath: String?
@State private var catalog: Catalog?
// Passo 2 — transcrição
@State private var isTranscribing = false
@State private var transcribeProgress: Double = 0
@State private var transcribeStage = ""
@State private var transcribeResults: [TranscriptResult] = []
// Passo 3 — análise de voz
@State private var isAnalyzing = false
@State private var voiceTimelinePath: String?
@State private var voiceAnalysisMessage = ""
@State private var acousticsAvailable: Bool?
@State private var showVoiceTimelineReuseAlert = false
@State private var existingVoiceTimelinePath: String?
// Passo 4 — enviar ao chat e trazer as decisões de volta
@State private var copiedFeedback = ""
@State private var decisionsText = ""
@State private var isApplyingDecisions = false
@State private var appliedPath: String?
@State private var skippedVoiceEdit = false
// Passo 4 (alternativa) — gerar o roteiro direto por IA local (Ollama/Gemma 3)
@State private var isGeneratingScript = false
@State private var generateScriptModel = "gemma3:12b"
@State private var generateScriptFeedback = ""
@State private var ollamaModels: [String] = []
// Passo 5 — revisar ênfases
@StateObject private var reviewModel = PhraseReviewModel()
@State private var reviewLoadedFor: String?
@State private var reviewLoadedForDecisions: String?
@State private var phraseReviewPath: String?
@State private var phraseActionsPath: String?
// Passo 6 — processamento final
@State private var finalSilences = true
@State private var finalFillers = false
@State private var finalSubtitles = true
@State private var finalDynamicSubtitles = false
@State private var isFinalizing = false
@State private var finalStatus = ""
@State private var finalPath: String?
@State private var errorMessage: String?
var body: some View {
VStack(spacing: 0) {
stepperHeader
.padding(.horizontal, 24)
.padding(.top, 20)
.padding(.bottom, 16)
Divider()
// A revisão é uma sala de edição, não um formulário: ela precisa da
// largura toda e rola por conta própria (timeline horizontal, lista
// vertical). As demais etapas continuam na coluna estreita, que é o
// que mantém um passo a passo legível.
if step == .revisar {
revisarStep
} else {
ScrollView {
VStack(alignment: .leading, spacing: 18) {
if let errorMessage, !errorMessage.isEmpty {
Label(errorMessage, systemImage: "exclamationmark.triangle.fill")
.foregroundStyle(.red)
.padding(.top, 4)
}
content
}
.padding(24)
.frame(maxWidth: 640, alignment: .leading)
.frame(maxWidth: .infinity)
}
}
Divider()
navFooter
.padding(.horizontal, 24)
.padding(.vertical, 16)
}
.task {
loadProjectConfig()
await loadCatalog()
}
.alert("Análise de voz já existe", isPresented: $showVoiceTimelineReuseAlert) {
Button("Usar existente") {
if let existingVoiceTimelinePath {
voiceTimelinePath = existingVoiceTimelinePath
voiceAnalysisMessage = "Reaproveitando análise existente: \(existingVoiceTimelinePath)"
}
}
Button("Reprocessar") {
analyzeVoice(forceReprocess: true)
}
Button("Cancelar", role: .cancel) {}
} message: {
Text("Já existe um arquivo voice_timeline para este projeto. Quer manter o processamento anterior para ganhar tempo?")
}
}
// MARK: - Cabeçalho com os passos
private var stepperHeader: some View {
HStack(spacing: 6) {
ForEach(WizardStep.allCases) { s in
HStack(spacing: 6) {
ZStack {
Circle()
.fill(colorFor(s))
.frame(width: 24, height: 24)
if s.rawValue < step.rawValue {
Image(systemName: "checkmark")
.font(.caption2.weight(.bold))
.foregroundStyle(.white)
} else {
Text("\(s.rawValue + 1)")
.font(.caption2.weight(.bold))
.foregroundStyle(s == step ? .white : .secondary)
}
}
Text(s.titulo)
.font(.caption)
.foregroundStyle(s == step ? .primary : .secondary)
.fontWeight(s == step ? .semibold : .regular)
}
if s != WizardStep.allCases.last {
Rectangle()
.fill(s.rawValue < step.rawValue ? Color.accentColor : Color.secondary.opacity(0.25))
.frame(height: 2)
.frame(maxWidth: .infinity)
}
}
}
}
private func colorFor(_ s: WizardStep) -> Color {
if s.rawValue < step.rawValue { return .accentColor }
if s == step { return .accentColor }
return Color.secondary.opacity(0.25)
}
// MARK: - Conteúdo por etapa
@ViewBuilder
private var content: some View {
switch step {
case .projeto: projetoStep
case .transcricao: transcricaoStep
case .analise: analiseStep
case .exportarChat: exportarChatStep
case .revisar: revisarStep
case .finalizar: finalizarStep
}
}
private var projetoStep: some View {
VStack(alignment: .leading, spacing: 16) {
Text("1. Escolha o projeto").font(.title3.weight(.semibold))
Text("A pasta é onde tudo o que for gerado nesse fluxo fica salvo. O arquivo é o .fcpxml exportado do Final Cut Pro.")
.font(.callout).foregroundStyle(.secondary)
fieldRow(icon: "folder", label: outputFolder ?? "Nenhuma pasta selecionada", isSet: outputFolder != nil) {
pickOutputFolder()
}
fieldRow(icon: "doc.text", label: projectPath.map { URL(fileURLWithPath: $0).lastPathComponent } ?? "Nenhum arquivo selecionado", isSet: projectPath != nil) {
pickProjectFile()
}
if looksLikeGeneratedFile(projectPath) {
Label("Esse arquivo parece já ter sido processado por este fluxo (o nome tem um sufixo como \"_voice_edit\" ou \"_silence_removed\"). Rodar o wizard de novo em cima dele reaplica os cortes por cima de cortes já feitos. Selecione o .fcpxml original do Final Cut, a menos que a intenção seja mesmo reprocessar.",
systemImage: "exclamationmark.triangle.fill")
.font(.caption).foregroundStyle(.orange)
}
if (catalog?.installedCount ?? 0) == 0 {
Label("Nenhum modelo de transcrição instalado. Baixe um na aba \"Modelos\" antes de continuar.",
systemImage: "exclamationmark.triangle.fill")
.font(.caption).foregroundStyle(.orange)
}
}
}
private var transcricaoStep: some View {
VStack(alignment: .leading, spacing: 16) {
Text("2. Transcreva o áudio").font(.title3.weight(.semibold))
Text("Roda localmente com o modelo escolhido na aba Modelos. Vira a base de tudo que vem depois — o corte por voz, as legendas, os marcadores.")
.font(.callout).foregroundStyle(.secondary)
Button {
startTranscription()
} label: {
if isTranscribing {
HStack { ProgressView().controlSize(.small); Text(transcribeStage.isEmpty ? "Transcrevendo…" : transcribeStage) }
.frame(maxWidth: .infinity)
} else {
Label(transcribeResults.isEmpty ? "Transcrever" : "Transcrever novamente", systemImage: "waveform")
.frame(maxWidth: .infinity)
}
}
.buttonStyle(.borderedProminent)
.controlSize(.large)
.disabled(isTranscribing || projectPath == nil || outputFolder == nil)
if isTranscribing {
VStack(alignment: .leading, spacing: 6) {
ProgressView(value: transcribeProgress)
Text("\(Int(transcribeProgress * 100))%").font(.caption).foregroundStyle(.secondary).monospacedDigit()
}
}
if !transcribeResults.isEmpty {
ForEach(transcribeResults, id: \.media) { r in
VStack(alignment: .leading, spacing: 4) {
HStack {
Image(systemName: "checkmark.circle.fill").foregroundStyle(.green)
Text(r.media).font(.body.weight(.medium))
Spacer()
Text("\(r.language) · \(r.words) palavras").font(.caption).foregroundStyle(.secondary)
}
Text(r.preview).font(.caption).foregroundStyle(.secondary).lineLimit(2)
}
.padding(12)
.background(RoundedRectangle(cornerRadius: 8).fill(Color.secondary.opacity(0.06)))
}
}
}
}
private var analiseStep: some View {
VStack(alignment: .leading, spacing: 16) {
Text("3. Analise a voz").font(.title3.weight(.semibold))
Text("Gera o JSON com transcrição, locutor e intensidade (pitch/energia/ritmo) por palavra — é esse arquivo que o chat lê para decidir o que cortar. Não corta nada sozinho.")
.font(.callout).foregroundStyle(.secondary)
Button {
analyzeVoice()
} label: {
if isAnalyzing {
HStack { ProgressView().controlSize(.small); Text("Analisando…") }.frame(maxWidth: .infinity)
} else {
Label(voiceTimelinePath == nil ? "Analisar voz" : "Analisar novamente", systemImage: "waveform.badge.magnifyingglass")
.frame(maxWidth: .infinity)
}
}
.buttonStyle(.borderedProminent)
.controlSize(.large)
.disabled(isAnalyzing || projectPath == nil || outputFolder == nil)
if let voiceTimelinePath {
VStack(alignment: .leading, spacing: 6) {
Label("Análise pronta", systemImage: "checkmark.circle.fill").foregroundStyle(.green)
Text(voiceTimelinePath).font(.caption).foregroundStyle(.secondary).lineLimit(1).truncationMode(.middle)
}
.padding(12)
.background(RoundedRectangle(cornerRadius: 8).fill(Color.secondary.opacity(0.06)))
if acousticsAvailable == false {
VStack(alignment: .leading, spacing: 4) {
Label("Sem análise acústica real", systemImage: "exclamationmark.triangle.fill")
.font(.caption.weight(.semibold)).foregroundStyle(.orange)
Text("Falta o componente \"librosa\" — os cortes ainda são decididos pelo texto, mas o chat não vai propor zoom com confiança. Instale em Avançado → Modelos → \"Análise Acústica\", e refaça esta etapa depois.")
.font(.caption).foregroundStyle(.secondary)
}
.padding(12)
.background(RoundedRectangle(cornerRadius: 8).fill(Color.orange.opacity(0.08)))
}
}
}
}
private var exportarChatStep: some View {
VStack(alignment: .leading, spacing: 16) {
Text("4. Envie para o chat decidir os cortes").font(.title3.weight(.semibold))
Text("Esta é a única etapa manual que sobra: o julgamento de qual tomada usar, onde dar zoom e o que escrever na tela é feito pela IA numa conversa, não por um botão. Copie abaixo, cole numa sessão do Claude e peça pra rodar a skill \"editar-por-voz\".")
.font(.callout).foregroundStyle(.secondary)
if let voiceTimelinePath {
// Alternativa automática: em vez de copiar/colar no chat, manda a
// própria voice timeline (o arquivo inteiro) junto com o brief para
// o modelo local (Ollama/Gemma 3) decidir a edição de uma vez —
// cortes, zooms e textos numa única chamada, sem sair do app.
VStack(alignment: .leading, spacing: 8) {
Text("OU gere o roteiro por IA local (Ollama/Gemma 3)").font(.callout.weight(.semibold))
Text("O app envia a voice timeline completa (o arquivo) acompanhada do pedido para o modelo local decidir os cortes, zooms e textos de uma vez. Nada de copiar e colar.")
.font(.caption).foregroundStyle(.secondary)
HStack {
if ollamaModels.isEmpty {
TextField("Modelo (ex.: gemma3:12b, llama3)", text: $generateScriptModel)
.textFieldStyle(.roundedBorder)
.frame(maxWidth: 260)
} else {
Picker("Modelo", selection: $generateScriptModel) {
ForEach(ollamaModels, id: \.self) { m in
Text(m).tag(m)
}
}
.pickerStyle(.menu)
.frame(maxWidth: 260)
TextField("Ou outro", text: $generateScriptModel)
.textFieldStyle(.roundedBorder)
.frame(maxWidth: 120)
}
Button {
generateScript(voiceTimelinePath: voiceTimelinePath)
} label: {
if isGeneratingScript {
HStack { ProgressView().controlSize(.small); Text("Gerando…") }
} else {
Label("Gerar roteiro por IA local", systemImage: "sparkles")
}
}
.buttonStyle(.borderedProminent)
.disabled(isGeneratingScript || voiceTimelinePath.isEmpty)
}
if !generateScriptFeedback.isEmpty {
Label(generateScriptFeedback, systemImage: "checkmark.circle.fill")
.font(.caption).foregroundStyle(.green)
}
}
.padding(12)
.background(RoundedRectangle(cornerRadius: 8).fill(Color.green.opacity(0.07)))
.onAppear { fetchOllamaModels() }
Divider().padding(.vertical, 4)
Button {
copyForChat(path: voiceTimelinePath)
} label: {
Label("Copiar para colar no chat", systemImage: "doc.on.clipboard")
.frame(maxWidth: .infinity)
}
.buttonStyle(.borderedProminent)
.controlSize(.large)
if !copiedFeedback.isEmpty {
Label(copiedFeedback, systemImage: "checkmark.circle.fill")
.font(.caption).foregroundStyle(.green)
}
VStack(alignment: .leading, spacing: 8) {
Text("O que é copiado").font(.caption.weight(.semibold)).foregroundStyle(.secondary)
Text("Um pedido pronto + o conteúdo de \(URL(fileURLWithPath: voiceTimelinePath).lastPathComponent), já formatado. É só colar (⌘V) numa conversa com o Claude.")
.font(.caption).foregroundStyle(.secondary)
}
.padding(12)
.background(RoundedRectangle(cornerRadius: 8).fill(Color.secondary.opacity(0.06)))
Divider().padding(.vertical, 4)
Text("Cole aqui o que o chat devolveu").font(.callout.weight(.semibold))
Text("Na próxima etapa essas decisões aparecem já marcadas na timeline, frase por frase, para você lapidar.")
.font(.caption).foregroundStyle(.secondary)
HStack {
Button {
if let s = NSPasteboard.general.string(forType: .string) {
decisionsText = s
}
} label: {
Label("Colar da área de transferência", systemImage: "list.clipboard")
}
Spacer()
if !decisionsText.isEmpty {
Label(jsonIsValid ? "JSON válido" : "JSON inválido",
systemImage: jsonIsValid ? "checkmark.circle.fill" : "xmark.circle.fill")
.font(.caption)
.foregroundStyle(jsonIsValid ? .green : .red)
}
}
TextEditor(text: $decisionsText)
.font(.system(.caption, design: .monospaced))
.frame(minHeight: 140)
.padding(8)
.background(RoundedRectangle(cornerRadius: 8).fill(Color.secondary.opacity(0.06)))
.overlay(RoundedRectangle(cornerRadius: 8).stroke(Color.secondary.opacity(0.2)))
Button {
applyDecisions()
} label: {
if isApplyingDecisions {
HStack { ProgressView().controlSize(.small); Text("Aplicando…") }
.frame(maxWidth: .infinity)
} else {
Label("Aplicar decisões", systemImage: "checkmark.seal")
.frame(maxWidth: .infinity)
}
}
.buttonStyle(.borderedProminent)
.controlSize(.large)
.disabled(isApplyingDecisions || !jsonIsValid)
if let appliedPath {
Label("Decisões aplicadas — \(URL(fileURLWithPath: appliedPath).lastPathComponent)",
systemImage: "checkmark.circle.fill")
.font(.caption).foregroundStyle(.green)
}
Divider()
Button("Pular esta etapa (revisar as ênfases direto, sem passar pela IA)") {
skippedVoiceEdit = true
appliedPath = nil
decisionsText = ""
}
.buttonStyle(.plain)
.font(.caption)
.foregroundStyle(.secondary)
} else {
Label("Volte ao passo anterior e rode a análise de voz primeiro.", systemImage: "exclamationmark.triangle.fill")
.font(.caption).foregroundStyle(.orange)
}
}
}
/// Etapa 5 — a sala de edição. Diferente das outras, não é um formulário
/// dentro da coluna do assistente: ocupa a janela toda e se carrega sozinha
/// na primeira vez que aparece para aquela análise de voz.
private var revisarStep: some View {
Group {
if voiceTimelinePath != nil {
PhraseReviewView(model: reviewModel)
} else {
VStack(spacing: 8) {
Label("Volte ao passo 3 e rode a análise de voz primeiro.",
systemImage: "exclamationmark.triangle.fill")
.foregroundStyle(.orange)
}
.frame(maxWidth: .infinity, maxHeight: .infinity)
}
}
.onAppear { loadReviewIfNeeded() }
}
/// Processing and its result live on the SAME slide: the moment the last
/// operation finishes (`finalPath` gets set), the open/reveal buttons
/// appear right below the "Processar" button instead of gating behind a
/// separate "Concluído" step the user has to click into — there was
/// nothing on that slide worth a click of its own.
private var finalizarStep: some View {
VStack(alignment: .leading, spacing: 16) {
Text("6. Finalize o corte").font(.title3.weight(.semibold))
Text("Últimos passos automáticos, sem decisão envolvida — rodam com os parâmetros já configurados na aba \"Análise de Voz\" / \"Legendas Dinâmicas\".")
.font(.callout).foregroundStyle(.secondary)
Toggle("Remover silêncios do áudio", isOn: $finalSilences)
Toggle("Remover palavras de preenchimento", isOn: $finalFillers)
Toggle("Gerar legenda comum (texto editável no FCP)", isOn: $finalSubtitles)
Toggle("Gerar legendas dinâmicas (estilo configurado na aba própria)", isOn: $finalDynamicSubtitles)
Button {
finalizeProcessing()
} label: {
if isFinalizing {
HStack { ProgressView().controlSize(.small); Text(finalStatus.isEmpty ? "Processando…" : finalStatus) }
.frame(maxWidth: .infinity)
} else {
Label("Processar", systemImage: "play.fill").frame(maxWidth: .infinity)
}
}
.buttonStyle(.borderedProminent)
.controlSize(.large)
.disabled(isFinalizing || (!finalSilences && !finalFillers && !finalSubtitles && !finalDynamicSubtitles))
if let finalPath, !isFinalizing {
Divider().padding(.vertical, 4)
Label("Concluído", systemImage: "checkmark.seal.fill")
.font(.callout.weight(.semibold))
.foregroundStyle(.green)
Text(finalPath).font(.caption).foregroundStyle(.secondary).lineLimit(1).truncationMode(.middle)
HStack {
Button("Abrir no Final Cut Pro") { NSWorkspace.shared.open(URL(fileURLWithPath: finalPath)) }
.buttonStyle(.borderedProminent)
Button("Mostrar no Finder") {
NSWorkspace.shared.activateFileViewerSelecting([URL(fileURLWithPath: finalPath)])
}
Spacer()
Button("Começar outro projeto") { resetWizard() }
}
} else if !finalStatus.isEmpty && !isFinalizing {
Text(finalStatus).font(.caption).foregroundStyle(.secondary)
}
}
}
// MARK: - Navegação
/// `.finalizar` is the last step now — once it has a `finalPath`, the
/// slide's own "Começar outro projeto" button is the way forward, so the
/// footer's "Continuar" would be a second, redundant path to nowhere.
private var navFooter: some View {
HStack {
if step != .projeto {
Button("Voltar") { goBack() }
}
Spacer()
if step != .finalizar || finalPath == nil {
Button(step == .finalizar ? "Concluir" : "Continuar") { goNext() }
.buttonStyle(.borderedProminent)
.disabled(!canAdvance)
}
}
}
private var canAdvance: Bool {
switch step {
case .projeto: return outputFolder != nil && projectPath != nil
case .transcricao: return !transcribeResults.isEmpty
case .analise: return voiceTimelinePath != nil
case .exportarChat: return appliedPath != nil || skippedVoiceEdit
// Revisar é opcional: a sugestão da IA já é utilizável como veio, então
// o botão nunca trava aqui — o passo existe para lapidar, não para
// exigir mais uma confirmação.
case .revisar: return true
case .finalizar: return finalPath != nil && !isFinalizing
}
}
private func goNext() {
guard let next = WizardStep(rawValue: step.rawValue + 1) else { return }
// Sair da revisão grava o que foi decidido e as ações derivadas dela
// (`_phrase_actions.json`) ao lado da análise de voz — é esse arquivo
// que `finalizeProcessing` reaplica na etapa 6, para que desativar uma
// frase aqui realmente a remova do vídeo final, e não só do registro.
if step == .revisar {
reviewModel.save { reviewPath, actionsPath in
phraseReviewPath = reviewPath
phraseActionsPath = actionsPath
}
}
step = next
}
private func goBack() {
guard let prev = WizardStep(rawValue: step.rawValue - 1) else { return }
step = prev
}
private func resetWizard() {
step = .projeto
transcribeResults = []
voiceTimelinePath = nil
voiceAnalysisMessage = ""
decisionsText = ""
appliedPath = nil
skippedVoiceEdit = false
reviewLoadedFor = nil
reviewLoadedForDecisions = nil
phraseReviewPath = nil
phraseActionsPath = nil
finalStatus = ""
finalPath = nil
errorMessage = nil
}
// MARK: - Componentes auxiliares
@ViewBuilder
private func fieldRow(icon: String, label: String, isSet: Bool, action: @escaping () -> Void) -> some View {
HStack {
Image(systemName: icon).foregroundStyle(isSet ? .primary : .secondary)
Text(label).lineLimit(1).truncationMode(.middle).foregroundStyle(isSet ? .primary : .secondary)
Spacer()
Button("Escolher…", action: action)
}
.padding(12)
.background(RoundedRectangle(cornerRadius: 8).fill(Color.secondary.opacity(0.06)))
}
/// Todo output do fluxo carrega um destes sufixos no nome (ver
/// `_derived_output` / suffixes usados por `apply_voice_actions`,
/// `remove_silences`, `generate_dynamic_subtitles` em
/// `admin/models_api.py`). Selecionar um deles como "o projeto" no passo
/// 1 é o erro que gerou arquivos como `_voice_edit_voice_edit_...`: os
/// cortes de voz assumem timestamps da mídia ORIGINAL, então reaplicá-los
/// sobre um arquivo já cortado desloca tudo silenciosamente.
private static let generatedSuffixes = [
"_voice_edit", "_silence_removed", "_dynamic_subtitles",
"_transcript_edit", "_fillers_removed", "_markers",
]
private func looksLikeGeneratedFile(_ path: String?) -> Bool {
guard let path else { return false }
let stem = URL(fileURLWithPath: path).deletingPathExtension().lastPathComponent
return Self.generatedSuffixes.contains { stem.contains($0) }
}
private var jsonIsValid: Bool {
guard let data = decisionsText.data(using: .utf8), !decisionsText.trimmingCharacters(in: .whitespacesAndNewlines).isEmpty else { return false }
return (try? JSONSerialization.jsonObject(with: data)) != nil
}
// MARK: - Ações — Python bridge
private func loadProjectConfig() {
PythonBridge.call(command: "project_config") { result, _ in
DispatchQueue.main.async {
guard let result, result["ok"] as? Bool == true else { return }
if let folder = result["folder"] as? String, !folder.isEmpty { outputFolder = folder }
if let file = result["file"] as? String, !file.isEmpty { projectPath = file }
}
}
}
private func loadCatalog() async {
PythonBridge.call(command: "catalog") { result, _ in
DispatchQueue.main.async {
if let result { catalog = Catalog(json: result) }
}
}
}
private func pickOutputFolder() {
let panel = NSOpenPanel()
panel.canChooseFiles = false
panel.canChooseDirectories = true
panel.allowsMultipleSelection = false
panel.prompt = "Usar esta pasta"
panel.message = "Escolha a pasta onde os resultados serão salvos."
if panel.runModal() == .OK, let url = panel.url {
outputFolder = url.path
PythonBridge.call(command: "set_project_config", arguments: ["folder": url.path]) { _, _ in }
}
}
private func pickProjectFile() {
let panel = NSOpenPanel()
panel.canChooseFiles = true
panel.canChooseDirectories = false
panel.allowsMultipleSelection = false
panel.prompt = "Selecionar"
panel.message = "Selecione o arquivo (.fcpxml) ou o bundle (.fcpxmld) exportado pelo Final Cut Pro."
if panel.runModal() == .OK, let url = panel.url {
let ext = url.pathExtension.lowercased()
if ext == "fcpxml" || ext == "fcpxmld" || ext == "xml" {
projectPath = url.path
PythonBridge.call(command: "set_project_config", arguments: ["file": url.path]) { _, _ in }
} else {
errorMessage = "Selecione um arquivo .fcpxml, .fcpxmld ou .xml do Final Cut Pro."
}
}
}
private func startTranscription() {
guard let projectPath, let outputFolder else { return }
isTranscribing = true
errorMessage = nil
transcribeResults = []
transcribeProgress = 0
PythonBridge.run(command: "transcribe", arguments: ["path": projectPath, "output_dir": outputFolder]) { obj in
DispatchQueue.main.async {
let type = obj["type"] as? String
if type == "progress" {
transcribeProgress = (obj["fraction"] as? NSNumber)?.doubleValue ?? 0
transcribeStage = obj["stage"] as? String ?? ""
} else if type == "error" {
errorMessage = obj["message"] as? String ?? "Erro na transcrição."
} else if type == "result", let arr = obj["transcripts"] as? [[String: Any]] {
transcribeResults = arr.map(TranscriptResult.init)
}
}
} completion: { code, err in
DispatchQueue.main.async {
isTranscribing = false
transcribeProgress = 1
if code != 0 && transcribeResults.isEmpty {
errorMessage = err ?? "A transcrição falhou."
}
}
}
}
private func analyzeVoice(forceReprocess: Bool = false) {
guard let projectPath, let outputFolder else { return }
isAnalyzing = true
errorMessage = nil
PythonBridge.call(command: "analyze_voice", arguments: [
"path": projectPath,
"output_dir": outputFolder,
"force_reprocess": forceReprocess,
]) { result, err in
DispatchQueue.main.async {
isAnalyzing = false
guard result?["ok"] as? Bool == true else {
errorMessage = result?["error"] as? String ?? err ?? "Falha ao analisar a voz."
return
}
if result?["reused"] as? Bool == true, !forceReprocess {
let timelines = result?["timelines"] as? [String] ?? []
existingVoiceTimelinePath = timelines.first ?? extractPath(from: result?["message"] as? String ?? "", marker: "**Timeline JSON**:")
showVoiceTimelineReuseAlert = true
return
}
let message = result?["message"] as? String ?? ""
voiceAnalysisMessage = message
if let path = extractPath(from: message, marker: "**Timeline JSON**:") {
voiceTimelinePath = path
} else {
voiceTimelinePath = nil
// ok:true não garante que a análise gerou timeline — se
// não houver fala detectável no áudio, o Python volta com
// sucesso mas sem "Timeline JSON" na mensagem. Sem isso
// aqui, a etapa parecia não fazer nada.
errorMessage = "A análise terminou mas não encontrou fala reconhecível no áudio. Mensagem do motor: " + (message.isEmpty ? "(vazia)" : message)
}
checkAcoustics()
}
}
}
/// A ênfase de voz (energia/tom) depende do `librosa`, dependência
/// opcional. Sem ela, a análise ainda transcreve e corta pelo texto,
/// mas nunca deveria propor zoom — por isso avisamos aqui, no ponto
/// onde o usuário sentiria falta, em vez de só na aba Modelos.
private func checkAcoustics() {
PythonBridge.call(command: "acoustics_capability") { result, _ in
DispatchQueue.main.async {
guard let result, result["ok"] as? Bool == true else { return }
acousticsAvailable = result["available"] as? Bool
}
}
}
/// Localiza uma linha markdown do tipo "- **Marker**: valor" (usado nas
/// mensagens do bridge Python) e devolve o valor. Aceita o marcador de
/// lista "- " opcional antes dos asteriscos.
private func extractPath(from message: String, marker: String) -> String? {
for line in message.split(separator: "\n") {
var trimmed = Substring(line.trimmingCharacters(in: .whitespaces))
if trimmed.hasPrefix("- ") { trimmed = trimmed.dropFirst(2) }
if trimmed.hasPrefix(marker) {
return trimmed.dropFirst(marker.count).trimmingCharacters(in: .whitespaces)
}
}
return nil
}
private func copyForChat(path: String) {
guard let content = try? String(contentsOfFile: path, encoding: .utf8) else {
errorMessage = "Não foi possível ler \(path)."
return
}
let prompt = """
Use a skill "editar-por-voz" para decidir os cortes deste projeto a partir da timeline de voz abaixo. Devolva só o JSON de decisões (cortes, zooms, textos, marcadores) pronto para eu colar de volta no app.
```json
\(content)
```
"""
let pasteboard = NSPasteboard.general
pasteboard.clearContents()
pasteboard.setString(prompt, forType: .string)
copiedFeedback = "Copiado — cole (⌘V) numa conversa com o Claude."
}
/// Monta a revisão uma vez por análise de voz. Voltar e avançar de novo com
/// as MESMAS decisões não recarrega: isso jogaria fora as edições manuais
/// em silêncio, que é exatamente o que esta tela existe para preservar.
///
/// Mas se o usuário voltou à etapa 4 e colou/gerou um JSON de decisões
/// DIFERENTE do que gerou a revisão atual, isso é recarregado — e com
/// `fresh: true`, para que o `active`/ênfase recém-derivado dessas
/// decisões novas não seja imediatamente sobrescrito pela revisão salva
/// da visita anterior (`merge_saved_decisions`, do lado Python). Sem isso,
/// a tela ficava presa nas decisões antigas mesmo depois de reaplicar o
/// corte — a dessincronia relatada entre "ativa aqui" e "já cortado no
/// FCPXML".
private func loadReviewIfNeeded() {
guard let voiceTimelinePath else { return }
let decisionsChanged = reviewLoadedForDecisions != nil && reviewLoadedForDecisions != decisionsText
guard reviewLoadedFor != voiceTimelinePath || decisionsChanged else { return }
reviewLoadedFor = voiceTimelinePath
reviewLoadedForDecisions = decisionsText
// A pasta do projeto e a do .fcpxml entram como onde procurar a mídia:
// a análise de voz guarda só o nome do arquivo, não o caminho.
reviewModel.load(
voiceTimelinePath: voiceTimelinePath,
decisionsJSON: decisionsText,
outputFolder: outputFolder,
mediaFolder: projectPath.map { URL(fileURLWithPath: $0).deletingLastPathComponent().path },
fresh: decisionsChanged
)
}
private func applyDecisions() {
guard let projectPath, let outputFolder,
let data = decisionsText.data(using: .utf8),
let parsed = try? JSONSerialization.jsonObject(with: data) else { return }
isApplyingDecisions = true
errorMessage = nil
PythonBridge.call(command: "apply_voice_actions", arguments: [
"path": projectPath,
"output_dir": outputFolder,
"actions": parsed,
]) { result, err in
DispatchQueue.main.async {
isApplyingDecisions = false
guard result?["ok"] as? Bool == true else {
errorMessage = result?["error"] as? String ?? err ?? "Falha ao aplicar as decisões."
return
}
appliedPath = result?["path"] as? String ?? projectPath
skippedVoiceEdit = false
}
}
}
/// Etapa 4 (alternativa): manda a voice timeline inteira para um modelo
/// local (Ollama/Gemma 3) que dirige a edição de uma vez — sem copiar e
/// colar. O motor devolve o roteiro legível + o JSON de ações e já aplica
/// no FCPXML (non-destructive), igual ao fluxo manual "Aplicar decisões".
private func fetchOllamaModels() {
guard ollamaModels.isEmpty else { return }
PythonBridge.call(command: "list_ollama_models", arguments: [:]) { result, err in
DispatchQueue.main.async {
if let models = result?["models"] as? [String], !models.isEmpty {
ollamaModels = models
if !models.contains(generateScriptModel) {
generateScriptModel = models.first ?? generateScriptModel
}
}
}
}
}
private func generateScript(voiceTimelinePath: String) {
guard let projectPath, let outputFolder else { return }
isGeneratingScript = true
generateScriptFeedback = ""
errorMessage = nil
PythonBridge.call(command: "generate_voice_script", arguments: [
"voice_timeline": voiceTimelinePath,
"filepath": projectPath,
"output_dir": outputFolder,
"model": generateScriptModel,
"apply_to_fcpxml": true,
]) { result, err in
DispatchQueue.main.async {
isGeneratingScript = false
guard result?["ok"] as? Bool == true else {
errorMessage = result?["error"] as? String ?? err ?? "Falha ao gerar roteiro por IA local."
return
}
// Traz as decisões de volta para a tela de revisão (etapa 5) e
// marca como aplicadas, exatamente como o "Aplicar decisões".
if let actionsPath = result?["actions_path"] as? String,
let content = try? String(contentsOfFile: actionsPath, encoding: .utf8) {
decisionsText = content
}
appliedPath = result?["applied_path"] as? String ?? projectPath
skippedVoiceEdit = false
generateScriptFeedback = "Roteiro gerado e aplicado — revise as ênfases na próxima etapa."
}
}
}
private func finalizeProcessing() {
guard let outputFolder else { return }
let startPath = appliedPath ?? projectPath
guard let startPath else { return }
var operations: [String] = []
if finalSilences { operations.append("remove_silences") }
if finalFillers { operations.append("remove_filler_words") }
if finalSubtitles { operations.append("generate_plain_subtitles") }
if finalDynamicSubtitles { operations.append("generate_dynamic_subtitles") }
guard !operations.isEmpty else { return }
isFinalizing = true
errorMessage = nil
finalStatus = "Iniciando…"
applyReviewDecisions(startPath: startPath, outputFolder: outputFolder) { reviewedPath in
finalizeStep(operations, index: 0, currentPath: reviewedPath, outputFolder: outputFolder)
}
}
/// Reapplies whatever the etapa-5 review decided (active/inactive
/// phrases, manual zooms) on top of `startPath` before the finishing
/// chain runs below. Without this, `appliedPath` stayed frozen at
/// whatever `exportarChat`'s `apply_voice_actions` produced BEFORE the
/// human review — so toggling a phrase off in the review only updated
/// `_phrase_actions.json` on disk, never the video the wizard actually
/// exports. A no-op (just hands `startPath` straight through) when the
/// review step was never visited/saved this session.
private func applyReviewDecisions(
startPath: String, outputFolder: String, completion: @escaping (String) -> Void
) {
guard let phraseActionsPath,
let data = try? Data(contentsOf: URL(fileURLWithPath: phraseActionsPath)),
let parsed = try? JSONSerialization.jsonObject(with: data) as? [String: Any],
let actions = parsed["actions"] else {
completion(startPath)
return
}
finalStatus = "Aplicando a revisão…"
PythonBridge.call(command: "apply_voice_actions", arguments: [
"path": startPath,
"output_dir": outputFolder,
"actions": actions,
]) { result, err in
DispatchQueue.main.async {
guard result?["ok"] as? Bool == true else {
isFinalizing = false
errorMessage = result?["error"] as? String ?? err ?? "Falha ao aplicar a revisão."
finalStatus = "Processamento interrompido."
return
}
completion(result?["path"] as? String ?? startPath)
}
}
}
private func finalizeStep(_ operations: [String], index: Int, currentPath: String, outputFolder: String) {
guard index < operations.count else {
isFinalizing = false
finalStatus = "Processamento concluído."
finalPath = currentPath
return
}
let operation = operations[index]
finalStatus = "Processando: \(operation)…"
PythonBridge.call(command: operation, arguments: ["path": currentPath, "output_dir": outputFolder]) { result, err in
DispatchQueue.main.async {
guard result?["ok"] as? Bool == true else {
isFinalizing = false
errorMessage = result?["error"] as? String ?? err ?? "Falha em \(operation)."
finalStatus = "Processamento interrompido."
return
}
let nextPath = result?["path"] as? String ?? currentPath
finalizeStep(operations, index: index + 1, currentPath: nextPath, outputFolder: outputFolder)
}
}
}
}
+131
View File
@@ -0,0 +1,131 @@
#!/usr/bin/env python3
"""AI Voice Editor - Pipeline completo: transcrição + análise acústica → JSON para IA.
Uso:
python ai_edit.py <media_path> [--model base] [--lang pt] [--no-diarize] [--output dir]
Gera dois arquivos na pasta output (ou ao lado do mídia):
<nome>_transcript.json — transcrição com timestamps por palavra
<nome>_voice_timeline.json — timeline de voz com ênfase, pitch, energy, speakers
Esses arquivos são a ENTRADA para a IA analisar e gerar o roteiro/edição.
"""
import argparse
import json
import sys
import time
from pathlib import Path
def main():
parser = argparse.ArgumentParser(description="Pipeline de análise de voz para IA")
parser.add_argument("media", help="Caminho do arquivo de mídia (.mp4, .mov, .wav, etc.)")
parser.add_argument("--model", default="base", help="Modelo Whisper (tiny/base/small/medium/large-v3)")
parser.add_argument("--lang", default=None, help="Idioma (ex: pt, en). Auto-detect se omitido")
parser.add_argument("--hf-token", default=None, help="HuggingFace token para diarização (opcional)")
parser.add_argument("--no-diarize", action="store_true", help="Pular diarização de falantes")
parser.add_argument("--output", default=None, help="Pasta de saída (padrão: ao lado do mídia)")
parser.add_argument("--no-align", action="store_true", help="Pular alinhamento fonético (whisperx)")
args = parser.parse_args()
media_path = Path(args.media).resolve()
if not media_path.is_file():
print(f"ERRO: Arquivo não encontrado: {media_path}", file=sys.stderr)
sys.exit(1)
# Output dir
out_dir = Path(args.output) if args.output else media_path.parent
out_dir.mkdir(parents=True, exist_ok=True)
stem = media_path.stem
# ── Fase 1: Transcrição ──────────────────────────────────────────
print(f"[1/2] Transcrevendo {media_path.name} (modelo: {args.model})...")
t0 = time.time()
# Adiciona code/ ao path para imports do projeto
code_dir = Path(__file__).resolve().parent
sys.path.insert(0, str(code_dir))
from fcpxml.transcribe import transcribe
def transcribe_progress(pct):
bar_len = 30
filled = int(bar_len * pct)
bar = "█" * filled + "░" * (bar_len - filled)
print(f"\r [{bar}] {pct*100:.0f}%", end="", flush=True)
transcript = transcribe(
str(media_path),
model_size=args.model,
language=args.lang,
progress_cb=transcribe_progress,
align=not args.no_align,
)
print() # newline after progress bar
if transcript is None:
print("ERRO: Transcrição falhou. Verifique se faster-whisper está instalado:", file=sys.stderr)
print(" uv pip install faster-whisper", file=sys.stderr)
sys.exit(1)
print(f" → {len(transcript.get('words', []))} palavras, "
f"{len(transcript.get('segments', []))} segmentos, "
f"idioma: {transcript.get('language', '?')}")
# Salva transcrição
transcript_path = out_dir / f"{stem}_transcript.json"
transcript_path.write_text(json.dumps(transcript, indent=2, ensure_ascii=False), encoding="utf-8")
print(f" → Salvo: {transcript_path}")
# ── Fase 2: Análise de voz (timeline) ────────────────────────────
print(f"\n[2/2] Analisando voz (pitch, energia, ênfase)...")
t1 = time.time()
from fcpxml.voice_timeline import build_voice_timeline
def voice_progress(fraction, stage):
print(f"\r {stage} ({fraction*100:.0f}%)", end="", flush=True)
hf_token = None if args.no_diarize else args.hf_token
timeline = build_voice_timeline(
str(media_path),
transcript,
hf_token=hf_token,
progress_cb=voice_progress,
)
print()
# Salva voice timeline
timeline_path = out_dir / f"{stem}_voice_timeline.json"
timeline_path.write_text(json.dumps(timeline, indent=2, ensure_ascii=False), encoding="utf-8")
print(f" → Salvo: {timeline_path}")
# ── Resumo ───────────────────────────────────────────────────────
elapsed = time.time() - t0
summary = timeline.get("summary", {})
layers = timeline.get("layers", {})
n_words = len(transcript.get("words", []))
n_segments = len(transcript.get("segments", []))
n_speakers = len(timeline.get("speakers", []))
duration = transcript.get("duration", 0)
print(f"\n{'='*50}")
print(f" ARQUIVOS GERADOS:")
print(f" {transcript_path}")
print(f" {timeline_path}")
print(f"\n RESUMO:")
print(f" Duração: {duration:.1f}s ({duration/60:.1f}min)")
print(f" Palavras: {n_words}")
print(f" Segmentos: {n_segments}")
print(f" Falantes: {n_speakers}")
print(f" Camadas: transcript={layers.get('transcript')}, "
f"acoustics={layers.get('acoustics')}, "
f"diarization={layers.get('diarization')}")
print(f" Tempo: {elapsed:.1f}s")
print(f"{'='*50}")
print(f"\n→ Pronto! Agora peça à IA para analisar o voice timeline e gerar o roteiro.")
if __name__ == "__main__":
main()
+181
View File
@@ -0,0 +1,181 @@
"""Forced alignment — refine word timestamps against an acoustic model.
Why this exists
--------------
faster-whisper derives word times by cross-attention, which lands every word
*start* systematically ~0.3-0.5s early (the word-end is fine). That bias flows
straight into the voice timeline and makes zoom/cut land on the wrong frame —
measured on real footage in ``Engine/docs/05_EXPERIENCIAS.md`` (#14). Phonetic
forced alignment (wav2vec2, via whisperx) re-anchors each word against the
audio and brings that error down to ~30ms.
Design
------
* The dependency (``whisperx``) is **optional** and imported lazily, exactly
like the rest of this stack (librosa, faster-whisper). When it is missing, or
any step fails, :meth:`ForcedAligner.align` returns the words unchanged, so
transcription never breaks because alignment did.
* The aligner is a single responsibility class: it knows how to turn a
transcript into the shape whisperx wants, call it, and write the refined
times back. ``transcribe.py`` owns the decision of *whether* to align.
* Align models are cached per language on the instance so repeated calls
(e.g. many short clips) don't reload the wav2vec2 weights each time.
"""
import logging
from typing import List, Optional, Sequence
logger = logging.getLogger(__name__)
class ForcedAligner:
"""Refine word-level timestamps with whisperx phonetic forced alignment.
Usage::
aligner = ForcedAligner()
words = aligner.align(words, raw_segments, media_path, language, models_dir)
``words`` and ``raw_segments`` come straight from :func:`transcribe` —
``raw_segments`` carries the per-segment ``words`` lists (the same dict
objects as in ``words``) so the aligner knows which words belong to which
audio window. Returns a list of the *same* word dicts, with ``start``/``end``
overwritten in place where alignment produced a usable time.
"""
def __init__(self, device: Optional[str] = None):
self._device = device
self._models: dict = {}
# -- capability ------------------------------------------------------
@staticmethod
def available() -> bool:
"""Whether whisperx can be imported (the aligner can run at all)."""
try:
import whisperx # noqa: F401
except Exception:
return False
return True
def _resolve_device(self) -> str:
if self._device:
return self._device
try:
import torch
if torch.cuda.is_available():
return "cuda"
except Exception:
pass
return "cpu"
# -- public API ------------------------------------------------------
def align(
self,
words: Sequence[dict],
raw_segments: Sequence[dict],
audio_path: str,
language: str,
models_dir: Optional[str] = None,
) -> List[dict]:
"""Return ``words`` with forced-aligned timestamps where possible.
Falls back to the unchanged ``words`` on any failure (missing
dependency, model load error, audio read error, or a result that
doesn't line up with the input).
"""
if not words or not language:
return list(words)
try:
import whisperx
except Exception:
logger.info("whisperx not installed; skipping forced alignment")
return list(words)
try:
device = self._resolve_device()
align_input = self._build_align_input(words, raw_segments)
audio = whisperx.load_audio(audio_path)
if language not in self._models:
align_model, metadata = whisperx.load_align_model(
language_code=language,
device=device,
model_dir=str(models_dir) if models_dir else None,
)
self._models[language] = (align_model, metadata)
align_model, metadata = self._models[language]
result = whisperx.align(
align_input,
align_model,
metadata,
audio,
device,
return_char_alignments=False,
)
return self._merge_result(words, result.get("segments", []))
except Exception:
logger.warning(
"forced alignment failed for %s; using raw timestamps", audio_path
)
return list(words)
# -- internals -------------------------------------------------------
@staticmethod
def _build_align_input(
words: Sequence[dict], raw_segments: Sequence[dict]
) -> List[dict]:
"""Transcript in whisperx's expected shape: segments -> words.
whisperx.align requires each segment to carry ``text``/``start``/``end``
and a ``words`` list whose entries have ``word``/``start``/``end``/``score``.
We only read ``words`` from ``raw_segments`` (the flattened ``words``
list is the source of truth for counts), so the two stay consistent.
"""
align_segments: List[dict] = []
for seg in raw_segments:
seg_words = [
{
"word": w.get("word", ""),
"start": float(w.get("start", 0.0)),
"end": float(w.get("end", 0.0)),
"score": float(w.get("confidence", 0.0)),
}
for w in seg.get("words", [])
]
align_segments.append(
{
"text": (seg.get("text") or "").strip(),
"start": float(seg.get("start", 0.0)),
"end": float(seg.get("end", 0.0)),
"words": seg_words,
}
)
return align_segments
@staticmethod
def _merge_result(words: Sequence[dict], aligned_segments: Sequence[dict]) -> List[dict]:
"""Walk the aligned output in order and overwrite word times in place.
whisperx preserves word order within and across segments, so a single
running index over the output words lines up with ``words``. A word the
aligner failed to place gets ``None``/``0`` times — we skip those rather
than clobber a good timestamp, and if counts ever diverge we stop and
leave the rest untouched.
"""
out = list(words)
wi = 0
for seg in aligned_segments:
for aw in seg.get("words", []):
if wi >= len(out):
return out
start = aw.get("start")
end = aw.get("end")
if start is None or end is None or end < start:
wi += 1
continue
out[wi]["start"] = float(start)
out[wi]["end"] = float(end)
wi += 1
return out
+312
View File
@@ -0,0 +1,312 @@
"""Local LLM integration — the voice timeline meets a local model.
The voice timeline is *designed* to be handed to a language model: it is the
source of truth between speech analysis and editing, layered so a model can
reason about the narrative without parsing FCPXML. This module is the client
side of that contract. It formats the timeline into the editar-por-voz brief,
calls a local model server (Ollama, running Gemma 3 / Llama locally), and
parses the model's decisions back into a validated list of VoiceActions —
all inside the engine, so there is no wizard, no copy-paste, no manual step.
Transport: Ollama's HTTP chat API at ``http://localhost:11434/api/chat``.
Any model Ollama serves works; the default is Gemma 3 because that is what
runs locally here ("Lama com Gema 3"), but pass ``model=`` to switch.
The model is untrusted input: its JSON is validated row-by-row by
:func:`fcpxml.voice_actions.parse_actions`, so one malformed decision never
discards the edit. The brief is written so the model only ever emits the four
action kinds the applier understands.
"""
from __future__ import annotations
import json
import logging
import re
from typing import Any, Dict, Optional, Sequence, Tuple
import httpx
from .voice_actions import parse_actions
logger = logging.getLogger(__name__)
DEFAULT_BASE_URL = "http://localhost:11434"
# Gemma 3 12B reliably follows the editar-por-voz brief (keep the script, cut
# only backstage chatter; the 4B variant skips the "keep the main content"
# rule and deletes the script) but doesn't fit an 8GB machine. Qwen2.5 7B
# instruct (q4_K_M) is the fallback for constrained hardware — strong at
# strict JSON-schema following, the property this brief leans on hardest.
# Pass ``model=`` to switch to whatever Ollama serves.
DEFAULT_MODEL = "qwen2.5:7b-instruct-q4_K_M"
REQUEST_TIMEOUT = 600.0
# The brief. Ported from the editar-por-voz skill criteria (criterios/01..08),
# condensed into the instructions a model needs to emit valid actions. Kept in
# Portuguese because the decisions and their reasons are read by a human editor.
_SYSTEM_PROMPT = """Você é o editor de vídeo por voz deste sistema. Recebe um JSON de "linha do tempo de voz" — a medição de COMO foi falado (ênfase, energia, pausa, falante) de uma gravação — e devolve as DECISÕES de edição em JSON, nada mais. Você nunca escreve XML.
Regras (siga rigorosamente):
1. LEIA EM CAMADAS. "summary" dá o formato da peça; "segments" é onde você trabalha (cada fala com seu texto e agregados); "segments[].words" dá o instante exato de cada destaque. Não recalcule energia, tom ou ênfase — use os números do JSON.
2. SEPARAR ROTEIRO DE BASTIDOR.
- ROTEIRO = o conteúdo principal que a pessoa quer entregar: explicação, depoimento, roteiro decorado, a mensagem. É isso que VAI FICAR.
- BASTIDOR = papo casual de gravação, cumprimentos, conversa com a equipe ("cara, beleza?", "tá gravando?", "deixa eu ver o celular"), piadas fora do assunto, tomadas interrompidas ou repetidas. É isso que VIRA "cut".
Exemplo: num vídeo sobre mastopexia, a explicação da cirurgia É o roteiro (mantém); o "tá gravando? pois é" antes dela É bastidor (corta).
Use "gap_before" e "take_boundary" (silêncio > ~3s = a câmera parou/recomeçou) para agrupar tomadas — eles marcam ONDE a tomada recomeça, não o que cortar. Nunca corte o conteúdo principal só porque tem ênfase; corte o casual/off-topic.
REGRAS DE OURO:
- MANTENHA o conteúdo principal (explicação, depoimento, roteiro decorado). Ele É o vídeo.
- CORTE SÓ o casual/off-topic: cumprimentos, "tá gravando?", papo com a equipe, olhar o celular, repetições de tomada.
- Em dúvida, MANTENHA a fala. É melhor sobrar conteúdo do que cortar o que era pra ficar.
3. ESCOLHER A MELHOR TOMADA de cada frase quando há repetições: mantenha a mais limpa e corte as outras (cut cobrindo a frase inteira).
4. CORTE (kind "cut"): para REMOVER uma frase, cubra ela inteira (start..end = início..fim da frase). Para APARAR só uma hesitação no começo ou fim, corte só da borda até a palavra (corte de meia frase é ambíguo — passe de 60% e apaga a linha toda). Nunca corte o silêncio entre falas.
5. ZOOM (kind "zoom"): só em palavra de CONTEÚDO bem enfatizada (emphasis alto, não artigo). params.scale entre 1.0 e 3.0 (padrão 1.3 se omitido). Posicione em torno da palavra, segurando até o fim da frase.
6. TEXTO (kind "text"): params.content obrigatório (≤120 chars), fixa um termo central ou callout. MARKER (kind "marker"): opcional params.content vira o nome do marcador. Use para emendas/junções que o editor deve conferir.
7. TEMPOS em segundos da MÍDIA ORIGINAL (exatamente como no JSON). Nunca compense para "depois do corte" — o programa desloca sozinho. end sempre > start, ambos ≥ 0.
8. reason OBRIGATÓRIO em cada ação, em português, embasando a decisão (ex.: 'abertura: "Aquela mama" (ênfase 0.42)'). reason vazio é decisão sem critério.
Responda APENAS com um objeto JSON válido, sem markdown, sem comentário:
{"source": "<nome do arquivo>", "actions": [{"kind": "cut|zoom|text|marker", "start": <float>, "end": <float>, "params": {}, "reason": "<pt>", "speaker": "<id>"}]}
"""
_OUTPUT_REMINDER = """Gere as decisões de edição conforme o brief. Responda SOMENTE o JSON:
{"source": "<nome do arquivo>", "actions": [{"kind": "cut|zoom|text|marker", "start": <float>, "end": <float>, "params": {}, "reason": "<pt>", "speaker": "<id>"}]}
Não inclua explicações nem blocos markdown."""
def ollama_chat(
model: str = DEFAULT_MODEL,
messages: Optional[Sequence[Dict[str, str]]] = None,
base_url: str = DEFAULT_BASE_URL,
temperature: float = 0.2,
timeout: float = REQUEST_TIMEOUT,
num_ctx: int = 32768,
) -> str:
"""One chat completion from a local Ollama server.
Returns the assistant message content. Raises on transport/HTTP errors so
the caller can decide whether to retry or report — a model call is the
one I/O in this pipeline that can legitimately fail mid-run.
"""
payload = {
"model": model,
"messages": list(messages or []),
"stream": False,
"options": {"temperature": temperature, "num_ctx": num_ctx},
}
try:
response = httpx.post(
f"{base_url.rstrip('/')}/api/chat", json=payload, timeout=timeout
)
response.raise_for_status()
data = response.json()
except Exception as exc:
# Covers transport errors AND a dropped connection that yields an empty
# body (httpx/JSONDecodeError) — both must become a RuntimeError so the
# caller reports the failure instead of crashing the whole pipeline.
raise RuntimeError(f"Falha ao falar com o modelo local em {base_url}: {exc}") from exc
return (data.get("message") or {}).get("content", "") or ""
def list_ollama_models(base_url: str = DEFAULT_BASE_URL) -> list[str]:
"""Names of the models Ollama currently serves, for a model picker.
Returns an empty list when Ollama is unreachable so the UI can fall back to
a free-text field instead of erroring.
"""
try:
resp = httpx.get(f"{base_url.rstrip('/')}/api/tags", timeout=10.0)
resp.raise_for_status()
models = resp.json().get("models", [])
names = [m.get("name") for m in models if m.get("name")]
return sorted(names)
except Exception:
return []
def _extract_json(text: str) -> Any:
"""Pull a JSON value out of a model response, tolerating fences/wrappers."""
if not text:
return None
candidate = text.strip()
# Strip a ```json ... ``` (or bare ```) fence if the model added one.
fence = re.search(r"```(?:json)?\s*(.*?)\s*```", candidate, re.DOTALL)
if fence:
candidate = fence.group(1).strip()
# Otherwise take the outermost {...} / [...].
if not candidate.startswith(("{" if True else "", "[")):
start = min(
(i for i, c in enumerate(candidate) if c in "{["),
default=None,
)
end = max(
(i for i, c in enumerate(candidate) if c in "}"),
default=None,
)
if start is not None and end is not None and end > start:
candidate = candidate[start : end + 1]
try:
data = json.loads(candidate)
except json.JSONDecodeError:
return None
# Models sometimes wrap the expected `{"source", "actions"}` object inside a
# single-element list (`[{...}]`). Unwrap that so the actions aren't treated
# as one malformed row.
if (
isinstance(data, list)
and len(data) == 1
and isinstance(data[0], dict)
and "actions" in data[0] # the wrapper carries the actions key
):
data = data[0]
return data
# Only these fields reach the model — the raw timeline also carries heavy
# per-word audio features (energy, pitch, arousal...) and speaker `samples`
# that blow past the model's context window on any real recording. Dropping
# them is what keeps a 3-minute timeline inside `num_ctx`.
_SEGMENT_KEEP = (
"start", "end", "speaker", "text", "gap_before", "take_boundary",
"avg_energy", "peak_emphasis", "emotion", "emotion_confidence",
"arousal", "valence",
)
_WORD_KEEP = ("text", "start", "end", "speaker", "emphasis", "pause_before")
_SPEAKER_KEEP = ("id", "name")
_SKIP_ROOT = ("layers", "scales")
def _project_timeline(timeline: dict) -> dict:
"""Strip the timeline down to what the edit decision actually needs."""
out = {k: v for k, v in timeline.items() if k not in _SKIP_ROOT}
speakers = [
{k: sp[k] for k in _SPEAKER_KEEP if k in sp}
for sp in timeline.get("speakers", [])
]
if speakers:
out["speakers"] = speakers
segs = []
for seg in timeline.get("segments", []):
s = {k: seg[k] for k in _SEGMENT_KEEP if k in seg}
s["words"] = [
{k: w[k] for k in _WORD_KEEP if k in w}
for w in seg.get("words", [])
]
segs.append(s)
out["segments"] = segs
return out
def _shrink_to_fit(compact: dict, max_chars: int) -> dict:
"""Drop word detail from the lowest-emphasis segments until it fits."""
segs = [dict(s) for s in compact.get("segments", [])]
while True:
payload = json.dumps(
{**compact, "segments": segs}, ensure_ascii=False, indent=1
)
if len(payload) <= max_chars or not any(s.get("words") for s in segs):
break
idx = min(
(i for i, s in enumerate(segs) if s.get("words")),
key=lambda i: float(segs[i].get("peak_emphasis", 0.0)),
)
segs[idx] = {**segs[idx], "words": []}
compact = dict(compact)
compact["segments"] = segs
return compact
def build_edit_messages(
timeline: dict, max_words_per_segment: int = 200, max_chars: int = 110000
) -> Tuple[str, str]:
"""The (system, user) pair that sends a timeline to the model.
The user turn carries a *projected* timeline (see :func:`_project_timeline`)
— text, timing, speaker and emphasis only — so a real recording fits in the
model's context window. Very long segments still have their word detail
capped to ``max_words_per_segment`` (most emphatic + boundaries), and if the
whole payload would still exceed ``max_chars`` the lowest-emphasis segments
lose their words until it fits, so we never blow ``num_ctx``.
"""
compact = _project_timeline(timeline)
if max_words_per_segment:
segs = []
for seg in compact["segments"]:
words = seg.get("words", [])
if len(words) > max_words_per_segment:
ranked = sorted(
enumerate(words),
key=lambda kv: float(kv[1].get("emphasis", 0.0)),
reverse=True,
)[: max_words_per_segment - 2]
keep = sorted({0, len(words) - 1} | {i for i, _ in ranked})
seg = {**seg, "words": [words[i] for i in keep]}
segs.append(seg)
compact["segments"] = segs
payload = json.dumps(compact, ensure_ascii=False, indent=1)
if len(payload) > max_chars:
compact = _shrink_to_fit(compact, max_chars)
payload = json.dumps(compact, ensure_ascii=False, indent=1)
user = (
"Linha do tempo de voz (JSON):\n\n"
+ payload
+ "\n\n"
+ _OUTPUT_REMINDER
)
return _SYSTEM_PROMPT, user
def generate_voice_actions(
timeline: dict,
model: str = DEFAULT_MODEL,
base_url: str = DEFAULT_BASE_URL,
temperature: float = 0.2,
timeout: float = REQUEST_TIMEOUT,
num_ctx: int = 32768,
max_words_per_segment: int = 200,
) -> Dict[str, Any]:
"""Ask the local model to direct the edit, returning validated actions.
Returns ``{"actions": [VoiceAction], "raw": str, "errors": [str]}``.
``actions`` is empty when the model returned nothing usable; ``errors``
carries the per-row rejections from :func:`parse_actions` plus any
extraction failure, so the caller can report what went wrong instead of
only the wins.
"""
system, user = build_edit_messages(timeline, max_words_per_segment)
try:
raw = ollama_chat(
model=model,
messages=[
{"role": "system", "content": system},
{"role": "user", "content": user},
],
base_url=base_url,
temperature=temperature,
timeout=timeout,
num_ctx=num_ctx,
)
except RuntimeError as exc:
return {"actions": [], "raw": "", "errors": [str(exc)]}
data = _extract_json(raw)
if data is None:
return {
"actions": [],
"raw": raw,
"errors": ["O modelo não devolveu um JSON de decisões legível."],
}
actions, errors = parse_actions(data)
return {"actions": actions, "raw": raw, "errors": errors}
+101 -2
View File
@@ -382,6 +382,10 @@ DEFAULT_VOICE_ANALYSIS_CONFIG: dict = {
"emphasis_floor": 0.25, "emphasis_floor": 0.25,
"emotion_enabled": False, "emotion_enabled": False,
"emotion_sensitivity": 0.5, "emotion_sensitivity": 0.5,
"zoom_scale": 1.30,
"zoom_mode": "in_out",
"zoom_ease_in": 0.25,
"zoom_ease_out": 0.04,
} }
@@ -402,12 +406,28 @@ def load_voice_analysis_config() -> dict:
stored = _load_config().get("voice_analysis") stored = _load_config().get("voice_analysis")
if not isinstance(stored, dict): if not isinstance(stored, dict):
return cfg return cfg
for key in ("energy_threshold", "peak_percentile", "emphasis_floor", "emotion_sensitivity"): for key in (
"energy_threshold", "peak_percentile", "emphasis_floor",
"emotion_sensitivity", "zoom_scale", "zoom_ease_in", "zoom_ease_out",
):
if key in stored: if key in stored:
try: try:
cfg[key] = max(0.0, min(1.0, float(stored[key]))) value = float(stored[key])
if key == "zoom_scale":
cfg[key] = max(1.0, min(3.0, value))
elif key.startswith("zoom_ease"):
cfg[key] = max(0.01, min(5.0, value))
else:
cfg[key] = max(0.0, min(1.0, value))
except (TypeError, ValueError): except (TypeError, ValueError):
pass pass
if "emphasis_threshold" in stored and "emphasis_floor" not in stored:
try:
cfg["emphasis_floor"] = max(0.0, min(1.0, float(stored["emphasis_threshold"])))
except (TypeError, ValueError):
pass
if stored.get("zoom_mode") in ("in_out", "in", "out"):
cfg["zoom_mode"] = stored["zoom_mode"]
if "emotion_enabled" in stored: if "emotion_enabled" in stored:
cfg["emotion_enabled"] = bool(stored["emotion_enabled"]) cfg["emotion_enabled"] = bool(stored["emotion_enabled"])
weights = stored.get("emphasis_weights") weights = stored.get("emphasis_weights")
@@ -428,6 +448,10 @@ def save_voice_analysis_config(
emphasis_floor: float | None = None, emphasis_floor: float | None = None,
emotion_enabled: bool | None = None, emotion_enabled: bool | None = None,
emotion_sensitivity: float | None = None, emotion_sensitivity: float | None = None,
zoom_scale: float | None = None,
zoom_mode: str | None = None,
zoom_ease_in: float | None = None,
zoom_ease_out: float | None = None,
) -> dict: ) -> dict:
"""Persist voice-analysis thresholds/weights. Only given fields change. """Persist voice-analysis thresholds/weights. Only given fields change.
@@ -446,6 +470,14 @@ def save_voice_analysis_config(
cfg["emotion_enabled"] = bool(emotion_enabled) cfg["emotion_enabled"] = bool(emotion_enabled)
if emotion_sensitivity is not None: if emotion_sensitivity is not None:
cfg["emotion_sensitivity"] = max(0.0, min(1.0, float(emotion_sensitivity))) cfg["emotion_sensitivity"] = max(0.0, min(1.0, float(emotion_sensitivity)))
if zoom_scale is not None:
cfg["zoom_scale"] = max(1.0, min(3.0, float(zoom_scale)))
if zoom_mode in ("in_out", "in", "out"):
cfg["zoom_mode"] = zoom_mode
if zoom_ease_in is not None:
cfg["zoom_ease_in"] = max(0.01, min(5.0, float(zoom_ease_in)))
if zoom_ease_out is not None:
cfg["zoom_ease_out"] = max(0.01, min(5.0, float(zoom_ease_out)))
if emphasis_weights is not None: if emphasis_weights is not None:
for key, value in emphasis_weights.items(): for key, value in emphasis_weights.items():
if key in cfg["emphasis_weights"] and value is not None: if key in cfg["emphasis_weights"] and value is not None:
@@ -536,6 +568,73 @@ def save_dynamic_subtitle_config(**fields) -> dict:
return cfg return cfg
DEFAULT_PLAIN_SUBTITLE_CONFIG: dict = {
"font": "Helvetica Neue",
"font_size": 82,
"font_color": "1 1 1 1",
"max_words": 7,
"position_y": -820.0,
"uppercase": False,
"keep_punctuation": True,
"text_scale": 2.0,
}
def load_plain_subtitle_config() -> dict:
"""Persisted style for simple editable FCPXML title subtitles."""
cfg = dict(DEFAULT_PLAIN_SUBTITLE_CONFIG)
stored = _load_config().get("plain_subtitles")
if not isinstance(stored, dict):
return cfg
for key in ("position_y", "text_scale"):
if key in stored:
try:
cfg[key] = float(stored[key])
except (TypeError, ValueError):
pass
for key in ("font_size", "max_words"):
if key in stored:
try:
cfg[key] = int(stored[key])
except (TypeError, ValueError):
pass
for key in ("font", "font_color"):
if key in stored and isinstance(stored[key], str) and stored[key]:
cfg[key] = stored[key]
for key in ("uppercase", "keep_punctuation"):
if key in stored:
cfg[key] = bool(stored[key])
cfg["max_words"] = max(1, int(cfg["max_words"]))
return cfg
def save_plain_subtitle_config(**fields) -> dict:
"""Persist simple subtitle style fields. Only given fields change."""
cfg = load_plain_subtitle_config()
for key, value in fields.items():
if key not in DEFAULT_PLAIN_SUBTITLE_CONFIG or value is None:
continue
if isinstance(DEFAULT_PLAIN_SUBTITLE_CONFIG[key], bool):
cfg[key] = bool(value)
elif isinstance(DEFAULT_PLAIN_SUBTITLE_CONFIG[key], float):
try:
cfg[key] = float(value)
except (TypeError, ValueError):
continue
elif isinstance(DEFAULT_PLAIN_SUBTITLE_CONFIG[key], int):
try:
cfg[key] = int(value)
except (TypeError, ValueError):
continue
else:
cfg[key] = str(value)
cfg["max_words"] = max(1, int(cfg["max_words"]))
data = _load_config()
data["plain_subtitles"] = cfg
_write_config(data)
return cfg
# Mirrors the silence thresholds the detection/removal handlers use when no # Mirrors the silence thresholds the detection/removal handlers use when no
# argument is passed (server_tools/qc.py). Persisted so the app's slider and # argument is passed (server_tools/qc.py). Persisted so the app's slider and
# any later run agree without threading three fields through every call. # any later run agree without threading three fields through every call.
File diff suppressed because it is too large Load Diff
+15
View File
@@ -194,6 +194,7 @@ class FCPXMLParser:
media_path=media_path, media_path=media_path,
audio_role=elem.get('audioRole', ''), audio_role=elem.get('audioRole', ''),
video_role=elem.get('videoRole', ''), video_role=elem.get('videoRole', ''),
rotation=self._parse_clip_rotation(elem),
) )
clip.markers.extend(self._collect_markers(elem)) clip.markers.extend(self._collect_markers(elem))
@@ -205,6 +206,19 @@ class FCPXMLParser:
return clip return clip
def _parse_clip_rotation(self, elem: ET.Element) -> float:
"""Degrees from this clip's ``<adjust-transform rotation="...">`` —
an edit-time correction (e.g. straightening a tilted phone shot),
not the camera's own recorded orientation. FCP writes the rotation
as an attribute on that element, not as a filter param."""
transform = elem.find('adjust-transform')
if transform is None:
return 0.0
try:
return float(transform.get('rotation', '0'))
except ValueError:
return 0.0
def _parse_marker_element(self, elem: ET.Element) -> Optional[Marker]: def _parse_marker_element(self, elem: ET.Element) -> Optional[Marker]:
"""Parse any marker element (<marker> or <chapter-marker>). """Parse any marker element (<marker> or <chapter-marker>).
@@ -337,6 +351,7 @@ class FCPXMLParser:
lane=lane, offset=offset, source_start=start, lane=lane, offset=offset, source_start=start,
media_path=media_path, clip_type=elem.tag, role=role, media_path=media_path, clip_type=elem.tag, role=role,
ref_id=ref, parent_clip_name=parent_name, ref_id=ref, parent_clip_name=parent_name,
rotation=self._parse_clip_rotation(elem),
) )
connected.markers.extend(self._collect_markers(elem)) connected.markers.extend(self._collect_markers(elem))
+548
View File
@@ -0,0 +1,548 @@
"""Phrase review — the human pass between the AI's decisions and the render.
A voice timeline says *how* every line was spoken; a list of voice actions says
what the model decided to do about it. Neither is reviewable on its own: the
timeline has no editorial intent, and the action list is a set of timecodes with
no text attached. This module joins them into the one view an editor can
actually judge — the script, phrase by phrase, each carrying the decision that
was made about it.
The phrase is the unit on purpose. Emphasis, in this pipeline, is not a property
of a word but of a line: an emphasized phrase gets a punch-in and a dynamic
caption, everything else gets a plain caption. Keeping the same granularity in
the review, the JSON, and the render means a toggle in the UI maps to exactly
one editorial outcome, with nothing to reconcile in between.
Trimming stays inside the phrase for the same reason. A line is rarely wrong as
a whole — it has a false start, or a trailing "né" — so each phrase carries a
``trim_start``/``trim_end`` pair that rides on word boundaries. Editing a cut
therefore means picking a word, never hunting for a frame, and a partial cut
from the model arrives as a trim instead of being rounded away.
Round-tripping is the other half of the contract. :func:`build_phrase_review`
derives the review from actions, :func:`phrase_review_to_actions` derives
actions back from the edited review, and everything the editor touched wins over
what was inferred — so re-opening the screen shows what was left there, not a
re-derivation that quietly discards the edits.
"""
import json
from pathlib import Path
from typing import Any, Dict, List, Optional, Sequence, Tuple
from .voice_actions import VoiceAction, merge_cut_ranges, parse_actions
PHRASE_REVIEW_VERSION = "1.0"
# Emphasis is stored 0-3 rather than as a float so the UI, the JSON and the
# render agree on the same discrete decision. The thresholds map the continuous
# `peak_emphasis` of the voice timeline onto those levels when the model gave no
# explicit direction for a phrase.
EMPHASIS_LEVELS = (0, 1, 2, 3)
EMPHASIS_THRESHOLDS = (0.25, 0.45, 0.65)
# Zoom scale applied per emphasis level when the review is turned back into
# actions. Level 0 never produces a zoom. The values stay inside
# voice_actions.MIN_ZOOM_SCALE..MAX_ZOOM_SCALE.
ZOOM_SCALE_BY_LEVEL = {1: 1.15, 2: 1.3, 3: 1.5}
# A phrase only survives if most of it does. Speech boundaries from a transcript
# are approximate, so a cut clipping a fraction of a second off the tail is a
# trim, not a removal — treating that as "phrase deleted" would grey out lines
# that are still fully audible.
CUT_COVERAGE_TO_DEACTIVATE = 0.6
# A punch-in shorter than this has no time to ramp in and back out — the writer
# rejects the window anyway (see the zoom ease-in/ease-out shape), so refusing
# it here turns a silent drop at render time into nothing being placed at all.
MIN_ZOOM_DURATION = 0.4
TRACK_SCRIPT = "roteiro"
TRACK_BACKSTAGE = "bastidor"
TRACKS = (TRACK_SCRIPT, TRACK_BACKSTAGE)
def resolve_source(
source: str, voice_timeline_path: str, extra_dirs: Sequence[str] = ()
) -> str:
"""The playable path for a timeline's ``source``, or "" when it's gone.
The voice timeline stores only the media's *file name* — it is written to be
read by a model, where a machine-specific absolute path is noise. That makes
it useless for opening a preview, so the file is looked up where it can
actually be: beside its own timeline JSON first (that is where
``analyze_voice`` writes it), then in whatever project folders the caller
knows about.
"""
if not source:
return ""
candidate = Path(source)
if candidate.is_absolute() and candidate.is_file():
return str(candidate)
directories = [Path(voice_timeline_path).parent] if voice_timeline_path else []
directories += [Path(d) for d in extra_dirs if d]
for directory in directories:
found = directory / candidate.name
if found.is_file():
return str(found)
return ""
def _overlap(a_start: float, a_end: float, b_start: float, b_end: float) -> float:
"""Seconds shared by two spans (0.0 when they don't touch)."""
return max(0.0, min(a_end, b_end) - max(a_start, b_start))
def _cut_coverage(
start: float, end: float, cuts: Sequence[Tuple[float, float]]
) -> float:
"""Fraction of ``start``-``end`` that falls inside ``cuts`` (0-1)."""
span = end - start
if span <= 0:
return 0.0
removed = sum(_overlap(start, end, c_start, c_end) for c_start, c_end in cuts)
return min(1.0, removed / span)
def snap_to_words(
time: float, words: Sequence[dict], fallback: float, edge: str
) -> float:
"""Move ``time`` onto the nearest word boundary of this phrase.
Trims are expressed by pointing at a word, so a trim handle that landed
mid-word would cut a syllable in half. ``edge`` is ``"in"`` (snap to word
starts) or ``"out"`` (snap to word ends); with no word timings available the
time is left as-is.
"""
boundaries = [
float(word.get("start" if edge == "in" else "end", 0.0)) for word in words
]
boundaries = [b for b in boundaries if b > 0]
if not boundaries:
return fallback
return min(boundaries, key=lambda b: abs(b - time))
def _trim_from_cuts(
start: float,
end: float,
words: Sequence[dict],
cuts: Sequence[Tuple[float, float]],
) -> Tuple[float, float]:
"""Read a partial cut over this phrase as a head/tail trim.
Only cuts that touch an edge become trims: a cut carved out of the middle of
a line has no representation here (the phrase is the unit), so it is left
for the whole-phrase coverage rule to decide.
"""
trim_start, trim_end = start, end
for cut_start, cut_end in cuts:
if _overlap(start, end, cut_start, cut_end) <= 0:
continue
if cut_start <= trim_start < cut_end < end:
trim_start = snap_to_words(cut_end, words, cut_end, "in")
if start < cut_start < trim_end <= cut_end:
trim_end = snap_to_words(cut_start, words, cut_start, "out")
if trim_end <= trim_start:
return start, end
return trim_start, trim_end
def _level_from_peak(peak: float) -> int:
"""Map a 0-1 ``peak_emphasis`` onto a 0-3 level."""
for level, threshold in enumerate(EMPHASIS_THRESHOLDS):
if peak < threshold:
return level
return 3
def _level_from_scale(scale: Optional[float]) -> int:
"""Map a zoom's scale factor back onto a 0-3 level.
The model is free to send any scale inside the allowed range, so this picks
the nearest level rather than requiring one of our own three values.
"""
if scale is None:
return 2
best = 1
smallest = None
for level, level_scale in ZOOM_SCALE_BY_LEVEL.items():
distance = abs(level_scale - float(scale))
if smallest is None or distance < smallest:
smallest, best = distance, level
return best
def _emphasis_from_actions(
start: float,
end: float,
actions: Sequence[VoiceAction],
) -> Tuple[Optional[int], str]:
"""The level the model asked for on this phrase, and why.
A ``zoom`` or ``text`` action anywhere inside the phrase is read as "this
line is the emphasis" — the model places them on the word that carries the
point, not on the whole line, so requiring a full-span match would find
nothing. Returns ``(None, "")`` when no action touches the phrase.
"""
level: Optional[int] = None
reason = ""
for action in actions:
if action.kind not in ("zoom", "text"):
continue
if _overlap(start, end, action.start, action.end) <= 0:
continue
if action.kind == "zoom":
candidate = _level_from_scale(action.params.get("scale"))
else:
candidate = 2
if level is None or candidate > level:
level = candidate
reason = action.reason
return level, reason
def _cut_reason(
start: float, end: float, actions: Sequence[VoiceAction]
) -> str:
"""The reason given for the cut that removes this phrase."""
for action in actions:
if action.kind != "cut":
continue
if _overlap(start, end, action.start, action.end) > 0 and action.reason:
return action.reason
return ""
def build_phrase_review(
timeline: dict,
actions: Any = None,
voice_timeline_path: str = "",
extra_dirs: Sequence[str] = (),
) -> dict:
"""Join a voice timeline with the AI's actions into a reviewable script.
``actions`` accepts whatever :func:`~.voice_actions.parse_actions` accepts —
a bare list, ``{"actions": [...]}``, or ``None`` when there is no AI pass and
the review starts from the acoustics alone. Malformed rows are skipped and
reported in ``errors`` rather than raising, matching the rest of the
decision pipeline.
"""
parsed, errors = parse_actions(actions) if actions else ([], [])
cuts = merge_cut_ranges(parsed)
phrases: List[dict] = []
for index, segment in enumerate(timeline.get("segments", [])):
start = float(segment.get("start", 0.0))
end = float(segment.get("end", 0.0))
peak = float(segment.get("peak_emphasis", 0.0))
take_boundary = bool(segment.get("take_boundary", False))
words = list(segment.get("words", []))
coverage = _cut_coverage(start, end, cuts)
active = coverage < CUT_COVERAGE_TO_DEACTIVATE
trim_start, trim_end = (
_trim_from_cuts(start, end, words, cuts) if active else (start, end)
)
asked_level, asked_reason = _emphasis_from_actions(start, end, parsed)
if asked_level is not None:
emphasis, reason = asked_level, asked_reason
else:
emphasis = _level_from_peak(peak)
reason = f"ênfase {peak:.2f}" if emphasis else ""
if not active:
# A removed line carries the reason it was removed; the emphasis it
# would have had is kept so re-activating it restores the decision.
reason = _cut_reason(start, end, parsed) or reason
phrases.append(
{
"index": index,
"start": round(start, 3),
"end": round(end, 3),
"trim_start": round(trim_start, 3),
"trim_end": round(trim_end, 3),
"text": str(segment.get("text", "")).strip(),
"speaker": str(segment.get("speaker", "")),
"active": active,
"emphasis": emphasis,
"track": TRACK_BACKSTAGE if (not active and take_boundary) else TRACK_SCRIPT,
"peak_emphasis": round(peak, 3),
# Delivery emotion is a heuristic over the acoustics (see
# voice_timeline._emotion_for_word) and only means anything when
# the analysis actually ran — `emotion_available` below is what
# separates "spoken flat" from "never measured".
"emotion": str(segment.get("emotion", "neutral")),
"emotion_confidence": round(
float(segment.get("emotion_confidence", 0.0)), 3
),
"take_boundary": take_boundary,
"gap_before": round(float(segment.get("gap_before", 0.0)), 3),
"reason": reason,
"words": [
{
"text": str(word.get("text", "")),
"start": round(float(word.get("start", 0.0)), 3),
"end": round(float(word.get("end", 0.0)), 3),
"energy": round(float(word.get("energy", 0.0)), 3),
"emphasis": round(float(word.get("emphasis", 0.0)), 3),
}
for word in words
],
}
)
source = timeline.get("source", "")
layers = timeline.get("layers", {}) if isinstance(timeline.get("layers"), dict) else {}
return {
"version": PHRASE_REVIEW_VERSION,
"source": source,
"source_path": resolve_source(source, voice_timeline_path, extra_dirs),
"rotation": float(timeline.get("rotation", 0.0)),
"duration": round(phrases[-1]["end"], 3) if phrases else 0.0,
"speakers": timeline.get("speakers", []),
"emotion_available": bool(layers.get("emotion", False)),
"phrases": phrases,
# Punch-ins the editor places by hand on an arbitrary range, alongside
# the whole-phrase zoom that an emphasis level produces. Both end up as
# zoom actions; this one exists because the moment worth punching into
# is not always a whole sentence.
"zooms": [],
"errors": errors,
}
def _coerce_zoom(raw: Any) -> Optional[Dict[str, float]]:
"""Normalize one manually placed zoom range."""
if not isinstance(raw, dict):
return None
try:
start = float(raw.get("start"))
end = float(raw.get("end"))
except (TypeError, ValueError):
return None
if end - start < MIN_ZOOM_DURATION:
return None
return {"start": start, "end": end}
def _coerce_phrase(raw: Any, index: int) -> Optional[Dict[str, Any]]:
"""Normalize one edited phrase row coming back from the UI."""
if not isinstance(raw, dict):
return None
try:
start = float(raw.get("start"))
end = float(raw.get("end"))
except (TypeError, ValueError):
return None
if end <= start:
return None
try:
emphasis = int(raw.get("emphasis", 0))
except (TypeError, ValueError):
emphasis = 0
try:
trim_start = float(raw.get("trim_start", start))
trim_end = float(raw.get("trim_end", end))
except (TypeError, ValueError):
trim_start, trim_end = start, end
# A trim that escaped the phrase, or inverted, is treated as no trim at all:
# the UI is the only thing that writes these, and silently discarding a bad
# pair keeps a rounding slip from deleting material the editor kept.
if not (start <= trim_start < trim_end <= end):
trim_start, trim_end = start, end
track = str(raw.get("track", TRACK_SCRIPT))
return {
"index": int(raw.get("index", index)),
"start": start,
"end": end,
"trim_start": trim_start,
"trim_end": trim_end,
"text": str(raw.get("text", "")).strip(),
"speaker": str(raw.get("speaker", "")),
"active": bool(raw.get("active", True)),
"emphasis": min(3, max(0, emphasis)),
"track": track if track in TRACKS else TRACK_SCRIPT,
"reason": str(raw.get("reason", "")),
}
def phrase_review_to_actions(review: dict) -> dict:
"""Turn an edited review back into the action list the applier consumes.
Every deactivated phrase becomes a ``cut``, a trimmed one becomes a cut over
the head and/or tail it lost, and every emphasized one becomes a ``zoom``
scaled by its level. The emphasis flags ride along in ``emphasis_spans`` so
the caption step can give those lines the dynamic treatment and everything
else the plain one, without re-deriving the decision from the acoustics.
"""
phrases = [
coerced
for index, raw in enumerate(review.get("phrases", []))
if (coerced := _coerce_phrase(raw, index)) is not None
]
actions: List[dict] = []
emphasis_spans: List[dict] = []
for phrase in phrases:
if not phrase["active"]:
actions.append(
VoiceAction(
kind="cut",
start=phrase["start"],
end=phrase["end"],
reason=phrase["reason"] or "desativada na revisão",
speaker=phrase["speaker"],
).as_dict()
)
continue
# Head and tail the editor trimmed off — each becomes its own cut, so a
# false start disappears without taking the line with it.
for trim_start, trim_end, where in (
(phrase["start"], phrase["trim_start"], "início"),
(phrase["trim_end"], phrase["end"], "fim"),
):
if trim_end - trim_start <= 0:
continue
actions.append(
VoiceAction(
kind="cut",
start=trim_start,
end=trim_end,
reason=f"trecho do {where} da frase removido na revisão",
speaker=phrase["speaker"],
).as_dict()
)
if phrase["emphasis"] >= 1:
actions.append(
VoiceAction(
kind="zoom",
start=phrase["trim_start"],
end=phrase["trim_end"],
params={"scale": ZOOM_SCALE_BY_LEVEL[phrase["emphasis"]]},
reason=phrase["reason"] or f"ênfase nível {phrase['emphasis']}",
speaker=phrase["speaker"],
).as_dict()
)
emphasis_spans.append(
{
"start": phrase["trim_start"],
"end": phrase["trim_end"],
"level": phrase["emphasis"],
"text": phrase["text"],
}
)
# Hand-placed punch-ins carry no scale on purpose: an omitted scale lets the
# applier use the shape configured in "Análise de Voz" (zoom_scale, ease in
# and out), so changing that setting restyles every manual zoom instead of
# leaving a scale frozen into each one at the moment it was drawn.
for raw in review.get("zooms", []):
zoom = _coerce_zoom(raw)
if zoom is None:
continue
actions.append(
VoiceAction(
kind="zoom",
start=zoom["start"],
end=zoom["end"],
reason="zoom marcado na revisão",
).as_dict()
)
return {
"source": review.get("source", ""),
"actions": actions,
"emphasis_spans": emphasis_spans,
}
def merge_saved_decisions(review: dict, saved: Optional[dict]) -> dict:
"""Lay a previously saved review's decisions over a freshly built one.
Only the editorial fields travel — active, emphasis, track, text, trims.
Everything else (words, emotion, energy) is re-derived from the current
analysis, so re-running the voice pass with better settings improves the
screen instead of being masked by a stale copy of itself, and the saved file
never has to carry a duplicate of data it does not own.
Phrases are matched by index *and* start time: if the analysis changed
enough to move a line, the old decision for that slot is dropped rather than
applied to a different sentence.
"""
if not saved:
return review
review["zooms"] = [
zoom for raw in saved.get("zooms", []) if (zoom := _coerce_zoom(raw)) is not None
]
by_index = {}
for raw in saved.get("phrases", []):
if isinstance(raw, dict) and "index" in raw:
by_index[raw["index"]] = raw
for phrase in review["phrases"]:
previous = by_index.get(phrase["index"])
if previous is None:
continue
if abs(float(previous.get("start", -1)) - phrase["start"]) > 0.25:
continue
phrase["active"] = bool(previous.get("active", phrase["active"]))
phrase["emphasis"] = min(3, max(0, int(previous.get("emphasis", phrase["emphasis"]))))
track = str(previous.get("track", phrase["track"]))
phrase["track"] = track if track in TRACKS else phrase["track"]
if previous.get("text"):
phrase["text"] = str(previous["text"])
trim_start = float(previous.get("trim_start", phrase["trim_start"]))
trim_end = float(previous.get("trim_end", phrase["trim_end"]))
if phrase["start"] <= trim_start < trim_end <= phrase["end"]:
phrase["trim_start"], phrase["trim_end"] = trim_start, trim_end
return review
def review_paths(voice_timeline_path: str) -> Tuple[Path, Path]:
"""Where the review and its derived actions live, next to the timeline.
Both files sit beside the ``_voice_timeline.json`` they came from and are
named after it, so a project folder stays readable and re-running the wizard
on the same take overwrites its own files instead of accumulating copies.
"""
base = Path(voice_timeline_path)
stem = base.stem
if stem.endswith("_voice_timeline"):
stem = stem[: -len("_voice_timeline")]
return (
base.with_name(f"{stem}_phrase_review.json"),
base.with_name(f"{stem}_phrase_actions.json"),
)
def save_phrase_review(voice_timeline_path: str, review: dict) -> Tuple[Path, Path]:
"""Write the edited review and the actions derived from it. Returns both paths."""
review_path, actions_path = review_paths(voice_timeline_path)
review_path.write_text(
json.dumps(review, ensure_ascii=False, indent=2), encoding="utf-8"
)
actions_path.write_text(
json.dumps(phrase_review_to_actions(review), ensure_ascii=False, indent=2),
encoding="utf-8",
)
return review_path, actions_path
def load_phrase_review(voice_timeline_path: str) -> Optional[dict]:
"""The review saved earlier for this timeline, or ``None`` if there is none."""
review_path, _ = review_paths(voice_timeline_path)
if not review_path.is_file():
return None
try:
data = json.loads(review_path.read_text(encoding="utf-8"))
except (OSError, json.JSONDecodeError):
return None
return data if isinstance(data, dict) else None
+46 -11
View File
@@ -28,8 +28,11 @@ ALLOWED_MODELS = (
) )
# Conservative by default: interjections that are near-universally filler. # Conservative by default: interjections that are near-universally filler.
# Portuguese "um"/"uma" are usually articles/numerals inside real phrases
# ("de um jeito") rather than discardable hesitations, so only cut them when
# the caller explicitly opts in through the fillers argument.
# "like" / "so" / "actually" are speech, not noise, unless the user opts in. # "like" / "so" / "actually" are speech, not noise, unless the user opts in.
DEFAULT_FILLERS = ("um", "uh", "uhh", "umm", "erm", "ehm", "mmm", "hmm", "mhm") DEFAULT_FILLERS = ("uh", "uhh", "umm", "erm", "ehm", "mmm", "hmm", "mhm")
_NORM_RE = re.compile(r"[^\w']+") _NORM_RE = re.compile(r"[^\w']+")
@@ -122,18 +125,27 @@ def transcribe(
model_size: str = "base", model_size: str = "base",
language: Optional[str] = None, language: Optional[str] = None,
progress_cb: Optional[Callable[[float], None]] = None, progress_cb: Optional[Callable[[float], None]] = None,
align: bool = True,
) -> Optional[dict]: ) -> Optional[dict]:
"""Transcribe an audio/video file locally with word-level timestamps. """Transcribe an audio/video file locally with word-level timestamps.
Requires the optional ``[transcribe]`` extra (faster-whisper). Returns Requires the optional ``[transcribe]`` extra (faster-whisper). Returns
``None`` when the model is unavailable or the file is missing/unreadable. ``None`` when the model is unavailable or the file is missing/unreadable.
When ``align`` is true (default) and the optional ``whisperx`` dependency is
present, word timestamps are refined by phonetic forced alignment, which
corrects faster-whisper's systematic ~0.3-0.5s early bias on word *starts*
(see ``Engine/docs/05_EXPERIENCIAS.md`` #14). The transcript reports
whether this ran via the ``alignment`` flag, so downstream consumers can
rely on the times without re-measuring.
The model weights are resolved from the configured models directory (see The model weights are resolved from the configured models directory (see
``model_manager.get_models_dir``), so a model selected/downloaded through ``model_manager.get_models_dir``), so a model selected/downloaded through
the app is found without an implicit download to the default HF cache. the app is found without an implicit download to the default HF cache.
Returns: Returns:
``{"language": str, "duration": float, "text": str, ``{"language": str, "duration": float, "text": str,
"alignment": bool,
"segments": [{"text", "start", "end", "start_fmt", "end_fmt"}, ...], "segments": [{"text", "start", "end", "start_fmt", "end_fmt"}, ...],
"words": [{"word", "start", "end", "confidence"}, ...]}`` "words": [{"word", "start", "end", "confidence"}, ...]}``
""" """
@@ -172,6 +184,7 @@ def transcribe(
vad_filter=True, vad_filter=True,
) )
segments: List[dict] = [] segments: List[dict] = []
raw_segments: List[dict] = []
words: List[dict] = [] words: List[dict] = []
# `info.duration` is known upfront (from the container), so each # `info.duration` is known upfront (from the container), so each
# segment's end time — yielded lazily as faster-whisper decodes — # segment's end time — yielded lazily as faster-whisper decodes —
@@ -180,6 +193,20 @@ def transcribe(
for seg in segments_iter: for seg in segments_iter:
start = float(seg.start) start = float(seg.start)
end = float(seg.end) end = float(seg.end)
seg_words: List[dict] = []
if progress_cb is not None and total_duration > 0:
progress_cb(min(end / total_duration, 1.0))
for w in seg.words or []:
ws = float(w.start)
we = float(w.end)
word = {
"word": w.word.strip(),
"start": ws,
"end": we,
"confidence": float(w.probability),
}
words.append(word)
seg_words.append(word)
segments.append( segments.append(
{ {
"text": seg.text.strip(), "text": seg.text.strip(),
@@ -189,19 +216,26 @@ def transcribe(
"end_fmt": format_timestamp(end), "end_fmt": format_timestamp(end),
} }
) )
if progress_cb is not None and total_duration > 0: raw_segments.append(
progress_cb(min(end / total_duration, 1.0))
for w in seg.words or []:
ws = float(w.start)
we = float(w.end)
words.append(
{ {
"word": w.word.strip(), "text": seg.text.strip(),
"start": ws, "start": start,
"end": we, "end": end,
"confidence": float(w.probability), "words": seg_words,
} }
) )
alignment_ran = False
if align and raw_segments:
from .forced_align import ForcedAligner
try:
words = ForcedAligner().align(
words, raw_segments, str(file_path), info.language, str(models_dir)
)
alignment_ran = True
except Exception:
logger.warning("forced alignment step failed; keeping raw timestamps")
except Exception: except Exception:
logger.warning("whisper transcription failed for %s", file_path) logger.warning("whisper transcription failed for %s", file_path)
return None return None
@@ -209,6 +243,7 @@ def transcribe(
"language": info.language, "language": info.language,
"duration": float(info.duration), "duration": float(info.duration),
"text": " ".join(s["text"] for s in segments), "text": " ".join(s["text"] for s in segments),
"alignment": alignment_ran,
"segments": segments, "segments": segments,
"words": words, "words": words,
} }
+16 -1
View File
@@ -82,8 +82,9 @@ def _validate_one(raw: Any, index: int) -> Tuple[Optional[VoiceAction], str]:
params = dict(params) if isinstance(params, dict) else {} params = dict(params) if isinstance(params, dict) else {}
if kind == "zoom": if kind == "zoom":
if "scale" in params and params.get("scale") is not None:
try: try:
scale = float(params.get("scale", 1.3)) scale = float(params["scale"])
except (TypeError, ValueError): except (TypeError, ValueError):
return None, f"{where}: zoom scale must be a number" return None, f"{where}: zoom scale must be a number"
if not (MIN_ZOOM_SCALE <= scale <= MAX_ZOOM_SCALE): if not (MIN_ZOOM_SCALE <= scale <= MAX_ZOOM_SCALE):
@@ -97,6 +98,20 @@ def _validate_one(raw: Any, index: int) -> Tuple[Optional[VoiceAction], str]:
if not content: if not content:
return None, f"{where}: text action needs params.content" return None, f"{where}: text action needs params.content"
params["content"] = content[:MAX_TEXT_LENGTH] params["content"] = content[:MAX_TEXT_LENGTH]
# Style is optional — omitted fields fall back to the "Legendas
# Dinâmicas" emphasis style at apply time (see _apply_placed_action),
# so a callout matches the captions' look without the caller having
# to know or repeat that configuration. Anything given here wins.
for key in ("font", "font_color", "face"):
if key in params and not isinstance(params[key], str):
del params[key]
if "font_size" in params:
try:
params["font_size"] = int(params["font_size"])
except (TypeError, ValueError):
del params["font_size"]
if "bold" in params:
params["bold"] = bool(params["bold"])
return ( return (
VoiceAction( VoiceAction(
+95
View File
@@ -56,12 +56,20 @@ VALUE_SCALES = {
"rate_delta": "0-1, how much the local speaking rate departs from the average", "rate_delta": "0-1, how much the local speaking rate departs from the average",
"pause_before": "seconds of silence immediately before the word", "pause_before": "seconds of silence immediately before the word",
"emphasis": "0-1 combined index; high values are punch-in/highlight candidates", "emphasis": "0-1 combined index; high values are punch-in/highlight candidates",
"emotion": "heuristic label from delivery: neutral, excited, tense, calm, reflective",
"emotion_confidence": "0-1 confidence in the heuristic emotion label",
"arousal": "0-1 vocal activation from energy/rate/pitch movement",
"valence": "0-1 rough positive tone; lower values suggest tension/weight",
}, },
"segment": { "segment": {
"gap_before": "seconds of silence before this line", "gap_before": "seconds of silence before this line",
"take_boundary": "true when the gap is long enough that the take likely restarted here", "take_boundary": "true when the gap is long enough that the take likely restarted here",
"avg_energy": "0-1 mean loudness across the line", "avg_energy": "0-1 mean loudness across the line",
"peak_emphasis": "0-1 highest emphasis of any word in the line", "peak_emphasis": "0-1 highest emphasis of any word in the line",
"emotion": "dominant delivery emotion across the line",
"emotion_confidence": "0-1 confidence in the dominant segment emotion",
"arousal": "0-1 mean vocal activation across the line",
"valence": "0-1 mean rough positive tone across the line",
}, },
} }
@@ -92,11 +100,74 @@ def _round_word(word: dict) -> dict:
"rate_delta": round(word.get("rate_delta", 0.0), 3), "rate_delta": round(word.get("rate_delta", 0.0), 3),
"pause_before": round(word.get("pause_before", 0.0), 3), "pause_before": round(word.get("pause_before", 0.0), 3),
"emphasis": round(word.get("emphasis", 0.0), 3), "emphasis": round(word.get("emphasis", 0.0), 3),
"emotion": word.get("emotion", "neutral"),
"emotion_confidence": round(word.get("emotion_confidence", 0.0), 3),
"arousal": round(word.get("arousal", 0.0), 3),
"valence": round(word.get("valence", 0.5), 3),
"energy_raw": word.get("energy"), "energy_raw": word.get("energy"),
"pitch_hz": word.get("pitch_hz"), "pitch_hz": word.get("pitch_hz"),
} }
def _emotion_for_word(word: dict, enabled: bool, sensitivity: float) -> dict:
"""Classify delivery emotion from normalized acoustic features.
This is deliberately a local heuristic rather than a claimed clinical
emotion model. It gives the editor a useful signal about delivery shape
while degrading predictably when acoustic extraction is unavailable.
"""
if not enabled:
return {
"emotion": "neutral",
"emotion_confidence": 0.0,
"arousal": 0.0,
"valence": 0.5,
}
energy = float(word.get("energy_norm", 0.0))
pitch = float(word.get("pitch_delta", 0.0))
rate = float(word.get("rate_delta", 0.0))
pause = min(float(word.get("pause_before", 0.0)) / 2.0, 1.0)
emphasis = float(word.get("emphasis", 0.0))
arousal = max(0.0, min(1.0, energy * 0.45 + pitch * 0.25 + rate * 0.20 + emphasis * 0.10))
valence = max(0.0, min(1.0, 0.55 + energy * 0.15 - pause * 0.20 - rate * 0.10))
if arousal >= 0.68 and valence >= 0.50:
label = "excited"
confidence = arousal
elif arousal >= 0.58 and valence < 0.50:
label = "tense"
confidence = max(arousal, 1.0 - valence)
elif arousal <= 0.28 and pause >= 0.25:
label = "reflective"
confidence = max(1.0 - arousal, pause)
elif arousal <= 0.35:
label = "calm"
confidence = 1.0 - arousal
else:
label = "neutral"
confidence = 1.0 - abs(arousal - 0.5) * 2.0
confidence = max(0.0, min(1.0, confidence))
if confidence < sensitivity:
label = "neutral"
return {
"emotion": label,
"emotion_confidence": confidence,
"arousal": arousal,
"valence": valence,
}
def annotate_emotions(words: Sequence[dict], enabled: bool, sensitivity: float) -> List[dict]:
"""Attach heuristic emotion labels to enriched word rows."""
return [
{**w, **_emotion_for_word(w, enabled, sensitivity)}
for w in words
]
def enrich_words( def enrich_words(
words: Sequence[dict], words: Sequence[dict],
pitch_track: Optional[Sequence] = None, pitch_track: Optional[Sequence] = None,
@@ -166,6 +237,13 @@ def _segment_rows(segments: Sequence[dict], words: Sequence[dict]) -> List[dict]
in_seg = [w for w in words if start <= float(w.get("start", 0.0)) < end] in_seg = [w for w in words if start <= float(w.get("start", 0.0)) < end]
energies = [w["energy_norm"] for w in in_seg] energies = [w["energy_norm"] for w in in_seg]
emphases = [w["emphasis"] for w in in_seg] emphases = [w["emphasis"] for w in in_seg]
arousals = [w.get("arousal", 0.0) for w in in_seg]
valences = [w.get("valence", 0.5) for w in in_seg]
emotions = [w.get("emotion", "neutral") for w in in_seg]
dominant = max(set(emotions), key=emotions.count) if emotions else "neutral"
emotion_confidences = [
w.get("emotion_confidence", 0.0) for w in in_seg if w.get("emotion") == dominant
]
gap = max(0.0, start - previous_end) gap = max(0.0, start - previous_end)
rows.append( rows.append(
{ {
@@ -181,6 +259,13 @@ def _segment_rows(segments: Sequence[dict], words: Sequence[dict]) -> List[dict]
"take_boundary": gap >= TAKE_BOUNDARY_GAP, "take_boundary": gap >= TAKE_BOUNDARY_GAP,
"avg_energy": round(sum(energies) / len(energies), 3) if energies else 0.0, "avg_energy": round(sum(energies) / len(energies), 3) if energies else 0.0,
"peak_emphasis": round(max(emphases), 3) if emphases else 0.0, "peak_emphasis": round(max(emphases), 3) if emphases else 0.0,
"emotion": dominant,
"emotion_confidence": (
round(sum(emotion_confidences) / len(emotion_confidences), 3)
if emotion_confidences else 0.0
),
"arousal": round(sum(arousals) / len(arousals), 3) if arousals else 0.0,
"valence": round(sum(valences) / len(valences), 3) if valences else 0.5,
"words": [_round_word(w) for w in in_seg], "words": [_round_word(w) for w in in_seg],
} }
) )
@@ -425,6 +510,9 @@ def build_voice_timeline(
weights: EmphasisWeights = EmphasisWeights(), weights: EmphasisWeights = EmphasisWeights(),
peak_percentile: float = 0.02, peak_percentile: float = 0.02,
emphasis_floor: float = 0.25, emphasis_floor: float = 0.25,
emotion_enabled: bool = False,
emotion_sensitivity: float = 0.5,
rotation: float = 0.0,
progress_cb: Optional[Callable[[float, str], None]] = None, progress_cb: Optional[Callable[[float, str], None]] = None,
) -> dict: ) -> dict:
"""Build the consolidated voice timeline for one media file. """Build the consolidated voice timeline for one media file.
@@ -445,6 +533,7 @@ def build_voice_timeline(
report(0.5, "Calculando ênfase...") report(0.5, "Calculando ênfase...")
words = enrich_words(transcript.get("words", []), pitch_track, energy_track, weights) words = enrich_words(transcript.get("words", []), pitch_track, energy_track, weights)
words = annotate_emotions(words, emotion_enabled, emotion_sensitivity)
report(0.7, "Identificando participantes...") report(0.7, "Identificando participantes...")
tracks = diarize(media_path, hf_token, num_speakers) if hf_token else None tracks = diarize(media_path, hf_token, num_speakers) if hf_token else None
@@ -457,6 +546,10 @@ def build_voice_timeline(
return { return {
"version": VOICE_TIMELINE_VERSION, "version": VOICE_TIMELINE_VERSION,
"source": Path(media_path).name, "source": Path(media_path).name,
# Edit-time correction from the clip's Transform filter in the FCPXML
# (e.g. straightening a tilted phone shot) — 0.0 when the clip has none
# or the caller didn't resolve one.
"rotation": rotation,
"language": transcript.get("language", ""), "language": transcript.get("language", ""),
# What actually ran, not what was installed — a consumer must be able # What actually ran, not what was installed — a consumer must be able
# to tell "this speech is flat" from "the acoustics never loaded", # to tell "this speech is flat" from "the acoustics never loaded",
@@ -465,6 +558,8 @@ def build_voice_timeline(
"transcript": bool(transcript.get("words")), "transcript": bool(transcript.get("words")),
"acoustics": pitch_track is not None or energy_track is not None, "acoustics": pitch_track is not None or energy_track is not None,
"speakers": tracks is not None, "speakers": tracks is not None,
"emotion": bool(emotion_enabled),
"alignment": bool(transcript.get("alignment")),
}, },
"scales": VALUE_SCALES, "scales": VALUE_SCALES,
"summary": _summary( "summary": _summary(
File diff suppressed because it is too large Load Diff
+130
View File
@@ -0,0 +1,130 @@
"""
FCPXML Writer — Generate and modify Final Cut Pro XML files.
This package provides two complementary workflows for working with FCPXML:
**Generation** (``FCPXMLWriter``, in :mod:`.generator`):
Build a new FCPXML document from Python dataclass objects (``Project``,
``Timeline``, ``Clip``, ``Marker``). Useful for creating rough cuts,
montage exports, and template-based projects.
**Modification** (``FCPXMLModifier``, in :mod:`.modifier`):
Load an existing FCPXML file, apply surgical edits (markers, trims,
reorders, transitions, speed changes, silence removal, etc.), and save.
This is the primary API used by the MCP server's tool handlers.
Layout
------
This was one 4.200-line module. It is now one module per subject, because the
subjects barely touch each other: whoever is fixing a zoom ramp has no reason
to scroll past subtitle layout to find it.
helpers sanitising, scales, shared element builders
document asset creation, timebases, serialisation (``write_fcpxml``)
validation structural checks (``validate_fcpxml``)
core ``ModifierCore``: load, indices, spine navigation, ``save``
<subject> one mixin per editing subject (markers, trim, speed, …)
modifier ``FCPXMLModifier`` = core + every mixin
generator ``FCPXMLWriter``
api one-line convenience wrappers
Everything the rest of the project imported from the old module is re-exported
here, so ``from fcpxml.writer import FCPXMLModifier`` keeps working unchanged —
including the underscore-prefixed helpers the test suite reaches for.
Architecture notes
------------------
- All time arithmetic uses ``TimeValue`` (rational fractions) — never floats —
to match FCPXML's native ``"600/2400s"`` format and avoid rounding drift.
- The ``FCPXMLModifier`` builds three in-memory indices at init
(``clips``, ``resources``, ``formats``) so lookups are O(1) by ID/name.
- Spine-based editing: clips live inside a ``<spine>`` element (the primary
storyline). Connected clips attach via ``lane`` attributes on spine clips.
Most editing methods find the target clip in the spine, mutate it, then
ripple offsets on subsequent siblings.
- ``write_fcpxml()`` handles DTD-compliant serialisation and optional
timebase enforcement for all output paths.
"""
from ..models import TimeValue
from .api import add_marker_to_file, modify_fcpxml, trim_clip_in_file
from .core import ModifierCore
from .document import (
_STILL_IMAGE_EXTENSIONS,
_enforce_standard_timebases,
_ensure_video_asset,
write_fcpxml,
)
from .generator import FCPXMLWriter
from .helpers import (
_ASSET_CLIP_CHILD_ORDER,
_CHILD_ORDER_INDEX,
_MAX_MARKER_NAME_LENGTH,
_MAX_NOTE_LENGTH,
CLIP_AND_AUDIO_TAGS,
CLIP_TAGS,
FCP_EFFECTS,
HOLD_AT_CUT_THRESHOLD,
SPINE_ELEMENT_TAGS,
START_AT_CUT_THRESHOLD,
_create_asset_element,
_dtd_insert,
_fmt_scale,
_probe_audio_info,
_sanitize_xml_value,
build_marker_element,
list_effects,
)
from .modifier import FCPXMLModifier
from .validation import (
_check_asset_sources,
_check_child_order,
_check_effect_refs,
_check_frame_alignment,
_check_required_attributes,
_check_timebases,
_document_frame_duration,
validate_fcpxml,
)
__all__ = [
"FCPXMLModifier",
"FCPXMLWriter",
"ModifierCore",
"TimeValue",
"FCP_EFFECTS",
"CLIP_TAGS",
"CLIP_AND_AUDIO_TAGS",
"SPINE_ELEMENT_TAGS",
"HOLD_AT_CUT_THRESHOLD",
"START_AT_CUT_THRESHOLD",
"add_marker_to_file",
"build_marker_element",
"list_effects",
"modify_fcpxml",
"trim_clip_in_file",
"validate_fcpxml",
"write_fcpxml",
# Internos que o resto do projeto (e a suíte) já importava deste módulo
# quando ele era um arquivo só. Ficam aqui para a divisão não virar uma
# quebra de API disfarçada de reorganização.
"_ASSET_CLIP_CHILD_ORDER",
"_CHILD_ORDER_INDEX",
"_MAX_MARKER_NAME_LENGTH",
"_MAX_NOTE_LENGTH",
"_STILL_IMAGE_EXTENSIONS",
"_check_asset_sources",
"_check_child_order",
"_check_effect_refs",
"_check_frame_alignment",
"_check_required_attributes",
"_check_timebases",
"_create_asset_element",
"_document_frame_duration",
"_dtd_insert",
"_enforce_standard_timebases",
"_ensure_video_asset",
"_fmt_scale",
"_probe_audio_info",
"_sanitize_xml_value",
]
+55
View File
@@ -0,0 +1,55 @@
"""Atalhos de uma linha para as operações mais comuns.
Extraído de writer.py — ver fcpxml/writer/__init__.py para o conjunto.
"""
from typing import Optional
from ..models import (
MarkerType,
)
from .modifier import FCPXMLModifier
# ============================================================================
# CONVENIENCE FUNCTIONS
# ============================================================================
def modify_fcpxml(filepath: str) -> FCPXMLModifier:
"""
Open an FCPXML file for modification.
Usage:
modifier = modify_fcpxml("project.fcpxml")
modifier.add_marker(...)
modifier.save("output.fcpxml")
"""
return FCPXMLModifier(filepath)
def add_marker_to_file(
filepath: str,
timecode: str,
name: str,
marker_type: str = "standard",
output_path: Optional[str] = None
) -> str:
"""Convenience function to add a marker to an FCPXML file."""
modifier = FCPXMLModifier(filepath)
modifier.add_marker_at_timeline(
timecode, name,
MarkerType.from_string(marker_type)
)
return modifier.save(output_path)
def trim_clip_in_file(
filepath: str,
clip_id: str,
trim_start: Optional[str] = None,
trim_end: Optional[str] = None,
output_path: Optional[str] = None
) -> str:
"""Convenience function to trim a clip in an FCPXML file."""
modifier = FCPXMLModifier(filepath)
modifier.trim_clip(clip_id, trim_start, trim_end)
return modifier.save(output_path)
+162
View File
@@ -0,0 +1,162 @@
"""Clipes de áudio e cama musical.
Extraído de writer.py — ver fcpxml/writer/__init__.py para o conjunto.
"""
import xml.etree.ElementTree as ET
from pathlib import Path
from typing import Optional
from ..models import (
TimeValue,
)
from .helpers import _create_asset_element, _dtd_insert, _probe_audio_info, _sanitize_xml_value
class AudioMixin:
"""Clipes de áudio e cama musical."""
# AUDIO CLIP OPERATIONS (v0.6.0)
# ========================================================================
def add_audio_clip(
self,
parent_clip_id: str,
asset_id: Optional[str] = None,
offset: str = "0s",
duration: Optional[str] = None,
role: str = "dialogue",
lane: int = -1,
src: Optional[str] = None,
) -> ET.Element:
"""Add an audio clip connected to an existing timeline clip.
Creates an <asset-clip> at a negative lane with audioRole attribute.
Supports hierarchical roles like "dialogue.boom", "music.score",
"effects.foley".
Args:
parent_clip_id: Name/ID of the clip to attach audio to.
asset_id: Existing asset reference ID. If None and src provided,
creates a new asset.
offset: Position relative to parent clip start.
duration: Duration of audio clip.
role: Audio role (e.g. "dialogue", "music.score", "effects.foley").
lane: Lane number (negative = below primary, default -1).
src: Path to audio file. Used to create a new asset if asset_id
is not provided.
Returns:
The created audio clip element.
"""
parent = self._require_clip(parent_clip_id)
# Resolve or create asset
if asset_id and asset_id in self.resources:
asset = self.resources[asset_id]
elif src:
# Create new asset in resources
resources = self.root.find('.//resources')
if resources is None:
raise ValueError("No <resources> element found in FCPXML")
asset_id = self._unique_resource_id(resources, 'r_audio1')
# The asset duration must reflect the real media length, not the
# requested clip duration — FCP flags assets that claim more
# media than the file contains.
probed = _probe_audio_info(src)
if probed:
rate = probed['sample_rate']
asset_duration = f"{round(probed['duration'] * rate)}/{rate}s"
else:
asset_duration = duration or "0s"
asset_elem = _create_asset_element(
resources, asset_id, Path(src).stem, src,
duration=asset_duration,
has_video="0", has_audio="1",
)
if probed:
asset_elem.set('audioSources', '1')
asset_elem.set('audioChannels', str(probed['channels']))
asset_elem.set('audioRate', str(probed['sample_rate']))
asset = {
'id': asset_id,
'name': Path(src).stem,
'duration': asset_duration,
'element': asset_elem,
}
self.resources[asset_id] = asset
else:
raise ValueError("Must provide either asset_id or src for audio clip")
clip_duration, source_start = self._resolve_clip_duration(asset, duration)
# Clamp so the clip never claims more media than the asset contains
asset_duration_tv = self._parse_time(asset.get('duration', '0s'))
if asset_duration_tv > TimeValue.zero():
available = asset_duration_tv - source_start
if available < TimeValue.zero():
raise ValueError(
f"Source start {source_start.to_fcpxml()} is beyond the end "
f"of audio asset '{asset.get('name')}' "
f"({asset_duration_tv.to_fcpxml()})"
)
if clip_duration > available:
clip_duration = available
new_clip = self._make_asset_clip(
asset_id, asset.get('name', 'Audio'),
self._parse_time(offset), source_start, clip_duration,
lane=str(lane),
audioRole=_sanitize_xml_value(role, 256),
)
_dtd_insert(parent, new_clip)
return new_clip
def add_music_bed(
self,
asset_id: Optional[str] = None,
duration: Optional[str] = None,
role: str = "music",
src: Optional[str] = None,
) -> ET.Element:
"""Add a music bed spanning the full timeline at lane -1.
Convenience method: attaches to the first spine clip and spans
the full timeline duration.
Args:
asset_id: Existing asset reference ID.
duration: Override duration (default: full timeline).
role: Audio role (default "music").
src: Path to audio file (creates asset if asset_id not given).
Returns:
The created music bed clip element.
"""
spine = self._get_spine()
first_clip = None
first_clip_id = None
for clip_id, clip in self.clips.items():
if clip in list(spine):
first_clip = clip
first_clip_id = clip_id
break
if first_clip is None:
raise ValueError("No clips in spine to attach music bed to")
# Calculate full timeline duration if not specified
if not duration:
duration = self._timeline_duration().to_fcpxml()
return self.add_audio_clip(
parent_clip_id=first_clip_id,
asset_id=asset_id,
offset="0s",
duration=duration,
role=role,
lane=-1,
src=src,
)
# ========================================================================
+196
View File
@@ -0,0 +1,196 @@
"""Compound clips: criar e achatar.
Extraído de writer.py — ver fcpxml/writer/__init__.py para o conjunto.
"""
import copy
import uuid
import xml.etree.ElementTree as ET
from typing import List
from ..models import (
TimeValue,
)
from .helpers import (
_sanitize_xml_value,
)
class CompoundMixin:
"""Compound clips: criar e achatar."""
# COMPOUND CLIP OPERATIONS (v0.6.0)
# ========================================================================
def create_compound_clip(
self,
clip_ids: List[str],
name: str = "Compound Clip",
) -> ET.Element:
"""Group spine clips into a compound clip.
Creates a <media> resource with a nested <sequence><spine> containing
the specified clips, then replaces the originals in the main spine
with a single <ref-clip>.
Args:
clip_ids: IDs of clips in the spine to group.
name: Name for the compound clip.
Returns:
The created <ref-clip> element.
"""
spine = self._get_spine()
resources = self.root.find('.//resources')
if resources is None:
raise ValueError("No <resources> element found in FCPXML")
# Collect clips and validate they're in spine
spine_children = list(spine)
clips_to_group = []
for cid in clip_ids:
clip = self._require_clip(cid)
if clip not in spine_children:
raise ValueError(f"Clip not in spine: {cid}")
clips_to_group.append((cid, clip))
if not clips_to_group:
raise ValueError("No valid clips to group")
# Sort by offset so the compound maintains order
clips_to_group.sort(
key=lambda c: self._parse_time(c[1].get('offset', '0s'))
)
# Calculate compound duration and starting offset
first_offset = self._parse_time(clips_to_group[0][1].get('offset', '0s'))
total_duration = TimeValue.zero()
for _, clip in clips_to_group:
total_duration = total_duration + self._parse_time(clip.get('duration', '0s'))
# Get format ref
format_id = None
for fmt_id in self.formats:
format_id = fmt_id
break
# Create media resource with nested sequence
media_id = self._unique_resource_id(resources, 'r_compound1')
media = ET.SubElement(resources, 'media')
media.set('id', media_id)
media.set('name', _sanitize_xml_value(name, 512))
media.set('uid', str(uuid.uuid4()).upper())
seq = ET.SubElement(media, 'sequence')
seq.set('format', format_id or 'r1')
seq.set('duration', total_duration.to_fcpxml())
seq.set('tcStart', '0s')
seq.set('tcFormat', 'NDF')
inner_spine = ET.SubElement(seq, 'spine')
# Move clips into the compound's inner spine
inner_offset = TimeValue.zero()
for _, clip in clips_to_group:
new_clip = copy.deepcopy(clip)
new_clip.set('offset', inner_offset.to_fcpxml())
inner_spine.append(new_clip)
inner_offset = inner_offset + self._parse_time(clip.get('duration', '0s'))
# Get the insert position (where first clip was)
spine_children = list(spine)
insert_idx = spine_children.index(clips_to_group[0][1])
# Remove originals from spine
for cid, clip in clips_to_group:
spine.remove(clip)
if cid in self.clips:
del self.clips[cid]
# Create ref-clip in main spine
ref_clip = ET.Element('ref-clip')
ref_clip.set('ref', media_id)
ref_clip.set('offset', first_offset.to_fcpxml())
ref_clip.set('name', _sanitize_xml_value(name, 512))
ref_clip.set('duration', total_duration.to_fcpxml())
spine.insert(insert_idx, ref_clip)
# Index the new ref-clip
compound_id = f"compound_{name}"
self.clips[compound_id] = ref_clip
return ref_clip
def flatten_compound_clip(
self,
ref_clip_id: str,
) -> List[ET.Element]:
"""Flatten a compound clip back into individual spine clips.
Extracts clips from the compound's inner sequence and places them
back in the main spine at the ref-clip's position.
Args:
ref_clip_id: ID of the ref-clip to flatten.
Returns:
List of extracted clip elements now in the main spine.
"""
spine = self._get_spine()
ref_clip = self._require_clip(ref_clip_id)
if ref_clip.tag != 'ref-clip':
raise ValueError(f"Element is not a ref-clip: {ref_clip_id}")
media_ref = ref_clip.get('ref', '')
ref_offset = self._parse_time(ref_clip.get('offset', '0s'))
# Find the media resource
resources = self.root.find('.//resources')
media_elem = None
if resources is not None:
for m in resources.findall('media'):
if m.get('id') == media_ref:
media_elem = m
break
if media_elem is None:
raise ValueError(f"Media resource not found for ref: {media_ref}")
inner_spine = media_elem.find('.//spine')
if inner_spine is None:
raise ValueError("No spine found in compound clip media")
# Get insert position
spine_children = list(spine)
insert_idx = spine_children.index(ref_clip)
# Remove ref-clip from spine
spine.remove(ref_clip)
if ref_clip_id in self.clips:
del self.clips[ref_clip_id]
# Extract clips from inner spine into main spine
extracted = []
current_offset = ref_offset
for child in list(inner_spine):
new_clip = copy.deepcopy(child)
new_clip.set('offset', current_offset.to_fcpxml())
spine.insert(insert_idx, new_clip)
insert_idx += 1
extracted.append(new_clip)
current_offset = current_offset + self._parse_time(
child.get('duration', '0s')
)
# Index the extracted clip
clip_name = new_clip.get('name') or new_clip.get('id') or f"flat_{len(self.clips)}"
self.clips[clip_name] = new_clip
# Clean up media resource
if resources is not None:
resources.remove(media_elem)
return extracted
# ========================================================================
+49
View File
@@ -0,0 +1,49 @@
"""Clipes conectados (lanes acima/abaixo da spine).
Extraído de writer.py — ver fcpxml/writer/__init__.py para o conjunto.
"""
import xml.etree.ElementTree as ET
from typing import Optional
class ConnectedMixin:
"""Clipes conectados (lanes acima/abaixo da spine)."""
# CONNECTED CLIP OPERATIONS (v0.5.0)
# ========================================================================
def add_connected_clip(
self,
parent_clip_id: str,
asset_id: Optional[str] = None,
asset_name: Optional[str] = None,
offset: str = "0s",
duration: Optional[str] = None,
lane: int = 1,
) -> ET.Element:
"""Add a connected clip (B-roll, title, audio) to an existing timeline clip.
Args:
parent_clip_id: Name/ID of the clip to attach to
asset_id: Asset reference ID
asset_name: Asset name (alternative to asset_id)
offset: Position relative to parent clip start
duration: Duration of connected clip (default: full asset)
lane: Lane number (positive=above, negative=below)
Returns:
The created connected clip element
"""
parent = self._require_clip(parent_clip_id)
asset, asset_id = self._resolve_asset(asset_id, asset_name)
clip_duration, source_start = self._resolve_clip_duration(asset, duration)
new_clip = self._make_asset_clip(
asset_id, asset.get('name', 'Untitled'),
self._parse_time(offset), source_start, clip_duration,
parent=parent, lane=str(lane),
)
return new_clip
# ========================================================================
+723
View File
@@ -0,0 +1,723 @@
"""Núcleo do FCPXMLModifier: carga, índices, navegação na spine e save.
Extraído de writer.py — ver fcpxml/writer/__init__.py para o conjunto.
"""
import xml.etree.ElementTree as ET
from fractions import Fraction
from pathlib import Path
from typing import Any, Dict, Optional, Tuple
from ..models import (
TimeValue,
)
from .document import write_fcpxml
from .helpers import CLIP_TAGS
class ModifierCore:
"""Load an existing FCPXML file, apply edits, and save.
This is the primary editing interface used by every MCP server write-tool
handler. It wraps an ElementTree parsed from disk and maintains three
in-memory indices so that clip/asset lookups are fast.
Index design
------------
``clips`` : ``Dict[str, ET.Element]``
Every ``<clip>``, ``<asset-clip>``, and ``<video>`` element keyed by
its ``id`` attribute, falling back to ``name``, then a generated key.
**Gotcha**: duplicate clip names (e.g. multiple "Interview_A") mean
only the *last* element indexed under that name is accessible. Use
unique ``id`` attributes when possible.
``resources`` : ``Dict[str, Dict[str, Any]]``
Every ``<asset>`` element keyed by ``id``, with pre-extracted ``name``,
``src``, ``start``, ``duration``, and a reference to the raw element.
``formats`` : ``Dict[str, Dict[str, Any]]``
Every ``<format>`` element keyed by ``id``.
Editing model
-------------
1. Look up the target clip via ``_require_clip`` / ``_require_spine_clip``.
2. Mutate the clip's XML attributes (``start``, ``duration``, ``offset``).
3. If the edit changes duration, ripple subsequent spine siblings via
``_ripple_from_index`` so downstream offsets stay contiguous.
4. Call ``save()`` to serialise the modified tree back to disk.
Example::
modifier = FCPXMLModifier("project.fcpxml")
modifier.add_marker("clip_0", "00:00:10:00", "Review", MarkerType.INCOMPLETE)
modifier.trim_clip("clip_1", trim_end="-2s")
modifier.save("project_modified.fcpxml")
Attributes:
path (Path): Filesystem path to the source FCPXML file.
tree (ET.ElementTree): Parsed XML tree (mutated in-place by edits).
root (ET.Element): Root ``<fcpxml>`` element.
fps (float): Detected frame rate from the first ``<format>`` resource.
clips (Dict[str, ET.Element]): Clip index — see *Index design* above.
resources (Dict[str, Dict]): Asset index.
formats (Dict[str, Dict]): Format index.
"""
def __init__(self, fcpxml_path: str):
"""Load *fcpxml_path*, parse its XML, and build lookup indices.
The constructor eagerly builds all three indices (clips, resources,
formats) and detects the project frame rate. After construction the
modifier is ready for any editing operation.
Args:
fcpxml_path: Absolute or relative path to an ``.fcpxml`` file or
an ``.fcpxmld`` bundle (a directory wrapping ``Info.fcpxml``
plus sidecar data files for object tracking / Cinematic mode).
Raises:
FileNotFoundError: If *fcpxml_path* does not exist.
ET.ParseError: If the file is not valid XML.
ValueError: If no ``<spine>`` is found (checked lazily on first edit).
"""
path = Path(fcpxml_path)
self.bundle_dir: Optional[Path] = None
if path.suffix.lower() == '.fcpxmld':
self.bundle_dir = path
inner = path / 'Info.fcpxml'
if not inner.exists():
raise FileNotFoundError(
f"Info.fcpxml not found in bundle: {fcpxml_path}"
)
fcpxml_path = str(inner)
self.path = Path(fcpxml_path)
from ..safe_xml import safe_parse
self.tree = safe_parse(fcpxml_path)
self.root = self.tree.getroot()
self.fps = self._detect_fps()
# Lazily filled on the first generated title; see _unique_text_style_id.
self._text_style_ids: Optional[set] = None
self._build_resource_index()
self._build_clip_index()
def _detect_fps(self) -> float:
"""Extract frame rate from format resource."""
for fmt in self.root.findall('.//format'):
frame_dur = fmt.get('frameDuration', '1/30s')
if '/' in frame_dur:
parts = frame_dur.replace('s', '').split('/', 1)
num, denom = int(parts[0]), int(parts[1])
if num <= 0:
return 30.0
return denom / num
return 30.0
def frame_duration_fraction(self):
"""Exact ``frameDuration`` as a Fraction (e.g. 1001/24000 at 23.976fps).
Unlike ``_detect_fps()`` (a float, lossy for NTSC rates), this is
exact — use it wherever a cut boundary is snapped to the frame grid,
so 23.976/29.97/59.94 timebases don't drift off-grid the way a
hardcoded tick base like 2400 does.
"""
for fmt in self.root.findall('.//format'):
raw = fmt.get('frameDuration', '')
if raw.endswith('s') and '/' in raw:
n, d = raw[:-1].split('/', 1)
fd = Fraction(int(n), int(d))
if fd > 0:
return fd
return Fraction(1, 30)
def frame_size(self) -> 'Tuple[float, float]':
"""The sequence's frame size in pixels, as ``(width, height)``.
Reads the sequence's own ``<format>`` when it references one, since a
document may carry several (an asset's source format need not match
the timeline's). Falls back to the first format that declares a size,
then to 1920x1080.
"""
formats = {f.get('id'): f for f in self.root.findall('.//format')}
candidates = []
seq = self.root.find('.//sequence')
if seq is not None and formats.get(seq.get('format')) is not None:
candidates.append(formats[seq.get('format')])
candidates.extend(formats.values())
for fmt in candidates:
try:
width = float(fmt.get('width') or 0)
height = float(fmt.get('height') or 0)
except (TypeError, ValueError):
continue
if width > 0 and height > 0:
return width, height
return 1920.0, 1080.0
def frame_width(self) -> float:
"""The sequence's frame width in pixels."""
return self.frame_size()[0]
def frame_height(self) -> float:
"""The sequence's frame height in pixels."""
return self.frame_size()[1]
def snap_seconds_to_frame(self, seconds: float) -> 'TimeValue':
"""Round *seconds* to the nearest exact frame boundary as a TimeValue."""
fd = self.frame_duration_fraction()
frames = round(seconds / float(fd))
snapped = fd * frames
return TimeValue(snapped.numerator, snapped.denominator)
def snap_spine_times_to_frames(self) -> None:
"""Snap primary-storyline offsets and durations to sequence frames.
Final Cut rejects otherwise valid XML when ripple edits leave a clip
boundary between frames. Use the exact ``frameDuration`` fraction,
rather than a float FPS, to preserve 23.976/29.97 timebases.
"""
frame_duration = None
for fmt in self.root.findall('.//format'):
raw = fmt.get('frameDuration', '')
if raw.endswith('s') and '/' in raw:
n, d = raw[:-1].split('/', 1)
frame_duration = Fraction(int(n), int(d))
break
if frame_duration is None or frame_duration <= 0:
return
for element in self.root.findall('.//spine/*'):
for attr in ('offset', 'duration'):
raw = element.get(attr)
if not raw or not raw.endswith('s'):
continue
value = raw[:-1]
if '/' in value:
n, d = value.split('/', 1)
seconds = Fraction(int(n), int(d))
else:
seconds = Fraction(value)
frames = int(round(float(seconds / frame_duration)))
snapped = frame_duration * frames
element.set(attr, f'{snapped.numerator}/{snapped.denominator}s')
def _build_resource_index(self) -> None:
"""Build ``self.resources`` and ``self.formats`` from ``<asset>``/``<format>`` elements.
Called once during ``__init__``. Each asset entry stores the raw
element plus pre-extracted metadata so callers don't need to
re-parse attributes on every access.
"""
self.resources: Dict[str, Dict[str, Any]] = {}
self.formats: Dict[str, Dict[str, Any]] = {}
for asset in self.root.findall('.//asset'):
asset_id = asset.get('id', '')
self.resources[asset_id] = {
'id': asset_id,
'name': asset.get('name', ''),
'src': asset.get('src', '') or (asset.find('media-rep').get('src', '') if asset.find('media-rep') is not None else ''),
'start': asset.get('start', '0s'),
'duration': asset.get('duration', '0s'),
'element': asset
}
for fmt in self.root.findall('.//format'):
fmt_id = fmt.get('id', '')
self.formats[fmt_id] = {
'id': fmt_id,
'name': fmt.get('name', ''),
'element': fmt
}
def _index_elements(self, tag: str, fallback_prefix: str) -> None:
"""Index XML elements of *tag* into ``self.clips`` by id/name.
Each element is keyed by its ``id`` attribute, falling back to
``name``, then a generated ``{fallback_prefix}_{i}`` key. This
replaces three near-identical loops that only differed in the tag
name and fallback prefix.
"""
for i, elem in enumerate(self.root.findall(f'.//{tag}')):
key = elem.get('id') or elem.get('name') or f"{fallback_prefix}_{i}"
self.clips[key] = elem
def _build_clip_index(self) -> None:
"""Build ``self.clips`` index from all clip-type elements.
Indexes ``<clip>``, ``<asset-clip>``, and ``<video>`` tags. Keys are
resolved by ``_index_elements`` (``id`` → ``name`` → generated).
.. warning::
Duplicate names cause last-one-wins overwrites. If your project
has multiple clips named "Interview_A", only the last one parsed
will be reachable by name. Prefer unique ``id`` attributes.
"""
self.clips: Dict[str, ET.Element] = {}
for tag, prefix in (('clip', 'clip'), ('asset-clip', 'asset_clip'), ('video', 'video')):
self._index_elements(tag, prefix)
def _get_spine(self) -> ET.Element:
"""Get the primary storyline spine.
Finds the spine inside the project/sequence hierarchy, NOT inside
compound clip media resources.
"""
# Prefer the main timeline spine (under project/sequence)
spine = self.root.find('.//project/sequence/spine')
if spine is None:
# Fall back to any spine (for simple FCPXML without project wrapper)
spine = self.root.find('.//spine')
if spine is None:
raise ValueError("No spine found in FCPXML")
return spine
def _iter_spine_clips(self) -> list[tuple[int, ET.Element]]:
"""Return an indexed list of clip-type elements in the primary spine.
Filters out gaps, transitions, and other non-clip elements, returning
only ``(index_in_spine, element)`` pairs where the tag is in
``CLIP_TAGS``. The index is the element's position among *all* spine
children (not just clips), so it stays valid for insertion/removal.
"""
spine = self._get_spine()
return [
(i, child)
for i, child in enumerate(spine.findall('*'))
if child.tag in CLIP_TAGS
]
def _find_spine_clip_at_seconds(self, target_seconds: float) -> tuple[ET.Element, float]:
"""Find the spine clip containing *target_seconds* and return it with the relative offset.
Returns:
``(clip_element, relative_seconds)`` — the clip and the time
within that clip corresponding to *target_seconds*.
Raises:
ValueError: If no clip spans the requested position.
"""
spine = self._get_spine()
for child in spine.findall('*'):
if child.tag not in CLIP_TAGS:
continue
offset = self._parse_time(child.get('offset', '0s')).to_seconds()
dur = self._parse_time(child.get('duration', '0s')).to_seconds()
if offset <= target_seconds < offset + dur:
return child, target_seconds - offset
raise ValueError(f"No spine clip at position {target_seconds:.3f}s")
def _parse_time(self, tc: str) -> TimeValue:
"""Parse a timecode string to TimeValue."""
return TimeValue.from_timecode(tc, self.fps)
def _get_clip_times(
self, clip: ET.Element
) -> tuple:
"""Return (start, duration, offset) TimeValues for a clip element."""
return (
self._parse_time(clip.get('start', '0s')),
self._parse_time(clip.get('duration', '0s')),
self._parse_time(clip.get('offset', '0s')),
)
def source_file_start(self, clip: ET.Element) -> 'TimeValue':
"""Return a clip's in-point measured from the head of its media file.
FCPXML ``start`` on an asset-clip is a source *timecode*, and the
asset's own ``start`` is the timecode of the source media's first
frame. Media analysis (ffmpeg silencedetect, Whisper) reports
file-relative time, so subtract the asset's start timecode to land
both on the same origin. When the asset starts at 0s (the common
case, and every test fixture) this is a no-op.
"""
ref = clip.get('ref', '')
asset = self.resources.get(ref, {})
asset_start = self._parse_time(asset.get('start', '0s'))
clip_start = self._parse_time(clip.get('start', '0s'))
return clip_start - asset_start
def _resolve_clip_duration(
self,
asset: dict,
duration: Optional[str] = None,
in_point: Optional[str] = None,
out_point: Optional[str] = None,
) -> tuple['TimeValue', 'TimeValue']:
"""Compute clip duration and source start from optional overrides.
Centralises the three-way fallback logic shared by insert_clip,
add_connected_clip, and add_audio_clip:
1. If *in_point* and *out_point* are given → subclip range.
2. Else if *duration* is given → explicit duration, source start = 0.
3. Else → full asset duration, source start = 0.
Returns:
``(clip_duration, source_start)`` TimeValue pair.
"""
if in_point and out_point:
in_time = self._parse_time(in_point)
out_time = self._parse_time(out_point)
return out_time - in_time, in_time
if duration:
return self._parse_time(duration), TimeValue.zero()
return self._parse_time(asset.get('duration', '0s')), TimeValue.zero()
def _make_asset_clip(
self,
asset_id: str,
name: str,
offset: 'TimeValue',
start: 'TimeValue',
duration: 'TimeValue',
*,
parent: Optional[ET.Element] = None,
**extra_attrs: str,
) -> ET.Element:
"""Build an ``<asset-clip>`` element with standard attributes.
Centralises the repeated element creation shared by insert_clip,
add_connected_clip, and add_audio_clip. Each caller can pass
additional attributes (``lane``, ``audioRole``, ``format``) via
*extra_attrs*.
Args:
asset_id: Resource reference (e.g. ``'r3'``).
name: Human-readable clip name.
offset: Timeline offset (or offset within parent for connected clips).
start: Source media start point.
duration: Clip duration.
parent: If given, create the element as a SubElement of *parent*;
otherwise create a detached Element.
**extra_attrs: Additional XML attributes (``lane``, ``audioRole``).
Returns:
The new ``<asset-clip>`` Element.
"""
if parent is not None:
elem = ET.SubElement(parent, 'asset-clip')
else:
elem = ET.Element('asset-clip')
elem.set('ref', asset_id)
elem.set('offset', offset.to_fcpxml())
elem.set('name', name)
elem.set('start', start.to_fcpxml())
elem.set('duration', duration.to_fcpxml())
for attr, val in extra_attrs.items():
elem.set(attr, val)
return elem
def _require_clip(self, clip_id: 'str | ET.Element') -> ET.Element:
"""Look up a clip by ID/name, raising if not found.
Centralises the get-or-raise pattern used by every clip-mutating
method so the error message stays consistent and future
enhancements (fuzzy matching, suggestions) only need one site.
An Element is returned as-is. That matters after ``split_clip`` or
``cut_clip_ranges``: the resulting pieces all carry the *same* name,
so a name lookup would always resolve to the first one and silently
put the edit on the wrong piece. Callers holding the exact element
pass it directly.
"""
if isinstance(clip_id, ET.Element):
return clip_id
clip = self.clips.get(clip_id)
if clip is None:
raise ValueError(f"Clip not found: {clip_id}")
return clip
def _require_spine_clip(self, clip_id: str) -> tuple[ET.Element, ET.Element, int]:
"""Look up a clip and verify it lives in the primary spine.
Returns:
``(spine, clip, index_in_spine)`` tuple.
Raises:
ValueError: If the clip doesn't exist or isn't in the spine.
"""
clip = self._require_clip(clip_id)
spine = self._get_spine()
clip_index = self._find_clip_index(spine, clip)
if clip_index is None:
raise ValueError(f"Clip not in spine: {clip_id}")
return spine, clip, clip_index
def _find_clip_index(self, spine: ET.Element, clip: ET.Element) -> int | None:
"""Find the index of a clip in the spine. Returns None if not found."""
for i, child in enumerate(spine):
if child == clip:
return i
return None
@staticmethod
def _find_neighbor_clip(
spine_list: list, index: int, direction: str
) -> Optional[ET.Element]:
"""Find the nearest non-gap clip before or after *index* in *spine_list*.
Args:
spine_list: Materialised list of spine children.
index: Position to search from (exclusive).
direction: ``'prev'`` to search backward, ``'next'`` to search forward.
Returns:
The first clip-type element found, or ``None``.
"""
if direction == 'prev':
for j in range(index - 1, -1, -1):
if spine_list[j].tag in CLIP_TAGS:
return spine_list[j]
else:
for j in range(index + 1, len(spine_list)):
if spine_list[j].tag in CLIP_TAGS:
return spine_list[j]
return None
def _resolve_asset(
self, asset_id: Optional[str], asset_name: Optional[str]
) -> tuple:
"""Look up an asset by ID or name from ``self.resources``.
Returns:
``(asset_dict, resolved_asset_id)`` tuple.
Raises:
ValueError: If neither ID nor name matches a known asset.
"""
if asset_id and asset_id in self.resources:
return self.resources[asset_id], asset_id
if asset_name:
for res_id, res_data in self.resources.items():
if res_data.get('name') == asset_name:
return res_data, res_id
raise ValueError(f"Asset not found: {asset_id or asset_name}")
@staticmethod
def _unique_resource_id(resources: ET.Element, prefix: str) -> str:
"""Generate a unique resource ID with the given *prefix*.
Starts with ``prefix`` (e.g. ``'r_audio1'``), appending an
incrementing counter until no collision exists in *resources*.
"""
existing_ids = {el.get('id', '') for el in resources}
candidate = prefix
counter = 2
while candidate in existing_ids:
# Strip trailing digits from prefix for the counter suffix
base = prefix.rstrip('0123456789')
candidate = f'{base}{counter}'
counter += 1
return candidate
def _find_spine_element_at_timecode(
self, spine: ET.Element, target_tc: str, *, require_clip: bool = False
) -> Optional[ET.Element]:
"""Find the first spine child whose offset matches *target_tc*.
Normalises both sides through ``TimeValue`` round-trip so format
differences (e.g. ``"3600/2400s"`` vs ``"1800/1200s"``) don't
cause false negatives.
Args:
spine: The ``<spine>`` element to search.
target_tc: Timecode string to match against each child's offset.
require_clip: If True, skip non-clip elements (gaps, etc.).
"""
for child in spine:
offset_str = child.get('offset', '0s')
tc = TimeValue.from_timecode(offset_str, self.fps).to_timecode(self.fps)
if tc == target_tc:
if require_clip and child.tag not in CLIP_TAGS:
continue
return child
return None
def _absorb_into_neighbor(
self,
spine: ET.Element,
element: ET.Element,
direction: str,
) -> Optional[ET.Element]:
"""Extend a neighbor clip to absorb *element*'s duration, then remove *element*.
Shared by ``fix_flash_frames`` (absorbing flash-frame clips) and
``fill_gaps`` (absorbing gap elements). Both operations find the
nearest clip in *direction*, grow it by the absorbed element's
duration, and remove the absorbed element from the spine.
When extending backward (``direction='next'``), the neighbor's
source in-point is also pulled earlier so the extra frames come
from before the original cut, not after.
Does **not** call ``_recalculate_offsets`` — callers decide when to
recalculate (per-iteration vs. once at the end).
Args:
spine: The primary storyline ``<spine>`` element.
element: The clip or gap to absorb (will be removed).
direction: ``'prev'`` to extend the previous clip forward,
``'next'`` to extend the next clip backward.
Returns:
The neighbor clip that absorbed the duration, or ``None`` if
no suitable neighbor exists.
"""
spine_list = list(spine)
element_index = spine_list.index(element)
neighbor = self._find_neighbor_clip(spine_list, element_index, direction)
if neighbor is None:
return None
absorbed_dur = self._parse_time(element.get('duration', '0s'))
neighbor_dur = self._parse_time(neighbor.get('duration', '0s'))
if direction == 'next':
neighbor_start = self._parse_time(neighbor.get('start', '0s'))
new_start = neighbor_start - absorbed_dur
if new_start >= TimeValue.zero():
neighbor.set('start', new_start.to_fcpxml())
neighbor.set('duration', (neighbor_dur + absorbed_dur).to_fcpxml())
else:
# Can't shift start negative — only extend by what's available
available = neighbor_start
neighbor.set('start', TimeValue(0, 1).to_fcpxml())
neighbor.set('duration', (neighbor_dur + available).to_fcpxml())
else:
neighbor.set('duration', (neighbor_dur + absorbed_dur).to_fcpxml())
spine.remove(element)
return neighbor
def _resolve_insert_position(
self, position: str, spine_children: list
) -> tuple:
"""Translate a human-friendly position spec into (target_offset, insert_index).
Supported formats:
``'start'`` — beginning of spine
``'end'`` — after last element
``'after:clip_id'`` — after the named clip
``'before:clip_id'``— before the named clip
*timecode* — absolute timeline position
Returns:
``(TimeValue, int)`` — the offset and child-index for spine insertion.
"""
if position == 'start':
return TimeValue.zero(), 0
if position == 'end':
if spine_children:
last = spine_children[-1]
last_offset = self._parse_time(last.get('offset', '0s'))
last_dur = self._parse_time(last.get('duration', '0s'))
return last_offset + last_dur, len(spine_children)
return TimeValue.zero(), len(spine_children)
if position.startswith('after:') or position.startswith('before:'):
is_after = position.startswith('after:')
ref_id = position.split(':', 1)[1]
ref_clip = self.clips.get(ref_id)
if ref_clip is None or ref_clip not in spine_children:
raise ValueError(f"Reference clip not found: {ref_id}")
idx = spine_children.index(ref_clip)
ref_offset = self._parse_time(ref_clip.get('offset', '0s'))
if is_after:
ref_dur = self._parse_time(ref_clip.get('duration', '0s'))
return ref_offset + ref_dur, idx + 1
return ref_offset, idx
# Assume timecode
target_offset = self._parse_time(position)
insert_index = 0
for i, child in enumerate(spine_children):
child_offset = self._parse_time(child.get('offset', '0s'))
if child_offset >= target_offset:
insert_index = i
break
insert_index = i + 1
return target_offset, insert_index
def _make_transition_element(
self,
effect_name: str,
trans_offset: 'TimeValue',
trans_duration: 'TimeValue',
effect_ref_id: str | None,
) -> ET.Element:
"""Build a <transition> element with optional filter-video child."""
transition = ET.Element('transition')
transition.set('name', effect_name)
transition.set('offset', trans_offset.to_fcpxml())
transition.set('duration', trans_duration.to_fcpxml())
if effect_ref_id:
fv = ET.SubElement(transition, 'filter-video')
fv.set('ref', effect_ref_id)
fv.set('name', effect_name)
return transition
def save(self, output_path: Optional[str] = None) -> str:
"""Serialise the modified XML tree to disk.
When the destination ends in ``.fcpxmld`` a bundle directory is
created and the XML lands in ``Info.fcpxml`` inside it. If the
source was also a bundle, every sidecar file (object-tracking /
Cinematic-mode ``dataLocator`` payloads — anything that isn't
``Info.fcpxml``) is copied across so the round-trip is lossless.
Writing a bundle source to a flat ``.fcpxml`` destination drops
those sidecars by definition.
Args:
output_path: Destination ``.fcpxml`` file or ``.fcpxmld``
bundle path. Defaults to overwriting the original
file/bundle loaded in ``__init__``.
Returns:
The absolute path written to (the bundle path when writing
a bundle, not the inner ``Info.fcpxml``).
"""
if output_path is None:
out = self.bundle_dir if self.bundle_dir is not None else self.path
else:
out = Path(output_path)
# Every write path goes through here, so snapping here (rather than
# in each handler) guarantees ripple edits never leave a spine clip
# off the frame grid — see snap_spine_times_to_frames() docstring.
# No-op (each value already equals its own snapped form) on content
# that was already frame-aligned.
self.snap_spine_times_to_frames()
if out.suffix.lower() == '.fcpxmld':
out.mkdir(exist_ok=True)
if (
self.bundle_dir is not None
and self.bundle_dir.resolve() != out.resolve()
):
self._copy_bundle_sidecars(self.bundle_dir, out)
write_fcpxml(self.root, str(out / 'Info.fcpxml'), fps=self.fps)
return str(out)
return write_fcpxml(self.root, str(out), fps=self.fps)
@staticmethod
def _copy_bundle_sidecars(src_bundle: Path, dst_bundle: Path) -> None:
"""Copy every sidecar entry of *src_bundle* into *dst_bundle*.
Sidecars are all bundle members except ``Info.fcpxml`` itself —
e.g. the external data files that ``locator``/``dataLocator``
elements reference for object tracking and Cinematic mode.
"""
import shutil
for entry in src_bundle.iterdir():
if entry.name == 'Info.fcpxml':
continue
target = dst_bundle / entry.name
if entry.is_dir():
shutil.copytree(entry, target, dirs_exist_ok=True)
else:
shutil.copy2(entry, target)
+333
View File
@@ -0,0 +1,333 @@
"""Dividir, cortar faixas e apagar clipes.
Extraído de writer.py — ver fcpxml/writer/__init__.py para o conjunto.
"""
import copy
import xml.etree.ElementTree as ET
from typing import List, Tuple
from ..models import (
TimeValue,
)
class CutMixin:
"""Dividir, cortar faixas e apagar clipes."""
# SPLIT & DELETE OPERATIONS
# ========================================================================
@staticmethod
def _filter_children_for_segment(
clip: ET.Element,
seg_start: 'TimeValue',
seg_duration: 'TimeValue',
) -> None:
"""Remove markers/keywords/titles from *clip* that fall outside the segment range.
After ``split_clip`` deepcopy's the original clip into each segment, every
segment inherits all child elements. Markers whose ``start`` falls outside
``[seg_start, seg_start + seg_duration)`` are phantom duplicates and must be
removed. Keywords that partially overlap get their ``start``/``duration``
clamped to the segment boundaries.
A lane-nested ``<title>`` (a "text" voice action's on-screen callout,
or a caption from an earlier `generate_dynamic_subtitles` pass) is
the same kind of phantom duplicate, just keyed on ``offset`` instead
of ``start`` — its offset lives in the same source-media coordinate
space as a marker's ``start`` (see ``add_text_title``/``add_marker``,
both anchored at ``parent.start``). Left unfiltered, every further
cut (silence removal, filler removal) duplicates it into every
resulting piece, so the same word shows up several times across the
edited timeline instead of once where it was placed.
"""
seg_end = seg_start + seg_duration
to_remove = []
for child in clip:
tag = child.tag
if tag in ('marker', 'chapter-marker'):
child_start = TimeValue.from_timecode(child.get('start', '0s'))
if child_start < seg_start or child_start >= seg_end:
to_remove.append(child)
elif tag == 'title':
title_offset = TimeValue.from_timecode(child.get('offset', '0s'))
if title_offset < seg_start or title_offset >= seg_end:
to_remove.append(child)
elif tag == 'keyword':
kw_start = TimeValue.from_timecode(child.get('start', '0s'))
kw_dur = TimeValue.from_timecode(child.get('duration', '0s'))
kw_end = kw_start + kw_dur
# Completely outside segment → remove
if kw_end <= seg_start or kw_start >= seg_end:
to_remove.append(child)
else:
# Clamp keyword range to segment boundaries
clamped_start = max(kw_start, seg_start)
clamped_end = min(kw_end, seg_end)
child.set('start', clamped_start.to_fcpxml())
child.set('duration', (clamped_end - clamped_start).to_fcpxml())
for child in to_remove:
clip.remove(child)
def split_clip(
self,
clip_id: str,
split_points: List[str]
) -> List[ET.Element]:
"""
Split a clip at specified timecodes.
Args:
clip_id: Clip to split
split_points: Timecodes within the clip to split at
Returns:
List of resulting clip elements
"""
spine, clip, clip_index = self._require_spine_clip(clip_id)
# Get clip properties
clip_start, clip_duration, clip_offset = self._get_clip_times(clip)
clip_name = clip.get('name', 'Clip')
# Sort split points
split_times = sorted([self._parse_time(sp) for sp in split_points])
# Remove original clip
spine.remove(clip)
# Create new clips
new_clips = []
current_offset = clip_offset
current_start = clip_start
all_points = split_times + [clip_duration]
for i, split_time in enumerate(all_points):
if i == 0:
segment_duration = split_time
else:
segment_duration = split_time - split_times[i - 1]
if segment_duration <= TimeValue.zero():
continue
# Create new clip
new_clip = copy.deepcopy(clip)
new_clip.set('name', clip_name)
new_clip.set('offset', current_offset.to_fcpxml())
new_clip.set('start', current_start.to_fcpxml())
new_clip.set('duration', segment_duration.to_fcpxml())
# Remove markers/keywords that belong to other segments
self._filter_children_for_segment(
new_clip, current_start, segment_duration
)
self._reassign_text_style_ids(new_clip)
spine.insert(clip_index + len(new_clips), new_clip)
new_clips.append(new_clip)
# Update for next iteration
current_offset = current_offset + segment_duration
current_start = current_start + segment_duration
# Update clip index: remove stale original entry, add split entries
self.clips.pop(clip_id, None)
for i, new_clip in enumerate(new_clips):
new_id = f"{clip_id}_split_{i}"
self.clips[new_id] = new_clip
return new_clips
def cut_clip_ranges(
self,
clip: ET.Element,
cut_ranges: List[Tuple['TimeValue', 'TimeValue']],
) -> 'TimeValue':
"""Remove clip-relative time ranges from a spine clip, rippling after.
Element-based on purpose: callers that walk the spine (e.g. media
silence removal) pass the exact element, so duplicate-named clips are
never ambiguous the way name-keyed operations are.
Args:
clip: The spine clip element to cut (must be a direct spine child).
cut_ranges: (start, end) TimeValue pairs measured from the clip's
own head. Overlapping/unsorted ranges are merged; portions
outside [0, clip duration] are clamped. A cut covering the
whole clip removes it entirely.
Returns:
Total removed duration (zero if no effective ranges).
"""
spine = self._get_spine()
clip_start, clip_duration, clip_offset = self._get_clip_times(clip)
clip_index = list(spine).index(clip)
zero = TimeValue.zero()
# Clamp, sort, merge.
clamped = []
for start, end in cut_ranges:
start = start if start > zero else zero
end = end if end < clip_duration else clip_duration
if end > start:
clamped.append((start, end))
clamped.sort(key=lambda r: r[0])
merged: List[Tuple[TimeValue, TimeValue]] = []
for start, end in clamped:
if merged and start <= merged[-1][1]:
if end > merged[-1][1]:
merged[-1] = (merged[-1][0], end)
else:
merged.append((start, end))
if not merged:
return zero
# Keep ranges = complement of the merged cuts.
keeps: List[Tuple[TimeValue, TimeValue]] = []
cursor = zero
for start, end in merged:
if start > cursor:
keeps.append((cursor, start))
cursor = end
if cursor < clip_duration:
keeps.append((cursor, clip_duration))
# A keep segment shorter than a couple frames at the very start or
# end of the clip is just leftover cut padding with no neighboring
# kept audio on its outer side (the silence butts against the clip's
# own edge) — not a real clip. Rather than emit it as its own
# near-invisible micro-clip, fold it into the adjacent real segment,
# which simply starts earlier / ends later to absorb it.
min_keep_seconds = 2 * float(self.frame_duration_fraction())
if len(keeps) > 1:
first_start, first_end = keeps[0]
if (first_end - first_start).to_seconds() < min_keep_seconds:
keeps[1] = (first_start, keeps[1][1])
keeps.pop(0)
if len(keeps) > 1:
last_start, last_end = keeps[-1]
if (last_end - last_start).to_seconds() < min_keep_seconds:
keeps[-2] = (keeps[-2][0], last_end)
keeps.pop()
spine.remove(clip)
new_clips: List[ET.Element] = []
current_offset = clip_offset
kept_total = zero
for keep_start, keep_end in keeps:
seg_duration = keep_end - keep_start
seg_start = clip_start + keep_start
new_clip = copy.deepcopy(clip)
new_clip.set('offset', current_offset.to_fcpxml())
new_clip.set('start', seg_start.to_fcpxml())
new_clip.set('duration', seg_duration.to_fcpxml())
self._filter_children_for_segment(new_clip, seg_start, seg_duration)
self._reassign_text_style_ids(new_clip)
spine.insert(clip_index + len(new_clips), new_clip)
new_clips.append(new_clip)
current_offset = current_offset + seg_duration
kept_total = kept_total + seg_duration
removed = clip_duration - kept_total
self._ripple_from_index(spine, clip_index + len(new_clips), zero - removed)
self._update_sequence_duration()
# Keep the name index coherent, mirroring delete_clip/split_clip.
name = clip.get('id') or clip.get('name') or ''
if name and self.clips.get(name) is clip:
if new_clips:
self.clips[name] = new_clips[0]
else:
remaining = [
sc for _, sc in self._iter_spine_clips()
if (sc.get('id') or sc.get('name') or '') == name
]
if remaining:
self.clips[name] = remaining[0]
else:
self.clips.pop(name, None)
return removed
def remove_trailing_gaps(self) -> None:
"""Remove empty ``<gap>`` elements at the end of the timeline.
Silence removal (and FCP round-trips) can leave a trailing gap holding
the timeline open past the last real clip. This removes only *trailing*
gaps — a gap in the middle is left untouched — and re-syncs the sequence
duration so the exported file ends where the content ends.
"""
spine = self._get_spine()
children = list(spine)
if not children:
return
last = children[-1]
if last.tag != 'gap':
return
spine.remove(last)
self._update_sequence_duration()
def delete_clip(
self,
clip_ids: List[str],
ripple: bool = True
) -> None:
"""
Delete clips from timeline.
Uses spine iteration instead of the name-indexed dict so that
duplicate-named clips (e.g. four ``Interview_A``) are resolved
correctly — always targeting the *first* spine match rather than
the last-indexed entry.
Args:
clip_ids: Clips to delete
ripple: If True, shift subsequent clips. If False, leave gaps.
"""
spine = self._get_spine()
for clip_id in clip_ids:
# Walk spine directly to find the first clip matching this name,
# avoiding the last-one-wins problem in self.clips.
target = None
for _spine_idx, spine_clip in self._iter_spine_clips():
name = spine_clip.get('id') or spine_clip.get('name') or ''
if name == clip_id:
target = spine_clip
break
if target is None:
continue
_, clip_duration, clip_offset = self._get_clip_times(target)
clip_index = list(spine).index(target)
if ripple:
spine.remove(target)
self._ripple_from_index(
spine, clip_index, TimeValue.zero() - clip_duration
)
else:
# Replace with gap
gap = ET.Element('gap')
gap.set('name', 'Gap')
gap.set('offset', clip_offset.to_fcpxml())
gap.set('duration', clip_duration.to_fcpxml())
spine.remove(target)
spine.insert(clip_index, gap)
# Re-index: if other spine clips share this name, point the
# dict entry at the next one; otherwise remove entirely.
remaining = [
sc for _, sc in self._iter_spine_clips()
if (sc.get('id') or sc.get('name') or '') == clip_id
]
if remaining:
self.clips[clip_id] = remaining[0]
else:
self.clips.pop(clip_id, None)
# ========================================================================
+170
View File
@@ -0,0 +1,170 @@
"""Escrita do documento FCPXML: assets de vídeo, timebases e serialização.
Extraído de writer.py — ver fcpxml/writer/__init__.py para o conjunto.
"""
import logging
import subprocess
import xml.etree.ElementTree as ET
from pathlib import Path
from typing import Optional
from ..models import (
TimeValue,
)
from .validation import validate_fcpxml
_log = logging.getLogger(__name__)
# ============================================================================
# STILL IMAGE AUTO-CONVERSION (v0.6.0)
# ============================================================================
_STILL_IMAGE_EXTENSIONS = {'.png', '.jpg', '.jpeg', '.tiff', '.tif', '.bmp'}
def _ensure_video_asset(
src_path: str,
duration: float = 10.0,
fps: int = 24,
width: int = 1920,
height: int = 1080,
) -> str:
"""Convert a still image to a video file if needed.
Detects still images by extension and converts them to MOV using ffmpeg.
Video files are returned as-is.
Args:
src_path: Path to the source media file.
duration: Duration in seconds for the still-to-video conversion.
fps: Frame rate for the output video.
width: Output width (even number).
height: Output height (even number).
Returns:
Path to the video file (original path if already video, new .mov path
if converted from still).
Raises:
FileNotFoundError: If ffmpeg is not installed.
"""
# Validate numeric parameters to prevent ffmpeg abuse / resource exhaustion.
if not isinstance(duration, (int, float)) or duration <= 0 or duration > 3600:
raise ValueError(f"duration must be 0 < d <= 3600, got {duration!r}")
if not isinstance(fps, int) or fps < 1 or fps > 240:
raise ValueError(f"fps must be 1–240, got {fps!r}")
if not isinstance(width, int) or width < 2 or width > 7680 or width % 2:
raise ValueError(f"width must be even, 2–7680, got {width!r}")
if not isinstance(height, int) or height < 2 or height > 4320 or height % 2:
raise ValueError(f"height must be even, 2–4320, got {height!r}")
path = Path(src_path)
if path.suffix.lower() not in _STILL_IMAGE_EXTENSIONS:
return src_path
output_path = path.with_suffix('.mov')
if output_path.exists():
return str(output_path)
# Build ffmpeg command: still image → video with specified duration
cmd = [
'ffmpeg', '-y',
'-loop', '1',
'-i', str(path),
'-c:v', 'prores_ks',
'-profile:v', '0',
'-t', str(duration),
'-r', str(fps),
'-vf', f'scale={width}:{height}:force_original_aspect_ratio=decrease,'
f'pad={width}:{height}:(ow-iw)/2:(oh-ih)/2',
'-pix_fmt', 'yuva444p10le',
str(output_path),
]
try:
subprocess.run(cmd, check=True, capture_output=True, timeout=120)
except FileNotFoundError:
raise FileNotFoundError(
"ffmpeg not found. Install ffmpeg to use still image auto-conversion: "
"brew install ffmpeg"
)
except subprocess.TimeoutExpired:
raise RuntimeError(
f"Image conversion timed out after 120s: {path}"
)
except subprocess.CalledProcessError as e:
stderr_msg = e.stderr.decode(errors='replace') if e.stderr else str(e)
raise RuntimeError(f"ffmpeg conversion failed: {stderr_msg}")
return str(output_path)
def _enforce_standard_timebases(root: ET.Element) -> None:
"""Walk all elements and snap time attributes to standard FCPXML timebases.
Targets offset, start, duration, and tcStart attributes. Values that
already use a standard denominator are left untouched.
"""
time_attrs = ('offset', 'start', 'duration', 'tcStart')
for elem in root.iter():
for attr in time_attrs:
val = elem.get(attr)
if val and val.endswith('s') and '/' in val:
try:
tv = TimeValue.from_timecode(val)
if not tv.is_standard_timebase():
# Snap to nearest frame at 2400 ticks/sec
snapped = tv.snap_to_frame(24)
elem.set(attr, snapped.to_fcpxml())
except (ValueError, ZeroDivisionError):
pass # Skip unparseable values
def write_fcpxml(
root: ET.Element,
filepath: str,
enforce_timebases: bool = False,
strict: bool = False,
fps: Optional[float] = None,
) -> str:
"""Format an ElementTree root as pretty-printed FCPXML and write to disk.
Handles XML declaration, DOCTYPE insertion, and blank-line cleanup
consistently across all FCPXML output paths (modifier, writer, rough cut).
Args:
root: The <fcpxml> root Element to serialize.
filepath: Destination file path.
enforce_timebases: If True, snap all time values to standard FCPXML
timebases before writing. Default False for backward compat.
strict: If True, raise ValueError on validation errors.
If False (default), log warnings.
fps: Frame rate for the frame-alignment validation check. Defaults
to 24 when omitted — pass the sequence's real (float) rate so
NTSC projects (23.976/29.97/59.94fps) don't get spurious
"not frame-aligned at 24fps" warnings for values that are
exactly aligned at their own true rate.
Returns:
The filepath written to.
"""
if enforce_timebases:
_enforce_standard_timebases(root)
# Auto-validate before writing
issues = validate_fcpxml(root, fps=fps if fps is not None else 24.0)
if issues:
errors = [i for i in issues if i.severity == "error"]
warnings = [i for i in issues if i.severity == "warning"]
for w in warnings:
_log.warning("FCPXML validation: %s", w.message)
if errors and strict:
msg = "; ".join(e.message for e in errors)
raise ValueError(f"FCPXML validation failed: {msg}")
for e in errors:
_log.error("FCPXML validation: %s", e.message)
from ..safe_xml import serialize_xml
return serialize_xml(root, filepath, doctype='<!DOCTYPE fcpxml>')
+147
View File
@@ -0,0 +1,147 @@
"""FCPXMLWriter: gera um documento novo a partir de objetos Python.
Extraído de writer.py — ver fcpxml/writer/__init__.py para o conjunto.
"""
import uuid
import xml.etree.ElementTree as ET
from datetime import datetime
from ..models import (
Marker,
Project,
Timecode,
)
from .document import write_fcpxml
from .helpers import build_marker_element
# ============================================================================
# FCPXML GENERATOR - Create from Python objects
# ============================================================================
class FCPXMLWriter:
"""Generate a new FCPXML document from Python dataclass objects.
Converts a ``Project`` (containing ``Timeline`` → ``Clip`` → ``Marker``
hierarchies) into a spec-compliant FCPXML v1.11 element tree and writes
it to disk. Used by ``RoughCutGenerator`` and the ``generate_*`` MCP
tools to create fresh timelines from scratch.
Unlike ``FCPXMLModifier`` (which mutates existing XML), this class
*creates* XML from structured Python objects.
Example::
from fcpxml.models import Project, Timeline, Clip, Timecode
project = Project(name="My Edit", timelines=[...])
writer = FCPXMLWriter()
writer.write_project(project, "output.fcpxml")
"""
def __init__(self, version: str = "1.13"):
"""Initialize writer targeting the given FCPXML version."""
self.version = version
self.resource_counter = 1
def _next_resource_id(self) -> str:
"""Return an auto-incrementing resource ID (r1, r2, ...)."""
rid = f"r{self.resource_counter}"
self.resource_counter += 1
return rid
def _generate_uid(self) -> str:
"""Generate a unique identifier for FCPXML elements."""
return str(uuid.uuid4()).upper()
def _tc_to_rational(self, tc: Timecode) -> str:
"""Convert a Timecode to FCPXML rational time string (e.g. '48/24s')."""
return f"{tc.frames}/{int(tc.frame_rate)}s"
def write_project(self, project: Project, filepath: str):
"""Write a project to an FCPXML file."""
root = self._build_fcpxml(project)
write_fcpxml(root, filepath)
def _build_fcpxml(self, project: Project) -> ET.Element:
"""Build the full FCPXML element tree: fcpxml > resources + library > event > project."""
root = ET.Element('fcpxml', version=self.version)
resources = ET.SubElement(root, 'resources')
resource_map = {}
if project.timelines:
timeline = project.timelines[0]
format_id = self._next_resource_id()
ET.SubElement(resources, 'format',
id=format_id,
name=f"FFVideoFormat{timeline.height}p{int(timeline.frame_rate)}",
frameDuration=f"1/{int(timeline.frame_rate)}s",
width=str(timeline.width), height=str(timeline.height))
resource_map['_format'] = format_id
library = ET.SubElement(root, 'library',
location=f"file:///Users/editor/Movies/{project.name}.fcpbundle/")
event = ET.SubElement(library, 'event', name=project.name, uid=self._generate_uid())
for timeline in project.timelines:
self._add_timeline(event, timeline, resources, resource_map)
return root
def _add_timeline(self, event, timeline, resources, resource_map):
"""Add a timeline as a project > sequence > spine structure under the event."""
project_elem = ET.SubElement(event, 'project',
name=timeline.name, uid=self._generate_uid(),
modDate=datetime.now().strftime("%Y-%m-%d %H:%M:%S -0500"))
format_id = resource_map.get('_format', 'r1')
sequence = ET.SubElement(project_elem, 'sequence',
format=format_id, duration=self._tc_to_rational(timeline.duration),
tcStart="0s", tcFormat="NDF", audioLayout="stereo", audioRate="48k")
spine = ET.SubElement(sequence, 'spine')
for clip in timeline.clips:
self._add_clip(spine, clip, resources, resource_map)
for marker in timeline.markers:
self._add_marker(sequence, marker)
def _add_clip(self, spine, clip, resources, resource_map):
"""Add a clip as an asset-clip element, creating its asset resource if needed."""
if clip.media_path and clip.media_path not in resource_map:
asset_id = self._next_resource_id()
ET.SubElement(resources, 'asset', id=asset_id, name=clip.name,
uid=self._generate_uid(), src=clip.media_path, start="0s",
duration=self._tc_to_rational(clip.duration), hasVideo="1", hasAudio="1")
resource_map[clip.media_path] = asset_id
asset_id = resource_map.get(clip.media_path, 'r1')
format_id = resource_map.get('_format', 'r1')
clip_elem = ET.SubElement(spine, 'asset-clip',
ref=asset_id, offset=self._tc_to_rational(clip.start), name=clip.name,
start=self._tc_to_rational(clip.source_start) if clip.source_start else "0s",
duration=self._tc_to_rational(clip.duration), format=format_id, tcFormat="NDF")
for marker in clip.markers:
self._add_marker(clip_elem, marker)
for keyword in clip.keywords:
self._add_keyword(clip_elem, keyword)
def _add_marker(self, parent: ET.Element, marker: Marker):
"""Add a marker or chapter-marker element to a parent clip or sequence."""
build_marker_element(
parent=parent,
marker_type=marker.marker_type,
start=self._tc_to_rational(marker.start),
duration=self._tc_to_rational(marker.duration) if marker.duration else "1/24s",
name=marker.name,
note=marker.note or None,
)
def _add_keyword(self, parent, keyword):
"""Add a keyword element with optional start/duration range to a parent clip."""
attrs = {'value': keyword.value}
if keyword.start:
attrs['start'] = self._tc_to_rational(keyword.start)
if keyword.duration:
attrs['duration'] = self._tc_to_rational(keyword.duration)
ET.SubElement(parent, 'keyword', **attrs)
+279
View File
@@ -0,0 +1,279 @@
"""Ajudantes de nível de módulo do writer: sanitização, escalas, elementos base.
Extraído de writer.py — ver fcpxml/writer/__init__.py para o conjunto.
"""
import subprocess
import xml.etree.ElementTree as ET
from pathlib import Path
from typing import Any, Dict, List, Optional
from ..models import (
MarkerType,
)
# Maximum lengths for XML attribute values to prevent memory abuse
_MAX_MARKER_NAME_LENGTH = 1024
_MAX_NOTE_LENGTH = 4096
# ============================================================================
# EFFECT RESOURCE REGISTRY (v0.6.0)
# ============================================================================
# FCP built-in transition/filter effect UUIDs extracted from Filters.bundle.
# Maps slug → (display_name, uuid).
FCP_EFFECTS: Dict[str, tuple] = {
# Dissolves
'cross-dissolve': ('Cross Dissolve', '4731E73A-8DAC-4113-9A30-AE85B1761265'),
'fade': ('Fade', '8154D0DA-C99B-4EF8-8FF8-006FE5ED57F1'),
'dip-to-color': ('Dip to Color', 'F779C565-486D-4633-8035-0374B4DB8F5C'),
'noise-dissolve': ('Noise Dissolve', 'ABFED81E-35D9-429C-AB47-438C1FB5D9DE'),
# Wipes
'edge-wipe': ('Edge Wipe', '857E2FBA-98DB-411B-A88C-CE6ABC1F65D8'),
'slide': ('Slide', '6AAB0D54-FCD8-4EBD-A62D-D352A5ED1648'),
'band-wipe': ('Band Wipe', 'A4E0B8E4-E916-474B-A14C-E3A9E0B1A3C1'),
'center-wipe': ('Center Wipe', 'B3F2D4A1-7C8E-4B9D-A5F6-D1E2C3B4A5D6'),
'checker-wipe': ('Checker Wipe', 'C4D3E2F1-8A7B-4C6D-B5E4-F2A1D3C4B5E6'),
'clock-wipe': ('Clock Wipe', 'D5E4F3A2-9B8C-4D7E-C6F5-A3B2E4D5C6F7'),
'gradient-wipe': ('Gradient Wipe', 'E6F5A4B3-AC9D-4E8F-D7A6-B4C3F5E6D7A8'),
'inset-wipe': ('Inset Wipe', 'F7A6B5C4-BD0E-4F9A-E8B7-C5D4A6F7E8B9'),
'star-wipe': ('Star Wipe', 'A8B7C6D5-CE1F-4A0B-F9C8-D6E5B7A8F9C0'),
# Legacy aliases — map common shorthand to canonical slugs
'fade-to-black': ('Fade', '8154D0DA-C99B-4EF8-8FF8-006FE5ED57F1'),
'fade-from-black': ('Fade', '8154D0DA-C99B-4EF8-8FF8-006FE5ED57F1'),
'wipe': ('Edge Wipe', '857E2FBA-98DB-411B-A88C-CE6ABC1F65D8'),
'dissolve': ('Cross Dissolve', '4731E73A-8DAC-4113-9A30-AE85B1761265'),
}
def list_effects() -> List[Dict[str, str]]:
"""Return a list of all available FCP transition effects.
Each entry contains slug, display_name, and uuid.
Legacy aliases are excluded to avoid duplicates.
"""
seen_uuids: set = set()
effects = []
for slug, (name, uid) in FCP_EFFECTS.items():
if uid in seen_uuids:
continue
seen_uuids.add(uid)
effects.append({'slug': slug, 'name': name, 'uuid': uid})
return effects
# Named constants for clip-tag sets used across operations.
# Using named tuples prevents inconsistent ad-hoc tag lists and ensures
# new clip types only need adding in one place.
CLIP_TAGS = ('clip', 'asset-clip', 'video', 'ref-clip')
CLIP_AND_AUDIO_TAGS = ('clip', 'asset-clip', 'video', 'audio', 'ref-clip')
SPINE_ELEMENT_TAGS = ('clip', 'asset-clip', 'video', 'audio', 'gap', 'transition', 'ref-clip')
def _sanitize_xml_value(value: str, max_length: int = _MAX_MARKER_NAME_LENGTH) -> str:
"""Sanitize a string value before writing it into an XML attribute.
Strips null bytes, control characters (except tab/newline/CR), and
enforces a length limit to prevent memory abuse or malformed XML.
"""
if not isinstance(value, str):
return str(value)
# Remove null bytes and non-printable control characters
cleaned = ''.join(
c for c in value
if c in ('\t', '\n', '\r') or ord(c) >= 32
)
if len(cleaned) > max_length:
cleaned = cleaned[:max_length]
return cleaned
# FCPXML DTD child element ordering for asset-clip / clip elements.
# Elements MUST appear in this order for DTD validation.
# See: https://developer.apple.com/documentation/professional-video-applications/fcpxml-reference
_ASSET_CLIP_CHILD_ORDER = [
'note',
'conform-rate', 'timeMap',
'adjust-crop', 'adjust-corners', 'adjust-conform', 'adjust-transform',
'adjust-blend', 'adjust-stabilization', 'adjust-rollingShutter',
'adjust-360-transform', 'adjust-reorient', 'adjust-orientation',
'adjust-volume', 'adjust-panner',
# anchor items (connected clips, titles, etc.)
'audio', 'video', 'clip', 'title', 'caption',
'mc-clip', 'ref-clip', 'sync-clip', 'asset-clip', 'audition', 'spine',
# marker items
'marker', 'chapter-marker', 'rating', 'keyword', 'analysis-marker',
# trailing
'audio-channel-source',
'filter-video', 'filter-video-mask',
'filter-audio',
'metadata',
]
# Build a priority lookup: tag → index for fast comparison
_CHILD_ORDER_INDEX = {tag: i for i, tag in enumerate(_ASSET_CLIP_CHILD_ORDER)}
# How close to the end of a clip a zoom must finish for the return to be
# skipped. Within this margin the cut arrives before the eye registers the
# move back, so the return reads as a twitch rather than a resolution.
HOLD_AT_CUT_THRESHOLD = 1.0
# How close to the start of a clip a zoom must begin for the ramp-in to be
# skipped and the shot to simply open already zoomed. Tighter than the end
# margin on purpose: at the end the cut hides an unfinished return, but at
# the start a ramp is visible from frame one and reads as the shot settling.
START_AT_CUT_THRESHOLD = 0.5
def _fmt_scale(value: float) -> str:
"""Format a scale factor without trailing float noise (1.0 -> "1")."""
return f"{value:.6f}".rstrip("0").rstrip(".") or "0"
def _dtd_insert(parent: ET.Element, child: ET.Element) -> ET.Element:
"""Insert a child element into parent at the correct DTD-ordered position.
Instead of blindly appending (which can violate DTD ordering),
this finds the right insertion point based on the FCPXML DTD's
required element sequence for asset-clip / clip elements.
Unknown tags are appended at the end.
"""
child_priority = _CHILD_ORDER_INDEX.get(child.tag, len(_ASSET_CLIP_CHILD_ORDER))
# Find the first existing child whose priority is greater than ours
insert_idx = len(parent)
for i, existing in enumerate(parent):
existing_priority = _CHILD_ORDER_INDEX.get(existing.tag, len(_ASSET_CLIP_CHILD_ORDER))
if existing_priority > child_priority:
insert_idx = i
break
parent.insert(insert_idx, child)
return child
def build_marker_element(
parent: ET.Element,
marker_type: MarkerType,
start: str,
duration: str,
name: str,
note: Optional[str] = None,
) -> ET.Element:
"""Create a marker or chapter-marker XML element under *parent*.
Single source of truth for marker element construction — used by both
FCPXMLModifier (edit-existing workflow) and FCPXMLWriter (generate-new
workflow). Centralises tag selection, type-specific attributes, note
guards, and input sanitization so changes only need to happen once.
"""
elem = ET.Element(marker_type.xml_tag)
elem.set('start', start)
elem.set('duration', duration)
elem.set('value', _sanitize_xml_value(name, _MAX_MARKER_NAME_LENGTH))
for attr, val in marker_type.xml_attrs.items():
elem.set(attr, val)
if note and marker_type != MarkerType.CHAPTER:
elem.set('note', _sanitize_xml_value(note, _MAX_NOTE_LENGTH))
_dtd_insert(parent, elem)
return elem
def _create_asset_element(
resources: ET.Element,
asset_id: str,
name: str,
src: str,
duration: str = "0s",
start: str = "0s",
has_video: str = "1",
has_audio: str = "1",
uid: Optional[str] = None,
) -> ET.Element:
"""Create an <asset> element with <media-rep> child instead of src attribute.
FCP's DTD prefers <media-rep kind="original-media" src="..."/> children
over the src attribute on <asset>. This helper produces the preferred form.
Args:
resources: Parent <resources> element to append to.
asset_id: Resource ID (e.g. "r3").
name: Human-readable asset name.
src: File path or URL for the media source.
duration: Asset duration in FCPXML rational format.
start: Asset start time.
has_video: "1" if asset has video track.
has_audio: "1" if asset has audio track.
uid: Optional UUID; auto-generated if not provided.
Returns:
The created <asset> Element.
"""
import uuid as _uuid
asset = ET.SubElement(resources, 'asset')
asset.set('id', asset_id)
asset.set('name', _sanitize_xml_value(name, 512))
asset.set('uid', uid or str(_uuid.uuid4()).upper())
asset.set('start', start)
asset.set('duration', duration)
asset.set('hasVideo', has_video)
asset.set('hasAudio', has_audio)
# Use media-rep child instead of src attribute
media_rep = ET.SubElement(asset, 'media-rep')
media_rep.set('kind', 'original-media')
media_rep.set('src', src)
return asset
def _probe_audio_info(src: str) -> Optional[Dict[str, Any]]:
"""Probe an audio file for its real duration, sample rate, and channels.
Tries ffprobe first, then falls back to the stdlib ``wave`` module for
.wav files. Returns ``None`` when the file can't be probed, so callers
can fall back to caller-supplied durations.
Returns:
``{'duration': float, 'sample_rate': int, 'channels': int}`` or None.
"""
path = Path(src)
if not path.is_file():
return None
try:
result = subprocess.run(
['ffprobe', '-v', 'error', '-select_streams', 'a:0',
'-show_entries', 'stream=sample_rate,channels,duration',
'-show_entries', 'format=duration',
'-of', 'json', str(path)],
capture_output=True, text=True, timeout=15,
)
if result.returncode == 0:
import json
data = json.loads(result.stdout)
streams = data.get('streams') or [{}]
stream = streams[0]
duration = stream.get('duration') or data.get('format', {}).get('duration')
if duration:
return {
'duration': float(duration),
'sample_rate': int(stream.get('sample_rate') or 48000),
'channels': int(stream.get('channels') or 2),
}
except (OSError, subprocess.TimeoutExpired, ValueError):
pass
if path.suffix.lower() == '.wav':
try:
import wave
with wave.open(str(path), 'rb') as wf:
rate = wf.getframerate()
if rate > 0:
return {
'duration': wf.getnframes() / rate,
'sample_rate': rate,
'channels': wf.getnchannels(),
}
except (OSError, wave.Error, EOFError):
pass
return None
+78
View File
@@ -0,0 +1,78 @@
"""Inserir clipes na spine.
Extraído de writer.py — ver fcpxml/writer/__init__.py para o conjunto.
"""
import xml.etree.ElementTree as ET
from typing import Optional
class InsertMixin:
"""Inserir clipes na spine."""
# INSERT CLIP OPERATIONS
# ========================================================================
def insert_clip(
self,
position: str,
asset_id: Optional[str] = None,
asset_name: Optional[str] = None,
duration: Optional[str] = None,
in_point: Optional[str] = None,
out_point: Optional[str] = None,
ripple: bool = True
) -> ET.Element:
"""
Insert a library clip onto the timeline.
Args:
position: Where to insert - 'start', 'end', timecode, or 'after:clip_id'
asset_id: Asset reference ID (e.g., 'r3')
asset_name: Asset name (alternative to asset_id)
duration: Duration of clip (if not using in/out points)
in_point: Source in-point for subclip
out_point: Source out-point for subclip
ripple: Whether to shift subsequent clips
Returns:
The created clip element
"""
asset, asset_id = self._resolve_asset(asset_id, asset_name)
clip_duration, source_start = self._resolve_clip_duration(
asset, duration, in_point, out_point
)
# Get spine and calculate insert position
spine = self._get_spine()
spine_children = list(spine)
target_offset, insert_index = self._resolve_insert_position(
position, spine_children
)
# Build extra attrs — include format from first available format
extra: dict[str, str] = {}
for fmt_id in self.formats:
extra['format'] = fmt_id
break
new_clip = self._make_asset_clip(
asset_id, asset.get('name', 'Untitled'),
target_offset, source_start, clip_duration,
**extra,
)
# Insert into spine
spine.insert(insert_index, new_clip)
# Ripple subsequent clips if needed
if ripple and insert_index < len(spine_children):
self._ripple_from_index(spine, insert_index + 1, clip_duration)
# Add to clip index
clip_id = f"inserted_{len(self.clips)}"
self.clips[clip_id] = new_clip
return new_clip
# ========================================================================
+165
View File
@@ -0,0 +1,165 @@
"""Marcadores: um, por timecode, e em lote.
Extraído de writer.py — ver fcpxml/writer/__init__.py para o conjunto.
"""
import xml.etree.ElementTree as ET
from typing import Any, Dict, List, Optional
from ..models import (
MarkerColor,
MarkerType,
TimeValue,
)
from .helpers import build_marker_element
class MarkersMixin:
"""Marcadores: um, por timecode, e em lote."""
# ========================================================================
# MARKER OPERATIONS
# ========================================================================
def add_marker(
self,
clip_id: 'str | ET.Element',
timecode: str,
name: str,
marker_type: "MarkerType | str" = MarkerType.STANDARD,
color: Optional[MarkerColor] = None,
note: Optional[str] = None
) -> ET.Element:
"""
Add a marker to a clip.
Args:
clip_id: Target clip identifier (name or ID)
timecode: Position within clip (relative to clip start)
name: Marker label
marker_type: STANDARD, TODO, COMPLETED, or CHAPTER (enum or string)
color: Optional marker color
note: Optional marker note
Returns:
The created marker element
"""
clip = self._require_clip(clip_id)
if isinstance(marker_type, str):
marker_type = MarkerType.from_string(marker_type)
time_value = self._parse_time(timecode)
return build_marker_element(
parent=clip,
marker_type=marker_type,
start=time_value.to_fcpxml(),
duration=f"1/{int(self.fps)}s",
name=name,
note=note,
)
def add_marker_at_timeline(
self,
timecode: str,
name: str,
marker_type: "MarkerType | str" = MarkerType.STANDARD,
color: Optional[MarkerColor] = None,
note: Optional[str] = None
) -> ET.Element:
"""Add a marker at a timeline position (finds the containing clip).
Uses ``_find_spine_clip_at_seconds`` to walk the spine directly,
avoiding the name-indexed ``self.clips`` dict which silently drops
duplicate-named clips.
"""
if isinstance(marker_type, str):
marker_type = MarkerType.from_string(marker_type)
time_value = self._parse_time(timecode)
target_seconds = time_value.to_seconds()
clip, relative_seconds = self._find_spine_clip_at_seconds(target_seconds)
relative_tc = TimeValue.from_seconds(relative_seconds, self.fps)
return build_marker_element(
parent=clip,
marker_type=marker_type,
start=relative_tc.to_fcpxml(),
duration=f"1/{int(self.fps)}s",
name=name,
note=note,
)
def batch_add_markers(
self,
markers: List[Dict[str, Any]],
auto_at_cuts: bool = False,
auto_at_intervals: Optional[str] = None
) -> List[ET.Element]:
"""
Add multiple markers at once.
Args:
markers: List of marker specs [{timecode, name, marker_type, color}]
auto_at_cuts: Add marker at every cut point
auto_at_intervals: Add markers at regular intervals (e.g., "00:00:30:00")
Returns:
List of created marker elements
"""
created = []
# Handle explicit markers
for m in markers:
marker = self.add_marker_at_timeline(
timecode=m['timecode'],
name=m['name'],
marker_type=MarkerType.from_string(m.get('marker_type', 'standard')),
color=MarkerColor[m['color'].upper()] if m.get('color') else None,
note=m.get('note')
)
created.append(marker)
# Auto-detect at cuts — add a marker at the start of every spine clip.
if auto_at_cuts:
for i, clip in self._iter_spine_clips():
clip_start = clip.get('start', '0s')
marker = build_marker_element(
parent=clip,
marker_type=MarkerType.STANDARD,
start=clip_start,
duration=f"1/{int(self.fps)}s",
name=f"Cut {i+1}",
)
created.append(marker)
# Auto-detect at intervals — place markers at regular time steps.
if auto_at_intervals:
interval = self._parse_time(auto_at_intervals).to_seconds()
total_duration = self._timeline_duration().to_seconds()
if total_duration > 0:
current = interval
count = 1
while current < total_duration:
try:
clip, relative = self._find_spine_clip_at_seconds(current)
except ValueError:
current += interval
count += 1
continue
rel_tv = TimeValue.from_seconds(relative, self.fps)
marker = build_marker_element(
parent=clip,
marker_type=MarkerType.STANDARD,
start=rel_tv.to_fcpxml(),
duration=f"1/{int(self.fps)}s",
name=f"Marker {count}",
)
created.append(marker)
current += interval
count += 1
return created
+64
View File
@@ -0,0 +1,64 @@
"""FCPXMLModifier — a edição de FCPXML montada a partir de um mixin por assunto.
A classe era um bloco de 3.300 linhas com dezoito assuntos dentro. Ela continua
sendo uma classe só para quem chama — `modifier.add_marker(...)` não mudou — mas
cada assunto agora mora no seu próprio arquivo e pode ser lido inteiro sem rolar
por marcadores, velocidade e legendas até achar o trecho procurado.
Mixins em vez de objetos separados por uma razão concreta: todas essas operações
mexem no *mesmo* documento e dependem dos mesmos índices e da mesma navegação na
spine (`_require_clip`, `_iter_spine_clips`, `_ripple_after_clip`). Separá-las em
objetos independentes obrigaria cada um a carregar uma referência de volta ao
documento e transformaria toda chamada interna em travessia de fronteira, sem
nada em troca — a divisão que importa aqui é de *leitura*, não de estado.
A ordem abaixo é irrelevante para o comportamento: nenhum mixin sobrescreve
método de outro; cada um contribui com um conjunto disjunto de operações.
"""
from .audio import AudioMixin
from .compound import CompoundMixin
from .connected import ConnectedMixin
from .core import ModifierCore
from .cut import CutMixin
from .insert import InsertMixin
from .markers import MarkersMixin
from .rapid import RapidMixin
from .reformat import ReformatMixin
from .relink import RelinkMixin
from .reorder import ReorderMixin
from .roles import RolesMixin
from .selection import SelectionMixin
from .silence import SilenceMixin
from .speed import SpeedMixin
from .titles import TitlesMixin
from .transitions import TransitionsMixin
from .trim import TrimMixin
class FCPXMLModifier(
RelinkMixin,
MarkersMixin,
TrimMixin,
ReorderMixin,
TransitionsMixin,
SpeedMixin,
CutMixin,
RapidMixin,
SelectionMixin,
InsertMixin,
ConnectedMixin,
TitlesMixin,
AudioMixin,
CompoundMixin,
RolesMixin,
ReformatMixin,
SilenceMixin,
ModifierCore,
):
"""Carrega um FCPXML, aplica edições cirúrgicas e salva.
Interface de escrita usada por todos os handlers do servidor MCP. A
documentação de cada operação está no mixin correspondente; o
carregamento, os índices e o `save` estão em `core.ModifierCore`.
"""
+240
View File
@@ -0,0 +1,240 @@
"""Corte rápido: flash frames, rapid trim, preencher buracos.
Extraído de writer.py — ver fcpxml/writer/__init__.py para o conjunto.
"""
from typing import Any, Dict, List, Optional
class RapidMixin:
"""Corte rápido: flash frames, rapid trim, preencher buracos."""
# SPEED CUTTING OPERATIONS (v0.3.0)
# ========================================================================
def fix_flash_frames(
self,
mode: str = 'auto',
threshold_frames: int = 6,
critical_threshold_frames: int = 2
) -> List[Dict[str, Any]]:
"""
Automatically fix flash frames (ultra-short clips).
Args:
mode: How to fix flash frames:
- 'extend_previous': Extend the previous clip to cover the flash frame
- 'extend_next': Extend the next clip backward to cover the flash frame
- 'delete': Remove the flash frame entirely (ripple)
- 'auto': Use smart logic (extend prev for critical, delete for warning)
threshold_frames: Frames below this are considered flash frames
critical_threshold_frames: Frames below this are critical (default: 2)
Returns:
List of fixed flash frames with details
"""
spine = self._get_spine()
fixed = []
# Collect flash frames first (can't modify while iterating)
flash_frames = []
for i, clip in self._iter_spine_clips():
duration = self._parse_time(clip.get('duration', '0s'))
duration_frames = duration.to_frames(self.fps)
if duration_frames < threshold_frames:
is_critical = duration_frames < critical_threshold_frames
flash_frames.append({
'index': i,
'clip': clip,
'clip_id': clip.get('name') or clip.get('id') or f"clip_{i}",
'duration_frames': duration_frames,
'is_critical': is_critical
})
# Process in reverse order to maintain indices
for ff in reversed(flash_frames):
clip = ff['clip']
_, _, clip_offset = self._get_clip_times(clip)
# Determine actual mode
actual_mode = mode
if mode == 'auto':
# Critical: try to extend previous, otherwise delete
# Warning: delete
actual_mode = 'extend_previous' if ff['is_critical'] else 'delete'
result = {
'clip_name': ff['clip_id'],
'duration_frames': ff['duration_frames'],
'was_critical': ff['is_critical'],
'action': actual_mode,
'timecode': clip_offset.to_timecode(self.fps)
}
direction = {'extend_previous': 'prev', 'extend_next': 'next'}.get(actual_mode)
if direction:
neighbor = self._absorb_into_neighbor(spine, clip, direction)
if neighbor is not None:
self._recalculate_offsets(spine)
result['extended_clip'] = neighbor.get('name', direction.title())
else:
spine.remove(clip)
self._recalculate_offsets(spine)
else: # delete
spine.remove(clip)
self._recalculate_offsets(spine)
fixed.append(result)
# Rebuild clip index
self._build_clip_index()
return fixed
def rapid_trim(
self,
max_duration: Optional[str] = None,
min_duration: Optional[str] = None,
keywords: Optional[List[str]] = None,
trim_from: str = 'end'
) -> List[Dict[str, Any]]:
"""
Batch trim clips to enforce duration limits.
Args:
max_duration: Maximum clip duration (e.g., '2s', '00:00:02:00')
min_duration: Minimum clip duration (clips shorter are extended/left alone)
keywords: Only trim clips with these keywords (None = all clips)
trim_from: Where to trim - 'start', 'end', or 'center'
Returns:
List of trimmed clips with before/after durations
"""
trimmed = []
max_dur = self._parse_time(max_duration) if max_duration else None
min_dur = self._parse_time(min_duration) if min_duration else None
for _i, clip in self._iter_spine_clips():
clip_name = clip.get('name') or clip.get('id') or 'Unknown'
# Check keyword filter
if keywords:
clip_keywords = set()
for kw_elem in clip.findall('keyword'):
clip_keywords.add(kw_elem.get('value', ''))
if not clip_keywords.intersection(set(keywords)):
continue
current_start, current_duration, _ = self._get_clip_times(clip)
original_duration = current_duration.to_seconds()
# Skip clips shorter than min_duration (leave them alone)
if min_dur and current_duration < min_dur:
continue
# Check max duration
if max_dur and current_duration > max_dur:
excess = current_duration - max_dur
if trim_from == 'end':
# Keep start, reduce duration
clip.set('duration', max_dur.to_fcpxml())
elif trim_from == 'start':
# Increase start, reduce duration
new_start = current_start + excess
clip.set('start', new_start.to_fcpxml())
clip.set('duration', max_dur.to_fcpxml())
elif trim_from == 'center':
# Trim equal amounts from both ends
half_excess = excess * 0.5
new_start = current_start + half_excess
clip.set('start', new_start.to_fcpxml())
clip.set('duration', max_dur.to_fcpxml())
trimmed.append({
'clip_name': clip_name,
'original_duration': original_duration,
'new_duration': max_dur.to_seconds(),
'trim_from': trim_from,
'action': 'trimmed'
})
# Recalculate offsets
self._recalculate_offsets(self._get_spine())
return trimmed
def fill_gaps(
self,
mode: str = 'extend_previous',
max_gap: Optional[str] = None
) -> List[Dict[str, Any]]:
"""
Fill gaps in the timeline.
Args:
mode: How to fill gaps:
- 'extend_previous': Extend previous clip to fill gap
- 'extend_next': Extend next clip backward to fill gap
- 'delete': Remove gap elements and ripple
max_gap: Only fill gaps smaller than this (None = all gaps)
Returns:
List of filled gaps with details
"""
spine = self._get_spine()
filled = []
max_gap_time = self._parse_time(max_gap) if max_gap else None
# Find all gaps
gaps_to_process = []
for i, child in enumerate(list(spine)):
if child.tag == 'gap':
gap_duration = self._parse_time(child.get('duration', '0s'))
gap_offset = self._parse_time(child.get('offset', '0s'))
# Check max_gap filter
if max_gap_time and gap_duration > max_gap_time:
continue
gaps_to_process.append({
'element': child,
'index': i,
'duration': gap_duration,
'offset': gap_offset
})
# Process in reverse to maintain indices
for gap_info in reversed(gaps_to_process):
gap = gap_info['element']
gap_duration = gap_info['duration']
gap_offset = gap_info['offset']
result = {
'timecode': gap_offset.to_timecode(self.fps),
'duration_frames': gap_duration.to_frames(self.fps),
'duration_seconds': gap_duration.to_seconds(),
'action': mode
}
direction = {'extend_previous': 'prev', 'extend_next': 'next'}.get(mode)
if direction:
neighbor = self._absorb_into_neighbor(spine, gap, direction)
if neighbor is not None:
result['extended_clip'] = neighbor.get('name', direction.title())
filled.append(result)
else: # delete
spine.remove(gap)
filled.append(result)
# Recalculate offsets
self._recalculate_offsets(spine)
return filled
# ========================================================================
+43
View File
@@ -0,0 +1,43 @@
"""Reenquadrar a resolução do projeto.
Extraído de writer.py — ver fcpxml/writer/__init__.py para o conjunto.
"""
class ReformatMixin:
"""Reenquadrar a resolução do projeto."""
# REFORMAT OPERATIONS (v0.5.0)
# ========================================================================
SOCIAL_FORMATS = {
"9:16": (1080, 1920),
"1:1": (1080, 1080),
"4:5": (1080, 1350),
"16:9": (1920, 1080),
"4:3": (1440, 1080),
}
def reformat_resolution(self, width: int, height: int) -> None:
"""Change the timeline format to a new resolution.
Updates the format resource dimensions. FCP handles spatial
conforming (letterbox/pillarbox) on import.
Args:
width: Target width in pixels
height: Target height in pixels
"""
for fmt in self.root.findall('.//format'):
fmt.set('width', str(width))
fmt.set('height', str(height))
old_name = fmt.get('name', '')
if old_name:
fmt.set('name', f"FFVideoFormat{width}x{height}")
sequence = self.root.find('.//sequence')
if sequence is not None and sequence.get('format'):
pass # format ref stays the same, dimensions updated in-place
# ========================================================================
+94
View File
@@ -0,0 +1,94 @@
"""Repontar a mídia de um projeto para novos arquivos.
Extraído de writer.py — ver fcpxml/writer/__init__.py para o conjunto.
"""
from pathlib import Path
from typing import Any, Dict
class RelinkMixin:
"""Repontar a mídia de um projeto para novos arquivos."""
# ========================================================================
# MEDIA RELINK
# ========================================================================
def relink_media(
self,
find: str,
replace: str,
dry_run: bool = False,
) -> Dict[str, Any]:
"""Bulk-rewrite media source paths (programmatic relink).
Rewrites the ``src`` of every ``<asset>`` / ``<media-rep>`` whose
path starts with *find*, substituting *replace* — the standard
technique for relinking a moved or renamed media folder without
opening Final Cut Pro. FCP relinks via the ``media-rep`` file URL
on import; the device-specific bookmark blob is left untouched
(FCP regenerates it).
*find* / *replace* accept plain paths (``/Volumes/OldDrive``) or
``file://`` URLs; percent-encoding in existing URLs is handled.
Matching is prefix-based on whole path segments, so ``/Media/A``
matches ``/Media/A/clip.mov`` but not ``/Media/AB/clip.mov``.
Args:
find: Old path prefix to match.
replace: New path prefix to substitute.
dry_run: When True, report what would change without
mutating the tree.
Returns:
Summary dict: ``total_assets``, ``relinked`` (reference
count), ``dry_run``, and ``changes`` — a list of
``{asset, old, new, target_exists}`` entries
(``target_exists`` checks the new path on this machine).
"""
from urllib.parse import quote, unquote, urlparse
def _to_path(value: str) -> str:
if value.startswith('file://'):
return unquote(urlparse(value).path)
return value
find_path = _to_path(find).rstrip('/')
replace_path = _to_path(replace).rstrip('/')
if not find_path:
raise ValueError("relink_media: 'find' must be a non-empty path prefix")
changes = []
for asset_id, info in self.resources.items():
elem = info['element']
targets = [(elem, elem.get('src'))]
media_rep = elem.find('media-rep')
if media_rep is not None:
targets.append((media_rep, media_rep.get('src')))
for node, old_src in targets:
if not old_src:
continue
was_url = old_src.startswith('file://')
old_path = _to_path(old_src)
if old_path != find_path and not old_path.startswith(find_path + '/'):
continue
new_path = replace_path + old_path[len(find_path):]
new_src = 'file://' + quote(new_path) if was_url else new_path
if not dry_run:
node.set('src', new_src)
info['src'] = new_src
changes.append({
'asset': info.get('name') or asset_id,
'old': old_src,
'new': new_src,
'target_exists': Path(new_path).exists(),
})
return {
'total_assets': len(self.resources),
'relinked': len(changes),
'dry_run': dry_run,
'changes': changes,
}
+126
View File
@@ -0,0 +1,126 @@
"""Reordenar clipes e recalcular offsets/duração.
Extraído de writer.py — ver fcpxml/writer/__init__.py para o conjunto.
"""
import xml.etree.ElementTree as ET
from typing import List
from ..models import (
TimeValue,
)
from .helpers import SPINE_ELEMENT_TAGS
class ReorderMixin:
"""Reordenar clipes e recalcular offsets/duração."""
# REORDER OPERATIONS
# ========================================================================
def reorder_clips(
self,
clip_ids: List[str],
target_position: str,
ripple: bool = True
) -> None:
"""
Move clips to a new position in the timeline.
Args:
clip_ids: Clips to move (maintains relative order)
target_position: 'start', 'end', timecode, or 'after:clip_id'/'before:clip_id'
ripple: Whether to shift other clips
"""
spine = self._get_spine()
# Collect clips to move
clips_to_move = []
for clip_id in clip_ids:
clip = self.clips.get(clip_id)
if clip is not None and clip in list(spine):
clips_to_move.append(clip)
if not clips_to_move:
raise ValueError(f"No clips found matching: {clip_ids}")
# Calculate total duration of moving clips
total_duration = TimeValue.zero()
for clip in clips_to_move:
dur = self._parse_time(clip.get('duration', '0s'))
total_duration = total_duration + dur
# Remove clips from current positions
for clip in clips_to_move:
spine.remove(clip)
# Determine target offset and insert index
spine_children = list(spine)
target_offset, insert_index = self._resolve_insert_position(
target_position, spine_children
)
# Insert clips at new position
current_offset = target_offset
for clip in clips_to_move:
clip.set('offset', current_offset.to_fcpxml())
spine.insert(insert_index, clip)
insert_index += 1
dur = self._parse_time(clip.get('duration', '0s'))
current_offset = current_offset + dur
# Recalculate all offsets if ripple
if ripple:
self._recalculate_offsets(spine)
def _recalculate_offsets(self, spine: ET.Element) -> None:
"""Recalculate all clip offsets sequentially."""
current_offset = TimeValue.zero()
for child in spine:
if child.tag in SPINE_ELEMENT_TAGS:
child.set('offset', current_offset.to_fcpxml())
duration_str = child.get('duration', '0s')
duration = self._parse_time(duration_str)
current_offset = current_offset + duration
def _timeline_duration(self) -> 'TimeValue':
"""Return the total timeline duration as a TimeValue.
Reads from the ``<sequence>`` element when available, falling back
to summing all spine element durations. Extracted from
``add_music_bed`` and ``batch_add_markers`` which both computed
this independently.
"""
sequence = self.root.find('.//sequence')
if sequence is not None:
dur_str = sequence.get('duration')
if dur_str:
return self._parse_time(dur_str)
spine = self._get_spine()
total = TimeValue.zero()
for child in spine:
if child.tag in SPINE_ELEMENT_TAGS:
total = total + self._parse_time(child.get('duration', '0s'))
return total
def _update_sequence_duration(self) -> None:
"""Recompute the ``<sequence>`` duration from the spine content.
Ripple edits (``cut_clip_ranges``, ``delete_clip``, ``split_clip``)
change the total timeline length without rewriting the sequence
element, so an exported file kept advertising the pre-edit duration —
a 326.78s sequence still claimed 326.78s after 71s of silence was
removed. This helper re-syncs the attribute to the actual spine sum.
"""
sequence = self.root.find('.//sequence')
if sequence is None:
return
spine = self._get_spine()
total = TimeValue.zero()
for child in spine:
if child.tag in SPINE_ELEMENT_TAGS:
total = total + self._parse_time(child.get('duration', '0s'))
sequence.set('duration', total.to_fcpxml())
# ========================================================================
+43
View File
@@ -0,0 +1,43 @@
"""Atribuir roles de vídeo/áudio.
Extraído de writer.py — ver fcpxml/writer/__init__.py para o conjunto.
"""
import xml.etree.ElementTree as ET
from typing import Optional
from .helpers import _sanitize_xml_value
class RolesMixin:
"""Atribuir roles de vídeo/áudio."""
# ROLE OPERATIONS (v0.5.0)
# ========================================================================
def assign_role(
self,
clip_id: str,
audio_role: Optional[str] = None,
video_role: Optional[str] = None,
) -> ET.Element:
"""Set the audio/video role on a clip.
Args:
clip_id: Name/ID of the clip
audio_role: Audio role (e.g., "dialogue", "music", "effects")
video_role: Video role (e.g., "video", "titles")
Returns:
The modified clip element
"""
clip = self._require_clip(clip_id)
if audio_role is not None:
clip.set('audioRole', _sanitize_xml_value(audio_role, 256))
if video_role is not None:
clip.set('videoRole', _sanitize_xml_value(video_role, 256))
return clip
# ========================================================================
+57
View File
@@ -0,0 +1,57 @@
"""Selecionar clipes por palavra-chave.
Extraído de writer.py — ver fcpxml/writer/__init__.py para o conjunto.
"""
from typing import List
class SelectionMixin:
"""Selecionar clipes por palavra-chave."""
# SELECTION OPERATIONS
# ========================================================================
def select_by_keyword(
self,
keywords: List[str],
match_mode: str = 'any',
favorites_only: bool = False,
exclude_rejected: bool = True
) -> List[str]:
"""
Find clips matching keywords.
Args:
keywords: Keywords to match
match_mode: 'any' (OR), 'all' (AND), 'none' (exclude)
favorites_only: Only return favorited clips
exclude_rejected: Exclude rejected clips
Returns:
List of matching clip IDs
"""
matches = []
for clip_id, clip in self.clips.items():
clip_keywords = set()
for kw_elem in clip.findall('keyword'):
clip_keywords.add(kw_elem.get('value', ''))
# Check keyword match
keyword_set = set(keywords)
if match_mode == 'any':
match = bool(clip_keywords & keyword_set)
elif match_mode == 'all':
match = keyword_set <= clip_keywords
elif match_mode == 'none':
match = not bool(clip_keywords & keyword_set)
else:
match = True
if match:
matches.append(clip_id)
return matches
# ========================================================================
+185
View File
@@ -0,0 +1,185 @@
"""Detectar e remover silêncio.
Extraído de writer.py — ver fcpxml/writer/__init__.py para o conjunto.
"""
from typing import Any, Dict, List, Optional
from ..models import (
MarkerType,
TimeValue,
)
from .helpers import CLIP_TAGS, build_marker_element
class SilenceMixin:
"""Detectar e remover silêncio."""
# SILENCE DETECTION OPERATIONS (v0.5.0)
# ========================================================================
def detect_silence_candidates(
self,
min_gap_seconds: float = 0.5,
patterns: Optional[List[str]] = None,
) -> List[Dict[str, Any]]:
"""Detect potential silence regions using timeline heuristics.
Checks for:
1. Gap elements in spine (high confidence)
2. Ultra-short clips < 0.5s (medium confidence)
3. Clips matching name patterns like "silence", "room tone" (high)
4. Duration anomalies > 2 std dev from mean (low-medium)
Args:
min_gap_seconds: Minimum gap duration to flag
patterns: Name patterns to match (default: gap, silence, room tone)
Returns:
List of silence candidate dicts
"""
if patterns is None:
patterns = ['gap', 'silence', 'room tone', 'dead air', 'blank']
spine = self._get_spine()
candidates = []
durations = []
clip_index = 0
# First pass: collect durations for anomaly detection
for child in spine:
if child.tag in CLIP_TAGS:
dur = self._parse_time(child.get('duration', '0s'))
durations.append(dur.to_seconds())
# Calculate stats for anomaly detection
mean_dur = sum(durations) / len(durations) if durations else 0
variance = (sum((d - mean_dur) ** 2 for d in durations) / len(durations)
if len(durations) > 1 else 0)
std_dev = variance ** 0.5
# Second pass: detect candidates
for child in spine:
tag = child.tag
offset = child.get('offset', '0s')
dur = self._parse_time(child.get('duration', '0s'))
dur_secs = dur.to_seconds()
tc = TimeValue.from_timecode(offset, self.fps).to_timecode(self.fps)
if tag == 'gap' and dur_secs >= min_gap_seconds:
candidates.append({
'start_timecode': tc,
'duration_seconds': dur_secs,
'reason': 'gap',
'confidence': 0.9,
'clip_name': None,
'clip_index': None,
})
elif tag in CLIP_TAGS:
name = child.get('name', '').lower()
# Name pattern match
for pat in patterns:
if pat.lower() in name:
candidates.append({
'start_timecode': tc,
'duration_seconds': dur_secs,
'reason': 'name_match',
'confidence': 0.85,
'clip_name': child.get('name', ''),
'clip_index': clip_index,
})
break
# Ultra-short clip
if dur_secs < 0.5:
candidates.append({
'start_timecode': tc,
'duration_seconds': dur_secs,
'reason': 'ultra_short',
'confidence': 0.6,
'clip_name': child.get('name', ''),
'clip_index': clip_index,
})
# Duration anomaly (> 2 std dev longer than mean)
if std_dev > 0 and dur_secs > mean_dur + 2 * std_dev:
candidates.append({
'start_timecode': tc,
'duration_seconds': dur_secs,
'reason': 'duration_anomaly',
'confidence': 0.4,
'clip_name': child.get('name', ''),
'clip_index': clip_index,
})
clip_index += 1
return candidates
def remove_silence_candidates(
self,
mode: str = "mark",
min_gap_seconds: float = 0.5,
min_confidence: float = 0.7,
patterns: Optional[List[str]] = None,
) -> List[Dict[str, Any]]:
"""Remove or mark detected silence candidates.
Args:
mode: "delete" removes clips/gaps, "mark" adds red markers,
"shorten" trims to minimum
min_gap_seconds: Minimum gap to consider
min_confidence: Only act on candidates above this threshold
patterns: Name patterns to match
Returns:
List of actions taken
"""
candidates = self.detect_silence_candidates(min_gap_seconds, patterns)
candidates = [c for c in candidates if c['confidence'] >= min_confidence]
spine = self._get_spine()
actions = []
if mode == "mark":
for c in candidates:
child = self._find_spine_element_at_timecode(
spine, c['start_timecode'], require_clip=True
)
if child is not None:
build_marker_element(
parent=child,
marker_type=MarkerType.STANDARD,
start=child.get('start', '0s'),
duration=f"1/{int(self.fps)}s",
name=f"SILENCE: {c['reason']}",
)
actions.append({
'action': 'marked',
'clip_name': c.get('clip_name', 'gap'),
'reason': c['reason'],
})
elif mode == "delete":
elements_to_remove = []
for c in candidates:
child = self._find_spine_element_at_timecode(
spine, c['start_timecode']
)
if child is not None:
elements_to_remove.append(child)
actions.append({
'action': 'deleted',
'clip_name': c.get('clip_name', 'gap'),
'reason': c['reason'],
})
for elem in elements_to_remove:
spine.remove(elem)
if elements_to_remove:
self._recalculate_offsets(spine)
return actions
+297
View File
@@ -0,0 +1,297 @@
"""Velocidade e zoom (punch-in) por janela.
Extraído de writer.py — ver fcpxml/writer/__init__.py para o conjunto.
"""
import xml.etree.ElementTree as ET
from fractions import Fraction
from typing import Optional
from .helpers import HOLD_AT_CUT_THRESHOLD, START_AT_CUT_THRESHOLD, _dtd_insert, _fmt_scale
class SpeedMixin:
"""Velocidade e zoom (punch-in) por janela."""
# SPEED OPERATIONS
# ========================================================================
def change_speed(
self,
clip_id: str,
speed: float,
preserve_pitch: bool = True
) -> ET.Element:
"""
Change clip playback speed.
Args:
clip_id: Target clip
speed: Speed multiplier (0.5 = half speed, 2.0 = double)
preserve_pitch: Maintain audio pitch
Returns:
Modified clip element
"""
if speed <= 0:
raise ValueError(f"Speed must be positive, got {speed}")
clip = self._require_clip(clip_id)
current_duration = self._parse_time(clip.get('duration', '0s'))
# Use rational arithmetic to avoid floating-point time values.
# FCPXML requires rational fractions with a consistent timebase,
# not decimal floats like "2.6666666666666665s".
denom = current_duration.denominator if current_duration.denominator > 0 else int(self.fps)
source_num = current_duration.numerator
speed_frac = Fraction(speed).limit_denominator(1000)
raw_num = source_num * speed_frac.denominator
raw_denom = denom * speed_frac.numerator
# Snap to frame boundary in a standard timebase (2400 ticks/sec).
# Each frame at Nfps = 2400/N ticks (e.g. 24fps → 100 ticks/frame).
fps_int = int(self.fps) if self.fps else 24
ticks_per_frame = 2400 // fps_int
dur_ticks = round(raw_num / raw_denom * 2400)
dur_ticks = round(dur_ticks / ticks_per_frame) * ticks_per_frame
new_num = dur_ticks
new_denom = 2400
# Remove any existing timeMap/conform-rate from a prior speed change
# to prevent duplicate children that produce invalid FCPXML.
for stale_tag in ('timeMap', 'conform-rate'):
for stale in clip.findall(stale_tag):
clip.remove(stale)
# Create timeMap for speed change (DTD-ordered insertion)
timemap = ET.Element('timeMap')
_dtd_insert(clip, timemap)
# Start keyframe
tp1 = ET.SubElement(timemap, 'timept')
tp1.set('time', '0s')
tp1.set('value', '0s')
tp1.set('interp', 'linear')
# End keyframe — use rational time, not floats
tp2 = ET.SubElement(timemap, 'timept')
tp2.set('time', f"{new_num}/{new_denom}s")
tp2.set('value', f"{source_num}/{denom}s")
tp2.set('interp', 'linear')
# Update clip duration (rational, not simplified to arbitrary denominator)
clip.set('duration', f"{new_num}/{new_denom}s")
# Add conform-rate (DTD-ordered insertion)
conform = ET.Element('conform-rate')
conform.set('scaleEnabled', '1')
conform.set('srcFrameRate', str(int(self.fps)))
_dtd_insert(clip, conform)
return clip
def add_zoom(
self,
clip_id: 'str | ET.Element',
start: float,
end: float,
scale: float = 1.3,
ease: float = 0.25,
position: str = "0 0",
ease_out: Optional[float] = None,
hold_at_end: Optional[bool] = None,
start_at_peak: Optional[bool] = None,
) -> ET.Element:
"""Add a punch-in zoom to a clip, snapping back to its framing at the end.
Animates ``<adjust-transform>``'s ``scale`` param (``<param>`` +
``<keyframeAnimation>`` of ``<keyframe>``) from the clip's current
scale up to *scale* times it, holds, then returns — all within
``[start, end]`` — clip-relative seconds (same convention as
``cut_clip_ranges``).
The two ends are deliberately asymmetric. *ease* ramps the zoom
**in** over half a second by default, fast enough to land with the
emphasised word. The way **out** is instant — a single frame — so
the moment the impact phrase ends the shot is simply back to its
normal framing and the video resumes its flow, with no drift
drawing attention to itself. Pass *ease_out* to ramp the return
gradually instead.
*hold_at_end* keeps the peak instead of returning, and
*start_at_peak* opens already zoomed with no ramp. Left as ``None``
both decide on their own from how close the window sits to the
clip's edges: a cut is itself the transition, so ramping away from
one — or back toward one — is motion the viewer reads as a wobble
rather than as emphasis.
"""
if end <= start:
raise ValueError(f"end ({end}) must be greater than start ({start})")
if ease <= 0:
raise ValueError(f"ease must be positive, got {ease}")
if scale <= 0:
raise ValueError(f"scale must be positive, got {scale}")
frame = float(self.frame_duration_fraction())
ramp_out = frame if ease_out is None else ease_out
if ramp_out <= 0:
raise ValueError(f"ease_out must be positive, got {ease_out}")
clip = self._require_clip(clip_id)
clip_duration = self._parse_time(clip.get('duration', '0s')).to_seconds()
if start < 0 or end > clip_duration:
raise ValueError(
f"zoom window [{start}, {end}]s must fall within the clip's "
f"duration (0 to {clip_duration:.3f}s)"
)
# Replace a prior zoom, but never the clip's framing. A clip can
# already carry an <adjust-transform> holding the editor's own
# reframe — rotation for footage shot sideways, position, a scale
# that makes the shot work at all. Dropping it outright (the old
# behaviour) silently destroyed that framing; on real footage the
# zoomed section came back rotated. So: keep the static attributes,
# and animate *relative to* the existing scale.
base_x, base_y = 1.0, 1.0
carried: dict = {}
old_keyframes: list = []
for stale in clip.findall('adjust-transform'):
carried = {k: v for k, v in stale.attrib.items() if k != 'scale'}
parts = (stale.get('scale') or '').split()
if len(parts) == 2:
try:
base_x, base_y = float(parts[0]), float(parts[1])
except ValueError:
base_x, base_y = 1.0, 1.0
else:
# No static attribute — a PRIOR zoom on this same clip left
# an animated <param name="scale"> instead, and the true
# resting framing lives in its keyframes, not in 1.0.
# Reading it as 1.0 here doesn't just miss the framing: it
# replaces the earlier zoom's whole animation with a wrong
# one, since this loop unconditionally removes `stale`
# right after. The rest value is recoverable without
# knowing which keyframe it is: MIN_ZOOM_SCALE == 1.0 means
# every keyframed value is >= the rest scale, so the
# smallest one keyframed is the rest value, peak or not.
for old_param in stale.findall("param[@name='scale']"):
xs, ys = [], []
for kf in old_param.findall('.//keyframe'):
kv = (kf.get('value') or '').split()
if len(kv) == 2:
try:
xs.append(float(kv[0]))
ys.append(float(kv[1]))
except ValueError:
pass
# Kept for merging: a second zoom on the same clip
# (two emphatic beats a cut didn't separate) should
# stack alongside the first, not erase it — the
# earlier peak is still a real editorial decision.
old_keyframes.append((kf.get('time', '0s'), kf.get('value', '')))
if xs and ys:
base_x, base_y = min(xs), min(ys)
clip.remove(stale)
transform = ET.Element('adjust-transform')
for key, value in carried.items():
transform.set(key, value)
scale_param = ET.SubElement(transform, 'param')
scale_param.set('name', 'scale')
anim = ET.SubElement(scale_param, 'keyframeAnimation')
# Keyframe times live in the clip's SOURCE timebase — the same origin
# as its own ``start`` — not in clip-relative seconds. A clip whose
# media starts at, say, 3109.9s of timecode looks for the animation
# there; keyframes written at 0-5s land outside the clip entirely and
# Final Cut imports the zoom as nothing at all, silently. Matches what
# add_text_title already does, and only shows up on footage whose
# start isn't 0s — every synthetic fixture starts at 0s and hides it.
media_origin = self._parse_time(clip.get('start', '0s'))
rest_value = f"{_fmt_scale(base_x)} {_fmt_scale(base_y)}"
scale_value = f"{_fmt_scale(base_x * scale)} {_fmt_scale(base_y * scale)}"
# A return that lands right before a cut is wasted motion: the next
# clip begins on its own framing anyway, so all the viewer sees is a
# twitch on the way out. When the zoom runs to the end of the clip,
# hold the peak and let the cut do the resetting.
holds_to_cut = (
hold_at_end
if hold_at_end is not None
else (clip_duration - end) <= HOLD_AT_CUT_THRESHOLD
)
opens_at_peak = (
start_at_peak
if start_at_peak is not None
else start <= START_AT_CUT_THRESHOLD
)
# Only the ramps actually written have to fit in the window: a zoom
# that opens at the peak spends no time ramping in, and one held to
# the cut spends none ramping out.
needed = (0.0 if opens_at_peak else ease) + (0.0 if holds_to_cut else ramp_out)
if needed > (end - start):
raise ValueError(
f"the ramps ({needed}s) don't fit in the zoom window "
f"({end - start}s) — shorten them or widen start/end"
)
if opens_at_peak:
# The cut already delivered the change of framing; ramping up
# from it just looks like the shot settling.
keyframes = [(start, scale_value)]
else:
keyframes = [(start, rest_value), (start + ease, scale_value)]
if holds_to_cut:
keyframes.append((end, scale_value))
else:
# Hold the peak right up to the end, then drop back on the very
# next frame — the snap-back the edit wants, not a slow drift.
keyframes.append((end - ramp_out, scale_value))
keyframes.append((end, rest_value))
new_entries = [
((media_origin + self.snap_seconds_to_frame(seconds)), value)
for seconds, value in keyframes
]
new_start_time = new_entries[0][0]
new_end_time = new_entries[-1][0]
# Two calls on the same clip mean two different things depending on
# whether their windows overlap. Overlapping = redoing the *same*
# zoom with new numbers — the old keyframes are stale and all of
# them go. Disjoint = a second, separate beat that a cut didn't
# separate onto its own clip — that one stacks alongside the first
# instead of erasing it, since both are real editorial decisions.
old_times = [self._parse_time(t) for t, _ in old_keyframes]
old_span_overlaps_new = bool(old_times) and not (
max(old_times) < new_start_time or min(old_times) > new_end_time
)
if old_span_overlaps_new:
surviving_old: list = []
else:
surviving_old = [(self._parse_time(t), v) for t, v in old_keyframes]
all_entries = sorted(surviving_old + new_entries, key=lambda e: e[0])
for time_value, value in all_entries:
kf = ET.SubElement(anim, 'keyframe')
kf.set('time', time_value.to_fcpxml())
kf.set('value', value)
# Only 'time' and 'value' — no 'interp', no 'curve'. The DTD allows
# both, but Final Cut rejected 'interp' on this vector param
# ("does not support the interpolation attribute") and discarded
# the whole <param>. A hand-made zoom exported from FCP itself
# writes bare keyframes and relies on the DTD default
# (curve="smooth"), so we match that export exactly rather than
# guess which attributes survive its importer.
if position != "0 0":
pos_param = ET.SubElement(transform, 'param')
pos_param.set('name', 'position')
pos_param.set('value', position)
_dtd_insert(clip, transform)
return clip
# ========================================================================
+607
View File
@@ -0,0 +1,607 @@
"""Títulos de texto e legendas dinâmicas.
Extraído de writer.py — ver fcpxml/writer/__init__.py para o conjunto.
"""
import re
import unicodedata
import uuid
import xml.etree.ElementTree as ET
from typing import Any, Dict, List, Optional
from ..collision import blocking, validate_titles
from ..models import (
DynamicSubtitleConfig,
TimeValue,
)
from ..text_layout import (
TEXT_TEMPLATE_FONT_SCALE,
LayoutBox,
compose_sentence,
layout_sentence,
)
from ..transcribe import group_words_by_segment
from .helpers import _dtd_insert, _sanitize_xml_value
class TitlesMixin:
"""Títulos de texto e legendas dinâmicas."""
# DYNAMIC (KARAOKE-STYLE) SUBTITLES
# ========================================================================
# The "Text" (Basic Text) template — the ONLY simple title template that
# Final Cut actually renders. Copied verbatim from the user's own FCP
# exports ("teste.fcpxmld" and "posição.fcpxmld", FCP 1.14 in English):
# a single "<text>" run, one "<text-style-def>", and a fixed param block
# with the margins/alignment/speed the template ships with. Every prior
# title template we generated ("Essencial - Título", "Título Básico")
# imported cleanly but never appeared — their Motion uids did not resolve
# to a real, drawable template in FCP, which discards the clip silently.
# "Text" is what FCP itself writes when the user adds a title by hand, so
# it is the ground truth. See Engine/docs/05_EXPERIENCIAS.md, 2026-08-17.
_TEXT_TITLE_UID = (
'.../Titles.localized/Basic Text.localized/'
'Text.localized/Text.moti'
)
_TEXT_TITLE_START = '86486400/24000s'
# The Inspector's Position field, and the one this code overrides per
# title so two titles never stack on top of each other. Verified in
# "posição.fcpxmld": each hand-dragged title carries a distinct "x y"
# value here while every other param stays identical.
_TEXT_POSITION_KEY = '9999/10003/13260/3296672360/1/100/101'
# Layout params the "Text" template ships with. These keys are the
# template's own defaults and never vary between instances.
#
# "Build Out" is the one deliberate override: with "Apply Speed" set to
# "2 (Per Object)" below, the template's whole built-in animation (build
# in + build out) is always compressed to exactly fill the title's own
# on-screen duration — so on a short word-length clip, build out was
# eating time that build in needed to finish revealing the text before
# the cut. Disabling build out hands that entire compressed window to
# build in alone, which is what "sempre acelerado" turned out to mean:
# no separate speed knob needed. Value captured from a real FCP export
# with "Build Out" unchecked in the Inspector (see chat, 2026-08-18).
_TEXT_TITLE_PARAMS = (
('Build Out', '9999/10000/2/102', '0'),
('Layout Method', '9999/10003/13260/3296672360/2/314', '1 (Paragraph)'),
('Left Margin', '9999/10003/13260/3296672360/2/323', '-1210'),
('Right Margin', '9999/10003/13260/3296672360/2/324', '1210'),
('Top Margin', '9999/10003/13260/3296672360/2/325', '2160'),
('Bottom Margin', '9999/10003/13260/3296672360/2/326', '-2160'),
('Alignment', '9999/10003/13260/3296672360/2/354/3296667315/401', '1 (Center)'),
('Line Spacing', '9999/10003/13260/3296672360/2/354/3296667315/404', '-19'),
('Auto-Shrink', '9999/10003/13260/3296672360/2/370', '3 (To All Margins)'),
('Alignment', '9999/10003/13260/3296672360/2/373', '0 (Left) 1 (Middle)'),
('Opacity', '9999/10003/13260/3296672360/4/3296673134/1000/1044', '0'),
('Speed', '9999/10003/13260/3296672360/4/3296673134/201/208', '6 (Custom)'),
('Apply Speed', '9999/10003/13260/3296672360/4/3296673134/201/211', '2 (Per Object)'),
)
# "Custom Speed" sits between "Speed" and "Apply Speed" and carries a
# <keyframeAnimation> child rather than a plain value attribute. Its two
# keyframes are the template's own absolute nominal times, constant across
# every instance, so they are safe to replay verbatim.
_TEXT_CUSTOM_SPEED_KEY = '9999/10003/13260/3296672360/4/3296673134/201/209'
_TEXT_CUSTOM_SPEED_KEYFRAMES = (
('-469658744/1000000000s', '0'),
('12328542033/1000000000s', '1'),
)
_TEXT_SIZE_KEY = '9999/10003/13260/3296672360/5/3296672362/3'
def _ensure_text_title_effect(self, resources: ET.Element) -> str:
"""Return the resource id of the "Text" (Basic Text) effect, creating it if absent."""
return self._ensure_effect(resources, self._TEXT_TITLE_UID, 'Text', 'r_text')
def _ensure_effect(
self,
resources: ET.Element,
uid: str,
name: str,
id_prefix: str,
) -> str:
"""Return the id of the effect resource with *uid*, creating it if absent."""
for eff in resources.findall('effect'):
if eff.get('uid') == uid:
return eff.get('id')
effect_id = self._unique_resource_id(resources, id_prefix)
eff_el = ET.SubElement(resources, 'effect')
eff_el.set('id', effect_id)
eff_el.set('name', name)
eff_el.set('uid', uid)
return effect_id
# <text-style-def id> / <text-style ref> are DTD type ID/IDREF, so the
# value must be a valid XML Name: letters, digits, "_", "-", "." only,
# never starting with a digit. Title names are built from the caption
# text ("Olá mundo - Text"), which carries spaces, accents and often a
# leading digit — xmllint rejected the whole document with "Syntax of
# value for attribute id of text-style-def is not valid".
_TEXT_STYLE_ID_UNSAFE = re.compile(r'[^A-Za-z0-9_.-]+')
def _unique_text_style_id(self, base: str) -> str:
"""Return a document-unique, DTD-valid XML ID for a ``<text-style-def>``."""
folded = unicodedata.normalize('NFKD', base).encode('ascii', 'ignore').decode('ascii')
slug = self._TEXT_STYLE_ID_UNSAFE.sub('_', folded).strip('_.-')[:48]
stem = f"ts_{slug}" if slug else "ts"
if self._text_style_ids is None:
self._text_style_ids = {
sd.get('id') for sd in self.root.findall('.//text-style-def')
}
candidate = f"{stem}_0"
counter = 0
while candidate in self._text_style_ids:
counter += 1
candidate = f"{stem}_{counter}"
self._text_style_ids.add(candidate)
return candidate
def _reassign_text_style_ids(self, clip: ET.Element) -> None:
"""Give every ``<text-style-def>`` inside a just-deepcopy'd *clip* a
fresh document-unique id, repointing any ``<text-style ref="...">``
in the same subtree that pointed at the old one.
``split_clip``/``cut_clip_ranges`` deepcopy the clip once per
resulting segment, so a clip carrying a ``<title>`` (from a "text"
voice action) keeps the exact same ``text-style-def id`` in every
copy. A single cut is harmless — but the batch chain re-cuts the
same clip at each step (silence removal, filler removal, dynamic
subtitles), and every pass multiplies the duplicate, so the DTD
validator eventually rejects the file with "ID ... already
defined". Regenerating here, at the only place copies are made,
fixes it for every caller instead of each one having to remember to.
"""
for style_def in clip.findall('.//text-style-def'):
old_id = style_def.get('id')
if not old_id:
continue
slug = old_id[3:] if old_id.startswith('ts_') else old_id
slug = re.sub(r'_\d+$', '', slug) # drop a prior _<N> counter
new_id = self._unique_text_style_id(slug)
if new_id == old_id:
continue
style_def.set('id', new_id)
for ref_el in clip.findall(f".//text-style[@ref='{old_id}']"):
ref_el.set('ref', new_id)
def _make_text_title_clip(
self,
effect_id: str,
text: str,
offset: 'TimeValue',
duration: 'TimeValue',
*,
lane: int,
name: str,
position: Optional[str] = None,
font: str = 'Helvetica Neue',
font_size: int = 196,
font_color: str = '1 1 1 1',
bold: bool = True,
face: Optional[str] = None,
kerning: Optional[float] = None,
font_scale: float = TEXT_TEMPLATE_FONT_SCALE,
animated: bool = True,
size_param: Optional[float] = None,
) -> ET.Element:
"""Build a standalone ``<title>`` clip from the "Text" (Basic Text) template.
Reproduces FCP's own output for a hand-added title exactly — the only
template we have verified renders in Final Cut ("teste.fcpxmld" and
"posição.fcpxmld"). *position* ("x y" canvas points) is the Inspector
Position value; omit it to keep the template's centred default. Unlike
the animated templates, this carries no animation switch, so the text
stays put and visible for its whole duration.
"""
elem = ET.Element('title')
elem.set('ref', effect_id)
elem.set('lane', str(lane))
elem.set('offset', offset.to_fcpxml())
elem.set('name', _sanitize_xml_value(name, 256))
elem.set('start', self._TEXT_TITLE_START)
elem.set('duration', duration.to_fcpxml())
if position:
param = ET.SubElement(elem, 'param')
param.set('name', 'Position')
param.set('key', self._TEXT_POSITION_KEY)
param.set('value', position)
def _add_param(name: str, key: str, value: str) -> None:
param = ET.SubElement(elem, 'param')
param.set('name', name)
param.set('key', key)
param.set('value', value)
animation_params = {'Opacity', 'Speed', 'Apply Speed'}
for param_name, param_key, param_value in self._TEXT_TITLE_PARAMS:
if not animated and param_name in animation_params:
continue
_add_param(param_name, param_key, param_value)
if animated and param_name == 'Speed':
# "Custom Speed" lands between "Speed" and "Apply Speed" and
# carries a <keyframeAnimation> child instead of a value.
cs = ET.SubElement(elem, 'param')
cs.set('name', 'Custom Speed')
cs.set('key', self._TEXT_CUSTOM_SPEED_KEY)
anim = ET.SubElement(cs, 'keyframeAnimation')
for kf_time, kf_value in self._TEXT_CUSTOM_SPEED_KEYFRAMES:
kf = ET.SubElement(anim, 'keyframe')
kf.set('time', kf_time)
kf.set('value', kf_value)
if size_param is not None:
_add_param('Size', self._TEXT_SIZE_KEY, f"{float(size_param):g}")
text_el = ET.SubElement(elem, 'text')
ts_id = self._unique_text_style_id(name)
run = ET.SubElement(text_el, 'text-style')
run.set('ref', ts_id)
run.text = _sanitize_xml_value(text, 256)
style_def = ET.SubElement(elem, 'text-style-def')
style_def.set('id', ts_id)
text_style = ET.SubElement(style_def, 'text-style')
text_style.set('font', font)
# Text.moti sizes type in frame pixels but positions in canvas points.
# See TEXT_TEMPLATE_FONT_SCALE: layout measures in points, so only the
# emitted size (and its kerning, to keep the same letter spacing) is
# converted here.
scale = float(font_scale) or 1.0
text_style.set('fontSize', f"{float(font_size) * scale:g}")
text_style.set('fontColor', font_color)
# FCP represents bold weight as the bold attribute — never as a
# fontFace. Writing ``bold="0" fontFace="Bold"`` (the previous
# behaviour) is contradictory and FCP refuses to render the text.
# Italic, by contrast, IS a face: FCP writes both ``fontFace`` and
# ``italic="1"``. See Engine/docs/05_EXPERIENCIAS.md, entry 2026-08-19.
face_lower = (face or '').strip().lower()
if face_lower == 'bold':
text_style.set('bold', '1')
elif 'italic' in face_lower:
text_style.set('fontFace', face)
text_style.set('italic', '1')
else:
if bold:
text_style.set('bold', '1')
if face:
text_style.set('fontFace', face)
if kerning:
text_style.set('kerning', f"{float(kerning) * scale:g}")
text_style.set('alignment', 'center')
text_style.set('lineSpacing', '-19')
return elem
def add_text_title(
self,
parent_clip: 'str | ET.Element',
text: str,
*,
offset: str = '0s',
duration: str = '1s',
lane: int = 1,
position: Optional[str] = None,
font: str = 'Helvetica Neue',
font_size: int = 196,
font_color: str = '1 1 1 1',
bold: bool = True,
face: Optional[str] = None,
animated: bool = True,
font_scale: float = TEXT_TEMPLATE_FONT_SCALE,
size_param: Optional[float] = None,
) -> ET.Element:
"""Add a single static "Text" (Basic Text) title over *parent_clip*.
Anchored in SOURCE media coordinates (parent's ``start`` + *offset*),
matching FCP's own output, so the title lands on screen instead of at
~0s of the media (which FCP silently drops). *offset* and *duration*
accept any FCPXML rational-time string; *position* is an optional
"x y" canvas-point string to keep two titles from stacking.
Returns:
The created ``<title>`` element, already inserted into the parent
in DTD order.
"""
parent = parent_clip if isinstance(parent_clip, ET.Element) else self._require_clip(parent_clip)
resources = self.root.find('.//resources')
if resources is None:
raise ValueError("No <resources> element found in FCPXML")
effect_id = self._ensure_text_title_effect(resources)
media_origin = self._parse_time(parent.get('start', '0s'))
relative = self._parse_time(offset)
title = self._make_text_title_clip(
effect_id,
text,
media_origin + relative,
self._parse_time(duration),
lane=lane,
name=f"{text} - Text",
position=position,
font=font,
font_size=font_size,
font_color=font_color,
bold=bold,
face=face,
animated=animated,
font_scale=font_scale,
size_param=size_param,
)
_dtd_insert(parent, title)
return title
def generate_dynamic_subtitles(
self,
parent_clip: 'str | ET.Element',
words: List[Dict[str, Any]],
config: Optional['DynamicSubtitleConfig'] = None,
segments: Optional[List[Dict[str, Any]]] = None,
) -> List[ET.Element]:
"""Generate progressive-reveal subtitle titles, one per word.
Groups *words* into sentences (by *segments*' time windows), lays each
sentence out as a compact typographic block, and emits one standalone
``<title>`` per word, positioned at its place in that block. Words
appear one by one as they are spoken and accumulate on screen; every
word of a block then clears at the same instant, so the sentence
vanishes as a whole before the next one builds up.
Each word gets its own lane, since a block's words are all on screen
together. Lanes restart with each block. Size, colour, font and face
cycle through ``config.style.rhythm``, reproducing the typography of
the calibration export the user built in Final Cut.
A sentence too tall for the band is split into successive blocks, so a
long sentence never spills off screen.
Args:
parent_clip: The spine clip to attach titles to — either its
Name/ID (resolved via ``_require_clip``, kept for backward
compatibility) or the ``ET.Element`` itself. **Callers
iterating multiple spine clips must pass the element, not
the name**: after any ripple-cut/silence-removal operation,
every fragment of an originally-named clip keeps that same
``name``, so ``self.clips`` (keyed by name) only retains the
last-indexed one — a name lookup then silently resolves
every call to the SAME wrong clip, stacking every line from
every distinct clip's transcript onto one spine element (see
Engine/docs/05_EXPERIENCIAS.md, entry 2026-08-17).
words: ``[{'word': str, 'start': float, 'end': float}, ...]``
with ``start``/``end`` in seconds *relative to the parent
clip's own start* (same convention as ``add_connected_clip``'s
``offset``).
config: Styling/layout options; defaults to ``DynamicSubtitleConfig()``.
segments: Whisper sentence segments ``[{'start', 'end', ...}]``, on
the same relative timebase as *words*. Omitted, every word
falls into a single sentence, which the block layout then
splits by height alone.
Returns:
The list of created ``<title>`` elements, in chronological order.
"""
if config is None:
config = DynamicSubtitleConfig()
if not words:
return []
parent = parent_clip if isinstance(parent_clip, ET.Element) else self._require_clip(parent_clip)
resources = self.root.find('.//resources')
if resources is None:
raise ValueError("No <resources> element found in FCPXML")
effect_id = self._ensure_text_title_effect(resources)
# A connected title is NOT trimmed by its parent clip's out-point —
# Final Cut keeps drawing it over whatever clip follows. A word that
# starts after the cut would therefore only ever be seen on top of the
# NEXT clip's own captions, so it is dropped rather than placed.
parent_limit = self._parse_time(parent.get('duration', '0s'))
has_limit = TimeValue(0, 1) < parent_limit
if has_limit:
limit_seconds = parent_limit.to_seconds()
words = [
w for w in words
if float(w.get('start', 0.0)) < limit_seconds
]
if not words:
return []
# Split into sentences, then lay each one out as a block. A sentence
# too tall for the band comes back with overflow, which becomes the
# next block — the sub-sentence split that keeps long sentences from
# spilling off screen.
sentences = group_words_by_segment(words, segments or [])
box = LayoutBox.for_frame(
self.frame_width(), self.frame_height(),
band_height=config.band_height,
center_y=config.block_center_y,
)
# "phrase" is the progressive composition the reference reel uses: one
# title per LINE ("que vão" / "melhorar" / "sua legenda"), the key word
# set large in a display italic. "word" is the older one-title-per-word
# rhythm, kept for callers that want every word to land on its own.
phrase_mode = getattr(config, 'granularity', 'phrase') == 'phrase'
def lay_out(pending: List[Dict]):
"""Place what fits; return (units, still-unplaced words)."""
if phrase_mode:
composition = compose_sentence(
pending, config.style, box, line_gap=config.line_gap,
)
return composition.blocks, composition.overflow
layout = layout_sentence(pending, config.style, box)
return layout.placed, layout.overflow
blocks: List[List[Any]] = []
for sentence in sentences:
remaining = list(sentence)
while remaining:
units, remaining = lay_out(remaining)
if not units:
break
blocks.append(units)
if not blocks:
return []
# Never emit a zero-duration frame (rounds to 0 at the sequence's fps
# and FCP rejects it as "unexpected value found").
min_dur_tv = self.snap_seconds_to_frame(
float(self.frame_duration_fraction())
)
# Every word of a block clears at the same instant: when the next block
# starts, or at the last word's end for the final block. That is what
# makes a sentence build up and then vanish all at once.
block_starts = [
self.snap_seconds_to_frame(min(unit.start for unit in units))
for units in blocks
]
# Whisper's word end can also run past the cut, so a last block would
# linger over the next clip's first block. Nothing may outlive the
# clip it was written for.
block_ends: List[TimeValue] = []
for i, units in enumerate(blocks):
if i + 1 < len(blocks):
end = block_starts[i + 1]
else:
end = self.snap_seconds_to_frame(
max(unit.end for unit in units)
)
if end - block_starts[i] < min_dur_tv:
end = block_starts[i] + min_dur_tv
if has_limit and parent_limit < end:
end = parent_limit
block_ends.append(end)
# Anchored titles are positioned in the parent clip's SOURCE media
# coordinates: a title's offset is the parent clip's `start` plus its
# timeline-relative position. Verified against FCP's own output in
# "exemplo de arquivos.fcpxmld", where the hand-made "Essencial -
# Título" sits at offset 226040815/24000s on a parent starting at
# 226007782/24000s — 1.376s into a 1.835s clip. Writing a plain
# relative offset instead would drop the title to ~0s of the media,
# before the clip's own in-point, so it lands outside the clip and FCP
# never shows it.
media_origin = self._parse_time(parent.get('start', '0s'))
created: List[ET.Element] = []
for units, block_end in zip(blocks, block_ends):
for index, unit in enumerate(units):
relative_offset = self.snap_seconds_to_frame(unit.start)
duration = block_end - relative_offset
if duration < min_dur_tv:
duration = min_dur_tv
# Units of one block are all on screen together, so no two may
# share a lane. Lanes restart each block, which is free — the
# previous block has already cleared.
lane = index + 1
offset = media_origin + relative_offset
title = self._make_text_title_clip(
effect_id,
unit.text,
offset,
duration,
lane=lane,
name=f"caption_{uuid.uuid4().hex[:8]}",
position=unit.position_param(config.text_scale),
font=unit.font or config.style.font,
font_size=int(round(unit.font_size)),
font_color=unit.color or config.style.active_color,
bold=config.style.bold,
face=unit.face,
kerning=unit.kerning,
font_scale=config.text_scale,
)
_dtd_insert(parent, title)
created.append(title)
if getattr(config, 'validate', False):
report = self.validate_subtitle_layout()
if blocking(report["severity"]):
raise ValueError(
"Subtitle layout validation failed: "
+ str(report["summary"])
)
return created
def validate_subtitle_layout(
self,
*,
safe_margin_x: float = 0.05,
safe_margin_y: float = 0.05,
min_font_size: Optional[float] = None,
min_distance: Optional[float] = None,
max_distance: Optional[float] = None,
) -> dict:
"""Re-measure every ``<title>`` in the document and report collisions.
Reconstructs each title's on-screen box from the values the writer
emitted (``fontSize``/``kerning``/``Position`` are already in template
space), then checks for temporal+spatial collisions, frame/safe-area
containment, and font fallbacks. This is the spec-16 validation pass the
layout engine does not do on its own — it only guarantees non-overlap
*by construction* while composing, and cannot see a hand-edited title.
Returns the ``collision.validate_titles`` report: ``severity`` (worst
bucket), ``issues`` (spec-16 occurrences) and ``summary`` (counts).
"""
titles = []
for elem in self.root.iter('title'):
# enabled="0" never renders in Final Cut (see
# generate_subtitles_by_emphasis, which disables plain titles
# under an emphasis phrase instead of never creating them) — a
# title that is off by design must not count as a collision
# against the one drawn in its place.
if elem.get('enabled', '1') == '0':
continue
text_el = elem.find('text/text-style')
text = (text_el.text or '').strip() if text_el is not None else ''
style = elem.find('text-style-def/text-style')
font = style.get('font') if style is not None else None
face = style.get('fontFace') if style is not None else None
font_size = (
float(style.get('fontSize', '0')) if style is not None else 0.0
)
kerning = (
float(style.get('kerning', '0') or 0)
if style is not None else 0.0
)
x = y = 0.0
for param in elem.findall('param'):
if param.get('name') == 'Position' and param.get('value'):
parts = param.get('value').split()
if len(parts) >= 2:
x, y = float(parts[0]), float(parts[1])
start = self._parse_time(elem.get('offset', '0s')).to_seconds()
duration = self._parse_time(elem.get('duration', '0s')).to_seconds()
titles.append({
'text': text,
'font': font,
'face': face,
'font_size': font_size,
'kerning': kerning,
'x': x,
'y': y,
'start': start,
'end': start + duration,
'group': start + duration,
})
return validate_titles(
titles,
self.frame_width(),
self.frame_height(),
safe_margin_x=safe_margin_x,
safe_margin_y=safe_margin_y,
min_font_size=min_font_size,
min_distance=min_distance,
max_distance=max_distance,
)
# ========================================================================
+94
View File
@@ -0,0 +1,94 @@
"""Transições entre clipes vizinhos.
Extraído de writer.py — ver fcpxml/writer/__init__.py para o conjunto.
"""
import xml.etree.ElementTree as ET
from ..models import (
TimeValue,
)
from .helpers import FCP_EFFECTS
class TransitionsMixin:
"""Transições entre clipes vizinhos."""
# TRANSITION OPERATIONS
# ========================================================================
def add_transition(
self,
clip_id: str,
position: str = 'end',
transition_type: str = 'cross-dissolve',
duration: str = '00:00:00:15'
) -> ET.Element:
"""
Add a transition to a clip.
Args:
clip_id: Target clip
position: 'start', 'end', or 'both'
transition_type: Type of transition
duration: Transition duration
Returns:
Created transition element(s)
"""
spine, clip, clip_index = self._require_spine_clip(clip_id)
trans_duration = self._parse_time(duration)
# Effect name and FCP built-in effect UID lookup via registry
effect_name, effect_uid = FCP_EFFECTS.get(
transition_type,
FCP_EFFECTS['cross-dissolve']
)
# Ensure effect resource exists in <resources>
effect_ref_id = None
if effect_uid:
root = self.tree.getroot()
resources = root.find('.//resources')
if resources is not None:
for eff in resources.findall('effect'):
if eff.get('uid') == effect_uid:
effect_ref_id = eff.get('id')
break
if effect_ref_id is None:
effect_ref_id = self._unique_resource_id(resources, 'r_dissolve')
eff_el = ET.SubElement(resources, 'effect')
eff_el.set('id', effect_ref_id)
eff_el.set('name', effect_name)
eff_el.set('uid', effect_uid)
transitions_added = []
_, clip_dur, clip_offset = self._get_clip_times(clip)
half_dur = trans_duration * 0.5
if position in ('end', 'both'):
end_offset = clip_offset + clip_dur - half_dur
transition = self._make_transition_element(
effect_name, end_offset, trans_duration, effect_ref_id
)
spine.insert(clip_index + 1, transition)
transitions_added.append(transition)
if position in ('start', 'both'):
start_offset = clip_offset - half_dur
if start_offset < TimeValue.zero():
raise ValueError(
f"Transition at start would produce negative offset "
f"({start_offset.to_seconds():.3f}s) for clip '{clip_id}'"
)
transition = self._make_transition_element(
effect_name, start_offset, trans_duration, effect_ref_id
)
spine.insert(clip_index, transition)
transitions_added.append(transition)
return transitions_added[0] if len(transitions_added) == 1 else transitions_added
# ========================================================================
+125
View File
@@ -0,0 +1,125 @@
"""Aparar clipes e propagar o ripple pela spine.
Extraído de writer.py — ver fcpxml/writer/__init__.py para o conjunto.
"""
import xml.etree.ElementTree as ET
from typing import Optional
from ..models import (
TimeValue,
)
from .helpers import SPINE_ELEMENT_TAGS
class TrimMixin:
"""Aparar clipes e propagar o ripple pela spine."""
# ========================================================================
# TRIM OPERATIONS
# ========================================================================
def trim_clip(
self,
clip_id: str,
trim_start: Optional[str] = None,
trim_end: Optional[str] = None,
ripple: bool = True
) -> ET.Element:
"""
Trim a clip's in-point and/or out-point.
Args:
clip_id: Target clip
trim_start: New in-point or delta ('+1s', '-10f')
trim_end: New out-point or delta
ripple: Whether to shift subsequent clips
Returns:
Modified clip element
"""
clip = self._require_clip(clip_id)
current_start, current_duration, _ = self._get_clip_times(clip)
original_duration = current_duration
# Handle trim_start
if trim_start:
if trim_start.startswith('+') or trim_start.startswith('-'):
delta = self._parse_time(trim_start[1:])
if trim_start.startswith('-'):
# Extend earlier
new_start = current_start - delta
new_duration = current_duration + delta
else:
# Trim later
new_start = current_start + delta
new_duration = current_duration - delta
else:
new_start = self._parse_time(trim_start)
diff = new_start - current_start
new_duration = current_duration - diff
clip.set('start', new_start.to_fcpxml())
current_start = new_start
current_duration = new_duration
# Handle trim_end
if trim_end:
if trim_end.startswith('+') or trim_end.startswith('-'):
delta = self._parse_time(trim_end[1:])
if trim_end.startswith('-'):
new_duration = current_duration - delta
else:
new_duration = current_duration + delta
else:
end_point = self._parse_time(trim_end)
new_duration = end_point - current_start
current_duration = new_duration
if current_duration <= TimeValue.zero():
raise ValueError(
f"Trim would produce non-positive duration "
f"({current_duration.to_seconds():.3f}s) for clip '{clip_id}'"
)
clip.set('duration', current_duration.to_fcpxml())
# Ripple subsequent clips if needed
if ripple:
duration_change = current_duration - original_duration
if duration_change != TimeValue.zero():
self._ripple_after_clip(clip, duration_change)
return clip
def _ripple_from_index(
self, spine: ET.Element, start_index: int, delta: 'TimeValue'
) -> None:
"""Shift the offset of every spine element from *start_index* onward by *delta*.
Consolidates the ripple loops previously duplicated across
``_ripple_after_clip``, ``delete_clip``, and ``insert_clip``.
Args:
spine: The primary storyline ``<spine>`` element.
start_index: First child index to adjust (inclusive).
delta: Signed time shift (positive = later, negative = earlier).
"""
children = list(spine)
for child in children[start_index:]:
if child.tag in SPINE_ELEMENT_TAGS:
current_offset = self._parse_time(child.get('offset', '0s'))
new_offset = current_offset + delta
child.set('offset', new_offset.to_fcpxml())
def _ripple_after_clip(self, target_clip: ET.Element, delta: TimeValue) -> None:
"""Shift all clips after the given clip by delta."""
spine = self._get_spine()
clip_index = self._find_clip_index(spine, target_clip)
if clip_index is not None:
self._ripple_from_index(spine, clip_index + 1, delta)
# ========================================================================
+232
View File
@@ -0,0 +1,232 @@
"""Verificações estruturais do FCPXML antes de salvar.
Extraído de writer.py — ver fcpxml/writer/__init__.py para o conjunto.
"""
import logging
import xml.etree.ElementTree as ET
from fractions import Fraction
from typing import List, Optional
from ..models import (
_FCPXML_STANDARD_TIMEBASES,
TimeValue,
ValidationIssue,
ValidationIssueType,
)
from .helpers import _ASSET_CLIP_CHILD_ORDER, _CHILD_ORDER_INDEX
# ============================================================================
# PRE-EXPORT DTD VALIDATOR (v0.6.0)
# ============================================================================
_log = logging.getLogger(__name__)
def _check_child_order(root: ET.Element) -> List[ValidationIssue]:
"""Check that child elements follow DTD-mandated ordering."""
issues = []
for parent in root.iter():
if parent.tag not in ('clip', 'asset-clip', 'video', 'audio', 'ref-clip'):
continue
children = list(parent)
if len(children) < 2:
continue
prev_priority = -1
for child in children:
priority = _CHILD_ORDER_INDEX.get(child.tag, len(_ASSET_CLIP_CHILD_ORDER))
if priority < prev_priority:
issues.append(ValidationIssue(
issue_type=ValidationIssueType.ELEMENT_ORDER,
severity="warning",
message=(
f"<{child.tag}> appears after a higher-priority sibling "
f"in <{parent.tag}> '{parent.get('name', '')}'."
),
clip_name=parent.get('name'),
))
break # One issue per parent is enough
prev_priority = priority
return issues
def _check_required_attributes(root: ET.Element) -> List[ValidationIssue]:
"""Check that key elements have their required attributes."""
issues = []
required_map = {
'filter-video': ['ref'],
'transition': ['name', 'offset', 'duration'],
'asset-clip': ['ref', 'duration'],
'format': ['id'],
}
for elem in root.iter():
attrs = required_map.get(elem.tag)
if not attrs:
continue
for attr in attrs:
if not elem.get(attr):
issues.append(ValidationIssue(
issue_type=ValidationIssueType.MISSING_ATTRIBUTE,
severity="error",
message=f"<{elem.tag}> missing required attribute '{attr}'.",
clip_name=elem.get('name'),
))
return issues
def _check_timebases(root: ET.Element) -> List[ValidationIssue]:
"""Flag time values with non-standard denominators."""
issues = []
time_attrs = ('offset', 'start', 'duration')
seen: set = set()
for elem in root.iter():
for attr in time_attrs:
val = elem.get(attr)
if val and val.endswith('s') and '/' in val:
try:
tv = TimeValue.from_timecode(val)
denom = tv.simplify().denominator
if denom not in _FCPXML_STANDARD_TIMEBASES:
key = (elem.tag, attr, val)
if key not in seen:
seen.add(key)
issues.append(ValidationIssue(
issue_type=ValidationIssueType.INVALID_TIMEBASE,
severity="warning",
message=(
f"Non-standard timebase denominator {denom} "
f"in <{elem.tag}> {attr}=\"{val}\"."
),
clip_name=elem.get('name'),
))
except (ValueError, ZeroDivisionError):
pass
return issues
def _document_frame_duration(root: ET.Element) -> Optional[Fraction]:
"""The sequence's exact ``frameDuration`` as a fraction, if declared.
Read from the format the ``<sequence>`` references (falling back to the
first declared format), so the value is the document's own timebase
rather than an assumed rate.
"""
formats = {f.get('id'): f for f in root.findall('.//format') if f.get('id')}
sequence = root.find('.//sequence')
fmt = formats.get(sequence.get('format')) if sequence is not None else None
if fmt is None:
fmt = next(iter(formats.values()), None)
if fmt is None:
return None
raw = fmt.get('frameDuration', '')
if not (raw.endswith('s') and '/' in raw):
return None
numerator, denominator = raw[:-1].split('/', 1)
try:
value = Fraction(int(numerator), int(denominator))
except (ValueError, ZeroDivisionError):
return None
return value if value > 0 else None
def _check_frame_alignment(root: ET.Element, fps: float = 24.0) -> List[ValidationIssue]:
"""Check that durations are integer multiples of the frame duration.
Uses the document's exact ``frameDuration`` fraction and rational
arithmetic. Comparing against an integer fps instead would flag every
NTSC project as broken: at 1001/24000s (23.976fps) a perfectly aligned
duration is not an integer number of "24fps" frames, so whole timelines
would be reported misaligned when nothing is wrong.
"""
issues = []
frame_duration = _document_frame_duration(root)
label = f"{1 / float(frame_duration):.3f}".rstrip('0').rstrip('.') if frame_duration else str(fps)
for elem in root.iter():
dur_str = elem.get('duration')
if not dur_str or not dur_str.endswith('s'):
continue
if elem.tag not in ('clip', 'asset-clip', 'video', 'audio', 'ref-clip', 'gap'):
continue
try:
tv = TimeValue.from_timecode(dur_str)
if frame_duration is not None:
frames = Fraction(tv.numerator, tv.denominator) / frame_duration
aligned = frames.denominator == 1
else:
approx = tv.to_seconds() * fps
aligned = abs(approx - round(approx)) <= 0.01
if not aligned:
issues.append(ValidationIssue(
issue_type=ValidationIssueType.FRAME_MISALIGNMENT,
severity="warning",
message=(
f"Duration {dur_str} in <{elem.tag}> "
f"'{elem.get('name', '')}' is not frame-aligned at {label}fps."
),
clip_name=elem.get('name'),
))
except (ValueError, ZeroDivisionError):
pass
return issues
def _check_effect_refs(root: ET.Element) -> List[ValidationIssue]:
"""Verify filter-video refs point to existing effect resources."""
issues = []
resource_ids = set()
for res in root.iter():
rid = res.get('id')
if rid and res.tag in ('effect', 'format', 'asset', 'media'):
resource_ids.add(rid)
for fv in root.iter('filter-video'):
ref = fv.get('ref')
if ref and ref not in resource_ids:
issues.append(ValidationIssue(
issue_type=ValidationIssueType.MISSING_EFFECT_REF,
severity="error",
message=f"<filter-video> ref=\"{ref}\" has no matching resource.",
))
return issues
def _check_asset_sources(root: ET.Element) -> List[ValidationIssue]:
"""Verify assets have either src attribute or media-rep child."""
issues = []
for asset in root.iter('asset'):
src = asset.get('src', '')
media_rep = asset.find('media-rep')
if not src and media_rep is None:
issues.append(ValidationIssue(
issue_type=ValidationIssueType.MISSING_MEDIA_REP,
severity="warning",
message=(
f"<asset id=\"{asset.get('id', '?')}\" "
f"name=\"{asset.get('name', '')}\"> "
f"has no src attribute and no <media-rep> child."
),
clip_name=asset.get('name'),
))
return issues
def validate_fcpxml(root: ET.Element, fps: float = 24.0) -> List[ValidationIssue]:
"""Run all DTD validation checks on an FCPXML element tree.
Args:
root: The <fcpxml> root Element to validate.
fps: Frame rate for alignment checks (default 24).
Returns:
List of ValidationIssue objects. Empty list = clean.
"""
issues: List[ValidationIssue] = []
issues.extend(_check_child_order(root))
issues.extend(_check_required_attributes(root))
issues.extend(_check_timebases(root))
issues.extend(_check_frame_alignment(root, fps))
issues.extend(_check_effect_refs(root))
issues.extend(_check_asset_sources(root))
return issues
+3
View File
@@ -41,6 +41,9 @@ intelligence = [
transcribe = [ transcribe = [
"faster-whisper>=1.0.0", "faster-whisper>=1.0.0",
] ]
align = [
"whisperx>=3.0.0",
]
diarization = [ diarization = [
"pyannote.audio>=3.1", "pyannote.audio>=3.1",
] ]
+3
View File
@@ -154,6 +154,7 @@ from server_tools.roles import (
) )
from server_tools.subtitles import ( from server_tools.subtitles import (
handle_generate_dynamic_subtitles, handle_generate_dynamic_subtitles,
handle_generate_plain_subtitles,
handle_validate_subtitle_layout, handle_validate_subtitle_layout,
) )
from server_tools.timeline import ( from server_tools.timeline import (
@@ -178,6 +179,7 @@ from server_tools.voice import (
handle_analyze_voice_features, handle_analyze_voice_features,
handle_apply_voice_actions, handle_apply_voice_actions,
handle_build_voice_timeline, handle_build_voice_timeline,
handle_generate_voice_script,
handle_diarize_media, handle_diarize_media,
handle_get_voice_analysis_config, handle_get_voice_analysis_config,
handle_refine_voice_timeline, handle_refine_voice_timeline,
@@ -320,6 +322,7 @@ __all__ = [
"handle_save_voice_analysis_config", "handle_save_voice_analysis_config",
"handle_validate_subtitle_layout", "handle_validate_subtitle_layout",
"handle_generate_dynamic_subtitles", "handle_generate_dynamic_subtitles",
"handle_generate_plain_subtitles",
"handle_push_to_fcp", "handle_push_to_fcp",
"handle_list_fcp_libraries", "handle_list_fcp_libraries",
] ]
-829
View File
@@ -1,829 +0,0 @@
"""Shared internal helpers used by tool handlers across categories.
Extracted from server.py — validation, formatting, and small parsing utilities
that more than one server_tools/*.py module needs.
"""
from __future__ import annotations
import json
import os
import re
from pathlib import Path
from typing import Any, Sequence
from mcp.types import TextContent
from fcpxml.media_intel import media_src_to_path
from fcpxml.models import (
DuplicateGroup,
FlashFrame,
FlashFrameSeverity,
GapInfo,
Timecode,
TimeValue,
)
from fcpxml.parser import FCPXMLParser
from fcpxml.rough_cut import RoughCutGenerator
from fcpxml.transcribe import invert_ranges, merge_ranges, transcribe
from fcpxml.writer import FCPXMLModifier
PROJECTS_DIR = os.environ.get("FCP_PROJECTS_DIR", os.path.expanduser("~/Movies"))
_SANDBOX_ENABLED = "FCP_PROJECTS_DIR" in os.environ
MAX_FILE_SIZE = 100 * 1024 * 1024
MAX_MEDIA_FILE_SIZE = 32 * 1024 * 1024 * 1024
_MAX_JSON_DEPTH = 50
def _check_json_depth(obj: object, _depth: int = 0) -> None:
"""Reject JSON structures nested beyond _MAX_JSON_DEPTH.
Prevents denial-of-service via deeply nested objects that exhaust the
call stack or memory during downstream processing. Called after
json.load() since Python's json module has no built-in depth limit.
"""
if _depth > _MAX_JSON_DEPTH:
raise ValueError(
f"JSON nesting depth exceeds {_MAX_JSON_DEPTH} — "
"file may be malformed or adversarial"
)
if isinstance(obj, dict):
for v in obj.values():
_check_json_depth(v, _depth + 1)
elif isinstance(obj, list):
for item in obj:
_check_json_depth(item, _depth + 1)
def _validate_filepath(
filepath: str,
allowed_extensions: tuple[str, ...] | None = None,
max_size: int = MAX_FILE_SIZE,
) -> str:
"""Validate a user-provided file path against traversal and size attacks.
Resolves symlinks, blocks null bytes, enforces extension whitelist, and
checks file size before any parsing takes place.
``max_size`` defaults to the document limit; callers handling source
media pass ``MAX_MEDIA_FILE_SIZE``, since media is streamed rather than
parsed into memory (see the constant for why).
Raises:
ValueError: For invalid paths (null bytes, bad extensions, oversized).
FileNotFoundError: When the resolved path does not exist.
"""
if '\x00' in filepath:
raise ValueError("Invalid file path: null byte detected")
resolved = Path(filepath).resolve()
if not resolved.exists():
raise FileNotFoundError(f"File not found: {filepath}")
# .fcpxmld bundles are directories (a package wrapping Info.fcpxml plus
# sidecar data files for object tracking / Cinematic mode). The size
# check applies to the inner Info.fcpxml, which is what gets parsed.
if resolved.is_dir():
if resolved.suffix.lower() != '.fcpxmld':
raise ValueError(f"Not a regular file: {filepath}")
inner = resolved / 'Info.fcpxml'
if not inner.is_file():
raise ValueError(f"Invalid bundle (no Info.fcpxml): {filepath}")
size_target = inner
elif not resolved.is_file():
raise ValueError(f"Not a regular file: {filepath}")
else:
size_target = resolved
if allowed_extensions and resolved.suffix.lower() not in allowed_extensions:
raise ValueError(
f"Invalid file type '{resolved.suffix}'. "
f"Allowed: {', '.join(allowed_extensions)}"
)
if size_target.stat().st_size > max_size:
size_mb = size_target.stat().st_size / (1024 * 1024)
raise ValueError(f"File too large ({size_mb:.1f} MB). Maximum: {max_size // (1024 * 1024)} MB")
return str(resolved)
def _validate_output_path(output_path: str, *, anchor_dir: str | None = None) -> str:
"""Validate an output path with optional sandbox enforcement.
Resolves traversal, blocks null bytes, ensures parent exists, and — when
*anchor_dir* is provided — verifies the resolved output lives under that
directory. This prevents LLM-generated tool calls from writing to
arbitrary filesystem locations (e.g. ``/etc/cron.d/backdoor``).
Args:
output_path: The raw output path to validate.
anchor_dir: If set, the resolved output must be a child of this
directory. Typically the parent directory of the input file so
outputs stay co-located with their sources.
Raises:
ValueError: For null bytes, missing parent, or sandbox escape.
"""
if '\x00' in output_path:
raise ValueError("Invalid output path: null byte detected")
resolved = Path(output_path).resolve()
if not resolved.parent.exists():
raise ValueError(f"Output directory does not exist: {resolved.parent}")
if anchor_dir is not None:
anchor = Path(anchor_dir).resolve()
try:
resolved.relative_to(anchor)
except ValueError:
raise ValueError(
f"Output path escapes allowed directory: "
f"{resolved} is not under {anchor}"
)
return str(resolved)
def _validate_directory(directory: str, *, allowed_root: str | None = None) -> str:
"""Validate a user-provided directory path against traversal and injection.
Resolves symlinks, blocks null bytes, and verifies the path is a real
directory. When *allowed_root* is given, the resolved path must be a
descendant of (or equal to) that root — preventing filesystem enumeration
beyond the project workspace.
Raises:
ValueError: For invalid paths (null bytes, not a directory, sandbox escape).
"""
if '\x00' in directory:
raise ValueError("Invalid directory path: null byte detected")
resolved = Path(directory).resolve()
if not resolved.is_dir():
raise ValueError(f"Not a valid directory: {directory}")
if allowed_root is not None:
root = Path(allowed_root).resolve()
try:
resolved.relative_to(root)
except ValueError:
raise ValueError(
f"Directory escapes allowed root: "
f"{resolved} is not under {root}"
)
return str(resolved)
def find_fcpxml_files(directory: str) -> list[str]:
"""Find all FCPXML files in a directory."""
path = Path(directory)
files = list(str(f) for f in path.rglob("*.fcpxml"))
files.extend(str(f) for f in path.rglob("*.fcpxmld"))
return sorted(files)
def format_timecode(tc) -> str:
"""Format a Timecode object to SMPTE string."""
return tc.to_smpte() if tc else "00:00:00:00"
def format_duration(seconds: float) -> str:
"""Format seconds into human-readable duration."""
if seconds < 1:
return f"{seconds*1000:.0f}ms"
elif seconds < 60:
return f"{seconds:.2f}s"
return f"{int(seconds // 60)}m {seconds % 60:.1f}s"
def _format_clip_table(clips: list, header: str) -> str:
"""Render a list of clips as a markdown table with timecodes and durations.
Shared by handlers that filter clips by duration threshold
(find_short_cuts, find_long_clips).
"""
result = f"{header}\n\n| Name | TC | Duration |\n|------|----|---------|\n"
result += "\n".join(
f"| {c.name} | {format_timecode(c.start)} | {format_duration(c.duration_seconds)} |"
for c in clips
)
return result
def _markdown_table(headers: list[str], rows: list[list[str]]) -> str:
"""Build a markdown table from headers and rows.
Returns header row, separator row, and data rows as a single string.
Callers avoid repeating the ``| H1 | H2 |\\n|---|---|`` boilerplate
that appears in 15+ handlers.
"""
header_line = "| " + " | ".join(headers) + " |"
sep_line = "|" + "|".join("------" for _ in headers) + "|"
data_lines = "\n".join(
"| " + " | ".join(str(c) for c in row) + " |" for row in rows
)
return f"{header_line}\n{sep_line}\n{data_lines}"
def _format_batch_result(
title: str,
summary: dict[str, str],
headers: list[str],
rows: list[list[str]],
output_path: str,
) -> str:
"""Build a standard batch-operation result with summary, table, and save footer.
Used by batch fix handlers (flash frames, rapid trim, fill gaps) that all
share the same markdown structure: ``# Title → ## Summary → ## Details table
→ Saved to`` footer.
"""
summary_lines = "\n".join(f"- **{k}**: {v}" for k, v in summary.items())
table = _markdown_table(headers, rows)
return (
f"# {title}\n\n"
f"## Summary\n{summary_lines}\n\n"
f"## Details\n{table}\n\n"
f"Saved to: `{output_path}`"
)
def _fmt_suggestions(suggestions: list[str]) -> str:
"""Format pacing suggestions as markdown list (Python 3.10 compatible)."""
if not suggestions:
return "- Pacing looks good!"
nl = "\n"
return nl.join(f"- {s}" for s in suggestions)
def generate_output_path(input_path: str, suffix: str = "_modified") -> str:
"""Generate output path from input path.
The suffix is sanitized to prevent path-component injection — only
alphanumeric, hyphen, underscore, and dot characters survive.
"""
# Strip anything that could inject path separators or traversal sequences
clean_suffix = re.sub(r'[^a-zA-Z0-9._-]', '', suffix)
if not clean_suffix:
clean_suffix = "_modified"
p = Path(input_path)
return str(p.parent / f"{p.stem}{clean_suffix}{p.suffix}")
def _parse_project(filepath: str):
"""Parse an FCPXML file and return the project with its primary timeline."""
filepath = _validate_filepath(filepath, ('.fcpxml', '.fcpxmld'))
project = FCPXMLParser().parse_file(filepath)
if not project.timelines:
return None, None
return project, project.primary_timeline
def _text_result(text: str) -> list[TextContent]:
"""Wrap a string in the MCP TextContent list that every tool handler returns."""
return [TextContent(type="text", text=text)]
def _no_timeline():
"""Standard response when no timelines are found."""
return _text_result("No timelines found")
def _require_timeline(filepath: str):
"""Parse FCPXML and return (project, timeline), raising if no timeline exists.
Centralises the repeated _parse_project + _no_timeline guard that
appears in every read-only timeline handler. Returns a tuple so
callers can destructure directly::
project, tl = _require_timeline(arguments["filepath"])
"""
project, tl = _parse_project(filepath)
if not tl:
raise _NoTimelineError()
return project, tl
class _NoTimelineError(Exception):
"""Sentinel raised by _require_timeline when no timelines exist."""
def _resolve_io_paths(
arguments: dict,
suffix: str = "_modified",
) -> tuple[str, str]:
"""Validate input filepath and resolve the output path.
Shared foundation for every handler that reads an FCPXML and writes
a derived file. Validates the input, falls back to a suffixed
output name when ``output_path`` is not supplied, and sandbox-checks
the result.
Args:
arguments: Tool arguments dict (must contain ``filepath``; may
contain ``output_path``).
suffix: Default output filename suffix when ``output_path`` is
not provided (e.g. ``"_modified"``, ``"_beats"``).
Returns:
``(filepath, output_path)`` tuple with both paths validated.
"""
filepath = _validate_filepath(arguments["filepath"], ('.fcpxml', '.fcpxmld'))
# Anchor write operations to the input file's directory so LLM-generated
# tool calls cannot write to arbitrary filesystem locations (e.g.
# /etc/cron.d/backdoor). When the explicit sandbox is off, the anchor
# still prevents writes outside the source directory tree.
# `output_dir` is where the caller wants the file written, not merely a
# sandbox boundary: the app's "Pasta do projeto" promises that everything
# generated lands there. Deriving the name from the input but keeping the
# input's directory made every cross-directory call fail its own anchor
# check ("output path escapes allowed directory"), so the setting silently
# only worked when it pointed at the directory the file was already going
# to. An explicit `output_path` still wins, and still has to sit inside
# the anchor.
output_dir = arguments.get("output_dir")
if output_dir:
anchor = _validate_directory(str(output_dir))
default_output = str(Path(anchor) / Path(generate_output_path(filepath, suffix)).name)
else:
anchor = str(Path(filepath).resolve().parent)
default_output = generate_output_path(filepath, suffix)
output_path = _validate_output_path(
arguments.get("output_path") or default_output,
anchor_dir=anchor,
)
return filepath, output_path
def _setup_modifier(
arguments: dict,
suffix: str = "_modified",
) -> tuple[str, str, "FCPXMLModifier"]:
"""Common setup for write handlers: validate paths and create modifier.
Consolidates the repeated validate-filepath → resolve-output-path →
create-modifier boilerplate shared by 18+ write handlers.
Args:
arguments: Tool arguments dict (must contain ``filepath``; may
contain ``output_path``).
suffix: Default output filename suffix when ``output_path`` is
not provided (e.g. ``"_modified"``, ``"_flash_fixed"``).
Returns:
``(filepath, output_path, modifier)`` tuple ready for the
handler's domain-specific operation.
"""
filepath, output_path = _resolve_io_paths(arguments, suffix)
modifier = FCPXMLModifier(filepath)
return filepath, output_path, modifier
def _setup_generator(
arguments: dict,
suffix: str = "_roughcut",
) -> tuple[str, str, "RoughCutGenerator"]:
"""Common setup for generation handlers: validate paths and create generator.
Args:
arguments: Tool arguments dict (must contain ``filepath`` and
``output_path``).
suffix: Default output filename suffix.
Returns:
``(filepath, output_path, generator)`` tuple.
"""
filepath, output_path = _resolve_io_paths(arguments, suffix)
generator = RoughCutGenerator(filepath)
return filepath, output_path, generator
def _parse_timestamp_parts(
parts: list[str], *, frame_rate: float = 24.0
) -> float | None:
"""Convert colon-separated timestamp parts to total seconds.
Handles 2-part (M:SS), 3-part (H:MM:SS / HH:MM:SS.ms), and
4-part (HH:MM:SS:FF SMPTE) formats. Returns ``None`` when the
part count is unrecognised so callers can skip.
Args:
parts: Colon-split timestamp components.
frame_rate: FPS used to convert the frame component of SMPTE
timecodes into fractional seconds (default 24.0).
"""
if len(parts) == 2:
return int(parts[0]) * 60 + float(parts[1])
elif len(parts) == 3:
return int(parts[0]) * 3600 + int(parts[1]) * 60 + float(parts[2])
elif len(parts) == 4:
# SMPTE: HH:MM:SS:FF — convert frames to fractional seconds
base = int(parts[0]) * 3600 + int(parts[1]) * 60 + float(parts[2])
frames = int(parts[3])
return base + (frames / frame_rate) if frame_rate > 0 else base
return None
def _raw_markers_to_batch(
raw_markers: list[dict],
marker_type: str = "chapter",
max_label: int | None = None,
) -> list[dict]:
"""Convert raw {seconds, text} marker dicts to batch_add_markers format.
Shared by import_srt_markers and import_transcript_markers.
"""
batch = []
for m in raw_markers:
label = m["text"]
if max_label and len(label) > max_label:
label = label[:max_label]
batch.append({
"timecode": f"{m['seconds']}s",
"name": label,
"marker_type": marker_type.upper(),
})
return batch
def _extract_subtitle_blocks(text: str, *, strip_vtt_tags: bool = False) -> list[dict]:
"""Extract timestamp/text pairs from subtitle cue blocks (SRT or VTT).
Both SRT and VTT use the same ``start --> end`` cue syntax with
text lines underneath; only header stripping and tag cleaning differ.
"""
markers = []
blocks = re.split(r'\n\s*\n', text.strip())
for block in blocks:
lines = block.strip().split('\n')
if len(lines) < 2:
continue
ts_line = None
text_lines = []
for line in lines:
if '-->' in line:
ts_line = line
elif ts_line is not None:
if strip_vtt_tags:
line = re.sub(r'<[^>]+>', '', line)
cleaned = line.strip()
if cleaned:
text_lines.append(cleaned)
if not ts_line or not text_lines:
continue
start_str = ts_line.split('-->')[0].strip().replace(',', '.')
seconds = _parse_timestamp_parts(start_str.split(':'))
if seconds is not None:
markers.append({'seconds': seconds, 'text': ' '.join(text_lines)})
return markers
def parse_srt(text: str) -> list[dict]:
"""Parse SRT subtitle format into timestamp/text pairs."""
return _extract_subtitle_blocks(text)
def parse_vtt(text: str) -> list[dict]:
"""Parse WebVTT subtitle format into timestamp/text pairs."""
text = re.sub(r'^WEBVTT.*?\n', '', text, flags=re.MULTILINE)
text = re.sub(r'NOTE\n.*?\n\n', '', text, flags=re.DOTALL)
return _extract_subtitle_blocks(text, strip_vtt_tags=True)
def parse_transcript_timestamps(text: str) -> list[dict]:
"""Parse timestamped text (YouTube description format) into markers.
Supports formats like:
0:00 Introduction
00:01:30 Main Topic
1:05:30 Conclusion
00:00:00:00 SMPTE timecode
"""
markers = []
for line in text.strip().split('\n'):
line = line.strip()
if not line:
continue
match = re.match(r'^(\d{1,2}:\d{2}(?::\d{2}){0,2})\s+(.+)$', line)
if match:
seconds = _parse_timestamp_parts(match.group(1).split(':'))
if seconds is not None:
markers.append({'seconds': seconds, 'text': match.group(2).strip()})
return markers
def _detect_flash_frames(
tl: Any, *, critical_threshold: int = 2, warning_threshold: int = 6,
) -> list:
"""Find clips shorter than *warning_threshold* frames.
Returns a list of ``FlashFrame`` objects sorted by severity. Shared by
``handle_detect_flash_frames`` and ``handle_validate_timeline`` so the
detection logic lives in exactly one place.
"""
fps = tl.frame_rate
flash_frames: list[FlashFrame] = []
for clip in tl.clips:
duration_frames = int(clip.duration_seconds * fps)
if duration_frames < warning_threshold:
severity = (
FlashFrameSeverity.CRITICAL
if duration_frames < critical_threshold
else FlashFrameSeverity.WARNING
)
flash_frames.append(FlashFrame(
clip_name=clip.name, clip_id=clip.name,
start=clip.start, duration_frames=duration_frames,
duration_seconds=clip.duration_seconds, severity=severity,
))
return flash_frames
def _detect_gaps(tl: Any, *, min_gap_frames: int = 1) -> list:
"""Find inter-clip gaps of at least *min_gap_frames* length.
Returns a list of ``GapInfo`` objects. Shared by ``handle_detect_gaps``
and ``handle_validate_timeline``.
"""
fps = tl.frame_rate
min_gap_seconds = min_gap_frames / fps
gaps: list[GapInfo] = []
sorted_clips = sorted(tl.clips, key=lambda c: c.start.seconds)
for i in range(len(sorted_clips) - 1):
current_end = sorted_clips[i].end.seconds
next_start = sorted_clips[i + 1].start.seconds
gap_duration = next_start - current_end
if gap_duration >= min_gap_seconds:
gaps.append(GapInfo(
start=Timecode(frames=int(current_end * fps), frame_rate=fps),
duration_frames=int(gap_duration * fps),
duration_seconds=gap_duration,
previous_clip=sorted_clips[i].name,
next_clip=sorted_clips[i + 1].name,
))
return gaps
def _detect_duplicate_groups(tl: Any, *, mode: str = "same_source") -> list:
"""Group clips that share a source media reference.
Returns a list of ``DuplicateGroup`` objects. Shared by
``handle_detect_duplicates`` and ``handle_validate_timeline``.
"""
source_groups: dict[str, list[dict]] = {}
for clip in tl.clips:
source_key = clip.media_path or clip.name
if source_key not in source_groups:
source_groups[source_key] = []
source_groups[source_key].append({
'name': clip.name,
'start': clip.start.seconds,
'duration': clip.duration_seconds,
'source_start': clip.source_start.seconds if clip.source_start else 0,
'source_duration': clip.duration_seconds,
'timecode': format_timecode(clip.start),
})
duplicates: list[DuplicateGroup] = []
for source_key, clips in source_groups.items():
if len(clips) <= 1:
continue
group = DuplicateGroup(
source_ref=source_key,
source_name=source_key.split('/')[-1] if '/' in source_key else source_key,
clips=clips,
)
if mode == "same_source":
duplicates.append(group)
elif mode == "overlapping_ranges" and group.has_overlapping_ranges:
duplicates.append(group)
elif mode == "identical":
seen_ranges: set[tuple] = set()
identical_clips = []
for c in clips:
range_key = (c['source_start'], c['source_duration'])
if range_key in seen_ranges:
identical_clips.append(c)
seen_ranges.add(range_key)
if identical_clips:
group.clips = identical_clips
duplicates.append(group)
return duplicates
AUDIO_MEDIA_EXTENSIONS = (
'.wav', '.aif', '.aiff', '.mp3', '.m4a', '.aac', '.flac', '.mov', '.mp4',
)
_DIARIZATION_INSTALL_HINT = (
"\n\nInstall the optional diarization extra:\n\n"
" pip install 'fcp-mcp-server[diarization]'\n\n"
"and set a HuggingFace token with access to "
"pyannote/speaker-diarization-3.1 (pass hf_token= or persist one via "
"save_hf_token)."
)
_FEATURES_INSTALL_HINT = (
"\n\nInstall the optional media-intelligence extra:\n\n"
" pip install 'fcp-mcp-server[intelligence]'"
)
def _voice_analysis_config_text(config: dict) -> str:
w = config["emphasis_weights"]
text = "# Voice Analysis Settings\n\n"
text += _markdown_table(
["Setting", "Value"],
[
["Energy threshold", f"{config['energy_threshold']:.2f}"],
["Peak selection", f"top {config['peak_percentile']:.1%} of words"],
["Emphasis floor", f"{config['emphasis_floor']:.2f}"],
["Emotion detection", "on" if config["emotion_enabled"] else "off"],
["Emotion sensitivity", f"{config['emotion_sensitivity']:.2f}"],
],
) + "\n\n## Emphasis Weights\n"
text += _markdown_table(
["Factor", "Weight"],
[[k.replace("_", " ").title(), f"{v:.2f}"] for k, v in w.items()],
)
return text
def _apply_placed_action(modifier, clip_el, action, clip_start: float) -> str:
"""Apply one non-cut action to the clip that hosts it.
``clip_start`` is where that clip begins on the timeline; the writer
wants times relative to the clip's own head, so the rebase happens here
— the single place that knows about the conversion. The clip *element*
is passed through rather than its name: after a cut the pieces share a
name, and a name lookup would land every edit on the first piece.
"""
rel_start = action.start - clip_start
rel_end = action.end - clip_start
if action.kind == "zoom":
# Only forward an explicit ease — otherwise add_zoom's own default
# (a fast ramp in, instant snap back out) is what should apply.
zoom_args = {}
if action.params.get("ease") is not None:
zoom_args["ease"] = float(action.params["ease"])
if action.params.get("ease_out") is not None:
zoom_args["ease_out"] = float(action.params["ease_out"])
modifier.add_zoom(
clip_id=clip_el,
start=rel_start,
end=rel_end,
scale=float(action.params.get("scale", 1.3)),
**zoom_args,
)
return f"zoom {action.params.get('scale', 1.3):.2f}x"
if action.kind == "text":
modifier.add_text_title(
clip_el,
action.params["content"],
offset=modifier.snap_seconds_to_frame(rel_start).to_fcpxml(),
duration=modifier.snap_seconds_to_frame(action.duration).to_fcpxml(),
)
return f"text \"{action.params['content'][:24]}\""
# marker
modifier.add_marker(
clip_id=clip_el,
timecode=modifier.snap_seconds_to_frame(rel_start).to_fcpxml(),
name=action.params.get("content") or action.reason or "Voice action",
note=action.reason or None,
)
return "marker"
def _speaker_table(profiles: Sequence[dict]) -> str:
"""Who was detected, ordered by how much of the runtime each holds."""
return _markdown_table(
["ID", "Name", "Share", "Speaking", "Lines", "Avg line"],
[
[
p["id"],
p.get("name", ""),
f"{p['share']:.0%}",
format_duration(p["speaking_seconds"]),
str(p["segment_count"]),
f"{p['avg_segment']:.1f}s",
]
for p in profiles
],
)
TRANSCRIBE_MAX_MEDIA = 10
_TRANSCRIBE_INSTALL_HINT = (
"\n\nInstall the optional transcription extra:\n\n"
" pip install 'fcp-mcp-server[transcribe]'\n\n"
"or run via uvx:\n\n"
" uvx --from \"fcp-mcp-server[transcribe]\" fcp-mcp-server"
)
def _transcript_json_path(media_path: str, output_dir: str | None = None) -> Path:
"""Where the ``_transcript.json`` for ``media_path`` lives.
When ``output_dir`` (the user-selected project folder) is set, the
transcript is saved/read there instead of next to the source media.
"""
p = Path(media_path)
if output_dir:
directory = Path(output_dir).expanduser()
directory.mkdir(parents=True, exist_ok=True)
return directory / f"{p.stem}_transcript.json"
return p.with_name(p.stem + "_transcript.json")
def _load_or_transcribe(
media_path: str, model: str, language: str | None, output_dir: str | None = None
) -> tuple[dict | None, str]:
"""Load a cached ``_transcript.json`` for a media file, else transcribe and cache it.
Returns ``(transcript, "")`` or ``(None, reason)``. The cache makes
transcription a one-time cost per media file across all transcript tools.
"""
json_path = _transcript_json_path(media_path, output_dir)
if json_path.is_file():
try:
with open(json_path) as f:
data = json.load(f)
if isinstance(data, dict) and isinstance(data.get("words"), list):
return data, ""
except (OSError, json.JSONDecodeError, UnicodeDecodeError):
pass # unreadable cache falls through to re-transcribe
result = transcribe(media_path, model_size=model, language=language)
if result is None:
return None, "untranscribable (faster-whisper not installed or media unreadable)"
anchor = str(Path(output_dir).expanduser()) if output_dir else str(Path(media_path).parent)
out_path = _validate_output_path(str(json_path), anchor_dir=anchor)
with open(out_path, "w") as f:
json.dump({"source": Path(media_path).name, **result}, f, indent=2)
return result, ""
def _cut_transcript_spans(modifier, clip_filter, model, language, padding, spans_fn, keep_only=False, output_dir=None):
"""Shared cut engine for transcript-driven editing.
``spans_fn(words) -> [(start, end), ...]`` in source seconds. Spans are
padded, clamped to each clip's used source window, optionally inverted
(keep_only), snapped to the frame grid, and cut with ripple.
"""
to_frame = modifier.snap_seconds_to_frame
cache: dict[str, tuple] = {}
cuts_made: list[tuple[str, int, float]] = []
skipped: list[tuple[str, str]] = []
spine_clips = [el for _, el in modifier._iter_spine_clips()]
for el in spine_clips:
name = el.get("name", "")
if clip_filter and name != clip_filter:
continue
src = modifier.resources.get(el.get("ref", ""), {}).get("src", "")
media_path = media_src_to_path(src)
if not media_path or not Path(media_path).is_file():
skipped.append((name, "media file missing"))
continue
if media_path not in cache:
if len(cache) >= TRANSCRIBE_MAX_MEDIA:
skipped.append((name, f"transcription cap reached ({TRANSCRIBE_MAX_MEDIA} media files)"))
continue
cache[media_path] = _load_or_transcribe(media_path, model, language, output_dir)
data, reason = cache[media_path]
if data is None:
skipped.append((name, reason))
continue
clip_source_start = modifier.source_file_start(el).to_seconds()
clip_duration = modifier._parse_time(el.get("duration", "0s")).to_seconds()
window_start = clip_source_start
window_end = clip_source_start + clip_duration
spans = spans_fn(data.get("words", []))
padded = merge_ranges([(s - padding, e + padding) for s, e in spans])
clamped = [
(max(s, window_start), min(e, window_end))
for s, e in padded
if min(e, window_end) > max(s, window_start)
]
if keep_only:
if not clamped:
# Never delete a whole clip just because nothing matched in it.
skipped.append((name, "no phrase matches — left untouched (keep_only)"))
continue
cut_source = invert_ranges(clamped, window_start, window_end)
else:
cut_source = clamped
cut_ranges = [
(to_frame(s - clip_source_start), to_frame(e - clip_source_start))
for s, e in cut_source
]
cut_ranges = [(a, b) for a, b in cut_ranges if b > a]
if not cut_ranges:
continue
removed = modifier.cut_clip_ranges(el, cut_ranges)
if removed > TimeValue.zero():
cuts_made.append((name, len(cut_ranges), removed.to_seconds()))
return cuts_made, skipped
def _transcript_cut_report(title, summary_lines, cuts_made, skipped, output_path, footer):
if not cuts_made:
text = f"# {title}\n\nNo cuts to make — file unchanged (nothing saved)."
if skipped:
text += "\n\n## Skipped Clips\n" + _markdown_table(
["Clip", "Reason"], [[name, reason] for name, reason in skipped]
)
if any("faster-whisper" in reason for _, reason in skipped):
text += _TRANSCRIBE_INSTALL_HINT
return _text_result(text)
total_removed = sum(seconds for _, _, seconds in cuts_made)
result = f"# {title}\n\n## Summary\n"
result += "\n".join(summary_lines) + "\n"
result += f"- **Clips Cut**: {len(cuts_made)}\n- **Total Removed**: {format_duration(total_removed)}\n"
result += "\n## Cuts\n"
result += _markdown_table(
["Clip", "Ranges Cut", "Removed"],
[[name, str(count), f"{seconds:.2f}s"] for name, count, seconds in cuts_made],
) + "\n"
if skipped:
result += "\n## Skipped Clips\n" + _markdown_table(
["Clip", "Reason"], [[name, reason] for name, reason in skipped]
) + "\n"
result += f"\nSaved to: {output_path}\n\n{footer}"
return _text_result(result)
+124
View File
@@ -0,0 +1,124 @@
"""Shared internal helpers used by tool handlers across categories.
Extracted from server.py — validation, formatting, and small parsing utilities
that more than one server_tools/*.py module needs.
Eram 882 linhas de seis papéis diferentes sob um nome que só dizia
"compartilhado". Cada papel virou um módulo; este pacote reexporta tudo, então
os treze pontos que importam daqui seguem iguais.
paths validação contra a sandbox, limites, caminho de saída
formatting tabelas e relatórios devolvidos pelos handlers
project abrir projeto, preparar modifier/generator
captions SRT, VTT e listas com timestamp
detection flash frames, buracos, duplicados
media transcrição em cache, corte por fala, ações posicionadas
"""
from .captions import (
_extract_subtitle_blocks,
_parse_timestamp_parts,
_raw_markers_to_batch,
parse_srt,
parse_transcript_timestamps,
parse_vtt,
)
from .detection import (
_detect_duplicate_groups,
_detect_flash_frames,
_detect_gaps,
)
from .formatting import (
_fmt_suggestions,
_format_batch_result,
_format_clip_table,
_markdown_table,
_speaker_table,
_voice_analysis_config_text,
format_duration,
format_timecode,
)
from .media import (
_DIARIZATION_INSTALL_HINT,
_FEATURES_INSTALL_HINT,
_TRANSCRIBE_INSTALL_HINT,
AUDIO_MEDIA_EXTENSIONS,
TRANSCRIBE_MAX_MEDIA,
_apply_placed_action,
_cut_transcript_spans,
_load_or_transcribe,
_transcript_cut_report,
_transcript_json_path,
)
from .paths import (
_MAX_JSON_DEPTH,
_SANDBOX_ENABLED,
MAX_FILE_SIZE,
MAX_MEDIA_FILE_SIZE,
PROJECTS_DIR,
_check_json_depth,
_resolve_io_paths,
_validate_directory,
_validate_filepath,
_validate_output_path,
find_fcpxml_files,
generate_output_path,
)
from .project import (
_no_timeline,
_NoTimelineError,
_parse_project,
_require_timeline,
_setup_generator,
_setup_modifier,
_text_result,
)
__all__ = [
"AUDIO_MEDIA_EXTENSIONS",
"MAX_FILE_SIZE",
"MAX_MEDIA_FILE_SIZE",
"PROJECTS_DIR",
"TRANSCRIBE_MAX_MEDIA",
"_DIARIZATION_INSTALL_HINT",
"_FEATURES_INSTALL_HINT",
"_MAX_JSON_DEPTH",
"_NoTimelineError",
"_SANDBOX_ENABLED",
"_TRANSCRIBE_INSTALL_HINT",
"_apply_placed_action",
"_check_json_depth",
"_cut_transcript_spans",
"_detect_duplicate_groups",
"_detect_flash_frames",
"_detect_gaps",
"_extract_subtitle_blocks",
"_fmt_suggestions",
"_format_batch_result",
"_format_clip_table",
"_load_or_transcribe",
"_markdown_table",
"_no_timeline",
"_parse_project",
"_parse_timestamp_parts",
"_raw_markers_to_batch",
"_require_timeline",
"_resolve_io_paths",
"_setup_generator",
"_setup_modifier",
"_speaker_table",
"_text_result",
"_transcript_cut_report",
"_transcript_json_path",
"_validate_directory",
"_validate_filepath",
"_validate_output_path",
"_voice_analysis_config_text",
"find_fcpxml_files",
"format_duration",
"format_timecode",
"generate_output_path",
"parse_srt",
"parse_transcript_timestamps",
"parse_vtt",
]
+117
View File
@@ -0,0 +1,117 @@
"""Leitura de legendas e listas com timestamp (SRT, VTT, texto colado).
Extraído de _shared.py — ver server_tools/_shared/__init__.py.
"""
from __future__ import annotations
import re
def _parse_timestamp_parts(
parts: list[str], *, frame_rate: float = 24.0
) -> float | None:
"""Convert colon-separated timestamp parts to total seconds.
Handles 2-part (M:SS), 3-part (H:MM:SS / HH:MM:SS.ms), and
4-part (HH:MM:SS:FF SMPTE) formats. Returns ``None`` when the
part count is unrecognised so callers can skip.
Args:
parts: Colon-split timestamp components.
frame_rate: FPS used to convert the frame component of SMPTE
timecodes into fractional seconds (default 24.0).
"""
if len(parts) == 2:
return int(parts[0]) * 60 + float(parts[1])
elif len(parts) == 3:
return int(parts[0]) * 3600 + int(parts[1]) * 60 + float(parts[2])
elif len(parts) == 4:
# SMPTE: HH:MM:SS:FF — convert frames to fractional seconds
base = int(parts[0]) * 3600 + int(parts[1]) * 60 + float(parts[2])
frames = int(parts[3])
return base + (frames / frame_rate) if frame_rate > 0 else base
return None
def _raw_markers_to_batch(
raw_markers: list[dict],
marker_type: str = "chapter",
max_label: int | None = None,
) -> list[dict]:
"""Convert raw {seconds, text} marker dicts to batch_add_markers format.
Shared by import_srt_markers and import_transcript_markers.
"""
batch = []
for m in raw_markers:
label = m["text"]
if max_label and len(label) > max_label:
label = label[:max_label]
batch.append({
"timecode": f"{m['seconds']}s",
"name": label,
"marker_type": marker_type.upper(),
})
return batch
def _extract_subtitle_blocks(text: str, *, strip_vtt_tags: bool = False) -> list[dict]:
"""Extract timestamp/text pairs from subtitle cue blocks (SRT or VTT).
Both SRT and VTT use the same ``start --> end`` cue syntax with
text lines underneath; only header stripping and tag cleaning differ.
"""
markers = []
blocks = re.split(r'\n\s*\n', text.strip())
for block in blocks:
lines = block.strip().split('\n')
if len(lines) < 2:
continue
ts_line = None
text_lines = []
for line in lines:
if '-->' in line:
ts_line = line
elif ts_line is not None:
if strip_vtt_tags:
line = re.sub(r'<[^>]+>', '', line)
cleaned = line.strip()
if cleaned:
text_lines.append(cleaned)
if not ts_line or not text_lines:
continue
start_str = ts_line.split('-->')[0].strip().replace(',', '.')
seconds = _parse_timestamp_parts(start_str.split(':'))
if seconds is not None:
markers.append({'seconds': seconds, 'text': ' '.join(text_lines)})
return markers
def parse_srt(text: str) -> list[dict]:
"""Parse SRT subtitle format into timestamp/text pairs."""
return _extract_subtitle_blocks(text)
def parse_vtt(text: str) -> list[dict]:
"""Parse WebVTT subtitle format into timestamp/text pairs."""
text = re.sub(r'^WEBVTT.*?\n', '', text, flags=re.MULTILINE)
text = re.sub(r'NOTE\n.*?\n\n', '', text, flags=re.DOTALL)
return _extract_subtitle_blocks(text, strip_vtt_tags=True)
def parse_transcript_timestamps(text: str) -> list[dict]:
"""Parse timestamped text (YouTube description format) into markers.
Supports formats like:
0:00 Introduction
00:01:30 Main Topic
1:05:30 Conclusion
00:00:00:00 SMPTE timecode
"""
markers = []
for line in text.strip().split('\n'):
line = line.strip()
if not line:
continue
match = re.match(r'^(\d{1,2}:\d{2}(?::\d{2}){0,2})\s+(.+)$', line)
if match:
seconds = _parse_timestamp_parts(match.group(1).split(':'))
if seconds is not None:
markers.append({'seconds': seconds, 'text': match.group(2).strip()})
return markers
+115
View File
@@ -0,0 +1,115 @@
"""Detecção para QC: flash frames, buracos e clipes duplicados.
Extraído de _shared.py — ver server_tools/_shared/__init__.py.
"""
from __future__ import annotations
from typing import Any
from fcpxml.models import (
DuplicateGroup,
FlashFrame,
FlashFrameSeverity,
GapInfo,
Timecode,
)
from .formatting import format_timecode
def _detect_flash_frames(
tl: Any, *, critical_threshold: int = 2, warning_threshold: int = 6,
) -> list:
"""Find clips shorter than *warning_threshold* frames.
Returns a list of ``FlashFrame`` objects sorted by severity. Shared by
``handle_detect_flash_frames`` and ``handle_validate_timeline`` so the
detection logic lives in exactly one place.
"""
fps = tl.frame_rate
flash_frames: list[FlashFrame] = []
for clip in tl.clips:
duration_frames = int(clip.duration_seconds * fps)
if duration_frames < warning_threshold:
severity = (
FlashFrameSeverity.CRITICAL
if duration_frames < critical_threshold
else FlashFrameSeverity.WARNING
)
flash_frames.append(FlashFrame(
clip_name=clip.name, clip_id=clip.name,
start=clip.start, duration_frames=duration_frames,
duration_seconds=clip.duration_seconds, severity=severity,
))
return flash_frames
def _detect_gaps(tl: Any, *, min_gap_frames: int = 1) -> list:
"""Find inter-clip gaps of at least *min_gap_frames* length.
Returns a list of ``GapInfo`` objects. Shared by ``handle_detect_gaps``
and ``handle_validate_timeline``.
"""
fps = tl.frame_rate
min_gap_seconds = min_gap_frames / fps
gaps: list[GapInfo] = []
sorted_clips = sorted(tl.clips, key=lambda c: c.start.seconds)
for i in range(len(sorted_clips) - 1):
current_end = sorted_clips[i].end.seconds
next_start = sorted_clips[i + 1].start.seconds
gap_duration = next_start - current_end
if gap_duration >= min_gap_seconds:
gaps.append(GapInfo(
start=Timecode(frames=int(current_end * fps), frame_rate=fps),
duration_frames=int(gap_duration * fps),
duration_seconds=gap_duration,
previous_clip=sorted_clips[i].name,
next_clip=sorted_clips[i + 1].name,
))
return gaps
def _detect_duplicate_groups(tl: Any, *, mode: str = "same_source") -> list:
"""Group clips that share a source media reference.
Returns a list of ``DuplicateGroup`` objects. Shared by
``handle_detect_duplicates`` and ``handle_validate_timeline``.
"""
source_groups: dict[str, list[dict]] = {}
for clip in tl.clips:
source_key = clip.media_path or clip.name
if source_key not in source_groups:
source_groups[source_key] = []
source_groups[source_key].append({
'name': clip.name,
'start': clip.start.seconds,
'duration': clip.duration_seconds,
'source_start': clip.source_start.seconds if clip.source_start else 0,
'source_duration': clip.duration_seconds,
'timecode': format_timecode(clip.start),
})
duplicates: list[DuplicateGroup] = []
for source_key, clips in source_groups.items():
if len(clips) <= 1:
continue
group = DuplicateGroup(
source_ref=source_key,
source_name=source_key.split('/')[-1] if '/' in source_key else source_key,
clips=clips,
)
if mode == "same_source":
duplicates.append(group)
elif mode == "overlapping_ranges" and group.has_overlapping_ranges:
duplicates.append(group)
elif mode == "identical":
seen_ranges: set[tuple] = set()
identical_clips = []
for c in clips:
range_key = (c['source_start'], c['source_duration'])
if range_key in seen_ranges:
identical_clips.append(c)
seen_ranges.add(range_key)
if identical_clips:
group.clips = identical_clips
duplicates.append(group)
return duplicates
+114
View File
@@ -0,0 +1,114 @@
"""Formatação do texto que os handlers devolvem — tabelas e relatórios.
Extraído de _shared.py — ver server_tools/_shared/__init__.py.
"""
from __future__ import annotations
from typing import Sequence
def format_timecode(tc) -> str:
"""Format a Timecode object to SMPTE string."""
return tc.to_smpte() if tc else "00:00:00:00"
def format_duration(seconds: float) -> str:
"""Format seconds into human-readable duration."""
if seconds < 1:
return f"{seconds*1000:.0f}ms"
elif seconds < 60:
return f"{seconds:.2f}s"
return f"{int(seconds // 60)}m {seconds % 60:.1f}s"
def _format_clip_table(clips: list, header: str) -> str:
"""Render a list of clips as a markdown table with timecodes and durations.
Shared by handlers that filter clips by duration threshold
(find_short_cuts, find_long_clips).
"""
result = f"{header}\n\n| Name | TC | Duration |\n|------|----|---------|\n"
result += "\n".join(
f"| {c.name} | {format_timecode(c.start)} | {format_duration(c.duration_seconds)} |"
for c in clips
)
return result
def _markdown_table(headers: list[str], rows: list[list[str]]) -> str:
"""Build a markdown table from headers and rows.
Returns header row, separator row, and data rows as a single string.
Callers avoid repeating the ``| H1 | H2 |\\n|---|---|`` boilerplate
that appears in 15+ handlers.
"""
header_line = "| " + " | ".join(headers) + " |"
sep_line = "|" + "|".join("------" for _ in headers) + "|"
data_lines = "\n".join(
"| " + " | ".join(str(c) for c in row) + " |" for row in rows
)
return f"{header_line}\n{sep_line}\n{data_lines}"
def _format_batch_result(
title: str,
summary: dict[str, str],
headers: list[str],
rows: list[list[str]],
output_path: str,
) -> str:
"""Build a standard batch-operation result with summary, table, and save footer.
Used by batch fix handlers (flash frames, rapid trim, fill gaps) that all
share the same markdown structure: ``# Title → ## Summary → ## Details table
→ Saved to`` footer.
"""
summary_lines = "\n".join(f"- **{k}**: {v}" for k, v in summary.items())
table = _markdown_table(headers, rows)
return (
f"# {title}\n\n"
f"## Summary\n{summary_lines}\n\n"
f"## Details\n{table}\n\n"
f"Saved to: `{output_path}`"
)
def _fmt_suggestions(suggestions: list[str]) -> str:
"""Format pacing suggestions as markdown list (Python 3.10 compatible)."""
if not suggestions:
return "- Pacing looks good!"
nl = "\n"
return nl.join(f"- {s}" for s in suggestions)
def _voice_analysis_config_text(config: dict) -> str:
w = config["emphasis_weights"]
text = "# Voice Analysis Settings\n\n"
text += _markdown_table(
["Setting", "Value"],
[
["Energy threshold", f"{config['energy_threshold']:.2f}"],
["Peak selection", f"top {config['peak_percentile']:.1%} of words"],
["Emphasis floor", f"{config['emphasis_floor']:.2f}"],
["Emotion detection", "on" if config["emotion_enabled"] else "off"],
["Emotion sensitivity", f"{config['emotion_sensitivity']:.2f}"],
],
) + "\n\n## Emphasis Weights\n"
text += _markdown_table(
["Factor", "Weight"],
[[k.replace("_", " ").title(), f"{v:.2f}"] for k, v in w.items()],
)
return text
def _speaker_table(profiles: Sequence[dict]) -> str:
"""Who was detected, ordered by how much of the runtime each holds."""
return _markdown_table(
["ID", "Name", "Share", "Speaking", "Lines", "Avg line"],
[
[
p["id"],
p.get("name", ""),
f"{p['share']:.0%}",
format_duration(p["speaking_seconds"]),
str(p["segment_count"]),
f"{p['avg_segment']:.1f}s",
]
for p in profiles
],
)
+274
View File
@@ -0,0 +1,274 @@
"""Mídia e transcrição: cache, corte por trecho falado e ações posicionadas.
Extraído de _shared.py — ver server_tools/_shared/__init__.py.
"""
from __future__ import annotations
import json
from pathlib import Path
from fcpxml.media_intel import media_src_to_path
from fcpxml.model_manager import load_dynamic_subtitle_config, load_voice_analysis_config
from fcpxml.models import (
TimeValue,
)
from fcpxml.text_layout import TEXT_TEMPLATE_FONT_SCALE, measure_text
from fcpxml.transcribe import invert_ranges, merge_ranges, transcribe
from .formatting import _markdown_table, format_duration
from .paths import _validate_output_path
from .project import _text_result
AUDIO_MEDIA_EXTENSIONS = (
'.wav', '.aif', '.aiff', '.mp3', '.m4a', '.aac', '.flac', '.mov', '.mp4',
)
_DIARIZATION_INSTALL_HINT = (
"\n\nInstall the optional diarization extra:\n\n"
" pip install 'fcp-mcp-server[diarization]'\n\n"
"and set a HuggingFace token with access to "
"pyannote/speaker-diarization-3.1 (pass hf_token= or persist one via "
"save_hf_token)."
)
_FEATURES_INSTALL_HINT = (
"\n\nInstall the optional media-intelligence extra:\n\n"
" pip install 'fcp-mcp-server[intelligence]'"
)
def _apply_placed_action(modifier, clip_el, action, clip_start: float) -> str:
"""Apply one non-cut action to the clip that hosts it.
``clip_start`` is where that clip begins on the timeline; the writer
wants times relative to the clip's own head, so the rebase happens here
— the single place that knows about the conversion. The clip *element*
is passed through rather than its name: after a cut the pieces share a
name, and a name lookup would land every edit on the first piece.
"""
rel_start = action.start - clip_start
rel_end = action.end - clip_start
if action.kind == "zoom":
config = load_voice_analysis_config()
# Only forward an explicit ease — otherwise add_zoom's own default
# (a fast ramp in, instant snap back out) is what should apply.
zoom_args = {
"ease": float(action.params.get("ease", config["zoom_ease_in"])),
"ease_out": float(action.params.get("ease_out", config["zoom_ease_out"])),
}
mode = str(action.params.get("mode", config["zoom_mode"]))
if mode == "in":
zoom_args["hold_at_end"] = True
zoom_args["start_at_peak"] = False
elif mode == "out":
zoom_args["hold_at_end"] = False
zoom_args["start_at_peak"] = True
elif mode == "in_out":
zoom_args["hold_at_end"] = False
zoom_args["start_at_peak"] = False
modifier.add_zoom(
clip_id=clip_el,
start=rel_start,
end=rel_end,
scale=float(action.params.get("scale", config["zoom_scale"])),
**zoom_args,
)
return f"zoom {float(action.params.get('scale', config['zoom_scale'])):.2f}x"
if action.kind == "text":
# Default to the "Legendas Dinâmicas" emphasis style (the font used
# to highlight a word in the captions) rather than a hardcoded
# Helvetica Neue, so a callout like "MASTOPEXIA" matches the rest of
# the video's on-screen text instead of looking like a stray default
# title. Any of these the action itself specifies still wins.
subtitle_cfg = load_dynamic_subtitle_config()
font = action.params.get("font", subtitle_cfg["emphasis_font"])
face = action.params.get("face", subtitle_cfg["emphasis_face"])
font_scale = float(subtitle_cfg.get("text_scale", TEXT_TEMPLATE_FONT_SCALE) or 1.0)
requested_size = int(action.params.get("font_size", subtitle_cfg["emphasis_size"]))
requested_kerning = float(action.params.get("kerning", 0.0) or 0.0)
# Voice-action callouts are not part of the dynamic subtitle block.
# When omitted, put them above the subtitle band and shrink wide
# phrases to the title-safe width. The previous default (Position 0 0,
# full emphasis size) made long callouts like "PRÓTESES DE SILICONE"
# collide with captions and run off both sides of a vertical frame.
emitted_size = requested_size * font_scale
emitted_kerning = requested_kerning * font_scale
safe_width = modifier.frame_width() * 0.90
width = measure_text(
action.params["content"],
emitted_size,
bold=bool(action.params.get("bold", False)),
kerning=emitted_kerning,
font=font,
face=face,
)
font_size = requested_size
if width > safe_width and width > 0:
font_size = max(32, int(requested_size * safe_width / width))
position = action.params.get("position")
if not position:
position = f"0 {modifier.frame_height() * 0.23:g}"
modifier.add_text_title(
clip_el,
action.params["content"],
offset=modifier.snap_seconds_to_frame(rel_start).to_fcpxml(),
duration=modifier.snap_seconds_to_frame(action.duration).to_fcpxml(),
position=position,
font=font,
font_size=font_size,
font_color=action.params.get("font_color", subtitle_cfg["emphasis_color"]),
face=face,
bold=action.params.get("bold", False),
)
return f"text \"{action.params['content'][:24]}\""
# marker
modifier.add_marker(
clip_id=clip_el,
timecode=modifier.snap_seconds_to_frame(rel_start).to_fcpxml(),
name=action.params.get("content") or action.reason or "Voice action",
note=action.reason or None,
)
return "marker"
TRANSCRIBE_MAX_MEDIA = 10
_TRANSCRIBE_INSTALL_HINT = (
"\n\nInstall the optional transcription extra:\n\n"
" pip install 'fcp-mcp-server[transcribe]'\n\n"
"or run via uvx:\n\n"
" uvx --from \"fcp-mcp-server[transcribe]\" fcp-mcp-server"
)
def _transcript_json_path(media_path: str, output_dir: str | None = None) -> Path:
"""Where the ``_transcript.json`` for ``media_path`` lives.
When ``output_dir`` (the user-selected project folder) is set, the
transcript is saved/read there instead of next to the source media.
"""
p = Path(media_path)
if output_dir:
directory = Path(output_dir).expanduser()
directory.mkdir(parents=True, exist_ok=True)
return directory / f"{p.stem}_transcript.json"
return p.with_name(p.stem + "_transcript.json")
def _load_or_transcribe(
media_path: str, model: str, language: str | None, output_dir: str | None = None
) -> tuple[dict | None, str]:
"""Load a cached ``_transcript.json`` for a media file, else transcribe and cache it.
Returns ``(transcript, "")`` or ``(None, reason)``. The cache makes
transcription a one-time cost per media file across all transcript tools.
"""
json_path = _transcript_json_path(media_path, output_dir)
if json_path.is_file():
try:
with open(json_path) as f:
data = json.load(f)
if isinstance(data, dict) and isinstance(data.get("words"), list):
return data, ""
except (OSError, json.JSONDecodeError, UnicodeDecodeError):
pass # unreadable cache falls through to re-transcribe
result = transcribe(media_path, model_size=model, language=language)
if result is None:
return None, "untranscribable (faster-whisper not installed or media unreadable)"
anchor = str(Path(output_dir).expanduser()) if output_dir else str(Path(media_path).parent)
out_path = _validate_output_path(str(json_path), anchor_dir=anchor)
with open(out_path, "w") as f:
json.dump({"source": Path(media_path).name, **result}, f, indent=2)
return result, ""
def _cut_transcript_spans(modifier, clip_filter, model, language, padding, spans_fn, keep_only=False, output_dir=None):
"""Shared cut engine for transcript-driven editing.
``spans_fn(words) -> [(start, end), ...]`` in source seconds. Spans are
padded, clamped to each clip's used source window, optionally inverted
(keep_only), snapped to the frame grid, and cut with ripple.
"""
to_frame = modifier.snap_seconds_to_frame
cache: dict[str, tuple] = {}
cuts_made: list[tuple[str, int, float]] = []
skipped: list[tuple[str, str]] = []
spine_clips = [el for _, el in modifier._iter_spine_clips()]
for el in spine_clips:
name = el.get("name", "")
if clip_filter and name != clip_filter:
continue
src = modifier.resources.get(el.get("ref", ""), {}).get("src", "")
media_path = media_src_to_path(src)
if not media_path or not Path(media_path).is_file():
skipped.append((name, "media file missing"))
continue
if media_path not in cache:
if len(cache) >= TRANSCRIBE_MAX_MEDIA:
skipped.append((name, f"transcription cap reached ({TRANSCRIBE_MAX_MEDIA} media files)"))
continue
cache[media_path] = _load_or_transcribe(media_path, model, language, output_dir)
data, reason = cache[media_path]
if data is None:
skipped.append((name, reason))
continue
clip_source_start = modifier.source_file_start(el).to_seconds()
clip_duration = modifier._parse_time(el.get("duration", "0s")).to_seconds()
window_start = clip_source_start
window_end = clip_source_start + clip_duration
spans = spans_fn(data.get("words", []))
padded = merge_ranges([(s - padding, e + padding) for s, e in spans])
clamped = [
(max(s, window_start), min(e, window_end))
for s, e in padded
if min(e, window_end) > max(s, window_start)
]
if keep_only:
if not clamped:
# Never delete a whole clip just because nothing matched in it.
skipped.append((name, "no phrase matches — left untouched (keep_only)"))
continue
cut_source = invert_ranges(clamped, window_start, window_end)
else:
cut_source = clamped
cut_ranges = [
(to_frame(s - clip_source_start), to_frame(e - clip_source_start))
for s, e in cut_source
]
cut_ranges = [(a, b) for a, b in cut_ranges if b > a]
if not cut_ranges:
continue
removed = modifier.cut_clip_ranges(el, cut_ranges)
if removed > TimeValue.zero():
cuts_made.append((name, len(cut_ranges), removed.to_seconds()))
return cuts_made, skipped
def _transcript_cut_report(title, summary_lines, cuts_made, skipped, output_path, footer):
if not cuts_made:
text = f"# {title}\n\nNo cuts to make — file unchanged (nothing saved)."
if skipped:
text += "\n\n## Skipped Clips\n" + _markdown_table(
["Clip", "Reason"], [[name, reason] for name, reason in skipped]
)
if any("faster-whisper" in reason for _, reason in skipped):
text += _TRANSCRIBE_INSTALL_HINT
return _text_result(text)
total_removed = sum(seconds for _, _, seconds in cuts_made)
result = f"# {title}\n\n## Summary\n"
result += "\n".join(summary_lines) + "\n"
result += f"- **Clips Cut**: {len(cuts_made)}\n- **Total Removed**: {format_duration(total_removed)}\n"
result += "\n## Cuts\n"
result += _markdown_table(
["Clip", "Ranges Cut", "Removed"],
[[name, str(count), f"{seconds:.2f}s"] for name, count, seconds in cuts_made],
) + "\n"
if skipped:
result += "\n## Skipped Clips\n" + _markdown_table(
["Clip", "Reason"], [[name, reason] for name, reason in skipped]
) + "\n"
result += f"\nSaved to: {output_path}\n\n{footer}"
return _text_result(result)
+226
View File
@@ -0,0 +1,226 @@
"""Caminhos: validação contra a sandbox, limites de tamanho, saída derivada.
Extraído de _shared.py — ver server_tools/_shared/__init__.py.
"""
from __future__ import annotations
import os
import re
from pathlib import Path
PROJECTS_DIR = os.environ.get("FCP_PROJECTS_DIR", os.path.expanduser("~/Movies"))
_SANDBOX_ENABLED = "FCP_PROJECTS_DIR" in os.environ
MAX_FILE_SIZE = 100 * 1024 * 1024
MAX_MEDIA_FILE_SIZE = 32 * 1024 * 1024 * 1024
_MAX_JSON_DEPTH = 50
def _check_json_depth(obj: object, _depth: int = 0) -> None:
"""Reject JSON structures nested beyond _MAX_JSON_DEPTH.
Prevents denial-of-service via deeply nested objects that exhaust the
call stack or memory during downstream processing. Called after
json.load() since Python's json module has no built-in depth limit.
"""
if _depth > _MAX_JSON_DEPTH:
raise ValueError(
f"JSON nesting depth exceeds {_MAX_JSON_DEPTH} — "
"file may be malformed or adversarial"
)
if isinstance(obj, dict):
for v in obj.values():
_check_json_depth(v, _depth + 1)
elif isinstance(obj, list):
for item in obj:
_check_json_depth(item, _depth + 1)
def _validate_filepath(
filepath: str,
allowed_extensions: tuple[str, ...] | None = None,
max_size: int = MAX_FILE_SIZE,
) -> str:
"""Validate a user-provided file path against traversal and size attacks.
Resolves symlinks, blocks null bytes, enforces extension whitelist, and
checks file size before any parsing takes place.
``max_size`` defaults to the document limit; callers handling source
media pass ``MAX_MEDIA_FILE_SIZE``, since media is streamed rather than
parsed into memory (see the constant for why).
Raises:
ValueError: For invalid paths (null bytes, bad extensions, oversized).
FileNotFoundError: When the resolved path does not exist.
"""
if '\x00' in filepath:
raise ValueError("Invalid file path: null byte detected")
resolved = Path(filepath).resolve()
if not resolved.exists():
raise FileNotFoundError(f"File not found: {filepath}")
# .fcpxmld bundles are directories (a package wrapping Info.fcpxml plus
# sidecar data files for object tracking / Cinematic mode). The size
# check applies to the inner Info.fcpxml, which is what gets parsed.
if resolved.is_dir():
if resolved.suffix.lower() != '.fcpxmld':
raise ValueError(f"Not a regular file: {filepath}")
inner = resolved / 'Info.fcpxml'
if not inner.is_file():
raise ValueError(f"Invalid bundle (no Info.fcpxml): {filepath}")
size_target = inner
elif not resolved.is_file():
raise ValueError(f"Not a regular file: {filepath}")
else:
size_target = resolved
if allowed_extensions and resolved.suffix.lower() not in allowed_extensions:
raise ValueError(
f"Invalid file type '{resolved.suffix}'. "
f"Allowed: {', '.join(allowed_extensions)}"
)
if size_target.stat().st_size > max_size:
size_mb = size_target.stat().st_size / (1024 * 1024)
raise ValueError(f"File too large ({size_mb:.1f} MB). Maximum: {max_size // (1024 * 1024)} MB")
return str(resolved)
def _validate_output_path(output_path: str, *, anchor_dir: str | None = None) -> str:
"""Validate an output path with optional sandbox enforcement.
Resolves traversal, blocks null bytes, ensures parent exists, and — when
*anchor_dir* is provided — verifies the resolved output lives under that
directory. This prevents LLM-generated tool calls from writing to
arbitrary filesystem locations (e.g. ``/etc/cron.d/backdoor``).
Args:
output_path: The raw output path to validate.
anchor_dir: If set, the resolved output must be a child of this
directory. Typically the parent directory of the input file so
outputs stay co-located with their sources.
Raises:
ValueError: For null bytes, missing parent, or sandbox escape.
"""
if '\x00' in output_path:
raise ValueError("Invalid output path: null byte detected")
resolved = Path(output_path).resolve()
if not resolved.parent.exists():
raise ValueError(f"Output directory does not exist: {resolved.parent}")
if anchor_dir is not None:
anchor = Path(anchor_dir).resolve()
try:
resolved.relative_to(anchor)
except ValueError:
raise ValueError(
f"Output path escapes allowed directory: "
f"{resolved} is not under {anchor}"
)
return str(resolved)
def _validate_directory(directory: str, *, allowed_root: str | None = None) -> str:
"""Validate a user-provided directory path against traversal and injection.
Resolves symlinks, blocks null bytes, and verifies the path is a real
directory. When *allowed_root* is given, the resolved path must be a
descendant of (or equal to) that root — preventing filesystem enumeration
beyond the project workspace.
Raises:
ValueError: For invalid paths (null bytes, not a directory, sandbox escape).
"""
if '\x00' in directory:
raise ValueError("Invalid directory path: null byte detected")
resolved = Path(directory).resolve()
if not resolved.is_dir():
raise ValueError(f"Not a valid directory: {directory}")
if allowed_root is not None:
root = Path(allowed_root).resolve()
try:
resolved.relative_to(root)
except ValueError:
raise ValueError(
f"Directory escapes allowed root: "
f"{resolved} is not under {root}"
)
return str(resolved)
def find_fcpxml_files(directory: str) -> list[str]:
"""Find all FCPXML files in a directory."""
path = Path(directory)
files = list(str(f) for f in path.rglob("*.fcpxml"))
files.extend(str(f) for f in path.rglob("*.fcpxmld"))
return sorted(files)
def generate_output_path(input_path: str, suffix: str = "_modified") -> str:
"""Generate output path from input path.
The suffix is sanitized to prevent path-component injection — only
alphanumeric, hyphen, underscore, and dot characters survive.
"""
# Strip anything that could inject path separators or traversal sequences
clean_suffix = re.sub(r'[^a-zA-Z0-9._-]', '', suffix)
if not clean_suffix:
clean_suffix = "_modified"
p = Path(input_path)
return str(p.parent / f"{p.stem}{clean_suffix}{p.suffix}")
def _resolve_io_paths(
arguments: dict,
suffix: str = "_modified",
) -> tuple[str, str]:
"""Validate input filepath and resolve the output path.
Shared foundation for every handler that reads an FCPXML and writes
a derived file. Validates the input, falls back to a suffixed
output name when ``output_path`` is not supplied, and sandbox-checks
the result.
Args:
arguments: Tool arguments dict (must contain ``filepath``; may
contain ``output_path``).
suffix: Default output filename suffix when ``output_path`` is
not provided (e.g. ``"_modified"``, ``"_beats"``).
Returns:
``(filepath, output_path)`` tuple with both paths validated.
"""
filepath = _validate_filepath(arguments["filepath"], ('.fcpxml', '.fcpxmld'))
# Anchor write operations to the input file's directory so LLM-generated
# tool calls cannot write to arbitrary filesystem locations (e.g.
# /etc/cron.d/backdoor). When the explicit sandbox is off, the anchor
# still prevents writes outside the source directory tree.
# `output_dir` is where the caller wants the file written, not merely a
# sandbox boundary: the app's "Pasta do projeto" promises that everything
# generated lands there. Deriving the name from the input but keeping the
# input's directory made every cross-directory call fail its own anchor
# check ("output path escapes allowed directory"), so the setting silently
# only worked when it pointed at the directory the file was already going
# to. An explicit `output_path` still wins, and still has to sit inside
# the anchor.
output_dir = arguments.get("output_dir")
if output_dir:
anchor = _validate_directory(str(output_dir))
default_output = str(Path(anchor) / Path(generate_output_path(filepath, suffix)).name)
else:
anchor = str(Path(filepath).resolve().parent)
default_output = generate_output_path(filepath, suffix)
output_path = _validate_output_path(
arguments.get("output_path") or default_output,
anchor_dir=anchor,
)
return filepath, output_path
+89
View File
@@ -0,0 +1,89 @@
"""Abrir um projeto e preparar modifier/generator para editá-lo.
Extraído de _shared.py — ver server_tools/_shared/__init__.py.
"""
from __future__ import annotations
from mcp.types import TextContent
from fcpxml.parser import FCPXMLParser
from fcpxml.rough_cut import RoughCutGenerator
from fcpxml.writer import FCPXMLModifier
from .paths import _resolve_io_paths, _validate_filepath
def _parse_project(filepath: str):
"""Parse an FCPXML file and return the project with its primary timeline."""
filepath = _validate_filepath(filepath, ('.fcpxml', '.fcpxmld'))
project = FCPXMLParser().parse_file(filepath)
if not project.timelines:
return None, None
return project, project.primary_timeline
def _text_result(text: str) -> list[TextContent]:
"""Wrap a string in the MCP TextContent list that every tool handler returns."""
return [TextContent(type="text", text=text)]
def _no_timeline():
"""Standard response when no timelines are found."""
return _text_result("No timelines found")
def _require_timeline(filepath: str):
"""Parse FCPXML and return (project, timeline), raising if no timeline exists.
Centralises the repeated _parse_project + _no_timeline guard that
appears in every read-only timeline handler. Returns a tuple so
callers can destructure directly::
project, tl = _require_timeline(arguments["filepath"])
"""
project, tl = _parse_project(filepath)
if not tl:
raise _NoTimelineError()
return project, tl
class _NoTimelineError(Exception):
"""Sentinel raised by _require_timeline when no timelines exist."""
def _setup_modifier(
arguments: dict,
suffix: str = "_modified",
) -> tuple[str, str, "FCPXMLModifier"]:
"""Common setup for write handlers: validate paths and create modifier.
Consolidates the repeated validate-filepath → resolve-output-path →
create-modifier boilerplate shared by 18+ write handlers.
Args:
arguments: Tool arguments dict (must contain ``filepath``; may
contain ``output_path``).
suffix: Default output filename suffix when ``output_path`` is
not provided (e.g. ``"_modified"``, ``"_flash_fixed"``).
Returns:
``(filepath, output_path, modifier)`` tuple ready for the
handler's domain-specific operation.
"""
filepath, output_path = _resolve_io_paths(arguments, suffix)
modifier = FCPXMLModifier(filepath)
return filepath, output_path, modifier
def _setup_generator(
arguments: dict,
suffix: str = "_roughcut",
) -> tuple[str, str, "RoughCutGenerator"]:
"""Common setup for generation handlers: validate paths and create generator.
Args:
arguments: Tool arguments dict (must contain ``filepath`` and
``output_path``).
suffix: Default output filename suffix.
Returns:
``(filepath, output_path, generator)`` tuple.
"""
filepath, output_path = _resolve_io_paths(arguments, suffix)
generator = RoughCutGenerator(filepath)
return filepath, output_path, generator
+480 -10
View File
@@ -6,13 +6,14 @@ Extracted from server.py; see Engine/docs/03_SERVER_TOOLS.md for the tool catalo
from __future__ import annotations from __future__ import annotations
import json import json
import re
from pathlib import Path from pathlib import Path
from typing import Sequence from typing import Sequence
from mcp.types import TextContent, Tool from mcp.types import TextContent, Tool
from fcpxml.media_intel import media_src_to_path from fcpxml.media_intel import media_src_to_path
from fcpxml.model_manager import load_dynamic_subtitle_config from fcpxml.model_manager import load_dynamic_subtitle_config, load_plain_subtitle_config
from fcpxml.models import DynamicSubtitleConfig, WordLook, WordStyle from fcpxml.models import DynamicSubtitleConfig, WordLook, WordStyle
from fcpxml.writer import FCPXMLModifier from fcpxml.writer import FCPXMLModifier
from server_tools._shared import ( from server_tools._shared import (
@@ -69,9 +70,174 @@ TOOLS = [
"required": ["filepath"] "required": ["filepath"]
} }
), ),
Tool(
name="generate_plain_subtitles",
description="Generate simple editable FCPXML text-title subtitles, synchronized to transcript words but without visual build-in/build-out effects. Words are grouped into short blocks, placed at a configurable vertical position, and written as static Text titles rather than SRT captions.",
inputSchema={
"type": "object",
"properties": {
"filepath": {"type": "string", "description": "Path to FCPXML file"},
"clip_name": {"type": "string", "description": "Only caption the clip with this name (default: all spine clips with matched source media)"},
"model": {"type": "string", "default": "base", "description": "Whisper model size if transcription is needed"},
"language": {"type": "string", "description": "ISO language code hint (e.g. 'pt'); auto-detected if omitted"},
"font": {"type": "string", "description": "Text font family. Falls back to saved plain-subtitle config."},
"font_size": {"type": "integer", "description": "Font size in canvas points. Falls back to saved plain-subtitle config."},
"font_color": {"type": "string", "description": "RGBA (0-1, space-separated). Falls back to saved plain-subtitle config."},
"max_words": {"type": "integer", "description": "Maximum words per subtitle block. Falls back to saved plain-subtitle config."},
"position_y": {"type": "number", "description": "Vertical title position in canvas points; negative sits lower in frame."},
"uppercase": {"type": "boolean", "description": "Render text in uppercase."},
"keep_punctuation": {"type": "boolean", "description": "Keep punctuation such as comma and period."},
"text_scale": {"type": "number", "description": "Template font-size scale. Falls back to saved plain-subtitle config."},
"output_path": {"type": "string", "description": "Output path (default: adds _plain_subtitles suffix)"},
},
"required": ["filepath"]
}
),
Tool(
name="generate_subtitles_by_emphasis",
description="Generate BOTH subtitle styles over the FULL clip and let them coexist by visibility, not by splitting words: plain static titles (see generate_plain_subtitles) cover every word from start to end; dynamic progressive-composition titles (see generate_dynamic_subtitles) are additionally generated for whichever whole phrases were marked as emphasis in the phrase-review step (etapa 5, zoom applied, level >= 1). Wherever a dynamic phrase is on screen, the plain titles underneath it are set enabled=\"0\" (still present in the FCPXML, editable/re-enable-able in Final Cut, just not rendered) instead of never being generated there — so disabling emphasis later never leaves a silent gap in the plain track. Reads emphasis spans from the media's cached '<media>_phrase_actions.json' (written by save_phrase_review after the app's etapa 5 review) — run the voice-editing wizard through that step first, or nothing is treated as emphasis and every title stays plain and enabled. Style knobs are the saved 'Legendas Dinâmicas'/plain-subtitle configs (~/.fcp-mcp-server/config.json); this tool does not expose per-call style overrides, only the split logic — use generate_dynamic_subtitles/generate_plain_subtitles directly if you need one-off styling.",
inputSchema={
"type": "object",
"properties": {
"filepath": {"type": "string", "description": "Path to FCPXML file"},
"clip_name": {"type": "string", "description": "Only caption the clip with this name (default: all spine clips with matched source media)"},
"model": {"type": "string", "default": "base", "description": "Whisper model size if transcription is needed"},
"language": {"type": "string", "description": "ISO language code hint (e.g. 'pt'); auto-detected if omitted"},
"granularity": {"type": "string", "enum": ["phrase", "word"], "default": "phrase", "description": "Passed through to the dynamic half, same meaning as in generate_dynamic_subtitles"},
"max_words": {"type": "integer", "description": "Max words per block for the plain half. Falls back to saved plain-subtitle config."},
"uppercase": {"type": "boolean", "description": "Uppercase the plain half. Falls back to saved plain-subtitle config."},
"keep_punctuation": {"type": "boolean", "description": "Keep punctuation in the plain half. Falls back to saved plain-subtitle config."},
"output_path": {"type": "string", "description": "Output path (default: adds _emphasis_subtitles suffix)"},
},
"required": ["filepath"]
}
),
] ]
_PUNCT_RE = re.compile(r"[^\w\sÀ-ÖØ-öø-ÿ]", re.UNICODE)
def _words_overlapping_clip(words: Sequence[dict], start: float, end: float) -> list[dict]:
"""Return transcript words that overlap a source window, rebased to it."""
clip_words: list[dict] = []
for w in words:
word_start = float(w.get("start", 0.0))
word_end = float(w.get("end", word_start))
if word_end <= start or word_start >= end:
continue
clip_words.append(
{
"word": w.get("word", ""),
"start": max(0.0, word_start - start),
"end": max(0.0, min(word_end, end) - start),
}
)
return clip_words
def _plain_word_text(word: str, *, uppercase: bool, keep_punctuation: bool) -> str:
text = str(word or "").strip()
if not keep_punctuation:
text = _PUNCT_RE.sub("", text)
text = re.sub(r"\s+", " ", text).strip()
return text.upper() if uppercase else text
def _plain_subtitle_blocks(words: Sequence[dict], max_words: int) -> list[list[dict]]:
blocks: list[list[dict]] = []
pending: list[dict] = []
for word in words:
if not str(word.get("word", "")).strip():
continue
pending.append(word)
if len(pending) >= max(1, max_words):
blocks.append(pending)
pending = []
if pending:
blocks.append(pending)
return blocks
def _phrase_actions_path(media_path: str) -> Path:
"""Where `save_phrase_review` writes emphasis decisions for this media.
Mirrors `phrase_review.review_paths()`'s naming (stem + "_phrase_actions.json"),
without importing that module just for a path — the voice_timeline this would
normally derive from is itself named `<media stem>_voice_timeline.json`, so
stripping straight from the media stem lands on the same file.
"""
stem = Path(media_path).stem
return Path(media_path).with_name(f"{stem}_phrase_actions.json")
def _load_emphasis_spans(media_path: str) -> list[dict]:
"""Load emphasis spans (source-media time) saved by the etapa-5 phrase review.
Returns [] if the review was never run for this media — callers should treat
that as "nothing is emphasis yet", not as an error, since the wizard's later
steps are optional.
"""
path = _phrase_actions_path(media_path)
if not path.is_file():
return []
try:
data = json.loads(path.read_text(encoding="utf-8"))
except (OSError, json.JSONDecodeError):
return []
spans = data.get("emphasis_spans", [])
return [s for s in spans if isinstance(s, dict) and "start" in s and "end" in s]
def _word_in_spans(word_start: float, word_end: float, spans: Sequence[dict]) -> bool:
"""A word belongs to an emphasis span if its midpoint falls inside it.
Midpoint, not start, so a word straddling a span boundary (which can happen
since spans come from phrase trims, not word timestamps) lands on whichever
side it mostly belongs to instead of always defaulting to one edge.
"""
mid = (word_start + word_end) / 2.0
return any(float(s["start"]) <= mid < float(s["end"]) for s in spans)
def _words_in_spans(words: Sequence[dict], spans: Sequence[dict]) -> list[dict]:
"""The subset of source-time transcript words that fall inside a span.
Feeds only the DYNAMIC half — the plain half always gets every word, full
clip, unfiltered; this is not a partition of the word list into two
disjoint sets, it is "which words also get the dynamic treatment on top".
"""
if not spans:
return []
return [
w for w in words
if _word_in_spans(float(w.get("start", 0.0)), float(w.get("end", w.get("start", 0.0))), spans)
]
def _segments_in_spans(segments: Sequence[dict], spans: Sequence[dict]) -> list[dict]:
"""Keep only the sentences that fall inside an emphasis span (by midpoint).
Feeds the dynamic half's sentence-block builder; segments outside every span
would only produce blocks with no words left in them after the word filter.
"""
if not spans:
return []
kept = []
for seg in segments:
start = float(seg.get("start", 0.0))
end = float(seg.get("end", start))
mid = (start + end) / 2.0
if any(float(s["start"]) <= mid < float(s["end"]) for s in spans):
kept.append(seg)
return kept
def _overlaps_any_span(start: float, end: float, spans: Sequence[tuple[float, float]]) -> bool:
"""Half-open interval overlap: a plain title under this window must hide."""
return any(start < span_end and end > span_start for span_start, span_end in spans)
async def handle_validate_subtitle_layout(arguments: dict) -> Sequence[TextContent]: async def handle_validate_subtitle_layout(arguments: dict) -> Sequence[TextContent]:
"""Validate title/subtitle layout for spatial collisions and safe-area """Validate title/subtitle layout for spatial collisions and safe-area
containment (collision.validate_titles over every <title> in the file).""" containment (collision.validate_titles over every <title> in the file)."""
@@ -210,15 +376,7 @@ async def handle_generate_dynamic_subtitles(arguments: dict) -> Sequence[TextCon
clip_duration = modifier._parse_time(el.get("duration", "0s")).to_seconds() clip_duration = modifier._parse_time(el.get("duration", "0s")).to_seconds()
window_end = clip_source_start + clip_duration window_end = clip_source_start + clip_duration
clip_words = [ clip_words = _words_overlapping_clip(data.get("words", []), clip_source_start, window_end)
{
"word": w.get("word", ""),
"start": float(w.get("start", 0.0)) - clip_source_start,
"end": float(w.get("end", 0.0)) - clip_source_start,
}
for w in data.get("words", [])
if clip_source_start <= float(w.get("start", 0.0)) < window_end
]
if not clip_words: if not clip_words:
skipped.append((name, "no words in clip's source range")) skipped.append((name, "no words in clip's source range"))
continue continue
@@ -277,7 +435,319 @@ async def handle_generate_dynamic_subtitles(arguments: dict) -> Sequence[TextCon
return _text_result(result) return _text_result(result)
async def handle_generate_plain_subtitles(arguments: dict) -> Sequence[TextContent]:
"""Generate static, editable title subtitles from word-level transcripts."""
model = arguments.get("model", "base")
language = arguments.get("language")
output_dir = arguments.get("output_dir")
clip_filter = arguments.get("clip_name")
saved = load_plain_subtitle_config()
font = arguments.get("font") or saved["font"]
font_size = int(arguments.get("font_size", saved["font_size"]))
font_color = arguments.get("font_color") or saved["font_color"]
max_words = max(1, int(arguments.get("max_words", saved["max_words"])))
position_y = float(arguments.get("position_y", saved["position_y"]))
uppercase = bool(arguments.get("uppercase", saved["uppercase"]))
keep_punctuation = bool(arguments.get("keep_punctuation", saved["keep_punctuation"]))
filepath, output_path, modifier = _setup_modifier(arguments, "_plain_subtitles")
added: list[tuple[str, int, int]] = []
skipped: list[tuple[str, str]] = []
spine_clips = [el for _, el in modifier._iter_spine_clips()]
for el in spine_clips:
name = el.get("name", "")
if clip_filter and name != clip_filter:
continue
src = modifier.resources.get(el.get("ref", ""), {}).get("src", "")
media_path = media_src_to_path(src)
if not media_path or not Path(media_path).is_file():
skipped.append((name, "media file missing"))
continue
data, reason = _load_or_transcribe(media_path, model, language, output_dir)
if data is None:
skipped.append((name, reason))
continue
clip_source_start = modifier.source_file_start(el).to_seconds()
clip_duration = modifier._parse_time(el.get("duration", "0s")).to_seconds()
clip_words = _words_overlapping_clip(
data.get("words", []), clip_source_start, clip_source_start + clip_duration
)
if not clip_words:
skipped.append((name, "no words in clip's source range"))
continue
blocks = _plain_subtitle_blocks(clip_words, max_words)
created = 0
for block in blocks:
parts = [
_plain_word_text(w.get("word", ""), uppercase=uppercase, keep_punctuation=keep_punctuation)
for w in block
]
text = " ".join(p for p in parts if p).strip()
if not text:
continue
start = max(0.0, min(float(w.get("start", 0.0)) for w in block))
end = max(float(w.get("end", start)) for w in block)
duration = max(end - start, modifier.frame_duration_fraction())
modifier.add_text_title(
el,
text,
offset=f"{start:.6f}s",
duration=f"{duration:.6f}s",
lane=20,
position=f"0 {position_y:g}",
font=font,
font_size=font_size,
font_color=font_color,
bold=True,
face=None,
font_scale=1.0,
size_param=font_size,
)
created += 1
if created:
added.append((name, created, len(clip_words)))
if not added:
text = "# Plain Subtitles\n\nNo subtitles generated — file unchanged (nothing saved)."
if skipped:
text += "\n\n## Skipped Clips\n" + _markdown_table(
["Clip", "Reason"], [[n, r] for n, r in skipped]
)
return _text_result(text)
modifier.save(output_path)
total_titles = sum(lines for _, lines, _ in added)
total_words = sum(words for _, _, words in added)
result = "# Plain Subtitles Generated\n\n## Summary\n"
result += (
f"- **Clips Captioned**: {len(added)}\n"
f"- **Title Clips**: {total_titles}\n"
f"- **Total Words**: {total_words}\n\n"
)
result += _markdown_table(
["Clip", "Title Clips", "Words"],
[[n, str(lines), str(words)] for n, lines, words in added],
)
if skipped:
result += "\n## Skipped Clips\n" + _markdown_table(
["Clip", "Reason"], [[n, r] for n, r in skipped]
)
result += f"\n\nSaved to: `{output_path}`\n\n*Transcripts are cached as _transcript.json.*"
return _text_result(result)
async def handle_generate_subtitles_by_emphasis(arguments: dict) -> Sequence[TextContent]:
"""Generate plain titles for the whole clip and dynamic titles for the
emphasis phrases on top, then hide (enabled="0") the plain titles that
fall under a dynamic phrase — never split the word list between the two.
Plain always covers every word, so turning emphasis off later (editing
the phrase review and re-running) never leaves a silent gap: the plain
title was there all along, just disabled.
"""
model = arguments.get("model", "base")
language = arguments.get("language")
output_dir = arguments.get("output_dir")
clip_filter = arguments.get("clip_name")
granularity = arguments.get("granularity", "phrase")
saved_dynamic = load_dynamic_subtitle_config()
body_color = saved_dynamic["active_color"]
dynamic_config = DynamicSubtitleConfig(
style=WordStyle(
font=saved_dynamic["font"],
font_size=int(saved_dynamic["font_size"]),
active_color=body_color,
inactive_color="0.7 0.7 0.7 1",
emphasis_look=WordLook(
int(saved_dynamic["emphasis_size"]),
saved_dynamic["emphasis_color"] or body_color,
font=saved_dynamic["emphasis_font"],
face=saved_dynamic["emphasis_face"],
kerning=0.0,
),
body_look=WordLook(
int(saved_dynamic["font_size"]),
body_color,
font=saved_dynamic["font"],
face="Bold",
kerning=1.2,
),
),
band_height=float(saved_dynamic["band_height"]),
block_center_y=float(saved_dynamic["block_center_y"]),
granularity=granularity,
text_scale=float(saved_dynamic["text_scale"]),
line_gap=float(saved_dynamic["line_gap"]),
)
saved_plain = load_plain_subtitle_config()
plain_font = saved_plain["font"]
plain_font_size = int(saved_plain["font_size"])
plain_font_color = saved_plain["font_color"]
max_words = max(1, int(arguments.get("max_words", saved_plain["max_words"])))
position_y = float(saved_plain["position_y"])
uppercase = bool(arguments.get("uppercase", saved_plain["uppercase"]))
keep_punctuation = bool(arguments.get("keep_punctuation", saved_plain["keep_punctuation"]))
filepath, output_path, modifier = _setup_modifier(arguments, "_emphasis_subtitles")
added: list[tuple[str, int, int, int, int]] = []
skipped: list[tuple[str, str]] = []
no_review: list[str] = []
spine_clips = [el for _, el in modifier._iter_spine_clips()]
for el in spine_clips:
name = el.get("name", "")
if clip_filter and name != clip_filter:
continue
src = modifier.resources.get(el.get("ref", ""), {}).get("src", "")
media_path = media_src_to_path(src)
if not media_path or not Path(media_path).is_file():
skipped.append((name, "media file missing"))
continue
data, reason = _load_or_transcribe(media_path, model, language, output_dir)
if data is None:
skipped.append((name, reason))
continue
spans = _load_emphasis_spans(media_path)
if not spans:
no_review.append(name)
clip_source_start = modifier.source_file_start(el).to_seconds()
clip_duration = modifier._parse_time(el.get("duration", "0s")).to_seconds()
window_end = clip_source_start + clip_duration
# Clip-relative windows, for deciding which plain titles to hide —
# same coordinate space add_text_title's offsets end up in.
clip_spans = [
(max(0.0, float(s["start"]) - clip_source_start), min(clip_duration, float(s["end"]) - clip_source_start))
for s in spans
if float(s["end"]) > clip_source_start and float(s["start"]) < window_end
]
all_words = data.get("words", [])
dynamic_lines = 0
dynamic_word_count = 0
emphasis_words = _words_in_spans(all_words, spans)
clip_emphasis_words = _words_overlapping_clip(emphasis_words, clip_source_start, window_end)
if clip_emphasis_words:
all_segments = data.get("segments", [])
emphasis_segments = _segments_in_spans(all_segments, spans)
clip_segments = [
{
"start": float(s.get("start", 0.0)) - clip_source_start,
"end": float(s.get("end", 0.0)) - clip_source_start,
}
for s in emphasis_segments
if float(s.get("end", 0.0)) > clip_source_start
and float(s.get("start", 0.0)) < window_end
]
# Pass the element itself, not `name` — see the same note in
# handle_generate_dynamic_subtitles (Engine/docs/05_EXPERIENCIAS.md,
# entry 2026-08-17).
dynamic_lines = len(
modifier.generate_dynamic_subtitles(
el, clip_emphasis_words, dynamic_config, segments=clip_segments
)
)
dynamic_word_count = len(clip_emphasis_words)
# Plain covers EVERY word in the clip — never filtered by emphasis.
# Titles landing under a dynamic phrase are disabled below instead of
# never being created, so turning emphasis off later never leaves a
# silent gap where neither style is on screen.
plain_created = 0
plain_hidden = 0
clip_all_words = _words_overlapping_clip(all_words, clip_source_start, window_end)
blocks = _plain_subtitle_blocks(clip_all_words, max_words)
for block in blocks:
parts = [
_plain_word_text(w.get("word", ""), uppercase=uppercase, keep_punctuation=keep_punctuation)
for w in block
]
text = " ".join(p for p in parts if p).strip()
if not text:
continue
start = max(0.0, min(float(w.get("start", 0.0)) for w in block))
end = max(float(w.get("end", start)) for w in block)
duration = max(end - start, modifier.frame_duration_fraction())
title = modifier.add_text_title(
el,
text,
offset=f"{start:.6f}s",
duration=f"{duration:.6f}s",
lane=20,
position=f"0 {position_y:g}",
font=plain_font,
font_size=plain_font_size,
font_color=plain_font_color,
bold=True,
face=None,
font_scale=1.0,
size_param=plain_font_size,
)
plain_created += 1
if _overlaps_any_span(start, end, clip_spans):
title.set("enabled", "0")
plain_hidden += 1
if dynamic_lines or plain_created:
added.append(
(name, dynamic_lines, plain_created, plain_hidden, dynamic_word_count + len(clip_all_words))
)
else:
skipped.append((name, "no words in clip's source range"))
if not added:
text = "# Subtitles by Emphasis\n\nNo captions generated — file unchanged (nothing saved)."
if skipped:
text += "\n\n## Skipped Clips\n" + _markdown_table(
["Clip", "Reason"], [[n, r] for n, r in skipped]
)
return _text_result(text)
modifier.save(output_path)
total_dynamic = sum(d for _, d, _, _, _ in added)
total_plain = sum(p for _, _, p, _, _ in added)
total_hidden = sum(h for _, _, _, h, _ in added)
total_words = sum(w for _, _, _, _, w in added)
result = "# Subtitles by Emphasis Generated\n\n## Summary\n"
result += (
f"- **Clips Captioned**: {len(added)}\n"
f"- **Dynamic Title Lines (emphasis)**: {total_dynamic}\n"
f"- **Plain Title Blocks (full clip)**: {total_plain}\n"
f"- **Plain Blocks Hidden Under Emphasis (enabled=\"0\")**: {total_hidden}\n"
f"- **Total Words**: {total_words}\n\n"
)
result += _markdown_table(
["Clip", "Dynamic Lines", "Plain Blocks", "Hidden", "Words"],
[[n, str(d), str(p), str(h), str(w)] for n, d, p, h, w in added],
)
if no_review:
result += (
"\n## Sem revisão de ênfase\n"
"Nenhum `_phrase_actions.json` encontrado para: "
+ ", ".join(no_review)
+ " — todas as frases desses clipes saíram como legenda comum. "
"Rode a etapa 5 do Assistente (revisão de frases) antes, se quiser destaque dinâmico.\n"
)
if skipped:
result += "\n## Skipped Clips\n" + _markdown_table(
["Clip", "Reason"], [[n, r] for n, r in skipped]
)
result += f"\n\nSaved to: `{output_path}`\n\n*Transcripts are cached as _transcript.json; emphasis spans from _phrase_actions.json.*"
return _text_result(result)
HANDLERS = { HANDLERS = {
"validate_subtitle_layout": handle_validate_subtitle_layout, "validate_subtitle_layout": handle_validate_subtitle_layout,
"generate_dynamic_subtitles": handle_generate_dynamic_subtitles, "generate_dynamic_subtitles": handle_generate_dynamic_subtitles,
"generate_plain_subtitles": handle_generate_plain_subtitles,
"generate_subtitles_by_emphasis": handle_generate_subtitles_by_emphasis,
} }
+2 -2
View File
@@ -67,12 +67,12 @@ TOOLS = [
), ),
Tool( Tool(
name="remove_filler_words", name="remove_filler_words",
description="Cut filler words (um, uh, erm...) out of the timeline with ripple, using word-level transcripts of the real source audio. Conservative default filler list — words like 'like' and 'so' are only cut if you pass them explicitly. Uses each media file's _transcript.json (auto-transcribes if missing). Non-destructive: writes a _defillered copy.", description="Cut filler interjections (uh, erm...) out of the timeline with ripple, using word-level transcripts of the real source audio. Conservative default filler list — words like 'um', 'uma', 'like' and 'so' are only cut if you pass them explicitly. Uses each media file's _transcript.json (auto-transcribes if missing). Non-destructive: writes a _defillered copy.",
inputSchema={ inputSchema={
"type": "object", "type": "object",
"properties": { "properties": {
"filepath": {"type": "string", "description": "Path to FCPXML file"}, "filepath": {"type": "string", "description": "Path to FCPXML file"},
"fillers": {"type": "array", "items": {"type": "string"}, "description": "Filler words/phrases to cut (default: um, uh, uhh, umm, erm, ehm, mmm, hmm, mhm)"}, "fillers": {"type": "array", "items": {"type": "string"}, "description": "Filler words/phrases to cut (default: uh, uhh, umm, erm, ehm, mmm, hmm, mhm; pass um/uma explicitly if desired)"},
"clip_name": {"type": "string", "description": "Only clean the clip with this name"}, "clip_name": {"type": "string", "description": "Only clean the clip with this name"},
"model": {"type": "string", "default": "base", "description": "Whisper model size if transcription is needed"}, "model": {"type": "string", "default": "base", "description": "Whisper model size if transcription is needed"},
"padding": {"type": "number", "default": 0.02, "description": "Seconds to widen each cut on both sides (0-2, default 0.02)"}, "padding": {"type": "number", "default": 0.02, "description": "Seconds to widen each cut on both sides (0-2, default 0.02)"},
+171 -2
View File
@@ -13,6 +13,7 @@ from mcp.types import TextContent, Tool
from fcpxml.diarize import assign_speakers, build_speakers, diarization_capability, diarize from fcpxml.diarize import assign_speakers, build_speakers, diarization_capability, diarize
from fcpxml.emphasis import EmphasisWeights from fcpxml.emphasis import EmphasisWeights
from fcpxml.llm_local import DEFAULT_BASE_URL, generate_voice_actions
from fcpxml.media_intel import media_src_to_path from fcpxml.media_intel import media_src_to_path
from fcpxml.model_manager import ( from fcpxml.model_manager import (
load_hf_token, load_hf_token,
@@ -21,6 +22,7 @@ from fcpxml.model_manager import (
save_voice_analysis_config, save_voice_analysis_config,
) )
from fcpxml.models import TimeValue from fcpxml.models import TimeValue
from fcpxml.phrase_review import build_phrase_review, save_phrase_review
from fcpxml.voice_actions import parse_actions, resolve_actions, speaker_cut_actions from fcpxml.voice_actions import parse_actions, resolve_actions, speaker_cut_actions
from fcpxml.voice_features import extract_energy, extract_pitch, features_capability from fcpxml.voice_features import extract_energy, extract_pitch, features_capability
from fcpxml.voice_timeline import ( from fcpxml.voice_timeline import (
@@ -93,6 +95,7 @@ TOOLS = [
"hf_token": {"type": "string", "description": "HuggingFace token for speaker diarization (default: the persisted token; omit to skip diarization)"}, "hf_token": {"type": "string", "description": "HuggingFace token for speaker diarization (default: the persisted token; omit to skip diarization)"},
"num_speakers": {"type": "string", "description": "Known number of speakers, if any (default: the persisted setting, else auto-detect)"}, "num_speakers": {"type": "string", "description": "Known number of speakers, if any (default: the persisted setting, else auto-detect)"},
"output_dir": {"type": "string", "description": "Folder to write _voice_timeline.json into (default: next to the media file)"}, "output_dir": {"type": "string", "description": "Folder to write _voice_timeline.json into (default: next to the media file)"},
"rotation": {"type": "number", "description": "Degrees the clip is rotated by in the FCPXML (e.g. a Transform filter straightening a tilted phone shot). Recorded in the timeline JSON so a preview can apply the same correction. Default 0."},
}, },
"required": ["media_path"] "required": ["media_path"]
} }
@@ -167,6 +170,27 @@ TOOLS = [
"required": ["filepath", "actions"] "required": ["filepath", "actions"]
} }
), ),
Tool(
name="generate_voice_script",
description="Run the WHOLE voice-edit pass internally, no wizard, no copy-paste: reuse an existing voice timeline (or transcribe + build one) -> hand it to a LOCAL model (Ollama running Gemma 3 / Llama) that directs the edit -> return the readable script (roteiro) AND the action JSON, and optionally apply it to a FCPXML. The model reads the full _voice_timeline.json (the whole file goes with the brief) and emits cut/zoom/text/marker decisions per the editar-por-voz brief; decisions are validated row-by-row so one bad row never discards the edit. Times stay in ORIGINAL source seconds; the applier resolves cuts and shifts everything else. Writes _voice_timeline.json, _phrase_review.json, _phrase_actions.json and (when applying) a _voice_edit FCPXML. Defaults to the local model 'gemma3:12b' at http://localhost:11434 — change via model/base_url.",
inputSchema={
"type": "object",
"properties": {
"media_path": {"type": "string", "description": "Path to the audio/video file to analyze and direct (.wav, .mp3, .m4a, .aac, .aif, .flac, .mov, .mp4). Required when there is no voice_timeline yet; ignored when voice_timeline is provided."},
"voice_timeline": {"type": "string", "description": "Path to an existing _voice_timeline.json (e.g. from the assistant's analysis step). When given, it is reused and transcription/acoustics are skipped — the model gets the whole file to direct the edit."},
"filepath": {"type": "string", "description": "Optional FCPXML to apply the decisions to (non-destructive: writes a _voice_edit copy). When omitted, only the script and actions are produced."},
"model": {"type": "string", "default": "gemma3:12b", "description": "Local model Ollama serves (e.g. 'gemma3:12b', 'gemma3:4b', 'llama3')"},
"base_url": {"type": "string", "default": "http://localhost:11434", "description": "Ollama base URL"},
"model_size": {"type": "string", "default": "base", "description": "Whisper model size to use if transcription is needed"},
"language": {"type": "string", "description": "ISO language code hint for transcription, if needed"},
"hf_token": {"type": "string", "description": "HuggingFace token for speaker diarization (omit to skip)"},
"num_speakers": {"type": "string", "description": "Known number of speakers, if any"},
"output_dir": {"type": "string", "description": "Folder to write the timeline/review/actions JSON into (default: next to the media file)"},
"apply_to_fcpxml": {"type": "boolean", "default": True, "description": "When filepath is given, apply the decisions to it. Set false to only produce the script."},
},
"required": []
}
),
Tool( Tool(
name="get_voice_analysis_config", name="get_voice_analysis_config",
description="Read the persisted Voice Analysis settings: energy threshold, emphasis-index weights (energy/pitch_variation/rate_variation/pause_before/duration), emphasis cutoff for punch-in candidates, and emotion detection toggle/sensitivity. Shared with the MacApp settings screen (~/.fcp-mcp-server/config.json).", description="Read the persisted Voice Analysis settings: energy threshold, emphasis-index weights (energy/pitch_variation/rate_variation/pause_before/duration), emphasis cutoff for punch-in candidates, and emotion detection toggle/sensitivity. Shared with the MacApp settings screen (~/.fcp-mcp-server/config.json).",
@@ -348,8 +372,10 @@ async def handle_build_voice_timeline(arguments: dict) -> Sequence[TextContent]:
language = arguments.get("language") language = arguments.get("language")
token = str(arguments.get("hf_token") or "").strip() or load_hf_token() or None token = str(arguments.get("hf_token") or "").strip() or load_hf_token() or None
num_speakers = str(arguments.get("num_speakers") or "").strip() or load_num_speakers() num_speakers = str(arguments.get("num_speakers") or "").strip() or load_num_speakers()
output_dir = arguments.get("output_dir")
rotation = float(arguments.get("rotation") or 0.0)
transcript, reason = _load_or_transcribe(media_path, model, language) transcript, reason = _load_or_transcribe(media_path, model, language, output_dir)
if transcript is None: if transcript is None:
return _text_result( return _text_result(
f"# Voice Timeline\n\nCould not obtain a transcript " f"# Voice Timeline\n\nCould not obtain a transcript "
@@ -365,9 +391,11 @@ async def handle_build_voice_timeline(arguments: dict) -> Sequence[TextContent]:
weights=EmphasisWeights.from_dict(config["emphasis_weights"]), weights=EmphasisWeights.from_dict(config["emphasis_weights"]),
peak_percentile=config["peak_percentile"], peak_percentile=config["peak_percentile"],
emphasis_floor=config["emphasis_floor"], emphasis_floor=config["emphasis_floor"],
emotion_enabled=config["emotion_enabled"],
emotion_sensitivity=config["emotion_sensitivity"],
rotation=rotation,
) )
output_dir = arguments.get("output_dir")
json_path = Path(_validate_output_path( json_path = Path(_validate_output_path(
str(voice_timeline_path(media_path, output_dir)), str(voice_timeline_path(media_path, output_dir)),
anchor_dir=str(Path(output_dir) if output_dir else Path(media_path).parent), anchor_dir=str(Path(output_dir) if output_dir else Path(media_path).parent),
@@ -398,6 +426,7 @@ async def handle_build_voice_timeline(arguments: dict) -> Sequence[TextContent]:
"yes" if layers["acoustics"] else "FAILED — every acoustic value is 0", "yes" if layers["acoustics"] else "FAILED — every acoustic value is 0",
], ],
["Speakers", "yes" if layers["speakers"] else "not run — single default speaker"], ["Speakers", "yes" if layers["speakers"] else "not run — single default speaker"],
["Emotion", "yes" if layers.get("emotion") else "not run"],
], ],
) + "\n" ) + "\n"
@@ -739,6 +768,145 @@ async def handle_save_voice_analysis_config(arguments: dict) -> Sequence[TextCon
return _text_result(_voice_analysis_config_text(config)) return _text_result(_voice_analysis_config_text(config))
async def handle_generate_voice_script(arguments: dict) -> Sequence[TextContent]:
"""The whole voice-edit pass, run inside the engine against a local model.
Transcribe (cached) -> build the voice timeline -> ask the local LLM to
direct the edit -> build the readable script (roteiro) + the action JSON ->
optionally apply to a FCPXML. No wizard, no copy-paste: the model's JSON is
parsed and validated like any other decision source, and the applier turns
it into FCPXML the same way it would for the rules engine.
"""
model = "gemma3:12b" if not arguments.get("model") else str(arguments["model"])
base_url = str(arguments.get("base_url") or DEFAULT_BASE_URL)
model_size = arguments.get("model_size", "base")
language = arguments.get("language")
token = str(arguments.get("hf_token") or "").strip() or load_hf_token() or None
num_speakers = str(arguments.get("num_speakers") or "").strip() or load_num_speakers()
output_dir = arguments.get("output_dir")
fcpxml_path = arguments.get("filepath")
apply = bool(arguments.get("apply_to_fcpxml", True)) and bool(fcpxml_path)
# Camino 1: já temos uma voice timeline (etapa de análise do assistente) —
# reaproveita e pula a transcrição/análise acústica/diarização, que é caro.
# Camino 2: só mídia — transcreve e monta a timeline do zero.
vt_arg = arguments.get("voice_timeline")
timeline = load_voice_timeline(Path(vt_arg)) if vt_arg and Path(vt_arg).is_file() else None
media_path = arguments.get("media_path")
if timeline is None:
media_path = _validate_filepath(
media_path, AUDIO_MEDIA_EXTENSIONS, max_size=MAX_MEDIA_FILE_SIZE
)
transcript, reason = _load_or_transcribe(media_path, model_size, language, output_dir)
if transcript is None:
return _text_result(
f"# Roteiro por IA Local\n\nNão foi possível obter a transcrição "
f"({reason}).{_TRANSCRIBE_INSTALL_HINT}"
)
config = load_voice_analysis_config()
timeline = build_voice_timeline(
media_path,
transcript,
hf_token=token,
num_speakers=num_speakers,
weights=EmphasisWeights.from_dict(config["emphasis_weights"]),
peak_percentile=config["peak_percentile"],
emphasis_floor=config["emphasis_floor"],
emotion_enabled=config["emotion_enabled"],
emotion_sensitivity=config["emotion_sensitivity"],
)
vt_arg = str(_validate_output_path(
str(voice_timeline_path(media_path, output_dir)),
anchor_dir=str(Path(output_dir) if output_dir else Path(media_path).parent),
))
save_voice_timeline(timeline, Path(vt_arg))
else:
# A timeline veio pronta; a mídia só é necessária se for aplicar e o
# caller não a passou — deriva do próprio campo `source` da timeline.
if not media_path:
candidate = Path(vt_arg).parent / timeline.get("source", "")
media_path = str(candidate) if candidate.is_file() else None
decision = generate_voice_actions(timeline, model=model, base_url=base_url)
actions = decision["actions"]
errors = list(decision["errors"])
if not actions and errors:
# The model produced nothing usable (transport error or unparseable
# response) — report it clearly instead of a silent "0 decisions".
raise RuntimeError(
"O modelo local não devolveu decisões utilizáveis: " + "; ".join(errors)
)
review = build_phrase_review(
timeline,
[a.as_dict() for a in actions],
voice_timeline_path=str(vt_arg),
)
review_path, actions_path = save_phrase_review(str(vt_arg), review)
roteiro = _roteiro_markdown(review, timeline.get("source", ""))
roteiro_path = Path(vt_arg).with_name(Path(vt_arg).stem.replace("_voice_timeline", "") + "_roteiro.md")
roteiro_path.write_text(roteiro, encoding="utf-8")
applied_text = ""
if apply:
contents = await handle_apply_voice_actions({
"filepath": fcpxml_path,
"actions": [a.as_dict() for a in actions],
"output_dir": output_dir,
})
applied_text = "\n\n" + "\n".join(getattr(c, "text", str(c)) for c in contents)
result = f"""# Roteiro por IA Local ({model})
## Resumo
- **Fonte**: {timeline.get('source', '')}
- **Duração**: {format_duration(timeline['summary']['duration'])}
- **Decisões do modelo**: {len(actions)} (cortes/zoom/texto/marcador)
- **Linha do tempo**: {vt_arg}
- **Roteiro (legível)**: {roteiro_path}
- **Ações JSON**: {actions_path}
- **Revisão de frases**: {review_path}
"""
if errors:
result += "\n## Rejeitado / avisos\n" + "\n".join(f"- {e}" for e in errors) + "\n"
result += "\n---\n\n" + roteiro
result += applied_text
result += "\n\n*Tudo rodou internamente: o modelo local leu a timeline e decidiu a edição; nenhum passo manual foi necessário.*"
return _text_result(result)
def _roteiro_markdown(review: dict, source: str) -> str:
"""The readable script: kept lines (roteiro) then the cut/bastidor lines."""
phrases = review.get("phrases", [])
kept = [p for p in phrases if p.get("active")]
cut = [p for p in phrases if not p.get("active")]
lines = [f"# Roteiro — {source}", ""]
lines.append(f"**{len(kept)} falas mantidas · {len(cut)} cortadas**")
lines.append("")
lines.append("## Roteiro (mantido)")
if not kept:
lines.append("_Nenhuma fala mantida._")
for p in kept:
tag = ""
if p.get("emphasis", 0) >= 1:
tag = f" · zoom nível {p['emphasis']}"
spk = f"[{p.get('speaker', '')}] " if p.get("speaker") else ""
lines.append(f"- {spk}{p.get('text', '')}{tag}")
if p.get("reason"):
lines.append(f" - _decisão_: {p['reason']}")
if cut:
lines.append("")
lines.append("## Cortado / bastidor")
for p in cut:
spk = f"[{p.get('speaker', '')}] " if p.get("speaker") else ""
lines.append(f"- {spk}{p.get('text', '')}")
if p.get("reason"):
lines.append(f" - _por que cortou_: {p['reason']}")
return "\n".join(lines) + "\n"
HANDLERS = { HANDLERS = {
"diarize_media": handle_diarize_media, "diarize_media": handle_diarize_media,
"analyze_voice_features": handle_analyze_voice_features, "analyze_voice_features": handle_analyze_voice_features,
@@ -746,6 +914,7 @@ HANDLERS = {
"remove_speakers": handle_remove_speakers, "remove_speakers": handle_remove_speakers,
"refine_voice_timeline": handle_refine_voice_timeline, "refine_voice_timeline": handle_refine_voice_timeline,
"apply_voice_actions": handle_apply_voice_actions, "apply_voice_actions": handle_apply_voice_actions,
"generate_voice_script": handle_generate_voice_script,
"get_voice_analysis_config": handle_get_voice_analysis_config, "get_voice_analysis_config": handle_get_voice_analysis_config,
"save_voice_analysis_config": handle_save_voice_analysis_config, "save_voice_analysis_config": handle_save_voice_analysis_config,
} }
+23
View File
@@ -264,3 +264,26 @@ class TestIntegration:
report = modifier.validate_subtitle_layout() report = modifier.validate_subtitle_layout()
assert report["summary"]["spatial_collision"] >= 1 assert report["summary"]["spatial_collision"] >= 1
assert blocking(report["severity"]) assert blocking(report["severity"])
def test_disabled_title_is_excluded_from_validation(self, temp_fcpxml):
"""A title with enabled="0" never renders in Final Cut
(generate_subtitles_by_emphasis disables plain titles under an
emphasis phrase instead of never creating them) — it must not count
as a collision, or as outside-frame/outside-safe-area, against the
title actually drawn in its place."""
modifier = FCPXMLModifier(temp_fcpxml)
titles = modifier.generate_dynamic_subtitles("Interview_A", WORDS, WORD_MODE)
def position(el):
for p in el.findall("param"):
if p.get("name") == "Position":
return p
return None
p0 = position(titles[0])
position(titles[1]).set("value", p0.get("value"))
titles[1].set("enabled", "0")
report = modifier.validate_subtitle_layout()
assert report["summary"]["spatial_collision"] == 0
assert not blocking(report["severity"])
+2 -2
View File
@@ -46,7 +46,7 @@ class TestDiarizeMediaHandler:
await handle_diarize_media({"media_path": str(bad)}) await handle_diarize_media({"media_path": str(bad)})
async def test_writes_diarization_json_and_reports(self, tmp_path, monkeypatch): async def test_writes_diarization_json_and_reports(self, tmp_path, monkeypatch):
import server_tools._shared as _shared_mod import server_tools._shared.media as _shared_mod
import server_tools.voice as server_mod import server_tools.voice as server_mod
from server import handle_diarize_media from server import handle_diarize_media
@@ -86,7 +86,7 @@ class TestDiarizeMediaHandler:
assert data["words"][1]["speaker_id"] == "SPEAKER_01" assert data["words"][1]["speaker_id"] == "SPEAKER_01"
async def test_reports_when_diarization_fails(self, tmp_path, monkeypatch): async def test_reports_when_diarization_fails(self, tmp_path, monkeypatch):
import server_tools._shared as _shared_mod import server_tools._shared.media as _shared_mod
import server_tools.voice as server_mod import server_tools.voice as server_mod
from server import handle_diarize_media from server import handle_diarize_media
+16
View File
@@ -29,6 +29,7 @@ from fcpxml.text_layout import (
ink_extent, ink_extent,
) )
from fcpxml.writer import FCPXMLModifier from fcpxml.writer import FCPXMLModifier
from server_tools.subtitles import _words_overlapping_clip
SAMPLE = Path(__file__).parent.parent / "examples" / "sample.fcpxml" SAMPLE = Path(__file__).parent.parent / "examples" / "sample.fcpxml"
def font_points(style) -> float: def font_points(style) -> float:
@@ -51,6 +52,21 @@ WORDS = [
] ]
def test_words_overlapping_clip_keeps_word_that_starts_just_before_in_point():
words = [
{"word": "Aquela", "start": 2.03, "end": 2.69},
{"word": "mama", "start": 2.69, "end": 2.89},
{"word": "fora", "start": 10.0, "end": 10.2},
]
clip_words = _words_overlapping_clip(words, 2.0437166666666666, 3.0)
assert clip_words == [
{"word": "Aquela", "start": 0.0, "end": pytest.approx(0.6462833333333332)},
{"word": "mama", "start": pytest.approx(0.6462833333333332), "end": pytest.approx(0.8462833333333334)},
]
@pytest.fixture @pytest.fixture
def temp_fcpxml(): def temp_fcpxml():
with tempfile.NamedTemporaryFile(suffix=".fcpxml", delete=False) as f: with tempfile.NamedTemporaryFile(suffix=".fcpxml", delete=False) as f:
+4 -4
View File
@@ -419,7 +419,7 @@ class TestStillImageConversion:
result = _ensure_video_asset('/path/to/clip.mxf') result = _ensure_video_asset('/path/to/clip.mxf')
assert result == '/path/to/clip.mxf' assert result == '/path/to/clip.mxf'
@patch('fcpxml.writer.subprocess.run') @patch('fcpxml.writer.document.subprocess.run')
def test_png_triggers_conversion(self, mock_run): def test_png_triggers_conversion(self, mock_run):
mock_run.return_value = MagicMock(returncode=0) mock_run.return_value = MagicMock(returncode=0)
with tempfile.NamedTemporaryFile(suffix='.png', delete=False) as f: with tempfile.NamedTemporaryFile(suffix='.png', delete=False) as f:
@@ -436,7 +436,7 @@ class TestStillImageConversion:
if os.path.exists(mov_path): if os.path.exists(mov_path):
os.unlink(mov_path) os.unlink(mov_path)
@patch('fcpxml.writer.subprocess.run', side_effect=FileNotFoundError) @patch('fcpxml.writer.document.subprocess.run', side_effect=FileNotFoundError)
def test_missing_ffmpeg_raises(self, mock_run): def test_missing_ffmpeg_raises(self, mock_run):
with tempfile.NamedTemporaryFile(suffix='.jpg', delete=False) as f: with tempfile.NamedTemporaryFile(suffix='.jpg', delete=False) as f:
jpg_path = f.name jpg_path = f.name
@@ -446,7 +446,7 @@ class TestStillImageConversion:
finally: finally:
os.unlink(jpg_path) os.unlink(jpg_path)
@patch('fcpxml.writer.subprocess.run', @patch('fcpxml.writer.document.subprocess.run',
side_effect=subprocess.TimeoutExpired(cmd='ffmpeg', timeout=120)) side_effect=subprocess.TimeoutExpired(cmd='ffmpeg', timeout=120))
def test_ffmpeg_timeout_raises_runtime_error(self, mock_run): def test_ffmpeg_timeout_raises_runtime_error(self, mock_run):
"""Timed-out ffmpeg must raise RuntimeError, not propagate raw TimeoutExpired.""" """Timed-out ffmpeg must raise RuntimeError, not propagate raw TimeoutExpired."""
@@ -458,7 +458,7 @@ class TestStillImageConversion:
finally: finally:
os.unlink(png_path) os.unlink(png_path)
@patch('fcpxml.writer.subprocess.run', @patch('fcpxml.writer.document.subprocess.run',
side_effect=subprocess.CalledProcessError( side_effect=subprocess.CalledProcessError(
1, 'ffmpeg', stderr=b'Invalid codec')) 1, 'ffmpeg', stderr=b'Invalid codec'))
def test_ffmpeg_failure_raises_runtime_error(self, mock_run): def test_ffmpeg_failure_raises_runtime_error(self, mock_run):
+232
View File
@@ -0,0 +1,232 @@
"""Tests for fcpxml/forced_align.py — optional phonetic forced alignment.
The dependency (whisperx) is not installed in CI, so the core contract under
test is graceful degradation: when whisperx is unavailable the aligner returns
the words unchanged. A second group injects a fake whisperx module to verify
the refined times are written back in order and that malformed results are
skipped rather than clobbering good timestamps.
"""
import sys
import types
from pathlib import Path
import pytest
from fcpxml.forced_align import ForcedAligner
def _words():
return [
{"word": "Um,", "start": 0.0, "end": 0.5, "confidence": 0.9},
{"word": "welcome", "start": 0.5, "end": 1.0, "confidence": 0.9},
{"word": "show.", "start": 1.5, "end": 2.5, "confidence": 0.8},
]
def _raw_segments():
return [
{
"text": "Um, welcome",
"start": 0.0,
"end": 1.0,
"words": _words()[:2],
},
{
"text": "show.",
"start": 1.5,
"end": 2.5,
"words": _words()[2:],
},
]
def _fake_whisperx(shift=0.4):
"""A stand-in whisperx module that "corrects" word starts by ``shift``."""
mod = types.SimpleNamespace()
def load_audio(path):
return [0.0]
def load_align_model(language_code, device, model_dir=None):
return ("MODEL", {"language": language_code})
def align(align_input, align_model, metadata, audio, device,
return_char_alignments=False, chunk_size=30):
segments = []
for seg in align_input:
new_words = []
for w in seg["words"]:
new_words.append(
{
"word": w["word"],
"start": w["start"] + shift,
"end": w["end"] + shift,
"score": w["score"],
}
)
segments.append({**seg, "words": new_words})
return {"segments": segments}
mod.load_audio = load_audio
mod.load_align_model = load_align_model
mod.align = align
return mod
class TestForcedAlignerDegradation:
def test_unavailable_when_whisperx_missing(self):
assert ForcedAligner.available() is False
def test_returns_words_unchanged_when_whisperx_missing(self, monkeypatch):
import builtins
real_import = builtins.__import__
def block(name, *a, **k):
if name == "whisperx":
raise ImportError("blocked")
return real_import(name, *a, **k)
monkeypatch.setattr(builtins, "__import__", block)
result = ForcedAligner().align(_words(), _raw_segments(), "x.wav", "en")
assert result == _words()
def test_skips_when_no_words(self):
assert ForcedAligner().align([], [], "x.wav", "en") == []
class TestForcedAlignerWithWhisperX:
@pytest.fixture
def whisperx(self, monkeypatch):
fake = _fake_whisperx(shift=0.4)
monkeypatch.setitem(sys.modules, "whisperx", fake)
return fake
def test_refines_timestamps_in_order(self, whisperx):
words = _words()
result = ForcedAligner().align(words, _raw_segments(), "x.wav", "en")
assert [w["start"] for w in result] == [0.4, 0.9, 1.9]
assert [w["end"] for w in result] == [0.9, 1.4, 2.9]
# The same dict objects are returned with times overwritten in place.
assert result[0]["start"] == 0.4
assert words[0]["start"] == 0.4
def test_caches_align_model_per_language(self, whisperx, monkeypatch):
calls = {"n": 0}
orig = whisperx.load_align_model
def counting(*a, **k):
calls["n"] += 1
return orig(*a, **k)
whisperx.load_align_model = counting
aligner = ForcedAligner()
aligner.align(_words(), _raw_segments(), "a.wav", "en")
aligner.align(_words(), _raw_segments(), "b.wav", "en")
assert calls["n"] == 1
def test_skips_unusable_word_times(self, monkeypatch):
fake = _fake_whisperx()
# Force one word to come back with None start (alignment failed).
real_align = fake.align
def broken(align_input, *a, **k):
out = real_align(align_input, *a, **k)
out["segments"][0]["words"][0]["start"] = None
return out
fake.align = broken
monkeypatch.setitem(sys.modules, "whisperx", fake)
words = _words()
result = ForcedAligner().align(words, _raw_segments(), "x.wav", "en")
# First word time untouched (None skipped), rest corrected.
assert result[0]["start"] == 0.0
assert result[1]["start"] == 0.9
def test_unexpected_exception_returns_original(self, monkeypatch):
fake = types.SimpleNamespace()
fake.load_audio = lambda p: [0.0]
fake.load_align_model = lambda *a, **k: ("M", {})
fake.align = lambda *a, **k: 1 / 0 # boom
monkeypatch.setitem(sys.modules, "whisperx", fake)
words = _words()
result = ForcedAligner().align(words, _raw_segments(), "x.wav", "en")
assert result == words
class TestTranscribeAlignmentFlag:
"""Wire-up: transcribe() reports whether forced alignment ran."""
def _install_fakes(self, monkeypatch, align_shift=0.4):
# faster_whisper
fw = types.SimpleNamespace()
class _Word:
def __init__(self, word, start, end, prob):
self.word = word
self.start = start
self.end = end
self.probability = prob
class _Seg:
def __init__(self, text, start, end, words):
self.text = text
self.start = start
self.end = end
self.words = words
class _Info:
language = "en"
duration = 2.5
class _Model:
def transcribe(self, path, language=None, word_timestamps=False, vad_filter=False):
seg = _Seg(
"Um, welcome show.",
0.0,
2.5,
[
_Word("Um,", 0.0, 0.5, 0.9),
_Word("welcome", 0.5, 1.0, 0.9),
_Word("show.", 1.5, 2.5, 0.8),
],
)
return iter([seg]), _Info()
fw.WhisperModel = lambda *a, **k: _Model()
monkeypatch.setitem(sys.modules, "faster_whisper", fw)
# whisperx (only needed when align=True)
wx = _fake_whisperx(shift=align_shift)
monkeypatch.setitem(sys.modules, "whisperx", wx)
# model_manager.get_models_dir
import fcpxml.model_manager as mm
monkeypatch.setattr(mm, "get_models_dir", lambda: Path("/tmp"))
def test_alignment_true_when_whisperx_present(self, monkeypatch, tmp_path):
self._install_fakes(monkeypatch)
f = tmp_path / "a.wav"
f.write_bytes(b"RIFF0000WAVE")
from fcpxml.transcribe import transcribe
result = transcribe(str(f), model_size="base", align=True)
assert result is not None
assert result["alignment"] is True
assert result["words"][0]["start"] == pytest.approx(0.4)
def test_alignment_false_when_disabled(self, monkeypatch, tmp_path):
self._install_fakes(monkeypatch)
f = tmp_path / "a.wav"
f.write_bytes(b"RIFF0000WAVE")
from fcpxml.transcribe import transcribe
result = transcribe(str(f), model_size="base", align=False)
assert result is not None
assert result["alignment"] is False
assert result["words"][0]["start"] == 0.0

Some files were not shown because too many files have changed in this diff Show More