Compare commits

19 Commits
Author SHA1 Message Date
João HenriqueandClaude Sonnet 5 e9a17c1b62 feat(rag): provisiona busca RAG do G-ART e corrige indexação que abortava em chunk grande
Cria rag/ (schema, busca híbrida densa+lexical com RRF em search.py/
search_gart.sh, SETUP.md) — o projeto já tinha admin/update_rag.py para
indexar, mas nenhuma forma de consultar o índice. Corrige admin/update_rag.py:
um chunk denso em tokens (code/fcpxml/font_metrics.py) estourava o contexto
do modelo de embedding e derrubava a transação inteira; agora só aquele
chunk é pulado. Banco rag_gart provisionado no rag-hub-db compartilhado e
primeira indexação completa rodada (304 arquivos, 1702 chunks).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-23 09:38:46 -04:00
João HenriqueandClaude Sonnet 5 d13f643ebc chore(fase0): higiene do repositório + corrige gitignore que escondia fcpxml/models/
Fase 0 do roteiro de reestruturação (Engine/docs/10_MAPA_REESTRUTURACAO.md):
move code/WHISPERX (2,6 GB de backups órfãos, sem uso ativo, sem
.gitmodules) para ~/Archives/G-ART-WHISPERX-backup fora do workspace git;
traz admin/ para o gate de lint de run_after_fix.sh; corrige
fcpxml/writer/adjustment.py, que gerava um wrapper <adjustment> inexistente
no DTD 1.13 (filtros agora vão direto no <clip>, na ordem exigida), com
teste de regressão novo.

Achado à parte: .gitignore tinha uma regra solta "models/" (pensada só
para o cache do Whisper em code/models/) que também escondia do git todo o
pacote fcpxml/models/ — nunca commitado, sem proteção nenhuma. Corrigida
para /code/models/, ancorada na raiz.

Docs atualizados no mesmo commit (02_MODULES, 09_MANUTENCAO,
10_MAPA_REESTRUTURACAO, 05_EXPERIENCIAS #34 e #36), conforme a regra do
CLAUDE.md.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-23 08:28:44 -04:00
João HenriqueandClaude Sonnet 5 0fdfe33613 fix(legendas): impede legenda comum sob composição dinâmica e duplicação ao regerar
Marca cada título gerado (dynamic/plain) em metadata para que regenerar
substitua a saída anterior em vez de empilhar, e usa os spans de ênfase
revisados (não os segmentos brutos do Whisper) como janela da composição
dinâmica, evitando que ela invada o trecho de legenda comum seguinte.
suppress_plain_under_dynamic corta qualquer sobra visível como rede de
segurança.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-22 21:51:02 -04:00
João HenriqueandClaude Sonnet 5 32d78d0f8d fix(qc): padding padrão do remove_media_silence de 0,05s para 0,2s
0,05s existia como margem de segurança contra cortar a palavra em cima,
mas um silêncio que essa ferramenta encontra costuma ser o respiro
natural antes de uma frase nova, não sujeira de edição — e 0,05s raspava
esse respiro quase todo.

Caso real (projeto Mastopexia): a pausa antes de "Com" tinha 0,567s no
áudio original; com padding 0,05 sobrou só ~0,1s no total (0,05 de cada
lado), colando o clipe seguinte a 5ms da palavra em vez de deixar uma
pausa perceptível. 0,2s alinha com a convenção já documentada para folga
em corte de fronteira de frase (editar-por-voz/06-texto-corte-marcador.md).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-26 17:15:44 -04:00
João HenriqueandClaude Sonnet 5 c99274895c feat(legendas): liga compound_subphrases por padrão no pipeline
generate_dynamic_subtitles e a metade dinâmica de generate_subtitles_by_emphasis
passam a empacotar cada sub-frase da legenda dinâmica num compound clip por
padrão (compound_subphrases=True), completando o wrap_titles_in_compound
e split_into_subphrases do commit anterior — que ainda não tinham chamador
em produção.

Também torna validate_subtitle_layout ciente de compound clips: media cada
grupo (spine principal + cada <media> de compound) no seu próprio espaço de
tempo, em vez de uma varredura .//title global — sem isso, âncoras de
compounds diferentes liam offset "0s" e acusavam colisão espacial entre
frases que nunca dividem a tela, só porque compartilham o mesmo zero de
tempo local.

Testado ponta a ponta na gravação real (Mastopexia): 12 compounds, 41
títulos todos empacotados, zero soltos, zero IDs duplicados, DTD válida.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-26 17:05:20 -04:00
João HenriqueandClaude Opus 5 688bdeddb6 feat(legendas): sub-frases por vírgula e empacotamento em compound clip
split_into_subphrases divide a frase na vírgula — onde a fala respira —
mas funde de volta o pedaço curto ("né?", "Então..."), que lê como parte
da frase anterior e não como bloco próprio.

wrap_titles_in_compound empacota os títulos de uma sub-frase num compound
clip, replicando a estrutura que o próprio Final Cut produz: o primeiro
título vira âncora do spine em offset 0, os demais penduram nele por lane,
e um ref-clip toma o lugar deles na lane original. Os offsets dos filhos
são rebaseados para o espaço de tempo da âncora, senão cada palavra
escorregaria pela diferença entre os dois start.

Junto: _filter_children_for_segment passa a filtrar também o <video> do
Clipe de Ajuste. Sem isso, cada corte subsequente duplicava o zoom em
todos os pedaços resultantes com o offset original intacto, e as cópias
desenhavam empilhadas na mesma posição da timeline.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-26 16:35:21 -04:00
João HenriqueandClaude Sonnet 5 635d1bb553 fix(fcpxml): tracking-shape id duplicado ao cortar clipe com Cinematic
Object-tracker/tracking-shape (dado de rastreamento de objeto preservado
do asset original) mantinha o mesmo id em cada deepcopy feito por
split_clip/cut_clip_ranges, e o FCP acabava rejeitando o arquivo com "ID
tr1 already defined" depois de vários cortes. Mesmo mecanismo do bug já
corrigido para text-style-def, agora coberto também para tracking-shape.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-26 16:03:31 -04:00
João HenriqueandClaude Sonnet 5 2ad5854570 fix(voz): sobra de fatia interior no corte, e legenda comum poluindo com clipe desativado
Dois problemas reais vistos no projeto Mastopexia:

1. cut_clip_ranges só absorvia um keep-segment curto no INÍCIO/FIM do
   clipe (a lógica já existente do #6). Um keep curto no MEIO (entre dois
   cuts, sem nenhum vizinho mantido pra herdar) nunca era absorvido —
   sobrava como clipe de vídeo de 0,07-0,23s na timeline. Generalizado
   pra qualquer posição, com limiar maior (6 frames / 0,3s, medido no
   material real) — no meio, o pedacinho é descartado (vira parte do
   corte ao redor), nas bordas continua sendo herdado pelo vizinho.

2. generate_subtitles_by_emphasis gerava a legenda comum inteira e
   desativava (enabled="0") onde a dinâmica cobre. Título desativado
   continua aparecendo como clipe riscado na timeline do Final Cut mesmo
   sem renderizar — um corte com bastante ênfase virava dezenas de clipes
   mortos poluindo a trilha (visto ao vivo pelo usuário: "ficou uma
   bosta"). Trocado por não gerar o bloco comum ali, em vez de gerar e
   desativar. Custo: reativar ênfase manualmente depois exige regenerar a
   legenda comum daquele trecho, não só reabilitar.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-21 18:59:21 -04:00
João HenriqueandClaude Sonnet 5 8257155fd3 fix(voz): frases desativadas em sequência deixavam fatias sobrando no corte
phrase_review_to_actions() cortava cada frase desativada isoladamente
(start..end da própria frase) — quando várias seguidas estavam desativadas,
a pausa ENTRE elas não pertencia a nenhuma frase e sobrevivia como um
clipe minúsculo (0,1-0,5s) na timeline final. Confirmado no projeto
Mastopexia real: 29 cuts individuais geravam mais de uma dezena de fatias
sub-segundo; agrupar frases desativadas consecutivas num único cut (do
início da primeira ao fim da última) reduziu para 3 cuts e 4 fatias
residuais (menores, provavelmente do padding do remove_media_silence —
registrado como dívida separada em 09_MANUTENCAO.md §2.5).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-21 18:49:13 -04:00
João HenriqueandClaude Sonnet 5 fd791e116a docs(skill): corte deve deixar folga na borda que encosta em fala mantida
Cortes escritos rente ao timestamp da palavra soavam secos (relatado no
projeto Mastopexia) — o critério e o prompt do modelo local mandavam cobrir
a frase inteira sem orientar a borda que toca fala mantida. Adiciona a regra
de recuar ~0,15-0,25s nas duas pontas quando o corte encosta em conteúdo
que fica, tanto no skill (06-texto-corte-marcador.md) quanto no prompt
embutido do Ollama (llm_local.py) — pra não precisar ajustar na mão de novo.

Também registra em 05_EXPERIENCIAS.md/09_MANUTENCAO.md a dívida de
resolve_actions não tolerar margem quando zoom/marker encosta na borda
de um corte (contornado manualmente, não corrigido em código ainda), e
atualiza a lista de dívidas abertas (etapa 6/offset de whisper já resolvidos
nesta sessão).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-21 18:31:27 -04:00
João HenriqueandClaude Sonnet 5 7b5aed79ee feat(voz): legenda por ênfase, forced align, IA local e correções de zoom/revisão
Trabalho da branch feat/revisao-enfases: pipeline de edição por voz ganha
alinhamento forçado (whisperx), roteirização por LLM local (Ollama), e a
etapa 5 (revisão de frases) passa a refletir de verdade o que é aplicado.

- generate_subtitles_by_emphasis: legenda comum cobre o clipe inteiro,
  legenda dinâmica só nas frases de ênfase, e a comum é desativada
  (enabled="0") onde a dinâmica cobre, em vez de nunca ser gerada ali.
- validate_subtitle_layout ignora títulos com enabled="0" — corrige falso
  positivo de colisão contra o que está desativado no lugar dele.
- Corrige zoom/marcador sendo descartado quando a borda encosta exatamente
  no início de um corte.
- Etapa 5 do Assistente: recarrega quando as decisões da IA mudam (com
  fresh=true, ignorando a revisão salva antiga) — resolve a dessincronia
  entre "ativa" na tela e o que já foi cortado no FCPXML.
- Etapa "Processar" reaplica as decisões da revisão (_phrase_actions.json)
  antes da cadeia de remoção de silêncio/legendas — antes, desativar uma
  frase na etapa 5 não tinha efeito nenhum no vídeo final.
- Etapa "Concluído" fundida em "Processar" — abrir no Final Cut/Finder
  aparece assim que termina, sem slide extra.
- Palavra clicável na etapa 5 agora funciona como toggle (clique de novo
  desfaz) e mostra a própria ênfase (sublinhado colorido + peso da fonte).
- fcpxml/forced_align.py, fcpxml/llm_local.py, ai_edit.py: alinhamento
  fonético via whisperx e roteirização local via Ollama/Gemma.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-21 18:26:04 -04:00
João HenriqueandClaude Opus 5 711c397dfe fix: admin/api apontava para admin/code (inexistente) — crash no app
Ao dividir _shared.py em admin/api/*.py ontem, o cálculo
`Path(__file__).resolve().parent.parent / "code"` foi copiado sem ajustar
para o nível de diretório novo. No arquivo original (admin/models_api.py,
direto em admin/) dois `.parent` chegavam na raiz do repo. Em
admin/api/shared.py, um nível mais fundo, dois `.parent` param em admin/ —
e admin/code nunca existiu. sys.path nunca recebia code/, então toda ação
que passa por `server` (analisar voz, aplicar decisões) crashava o app com
ModuleNotFoundError: server_tools.

O bug sobreviveu a duas rodadas de validação da sessão anterior — lint
zero, 1454 testes verdes, comando testado manualmente pela ponte — porque
todos rodam num venv com install editável (__editable__.fcp_mcp_server.pth)
que já deixa fcpxml/server_tools importáveis por conta própria, mascarando
qualquer erro no cálculo manual de sys.path. Só o app real, no fallback sem
uv, expõe o bug.

Correção: o cálculo de sys.path sai de cada módulo de comando (estava
duplicado em nove arquivos) e passa a existir uma única vez em
admin/api/__init__.py, que roda antes de qualquer submódulo — nenhum
precisa mais da própria cópia.

O teste de regressão precisou de duas tentativas pelo mesmo motivo do bug:
a primeira versão também passava com o bug presente, por rodar no mesmo
venv "de sorte". Só ficou confiável isolando um subprocess que remove
site-packages do sys.path antes de importar — confirmado nos dois sentidos,
falha com o bug reintroduzido e passa com a correção
(TestCodeDirResolution).

Detalhe completo, incluindo por que o comando manual não pegou:
Engine/docs/05_EXPERIENCIAS.md #25.

Lint zerado, 1457 testes passando (3 novos), app compilado.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-20 09:50:04 -04:00
João HenriqueandClaude Opus 5 cbd9297751 docs(skill): alinhar editar-por-voz com a revisão humana da etapa 5
A skill decidia a edição sem saber que o JSON dela agora passa por uma tela
de revisão antes de virar FCPXML. Isso não é detalhe de fluxo: a etapa 5
traduz cada ação para o vocabulário dela, e sem conhecer essa tradução a
intenção da IA se perde no caminho — que é exatamente como uma decisão vira
"arbitrária" aos olhos de quem revisa.

Novo criterios/10-revisao-humana.md, com o que o app faz com cada ação:

- cut cobrindo >=60% da frase remove a linha; tocando só uma borda vira trim
  encaixado na fronteira de palavra. Corte de meia frase é ambíguo — passa
  do limiar e apaga a linha toda quando a intenção era aparar a hesitação.
- zoom ou text sobre uma frase marca ênfase, e ênfase significa DUAS coisas:
  zoom mais legenda dinâmica; as demais frases ficam com legenda comum. A
  escala vira o nível (1.15→leve, 1.3→média, 1.5→forte).
- sem ação, o nível é derivado do peak_emphasis; a decisão da IA sempre ganha.
- reason é exibido ao lado da frase na tela — é o que o editor lê antes de
  manter ou desfazer. Deixou de ser campo de log.

Consequência prática que faltava em 05-zoom.md: não espalhar zoom "por
segurança", porque cada um promove a frase em duas dimensões ao mesmo tempo.
Na dúvida, deixar sem — promover custa uma tecla, despromover custa mais.

Cada arquivo de critério ganhou cabeçalho de escopo (o que cobre, em que
fase), no mesmo padrão dos docs do Engine, para ler só o necessário.

Todas as afirmações numéricas do novo critério foram verificadas contra
fcpxml/phrase_review.py rodando, não assumidas.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-19 22:54:51 -04:00
João HenriqueandClaude Opus 5 dcdd73edb5 docs: varredura geral, documentação por função e regra de atualização
A documentação descrevia um sistema que não existe mais: 62/73 ferramentas
(são 74), writer.py e models.py como arquivos (viraram pacotes), 1032 testes
(são 1454), models_api.py descrito como "API FastAPI" (é ponte JSON) e o app
SwiftUI ausente por completo — 5.500 linhas que o usuário opera todo dia sem
uma linha de documentação.

Cada arquivo passa a ter uma função específica, com cabeçalho de escopo
dizendo o que cobre e o que NÃO cobre (com a seta para quem cobre). O objetivo
é ler só o necessário: doc fora do assunto custa tempo e processamento sem
entregar nada.

    01 arquitetura   camadas, duas portas de entrada, regras transversais
    02 módulos       mapa do engine, incluindo o pipeline de voz
    03 server/tools  as 74 tools, helpers e como criar uma nova
    08 app macOS     NOVO — build por swiftc, telas, ponte, etapa 5
    09 manutenção    NOVO — por onde começar, o que está aberto, sintoma→arquivo

CLAUDE.md ganha a seção "Documentação (MANDATORY)": tabela de roteamento
(qual arquivo abrir para cada tarefa) e a regra de que toda alteração de
código atualiza a doc no mesmo commit, com o mapa de o-que-mexeu → o-que-
atualizar. Doc velha engana mais que doc ausente.

O índice do 05_EXPERIENCIAS subiu para o topo: consultar "isso já quebrou
antes?" custava carregar 1.281 linhas antes de chegar na tabela.

Dívidas levantadas na varredura e registradas em 09 §2: etapa 6 ainda ignora
o phrase_review.json, offset de ~400ms do Whisper, MacApp sem teste, admin/
fora do lint, confirmações visuais pendentes no FCP, submódulo WHISPERX sujo.

Também corrigidos dois links quebrados no Engine/README que apontavam um
nível acima do certo.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-19 22:51:33 -04:00
João HenriqueandClaude Opus 5 ffaebb3f72 refactor: _shared.py vira subpacote, um módulo por papel
Eram 882 linhas de seis papéis sem relação, sob um nome que só dizia
"compartilhado" — o depósito onde tudo que servia a mais de um handler
acabava caindo.

    media       316   transcrição em cache, corte por fala, relatório
    paths       206   sandbox, limites, caminho de saída
    project     116   abrir projeto, preparar modifier/generator
    captions    112   SRT, VTT, listas com timestamp
    detection    99   flash frames, buracos, duplicados
    formatting   86   tabelas e relatórios dos handlers

O __init__ reexporta os 46 nomes, então os treze pontos que importam daqui
não mudaram.

_transcript_cut_report saiu de formatting para media: ele precisa do hint de
instalação e do _text_result, ou seja, é relatório de transcrição e não
formatação genérica — mover foi mais honesto que cruzar imports entre os
dois módulos.

Quatro testes patchavam `server_tools._shared.transcribe`; o nome agora é
ligado por _shared/media.py, então o patch passou a apontar para lá — mesmo
padrão da experiência #23.

Lint zerado, 1454 testes passando.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-19 22:35:02 -04:00
João HenriqueandClaude Opus 5 368bb62706 refactor: models.py vira pacote, um módulo por família
Eram 1.091 linhas com seis famílias de modelo sem relação entre si —
enumerações, tempo racional, timeline, geração, QC e legendas.

    timing     304   TimeValue e Timecode
    timeline   217   clipes, marcadores, lanes, projeto
    enums      183   tipos/cores de marcador, transições, ritmo
    subtitles  157   paleta e look das legendas dinâmicas
    qc         121   achados de QC e resultado de validação
    planning    93   rough cut, ritmo, montagem

O __init__ reexporta os 43 nomes, incluindo os com underscore que o writer
e a suíte já importavam, então nenhum ponto de uso mudou.

Lint zerado, 1454 testes passando.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-19 22:31:18 -04:00
João HenriqueandClaude Opus 5 6090e229e9 refactor: models_api.py vira ponto de entrada sobre admin/api/
A ponte JSON do app tinha 1.395 linhas e 37 comandos de oito assuntos
diferentes num arquivo só. Agora models_api.py guarda apenas a referência
dos comandos, a tabela de despacho e o main(); cada assunto virou um módulo
em admin/api/ (models, project, editing, zoom, subtitles, transcription,
voice, review), com a base comum em shared.py.

Nada muda para o app: ele continua chamando admin/models_api.py por caminho,
e os 37 comandos respondem igual — verificado rodando a ponte de verdade.

Duas coisas que a divisão obrigou a arrumar:

- A saída passa por `shared.emit` chamada pelo módulo, não pelo nome
  importado. Isso preserva a propriedade de que trocar `emit` num lugar só
  captura a saída de todos os comandos — que era acidental quando tudo
  morava no mesmo arquivo, e vira intencional agora.
- `_CANCEL` e o lock eram globais compartilhados. O registro de downloads
  foi para models.py, junto de quem o usa, com lock próprio: o antigo
  protegia ao mesmo tempo o dicionário e a escrita em stdout, duas coisas
  sem relação.

Também: admin/test_models_api.py estava fora de `testpaths` e nunca rodava.
Movido para code/tests/ e ligado ao gate — 1441 → 1454 testes
(ver Engine/docs/05_EXPERIENCIAS.md #24).

Lint zerado, 1454 testes passando.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-19 21:46:47 -04:00
João HenriqueandClaude Opus 5 4f5cf94443 refactor: writer.py vira pacote, um módulo por assunto
O writer tinha 4.199 linhas, das quais 3.300 numa única classe com dezoito
assuntos dentro. Achar o trecho de zoom exigia rolar por marcadores,
velocidade e legendas.

Agora é o pacote fcpxml/writer/, com um arquivo por assunto e o
FCPXMLModifier montado por composição de mixins. Mixins, e não objetos
separados, porque todas essas operações mexem no mesmo documento e nos
mesmos índices — separá-las em objetos independentes transformaria toda
chamada interna em travessia de fronteira sem nada em troca. A divisão que
importa aqui é de leitura, não de estado.

Nenhuma mudança de comportamento e nenhuma alteração nos ~50 pontos que
importam do writer: o __init__ re-exporta tudo, inclusive os nomes com
underscore que a suíte já usava.

    core      723   carga, índices, navegação na spine, save
    titles    600   títulos e legendas dinâmicas
    cut       333   dividir, cortar faixas, apagar
    speed     297   velocidade e zoom
    (+ 20 módulos menores)

Único ajuste de chamada: quatro testes faziam patch em
fcpxml.writer.subprocess, que agora mora em writer.document (ver
Engine/docs/05_EXPERIENCIAS.md #23).

Lint zerado, 1441 testes passando.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-19 21:38:49 -04:00
João HenriqueandClaude Opus 5 1bebee4359 feat: etapa 5 do assistente — revisão de ênfases com timeline
Transforma a etapa "colar decisões" numa tela de lapidação: a sugestão da
IA chega carregada e o editor afina frase a frase o que é ênfase e o que
fica fora. Essa marcação é o norte da etapa 6 — só as frases com ênfase
recebem zoom e legenda dinâmica; as demais ficam com legenda comum.

O campo de colar o JSON sobe para a etapa 4, então a numeração das etapas
não muda e a etapa 6 segue intacta.

Backend (fcpxml/phrase_review.py):
- build_phrase_review funde o _voice_timeline.json com as actions da IA
- trim por frase que anda em fronteira de palavra; corte parcial da IA
  chega como trim em vez de ser arredondado fora
- phrase_review_to_actions volta a cuts/zooms + emphasis_spans
- merge_saved_decisions reaplica só as decisões salvas sobre uma revisão
  remontada da análise atual, para reprocessar a voz não ficar mascarado
- resolve_source acha a mídia: o voice timeline guarda só o nome do arquivo

App (SwiftUI):
- layout de sala de edição: preview em cima, inspector à direita, timeline
  atravessando embaixo com seis trilhas rotuladas
- preview enquadra no formato de entrega lido do .fcpxml (fonte horizontal,
  projeto vertical), com alternância para a mídia original
- reprodução pula os trechos removidos e para no fim do trecho
- zoom manual por trecho marcado, sem guardar escala: a forma vem das
  configurações de Análise de Voz no render
- emoção da fala exposta por frase

Correções encontradas no caminho:
- VideoPlayer (AVKit) aborta em runtime no app compilado por swiftc;
  trocado por AVPlayerLayer (ver Engine/docs/05_EXPERIENCIAS.md #22)
- teste que ainda afirmava o default zoom scale=1.3 removido do parser (#21)

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-19 21:29:27 -04:00
139 changed files with 19782 additions and 8004 deletions
+25 -2
View File
@@ -28,6 +28,22 @@ minutos.
`apply_voice_actions` aplica direto, mas é para **teste**. O produto do seu `apply_voice_actions` aplica direto, mas é para **teste**. O produto do seu
trabalho é a lista de decisões. trabalho é a lista de decisões.
**Caminho automatizado (sem wizard, sem copiar-e-colar):** a tool
`generate_voice_script` (MCP) / comando `generate_voice_script` (ponte do app)
corre o fluxo fechado: transcreve → `build_voice_timeline` → entrega a timeline
a um **modelo local Ollama (Gemma 3 / Llama)** que age exatamente como este
skill descreve (separa roteiro de bastidor, escolhe tomadas, decide zoom/corte)
→ devolve o roteiro legível **e** o JSON de ações, e opcionalmente aplica no
FCPXML. O cliente fica em `code/fcpxml/llm_local.py`; o prompt que embute este
contrato está em `_SYSTEM_PROMPT`. Use essa tool quando o usuário pedir para
"rodar tudo internamente" ou "gerar o roteiro por IA local".
**Para onde ela vai (modo manual):** o usuário cola o seu JSON no app, e ele
abre na etapa 5 do Assistente — uma tela onde cada frase do roteiro aparece com
a sua decisão já marcada, para ser revisada antes de gerar. Você é o **ponto de
partida** da edição, não a palavra final; escreva decisões defensáveis e motivos
legíveis. Como o app traduz cada ação sua: `criterios/10-revisao-humana.md`.
## Ordem de trabalho ## Ordem de trabalho
Siga nesta ordem. Pular a Fase 2 ou a 3 leva a decisões erradas. Siga nesta ordem. Pular a Fase 2 ou a 3 leva a decisões erradas.
@@ -44,6 +60,9 @@ Siga nesta ordem. Pular a Fase 2 ou a 3 leva a decisões erradas.
| **7** | Cortar a lista pelo ritmo | `criterios/07-ritmo.md` | | **7** | Cortar a lista pelo ritmo | `criterios/07-ritmo.md` |
| **8** | Montar o JSON de saída | `criterios/08-formato-de-saida.md` | | **8** | Montar o JSON de saída | `criterios/08-formato-de-saida.md` |
**Antes da Fase 5, leia `criterios/10-revisao-humana.md`.** Ele descreve o que
o app faz com o seu JSON — e muda *como* escrever cortes e zooms, não só quais.
## As três armadilhas ## As três armadilhas
Cada uma já causou erro silencioso em material real: Cada uma já causou erro silencioso em material real:
@@ -69,10 +88,14 @@ disso e o efeito cai no frame errado — sem erro visível.
``` ```
build_voice_timeline → [você decide] → refine_voice_timeline → [você corta build_voice_timeline → [você decide] → refine_voice_timeline → [você corta
pelo ritmo] → apply_voice_actions → remove_media_silence → pelo ritmo] → [revisão humana na etapa 5 do app] → apply_voice_actions →
generate_dynamic_subtitles remove_media_silence → generate_dynamic_subtitles
``` ```
A revisão humana entra entre a sua decisão e a aplicação. É por isso que o
`reason` importa tanto: ele é lido ali, na hora de decidir se a sua escolha
fica.
Vícios de linguagem e lacunas longas entram na **sua** lista, num ripple só Vícios de linguagem e lacunas longas entram na **sua** lista, num ripple só
(`06-texto-corte-marcador.md`). Silêncio fino e legendas vêm depois. (`06-texto-corte-marcador.md`). Silêncio fino e legendas vêm depois.
@@ -1,5 +1,8 @@
# 01 — Leitura do JSON # 01 — Leitura do JSON
> **Escopo:** Como ler o voice_timeline em camadas, sem recalcular o que já foi medido.
> **Quando:** Fase 1 — ver a ordem de trabalho em `../SKILL.md`.
O arquivo `<mídia>_voice_timeline.json` é a entrada de todo o trabalho. O arquivo `<mídia>_voice_timeline.json` é a entrada de todo o trabalho.
Leia em camadas, de cima para baixo, e só desça quando precisar. Leia em camadas, de cima para baixo, e só desça quando precisar.
@@ -50,14 +53,19 @@ use para decidir; existem para permitir a reanálise da Fase 2.
## O timestamp por palavra tem um viés conhecido ## O timestamp por palavra tem um viés conhecido
O início de cada palavra vem sistematicamente **adiantado em ~0,3-0,5s** em O início de cada palavra vem sistematicamente **adiantado em ~0,3-0,5s** em
relação ao ataque real da fala — medido em material real com ffmpeg relação ao ataque real da fala — medido em material real com ffmpeg (`astats`),
(`astats`), consistente em 6 pontos do mesmo vídeo. O fim da palavra não consistente em 6 pontos do mesmo vídeo. O fim da palavra não tem esse problema
tem esse problema (erro de poucos centésimos). Causa: `word_timestamps` do (erro de poucos centésimos). Causa: `word_timestamps` do faster-whisper deriva
faster-whisper deriva por atenção cruzada, sem alinhamento forçado — ver por atenção cruzada, sem alinhamento forçado — ver `05_EXPERIENCIAS.md`, entrada
`05_EXPERIENCIAS.md`, entrada de 2026-08-19. de 2026-08-19.
**Quando o pipeline já corrigiu isso:** se `layers.alignment` for `true`
(transcript gerado com alinhamento forçado fonético via whisperx, implementado
depois desse aviso), o viés foi removido na origem — **não aplique o offset
manual** abaixo. O aviso vale só para transcripts antigos sem `layers.alignment`.
Isso não é "reestimar no olho" — é um bug de medição na fonte, não um Isso não é "reestimar no olho" — é um bug de medição na fonte, não um
julgamento seu. Na prática: julgamento seu. Na prática (somente sem `layers.alignment`):
- Ao posicionar um `zoom` cujo `start` precisa cair exatamente na palavra - Ao posicionar um `zoom` cujo `start` precisa cair exatamente na palavra
(não uma frase inteira), some **+0,3 a +0,4s** ao timestamp do JSON antes (não uma frase inteira), some **+0,3 a +0,4s** ao timestamp do JSON antes
@@ -66,6 +74,3 @@ julgamento seu. Na prática:
- **Não aplique essa correção a `gap_before` para decidir corte** — a régua - **Não aplique essa correção a `gap_before` para decidir corte** — a régua
de silêncio (`06-texto-corte-marcador.md`) já é conservadora o bastante de silêncio (`06-texto-corte-marcador.md`) já é conservadora o bastante
para absorver esse erro; corrigir os dois ao mesmo tempo é redundante. para absorver esse erro; corrigir os dois ao mesmo tempo é redundante.
- Se um dia o pipeline ganhar alinhamento forçado (WhisperX), este aviso
perde a razão de existir — confira se `layers` ou a versão do documento
já indicam isso antes de aplicar o offset manualmente.
@@ -1,5 +1,8 @@
# 02 — Triagem: roteiro vs. conversa de bastidor # 02 — Triagem: roteiro vs. conversa de bastidor
> **Escopo:** Separar o texto do roteiro da conversa de bastidor — tarefa de texto, nunca de limiar.
> **Quando:** Fase 2 — ver a ordem de trabalho em `../SKILL.md`.
**Primeira coisa a fazer, antes de qualquer decisão de efeito.** **Primeira coisa a fazer, antes de qualquer decisão de efeito.**
Material bruto de gravação quase nunca é uma tomada só. A pessoa lê o Material bruto de gravação quase nunca é uma tomada só. A pessoa lê o
@@ -1,5 +1,8 @@
# 03 — Escolha da melhor tomada # 03 — Escolha da melhor tomada
> **Escopo:** Qual tomada de cada frase sobrevive, e o que fazer em caso de empate.
> **Quando:** Fase 3 — ver a ordem de trabalho em `../SKILL.md`.
A mesma frase costuma aparecer 2, 3, 4 vezes. Seu trabalho é ficar com A mesma frase costuma aparecer 2, 3, 4 vezes. Seu trabalho é ficar com
**uma**. **uma**.
@@ -1,5 +1,8 @@
# 04 — Reanálise do material que sobrou # 04 — Reanálise do material que sobrou
> **Escopo:** Renormalizar a ênfase sobre o que sobrou, antes de escolher zooms.
> **Quando:** Fase 4 — ver a ordem de trabalho em `../SKILL.md`.
**Não escolha zooms com os números da análise bruta.** **Não escolha zooms com os números da análise bruta.**
## O problema ## O problema
@@ -1,5 +1,8 @@
# 05 — Zoom (punch-in) # 05 — Zoom (punch-in)
> **Escopo:** Onde dar punch-in, qual janela e qual escala — e o que a escala significa além do zoom.
> **Quando:** Fase 5 — ver a ordem de trabalho em `../SKILL.md`.
## Quando usar ## Quando usar
No momento em que o argumento vira. Um pico acústico só merece zoom se for No momento em que o argumento vira. Um pico acústico só merece zoom se for
@@ -46,14 +49,24 @@ automático acerta na quase totalidade dos casos.
## Escala ## Escala
| Valor | Uso | | Valor | Uso | Vira, na tela de revisão |
|---|---| |---|---|---|
| 1,15 | sutil | | 1,15 | sutil | ênfase **1 — Leve** |
| 1,18 – 1,3 | padrão | | 1,18 – 1,3 | padrão | ênfase **2 — Média** |
| 1,5 | forte | | 1,5 | forte | ênfase **3 — Forte** |
Em vídeo institucional, fique na faixa baixa. Acima de 3,0 é rejeitado. Em vídeo institucional, fique na faixa baixa. Acima de 3,0 é rejeitado.
**A escala tem um segundo efeito, e ele é maior que o zoom.** A frase que
recebe um zoom é marcada como **ênfase** na etapa 5, e frase de ênfase recebe
**legenda dinâmica**; as demais ficam com legenda comum. Ou seja: escolher onde
dar zoom é também escolher onde o texto ganha tratamento tipográfico.
Consequência prática: **não espalhe zoom "por segurança"**. Cada um promove uma
frase a destaque em duas dimensões ao mesmo tempo. Na dúvida, deixe sem — o
editor promove numa tecla, e despromover custa mais que promover.
Detalhe: `10-revisao-humana.md`.
O zoom é **relativo ao enquadramento existente**: se o clipe já tem escala O zoom é **relativo ao enquadramento existente**: se o clipe já tem escala
1,77 (material gravado de lado e reenquadrado), um zoom 1,18 anima de 1,77 1,77 (material gravado de lado e reenquadrado), um zoom 1,18 anima de 1,77
para 2,09 e preserva rotação e posição. para 2,09 e preserva rotação e posição.
@@ -1,5 +1,8 @@
# 06 — Texto, corte e marcador # 06 — Texto, corte e marcador
> **Escopo:** Texto na tela, o que cortar (inclui muletas e lacunas) e quando marcar.
> **Quando:** Fase 6 — ver a ordem de trabalho em `../SKILL.md`.
## Texto ## Texto
Para fixar um **conceito, número ou nome** que o espectador precisa reter. Para fixar um **conceito, número ou nome** que o espectador precisa reter.
@@ -50,6 +53,34 @@ Acima de 3s a pausa deixa de contar como ênfase por construção — medido em
material real, lacunas de 6–9s rankeavam como os momentos mais enfáticos da material real, lacunas de 6–9s rankeavam como os momentos mais enfáticos da
gravação só porque a escala saturava. gravação só porque a escala saturava.
### Nunca corte rente à palavra — deixe uma folga
Um `cut` cujo `start`/`end` cai exatamente no timestamp da palavra (fim da
última palavra mantida = início do corte) produz um corte seco: a palavra é
engolida antes de terminar de soar, e a fala seguinte começa sem nenhum ar.
Isso é diferente de cortar a pausa curta (que seria apagar a própria ênfase,
proibido acima) — aqui a pausa **já existe** entre o fim de um bloco mantido
e o início do próximo, e o corte está comendo justamente essa margem.
Ao escrever a borda de um `cut` que encosta em fala mantida (não em silêncio
puro), recue **~0,15–0,25s** para dentro do próprio corte, nos dois lados:
- o `start` do corte fica ~0,2s **depois** do fim real da última palavra
mantida;
- o `end` do corte fica ~0,2s **antes** do início real da próxima palavra
mantida.
Caso real (projeto Mastopexia): um corte escrito rente (`10.77 → 95.50`,
exatamente nos timestamps de palavra) soava abrupto nas duas emendas.
Recuado para `10.97 → 95.30`, cada lado ganhou ~0,2s de respiro sem alterar
o que é dito — e não empurra o próximo zoom/marcador contra a borda do corte
(ver `05-zoom.md` sobre janelas encostadas em corte).
Isso vale também para o **início e o fim do vídeo**: ar morto antes da
primeira palavra e depois da última também leva `cut`, com a mesma folga —
não é "silêncio dentro da fala" (isso é `remove_media_silence`), é o mesmo
corte de tomada/bastidor que você já está decidindo.
### O que continua NÃO sendo seu trabalho ### O que continua NÃO sendo seu trabalho
| Tarefa | Ferramenta | Por quê | | Tarefa | Ferramenta | Por quê |
@@ -1,5 +1,8 @@
# 07 — Ritmo # 07 — Ritmo
> **Escopo:** Quantos efeitos cabem: os tetos e como escolher o que fica.
> **Quando:** Fase 7 — ver a ordem de trabalho em `../SKILL.md`.
**O erro mais comum é efeito demais.** Cansa mais que efeito de menos, e **O erro mais comum é efeito demais.** Cansa mais que efeito de menos, e
denuncia edição automática. denuncia edição automática.
@@ -1,5 +1,8 @@
# 08 — Formato de saída # 08 — Formato de saída
> **Escopo:** O JSON de entrega: estrutura, regras e como o programa trata erros.
> **Quando:** Fase 8 — ver a ordem de trabalho em `../SKILL.md`.
O produto do seu trabalho é **este JSON**. É ele que vai para o programa O produto do seu trabalho é **este JSON**. É ele que vai para o programa
gerar o FCPXML. Você nunca escreve XML. gerar o FCPXML. Você nunca escreve XML.
@@ -53,6 +56,19 @@ uma. Um `reason` vazio é sinal de decisão sem critério.
Inclua o dado que embasou: *"abertura: 'Aquela mama' (ênfase 0.42)"* é útil; Inclua o dado que embasou: *"abertura: 'Aquela mama' (ênfase 0.42)"* é útil;
*"zoom"* não é. *"zoom"* não é.
Não é campo de log: o texto é **exibido na tela de revisão**, ao lado da frase,
e é o que o editor lê antes de manter ou desfazer o que você decidiu.
### 6. Corte: alinhe à intenção
A tela lê cada `cut` contra as frases da transcrição:
- cobre **≥ 60%** de uma frase → aquela frase é **removida**;
- toca só o **começo** ou só o **fim** → vira **trim** (a frase fica, aparada).
Então corte a frase **inteira** quando quiser removê-la, e corte **só da borda
até a palavra** quando quiser aparar uma hesitação. Um corte de meia frase é
ambíguo — passa de 60% e apaga a linha toda. Detalhe: `10-revisao-humana.md`.
## Como o programa trata erros ## Como o programa trata erros
- **Ação inválida** → rejeitada e reportada **individualmente**. Uma linha - **Ação inválida** → rejeitada e reportada **individualmente**. Uma linha
@@ -1,5 +1,8 @@
# 09 — Quando a análise veio incompleta # 09 — Quando a análise veio incompleta
> **Escopo:** O que fazer quando uma camada da análise não rodou.
> **Quando:** Fase 0 — ver a ordem de trabalho em `../SKILL.md`.
O bloco `layers` no topo do JSON diz **o que de fato rodou**. Leia antes de O bloco `layers` no topo do JSON diz **o que de fato rodou**. Leia antes de
qualquer outra coisa. qualquer outra coisa.
@@ -0,0 +1,121 @@
# 10 — A revisão humana: o que acontece com o seu JSON
> **Escopo:** O que o app faz com o seu JSON na etapa 5 — muda como escrever as ações.
> **Quando:** ler antes da Fase 5 — ver a ordem de trabalho em `../SKILL.md`.
> Leia antes de decidir cortes e zooms. Muda **como** escrever as ações, não
> apenas quais.
Seu JSON não vai direto para o FCPXML. Ele é colado no app e abre na **etapa 5
do Assistente**, uma tela onde o editor vê cada frase do roteiro com a sua
decisão já aplicada e lapida antes de gerar.
Isso tem duas consequências práticas:
1. **Suas decisões são lidas por uma pessoa, frase a frase.** Uma decisão sem
motivo explícito parece arbitrária — e será desfeita.
2. **A tela traduz suas ações para o vocabulário dela.** Se você não escrever
as ações do jeito que essa tradução espera, a intenção se perde no caminho.
---
## Como cada ação sua é lida
O app quebra a gravação em **frases** (os segmentos do voice timeline) e
projeta suas ações sobre elas.
### `cut`
| O corte cobre… | Vira | Na tela |
|---|---|---|
| **≥ 60%** da frase | frase **desativada** | apagada, riscada, reativável num clique |
| só o **começo** ou só o **fim** | **trim** da frase | a frase fica, aparada nas pontas |
| um pedaço no **meio** | nada em si | só conta para a regra dos 60% |
O trim é **encaixado na fronteira de palavra** mais próxima. Você não precisa
acertar o frame: mire na palavra onde a frase deve começar ou terminar.
**O que isso pede de você:** decida se está removendo *a linha* ou *aparando*
uma ponta, e escreva o corte de acordo.
- Removendo a linha → corte a frase inteira, de ponta a ponta.
- Aparando um falso começo → corte só da borda até a palavra onde a fala
engata. Um corte que cobre meia frase é ambíguo: passa de 60% e apaga a linha
toda, quando você só queria tirar a hesitação.
### `zoom` e `text`
Qualquer `zoom` ou `text` que toque uma frase marca aquela frase como
**ênfase** — e ênfase, nesta tela, significa **duas coisas**:
> **A frase de ênfase recebe zoom E legenda dinâmica. As demais recebem
> legenda comum.**
O nível vem da sua `scale`:
| `scale` | Nível na tela | |
|---|---|---|
| 1,15 | 1 — Leve | |
| 1,3 | 2 — Média | |
| 1,5 | 3 — Forte | |
| omitida, ou uma ação `text` | 2 — Média | padrão |
Sem nenhuma ação sua, a tela deriva o nível do `peak_emphasis` da frase
(< 0,25 → sem ênfase; < 0,45 → leve; < 0,65 → média; acima → forte). **A sua
decisão sempre ganha da derivação automática.**
**O que isso pede de você:** escolher a escala com intenção. Ela não é só
"quanto amplia" — é o peso que aquela frase terá no vídeo inteiro, incluindo o
tratamento da legenda. Um zoom leve numa frase de apoio não é neutro: promove
aquela frase a destaque tipográfico também.
### `marker`
Não altera a frase. Continua sendo o seu recado para o editor conferir uma
emenda — e é a ferramenta certa quando você está em dúvida (ver
`03-escolha-da-melhor-tomada.md`).
---
## `reason` aparece na tela
Não é campo de log. O texto que você escreve em `reason` é exibido para o
editor ao lado da frase selecionada, e é o que ele lê antes de manter ou
desfazer a sua decisão.
Escreva para quem está com pressa e vai decidir na hora:
- **Bom:** `"fecho, pico em 'devolver' (ênfase 0.34) — escala mais forte por ser o fechamento da peça"`
- **Ruim:** `"zoom"` · `"corte necessário"` · `"melhor tomada"`
A regra prática: se o `reason` não contém **o dado** que embasou (a palavra, o
número, a comparação entre tomadas), você provavelmente não tinha critério —
tinha impressão.
---
## O que a tela NÃO desfaz por você
- **Tempo errado continua errado.** A tela mostra suas ações no eixo da mídia
original; se você compensou para pós-corte, tudo aparece no lugar errado e o
editor não tem como adivinhar o que você quis dizer.
- **Excesso de zoom continua excesso.** A tela não impõe o teto de 2–4 por
minuto (`07-ritmo.md`) — ela mostra o que você mandou. Efeito demais chega
ao editor como trabalho de limpeza.
- **Frase promovida a ênfase sem querer.** Como zoom e legenda dinâmica andam
juntos, espalhar zooms "de segurança" enche o vídeo de legenda dinâmica. Na
dúvida, deixe sem — o editor promove; é mais barato que despromover.
---
## Depois da revisão
O editor pode, na tela: mudar o nível de ênfase (0–3), desativar ou reativar
frases, corrigir o texto, aparar as pontas por palavra, reclassificar entre
roteiro e bastidor e acrescentar zooms manuais em trechos arbitrários.
O resultado vira um `_phrase_review.json` e o `_phrase_actions.json` derivado —
e é esse que a geração usa. **Seu JSON é o ponto de partida da conversa, não a
palavra final.** Trabalhe para ser um bom ponto de partida: decisões
defensáveis, motivos legíveis e nenhuma escolha que o editor precise desfazer
antes de começar.
+6 -2
View File
@@ -30,6 +30,7 @@ Thumbs.db
# Env files (NUNCA commitar — contêm segredos) # Env files (NUNCA commitar — contêm segredos)
*.env *.env
.env .env
admin/gart-rag.env
# Graphify output (gerado, não rastrear) # Graphify output (gerado, não rastrear)
graphify-out/ graphify-out/
@@ -37,6 +38,9 @@ graphify-out/
# FCPXML bundles de exemplo (podem ser grandes) # FCPXML bundles de exemplo (podem ser grandes)
*.fcpxmld/ *.fcpxmld/
# WhisperX models cache # Cache de modelos Whisper baixados (código/models, ~11 GB, HuggingFace hub
models/ # format). Âncora em /code/models/ — NUNCA "models/" solto: isso também
# ignorava fcpxml/models/, o pacote de dados do engine (ver
# Engine/docs/05_EXPERIENCIAS.md #36).
/code/models/
whisper/ whisper/
+88 -15
View File
@@ -9,25 +9,91 @@ normalmente; a regra é sobre a comunicação com o usuário.
## What This Is ## What This Is
MCP server that reads/writes Final Cut Pro XML (FCPXML) files. 73 tools for timeline analysis, batch editing, QC, generation, multi-track support, media relink, NLE export, transcript-based editing (local Whisper), and LIVE FCP control (push_to_fcp / list_fcp_libraries via Apple events). Reads FCPXML 1.8–1.14 (incl. `.fcpxmld` bundles with sidecar preservation), writes 1.13 by default. Dual-mode (XML + Live) direction: `code/docs/CAPABILITY-AUDIT-2026-06.md`. MCP server that reads/writes Final Cut Pro XML (FCPXML) files. 77 tools for timeline analysis, batch editing, QC, generation, multi-track support, media relink, NLE export, transcript-based editing (local Whisper), LIVE FCP control (push_to_fcp / list_fcp_libraries via Apple events), and local-LLM voice scripting (editar-por-voz against Ollama/Gemma 3). Reads FCPXML 1.8–1.14 (incl. `.fcpxmld` bundles with sidecar preservation), writes 1.13 by default. Dual-mode (XML + Live) direction: `code/docs/CAPABILITY-AUDIT-2026-06.md`.
## Architecture ## Architecture
Toda a estrutura do projeto fica em `code/`. A pasta `admin/` fica na raiz (fora de `code/`). Toda a estrutura do projeto fica em `code/`. A pasta `admin/` fica na raiz (fora de `code/`).
``` Há **duas portas de entrada** para o mesmo engine: o MCP (Claude decide a
code/server.py — MCP server entry point. All 62 tool definitions, handlers, resources, prompts. edição) e a ponte JSON (o app macOS opera). Nenhuma das duas tem lógica de
Dispatch dict pattern: TOOL_HANDLERS maps tool names → async handler functions. timeline — as duas delegam a `fcpxml/`.
code/fcpxml/parser.py — Reads FCPXML → Python objects (Timeline, Clip, ConnectedClip, Marker, etc.)
code/fcpxml/writer.py — Writes modifications back to FCPXML. Handles markers, trimming, gaps, transitions.
code/fcpxml/rough_cut.py — Generates new timelines from source clips (rough cuts, montages, A/B rolls).
code/fcpxml/diff.py — Timeline comparison engine. Detects added/removed/moved/trimmed clips & markers.
code/fcpxml/export.py — DaVinci Resolve FCPXML v1.9 export + FCP7 XMEML v5 export for cross-NLE workflows.
code/fcpxml/models.py — Data classes: TimeValue, Timecode, Clip, ConnectedClip, CompoundClip, Timeline, etc.
code/fcpxml/media_intel.py — Real media analysis. Audio silence detection + beat detection.
code/fcpxml/dtd.py — Validates output against Apple's official DTDs.
``` ```
code/server.py — MCP entry point (592 linhas). Só dispatch: TOOL_HANDLERS.
code/server_tools/ — Os handlers das 77 tools, um módulo por categoria.
code/server_tools/_shared/ — Helpers compartilhados (paths, project, formatting,
captions, detection, media).
code/fcpxml/parser.py — Reads FCPXML → Python objects (Timeline, Clip, Marker…)
code/fcpxml/writer/ — PACOTE. Edição/escrita de FCPXML. FCPXMLModifier é
montado por mixins, um por assunto (markers, trim,
speed, titles, cut, silence…). Ver writer/modifier.py.
code/fcpxml/models/ — PACOTE. Data classes por família: timing, timeline,
enums, subtitles, qc, planning.
code/fcpxml/rough_cut.py — Generates new timelines (rough cuts, montages, A/B rolls).
code/fcpxml/diff.py — Timeline comparison engine.
code/fcpxml/export.py — DaVinci Resolve v1.9 + FCP7 XMEML v5 export.
code/fcpxml/media_intel.py — Silence detection + beat detection.
code/fcpxml/dtd.py — Validates output against Apple's official DTDs.
code/fcpxml/voice_*.py — Pipeline de voz: features → emphasis → voice_timeline
→ voice_actions → phrase_review. Ver Engine/docs/02.
admin/models_api.py — Ponte JSON com o app: docstring de comandos + dispatch.
admin/api/ — Os 37 comandos, um módulo por assunto.
code/MacApp/Sources/ — App SwiftUI. Compilado por swiftc (sem Xcode/SPM).
```
Os dois `__init__.py` de pacote (`writer/`, `models/`) reexportam tudo, então
`from .writer import FCPXMLModifier` e `from .models import TimeValue` seguem
valendo em todo o projeto.
## Documentação (MANDATORY)
A documentação viva fica em `code/Engine/docs/`. Cada arquivo tem **uma função
específica** — leia só o que a tarefa exige, não o conjunto. Carregar
documentação que não é do assunto custa tempo e processamento sem entregar nada.
### Qual arquivo abrir
| Sua tarefa | Abra | Não precisa de |
|-----------|------|----------------|
| Entender como o sistema é dividido | `01_ARCHITECTURE.md` | o resto |
| Achar onde mora uma função do engine | `02_MODULES.md` | 01, 03 |
| Criar/alterar uma ferramenta MCP | `03_SERVER_TOOLS.md` | 08 |
| Entender ou rodar os testes | `04_TESTS_AND_WORKFLOW.md` | — |
| "Isso já quebrou antes?" | `05_EXPERIENCIAS.md` — **só o índice no topo** | as entradas que não são a sua |
| Checklist antes de fechar | `06_BOAS_PRATICAS.md` | — |
| Mexer no app / no Assistente | `08_APP_MACOS.md` | 02, 03 |
| Escolher o que fazer, ver o que está aberto | `09_MANUTENCAO.md` | — |
Quando não souber por onde começar: `09_MANUTENCAO.md`. Ele roteia para o resto.
### Regra de atualização (obrigatória)
**Toda alteração de código atualiza a documentação no mesmo commit.** Doc velha
engana mais do que doc ausente — quem lê confia nela e erra com confiança.
| Você alterou | Atualize |
|--------------|----------|
| Estrutura de pastas, camadas ou dependências | `01_ARCHITECTURE.md` |
| Criou/moveu/dividiu módulo em `fcpxml/` | `02_MODULES.md` (tabela + linhas) |
| Criou/removeu ferramenta MCP | `03_SERVER_TOOLS.md` + contagem no `CLAUDE.md` |
| Comando da ponte | docstring de `admin/models_api.py` + `08_APP_MACOS.md` |
| Tela ou fluxo do app | `08_APP_MACOS.md` |
| Resolveu ou abriu uma dívida | `09_MANUTENCAO.md` §2 |
| Bateu num problema estrutural ou erro recorrente | `05_EXPERIENCIAS.md` + **índice no topo** |
Se um número (tools, testes, linhas) mudou, corrija onde ele aparece. Se um
documento divergir do código, **o código está certo** — conserte o documento.
### Ao escrever documentação
- **Um assunto por arquivo.** Se um doc começar a cobrir dois, divida.
- **Diga o que não está ali** e para onde ir — economiza a leitura seguinte.
- **Fatos verificados**, não suposições: rode o comando e use o número real.
- **Registre o porquê**, não só o quê. O "o quê" está no código; o "por quê"
se perde, e é o que evita alguém desfazer uma decisão por engano.
## Key Patterns ## Key Patterns
@@ -69,12 +135,19 @@ não passar. Equivalente a rodar manualmente os dois comandos abaixo.
Sempre que uma alteração for feita no app (MacApp/) durante o período de Sempre que uma alteração for feita no app (MacApp/) durante o período de
implementação, **compile e rode o programa localmente no computador** para implementação, **compile e rode o programa localmente no computador** para
validar visualmente a alteração, além de rodar os testes: validar visualmente a alteração, além de rodar os testes. O comando padrão
para isso — que fecha a instância anterior, recompila e abre o app para
conferência — é:
```bash ```bash
cd code && ./MacApp/build_app.sh --run # compila e abre o app localmente admin/run_app.command # compila e abre o app localmente (padrão de revisão)
``` ```
Equivalente a `cd code && ./MacApp/build_app.sh --run`, mas desacoplado do
Terminal. **Toda vez que uma alteração for concluída, rode este arquivo
automaticamente** para já conseguirmos revisar o que foi feito antes de
fechar a tarefa.
Regra geral: após qualquer alteração, o app deve ser executado localmente Regra geral: após qualquer alteração, o app deve ser executado localmente
antes de concluir a tarefa. Se houver erro de compilação, corrija antes de antes de concluir a tarefa. Se houver erro de compilação, corrija antes de
seguir. seguir.
@@ -88,7 +161,7 @@ CI runs both on every push to main. If either fails, the commit gets an X on Git
## Testing ## Testing
1342 tests across 34 files. `test_models.py` covers TimeValue arithmetic, Timecode parsing/formatting, Clip properties, validation models, and Timeline helpers. `test_writer.py` covers insert_clip, add_marker (all types), trim_clip, delete_clip, split_clip, and change_speed operations. `test_server.py` covers MCP tool handlers, parsers, and dispatch. `test_rough_cut.py` covers RoughCutGenerator. `test_features_v05.py` covers connected clips, roles, timeline diff, reformat, silence detection, export, and backward compatibility. `test_marker_pipeline.py` covers build_marker_element shared builder, batch auto-modes, clip index duplicate-name behavior, and write_fcpxml output format. `test_refactored_helpers.py` covers _index_elements, _iter_spine_clips, _find_spine_clip_at_seconds, _resolve_clip_duration, _make_asset_clip, _format_batch_result, and serialize_xml edge cases. `test_transcribe.py` covers phrase/filler span matching, range merge/invert algebra, whisper graceful degradation, and transcript-driven handler cuts against cached transcripts. `test_media_intel.py` covers silencedetect stderr parsing, source-to-timeline mapping, parameter bounds, and real-WAV ffmpeg integration (skips without ffmpeg; CI installs it). Tests use `examples/sample.fcpxml` as fixture data and inline XML fixtures. Tests create temp files and clean up after. 1498 tests across 43 files, all under `code/tests/`. Um teste fora dessa pasta não roda (`testpaths = ["tests"]`) — se você criar um em outro lugar, confirme que a contagem total subiu. Cobertura por área: `test_models.py` (TimeValue/Timecode/Clip/Timeline), `test_writer.py` (insert/marker/trim/delete/split/speed), `test_server.py` (handlers e dispatch), `test_rough_cut.py`, `test_features_v05.py` (connected clips, roles, diff, reformat, silêncio, export), `test_marker_pipeline.py`, `test_refactored_helpers.py`, `test_transcribe.py`, `test_media_intel.py` (pula sem ffmpeg; o CI instala), `test_phrase_review.py` (revisão de frases da etapa 5) e `test_models_api.py` (comandos da ponte). Fixtures: `examples/sample.fcpxml` e XML inline. Os testes criam temporários e limpam depois.
## FCPXML Gotchas ## FCPXML Gotchas
+20
View File
@@ -0,0 +1,20 @@
"""Comandos da ponte JSON usada pelo app, agrupados por assunto.
O setup de sys.path mora aqui, e só aqui, porque o pacote é importado antes de
qualquer um dos seus módulos (`from admin.api import models, voice, ...`
dispara este arquivo primeiro). Cada módulo de comando importa `fcpxml.*`
antes de importar `.shared` — sem o path já pronto neste ponto, o primeiro
desses imports falha com `ModuleNotFoundError`. Repetir o cálculo em cada
módulo (como era antes) é frágil por ordem: o app roda `admin/models_api.py`
por caminho absoluto, então `__file__` está sempre correto, mas cada arquivo
que refizesse essa conta um nível de diretório errado — como aconteceu quando
`_shared.py` virou este pacote e `admin/code` (inexistente) saiu no lugar de
`code/` — quebrava em silêncio até alguém rodar o comando de verdade.
"""
import sys
from pathlib import Path
_CODE_DIR = str(Path(__file__).resolve().parent.parent.parent / "code")
if _CODE_DIR not in sys.path:
sys.path.insert(0, _CODE_DIR)
+122
View File
@@ -0,0 +1,122 @@
"""Edições no projeto: silêncio, corte por texto, preenchimento, marcadores.
Extraído de models_api.py — a tabela de comandos segue lá.
"""
from __future__ import annotations
import asyncio
from pathlib import Path
from fcpxml.model_manager import (
load_silence_config,
save_silence_config,
)
from . import shared
from .shared import (
_derived_output,
_emit_no_change_or_error,
)
def cmd_remove_silences(args: dict) -> int:
"""Run the canonical server silence remover into a suffixed copy."""
path = str(args.get("path", ""))
if not path or not Path(path).exists():
shared.emit({"ok": False, "error": "Arquivo de projeto não encontrado."})
return 1
try:
from server import handle_remove_media_silence
output = _derived_output(path, "_silence_removed", args)
contents = asyncio.run(handle_remove_media_silence({**args, "filepath": path, "output_path": output}))
message = "\n".join(getattr(content, "text", str(content)) for content in contents)
if not Path(output).exists():
return _emit_no_change_or_error(path, message)
shared.emit({"ok": True, "path": output, "message": message})
return 0
except Exception as exc:
shared.emit({"ok": False, "error": str(exc)})
return 1
def cmd_edit_by_transcript(args: dict) -> int:
"""Cut (or keep only) spoken phrases, using each media's cached transcript."""
path = str(args.get("path", ""))
phrases = args.get("phrases") or []
if not path or not Path(path).exists():
shared.emit({"ok": False, "error": "Arquivo de projeto não encontrado."})
return 1
if not isinstance(phrases, list) or not [p for p in phrases if str(p).strip()]:
shared.emit({"ok": False, "error": "Informe ao menos uma frase para cortar."})
return 1
try:
from server import handle_edit_by_transcript
output = _derived_output(path, "_transcript_edit", args)
contents = asyncio.run(handle_edit_by_transcript({**args, "filepath": path, "output_path": output}))
message = "\n".join(getattr(content, "text", str(content)) for content in contents)
if not Path(output).exists():
shared.emit({"ok": False, "error": message})
return 1
shared.emit({"ok": True, "path": output, "message": message})
return 0
except Exception as exc:
shared.emit({"ok": False, "error": str(exc)})
return 1
def cmd_remove_filler_words(args: dict) -> int:
"""Cut filler words (um, uh, ...) out, using each media's cached transcript."""
path = str(args.get("path", ""))
if not path or not Path(path).exists():
shared.emit({"ok": False, "error": "Arquivo de projeto não encontrado."})
return 1
try:
from server import handle_remove_filler_words
output = _derived_output(path, "_defillered", args)
contents = asyncio.run(handle_remove_filler_words({**args, "filepath": path, "output_path": output}))
message = "\n".join(getattr(content, "text", str(content)) for content in contents)
if not Path(output).exists():
return _emit_no_change_or_error(path, message)
shared.emit({"ok": True, "path": output, "message": message})
return 0
except Exception as exc:
shared.emit({"ok": False, "error": str(exc)})
return 1
def cmd_transcript_markers(args: dict) -> int:
"""Add a marker per transcribed segment, using each media's cached transcript."""
path = str(args.get("path", ""))
if not path or not Path(path).exists():
shared.emit({"ok": False, "error": "Arquivo de projeto não encontrado."})
return 1
try:
from server import handle_transcript_markers
output = _derived_output(path, "_transcript_markers", args)
contents = asyncio.run(handle_transcript_markers({**args, "filepath": path, "output_path": output}))
message = "\n".join(getattr(content, "text", str(content)) for content in contents)
if not Path(output).exists():
shared.emit({"ok": False, "error": message})
return 1
shared.emit({"ok": True, "path": output, "message": message})
return 0
except Exception as exc:
shared.emit({"ok": False, "error": str(exc)})
return 1
def cmd_silence_config(args: dict) -> int:
"""Read the persisted silence thresholds (noise floor, duration, padding)."""
shared.emit({"ok": True, **load_silence_config()})
return 0
def cmd_set_silence_config(args: dict) -> int:
"""Persist silence thresholds. Only the given fields change."""
config = save_silence_config(
noise_db=args.get("noise_db"),
min_silence=args.get("min_silence"),
padding=args.get("padding"),
)
shared.emit({"ok": True, **config})
return 0
+137
View File
@@ -0,0 +1,137 @@
"""Catálogo de modelos: listar, baixar, escolher, apagar.
Extraído de models_api.py — a tabela de comandos segue lá.
"""
from __future__ import annotations
import shutil
import subprocess
import threading
from fcpxml.diarize import (
diarization_capability,
)
from fcpxml.model_manager import (
download_model,
get_models_dir,
is_model_downloaded,
list_installed_models,
load_catalog,
load_hf_token,
load_num_speakers,
load_selected_model,
load_transcript_language,
model_cache_dir,
save_models_dir,
save_selected_model,
save_transcript_language,
)
from . import shared
from .shared import (
RECOMMENDED,
)
# Downloads em andamento, para o comando `cancel` conseguir interrompê-los.
# Mora aqui, e não no shared, porque só `download` e `cancel` o tocam — e o
# lock é próprio: ele protege este dicionário, não a saída em stdout.
_CANCEL: dict[str, threading.Event] = {}
_CANCEL_LOCK = threading.Lock()
def cmd_catalog() -> None:
catalog = load_catalog()
installed = list_installed_models()
diar_ok, diar_msg = diarization_capability(load_hf_token())
shared.emit(
{
"models": catalog,
"installed": installed,
"selected": load_selected_model(),
"language": load_transcript_language(),
"models_dir": str(get_models_dir()),
"installed_count": len(installed),
"recommended": list(RECOMMENDED),
"diarization": diar_ok,
"diarization_message": diar_msg,
"hf_token_set": bool(load_hf_token()),
"num_speakers": load_num_speakers(),
}
)
def cmd_download(args: dict) -> int:
model = str(args.get("model", ""))
if model not in _model_names():
shared.emit({"type": "error", "message": f"Modelo desconhecido: {model}"})
return 1
ev = threading.Event()
with _CANCEL_LOCK:
_CANCEL[model] = ev
try:
download_model(model, progress_cb=lambda f: shared.emit({"type": "progress", "fraction": f}), cancel_event=ev)
installed = is_model_downloaded(model)
shared.emit({"type": "done", "installed": installed})
if installed:
save_selected_model(model)
return 0 if installed else 1
except Exception as exc:
shared.emit({"type": "error", "message": str(exc)})
return 1
finally:
with _CANCEL_LOCK:
_CANCEL.pop(model, None)
def cmd_cancel(args: dict) -> None:
model = str(args.get("model", ""))
ev = _CANCEL.get(model)
if ev is not None:
ev.set()
shared.emit({"ok": True})
def cmd_select(args: dict) -> None:
model = str(args.get("model", ""))
if not is_model_downloaded(model):
shared.emit({"ok": False, "error": "Modelo não está instalado."})
return
save_selected_model(model)
shared.emit({"ok": True, "selected": load_selected_model()})
def cmd_set_language(args: dict) -> int:
"""Persist the transcription language (the default for every transcription)."""
lang = str(args.get("language", "auto"))
try:
saved = save_transcript_language(lang)
except ValueError as exc:
shared.emit({"ok": False, "error": str(exc)})
return 1
shared.emit({"ok": True, "language": saved})
return 0
def cmd_delete(args: dict) -> None:
model = str(args.get("model", ""))
try:
shutil.rmtree(model_cache_dir(model), ignore_errors=True)
except Exception:
pass
shared.emit({"ok": True})
def cmd_open_finder(args: dict) -> None:
target = str(args.get("path") or model_cache_dir(str(args.get("model", ""))))
try:
subprocess.Popen(["open", target])
except OSError:
pass
shared.emit({"ok": True})
def cmd_set_models_dir(args: dict) -> int:
try:
d = save_models_dir(str(args.get("dir", "")))
shared.emit({"ok": True, "models_dir": d})
return 0
except ValueError as exc:
shared.emit({"ok": False, "error": str(exc)})
return 1
def _model_names() -> list[str]:
return [m["internal_name"] for m in load_catalog()]
+69
View File
@@ -0,0 +1,69 @@
"""Projeto: inspecionar o .fcpxml e lembrar a pasta/arquivo em uso.
Extraído de models_api.py — a tabela de comandos segue lá.
"""
from __future__ import annotations
from pathlib import Path
from fcpxml.model_manager import (
load_project_config,
save_project_config,
)
from fcpxml.parser import parse_fcpxml
from . import shared
def cmd_inspect(args: dict) -> int:
"""Validate an FCPXML file and return a summary of its projects/timelines."""
path = str(args.get("path", ""))
if not path:
shared.emit({"ok": False, "error": "Nenhum arquivo informado."})
return 1
if not Path(path).exists():
shared.emit({"ok": False, "error": "Arquivo não encontrado."})
return 1
try:
proj = parse_fcpxml(path)
except Exception as exc:
shared.emit({"ok": False, "error": f"Erro ao ler o projeto: {exc}"})
return 1
timelines = []
for tl in proj.timelines:
timelines.append(
{
"name": tl.name,
"duration_seconds": round(tl.duration.seconds, 3),
"frame_rate": round(tl.frame_rate, 3),
"width": tl.width,
"height": tl.height,
"clips": tl.total_clips,
"cuts": tl.total_cuts,
"connected": len(tl.connected_clips),
"markers": len(tl.markers),
}
)
shared.emit(
{
"ok": True,
"path": path,
"name": proj.name,
"fcpxml_version": proj.fcpxml_version,
"timelines": timelines,
}
)
return 0
def cmd_project_config(args: dict) -> int:
"""Read the last project folder/file the app was working on."""
shared.emit({"ok": True, **load_project_config()})
return 0
def cmd_set_project_config(args: dict) -> int:
"""Persist the last project folder/file. Only the given fields change."""
config = save_project_config(folder=args.get("folder"), file=args.get("file"))
shared.emit({"ok": True, **config})
return 0
+90
View File
@@ -0,0 +1,90 @@
"""Revisão de frases: montar a tela de ênfases e salvar o que foi decidido.
Extraído de models_api.py — a tabela de comandos segue lá.
"""
from __future__ import annotations
import json
from pathlib import Path
from . import shared
def cmd_build_phrase_review(args: dict) -> int:
"""Build the reviewable script (phrases + the AI's decisions) for the wizard.
`voice_timeline` points at the _voice_timeline.json; `actions` carries the
decision list the model returned (inline, in any of the shapes the skill
emits). The review is always rebuilt from the current analysis, then the
decisions saved on a previous visit are laid back over it — reopening the
step must show the edits the user left there without freezing the acoustics
as they were when they left.
"""
from fcpxml.phrase_review import (
build_phrase_review,
load_phrase_review,
merge_saved_decisions,
)
timeline_path = str(args.get("voice_timeline", ""))
if not timeline_path or not Path(timeline_path).exists():
shared.emit({"ok": False, "error": "Análise de voz (voice_timeline.json) não encontrada."})
return 1
try:
with open(timeline_path, encoding="utf-8") as fh:
timeline = json.load(fh)
except (OSError, ValueError) as exc:
shared.emit({"ok": False, "error": f"Erro ao ler a análise de voz: {exc}"})
return 1
extra = [d for d in (args.get("output_dir"), args.get("media_dir")) if d]
review = build_phrase_review(
timeline,
args.get("actions"),
voice_timeline_path=timeline_path,
extra_dirs=extra,
)
saved = None if args.get("fresh") else load_phrase_review(timeline_path)
review = merge_saved_decisions(review, saved)
shared.emit({"ok": True, "reused": saved is not None, **review})
return 0
def cmd_save_phrase_review(args: dict) -> int:
"""Persist the edited review and the actions derived from it."""
from fcpxml.phrase_review import save_phrase_review
timeline_path = str(args.get("voice_timeline", ""))
if not timeline_path:
shared.emit({"ok": False, "error": "Caminho da análise de voz não informado."})
return 1
phrases = args.get("phrases")
if not isinstance(phrases, list):
shared.emit({"ok": False, "error": "Nenhuma frase para salvar."})
return 1
review = {
"version": args.get("version", "1.0"),
"source": args.get("source", ""),
"duration": args.get("duration", 0.0),
"speakers": args.get("speakers", []),
"phrases": phrases,
"zooms": args.get("zooms", []),
}
try:
review_path, actions_path = save_phrase_review(timeline_path, review)
except OSError as exc:
shared.emit({"ok": False, "error": f"Erro ao salvar a revisão: {exc}"})
return 1
shared.emit({
"ok": True,
"review_path": str(review_path),
"actions_path": str(actions_path),
"emphasis_count": sum(1 for p in phrases if int(p.get("emphasis", 0) or 0) >= 1),
"removed_count": sum(1 for p in phrases if not p.get("active", True)),
})
return 0
+156
View File
@@ -0,0 +1,156 @@
"""Base comum dos comandos da ponte: saída JSON, caminhos derivados e cache.
A saída passa toda por `emit`. Os módulos de comando chamam `shared.emit(...)`
pelo módulo, e não pelo nome importado, de propósito: assim trocar `emit` num
lugar só — como a suíte faz para capturar a saída — continua alcançando todos
os comandos, o que deixaria de valer se cada um tivesse ligado o nome no seu
próprio import.
"""
from __future__ import annotations
import json
import os
import sys
import threading
from pathlib import Path
from typing import Any
from fcpxml.diarize import build_speakers
from fcpxml.media_intel import media_src_to_path
from fcpxml.parser import parse_fcpxml
RECOMMENDED = ("large-v3", "distil-large-v3", "small", "base")
def _derived_output(path: str, suffix: str, args: dict) -> str:
"""Resolve a derived XML path, optionally inside the chosen output folder."""
output_dir = str(args.get("output_dir", "")).strip()
if output_dir:
directory = Path(output_dir).expanduser()
directory.mkdir(parents=True, exist_ok=True)
source = Path(path)
extension = ".fcpxmld" if source.is_dir() else source.suffix
return str(directory / f"{source.stem}{suffix}{extension}")
from server import generate_output_path
return generate_output_path(path, suffix)
def _is_no_change_message(message: str) -> bool:
"""Whether a tool completed cleanly without needing to save a new file."""
text = message.lower()
return any(
token in text
for token in (
"no cuts to make",
"no silence",
"file unchanged",
"nothing saved",
)
)
def _emit_no_change_or_error(path: str, message: str) -> int:
if _is_no_change_message(message):
emit({"ok": True, "path": path, "unchanged": True, "message": message})
return 0
emit({"ok": False, "error": message})
return 1
# Serializa a escrita em stdout. A ponte é JSON-lines: dois comandos
# escrevendo ao mesmo tempo entrelaçariam documentos e o app leria lixo.
_OUT_LOCK = threading.Lock()
def emit(obj: Any) -> None:
sys.stdout.write(json.dumps(obj, ensure_ascii=False) + "\n")
sys.stdout.flush()
def _transcript_json_path(media_path: str, output_dir: str = "") -> Path:
"""Where the ``_transcript.json`` for ``media_path`` lives.
When ``output_dir`` (the user-selected project folder) is set, the
transcript is saved/read there — never next to the source media, which
may sit on a read-only volume or a Final Cut Library the user never
browses. Falls back to the media's own folder only when no project
folder has been chosen (legacy/MCP callers).
"""
p = Path(media_path)
if output_dir:
directory = Path(output_dir).expanduser()
directory.mkdir(parents=True, exist_ok=True)
return directory / f"{p.stem}_transcript.json"
return p.with_name(p.stem + "_transcript.json")
def _save_json_atomic(path: Path, data: Any) -> None:
"""Write ``data`` to ``path`` atomically and validate the result on disk.
Mirrors the reference WHISPERX save path: write a ``.tmp``, ``os.replace``
into place, then confirm the file exists, is non-empty, and parses as JSON.
"""
tmp_path = str(path) + ".tmp"
with open(tmp_path, "w", encoding="utf-8") as fh:
json.dump(data, fh, ensure_ascii=False, indent=2)
os.replace(tmp_path, path)
if not path.exists() or os.path.getsize(path) == 0:
raise RuntimeError("O arquivo salvo está vazio ou não foi encontrado.")
with open(path, encoding="utf-8") as fh:
json.load(fh)
def _project_media_paths(path: str) -> list[str]:
proj = parse_fcpxml(path)
tl = proj.primary_timeline or (proj.timelines[0] if proj.timelines else None)
media_paths: list[str] = []
if tl is not None:
for clip in getattr(tl, "clips", []):
mp = media_src_to_path(clip.media_path or "")
if mp and Path(mp).is_file() and mp not in media_paths:
media_paths.append(mp)
return media_paths
def _project_media_rotations(path: str) -> dict[str, float]:
"""Degrees each source media was rotated by via a Transform filter on its
clip in the FCPXML — keyed by the same resolved media path
``_project_media_paths`` returns, so the two can be joined by media_path."""
proj = parse_fcpxml(path)
tl = proj.primary_timeline or (proj.timelines[0] if proj.timelines else None)
rotations: dict[str, float] = {}
if tl is not None:
for clip in getattr(tl, "clips", []):
mp = media_src_to_path(clip.media_path or "")
if mp and clip.rotation:
rotations[mp] = clip.rotation
return rotations
def _voice_timeline_json_path(media_path: str, output_dir: str = "") -> Path:
p = Path(media_path)
if output_dir:
directory = Path(output_dir).expanduser()
directory.mkdir(parents=True, exist_ok=True)
return directory / f"{p.stem}_voice_timeline.json"
return p.with_name(p.stem + "_voice_timeline.json")
def _load_cached_voice_timeline(json_path: Path, media_path: str) -> dict | None:
try:
with open(json_path, encoding="utf-8") as fh:
data = json.load(fh)
except (OSError, json.JSONDecodeError, UnicodeDecodeError):
return None
if not isinstance(data, dict):
return None
if data.get("source") != Path(media_path).name:
return None
if not isinstance(data.get("segments"), list):
return None
return data
def _load_cached_transcript(json_path: Path) -> dict | None:
"""Return a valid cached transcript dict, or ``None`` if absent/unreadable."""
if not json_path.is_file():
return None
try:
data = json.loads(json_path.read_text(encoding="utf-8"))
except (OSError, ValueError):
return None
if isinstance(data, dict) and isinstance(data.get("words"), list):
if "speakers" not in data:
data["speakers"] = build_speakers(data.get("segments", []))
return data
return None
+234
View File
@@ -0,0 +1,234 @@
"""Legendas: dinâmicas, comuns, SRT e as configurações de estilo.
Extraído de models_api.py — a tabela de comandos segue lá.
"""
from __future__ import annotations
import asyncio
from pathlib import Path
from fcpxml.media_intel import media_src_to_path
from fcpxml.model_manager import (
load_dynamic_subtitle_config,
load_plain_subtitle_config,
save_dynamic_subtitle_config,
save_plain_subtitle_config,
)
from fcpxml.writer import FCPXMLModifier
from . import shared
from .shared import (
_derived_output,
_emit_no_change_or_error,
_load_cached_transcript,
_transcript_json_path,
)
def cmd_generate_dynamic_subtitles(args: dict) -> int:
"""Generate word-by-word ("karaoke") caption compound clips, one per line,
using each media's cached transcript."""
path = str(args.get("path", ""))
if not path or not Path(path).exists():
shared.emit({"ok": False, "error": "Arquivo de projeto não encontrado."})
return 1
try:
from server import handle_generate_dynamic_subtitles
output = _derived_output(path, "_dynamic_subtitles", args)
contents = asyncio.run(
handle_generate_dynamic_subtitles({**args, "filepath": path, "output_path": output})
)
message = "\n".join(getattr(content, "text", str(content)) for content in contents)
if not Path(output).exists():
shared.emit({"ok": False, "error": message})
return 1
shared.emit({"ok": True, "path": output, "message": message})
return 0
except Exception as exc:
shared.emit({"ok": False, "error": str(exc)})
return 1
def cmd_generate_plain_subtitles(args: dict) -> int:
"""Generate simple static editable subtitle title clips."""
path = str(args.get("path", ""))
if not path or not Path(path).exists():
shared.emit({"ok": False, "error": "Arquivo de projeto não encontrado."})
return 1
try:
from server import handle_generate_plain_subtitles
output = _derived_output(path, "_plain_subtitles", args)
contents = asyncio.run(
handle_generate_plain_subtitles({**args, "filepath": path, "output_path": output})
)
message = "\n".join(getattr(content, "text", str(content)) for content in contents)
if not Path(output).exists():
return _emit_no_change_or_error(path, message)
shared.emit({"ok": True, "path": output, "message": message})
return 0
except Exception as exc:
shared.emit({"ok": False, "error": str(exc)})
return 1
def cmd_export_srt(args: dict) -> int:
"""Write a captions .srt synced to the edited timeline.
Each transcribed segment is mapped from its SOURCE-media timestamp to its
real TIMELINE position (``clip_offset + (seg_start - clip_source_start)``),
so captions only cover the frames that remain after cuts/silence removal —
not the whole source file. One .srt is produced per media, in timeline order.
"""
path = str(args.get("path", ""))
output_dir = str(args.get("output_dir", "")).strip()
if not path or not Path(path).exists():
shared.emit({"ok": False, "error": "Arquivo de projeto não encontrado."})
return 1
try:
modifier = FCPXMLModifier(path)
except Exception as exc:
shared.emit({"ok": False, "error": f"Erro ao ler o projeto: {exc}"})
return 1
# Group spine clips by media so each transcript is loaded once.
by_media: dict[str, list] = {}
for _, el in modifier._iter_spine_clips():
src = modifier.resources.get(el.get("ref", ""), {}).get("src", "")
mp = media_src_to_path(src)
if not mp or not Path(mp).is_file():
continue
by_media.setdefault(mp, []).append(el)
# Never emit a caption past the end of the project — Final Cut rejects an
# SRT whose last cue overruns the timeline ("subtitle extends beyond project
# duration"). Clamp every mapped cue end to this ceiling.
timeline_total = modifier._timeline_duration().to_seconds()
srt_paths: list[str] = []
for mp, clips in by_media.items():
cached = _load_cached_transcript(_transcript_json_path(mp, output_dir))
if cached is None:
continue
segments = cached.get("segments") or []
if not segments:
continue
rows: list[tuple[float, float, str, int]] = []
for el in clips:
clip_source_start = modifier.source_file_start(el).to_seconds()
clip_duration = modifier._parse_time(el.get("duration", "0s")).to_seconds()
clip_offset = modifier._parse_time(el.get("offset", "0s")).to_seconds()
window_end = clip_source_start + clip_duration
for seg_index, seg in enumerate(segments):
seg_start = float(seg.get("start", 0.0))
seg_end = float(seg.get("end", seg_start))
text = seg.get("text", "").strip()
if not text or seg_end <= seg_start:
continue
# Intersect the complete source segment with this kept clip.
# Testing only seg_start loses speech whose first words fall in
# a removed range; interval intersection preserves the part
# that remains and avoids duplicating a segment wholesale.
source_start = max(seg_start, clip_source_start)
source_end = min(seg_end, window_end)
if source_end <= source_start:
continue
tl_start = clip_offset + (source_start - clip_source_start)
tl_end = clip_offset + (source_end - clip_source_start)
tl_start = max(0.0, min(tl_start, timeline_total))
tl_end = max(0.0, min(tl_end, timeline_total))
if tl_end > tl_start:
rows.append((tl_start, tl_end, text, seg_index))
if not rows:
continue
rows.sort(key=lambda r: (r[0], r[1], r[3]))
# Merge only pieces from the same original Whisper segment when their
# mapped intervals touch. Never merge unrelated speech or invent time.
merged: list[tuple[float, float, str, int]] = []
for row in rows:
if merged and row[3] == merged[-1][3] and row[0] <= merged[-1][1] + 0.001:
prev = merged[-1]
merged[-1] = (prev[0], max(prev[1], row[1]), prev[2], prev[3])
else:
merged.append(row)
blocks = []
for index, (s, e, text, _) in enumerate(merged, 1):
start_stamp = srt_stamp(s)
end_stamp = srt_stamp(e)
# Millisecond SRT precision can collapse a sub-millisecond span;
# omit it rather than emit an invalid zero-duration cue.
if start_stamp == end_stamp:
continue
blocks.append(f"{index}\n{start_stamp} --> {end_stamp}\n{text}\n")
if not blocks:
continue
out = (
Path(output_dir).expanduser() / f"{Path(mp).stem}_captions.srt"
if output_dir
else Path(mp).with_name(Path(mp).stem + "_captions.srt")
)
if output_dir:
out.parent.mkdir(parents=True, exist_ok=True)
try:
out.write_text("\n".join(blocks), encoding="utf-8")
except OSError as exc:
shared.emit({"ok": False, "error": f"Não foi possível salvar a legenda: {exc}"})
return 1
srt_paths.append(str(out))
if not srt_paths:
shared.emit({"ok": False, "error": "Nenhuma transcrição encontrada. Transcreva o projeto primeiro."})
return 1
shared.emit({"ok": True, "paths": srt_paths, "message": f"{len(srt_paths)} legenda(s) .srt sincronizada(s) com o corte."})
return 0
def srt_stamp(seconds: float) -> str:
"""Format float seconds as ``HH:MM:SS,mmm`` (SRT uses a comma).
Uses ``floor`` (not ``round``) so a timestamp never rounds up past a frame
boundary — an SRT cue ending on the last frame must not overrun the
project duration, or Final Cut flags it as extending beyond the project.
"""
ms = int((seconds if seconds > 0 else 0.0) * 1000)
h, rem = divmod(ms, 3600000)
m, rem = divmod(rem, 60000)
s, ms = divmod(rem, 1000)
return f"{h:02d}:{m:02d}:{s:02d},{ms:03d}"
def cmd_dynamic_subtitle_config(args: dict) -> int:
"""Read the persisted dynamic-subtitle style (font, size, color, layout)."""
shared.emit({"ok": True, **load_dynamic_subtitle_config()})
return 0
def cmd_set_dynamic_subtitle_config(args: dict) -> int:
"""Persist dynamic-subtitle style fields. Only the given fields change."""
config = save_dynamic_subtitle_config(**{
k: args.get(k) for k in (
"band_height", "block_center_y", "line_gap", "font", "font_size",
"emphasis_font", "emphasis_face", "emphasis_size",
"active_color", "emphasis_color", "text_scale",
)
})
shared.emit({"ok": True, **config})
return 0
def cmd_plain_subtitle_config(args: dict) -> int:
"""Read the persisted simple subtitle style."""
shared.emit({"ok": True, **load_plain_subtitle_config()})
return 0
def cmd_set_plain_subtitle_config(args: dict) -> int:
"""Persist simple subtitle style fields. Only the given fields change."""
config = save_plain_subtitle_config(**{
k: args.get(k) for k in (
"font", "font_size", "font_color", "max_words",
"position_y", "uppercase", "keep_punctuation", "text_scale",
)
})
shared.emit({"ok": True, **config})
return 0
+185
View File
@@ -0,0 +1,185 @@
"""Transcrição e locutores.
Extraído de models_api.py — a tabela de comandos segue lá.
"""
from __future__ import annotations
import json
from pathlib import Path
from fcpxml.diarize import (
assign_speakers,
build_speakers,
diarization_capability,
diarize,
)
from fcpxml.media_intel import media_src_to_path
from fcpxml.model_manager import (
is_model_downloaded,
load_hf_token,
load_num_speakers,
load_selected_model,
load_transcript_language,
save_hf_token,
save_num_speakers,
)
from fcpxml.parser import parse_fcpxml
from fcpxml.transcribe import transcribe
from . import shared
from .shared import (
_load_cached_transcript,
_save_json_atomic,
_transcript_json_path,
)
def cmd_transcribe(args: dict) -> int:
proj_path = str(args.get("path", ""))
output_dir = str(args.get("output_dir", "")).strip()
# Honra o modelo selecionado no programa quando nenhum é passado.
model = str(args.get("model", "") or load_selected_model() or "")
language = args.get("language")
if language is None:
language = load_transcript_language()
if language == "auto":
language = None
if not proj_path:
shared.emit({"type": "error", "message": "Nenhum projeto selecionado."})
return 1
if not output_dir:
shared.emit({"type": "error", "message": "Selecione a pasta do projeto antes de transcrever."})
return 1
if not model or not is_model_downloaded(model):
shared.emit(
{
"type": "error",
"message": "Nenhum modelo de transcrição instalado. Baixe e selecione um modelo na aba Modelos.",
}
)
return 1
token = str(args.get("hf_token") or load_hf_token() or "")
if args.get("num_speakers") is not None:
num_speakers = str(args.get("num_speakers"))
else:
num_speakers = load_num_speakers()
# Load project.
try:
proj = parse_fcpxml(proj_path)
except Exception as exc:
shared.emit({"type": "error", "message": f"Erro ao ler o projeto: {exc}"})
return 1
tl = proj.primary_timeline or (proj.timelines[0] if proj.timelines else None)
media_paths: list[str] = []
if tl is not None:
for clip in getattr(tl, "clips", []):
mp = media_src_to_path(clip.media_path or "")
if mp and Path(mp).is_file() and mp not in media_paths:
media_paths.append(mp)
if not media_paths:
shared.emit({"type": "error", "message": "Nenhum arquivo de mídia acessível encontrado."})
return 1
total = len(media_paths)
results: list[dict] = []
for i, mp in enumerate(media_paths, 1):
stage = f"Transcrevendo {Path(mp).name} ({i}/{total})…"
shared.emit({"type": "progress", "fraction": (i - 1) / total, "stage": stage})
json_path = _transcript_json_path(mp, output_dir)
cached = _load_cached_transcript(json_path)
if cached is not None:
shared.emit({"type": "progress", "fraction": i / total, "stage": stage})
results.append(_result_row(mp, cached))
continue
def _on_progress(file_fraction: float, _i: int = i, _stage: str = stage) -> None:
# Blend this file's own progress into the overall fraction so a
# single-media project doesn't jump straight to 100% before the
# actual (slow) decoding work has even started.
overall = (_i - 1 + file_fraction) / total
shared.emit({"type": "progress", "fraction": overall, "stage": _stage})
data = transcribe(mp, model_size=model, language=language, progress_cb=_on_progress)
if data is None:
shared.emit({"type": "error", "message": f"Não foi possível transcrever: {Path(mp).name}"})
return 1
# Diarização opcional (necessita token HF): assina speaker por segmento/palavra.
if token:
tracks = diarize(mp, token, num_speakers)
segments, words = assign_speakers(
data.get("segments", []), data.get("words", []), tracks
)
data = {**data, "segments": segments, "words": words}
data["speakers"] = build_speakers(data.get("segments", []))
payload = {
"schema_version": "1.0",
"source": Path(mp).name,
"model": model,
**data,
}
try:
_save_json_atomic(json_path, payload)
except (OSError, RuntimeError, ValueError) as exc:
shared.emit({"type": "error", "message": f"Não foi possível salvar o JSON: {exc}"})
return 1
results.append(_result_row(mp, data))
shared.emit({"type": "result", "transcripts": results})
return 0
def cmd_rename_speakers(args: dict) -> int:
"""Apply real names to speakers already saved in a transcript JSON."""
json_path = Path(str(args.get("path", "")))
names = args.get("speakers") or {}
if not json_path.is_file():
shared.emit({"type": "error", "message": "Transcrição não encontrada."})
return 1
try:
data = json.loads(json_path.read_text(encoding="utf-8"))
except (OSError, ValueError) as exc:
shared.emit({"type": "error", "message": f"Não foi possível ler o JSON: {exc}"})
return 1
mapping = {str(sid): str(name).strip() for sid, name in (names or {}).items()}
for sp in data.get("speakers", []):
sid = str(sp.get("id", ""))
if mapping.get(sid):
sp["name"] = mapping[sid]
try:
_save_json_atomic(json_path, data)
except (OSError, RuntimeError, ValueError) as exc:
shared.emit({"type": "error", "message": f"Não foi possível salvar: {exc}"})
return 1
shared.emit({"ok": True, "speakers": data.get("speakers", [])})
return 0
def cmd_set_diarization(args: dict) -> int:
"""Persist the HuggingFace token and expected speaker count for diarization."""
token = args.get("token")
num = args.get("num_speakers")
if token is not None:
save_hf_token(str(token))
if num is not None:
save_num_speakers(str(num))
ok, msg = diarization_capability(load_hf_token())
shared.emit({"ok": True, "diarization": ok, "diarization_message": msg, "num_speakers": load_num_speakers()})
return 0
def _result_row(mp: str, data: dict) -> dict:
words = data.get("words", [])
preview = (data.get("text", "") or "")[:160]
speakers = data.get("speakers") or []
return {
"media": Path(mp).name,
"language": data.get("language", "?"),
"words": len(words),
"duration": float(data.get("duration", 0.0)),
"preview": preview,
"saved": str(_transcript_json_path(mp)),
"speakers": [s.get("name", s.get("id", "")) for s in speakers],
}
+300
View File
@@ -0,0 +1,300 @@
"""Análise de voz e aplicação das decisões de edição.
Extraído de models_api.py — a tabela de comandos segue lá.
"""
from __future__ import annotations
import asyncio
import json
import re
from pathlib import Path
from fcpxml.model_manager import (
load_hf_token,
load_num_speakers,
load_selected_model,
load_transcript_language,
load_voice_analysis_config,
save_voice_analysis_config,
)
from . import shared
from .shared import (
_load_cached_transcript,
_load_cached_voice_timeline,
_project_media_paths,
_project_media_rotations,
_transcript_json_path,
_voice_timeline_json_path,
)
def cmd_analyze_voice(args: dict) -> int:
"""Build the voice timeline (transcript+diarization+acoustics -> emphasis)
for every unique source media in the project, so `refine_voice_timeline`
and friends have something to read without ever reopening the audio.
Analysis only — writes _voice_timeline.json next to each media, doesn't
touch the project XML. `path` passes through unchanged so it composes
with the other batch steps (silence removal, captions) regardless of
where in the list it runs.
"""
path = str(args.get("path", ""))
if not path or not Path(path).exists():
shared.emit({"ok": False, "error": "Arquivo de projeto não encontrado."})
return 1
model = str(args.get("model", "") or load_selected_model() or "")
language = args.get("language")
if language is None:
language = load_transcript_language()
if language == "auto":
language = None
token = str(args.get("hf_token") or load_hf_token() or "")
num_speakers = str(args.get("num_speakers") or load_num_speakers() or "")
try:
media_paths = _project_media_paths(path)
rotations = _project_media_rotations(path)
except Exception as exc:
shared.emit({"ok": False, "error": f"Erro ao ler o projeto: {exc}"})
return 1
if not media_paths:
shared.emit({"ok": False, "error": "Nenhum arquivo de mídia acessível encontrado."})
return 1
from server import handle_build_voice_timeline
messages: list[str] = []
output_dir = str(args.get("output_dir") or "").strip()
existing: list[Path] = []
for mp in media_paths:
timeline_path = _voice_timeline_json_path(mp, output_dir)
if _load_cached_voice_timeline(timeline_path, mp) is not None:
existing.append(timeline_path)
if existing and len(existing) == len(media_paths) and not bool(args.get("force_reprocess", False)):
message = "# Voice Timeline Cache\n\n"
message += "Reaproveitando análise de voz existente. Nada foi reprocessado.\n\n"
for timeline_path in existing:
message += f"- **Timeline JSON**: {timeline_path}\n"
shared.emit({
"ok": True,
"path": path,
"reused": True,
"timelines": [str(p) for p in existing],
"message": message,
})
return 0
for mp in media_paths:
transcript_path = _transcript_json_path(mp, output_dir)
reused_prefix = ""
if _load_cached_transcript(transcript_path) is not None:
reused_prefix = f"# Cache\n\nReaproveitando transcrição existente: `{transcript_path}`\n\n"
try:
contents = asyncio.run(handle_build_voice_timeline({
"media_path": mp, "model": model, "language": language,
"hf_token": token, "num_speakers": num_speakers,
"output_dir": output_dir, "rotation": rotations.get(mp, 0.0),
}))
except Exception as exc:
shared.emit({"ok": False, "error": f"Falha analisando {Path(mp).name}: {exc}"})
return 1
messages.append(reused_prefix + "\n".join(getattr(c, "text", str(c)) for c in contents))
shared.emit({"ok": True, "path": path, "message": "\n\n---\n\n".join(messages)})
return 0
def cmd_acoustics_capability(args: dict) -> int:
"""Whether librosa (pitch/energy extraction) is installed in this venv.
Surfaces `features_capability()` — previously computed but never
exposed to the app, so `layers.acoustics: false` in a voice timeline
had no explanation the user could act on.
"""
from fcpxml.voice_features import features_capability
ok, msg = features_capability()
shared.emit({"ok": True, "available": ok, "message": msg})
return 0
def cmd_voice_analysis(args: dict) -> int:
"""Read the persisted voice-analysis settings (energy/emphasis/emotion)."""
config = load_voice_analysis_config()
shared.emit({"ok": True, **config, "emphasis_threshold": config["emphasis_floor"]})
return 0
def cmd_set_voice_analysis(args: dict) -> int:
"""Persist voice-analysis settings. Only the given fields change."""
weights = args.get("emphasis_weights")
config = save_voice_analysis_config(
energy_threshold=args.get("energy_threshold"),
emphasis_weights=weights if isinstance(weights, dict) else None,
emphasis_floor=args.get("emphasis_threshold"),
emotion_enabled=args.get("emotion_enabled"),
emotion_sensitivity=args.get("emotion_sensitivity"),
zoom_scale=args.get("zoom_scale"),
zoom_mode=args.get("zoom_mode"),
zoom_ease_in=args.get("zoom_ease_in"),
zoom_ease_out=args.get("zoom_ease_out"),
)
shared.emit({"ok": True, **config})
return 0
def cmd_apply_voice_actions(args: dict) -> int:
"""Apply a decision list (cuts/zooms/texts/markers) to the project XML.
The list is produced by a model reading the _voice_timeline.json — this
is the step that turns those decisions into an edit, and the one the
batch chain was missing: without it the app could measure the voice and
caption the result, but never cut by it.
`actions_path` points at the JSON; either a bare list or the
``{"actions": [...]}`` wrapper the skill emits is accepted. Times stay in
ORIGINAL source seconds — the handler resolves cuts first and shifts
everything else itself.
"""
path = str(args.get("path", ""))
if not path or not Path(path).exists():
shared.emit({"ok": False, "error": "Arquivo de projeto não encontrado."})
return 1
actions = args.get("actions")
if actions is None:
actions_path = str(args.get("actions_path", ""))
if not actions_path or not Path(actions_path).exists():
shared.emit({"ok": False, "error": "Arquivo de decisões (JSON) não encontrado."})
return 1
try:
with open(actions_path, encoding="utf-8") as fh:
loaded = json.load(fh)
except (OSError, ValueError) as exc:
shared.emit({"ok": False, "error": f"Erro ao ler as decisões: {exc}"})
return 1
actions = loaded.get("actions") if isinstance(loaded, dict) else loaded
# The documented output format is {"source": ..., "actions": [...]} —
# callers passing that whole object inline (e.g. the wizard pasting the
# skill's JSON verbatim) need the same unwrap the actions_path branch
# above already does, or a well-formed payload gets rejected as
# "malformed" for having one extra layer of nesting.
if isinstance(actions, dict):
actions = actions.get("actions")
if not isinstance(actions, list) or not actions:
shared.emit({"ok": False, "error": "A lista de decisões está vazia ou malformada."})
return 1
from server import handle_apply_voice_actions
try:
contents = asyncio.run(handle_apply_voice_actions({
"filepath": path,
"actions": actions,
"output_dir": args.get("output_dir"),
}))
except Exception as exc:
shared.emit({"ok": False, "error": f"Falha ao aplicar as decisões: {exc}"})
return 1
message = "\n".join(getattr(c, "text", str(c)) for c in contents)
# The handler reports dropped/rejected actions individually; hand the
# whole report back so the app can surface them instead of only the count.
out_path = path
for line in message.splitlines():
if line.startswith("- **Saved to**:"):
out_path = line.split("`")[1] if "`" in line else path
break
shared.emit({"ok": True, "path": out_path, "message": message})
return 0
def cmd_generate_voice_script(args: dict) -> int:
"""Run the ENTIRE voice-edit pass against a LOCAL model, inside the engine.
Transcribe (cached) -> build the voice timeline -> hand it to a local
Ollama model (Gemma 3 / Llama) that directs the edit -> return the readable
script (roteiro) and the action JSON, and optionally apply to a FCPXML. No
wizard, no copy-paste: the model's decisions are validated and applied by
the same pipeline the rules engine uses.
Args (all optional except one of ``media_path`` / ``voice_timeline``):
media_path audio/video to analyze and direct (required when there is
no voice_timeline yet)
voice_timeline path to an existing _voice_timeline.json; when given the
analysis is reused and media_path is not required
filepath optional FCPXML to apply the decisions to (non-destructive)
model local model Ollama serves (default gemma3:12b)
base_url Ollama base URL (default http://localhost:11434)
model_size whisper size if transcription is needed
language ISO language hint for transcription
hf_token HuggingFace token for diarization
num_speakers known speaker count, if any
output_dir folder for the timeline/review/actions JSON
apply_to_fcpxml apply to filepath when given (default true)
-> {"ok": true, "message": "...", "roteiro_path", "actions_path",
"applied_path"} or {"ok": false, "error": "..."}
"""
media_path = str(args.get("media_path", ""))
voice_timeline = str(args.get("voice_timeline", ""))
if not voice_timeline and (not media_path or not Path(media_path).exists()):
shared.emit({"ok": False, "error": "Arquivo de mídia não encontrado (informe media_path ou voice_timeline)."})
return 1
from server import handle_generate_voice_script
try:
contents = asyncio.run(handle_generate_voice_script({
"media_path": media_path,
"voice_timeline": args.get("voice_timeline"),
"filepath": args.get("filepath"),
"model": args.get("model"),
"base_url": args.get("base_url"),
"model_size": args.get("model_size"),
"language": args.get("language"),
"hf_token": args.get("hf_token"),
"num_speakers": args.get("num_speakers"),
"output_dir": args.get("output_dir"),
"apply_to_fcpxml": args.get("apply_to_fcpxml", True),
}))
except Exception as exc:
shared.emit({"ok": False, "error": f"Falha ao gerar roteiro por IA local: {exc}"})
return 1
message = "\n".join(getattr(c, "text", str(c)) for c in contents)
def _path_after(label: str) -> str:
m = re.search(rf"\*\*{label}\*\*: (.+)", message)
return m.group(1).strip() if m else ""
roteiro_path = _path_after(r"Roteiro \(legível\)")
actions_path = _path_after("Ações JSON")
applied_path = ""
for line in message.splitlines():
if line.startswith("- **Saved to**:"):
applied_path = line.split("`")[1] if "`" in line else ""
break
shared.emit({
"ok": True,
"message": message,
"roteiro_path": roteiro_path,
"actions_path": actions_path,
"applied_path": applied_path,
})
return 0
def cmd_list_ollama_models(args: dict) -> int:
"""List the models Ollama currently serves, for the app's model picker.
Args:
base_url Ollama base URL (default http://localhost:11434)
-> {"ok": true, "models": ["gemma3:12b", ...]} (empty list if Ollama
is unreachable, so the UI can fall back to a text field)
"""
from fcpxml.llm_local import list_ollama_models
base_url = str(args.get("base_url") or "http://localhost:11434")
models = list_ollama_models(base_url=base_url)
shared.emit({"ok": True, "models": models})
return 0
+107
View File
@@ -0,0 +1,107 @@
"""Zoom (punch-in): por janela, por clipe e por trecho da transcrição.
Extraído de models_api.py — a tabela de comandos segue lá.
"""
from __future__ import annotations
import asyncio
from pathlib import Path
from . import shared
from .shared import (
_derived_output,
_load_cached_transcript,
_transcript_json_path,
)
def cmd_add_zoom(args: dict) -> int:
"""Add an ease-in/ease-out punch-in zoom to one clip."""
path = str(args.get("path", ""))
clip_id = str(args.get("clip_id", "")).strip()
if not path or not Path(path).exists():
shared.emit({"ok": False, "error": "Arquivo de projeto não encontrado."})
return 1
if not clip_id:
shared.emit({"ok": False, "error": "Informe o nome do clipe."})
return 1
try:
from server import handle_add_zoom
output = _derived_output(path, "_zoom", args)
contents = asyncio.run(handle_add_zoom({**args, "filepath": path, "output_path": output}))
message = "\n".join(getattr(content, "text", str(content)) for content in contents)
if not Path(output).exists():
shared.emit({"ok": False, "error": message})
return 1
shared.emit({"ok": True, "path": output, "message": message})
return 0
except Exception as exc:
shared.emit({"ok": False, "error": str(exc)})
return 1
def cmd_zoom_clips(args: dict) -> int:
"""Return timeline clips with enough identity for the zoom picker."""
path = Path(str(args.get("path", "")))
output_dir = str(args.get("output_dir", "")).strip()
if not path.exists():
shared.emit({"ok": False, "error": "Arquivo de projeto não encontrado."})
return 1
try:
from server import _require_timeline
_, timeline = _require_timeline(str(path))
clips = []
for index, clip in enumerate(timeline.clips):
media = clip.media_path or ""
cached = _load_cached_transcript(_transcript_json_path(media, output_dir)) if media else None
clips.append({
"id": f"{index}:{clip.start.seconds:.6f}",
"index": index,
"name": clip.name,
"start": clip.start.seconds,
"duration": clip.duration_seconds,
"media": Path(media).name if media else "",
"preview": ((cached or {}).get("text", "") or "")[:180],
"has_transcript": cached is not None,
})
shared.emit({"ok": True, "clips": clips})
return 0
except Exception as exc:
shared.emit({"ok": False, "error": str(exc)})
return 1
def cmd_zoom_segments(args: dict) -> int:
"""Return sentence/word ranges for one timeline clip."""
path = Path(str(args.get("path", "")))
output_dir = str(args.get("output_dir", "")).strip()
try:
from server import _require_timeline
_, timeline = _require_timeline(str(path))
index = int(args.get("index", -1))
if index < 0 or index >= len(timeline.clips):
raise ValueError("Clipe selecionado não existe.")
clip = timeline.clips[index]
if not clip.media_path:
raise ValueError("Este clipe não possui mídia associada.")
data = _load_cached_transcript(_transcript_json_path(clip.media_path, output_dir))
if data is None:
shared.emit({"ok": True, "segments": [], "message": "Transcreva este clipe primeiro."})
return 0
segments = []
for number, segment in enumerate(data.get("segments", [])):
text = str(segment.get("text", "")).strip()
if text:
segments.append({
"id": number,
"start": float(segment.get("start", 0)),
"end": float(segment.get("end", 0)),
"text": text,
})
shared.emit({"ok": True, "segments": segments})
return 0
except Exception as exc:
shared.emit({"ok": False, "error": str(exc)})
return 1
+12
View File
@@ -0,0 +1,12 @@
# Copie para admin/gart-rag.env e preencha a senha. Este arquivo é apenas um
# modelo; admin/gart-rag.env é ignorado pelo git.
RAG_DB_HOST=127.0.0.1
RAG_DB_PORT=55435
RAG_DB_NAME=rag_gart
RAG_DB_SCHEMA=gart
RAG_DB_USER=gart_rag_indexer
RAG_DB_PASSWORD=
# Ollama que fornece nomic-embed-text.
OLLAMA_URL=http://127.0.0.1:11434
RAG_EMBED_MODEL=nomic-embed-text
+174 -980
View File
File diff suppressed because it is too large Load Diff
+3 -4
View File
@@ -20,7 +20,6 @@ import subprocess
import sys import sys
import threading import threading
from pathlib import Path from pathlib import Path
from typing import Optional
import flet as ft import flet as ft
@@ -101,9 +100,9 @@ class ModelManagerApp:
def __init__(self, page: ft.Page) -> None: def __init__(self, page: ft.Page) -> None:
self.page = page self.page = page
self.selected = load_selected_model() self.selected = load_selected_model()
self.downloading: Optional[str] = None self.downloading: str | None = None
self._cancel_events: dict[str, threading.Event] = {} self._cancel_events: dict[str, threading.Event] = {}
self._picker: Optional[ft.FilePicker] = None self._picker: ft.FilePicker | None = None
# ── helpers ──────────────────────────────────────────────────────────── # ── helpers ────────────────────────────────────────────────────────────
@@ -121,7 +120,7 @@ class ModelManagerApp:
self._picker = ft.FilePicker() self._picker = ft.FilePicker()
self._picker.on_result = self._on_file_picked self._picker.on_result = self._on_file_picked
self.page.overlay.append(self._picker) self.page.overlay.append(self._picker)
self._pending_target: Optional[dict] = None self._pending_target: dict | None = None
def _on_file_picked(self, e) -> None: def _on_file_picked(self, e) -> None:
if self._pending_target == "project": if self._pending_target == "project":
-243
View File
@@ -1,243 +0,0 @@
"""Tests for admin/models_api.py — the SwiftUI JSON bridge commands.
Focused on the transcription-flow changes: atomic save, speaker renaming, and
the "use the selected model" default plus model-availability guard.
"""
import json
import admin.models_api as api
def _capture(monkeypatch):
captured: list[dict] = []
def _emit(obj):
captured.append(obj)
monkeypatch.setattr(api, "_emit", _emit)
return captured
def test_save_json_atomic(tmp_path):
p = tmp_path / "t.json"
api._save_json_atomic(p, {"a": [1, 2], "text": "olá"})
assert p.exists()
assert not (tmp_path / "t.json.tmp").exists()
assert json.loads(p.read_text(encoding="utf-8"))["text"] == "olá"
def test_rename_speakers(tmp_path, monkeypatch):
captured = _capture(monkeypatch)
p = tmp_path / "t.json"
p.write_text(
json.dumps(
{
"speakers": [
{"id": "SPEAKER_00", "name": "Speaker 1"},
{"id": "SPEAKER_01", "name": "Speaker 2"},
]
}
),
encoding="utf-8",
)
assert api.cmd_rename_speakers({"path": str(p), "speakers": {"SPEAKER_01": "Erika"}}) == 0
assert captured[0]["ok"] is True
saved = json.loads(p.read_text(encoding="utf-8"))
assert saved["speakers"][0]["name"] == "Speaker 1"
assert saved["speakers"][1]["name"] == "Erika"
def test_rename_speakers_missing_file(monkeypatch):
captured = _capture(monkeypatch)
assert api.cmd_rename_speakers({"path": "/nonexistent/x.json"}) == 1
assert captured[0]["type"] == "error"
def test_transcribe_requires_output_dir(monkeypatch):
captured = _capture(monkeypatch)
monkeypatch.setattr(api, "load_selected_model", lambda: "small")
monkeypatch.setattr(api, "is_model_downloaded", lambda m: True)
assert api.cmd_transcribe({"path": "/some/project.fcpxml"}) == 1
assert captured[0]["type"] == "error"
assert "pasta do projeto" in captured[0]["message"]
def test_transcribe_requires_installed_model(monkeypatch, tmp_path):
captured = _capture(monkeypatch)
monkeypatch.setattr(api, "load_selected_model", lambda: "")
monkeypatch.setattr(api, "is_model_downloaded", lambda m: False)
assert api.cmd_transcribe({"path": "/some/project.fcpxml", "output_dir": str(tmp_path)}) == 1
assert captured[0]["type"] == "error"
assert "instalado" in captured[0]["message"]
def test_transcribe_defaults_to_selected_model(monkeypatch, tmp_path):
captured = _capture(monkeypatch)
monkeypatch.setattr(api, "load_selected_model", lambda: "small")
monkeypatch.setattr(api, "is_model_downloaded", lambda m: m == "small")
class FakeTL:
clips = []
class FakeProject:
primary_timeline = None
timelines = [FakeTL()]
monkeypatch.setattr(api, "parse_fcpxml", lambda p: FakeProject())
# No media accessible -> reaches the media-path check (past model validation).
assert api.cmd_transcribe({"path": "/some/project.fcpxml", "output_dir": str(tmp_path)}) == 1
assert captured[0]["type"] == "error"
assert "mídia" in captured[0]["message"]
def test_set_language_persists(monkeypatch):
captured = _capture(monkeypatch)
assert api.cmd_set_language({"language": "pt"}) == 0
assert captured[0]["ok"] is True
assert captured[0]["language"] == "pt"
assert api.load_transcript_language() == "pt"
def test_set_language_rejects_unknown(monkeypatch):
captured = _capture(monkeypatch)
assert api.cmd_set_language({"language": "xx"}) == 1
assert captured[0]["ok"] is False
assert "language" in captured[0]["error"]
def test_transcribe_defaults_language_to_persisted(monkeypatch, tmp_path):
monkeypatch.setattr(api, "load_selected_model", lambda: "small")
monkeypatch.setattr(api, "is_model_downloaded", lambda m: m == "small")
monkeypatch.setattr(api, "load_transcript_language", lambda: "pt")
media = tmp_path / "clip.mov"
media.write_bytes(b"fake")
class FakeClip:
media_path = ""
class FakeTL:
clips = [FakeClip()]
class FakeProject:
primary_timeline = None
timelines = [FakeTL()]
monkeypatch.setattr(api, "parse_fcpxml", lambda p: FakeProject())
monkeypatch.setattr(api, "media_src_to_path", lambda mp: str(media))
called = {}
monkeypatch.setattr(
api, "transcribe", lambda mp, model_size, language, **kw: called.update(lang=language)
)
assert api.cmd_transcribe({"path": "/some/project.fcpxml", "output_dir": str(tmp_path / "out")}) == 1
assert called["lang"] == "pt"
def test_srt_stamp_format():
assert api.srt_stamp(0.0) == "00:00:00,000"
assert api.srt_stamp(1.5) == "00:00:01,500"
assert api.srt_stamp(3661.234) == "01:01:01,234"
_FCPXML_SAMPLE = """<?xml version="1.0" encoding="UTF-8"?>
<fcpxml version="1.13">
<resources>
<asset id="r1" name="clip" uid="u1" start="0s" duration="100s"
hasVideo="1" format="f1" hasAudio="1">
<media-rep kind="original-media" src="file:///tmp/clip.mp4"/>
</asset>
<format id="f1" name="FFVideoFormat1080p25" frameDuration="1/25s" width="1920" height="1080"/>
</resources>
<library>
<event name="Event">
<project name="P">
<sequence format="f1">
<spine>
<asset-clip ref="r1" offset="0s" start="10s" duration="10s" name="clip"/>
<gap name="Espaço" offset="10s" duration="90s" start="10s"/>
</spine>
</sequence>
</project>
</event>
</library>
</fcpxml>
"""
def test_cmd_export_srt_maps_to_edited_timeline(tmp_path, monkeypatch):
"""Captions must reflect the EDITED timeline, not the whole source file."""
captured = _capture(monkeypatch)
project = tmp_path / "proj.fcpxml"
project.write_text(_FCPXML_SAMPLE, encoding="utf-8")
media = tmp_path / "clip.mp4"
media.write_bytes(b"fake")
# Transcript covers 0..100s; the clip only USES source 10..20s -> timeline 0..10s.
transcript = {
"words": [],
"segments": [
{"start": 5.0, "end": 6.0, "text": "antes do corte"},
{"start": 12.0, "end": 14.0, "text": "dentro do corte"},
{"start": 50.0, "end": 51.0, "text": "depois do corte"},
]
}
tj = api._transcript_json_path(media)
tj.parent.mkdir(parents=True, exist_ok=True)
api._save_json_atomic(tj, transcript)
monkeypatch.setattr(api, "media_src_to_path", lambda src: str(media))
assert api.cmd_export_srt({"path": str(project)}) == 0
assert captured[0]["ok"] is True
srt = tmp_path / "clip_captions.srt"
assert srt.exists()
text = srt.read_text(encoding="utf-8")
# Only the segment inside the used source window (12s) survives.
assert "dentro do corte" in text
assert "antes do corte" not in text
assert "depois do corte" not in text
# Mapped to timeline 0..10s -> the 12s source segment lands at 2s.
assert "00:00:02,000 --> 00:00:04,000" in text
def test_cmd_export_srt_no_transcript(tmp_path, monkeypatch):
captured = _capture(monkeypatch)
project = tmp_path / "proj.fcpxml"
project.write_text(_FCPXML_SAMPLE, encoding="utf-8")
media = tmp_path / "clip.mp4"
media.write_bytes(b"fake")
monkeypatch.setattr(api, "media_src_to_path", lambda src: str(media))
assert api.cmd_export_srt({"path": str(project)}) == 1
assert captured[0]["ok"] is False
def test_cmd_export_srt_clamps_past_project_duration(tmp_path, monkeypatch):
"""A segment ending after the last clip must be clamped to the project end.
Final Cut rejects an SRT whose final cue overruns the timeline
("subtitle extends beyond project duration").
"""
captured = _capture(monkeypatch)
project = tmp_path / "proj.fcpxml"
project.write_text(_FCPXML_SAMPLE, encoding="utf-8")
media = tmp_path / "clip.mp4"
media.write_bytes(b"fake")
# Clip uses source 10..20s -> timeline 0..10s. A segment 12..30s maps to
# timeline 2..20s, but the project only lasts 10s: must clamp end to 10s.
transcript = {
"words": [],
"segments": [
{"start": 12.0, "end": 30.0, "text": "longa fala"},
]
}
tj = api._transcript_json_path(media)
tj.parent.mkdir(parents=True, exist_ok=True)
api._save_json_atomic(tj, transcript)
monkeypatch.setattr(api, "media_src_to_path", lambda src: str(media))
assert api.cmd_export_srt({"path": str(project)}) == 0
assert captured[0]["ok"] is True
srt = tmp_path / "clip_captions.srt"
text = srt.read_text(encoding="utf-8")
# Timeline is 10s; the cue must not end past it.
assert "00:00:02,000 --> 00:00:10,000" in text
assert "00:00:20,000" not in text
+68
View File
@@ -0,0 +1,68 @@
#!/bin/zsh
# Atualiza incrementalmente a RAG do G-ART usando o banco compartilhado.
#
# Credenciais: defina RAG_DB_PASSWORD no ambiente ou crie
# admin/gart-rag.env (ignorado pelo git). O arquivo pode conter também
# RAG_DB_USER, RAG_DB_PORT, OLLAMA_URL e RAG_EMBED_MODEL.
set -euo pipefail
ROOT="$(cd "$(dirname "$0")/.." && pwd -P)"
ENV_FILE="$ROOT/admin/gart-rag.env"
if [[ -f "$ENV_FILE" ]]; then
set -a
source "$ENV_FILE"
set +a
fi
PYTHON="${RAG_PYTHON:-}"
if [[ -z "$PYTHON" ]]; then
for candidate in "$ROOT/admin/.venv/bin/python3" "$ROOT/rag/.venv/bin/python3"; do
if [[ -x "$candidate" ]]; then PYTHON="$candidate"; break; fi
done
fi
PYTHON="${PYTHON:-$(command -v python3)}"
if ! "$PYTHON" -c 'import psycopg2, requests' >/dev/null 2>&1; then
echo "ERRO: o Python da RAG precisa dos pacotes psycopg2 e requests." >&2
echo "Instale-os no ambiente indicado por RAG_PYTHON e tente novamente." >&2
exit 1
fi
if [[ -z "${RAG_DB_PASSWORD:-}" ]]; then
echo "ERRO: defina RAG_DB_PASSWORD ou configure $ENV_FILE" >&2
exit 1
fi
HOST="${RAG_VPS_HOST:-179.197.228.240}"
LOCAL_PORT="${RAG_DB_PORT:-55435}"
REMOTE_PORT="${RAG_REMOTE_PORT:-55435}"
TUNNEL_PID=""
cleanup() {
if [[ -n "$TUNNEL_PID" ]] && kill -0 "$TUNNEL_PID" 2>/dev/null; then
kill "$TUNNEL_PID" 2>/dev/null || true
wait "$TUNNEL_PID" 2>/dev/null || true
fi
}
trap cleanup EXIT
if ! nc -z 127.0.0.1 "$LOCAL_PORT" 2>/dev/null; then
echo "==> Abrindo túnel RAG (127.0.0.1:$LOCAL_PORT)..."
ssh -N -o ExitOnForwardFailure=yes -o ServerAliveInterval=30 \
-o ServerAliveCountMax=3 -L "127.0.0.1:$LOCAL_PORT:127.0.0.1:$REMOTE_PORT" \
"${RAG_VPS_USER:-root}@$HOST" &
TUNNEL_PID=$!
for _ in {1..20}; do
nc -z 127.0.0.1 "$LOCAL_PORT" 2>/dev/null && break
kill -0 "$TUNNEL_PID" 2>/dev/null || break
sleep 0.25
done
fi
if ! nc -z 127.0.0.1 "$LOCAL_PORT" 2>/dev/null; then
echo "ERRO: não foi possível abrir o túnel RAG." >&2
exit 1
fi
echo "==> Atualizando RAG do G-ART (incremental)..."
cd "$ROOT"
exec "$PYTHON" "$ROOT/admin/update_rag.py"
+216
View File
@@ -0,0 +1,216 @@
#!/usr/bin/env python3
"""Atualiza incrementalmente o índice RAG do G-ART.
As credenciais são fornecidas pelo ambiente; este arquivo nunca deve conter
senha. O indexador usa o banco ``rag_gart`` e o schema ``gart`` por padrão.
"""
from __future__ import annotations
import hashlib
import os
import re
import sys
from pathlib import Path
import psycopg2
import requests
ROOT = Path(__file__).resolve().parents[1]
DB_NAME = os.environ.get("RAG_DB_NAME", "rag_gart")
DB_SCHEMA = os.environ.get("RAG_DB_SCHEMA", "gart")
OLLAMA_URL = os.environ.get("OLLAMA_URL", "http://127.0.0.1:11434")
EMBED_MODEL = os.environ.get("RAG_EMBED_MODEL", "nomic-embed-text")
EMBED_DIM = int(os.environ.get("RAG_EMBED_DIM", "768"))
INCLUDE_EXTENSIONS = {
".command", ".md", ".py", ".sh", ".sql", ".swift", ".txt", ".yml", ".yaml",
}
EXCLUDE_DIRS = {
".git", ".venv", ".pytest_cache", ".ruff_cache", "__pycache__", "build",
"dist", "node_modules", "graphify-out", "bm", "models", "whisper",
}
EXCLUDE_FILES = {".env", "admin/genial-crm.env", "admin/genial-crm.local.env"}
CHUNK_LINES = 60
CHUNK_OVERLAP = 10
CHUNK_MAX_CHARS = 5000
def _sql_id(value: str) -> str:
return '"' + value.replace('"', '""') + '"'
def _connect():
password = os.environ.get("RAG_DB_PASSWORD")
if not password:
raise RuntimeError("RAG_DB_PASSWORD não foi definida")
return psycopg2.connect(
host=os.environ.get("RAG_DB_HOST", "127.0.0.1"),
port=os.environ.get("RAG_DB_PORT", "55435"),
dbname=DB_NAME,
user=os.environ.get("RAG_DB_USER", "gart_rag_indexer"),
password=password,
connect_timeout=5,
)
def _iter_files():
for path in ROOT.rglob("*"):
if not path.is_file() or path.suffix.lower() not in INCLUDE_EXTENSIONS:
continue
rel = path.relative_to(ROOT).as_posix()
parts = set(path.relative_to(ROOT).parts)
if parts & EXCLUDE_DIRS or rel in EXCLUDE_FILES or path.name in EXCLUDE_FILES:
continue
if any(part.startswith(".") for part in path.relative_to(ROOT).parts[:-1]):
continue
yield path, rel
def _chunks(text: str):
lines = text.splitlines()
if not lines:
return []
step = max(1, CHUNK_LINES - CHUNK_OVERLAP)
result = []
for start in range(0, len(lines), step):
window_start = start
buffer = []
size = 0
for offset, line in enumerate(lines[start:start + CHUNK_LINES]):
if buffer and size + len(line) + 1 > CHUNK_MAX_CHARS:
result.append((window_start + 1, window_start + len(buffer), "\n".join(buffer).strip()))
buffer = []
window_start = start + offset
size = 0
buffer.append(line)
size += len(line) + 1
if buffer:
result.append((window_start + 1, window_start + len(buffer), "\n".join(buffer).strip()))
if start + CHUNK_LINES >= len(lines):
break
return [(start, end, content) for start, end, content in result if content]
def _facts(text: str, rel_path: str):
lines = text.splitlines()
summary = next(
(line.strip().lstrip("#! ").strip() for line in lines[:30] if line.strip()),
None,
)
symbols = re.findall(
r"^\s*(?:class|def|async\s+def|func|struct|enum|protocol|actor|interface)\s+([A-Za-z_]\w*)",
text,
re.MULTILINE,
)
parts = Path(rel_path).parts
module = parts[0] if len(parts) > 1 else None
return module, Path(rel_path).stem, summary, sorted(set(symbols)), len(lines)
class ChunkTooLargeError(Exception):
"""Chunk excede o contexto do modelo de embedding (ver EXCLUDE_FILES/CHUNK_MAX_CHARS)."""
def _embed(text: str):
response = requests.post(
f"{OLLAMA_URL.rstrip('/')}/api/embeddings",
json={"model": EMBED_MODEL, "prompt": f"search_document: {text}"},
timeout=60,
)
if response.status_code == 500 and "context length" in response.text.lower():
raise ChunkTooLargeError(response.text)
response.raise_for_status()
vector = response.json()["embedding"]
if len(vector) != EMBED_DIM:
raise ValueError(f"embedding com {len(vector)} dimensões; esperado {EMBED_DIM}")
return vector
def _hash(text: str) -> str:
return hashlib.md5(text.encode("utf-8")).hexdigest()
def index():
schema = _sql_id(DB_SCHEMA)
conn = _connect()
conn.autocommit = False
indexed = skipped = deleted = chunks_written = 0
seen = set()
try:
with conn.cursor() as cur:
for path, rel_path in sorted(_iter_files(), key=lambda item: item[1]):
try:
text = path.read_text(encoding="utf-8", errors="ignore")
except OSError as exc:
print(f"[RAG] ignorado {rel_path}: {exc}", file=sys.stderr)
continue
seen.add(rel_path)
digest = _hash(text)
cur.execute(f"SELECT content_hash FROM {schema}.indexed_files WHERE file_path = %s", (rel_path,))
row = cur.fetchone()
if row and row[0] == digest:
skipped += 1
continue
module, main_type, summary, symbols, n_lines = _facts(text, rel_path)
cur.execute(f"DELETE FROM {schema}.code_chunks WHERE file_path = %s", (rel_path,))
for index_number, (start, end, content) in enumerate(_chunks(text)):
try:
vector = _embed(content)
except ChunkTooLargeError:
# Chunks densos em tokens (ex: tabelas de dados numéricas
# como font_metrics.py) podem passar de CHUNK_MAX_CHARS em
# caracteres mas estourar o contexto do modelo em tokens.
# Pular o chunk em vez de abortar a transação inteira.
print(f"[RAG] chunk grande demais, pulado: {rel_path}:{start}-{end}", file=sys.stderr)
continue
cur.execute(
f"""INSERT INTO {schema}.code_chunks
(file_path, content, chunk_index, embedding, content_hash,
file_mtime, start_line, end_line, symbols, module, kind)
VALUES (%s, %s, %s, %s, %s, %s, %s, %s, %s, %s, %s)""",
(rel_path, content, index_number, vector, digest,
path.stat().st_mtime, start, end, ", ".join(symbols), module, "window"),
)
chunks_written += 1
cur.execute(
f"""INSERT INTO {schema}.file_index
(file_path, module, main_type, public_symbols, summary, n_lines, content_hash)
VALUES (%s, %s, %s, %s, %s, %s, %s)
ON CONFLICT (file_path) DO UPDATE SET
module = EXCLUDED.module, main_type = EXCLUDED.main_type,
public_symbols = EXCLUDED.public_symbols, summary = EXCLUDED.summary,
n_lines = EXCLUDED.n_lines, content_hash = EXCLUDED.content_hash,
updated_at = CURRENT_TIMESTAMP""",
(rel_path, module, main_type, symbols, summary, n_lines, digest),
)
cur.execute(
f"""INSERT INTO {schema}.indexed_files (file_path, content_hash)
VALUES (%s, %s)
ON CONFLICT (file_path) DO UPDATE SET
content_hash = EXCLUDED.content_hash, updated_at = CURRENT_TIMESTAMP""",
(rel_path, digest),
)
indexed += 1
print(f"[RAG] {rel_path}")
cur.execute(f"SELECT file_path FROM {schema}.indexed_files")
for (rel_path,) in cur.fetchall():
if rel_path not in seen:
cur.execute(f"DELETE FROM {schema}.code_chunks WHERE file_path = %s", (rel_path,))
cur.execute(f"DELETE FROM {schema}.file_index WHERE file_path = %s", (rel_path,))
cur.execute(f"DELETE FROM {schema}.indexed_files WHERE file_path = %s", (rel_path,))
deleted += 1
print(f"[RAG] removido {rel_path}")
conn.commit()
except Exception:
conn.rollback()
raise
finally:
conn.close()
print(f"[RAG] concluído: {indexed} atualizado(s), {skipped} sem mudança, {deleted} removido(s), {chunks_written} chunk(s).")
if __name__ == "__main__":
index()
+1
View File
@@ -0,0 +1 @@
analysis/
+43 -32
View File
@@ -10,12 +10,20 @@ opera **fora** do Final Cut Pro: você exporta o XML, o servidor processa o
documento como dados estruturados e devolve um XML modificado para importação. documento como dados estruturados e devolve um XML modificado para importação.
Nada é patcheado, nenhuma API privada é usada. Nada é patcheado, nenhuma API privada é usada.
Toda a análise foi feita a partir do código-fonte em `server.py` e `fcpxml/`. Toda a análise foi feita a partir do código-fonte. Este README é a visão
geral; o detalhe módulo a módulo mora em `docs/02_MODULES.md`, que é o
documento a manter atualizado quando a estrutura mudar.
> **Guia rápido:** [01 Arquitetura](docs/01_ARCHITECTURE.md) · > **Começando agora?** Leia [01 Arquitetura](docs/01_ARCHITECTURE.md) e depois
> [09 Manutenção](docs/09_MANUTENCAO.md) — o primeiro diz como o sistema é
> dividido, o segundo diz por onde começar a mexer e o que está em aberto.
>
> **Guia completo:** [01 Arquitetura](docs/01_ARCHITECTURE.md) ·
> [02 Módulos](docs/02_MODULES.md) · [03 Server/Tools](docs/03_SERVER_TOOLS.md) · > [02 Módulos](docs/02_MODULES.md) · [03 Server/Tools](docs/03_SERVER_TOOLS.md) ·
> [04 Testes & Workflow](docs/04_TESTS_AND_WORKFLOW.md) · > [04 Testes & Workflow](docs/04_TESTS_AND_WORKFLOW.md) ·
> [05 Experiências](docs/05_EXPERIENCIAS.md) · [06 Boas Práticas](docs/06_BOAS_PRATICAS.md) > [05 Experiências](docs/05_EXPERIENCIAS.md) · [06 Boas Práticas](docs/06_BOAS_PRATICAS.md) ·
> [07 Projeto Ativo no FCP](docs/07_ESTUDO_PROJETO_ATIVO_FCP.md) ·
> [08 App macOS](docs/08_APP_MACOS.md) · [09 Manutenção](docs/09_MANUTENCAO.md)
--- ---
@@ -26,7 +34,7 @@ Toda a análise foi feita a partir do código-fonte em `server.py` e `fcpxml/`.
Python, e reescreve de volta sem perda de sidecars (object tracking, Python, e reescreve de volta sem perda de sidecars (object tracking,
Cinematic). Cinematic).
2. **Uma camada MCP de 62 ferramentas** — expõe análise, edição em lote, QC, 2. **Uma camada MCP de 74 ferramentas** — expõe análise, edição em lote, QC,
geração, exportação cross-NLE, inteligência de mídia (silêncio/beats) e geração, exportação cross-NLE, inteligência de mídia (silêncio/beats) e
edição baseada em transcrição, tudo acessível por um cliente MCP (Claude). edição baseada em transcrição, tudo acessível por um cliente MCP (Claude).
@@ -40,7 +48,7 @@ Toda a análise foi feita a partir do código-fonte em `server.py` e `fcpxml/`.
| Camada | Tecnologia | | Camada | Tecnologia |
|--------|-----------| |--------|-----------|
| Linguagem | **Python 3.10+** (~7.1k linhas em `server.py` + `fcpxml/`) | | Linguagem | **Python 3.10+** (~13k linhas em `server.py`, `server_tools/` e `fcpxml/`) |
| Protocolo MCP | **mcp** (`mcp` SDK), servidor por stdio | | Protocolo MCP | **mcp** (`mcp` SDK), servidor por stdio |
| Parsing XML | **defusedxml** em todos os 4 entry points + `lxml`/`ElementTree` | | Parsing XML | **defusedxml** em todos os 4 entry points + `lxml`/`ElementTree` |
| Tempo racional | frações `numerador/denominador` no formato `"600/2400s"` | | Tempo racional | frações `numerador/denominador` no formato `"600/2400s"` |
@@ -56,26 +64,29 @@ Toda a análise foi feita a partir do código-fonte em `server.py` e `fcpxml/`.
``` ```
G-ART/ G-ART/
├── server.py # MCP server — 62 tools, prompts, resources, dispatch ├── CLAUDE.md # Regras do projeto para o agente
├── fcpxml/ # "Engine" — biblioteca Python de núcleo ├── admin/ # Ponte com o app (fora de code/)
│ ├── models.py # TimeValue, Timecode, Clip, Timeline, enums, QC models │ ├── models_api.py # Entry point: docstring dos comandos + dispatch
│ ├── parser.py # FCPXML → objetos Python (spine, connected clips, roles) │ └── api/ # Os 37 comandos, um módulo por assunto
│ ├── writer.py # Modifica e grava FCPXML (markers, trim, gaps, speed) └── code/
│ ├── rough_cut.py # Gera timelines novas (rough cuts, montages, A/B) ├── server.py # MCP entry point — só dispatch
│ ├── diff.py # Motor de comparação de timelines ├── server_tools/ # Handlers das 74 tools + _shared/
│ ├── export.py # Export DaVinci Resolve v1.9 + FCP7 XMEML v5 ├── fcpxml/ # "Engine" — biblioteca Python de núcleo
│ ├── media_intel.py # Detecção real de silêncio (ffmpeg) e beats (librosa) │ ├── writer/ # PACOTE: edição/escrita (mixins por assunto)
│ ├── transcribe.py # Transcrição Whisper local + edição por transcrição │ ├── models/ # PACOTE: dados por família (timing, timeline…)
│ ├── templates.py # Templates de timeline (intro/outro, lower thirds) │ ├── parser.py # FCPXML → objetos Python
│ ├── live.py # Modo Live — push_to_fcp / list_fcp_libraries │ ├── rough_cut.py # Gera timelines novas
│ ├── safe_xml.py # Wrappers defusedxml + serialize_xml() │ ├── voice_*.py # Pipeline de voz (features → timeline → actions)
│ └── dtd.py # Validação contra DTDs oficiais da Apple │ ├── phrase_review.py # Revisão de frases da etapa 5
├── Engine/ # Esta documentação da arquitetura │ ├── text_layout.py # Diagramação das legendas
├── admin/ # Scripts de manutenção (graphify.sh, graphify.md) │ ├── live.py # Modo Live — push_to_fcp
├── docs/ # WORKFLOWS, CAPABILITY-AUDIT, specs │ ├── safe_xml.py # defusedxml + serialize_xml()
├── examples/ # Fixture de teste (sample.fcpxml) │ └── dtd.py # Validação contra DTDs da Apple
├── tests/ # 1032 testes / 24 suítes ├── MacApp/Sources/ # App SwiftUI (compilado por swiftc)
└── tools/ # Pacote Python (__init__) ├── Engine/ # Esta documentação
├── docs/ # WORKFLOWS, CAPABILITY-AUDIT, specs
├── examples/ # Fixture de teste (sample.fcpxml)
└── tests/ # 1.466 testes / 42 suítes
``` ```
--- ---
@@ -102,7 +113,7 @@ TimeValue(600, 2400) # "600/2400s" == 0.25s
- Soma/subtração compartilham um único caminho `_binop()` (fast-path de mesmo - Soma/subtração compartilham um único caminho `_binop()` (fast-path de mesmo
denominador + alinhamento por LCM). denominador + alinhamento por LCM).
### 4.2 Modelos principais — `models.py` ### 4.2 Modelos principais — `models/`
| Classe | Função | | Classe | Função |
|--------|--------| |--------|--------|
@@ -126,8 +137,8 @@ escrita. `from_xml_element` faz match estrito do atributo `completed`
| Subsistema | Módulo | Função | | Subsistema | Módulo | Função |
|-----------|--------|--------| |-----------|--------|--------|
| Parser | `parser.py` | FCPXML → objetos Python: espinha, connected clips, secondary storylines, roles | | Parser | `parser.py` | FCPXML → objetos Python: espinha, connected clips, secondary storylines, roles |
| Modifier | `writer.FCPXMLModifier` | Edição index-based (clips/resources/formats dicts) do documento existente | | Modifier | `writer/` (`FCPXMLModifier`) | Edição index-based (clips/resources/formats dicts) do documento existente |
| Writer | `writer.FCPXMLWriter` | Gera FCPXML novo a partir de objetos Python | | Writer | `writer/generator.py` | Gera FCPXML novo a partir de objetos Python |
| Rough cut | `rough_cut.py` | Gera timelines (rough cuts, montages, A/B roll) | | Rough cut | `rough_cut.py` | Gera timelines (rough cuts, montages, A/B roll) |
| Diff | `diff.py` | Compara timelines — detecta added/removed/moved/trimmed | | Diff | `diff.py` | Compara timelines — detecta added/removed/moved/trimmed |
| Export | `export.py` | DaVinci Resolve v1.9 + FCP7 XMEML v5 | | Export | `export.py` | DaVinci Resolve v1.9 + FCP7 XMEML v5 |
@@ -151,7 +162,7 @@ assíncrono:
TOOL_HANDLERS = { TOOL_HANDLERS = {
"analyze_timeline": handle_analyze_timeline, "analyze_timeline": handle_analyze_timeline,
"list_clips": handle_list_clips, "list_clips": handle_list_clips,
# ... 62 tools # ... 74 tools, todos em server_tools/
} }
``` ```
@@ -249,7 +260,7 @@ correção, sem depender de lembrar dos dois comandos no pre-commit.
- [docs/02_MODULES.md](docs/02_MODULES.md) — guia módulo a módulo do `fcpxml/` - [docs/02_MODULES.md](docs/02_MODULES.md) — guia módulo a módulo do `fcpxml/`
(responsabilidade, tamanho, APIs públicas). (responsabilidade, tamanho, APIs públicas).
- [docs/03_SERVER_TOOLS.md](docs/03_SERVER_TOOLS.md) — a camada MCP `server.py`, - [docs/03_SERVER_TOOLS.md](docs/03_SERVER_TOOLS.md) — a camada MCP `server.py`,
62 ferramentas, helpers e o padrão de handler. 74 ferramentas, helpers e o padrão de handler.
- [docs/04_TESTS_AND_WORKFLOW.md](docs/04_TESTS_AND_WORKFLOW.md) — suíte de testes, - [docs/04_TESTS_AND_WORKFLOW.md](docs/04_TESTS_AND_WORKFLOW.md) — suíte de testes,
fluxo de trabalho (lint + pytest), execução e estado atual do sistema. fluxo de trabalho (lint + pytest), execução e estado atual do sistema.
- [docs/05_EXPERIENCIAS.md](docs/05_EXPERIENCIAS.md) — **memória de projeto**: - [docs/05_EXPERIENCIAS.md](docs/05_EXPERIENCIAS.md) — **memória de projeto**:
@@ -259,10 +270,10 @@ correção, sem depender de lembrar dos dois comandos no pre-commit.
programação** a aplicar em toda alteração/correção; inclui checklist final. programação** a aplicar em toda alteração/correção; inclui checklist final.
### Outros documentos ### Outros documentos
- [../CLAUDE.md](../CLAUDE.md) — visão geral, key patterns, execução e pre-commit. - [../CLAUDE.md](../../CLAUDE.md) — visão geral, key patterns, execução e pre-commit.
- [../docs/CAPABILITY-AUDIT-2026-06.md](../docs/CAPABILITY-AUDIT-2026-06.md) — - [../docs/CAPABILITY-AUDIT-2026-06.md](../docs/CAPABILITY-AUDIT-2026-06.md) —
auditoria do ecossistema e roadmap dual-mode (XML + Live). auditoria do ecossistema e roadmap dual-mode (XML + Live).
- [../docs/WORKFLOWS.md](../docs/WORKFLOWS.md) — 8 receitas de workflow de produção. - [../docs/WORKFLOWS.md](../docs/WORKFLOWS.md) — 8 receitas de workflow de produção.
- [../docs/specs/](../docs/specs/) — schemas de tools, estrutura FCPXML, pseudocódigo - [../docs/specs/](../docs/specs/) — schemas de tools, estrutura FCPXML, pseudocódigo
do writer, algoritmo de rough cut, implementação do server, roadmap, modelos. do writer, algoritmo de rough cut, implementação do server, roadmap, modelos.
- [../admin/graphify.md](../admin/graphify.md) — pipeline de graphify do código. - [../admin/graphify.md](../../admin/graphify.md) — pipeline de graphify do código.
+132 -64
View File
@@ -1,109 +1,177 @@
# 01 — Arquitetura do Sistema (G-ART / fcp-mcp-server) # 01 — Arquitetura do Sistema (G-ART / fcp-mcp-server)
> Referência canônica de como o sistema está dividido e implementado. Leia este > **Escopo:** Como o sistema é dividido em camadas e onde cada responsabilidade mora.
> documento antes de qualquer mudança de código. > **Não cobre:** Detalhe módulo a módulo (→ 02) · ferramentas MCP (→ 03) · app (→ 08)
> Referência canônica de como o sistema está dividido. Leia antes de qualquer
> mudança de código. Se algo aqui divergir do código, **o código está certo e
> este documento está velho** — corrija-o no mesmo commit.
Última varredura: 2026-08-19 · 77 ferramentas MCP · 1.498 testes · versão `0.6.35`
---
## 1. Visão de cima (camadas) ## 1. Visão de cima (camadas)
O sistema é um **servidor MCP em Python** que lê/analisa/reescreve arquivos O sistema lê, analisa e reescreve **FCPXML** do Final Cut Pro. Ele opera *fora*
**FCPXML** do Final Cut Pro. Há **três camadas** bem separadas: do FCP: você exporta o XML, o programa processa como dados estruturados e
devolve um XML para importar. Nada é patcheado, nenhuma API privada é usada.
São **quatro camadas**, e o ponto importante é que existem **duas portas de
entrada diferentes** para o mesmo motor:
``` ```
┌─────────────────────────────────────────────────────────────┐ ┌──────────────────────────┐ ┌──────────────────────────────┐
│ admin/ — Aplicações complementares (fora do MCP) │ │ MacApp/ (SwiftUI) │ │ Cliente MCP (Claude) │
│ models_api.py API (FastAPI) p/ gerenciar modelos │ │ O app que o usuário usa │ │ Conversa, decide a edição │
│ models_gui.py UI desktop (Flet) p/ gerenciar modelos │ └───────────┬──────────────┘ └───────────────┬──────────────┘
│ graphify.sh/.md Pipeline de graphify do código │ │ subprocesso + JSON-lines │ JSON-RPC (stdio)
├─────────────────────────────────────────────────────────────┤ ▼ ▼
│ server.py — CAMADA MCP / TRANSPORTE (NÃO tem lógica) │ ┌──────────────────────────┐ ┌──────────────────────────────┐
│ 73 tools, handlers, prompts, resources, dispatch │ │ admin/models_api.py │ │ server.py + server_tools/ │
│ Só valida entrada/saída e traduz JSON-RPC → chamadas │ │ + admin/api/ │ │ 77 tools, dispatch, schemas │
├─────────────────────────────────────────────────────────────┤ │ 37 comandos da ponte │ │ NÃO tem lógica de timeline │
│ fcpxml/ — "ENGINE" = NÚCLEO PURO Python (desacoplado) │ └───────────┬──────────────┘ └───────────────┬──────────────┘
│ Não conhece MCP nem argumentos de tool. │ └───────────────┬────────────────────┘
│ Trabalha com objetos Python e XML. │ ▼
│ É o foco / onde quase tudo mora. │ ┌───────────────────────────────────┐
└─────────────────────────────────────────────────────────────┘ │ fcpxml/ — O ENGINE │
│ Núcleo puro Python, desacoplado. │
│ Não conhece MCP nem o app. │
│ É onde quase tudo mora. │
└───────────────────────────────────┘
``` ```
**Regra de arquitetura:** `server.py` NUNCA implementa lógica de timeline — **A regra que sustenta tudo:** nem `server.py` nem `admin/api/` implementam
ele delega ao `fcpxml/`. Tudo em `fcpxml/` é testável isoladamente (1032 testes). lógica de timeline. Os dois validam entrada, chamam o engine e formatam a
saída. Toda regra de negócio é testável sem MCP e sem app.
## 2. Regras transversais (convenções em todo o código) **Por que duas portas.** O MCP existe para o julgamento editorial — qual tomada
usar, onde dar zoom — que é conversa com uma IA. A ponte existe para o que o
usuário faz sozinho no app — transcrever, configurar, processar. As duas caem
no mesmo engine, então uma correção ali vale para as duas.
---
## 2. Regras transversais (valem em todo o código)
| Conceito | Regra | | Conceito | Regra |
|----------|-------| |----------|-------|
| **Tempo** | `TimeValue` fração racional `"600/2400s"`. Nunca use float p/ tempo. | | **Tempo** | `TimeValue`, fração racional `"600/2400s"`. **Nunca float para tempo.** |
| **I/O paths** | Sempre via helpers `_validate_filepath` / `_validate_output_path` (sandbox). | | **Tempo de decisão** | Ações de voz usam sempre segundos da **mídia original**, nunca pós-corte. |
| **Nome de saída** | Nunca sobrescrever original: `output_<suffix>.fcpxml`. | | **I/O paths** | Sempre via `_validate_filepath` / `_validate_output_path` (sandbox). |
| **Segurança XML** | Sempre `defusedxml` (via `safe_xml.py`). Nunca `xml.etree` direto. | | **Nome de saída** | Nunca sobrescrever o original: `generate_output_path()` gera `_suffix`. |
| **Segurança XML** | Sempre `defusedxml` via `safe_xml.py`. Nunca `xml.etree` direto para ler. |
| **Deps opcionais** | `librosa`/`ffmpeg`/`huggingface_hub` importados **lazy**, degradam com `None`. | | **Deps opcionais** | `librosa`/`ffmpeg`/`huggingface_hub` importados **lazy**, degradam com `None`. |
| **Lint** | `ruff check . --exclude docs/` — zero erros. | | **Idioma** | Comunicação com o usuário em português. Código e comentários em inglês. |
| **Validação pós-correção** | `./Engine/run_after_fix.sh` SEMPRE após cada correção. | | **Validação** | `./Engine/run_after_fix.sh` **sempre** após cada correção. |
| **App** | Alterou `MacApp/`? Compile e rode: `admin/run_app.command` (padrão de revisão; equivale a `./MacApp/build_app.sh --run`). |
## 3. Fluxo de um request (round-trip) ---
## 3. Fluxo de um request
### Pela porta MCP (Claude decidindo a edição)
``` ```
Cliente MCP (Claude) Cliente MCP ──JSON-RPC──► server.py
│ JSON-RPC (stdio) │ TOOL_HANDLERS[nome]
▼ ▼
server.py ── dispatcher (TOOL_HANDLERS) server_tools/<categoria>.py
│ valida path, parseia projeto, chama engine │ _shared/: valida path, parseia projeto
▼ ▼
fcpxml/parser.py XML → objetos fcpxml/ (parser → writer → safe_xml)
fcpxml/writer.py edita / grava
fcpxml/rough_cut.py gera novas timelines
fcpxml/export.py cross-NLE
▼ ▼
output_<suffix>.fcpxml (original intocado) projeto_<suffix>.fcpxml (original intocado)
▼
Final Cut Pro: File → Import → XML (ou push_to_fcp, sem cliques)
``` ```
### Pela porta do app (usuário operando)
```
MacApp ──Process + argv JSON──► admin/models_api.py
│ handlers[comando]
▼
admin/api/<assunto>.py
│ shared.emit() devolve JSON-lines
▼
fcpxml/ (ou chama um handler do server)
▼
arquivo gerado + caminho de volta ao app
```
A saída da ponte é **JSON-lines**: um documento JSON por linha, para que
comandos longos transmitam progresso enquanto rodam. Toda escrita passa por
`admin/api/shared.py::emit`, que serializa o acesso a stdout — dois comandos
escrevendo ao mesmo tempo entrelaçariam documentos.
---
## 4. Dual-mode: XML + Live ## 4. Dual-mode: XML + Live
O sistema opera em **dois modos complementares**:
- **Modo XML (principal):** exporta FCPXML, processa como dados, reimporta. - **Modo XML (principal):** exporta FCPXML, processa como dados, reimporta.
Roda fora do FCP. Nenhuma API privada. - **Modo Live (`fcpxml/live.py`):** *push* do FCPXML direto para o FCP em
- **Modo Live (`fcpxml/live.py`):** *push* do FCPXML direto p/ o FCP em execução via Apple events oficiais (`Open Document`). Leitura de bibliotecas
execução via Apple events oficiais (`Open Document`), com `import-options`. via AppleScript read-only.
Leitura de bibliotecas via AppleScript read-only.
**Assimetria estrutural:** import é scriptable, mas a Apple não oferece export **Assimetria estrutural:** import é scriptable, mas a Apple não oferece export
programático — round-trips voltam pelas ferramentas XML. programático. Round-trips sempre voltam pelas ferramentas XML.
---
## 5. Onde está cada responsabilidade ## 5. Onde está cada responsabilidade
| Responsabilidade | Fica em | | Responsabilidade | Fica em |
|------------------|---------| |------------------|---------|
| Modelos de dados (tempo, clips, markers) | `fcpxml/models.py` | | Modelos de dados (tempo, clips, markers, QC, legendas) | `fcpxml/models/` |
| Parse FCPXML → objetos | `fcpxml/parser.py` | | Parse FCPXML → objetos | `fcpxml/parser.py` |
| Editing/escrita (modifier + writer) | `fcpxml/writer.py` | | Edição e escrita de FCPXML | `fcpxml/writer/` |
| Geração de timeline nova | `fcpxml/rough_cut.py` | | Geração de timeline nova | `fcpxml/rough_cut.py` |
| Comparação de timelines | `fcpxml/diff.py` | | Comparação de timelines | `fcpxml/diff.py` |
| Export cross-NLE (Resolve, FCP7) | `fcpxml/export.py` | | Export cross-NLE (Resolve, FCP7) | `fcpxml/export.py` |
| Inteligência de mídia (silêncio/beats) | `fcpxml/media_intel.py` | | Silêncio e beats | `fcpxml/media_intel.py` |
| Transcrição Whisper local | `fcpxml/transcribe.py` | | Transcrição Whisper | `fcpxml/transcribe.py` |
| Diarização (quem falou) | `fcpxml/diarize.py` |
| Ênfase acústica | `fcpxml/emphasis.py`, `fcpxml/voice_features.py` |
| Timeline de voz (o JSON que a IA lê) | `fcpxml/voice_timeline.py` |
| Decisões de edição (cut/zoom/text/marker) | `fcpxml/voice_actions.py` |
| Revisão de frases da etapa 5 | `fcpxml/phrase_review.py` |
| Layout de legendas e métricas de fonte | `fcpxml/text_layout.py`, `font_metrics.py`, `collision.py` |
| Gestão de modelos Whisper | `fcpxml/model_manager.py` | | Gestão de modelos Whisper | `fcpxml/model_manager.py` |
| Templates de timeline | `fcpxml/templates.py` |
| Controle Live do FCP | `fcpxml/live.py` | | Controle Live do FCP | `fcpxml/live.py` |
| Segurança XML (`defusedxml`, `serialize_xml`) | `fcpxml/safe_xml.py` | | Segurança XML | `fcpxml/safe_xml.py` |
| Validação contra DTDs da Apple | `fcpxml/dtd.py` | | Validação contra DTDs da Apple | `fcpxml/dtd.py` |
| Transporte MCP (73 tools) | `server.py` | | Transporte MCP (77 tools) | `server.py` + `server_tools/` |
| Ponte com o app (37 comandos) | `admin/models_api.py` + `admin/api/` |
| Interface do usuário | `MacApp/Sources/` |
## 6. Mapa de dependências (você está aqui se for mexer no X → quem tocar) ---
## 6. Mapa de dependências
``` ```
server.py ──► fcpxml/parser, writer, rough_cut, export, diff, MacApp/ ──► admin/models_api.py (subprocesso, por caminho)
media_intel, transcribe, templates, live, dtd admin/api/ ──► fcpxml/* e, para algumas operações, server.py
admin/models_gui.py ──► fcpxml/media_intel, model_manager, server.py ──► server_tools/*
parser, transcribe server_tools/* ──► server_tools/_shared/ ──► fcpxml/*
admin/models_api.py ──► fcpxml/model_manager fcpxml/writer/ ──► fcpxml/models/, safe_xml, dtd, text_layout, collision
fcpxml/writer.py ──► fcpxml/models, safe_xml, dtd fcpxml/models/ ──► fcpxml/text_layout (só o pacote subtitles)
fcpxml/__init__.py ──► reexporta a API pública fcpxml/__init__.py ──► reexporta a API pública
``` ```
> Se você cria uma **nova ferramenta MCP**, o trabalho principal é em `fcpxml/` **A seta que não existe, e não deve existir:** `fcpxml/` nunca importa de
> (função pura + testes). O handler em `server.py` fica fino: validação de `server_tools/`, de `admin/` ou de qualquer coisa que saiba o que é uma tool.
> caminho → `_parse_project` → chama a função → `_text_result`. Se você precisar disso, a lógica está no lugar errado.
---
## 7. Criando algo novo — por onde começar
| Você quer… | Comece por |
|-----------|-----------|
| Uma **ferramenta MCP** nova | Função pura em `fcpxml/` + teste. O handler em `server_tools/` fica fino. |
| Um **comando do app** novo | Mesmo caminho, e exponha em `admin/api/<assunto>.py` + tabela em `models_api.py`. |
| Uma **tela** nova | `MacApp/Sources/`, consumindo comandos que já existem na ponte. |
| Uma **regra de edição** nova | `fcpxml/` sempre. Se você está escrevendo `if` sobre timeline fora de `fcpxml/`, pare. |
O trabalho principal é **sempre** no engine. As camadas de cima são finas de
propósito: é o que permite testar 1.498 casos sem abrir o app nem subir o MCP.
+232 -96
View File
@@ -1,114 +1,250 @@
# 02 — Módulos do Engine (`fcpxml/`) # 02 — Módulos do Engine (`fcpxml/`)
Guia módulo a módulo do núcleo Python. Tamanho em linhas, responsabilidade e as > **Escopo:** Mapa do engine `fcpxml/`: qual módulo faz o quê e onde mexer.
funções/classes públicas de cada um. APIs públicas são reexportadas em > **Não cobre:** Camadas e regras gerais (→ 01) · handlers MCP (→ 03) · o que está aberto (→ 09)
`fcpxml/__init__.py` (fonte da verdade para o `__all__`).
## Versão atual Mapa módulo a módulo do núcleo Python: onde cada coisa mora e o que ela faz.
`__version__ = "0.6.35"` — ver `fcpxml/__init__.py`. A API pública é reexportada em `fcpxml/__init__.py` — essa é a fonte da verdade
do `__all__`.
Versão: `0.13.1` · Última varredura: 2026-09-22
> **Por que existem pacotes aqui.** `writer.py` tinha 4.199 linhas e `models.py`
> 1.091, cada um com muitos assuntos dentro. Viraram pacotes com um módulo por
> assunto. Do lado de fora **nada mudou**: `from .writer import FCPXMLModifier`
> e `from .models import TimeValue` seguem valendo, porque os `__init__.py`
> reexportam tudo — inclusive os nomes com underscore que a suíte usa.
--- ---
| Módulo | Linhas | Papel | ## Visão geral
|--------|-------:|-------|
| `models.py` | 930 | Data classes e enums (tempo, clips, markers, QC) | | Módulo / pacote | Linhas | Papel |
| `parser.py` | 367 | FCPXML → objetos Python | |-----------------|-------:|-------|
| `writer.py` | 3154 | Edição e escrita de FCPXML (o maior) | | `writer/` | 5.377 | **Edição e escrita de FCPXML** — o coração |
| `models/` | 1.195 | Data classes e enums |
| `text_layout.py` | 901 | Diagramação das legendas dinâmicas |
| `rough_cut.py` | 798 | Geração de timelines novas | | `rough_cut.py` | 798 | Geração de timelines novas |
| `dtd.py` | 112 | Validação contra DTDs oficiais | | `model_manager.py` | 748 | Modelos Whisper: catálogo, download, config |
| `safe_xml.py` | 113 | Wrappers `defusedxml` + `serialize_xml()` | | `voice_timeline.py` | 600 | O JSON de voz que a IA lê |
| `media_intel.py` | 173 | Silêncio (ffmpeg) e beats (librosa) | | `analise.py` | 218 | `AnalisadorDeArquivo` — orquestra transcrição/diarização/ênfase/emoção e monta o `_voice_timeline.json`; usado por `voice_timeline.py` |
| `transcribe.py` | 184 | Transcrição Whisper + edição por transcrição | | `phrase_review.py` | 547 | Revisão de frases (etapa 5 do assistente) |
| `model_manager.py` | 298 | Gestão de modelos Whisper (cache/catálogo) | | `speaker_review.py` | 208 | Revisão de falantes (etapa 3 do assistente) |
| `export.py` | 226 | Export DaVinci Resolve v1.9 + FCP7 XMEML v5 | | `transcription/` | 325 | Pacote: `engine.py` (adapter faster-whisper), `segments.py`/`text.py`/`timestamps.py` (operações puras sobre transcript); `transcribe.py` é a fachada de compatibilidade |
| `diff.py` | 269 | Comparação de timelines | | `collision.py` | 472 | Colisão entre títulos na tela |
| `live.py` | 273 | Modo Live — push_to_fcp / list_fcp_libraries | | `font_metrics.py` | 445 | Largura real de glifos por fonte |
| `templates.py` | 387 | Templates de timeline | | `templates.py` | 387 | Templates de timeline |
| `__init__.py` | 139 | Reexporta API pública | | `parser.py` | 367 | FCPXML → objetos Python |
| `transcribe.py` | 332 | Transcrição Whisper e corte por texto |
| `forced_align.py` | 181 | Alinhamento forçado opcional (whisperx/wav2vec2) que corrige o viés de ~0,4s no início das palavras |
| `live.py` | 273 | Modo Live (push_to_fcp) |
| `diff.py` | 269 | Comparação de timelines |
| `voice_actions.py` | 319 | Decisões de edição (cut/zoom/text/marker) |
| `export.py` | 226 | Export Resolve v1.9 + FCP7 XMEML v5 |
| `voice_features.py` | 220 | Pitch, energia, ritmo, pausas |
| `diarize.py` | 180 | Quem falou (pyannote) |
| `media_intel.py` | 177 | Silêncio (ffmpeg) e beats (librosa) |
| `emphasis.py` | 133 | Índice de ênfase por palavra |
| `safe_xml.py` | 113 | `defusedxml` + `serialize_xml()` |
| `dtd.py` | 112 | Validação contra os DTDs da Apple |
--- ---
## `models.py` — modelos e enums ## `writer/` — edição e escrita
Single source of truth para estrutura de dados. NUNCA mexa aqui sem rodar
`test_models.py`.
- **Tempo:** `TimeValue` (fração racional), `Timecode`. O `FCPXMLModifier` é montado por **composição de mixins**: um mixin por assunto
- **Clips:** `Clip`, `VideoClip`, `AudioClip`, `ConnectedClip` (lane), editorial, todos operando sobre o mesmo documento e os mesmos índices.
`CompoundClip`, `Transition`.
- **Contêineres:** `Timeline`, `Project`, `Keyword`.
- **Markers:** `Marker`, `MarkerType`, `MarkerColor`, `MARKER_XML_TAGS`.
`MarkerType` é o dono da serialização (`from_string`/`from_xml_element`/`xml_attrs`).
Match estrito do atributo `completed` (`'0'`/`'1'`, sem padding).
- **QC:** `SilenceCandidate`, `FlashFrame`, `GapInfo`, `DuplicateGroup`,
`ValidationIssue`, `ValidationResult`.
- **Geração:** `SegmentSpec`, `PacingConfig`, `PacingStyle`, `RoughCutResult`.
## `parser.py` — leitura | Módulo | Linhas | Conteúdo |
- `parse_fcpxml(path)` → `Project`. |--------|-------:|----------|
- `FCPXMLParser` — lê spine, connected clips (lanes), secondary storylines, roles. | `core.py` | 723 | `ModifierCore`: carga, índices, navegação na spine, `save` |
| `titles.py` | 867 | Títulos de texto e legendas dinâmicas |
| `cut.py` | 333 | Dividir, cortar faixas, apagar |
| `speed.py` | 94 | Velocidade de reprodução |
| `zoom.py` | 204 | Zoom (punch-in) via clipe de ajuste conectado |
| `helpers.py` | 279 | Sanitização, escalas, construtores de elemento |
| `rapid.py` | 240 | Flash frames, rapid trim, preencher buracos |
| `validation.py` | 232 | Verificações estruturais antes de salvar |
| `compound.py` | 196 | Compound clips: criar e achatar |
| `silence.py` | 185 | Detectar e remover silêncio |
| `document.py` | 170 | Assets de vídeo, timebases, `write_fcpxml` |
| `markers.py` | 165 | Marcadores: um, por timecode, em lote |
| `audio.py` | 162 | Clipes de áudio e cama musical |
| `generator.py` | 94 | `FCPXMLWriter` — orquestra a criação do zero (estado + delegação) |
| `builders.py` | 174 | Um builder por tipo de elemento: `FormatBuilder`, `AssetBuilder`, `MarkerBuilder`, `KeywordBuilder`, `ClipBuilder`, `SequenceBuilder`, `LibraryBuilder` |
| `adjustment.py` | 140 | `ClipDeAjuste` — camada de ajuste (filtros `filter-video`/`filter-audio` direto no `<clip>`, sem uso ainda em `server_tools`/`admin/api`) |
| `reorder.py` | 126 | Reordenar e recalcular offsets |
| `trim.py` | 125 | Aparar e propagar o ripple |
| `transitions.py` | 94 | Transições entre vizinhos |
| `relink.py` | 94 | Repontar mídia |
| `insert.py` | 78 | Inserir clipes na spine |
| `modifier.py` | 64 | Monta a classe a partir dos mixins |
| `selection.py` | 57 | Selecionar por palavra-chave |
| `api.py` | 55 | Atalhos de uma linha |
| `connected.py` | 49 | Clipes conectados (lanes) |
| `roles.py` | 43 | Atribuir roles |
| `reformat.py` | 43 | Reenquadrar resolução |
## `writer.py` — o coração (3154 linhas) **Onde mexer:** ache o assunto na tabela e abra só aquele arquivo. Se a sua
Duas classes principais: mudança precisa de dois mixins ao mesmo tempo, provavelmente o que você quer
é um método novo no `core.py` que os dois chamem.
- **`FCPXMLModifier`** — edita documento existente de forma index-based **Cuidado:** os mixins compartilham `self`. Um método novo que colida de nome
(dicts de `clips`/`resources`/`formats`), imune a ambiguidade de nomes duplicados. com outro mixin sobrescreve em silêncio — a ordem em `modifier.py` decide quem
Métodos: `insert_clip`, `add_marker`, `trim_clip`, `delete_clip`, `split_clip`, ganha. Hoje nenhum colide; mantenha assim.
`change_speed`, `cut_clip_ranges` (usado pela remoção de silêncio), etc.
- **`FCPXMLWriter`** — gera FCPXML novo a partir de objetos Python.
Helpers de nível de arquivo: `modify_fcpxml`, `add_marker_to_file`,
`trim_clip_in_file`, `build_marker_element`, `write_fcpxml`, `validate_fcpxml`,
`list_effects`, `FCP_EFFECTS`.
## `rough_cut.py` — geração
- `RoughCutGenerator`, `generate_rough_cut`, `generate_segmented_rough_cut`.
## `media_intel.py` — inteligência de mídia (v0.10)
- Silêncio via `ffmpeg silencedetect` (subprocess limitado), `remove_silence_candidates`,
mapeamento source→timeline.
- Beats via `librosa` (import lazy, extra `[intelligence]`).
- Degrada para `None` quando `ffmpeg` ausente.
## `transcribe.py` — Whisper local
- `transcribe(media_path, model_size, language)` → dict com `words` (spans).
- `ALLOWED_MODELS` — allowlist de nomes de modelo (também usado por `model_manager`).
- Edição por transcrição: remove filler words, aparar por transcrição.
## `model_manager.py` — gestão de modelos
Catálogo `models.json` + cache no HF hub. Config em `~/.fcp-mcp-server/config.json`.
Funções: `get/save_models_dir`, `list_installed_models`, `download_model`,
`delete_model`, `get/load_selected_model`, `save_selected_model`, `load_catalog`.
Permite cancelamento de download via `threading.Event`. Segue convenções:
allowlist, lazy imports, degradação graciosa.
## `export.py` — cross-NLE
- `DaVinciExporter` — FCPXML v1.9 p/ DaVinci Resolve.
- Export FCP7 XMEML v5.
## `diff.py` — comparação
- `compare_timelines`, `TimelineDiff`, `ClipDiff`, `MarkerDiff`.
- Detecta added/removed/moved/trimmed clips & markers.
## `live.py` — FCP ao vivo (macOS)
- `push_to_fcp(path, library, options)` — Apple event *Open Document* + `<import-options>`.
Requer `.fcpbundle` p/ zero-click real.
- `list_fcp_libraries()` — AppleScript read-only.
## `templates.py`
- `Template`, `TemplateSlot`, `ClipSpec`, `BUILTIN_TEMPLATES`, `apply_template`,
`list_templates`. Estruturas prontas: intro/outro, lower thirds, music video.
## `safe_xml.py`
Wrappers `defusedxml` centralizados + `serialize_xml()`. Todo parse/escrita passa aqui.
## `dtd.py`
Valida output contra DTDs oficiais no bundle do FCP (via `xmllint`; exige o caminho
do DTD percent-encoded por causa dos espaços em "Final Cut Pro.app").
--- ---
## Como adicionar um módulo novo ## `models/` — dados e enums
1. Criar `fcpxml/<seu_modulo>.py` — função pura, sem conhecer MCP.
2. Reexportar classes/funções em `fcpxml/__init__.py` (`__all__`). Fonte única da estrutura de dados. **Nunca mexa aqui sem rodar `test_models.py`.**
3. Cobrir em `tests/test_<seu_modulo>.py`.
4. Rodar `./Engine/run_after_fix.sh`. | Módulo | Linhas | Conteúdo |
|--------|-------:|----------|
| `timing.py` | 304 | `TimeValue` (fração racional), `Timecode` |
| `timeline.py` | 217 | `Clip`, `ConnectedClip`, `CompoundClip`, `Timeline`, `Project`, `Marker` |
| `enums.py` | 183 | `MarkerType`, `MarkerColor`, `TransitionType`, `PacingStyle`… |
| `subtitles.py` | 157 | `WordLook`, `WordStyle`, `DynamicSubtitleConfig`, paleta |
| `qc.py` | 121 | `FlashFrame`, `GapInfo`, `DuplicateGroup`, `ValidationIssue` |
| `planning.py` | 93 | `SegmentSpec`, `PacingConfig`, `RoughCutResult`, `MontageConfig` |
`MarkerType` é o dono da serialização de marcador (`from_string`,
`from_xml_element`, `xml_attrs`) — não reimplemente isso em outro lugar.
---
## O caminho da voz (do áudio à decisão)
Estes seis módulos formam um pipeline. É o fluxo mais novo e o menos óbvio do
projeto, então vale ler nesta ordem:
```
transcribe.py áudio → palavras com tempo
+
diarize.py quem falou cada trecho
+
voice_features.py pitch, energia, ritmo, pausas
▼
emphasis.py combina tudo num índice 0–1 por palavra
▼
voice_timeline.py monta o _voice_timeline.json ◄── a análise crua
▼
speaker_review.py (opcional) filtra falante mutado + linha riscada
→ _voice_timeline_clean.json ◄── é isto que a IA prefere
▼
[decisão: skill "editar-por-voz", ou a mão do usuário]
▼
voice_actions.py valida a lista de cut/zoom/text/marker
▼
phrase_review.py funde tudo em frases revisáveis (etapa 5 do app)
▼
writer/ aplica no FCPXML
```
**Regra de ouro do pipeline:** toda ação carrega tempo da **mídia original**,
nunca pós-corte. Cortes deslocam tudo depois deles; resolver o deslocamento só
na hora de aplicar (`shift_after_cuts`) elimina uma classe inteira de bug.
### `voice_timeline.py` — o contrato com a IA
Saída em camadas, para um modelo raciocinar do topo e descer só onde importa:
```
{version, source, language,
layers: {transcript, acoustics, speakers, emotion} ← o que rodou de verdade
scales: {…} ← como ler cada número
summary: {…}
speakers: [...]
segments: [{start, end, speaker, text, gap_before, take_boundary,
avg_energy, peak_emphasis, emotion, emotion_confidence,
words: [{text, start, end, energy, pitch_delta, rate_delta,
pause_before, emphasis}]}]}
```
`layers` existe para separar *"a fala é monótona"* de *"a análise acústica nunca
carregou"* — os dois deixam os mesmos zeros nos dados.
### `speaker_review.py` — a triagem antes da IA
Roda logo após `analyze_voice` (etapa 3 do assistente, tela `SpeakerReviewView`
no app): lista quem foi detectado (`speaker_profiles`, com % de fala e falas
de amostra) e a transcrição segmento a segmento, para o usuário nomear cada
falante, mutar quem não interessa (ex.: o entrevistador) e riscar linhas soltas
antes de qualquer IA ver o arquivo. `build_speaker_review` nunca toca a
timeline crua; `apply_speaker_review`/`write_clean_voice_timeline` produzem
uma cópia separada, `_voice_timeline_clean.json`, reaproveitando
`enrich_words`/`_segment_rows`/`_summary` de `voice_timeline.py` para
recalcular a ênfase só sobre quem sobrou — mesma lógica de `restrict_to_kept`,
por falante/segmento em vez de por intervalo de tempo. O merge de decisões
salvas segue o padrão de `phrase_review.merge_saved_decisions`: sempre
reconstrói da análise atual, só as escolhas humanas persistem.
A skill "editar-por-voz", `generate_voice_script` e `cmd_build_phrase_review`
(etapa 4/5, `admin/api/review.py`) preferem o `_clean` quando ele existe; sem
revisão salva, seguem lendo o `_voice_timeline.json` normal — a etapa 3 é
sempre opcional. Os três pontos de leitura precisam concordar nessa
preferência: se um deles voltar a ler o arquivo cru direto, a revisão de
falantes vira letra morta sem nenhum erro visível (ver `05_EXPERIENCIAS.md`).
Cada linha em `build_speaker_review` carrega suas `words` originais (ênfase
por palavra), para a tela desenhar os mesmos chips da etapa 5 sem esperar um
recálculo. `apply_speaker_review` é reaproveitada por dois caminhos: gravar
(`write_clean_voice_timeline`, via `save_speaker_review`) e só **prever**
(`cmd_recalc_speaker_review`, sem tocar disco) — o botão "Recalcular" da tela
usa o segundo caminho para atualizar a ênfase só sobre quem sobreviveu ao
corte, sem reprocessar áudio.
### `phrase_review.py` — a revisão humana
Junta o timeline de voz com as ações da IA numa lista de frases editáveis, e
converte de volta. Frase inativa vira `cut`; ênfase ≥ 1 vira `zoom` mais um
`emphasis_spans` que a etapa de legendas usa. O trim de cada frase anda em
**fronteira de palavra** — cortar é apontar para uma palavra, nunca caçar frame.
---
## Legendas dinâmicas (três módulos que andam juntos)
| Módulo | Papel |
|--------|-------|
| `text_layout.py` | Quebra a frase em linhas e posiciona cada palavra |
| `font_metrics.py` | Largura real de cada glifo na fonte escolhida |
| `collision.py` | Detecta título saindo do quadro ou colidindo com outro |
Estes três não estão divididos porque **cada um já é um assunto só**. O
`text_layout.py` tem 901 linhas de um problema coeso: diagramação.
### Separação de role entre legendas dinâmicas e convencionais
As duas categorias são ambas `<title>` conectados, mas recebem **roles
diferentes** para ficarem didáticas na timeline do FCP (cada role ganha cor
própria no índice). O atributo usado em `<title>` é `role` (CDATA) — **nunca**
`videoRole`, que é DTD-inválido para títulos (ver `05_EXPERIENCIAS.md`,
entrada 32).
| Categoria | `role` | De onde vem |
|-----------|--------|-------------|
| Legendas dinâmicas | `titles.dinamicas` | `DynamicSubtitleConfig.role` / `load_dynamic_subtitle_config()["role"]` |
| Legendas convencionais | `titles.convencionais` | `load_plain_subtitle_config()["role"]` |
A cor do texto em si continua nos configs de fonte (abas do app), não no
role. Os geradores `generate_dynamic_subtitles` (writer/titles.py),
`handle_generate_plain_subtitles` e `handle_generate_subtitles_by_emphasis`
(server_tools/subtitles.py) aplicam o role em cada `<title>` criado; o
parâmetro `role` das ferramentas MCP sobrescreve o default.
---
## Armadilhas do FCPXML (custaram sessões de depuração)
- Tempo é fração: `"3600/2400s"` = 1,5 s.
- `offset` é posição na timeline; `start` é o in-point da mídia.
- `<asset-clip>` (biblioteca) é diferente de `<clip>` (timeline).
- Marcadores são **filhos** do clipe, não irmãos.
- `.fcpxmld` é um **diretório** — sidecars precisam ser copiados no save, ou
dados de object tracking e Cinematic são destruídos.
- Negrito no FCP é `bold="1"` (atributo); itálico é `fontFace` + `italic="1"`.
- `id` de `<text-style-def>` precisa ser XML Name válido — acento, espaço ou
dígito inicial fazem o FCP recusar o arquivo inteiro.
- `code/examples/sample.fcpxml` **não** é DTD-conformante. Não use como fixture
de validade.
+127 -24
View File
@@ -1,26 +1,49 @@
# 03 — Camada MCP (`server.py`) — 73 ferramentas # 03 — Camada MCP (`server.py` + `server_tools/`) — 78 ferramentas
`server.py` (3824 linhas) é a camada de transporte. Não tem lógica de timeline — > **Escopo:** As 78 ferramentas MCP: helpers, categorias e como criar uma nova.
mapeia nome → handler e delega ao Engine. O dispatch é um dicionário > **Não cobre:** Lógica de edição, que mora no engine (→ 02) · comandos do app (→ 08)
`TOOL_HANDLERS` (padrão de despacho, sem cadeias gigantes de if/elif).
`server.py` (592 linhas) é só o transporte: dispatch por dicionário
`TOOL_HANDLERS`, sem cadeia de if/elif e **sem lógica de timeline**. Os handlers
moram em `server_tools/`, um módulo por categoria, e os helpers que todos usam
em `server_tools/_shared/`.
```
server_tools/
editing.py (649) qc.py (696) voice.py (754) timeline.py (400)
subtitles.py markers_import generation.py transcript.py
export.py roles.py live.py
_shared/ ← helpers compartilhados, ver abaixo
```
## Helpers centrais (use-os, não reinvente) ## Helpers centrais (use-os, não reinvente)
| Helper | Linha | Função | Todos reexportados por `server_tools/_shared`, então `from ._shared import X`
|--------|------:|--------| continua funcionando. A coluna diz o módulo real, para quando você precisar
| `_check_json_depth()` | 83 | Rejeita payloads além de 50 níveis | **editar** o helper — ou apontar um `monkeypatch` para ele.
| `_validate_filepath()` | 103 | Sandbox de entrada |
| `_validate_output_path()` | 149 | Sandbox de saída |
| `_format_clip_table()` | 245 | Renderização de tabela |
| `_markdown_table()` | 259 | Renderização de tabela markdown |
| `_parse_project()` | 319 | Parseia FCPXML → `(tree, timeline, project)`; quase todos os handlers começam aqui |
| `_resolve_io_paths()` | 357 | Validação de caminho de entrada/saída |
| `_setup_modifier()` | 390 | Prepara modifier com validação |
| `_setup_generator()` | 414 | Prepara generator com validação |
| `_parse_timestamp_parts()` | 433 | Parse de timestamps (min:seg, H:MM:SS, SMPTE) |
| `_detect_flash_frames/gaps/duplicate_groups()` | 1667+ | Detectores de QC |
## As 73 ferramentas por categoria | Helper | Mora em | Função |
|--------|---------|--------|
| `_validate_filepath()` | `_shared/paths.py` | Sandbox de entrada |
| `_validate_output_path()` | `_shared/paths.py` | Sandbox de saída |
| `_check_json_depth()` | `_shared/paths.py` | Rejeita payloads além de 50 níveis |
| `generate_output_path()` | `_shared/paths.py` | Nome derivado, sem tocar no original |
| `_resolve_io_paths()` | `_shared/paths.py` | Entrada + saída de uma vez |
| `_parse_project()` | `_shared/project.py` | FCPXML → `(tree, timeline, project)`; quase todo handler começa aqui |
| `_setup_modifier()` | `_shared/project.py` | Prepara modifier já validado |
| `_setup_generator()` | `_shared/project.py` | Prepara generator já validado |
| `_text_result()` | `_shared/project.py` | Envolve o texto em `TextContent` MCP |
| `_markdown_table()` | `_shared/formatting.py` | Tabela markdown |
| `_format_clip_table()` | `_shared/formatting.py` | Tabela de clipes |
| `_format_batch_result()` | `_shared/formatting.py` | Relatório de operação em lote |
| `_parse_timestamp_parts()` | `_shared/captions.py` | min:seg, H:MM:SS, SMPTE |
| `parse_srt()` / `parse_vtt()` | `_shared/captions.py` | Legendas coladas |
| `_detect_flash_frames/gaps/duplicate_groups()` | `_shared/detection.py` | Detectores de QC |
| `_load_or_transcribe()` | `_shared/media.py` | Transcrição com cache em disco |
| `_cut_transcript_spans()` | `_shared/media.py` | Corte por trecho falado |
| `_apply_placed_action()` | `_shared/media.py` | Aplica zoom/text/marker já posicionado |
## As 77 ferramentas por categoria
### Timeline & análise (Projeto) ### Timeline & análise (Projeto)
`list_projects`, `analyze_timeline`, `list_clips`, `list_markers`, `list_connected_clips`, `list_projects`, `analyze_timeline`, `list_clips`, `list_markers`, `list_connected_clips`,
@@ -59,18 +82,44 @@ mapeia nome → handler e delega ao Engine. O dispatch é um dicionário
### Voz (análise → decisão → aplicação) ### Voz (análise → decisão → aplicação)
`analyze_voice_features`, `build_voice_timeline`, `refine_voice_timeline`, `analyze_voice_features`, `build_voice_timeline`, `refine_voice_timeline`,
`remove_speakers`, `apply_voice_actions`, `get_voice_analysis_config`, `remove_speakers`, `remove_speech_gaps`, `apply_voice_actions`,
`generate_voice_script`, `get_voice_analysis_config`,
`save_voice_analysis_config`. `save_voice_analysis_config`.
O fluxo é sempre o mesmo: `build_voice_timeline` mede (caro, roda uma vez) → `remove_speech_gaps` corta pelo que a **transcrição** já sabe que não tem
fala — lê `words[].start/end` do `_voice_timeline.json` (função
`speech_gap_cut_actions`, em `fcpxml/voice_actions.py`) em vez de medir
volume. É o complemento correto para o caso que `remove_media_silence`
(silêncio por dB, ver seção "Silêncio e beats") não cobre: um trecho sem
fala mas com som real acima do limiar (respiração, ruído de roupa, batida) —
`remove_media_silence` nunca vai cortar isso porque tecnicamente não é
silêncio. Não corta a lacuna antes da primeiríssima palavra (pode ser quase
o arquivo inteiro, antes da tomada realmente começar) — isso continua
decisão manual na Fase 6 do `apply_voice_actions`
(`.claude/skills/editar-por-voz/criterios/06-texto-corte-marcador.md`).
O fluxo manual é: `build_voice_timeline` mede (caro, roda uma vez) →
o modelo decide os cortes → **`refine_voice_timeline` renormaliza sobre o que o modelo decide os cortes → **`refine_voice_timeline` renormaliza sobre o que
sobrou** (barato, sem reabrir áudio) e propõe as janelas de zoom → o modelo sobrou** (barato, sem reabrir áudio) e propõe as janelas de zoom → o modelo
corta a lista pelo ritmo → `apply_voice_actions` aplica. Pular a renormalização corta a lista pelo ritmo → `apply_voice_actions` aplica. Pular a renormalização
faz o ranking de ênfase apontar para as palavras erradas (ver faz o ranking de ênfase apontar para as palavras erradas (ver
`05_EXPERIENCIAS.md`). `05_EXPERIENCIAS.md`).
`generate_voice_script` é o fluxo **automático e fechado** (sem wizard, sem
copiar-e-colar): transcreve (cache) → `build_voice_timeline` → entrega a
timeline a um **modelo local Ollama** que dirige a edição → devolve o roteiro
legível (markdown) **e** o JSON de ações, e opcionalmente aplica num FCPXML.
O cliente fica em `fcpxml/llm_local.py`; o modelo é tratado como entrada não
confiável e cada ação é validada por `parse_actions`. Padrão:
`qwen2.5:7b-instruct-q4_K_M` (troca de `gemma3:12b` — não cabia em máquina de
8GB de RAM; Gemma 3 4B foi testado antes e falhou por apagar o roteiro
principal em vez de só cortar bastidor). Passe `model=` para usar outro
servido pelo Ollama.
### Legendas dinâmicas (geração → validação → aplicação) ### Legendas dinâmicas (geração → validação → aplicação)
`generate_dynamic_subtitles`, `validate_subtitle_layout`, `transcript_markers`. `generate_dynamic_subtitles`, `generate_plain_subtitles`,
`generate_subtitles_by_emphasis`, `validate_subtitle_layout`,
`transcript_markers`.
**Sempre gere e depois valide — nunca dê a geração como pronta sem **Sempre gere e depois valide — nunca dê a geração como pronta sem
`validate_subtitle_layout`.** A composição garante "sem sobreposição" só `validate_subtitle_layout`.** A composição garante "sem sobreposição" só
@@ -88,6 +137,49 @@ severidade probable/severe → investigar CADA colisão pela fração exata do
XML antes de mudar código (ver checklist abaixo) XML antes de mudar código (ver checklist abaixo)
``` ```
**`generate_subtitles_by_emphasis`** gera as duas legendas numa passada só,
dividindo por palavra: a dinâmica cobre as frases marcadas como ênfase na
etapa 5 (zoom aplicado, nível ≥ 1); a comum cobre **todo o resto** — um bloco
comum simplesmente não é criado onde a dinâmica já cobre. A primeira versão
gerava a comum inteira e desativava (`enabled="0"`) o que ficava sob a
dinâmica, mas um título desativado continua aparecendo como clipe riscado na
timeline do Final Cut mesmo sem renderizar — um corte com bastante ênfase
enchia a trilha de clipes mortos. Trocado por não gerar ali: o preço é que,
se a ênfase for desativada à mão depois, a legenda comum daquele trecho
precisa ser regenerada, não só reativada. É a tradução de `10-revisao-humana.md`
(skill `editar-por-voz`): "a frase de ênfase recebe zoom E legenda dinâmica;
as demais recebem legenda comum". A decisão vem de
`<mídia>_phrase_actions.json["emphasis_spans"]`, escrito por
`save_phrase_review` quando o editor termina a etapa 5 — sem esse arquivo (ou
sem `zoom`/`text` marcados na revisão), a tool gera só a comum, tudo ligado,
e avisa no relatório ("Sem revisão de ênfase"). Não expõe overrides de estilo
por chamada — usa a config salva ("Legendas Dinâmicas"/plain); para estilo
pontual, use `generate_dynamic_subtitles`/`generate_plain_subtitles` direto.
**Separação por role (didática na timeline):** `generate_dynamic_subtitles` e
`generate_plain_subtitles` (e a metade dinâmica/comum do `by_emphasis`) aplicam
`role="titles.dinamicas"` e `role="titles.convencionais"` em cada `<title>`
criado — sub-roles de `titles`, **nunca** `subtitles.*` (que esconderia o título
atrás de Code). O parâmetro `role` de cada ferramenta MCP sobrescreve o default
(vindo de `load_dynamic_subtitle_config()["role"]` /
`load_plain_subtitle_config()["role"]`). Ver `02_MODULES.md` (seção "Separação
de role") e `05_EXPERIENCIAS.md` entrada 32 (DTD: `<title>` leva `role`, não
`videoRole`).
**Compound clip por sub-frase (padrão em `generate_dynamic_subtitles` e na
metade dinâmica do `by_emphasis`):** `compound_subphrases=True` divide cada
frase em sub-frases pela vírgula (`transcribe.split_into_subphrases`) e
empacota os `<title>` de cada uma num `<ref-clip>` — a dúzia de títulos
empilhados por lane que uma frase gera vira uma barra só, arrastável/mutável
como unidade. Exceção: um trecho curto depois da vírgula ("né?", "Então...",
< 3 palavras) funde de volta na sub-frase anterior em vez de virar compound
próprio — soa como parte da mesma respiração, não uma frase nova. A estrutura
replica o que o próprio Final Cut gera em "New Compound Clip": o título mais
cedo vira âncora do spine interno em offset 0, os demais penduram nele por
lane. `validate_subtitle_layout` mede cada compound no seu próprio espaço de
tempo — sem isso, âncoras de compounds diferentes leem "0s" e colidem no
papel mesmo estando segundos distantes na timeline real.
**Antes de atribuir uma colisão ao gerador, confirme que é o gerador.** **Antes de atribuir uma colisão ao gerador, confirme que é o gerador.**
Um `<title>` de nome estranho (`ref` diferente, params tipo `Auto-Shrink`/ Um `<title>` de nome estranho (`ref` diferente, params tipo `Auto-Shrink`/
`Left Margin` que `_make_text_title_clip` nunca escreve) é conteúdo humano `Left Margin` que `_make_text_title_clip` nunca escreve) é conteúdo humano
@@ -144,7 +236,18 @@ Regras:
- Sempre retornam via `_text_result(text)` (envolve o texto em `TextContent` MCP). - Sempre retornam via `_text_result(text)` (envolve o texto em `TextContent` MCP).
## Para adicionar uma ferramenta nova ## Para adicionar uma ferramenta nova
1. Escrever a função no módulo do Engine (`fcpxml/…`) + testes.
2. Criar `handle_<nome>` em `server.py` seguindo o padrão acima. 1. **Escrever a função no Engine** (`fcpxml/…`) com testes. É aqui que mora o
3. Registrar no dicionário `TOOL_HANDLERS`. trabalho de verdade; o resto é encanamento.
2. **Criar `handle_<nome>`** em `server_tools/<categoria>.py`, seguindo o padrão
acima. Escolha a categoria pelo assunto, não pelo tamanho do arquivo.
3. **Declarar o schema** (`Tool(...)`) no mesmo módulo.
4. **Registrar** no `TOOL_HANDLERS` de `server.py`.
5. Rodar `./Engine/run_after_fix.sh`.
Se a ferramenta também deve aparecer no app, exponha um comando equivalente em
`admin/api/<assunto>.py` e registre na tabela de `admin/models_api.py` — ver
`08_APP_MACOS.md`. Uma capacidade que só existe como tool MCP **não existe para
quem usa o app** (foi exatamente o que aconteceu com `apply_voice_actions`,
`05_EXPERIENCIAS.md` #20).
4. Rodar `./Engine/run_after_fix.sh`. 4. Rodar `./Engine/run_after_fix.sh`.
@@ -1,5 +1,8 @@
# 04 — Testes, Fluxo de Trabalho e Estado Atual # 04 — Testes, Fluxo de Trabalho e Estado Atual
> **Escopo:** Como rodar e escrever testes, e o gate antes de commitar.
> **Não cobre:** O que testar em cada módulo (→ 02) · checklist de qualidade (→ 06)
## 1. Suíte de testes ## 1. Suíte de testes
**1032 testes em 24 arquivos** em `tests/`. Rode com `uv run pytest tests/ -v`. **1032 testes em 24 arquivos** em `tests/`. Rode com `uv run pytest tests/ -v`.
+710 -24
View File
@@ -11,6 +11,56 @@ houver uma correção ou trabalho em torno dele, **adicione um registro aqui**
antes de prosseguir. Um problema que se repete em várias tentativas é sinal de antes de prosseguir. Um problema que se repete em várias tentativas é sinal de
que merece entrada. que merece entrada.
> **Como usar:** o índice abaixo é o ponto de entrada. Procure o sintoma
> aqui primeiro; só abra a entrada completa (mais abaixo) se ela for a sua.
> As entradas ficam em ordem cronológica depois do índice.
## Resumo rápido (índice)
| # | Data | Problema | Estado |
|---|------|----------|--------|
| 1 | 2026-08-14 | Início do registro de experiências | `resolvido` |
| 4 | 2026-08-14 | XML fora da grade de frame em NTSC (23.976/29.97fps), confirmado no FCP | `resolvido` |
| 5 | 2026-08-14 | `TimeValue.from_timecode` corrompia segundos decimais em NTSC (3º ponto do bug) | `resolvido` |
| 6 | 2026-08-17 | Clipe-fantasma de 1 frame no início/fim após remoção de silêncio (padding sem vizinho na borda) | `resolvido` |
| 7 | 2026-08-17 | Legendas dinâmicas sobrepondo entre clipes (título conectado não é aparado pelo out-point do pai) | `resolvido` |
| 8 | 2026-08-17 | Importação recusada: `id` de `<text-style-def>` derivado do texto (acentos/espaços/dígito inicial) não é XML Name válido | `resolvido` |
| 9 | 2026-08-17 | Legendas palavra a palavra centradas em vez da composição progressiva diagramada (bloco por trecho, palavra-chave em display italic) | `resolvido` |
| 10 | 2026-08-17 | Cedilha/acentos da display italic invadindo a linha vizinha: empilhamento passou a usar a tinta real por classe de glifo | `resolvido` |
| 11 | 2026-08-18 | Preview das legendas dinâmicas desproporcional ao render do FCP (stagger/gap/canvas divergentes) e `inactive_color` exposto sem efeito | `resolvido` |
| 12 | 2026-08-18 | Espaço de coordenadas do modelo "Text": `fontSize`, `kerning` e `Position` no espaço do quadro — converter só o tamanho descolou o espaçamento | `resolvido` |
| 13 | 2026-08-19 | Reanálise de ênfase implementada no Engine mas sem ferramenta MCP — Fase 4 da skill era inexecutável | `resolvido` |
| 14 | 2026-08-19 | Offset sistemático de ~0,4s no timing por palavra (faster-whisper sem alinhamento forçado) — agora corrigido em pipeline por alinhamento forçado opcional | `resolvido` |
| 15 | 2026-08-19 | `add_zoom` perdia o enquadramento real (voltava a 100%) quando dois zooms caiam no mesmo clipe pós-corte; agora empilha ou substitui conforme as janelas se sobrepõem | `resolvido` |
| 16 | 2026-08-19 | `validate_subtitle_layout` acusava colisão severa em títulos que só se tocam na borda, por não-associatividade de float; 7 de 8 colisões reportadas no teste real eram falso positivo | `resolvido` |
| 17 | 2026-08-19 | Linha de ênfase das legendas dinâmicas sem limite de largura — palavra longa/maiúscula estourava o frame inteiro; auto-fit encolhe até caber, nunca abaixo do corpo | `resolvido` |
| 18 | 2026-08-19 | Legendas dinâmicas geradas com `bold="0" fontFace="Bold"` não renderizam no FCP — negrito deve ser `bold="1"` (atributo) e itálico `fontFace`+`italic="1"` | `resolvido` |
| 19 | 2026-08-19 | `output_dir` usado só como cerca de validação e nunca como destino — toda chamada entre pastas falhava acusando o caminho que ela mesma gerou | `resolvido` |
| 20 | 2026-08-19 | `apply_voice_actions` ausente da ponte e do encadeamento do app — dava para analisar e legendar, não para cortar | `resolvido` |
| 21 | 2026-08-19 | Teste ainda afirmava o default `zoom scale=1.3` removido do parser (agora vem do `zoom_scale` do usuário) | `resolvido` |
| 22 | 2026-08-19 | `VideoPlayer` (AVKit) aborta em runtime no app compilado por `swiftc` — etapa 5 fechava o app; trocado por `AVPlayerLayer` | `resolvido` |
| 23 | 2026-08-19 | Dividir `writer.py` em pacote quebrou `@patch('fcpxml.writer.subprocess')` — a suíte protege comportamento, não localização | `resolvido` |
| 24 | 2026-08-19 | `admin/test_models_api.py` existia mas estava fora de `testpaths` — 13 testes que nunca rodaram | `resolvido` |
| 25 | 2026-08-20 | `admin/api/shared.py` apontava para `admin/code` (inexistente) após a divisão — install editável mascarou o bug em toda validação anterior | `resolvido` |
| 26 | 2026-08-21 | `generate_voice_script` (IA local/Ollama) caía com "Falha ao gerar roteiro por IA local" — prompt embutia a timeline inteira (47k tokens) e estourava `num_ctx`; e `response.json()` de conexão caída escapava como `JSONDecodeError` | `resolvido` |
| 27 | 2026-08-21 | Cortes escritos rente ao timestamp da palavra soam secos — critério da skill e prompt do modelo local não instruíam folga na borda | `resolvido` |
| 28 | 2026-08-21 | Frases desativadas em sequência deixavam fatias de 0,1-0,5s sobrando entre clipes | `resolvido` |
| 29 | 2026-08-21 | `remove_media_silence` (dB) não corta lacuna sem fala mas com som real — trecho sobrevivia intacto na timeline final | `resolvido` |
| 30 | 2026-08-24 | `add_zoom` animava `<adjust-transform>` direto no clipe, diferente de como o FCP realmente exporta zoom (clipe de ajuste conectado) | `resolvido` |
| 31 | 2026-08-24 | Revisão de falantes ("Quem fica na edição") salvava certo, mas etapa 4 (Revisão de frases) lia a timeline crua, ignorando falantes mutados/linhas riscadas | `resolvido` |
| 32 | 2026-08-24 | Separar legendas dinâmicas de convencionais por role: `<title>` aceita `role` (CDATA), NÃO `videoRole` — este último é DTD-inválido para títulos e quebra a validação | `resolvido` |
| 33 | 2026-09-22 | Legenda comum sobreposta à composição dinâmica em `generate_subtitles_by_emphasis`; regenerar acumulava títulos em vez de substituir | `resolvido` |
| 34 | 2026-09-22 | `ClipDeAjuste` gerava wrapper `<adjustment>` inválido no DTD; `code/WHISPERX` era 2,6 GB de backup órfão que inflava o lint quando rodado com `--exclude` explícito | `resolvido` |
| 35 | 2026-09-22 | Indexação RAG (`admin/update_rag.py`) abortava a transação inteira ao achar um chunk que estoura o contexto do modelo de embedding | `resolvido` |
| 36 | 2026-09-23 | `.gitignore` com regra `models/` solta escondia do git o pacote inteiro `fcpxml/models/` (dados do engine), não só o cache do Whisper | `resolvido` |
> Mantenha o índice acima sempre sincronizado com as entradas mais recentes.
---
## Entradas (ordem cronológica)
--- ---
## Como registrar (template de entrada) ## Como registrar (template de entrada)
@@ -79,9 +129,9 @@ Use o bloco abaixo como modelo. Uma entrada = um problema resolvido/reconhecido.
- **Por que isso importa mais do que parece:** o erro contamina toda decisão temporal a jusante — zoom disparava ~0,4s antes da palavra-alvo, `gap_before` subestimava pausas reais na mesma medida (o que afeta diretamente a régua de silêncio recém-adotada), e as folgas de corte saíam erradas nas emendas. - **Por que isso importa mais do que parece:** o erro contamina toda decisão temporal a jusante — zoom disparava ~0,4s antes da palavra-alvo, `gap_before` subestimava pausas reais na mesma medida (o que afeta diretamente a régua de silêncio recém-adotada), e as folgas de corte saíam erradas nas emendas.
- **Decisão tomada:** não rodei `remove_media_silence` bruto sobre o corte. A detecção (ffmpeg, limiar -30dB/0,5s) não distingue "batida entre frases dentro da régua de 1,5s" de "ar morto de emenda" — cortar ambos teria apertado frases fluidas. Corrigi os tempos manualmente medindo o ataque real nos pontos críticos (cabeça, 2 emendas, cauda, 3 zooms) e refiz o corte numa passada só. - **Decisão tomada:** não rodei `remove_media_silence` bruto sobre o corte. A detecção (ffmpeg, limiar -30dB/0,5s) não distingue "batida entre frases dentro da régua de 1,5s" de "ar morto de emenda" — cortar ambos teria apertado frases fluidas. Corrigi os tempos manualmente medindo o ataque real nos pontos críticos (cabeça, 2 emendas, cauda, 3 zooms) e refiz o corte numa passada só.
- **Solução adotada (paliativa, aplicada manualmente neste teste):** medir o RMS real com `ffmpeg -af astats=metadata=1:reset=1:length=0.05,ametadata=print` em janelas curtas ao redor de cada ponto crítico antes de fixar um corte ou zoom que dependa de precisão de frame. Não é o padrão do sistema — é o que cobre a lacuna até o alinhamento forçado existir. - **Solução adotada (paliativa, aplicada manualmente neste teste):** medir o RMS real com `ffmpeg -af astats=metadata=1:reset=1:length=0.05,ametadata=print` em janelas curtas ao redor de cada ponto crítico antes de fixar um corte ou zoom que dependa de precisão de frame. Não é o padrão do sistema — é o que cobre a lacuna até o alinhamento forçado existir.
- **Solução estrutural ainda pendente:** ligar o WhisperX (ou alinhamento forçado equivalente) em `transcribe.py`, o que levaria o erro de ~400ms para ~30ms e corrigiria zoom, corte e `gap_before` de uma vez, sem paliativo por projeto. Não implementado ainda — é mudança de pipeline, exige regerar todos os `_transcript.json`/`_voice_timeline.json` existentes. - **Solução estrutural implementada:** `transcribe.py` agora roda alinhamento forçado fonético (wav2vec2 via whisperx) como passo opcional pós-transcrição, em `fcpxml/forced_align.py` (classe `ForcedAligner`). O erro cai de ~400ms para ~30ms e corrige zoom, corte e `gap_before` de uma vez. É **dependência opcional** (`[align]` extra / pacote `whisperx` do PyPI) — quando ausente ou em qualquer falha, degrada e devolve os tempos brutos sem quebrar a transcrição. O `transcript` traz `"alignment": true/false` e o `voice_timeline` expõe `layers.alignment`, para quem lê o JSON saber se o offset manual ainda é necessário. Não reaproveitamos código da pasta `WHISPERX/` local (problemas conhecidos) — só a ideia documentada aqui. Exige regerar os `_transcript.json`/`_voice_timeline.json` existentes para aplicar nos caches antigos.
- **Aprendizado:** "não reestime tempos no olho" (critério 01) continua certo para decisão *editorial* — mas não cobre erro sistemático de *medição* na fonte dos tempos. Um offset constante e na mesma direção, em vários pontos do material, é sinal de bug no pipeline de transcrição, não de julgamento errado sobre o material. Vale conferir com uma amostra de áudio real antes de confiar cegamente em timestamp de word-level de qualquer fonte nova. - **Aprendizado:** "não reestime tempos no olho" (critério 01) continua certo para decisão *editorial* — mas não cobre erro sistemático de *medição* na fonte dos tempos. Um offset constante e na mesma direção, em vários pontos do material, é sinal de bug no pipeline de transcrição, não de julgamento errado sobre o material. Vale conferir com uma amostra de áudio real antes de confiar cegamente em timestamp de word-level de qualquer fonte nova.
- **Estado:** `parcialmente resolvido` — paliativo documentado e aplicado neste teste; correção estrutural (WhisperX) pendente de implementação. - **Estado:** `resolvido` — alinhamento forçado implementado em `transcribe.py`/`fcpxml/forced_align.py`; paliativo de medição manual mantido apenas para transcripts antigos sem `layers.alignment=true`.
--- ---
@@ -1183,27 +1233,663 @@ o outro; percentil entrega um punhado útil nos dois casos.
--- ---
## Resumo rápido (índice) ## 21 — 2026-08-19 — Teste travado no default antigo de `zoom scale`
| # | Data | Problema | Estado | - **Sintoma:** `tests/test_voice_actions.py::test_default_scale_when_absent`
|---|------|----------|--------| quebrando com `KeyError: 'scale'`, sem relação com a alteração em curso.
| 1 | 2026-08-14 | Início do registro de experiências | `resolvido` | - **Causa raiz:** `parse_actions` deixou de carimbar `scale=1.3` quando o
| 4 | 2026-08-14 | XML fora da grade de frame em NTSC (23.976/29.97fps), confirmado no FCP | `resolvido` | parâmetro vem ausente, justamente para que
| 5 | 2026-08-14 | `TimeValue.from_timecode` corrompia segundos decimais em NTSC (3º ponto do bug) | `resolvido` | `server_tools/_shared.py` use o `zoom_scale` configurado pelo usuário. O
| 6 | 2026-08-17 | Clipe-fantasma de 1 frame no início/fim após remoção de silêncio (padding sem vizinho na borda) | `resolvido` | teste continuou afirmando o default antigo, então passou a acusar como erro
| 7 | 2026-08-17 | Legendas dinâmicas sobrepondo entre clipes (título conectado não é aparado pelo out-point do pai) | `resolvido` | exatamente o comportamento desejado.
| 8 | 2026-08-17 | Importação recusada: `id` de `<text-style-def>` derivado do texto (acentos/espaços/dígito inicial) não é XML Name válido | `resolvido` | - **Solução adotada:** teste reescrito para o contrato novo — um `scale`
| 9 | 2026-08-17 | Legendas palavra a palavra centradas em vez da composição progressiva diagramada (bloco por trecho, palavra-chave em display italic) | `resolvido` | omitido tem que chegar ausente ao aplicador (`test_absent_scale_is_left_absent`).
| 10 | 2026-08-17 | Cedilha/acentos da display italic invadindo a linha vizinha: empilhamento passou a usar a tinta real por classe de glifo | `resolvido` | - **Aprendizado:** quando um default sai do parser e vira configuração, o teste
| 11 | 2026-08-18 | Preview das legendas dinâmicas desproporcional ao render do FCP (stagger/gap/canvas divergentes) e `inactive_color` exposto sem efeito | `resolvido` | que afirmava o valor antigo passa a defender o bug. Ao remover um default,
| 12 | 2026-08-18 | Espaço de coordenadas do modelo "Text": `fontSize`, `kerning` e `Position` no espaço do quadro — converter só o tamanho descolou o espaçamento | `resolvido` | procure o teste que o fixava no mesmo commit — senão ele fica dizendo o
| 13 | 2026-08-19 | Reanálise de ênfase implementada no Engine mas sem ferramenta MCP — Fase 4 da skill era inexecutável | `resolvido` | contrário do código, e a próxima pessoa perde tempo achando que quebrou algo.
| 14 | 2026-08-19 | Offset sistemático de ~0,4s no timing por palavra (faster-whisper sem alinhamento forçado) — corrigido manualmente no teste, WhisperX pendente | `parcialmente resolvido` | - **Estado:** `resolvido`
| 15 | 2026-08-19 | `add_zoom` perdia o enquadramento real (voltava a 100%) quando dois zooms caiam no mesmo clipe pós-corte; agora empilha ou substitui conforme as janelas se sobrepõem | `resolvido` |
| 16 | 2026-08-19 | `validate_subtitle_layout` acusava colisão severa em títulos que só se tocam na borda, por não-associatividade de float; 7 de 8 colisões reportadas no teste real eram falso positivo | `resolvido` |
| 17 | 2026-08-19 | Linha de ênfase das legendas dinâmicas sem limite de largura — palavra longa/maiúscula estourava o frame inteiro; auto-fit encolhe até caber, nunca abaixo do corpo | `resolvido` |
| 18 | 2026-08-19 | Legendas dinâmicas geradas com `bold="0" fontFace="Bold"` não renderizam no FCP — negrito deve ser `bold="1"` (atributo) e itálico `fontFace`+`italic="1"` | `resolvido` |
| 19 | 2026-08-19 | `output_dir` usado só como cerca de validação e nunca como destino — toda chamada entre pastas falhava acusando o caminho que ela mesma gerou | `resolvido` |
| 20 | 2026-08-19 | `apply_voice_actions` ausente da ponte e do encadeamento do app — dava para analisar e legendar, não para cortar | `resolvido` |
> Mantenha o índice acima sempre sincronizado com as entradas mais recentes. ---
## 22 — 2026-08-19 — `VideoPlayer` (AVKit) derruba o app compilado por `swiftc`
- **Sintoma:** "G-ART encerrou inesperadamente" (SIGABRT) toda vez que o
assistente entrava na etapa 5. Nada aparecia na tela antes do crash.
- **Causa raiz:** o app é montado invocando `swiftc` direto
(`MacApp/build_app.sh`), não pelo Xcode. Nesse modo o runtime não consegue
resolver a superclasse Objective-C de `VideoPlayer`:
`failed to demangle superclass of VideoPlayerView from mangled name
'So12AVPlayerViewC'` → `getSuperclassMetadata` chama `fatalError`. É erro de
runtime, então a compilação passa limpa e o problema só aparece ao abrir a
view.
- **Solução adotada:** trocar `VideoPlayer` por um `AVPlayerLayer` dentro de um
`NSViewRepresentable` (`PlayerSurface`/`PlayerLayerView` em
`PhraseReviewView.swift`). Só depende de AVFoundation, que linka normalmente.
Os controles de transporte já viviam na barra da timeline, então não se perde
nada com a chrome do AVKit.
- **Aprendizado:** compilar limpo não prova que um componente de framework
existe em runtime neste build. Ao usar uma view SwiftUI que embrulha uma
classe AppKit/ObjC (AVKit, WebKit, MapKit), abra a tela de fato antes de
concluir. Um harness pequeno (`swiftc` com os mesmos fontes + um `@main` que
monta só aquela view e sai) reproduz o crash em segundos, sem precisar
navegar o app inteiro até lá.
- **Estado:** `resolvido`
---
## 23 — 2026-08-19 — Dividir um módulo em pacote quebra quem faz `patch` nele
- **Sintoma:** ao transformar `fcpxml/writer.py` (4.199 linhas) no pacote
`fcpxml/writer/`, quatro testes passaram a falhar com
`AttributeError: module 'fcpxml.writer' has no attribute 'subprocess'` —
embora nenhuma linha de lógica tivesse mudado.
- **Causa raiz:** os testes usavam `@patch('fcpxml.writer.subprocess.run')`.
Isso não depende da API pública, e sim de *onde o import mora*: com o
módulo dividido, `subprocess` passou a ser importado por
`fcpxml/writer/document.py`, então o alvo do patch deixou de existir.
Re-exportar no `__init__` não resolveria — substituir
`fcpxml.writer.subprocess` não afeta a referência que `document` já tem.
- **Solução adotada:** apontar o patch para o módulo real
(`fcpxml.writer.document.subprocess.run`). Duas armadilhas do tipo foram
evitadas antes: imports relativos precisam de um ponto a mais ao descer um
nível (`from .models` → `from ..models`), inclusive os que ficam *dentro*
de funções, e o `__all__` precisa listar os nomes com underscore que o
resto do projeto já importava, senão a divisão vira quebra de API.
- **Aprendizado:** a suíte protege comportamento, não localização. Antes de
dividir um módulo, procure por `patch('<modulo>.` e por imports relativos
escondidos dentro de funções — são as duas coisas que uma refatoração
puramente mecânica quebra em silêncio, e as únicas que os testes pegam
tarde.
- **Estado:** `resolvido`
---
## 24 — 2026-08-19 — Teste existia, mas estava fora da suíte
- **Sintoma:** `admin/test_models_api.py` (13 testes) nunca rodava. Não
falhava — simplesmente não era coletado, então `models_api.py` figurava
como "coberto" sem que uma única asserção fosse executada em nenhum
commit.
- **Causa raiz:** `testpaths = ["tests"]` no `pyproject.toml`, com o pytest
rodando de `code/`. O arquivo morava em `admin/`, fora do alcance. Rodá-lo
à mão também falhava (`ModuleNotFoundError: admin`), porque a raiz do
repositório não entra no `sys.path` — ou seja, o único jeito de executá-lo
exigia saber de antemão que ele existia e como.
- **Solução adotada:** movido para `code/tests/test_models_api.py`, com o
insert da raiz do repositório no `sys.path` ao lado do import que precisa
dele. Passou a rodar no gate: 1441 → 1454 testes.
- **Aprendizado:** um teste fora de `testpaths` é pior que teste nenhum — ele
dá a sensação de rede sem ser rede. Ao mover ou criar teste fora da pasta
padrão, confirme que a contagem total subiu; se não subiu, ele não está
rodando. Vale também para o lint: `admin/` ainda não é coberto pelo
`run_after_fix.sh`, que roda só dentro de `code/`.
- **Estado:** `resolvido`
---
## 25 — 2026-08-20 — `admin/api/shared.py` apontava para `admin/code` (inexistente)
- **Sintoma:** app do usuário crashava em toda ação que passa por `server`
(ex: "Analisar voz"), com `ModuleNotFoundError: No module named
'server_tools'`. Sobreviveu a **duas rodadas de validação minha** na sessão
anterior — lint zero, 1454 testes verdes, comando testado manualmente pela
ponte — sem nenhuma delas pegar o bug.
- **Causa raiz:** ao dividir `admin/_shared.py` (#25 da sessão de refatoração,
commit `ffaebb3`) em `admin/api/*.py`, o cálculo
`Path(__file__).resolve().parent.parent / "code"` foi copiado sem ajuste.
No arquivo original (`admin/models_api.py`, direto em `admin/`), dois
`.parent` chegam na raiz do repo. Em `admin/api/shared.py`, um nível mais
fundo, dois `.parent` param em `admin/` — e `admin/code` nunca existiu.
`sys.path` nunca recebia `code/`, então `import server_tools` (que só
funciona com `code/` no path) falhava assim que qualquer handler tentava
`from server import ...`.
- **Por que passou pela validação anterior:** todo teste que exercitava esse
caminho importava `admin.api.*` **dentro do processo do pytest**, que já
roda com `cwd=code/` sob um venv com **install editável**
(`__editable__.fcp_mcp_server*.pth`) — isso já deixa `fcpxml`/`server_tools`
importáveis por conta própria, mascarando qualquer erro no cálculo manual
de `sys.path`. O teste manual pela ponte (`uv run python
admin/models_api.py analyze_voice ...`) tem o mesmo problema: `uv run`
ativa o mesmo venv com o mesmo install editável. **Só o app real, chamando
o fallback `python3` sem `uv` ou um venv sem o install editável, expõe o
bug** — que é exatamente a diferença entre o ambiente de teste e o do
usuário.
- **Solução adotada:** o cálculo de `sys.path` saiu de cada módulo de
comando e passou a existir **uma única vez**, em `admin/api/__init__.py`
— que roda antes de qualquer submódulo do pacote, então nenhum deles
precisa da própria cópia. `.parent.parent.parent` (três níveis: `api/` →
`admin/` → raiz → `code/`).
- **Como o teste de regressão foi validado (e por que precisou de duas
tentativas):** a primeira versão do teste também passava com o bug
presente, pelo mesmo motivo do parágrafo acima — rodava em processo com o
install editável ativo. Só ficou confiável rodando um `subprocess` limpo
que remove manualmente qualquer entrada `site-packages` de `sys.path`
antes de importar, isolando o mecanismo real que o `__init__.py` precisa
fornecer. Confirmado nos dois sentidos: falha com o bug reintroduzido,
passa com a correção (`tests/test_models_api.py::TestCodeDirResolution`).
- **Aprendizado:** um install editável no venv de teste é uma segunda fonte
de verdade que mascara bugs de `sys.path` — o mesmo defeito de "a suíte
passa mas o comportamento real não bate" da entrada #23, só que desta vez
nem *rodar o comando manualmente* pegou, porque o `uv run` usado para
testar caía no mesmo venv "de sorte" que o app não usa. Ao validar correção
de caminho/import, rodar num ambiente que não tenha as dependências
instaladas por fora do mecanismo sendo testado — ou o teste prova que o
ambiente de teste está bem configurado, não que o código está certo.
- **Estado:** `resolvido`
---
## Entrada #26 — Prompt da IA local estoura o contexto do Ollama (e erro de parse escapa)
- **Sintoma:** botão "Gerar roteiro por IA local" (etapa 4 do assistente)
devolvia "Falha ao gerar roteiro por IA local". Rodando a ponte direto, o
erro real aparecia como *"Server disconnected without sending a response"*
ou *"Connection refused"* do Ollama, e 0 decisões ("Decisões do modelo: 0").
- **Causa raiz (dupla):**
1. `build_edit_messages` embutia o JSON da voice timeline **inteiro** no
prompt. Uma gravação de 3min vira ~188KB / **~47k tokens** (cada palavra
carrega energia, pitch, arousal, valence, `samples`…). Como `num_ctx`
estava em 32768, o prompt estourava a janela e o Ollama **dropava a
conexão** sem resposta.
2. Quando a conexão cai sem resposta, `httpx` entrega um body vazio e
`response.json()` lançava `JSONDecodeError` — que **não** é
`httpx.HTTPError`, então escapava do `try/except` de `ollama_chat` e
virava a exceção genérica que o `cmd_generate_voice_script` transforma
em `ok:false` com a mensagem "Falha ao gerar roteiro por IA local: …".
- **Correção (em `fcpxml/llm_local.py` + `server_tools/voice.py`):**
- `build_edit_messages` agora projeta a timeline (**`_project_timeline`**):
mantém só `text`/`start`/`end`/`speaker`/`emphasis`/`pause_before` das
palavras e `id`/`name` dos locutores; descarta `layers`, `scales`,
`samples` e os floats de áudio. Caiu de ~47k para **~17k tokens** (69KB).
- Salvaguarda `_shrink_to_fit`: se ainda passar de `max_chars` (110k),
remove os `words` dos segmentos de menor `peak_emphasis` até caber.
- `ollama_chat` envolve `post`+`raise_for_status`+`json()` num único
`except Exception` que relança como `RuntimeError` claro — fim do
`JSONDecodeError` escapando.
- `_extract_json` agora desembrulha a lista de 1 elemento `[{source,
actions}]` que alguns modelos devolvem, senão o `parse_actions` tratava o
objeto-wrapper como uma ação sem `kind` e rejeitava tudo (0 decisões).
- `handle_generate_voice_script` levanta `RuntimeError` com a causa quando o
modelo não devolve nenhuma decisão utilizável, então o app mostra a
mensagem real ("O modelo local não devolveu decisões utilizáveis: …")
em vez do genérico.
- **Validação:** `tests/test_llm_local.py` ganhou `test_build_edit_messages_is_compact`
(prompt < raw, sem `samples`/`energy_raw`/`pitch_hz`) e
`test_ollama_chat_wraps_empty_response`. Ponte testada com Ollama mockado
nos dois sentidos (sucesso aplica; falha → `ok:false` com msg clara).
- **Estado:** `resolvido`
> **Aprendizado:** modelo local tem contexto finito — nunca embutir o objeto
> de análise cru no prompt; projetar só o que a decisão usa. E qualquer parse
> de resposta de servidor local deve tratar body vazio/quebrado como erro de
> transporte, não como sucesso mudo.
---
### 2026-08-21 — Cortes escritos rente ao timestamp da palavra soam secos
- **Sintoma:** usuário revisou o corte final (projeto Mastopexia) e reportou
"os cortes estão muito secos, principalmente no final de frase — falta um
tempinho a mais pra concluir as palavras". Também notou que o ar morto
antes da primeira fala do vídeo não tinha sido cortado.
- **Causa:** o critério `06-texto-corte-marcador.md` (e o prompt embutido do
modelo local em `fcpxml/llm_local.py`) instruíam cobrir a frase inteira
(`start..end = início..fim da frase`) ao escrever um `cut`, sem nenhuma
orientação sobre a borda que encosta em fala **mantida** (não em silêncio
puro). Um `cut` com `start` exatamente no fim da última palavra mantida
engole essa palavra antes dela terminar de soar; um `cut` com `end` no
início exato da próxima engole o ataque da fala seguinte. É um problema
diferente de cortar a pausa curta (proibido, é a própria ênfase) — aqui a
pausa natural entre os blocos já existe, e o corte estava comendo essa
margem sozinho.
- **Correção:**
- `06-texto-corte-marcador.md` ganhou a seção "Nunca corte rente à
palavra — deixe uma folga": recuar `start`/`end` do corte em ~0,15–0,25s
para dentro do próprio corte nas bordas que tocam fala mantida (não em
silêncio puro), incluindo o início/fim do vídeo.
- `fcpxml/llm_local.py::_SYSTEM_PROMPT` (item 4) recebeu a mesma
instrução, para o modelo local gerar decisões já com a folga.
- **Validação manual:** reaplicado no projeto Mastopexia real —
`10.77 → 95.50` (rente) virou `10.97 → 95.30` (folga de ~0,2s nas duas
pontas), e as 4 emendas seguintes receberam o mesmo tratamento; zoom/texto/
marcador continuaram longe o suficiente da nova borda do corte — a folga
também evita o problema relacionado (não corrigido em código, só
contornado manualmente nesta sessão): um `zoom`/`marker` cuja borda cai
exatamente em cima do início/fim de um `cut` é descartado por
`resolve_actions` como "apontando para material cortado", mesmo quando a
intenção era ficar bem ao lado. Vale registrar como dívida: `resolve_actions`
poderia tolerar uma margem de meio-frame antes de considerar a ação "dentro"
do corte.
- **Estado:** `resolvido`
> **Aprendizado:** "cobrir a frase inteira" não é a instrução completa para
> um corte — a frase que **sobra** ao lado do corte também precisa de uma
> borda que respire. Regra prática: só cortar rente ao timestamp quando a
> borda encosta em silêncio real (`gap_before` grande) ou em conteúdo que
> também será descartado; encostando em fala mantida, sempre recuar.
---
### 2026-08-21 — Frases desativadas em sequência deixavam fatias de 0,1-0,5s sobrando
- **Sintoma:** usuário viu, no Final Cut, um clipe minúsculo sobrando entre
dois clipes normais na timeline (projeto Mastopexia, confirmado por
screenshot). Investigação achou 29 `cut`s individuais no
`_phrase_actions.json` gerado pela etapa 5, e a timeline final saiu com
mais de uma dezena de fatias de 0,1-0,5s entre clipes.
- **Causa:** `phrase_review_to_actions()` (`fcpxml/phrase_review.py`) gerava
**um `cut` por frase desativada**, cobrindo só `[phrase.start, phrase.end]`.
Quando duas ou mais frases seguidas estão desativadas, a pausa **entre**
elas nunca pertence a nenhuma frase — não é coberta por nenhum `cut` — e
sobrevive como um clipe próprio, minúsculo, que ninguém pediu para manter.
- **Correção:** `phrase_review_to_actions()` agora agrupa frases desativadas
**consecutivas** (`flush_inactive_run()`) e emite um único `cut` cobrindo do
início da primeira ao fim da última do grupo, absorvendo as pausas entre
elas. Uma frase ativa no meio ainda quebra o grupo — cuts continuam
separados quando há conteúdo mantido entre eles.
- **Validação:** `tests/test_phrase_review.py` ganhou
`test_consecutive_inactive_phrases_merge_into_one_cut`,
`test_inactive_run_at_the_end_still_flushes` e
`test_isolated_inactive_phrases_stay_separate_cuts`. No projeto Mastopexia
real, 29 cuts individuais viraram 3 cuts mescladas; a contagem de fatias
sub-segundo na timeline final caiu de mais de uma dezena para 4 (resíduo
menor, provavelmente do padding do `remove_media_silence` na emenda entre
clipes — não investigado a fundo nesta sessão, ver `09_MANUTENCAO.md`).
- **Estado:** `resolvido` (a causa principal); a sobra residual do
`remove_media_silence` continua como dívida separada.
> **Aprendizado:** "cortar cada frase desativada" não é a mesma coisa que
> "cortar o trecho desativado" quando frases se sucedem sem conteúdo mantido
> entre elas — a pausa entre duas coisas descartadas também precisa ser
> descartada, e ninguém a cobre por definição se o corte for por frase.
---
### 2026-08-21 — `remove_media_silence` (dB) não pega lacuna sem fala com som real
- **Sintoma:** usuário viu, no projeto Mastopexia real, um trecho de ~1,9s
sem fala (imagem parada antes da tomada começar) que sobreviveu intacto
na timeline final — depois de `apply_voice_actions`, `remove_media_silence`
e `generate_dynamic_subtitles` já terem rodado. Achou que era bug de ordem
no encadeamento das etapas ("corta e depois volta").
- **Investigação:** não era ordem. Extraído o áudio real do trecho
(`ffmpeg -af volumedetect`): `mean_volume -21.4dB`, `max_volume 0.0dB` —
longe do limiar padrão de silêncio (-30dB). Rodado `detect_silence` nos
mesmos limiares do sistema (-30/-25/-20/-16dB): nenhum sinaliza o trecho.
O trecho tem som real (roupa, respiração, ambiente) mas nenhuma palavra —
exatamente o caso que `06-texto-corte-marcador.md` já descrevia
("ausência de fala não é ausência de som"), só que sem ferramenta para
agir sobre ele: `remove_media_silence` só enxerga volume, nunca vai
cortar algo que soa alto mas não tem fala.
- **Correção:** nova função pura `speech_gap_cut_actions()` em
`fcpxml/voice_actions.py` — gera `cut`s a partir dos gaps entre
`words[].start/end` do `_voice_timeline.json` (tempo de fonte, como todo
`VoiceAction`), com a mesma folga por dentro (`padding`) que
`speaker_cut_actions()` já usava. Nova tool MCP `remove_speech_gaps`
(`server_tools/voice.py`, mesmo molde de `remove_speakers`): resolve o
`media_path`, lê a timeline, gera as ações e reaplica via
`handle_apply_voice_actions` — não duplica a lógica de corte no FCPXML.
Deliberadamente não corta a lacuna antes da primeiríssima palavra (pode
ser quase o arquivo inteiro, antes da tomada começar de verdade).
- **Ordem revista:** `apply_voice_actions → remove_speech_gaps →
remove_media_silence → generate_dynamic_subtitles` — a lacuna "sem fala"
some primeiro (cobertura ampla, por transcrição), o que sobra de silêncio
técnico *dentro* da fala é apertado depois.
- **Validação:** `tests/test_voice_actions.py::TestSpeechGapCutActions`
(gap acima/abaixo do limiar, lacuna antes da 1ª palavra nunca cortada,
segmentos com palavras sobrepostas não quebram, timeline vazia). Suíte
completa (1508 testes) roda limpa.
- **Estado:** `resolvido`
> **Aprendizado:** um detector de silêncio por dB nunca vai cobrir "sem fala
> com som" — são categorias diferentes, não uma questão de calibrar o
> limiar. Quando já existe transcrição confiável, ela é a fonte melhor para
> "onde não tem fala": não depende de threshold nenhum, só da própria
> palavra existir ou não naquele instante.
---
### 2026-08-24 — Zoom era `<adjust-transform>` no próprio clipe; FCP exporta como clipe de ajuste
- **Sintoma:** usuário pediu para o zoom parar de mexer diretamente no
clipe da timeline e passar a usar um "adjustment clip" com crop
animado — o jeito como ele já fazia zoom manualmente no FCP.
- **Investigação:** não havia amostra real no projeto para confirmar a
forma exata do XML (`adjust-crop`? um `<clip>` com `<adjustment>` como
`fcpxml/writer/adjustment.py` já fazia para filtros?). O usuário enviou
um `.fcpxmld` exportado pelo próprio FCP com um zoom manual
(`exemplo zoom.fcpxmld`), que revelou a forma real: um `<video ref="...">`
referenciando o efeito nativo `FFAdjustmentEffect` ("Clipe de Ajuste"),
anexado numa lane acima do clipe, com seu **próprio** `<adjust-transform>`
animando `scale` de `1 1` até o pico — não `adjust-crop`, e não o wrapper
`<adjustment>` que `adjustment.py` usa (que, conferido contra o DTD real
da Apple, **não existe** — aquele módulo gera XML inválido; ver dívida
em `09_MANUTENCAO.md`). Cruzado com o DTD oficial (`FCPXMLv1_13.dtd`, uma
cópia local encontrada fora do projeto): `<video>` é `%anchor_item;`
válido sem precisar de asset, e `adjust-transform` é filho direto seu.
- **Correção:** `add_zoom` (extraído para `fcpxml/writer/zoom.py`, deixou
de compartilhar módulo com `change_speed`) agora cria um `<video>`
conectado em vez de animar o clipe base. Isso **simplificou** a lógica
antiga: como o clipe de ajuste composita por cima da imagem já
reenquadrada, não precisa mais ler/preservar rotação, posição ou escala
do clipe original (a classe de teste inteira sobre "preservar
enquadramento" — e o bug histórico #15 que ela cobria — deixou de fazer
sentido); e dois zooms disjuntos no mesmo clipe agora são dois `<video>`
irmãos, não um merge de keyframes num `<adjust-transform>` só.
- **Validação:** os 22 testes de zoom em `test_writer.py` reescritos contra
a nova forma (`clip.find('video').find('adjust-transform')...`), mais
`test_voice_actions_tool.py`. Offset/duration da timeline gerada
conferidos byte a byte contra os números reais do `.fcpxmld` de exemplo
(bateram exatamente). Suíte completa roda limpa.
- **Estado:** `resolvido`
> **Aprendizado:** para decisões de forma exata de XML, um exemplo real
> exportado pelo próprio FCP vale mais que qualquer inferência — a diferença
> entre `adjust-crop`, o wrapper inválido de `adjustment.py` e a forma real
> (`<video ref="FFAdjustmentEffect">`) não dava para cravar sem um dos dois
> (amostra real ou o DTD oficial da Apple, que também foi cruzado aqui).
> Peça o exemplo antes de implementar às cegas.
---
### 2026-08-24 — Revisão de falantes salvava certo, mas a etapa 4 nunca lia o resultado
- **Sintoma:** usuário desmarcou falas de bastidor na tela "Quem fica na edição"
(etapa 3, `SpeakerReviewView`) e clicou "Salvar seleção", mas as falas
desmarcadas continuavam voltando na revisão de frases (etapa 4) e no roteiro
final gerado a partir dela.
- **Causa raiz:** `save_speaker_review` (`fcpxml/speaker_review.py`) e o
`_voice_timeline_clean.json` que ela grava estavam **corretos** — conferido
num projeto real: 37 segmentos na timeline crua, 15 marcados `excluded` na
revisão salva, 22 sobrando no `_clean.json` (37-15=22, bate exato). O bug
estava um passo adiante: `cmd_build_phrase_review`
(`admin/api/review.py`), que monta a etapa 4, abria
`args.get("voice_timeline")` — o arquivo **cru** — direto, sem nunca checar
se existia um `_voice_timeline_clean.json` ao lado. `generate_voice_script`
(`server_tools/voice.py`) e `copyForChat` (`WizardView.swift`) já faziam
essa checagem corretamente; só a etapa 4 ficou de fora.
- **Onde:** `admin/api/review.py::cmd_build_phrase_review`.
- **Por que passou despercebido:** a tela de revisão de falantes em si
funcionava e mostrava "Salvo" — o problema só aparecia num passo seguinte
e sem nenhum erro, então parecia que "a seleção não estava sendo salva"
quando na verdade ela salvava certo e era ignorada mais adiante.
- **Solução adotada:** `cmd_build_phrase_review` agora resolve
`speaker_review.clean_voice_timeline_path(timeline_path)` primeiro e lê
esse arquivo quando ele existe, caindo para o cru só na ausência dele —
mesma checagem que os outros dois pontos já faziam.
- **Aprendizado:** quando existem **múltiplos pontos de leitura** de um
mesmo artefato derivado (aqui: três lugares que podem preferir
`_voice_timeline_clean.json` sobre o cru), adicionar a checagem em um novo
ponto de leitura não é opcional — ela precisa ser replicada em todos, ou o
comportamento diverge silenciosamente conforme o caminho que o app tomar.
Vale grepar por todo lugar que abre o arquivo "canônico" sempre que um
arquivo "_clean"/derivado for introduzido.
- **Estado:** `resolvido` — corrigido em `admin/api/review.py`, suíte
completa (1506 de 1508 testes; as 2 falhas restantes são de ambiente —
WhisperX/torchcodec sem libs de sistema, sem relação com a mudança) e
lint do arquivo alterado limpos.
---
### Entrada 32 — 2026-08-24: `<title>` leva `role`, nunca `videoRole`
**Sintoma:** ao atribuir role de vídeo a legendas geradas (para separar
legendas dinâmicas de convencionais na timeline), a validação contra o DTD
FCPXML v1.13 quebrou com `No declaration for attribute videoRole of element
title`.
**Causa:** no DTD da Apple, `<title>` (`<!ATTLIST title %clip_attrs;>` +
`<!ATTLIST title role CDATA #IMPLIED>`) **não** declara `videoRole`. Esse
atributo existe em `<video>`, `<asset-clip>`, `<clip>` etc., mas não em
títulos. `<title>` usa o atributo genérico `role` (CDATA). Confirmado no
`FCPXMLv1_13.dtd` linhas 566–569.
**Decisão:** legendas dinâmicas e convencionais recebem `role="titles.dinamicas"`
e `role="titles.convencionais"` (sub-roles de `titles`, NUNCA `subtitles.*` —
ver entrada sobre roteamento de captions). O campo de config e o parâmetro dos
geradores chama-se `role` (não `video_role`). `assign_role` (mixin `RolesMixin`)
continua correto para clips/vídeos, pois seta `videoRole` neles — não confundir
os dois caminhos.
**Lição:** antes de setar `videoRole` num elemento qualquer, conferir o DTD:
títulos usam `role`. Teste de regressão em `tests/test_dynamic_subtitles.py`
(`test_titles_carry_title_subrole`) garante `titles.*` e bloqueia `subtitles.*`.
> **Nota de reconciliação:** entradas antigas deste arquivo (2026-08-17)
> afirmavam "nenhum título gerado carrega `role`" e tinham o teste
> `test_titles_carry_no_caption_role`. Aquilo referia-se **especificamente**
> a `role="subtitles.*"` (que roteia o título para a pista de captions e o
> esconde). A regra continua válida: proibido `subtitles.*`. O que mudou é que
> agora aplicamos `role="titles.*"` (sub-role de título, válido no DTD e útil
> para separar dinâmicas de convencionais na timeline). O teste foi renomeado
> para `test_titles_carry_title_subrole` e passa a exigir `titles.*` + bloquear
> `subtitles.*`.
---
### Entrada 33 — 2026-09-22: legenda comum sob a composição dinâmica; regenerar acumulava títulos
**Sintoma:** num corte real (Mastopexia), aos 11s a legenda comum "mamas
também mudam. É" aparecia simultaneamente com a composição dinâmica de
ênfase, poluindo o quadro com texto duplicado. Gerar novamente as legendas
(dinâmica ou convencional) sobre um clipe já legendado empilhava um segundo
conjunto de títulos por cima do anterior em vez de substituí-lo.
**Causa raiz — duas falhas distintas:**
1. **Sem marcação de autoria.** Os três handlers de legenda
(`handle_generate_dynamic_subtitles`, `handle_generate_plain_subtitles`,
`handle_generate_subtitles_by_emphasis`) só *adicionavam* títulos —
nenhum removia o que uma chamada anterior tinha gerado. Sem uma forma de
distinguir "título que este programa gerou" de "título que o editor
inseriu manualmente no FCP", uma regeneração não tinha como saber o que é
seguro apagar.
2. **Janela da legenda de ênfase maior que a fala.** Em
`handle_generate_subtitles_by_emphasis`, o cálculo de fim de bloco usava
os segmentos brutos do Whisper (`data["segments"]`) para decidir até onde
a composição dinâmica se estende — não os spans de ênfase revisados
(`spans`). Um segmento do Whisper cobre a frase inteira; a ênfase cobre só
o trecho grifado. A dinâmica então ficava "seguindo" além do próprio
áudio que a originou, invadindo o intervalo onde a legenda comum já
deveria estar sozinha.
- **Onde:** `code/fcpxml/writer/titles.py` (`TitlesMixin`) e
`code/server_tools/subtitles.py` (os três handlers de geração).
- **Solução adotada:**
- Todo título/composição gerado por este programa carrega uma marca em
`<metadata><md key="com.gart.subtitle.kind" value="dynamic|plain">`
(`mark_generated_subtitle`). Um heurístico de compatibilidade
(`_generated_subtitle_kind`) reconhece a assinatura exata de exports
antigos sem a marca (efeito/uid/start de texto do G-ART + padrão de nome),
para não tratar título manual do editor como "nosso" por engano.
- Cada handler chama `remove_generated_subtitles(el, kinds)` no início,
apagando só os títulos com a marca do próprio tipo que está sendo
regerado — títulos manuais e do outro tipo ficam intactos.
- `generate_dynamic_subtitles` ganhou o parâmetro `hold_between_sentences`
(default `True`, preserva o comportamento anterior nas chamadas normais).
`handle_generate_subtitles_by_emphasis` passa `hold_between_sentences=False`
e usa os `spans` de ênfase revisados como `emphasis_segments` (em vez dos
segmentos brutos do Whisper) — a composição dinâmica agora encerra no fim
real da palavra falada quando o próximo bloco pertence a outra frase, e
nunca ultrapassa a janela de ênfase que a gerou.
- `suppress_plain_under_dynamic` recorta (fatiando o clipe do título, sem
duplicar `text-style`) qualquer legenda comum gerada cujo intervalo caia
dentro de uma composição dinâmica ainda ativa — mesmo que o cálculo de
janela de algum outro caminho volte a divergir no futuro, isso funciona
como rede de segurança contra sobreposição visível.
- **Aprendizado:** um gerador que pode ser chamado de novo sobre a mesma
timeline **precisa** de uma forma de reconhecer sua própria saída anterior
antes de decidir "substituir" — sem isso, "regerar" e "empilhar" são
indistinguíveis. E ao derivar o fim de uma janela temporal a partir de uma
fonte (segmentos do Whisper, spans de ênfase, etc.), confirme que a fonte
escolhida tem a granularidade do fenômeno que está sendo delimitado — usar
a fonte "mais larga disponível" por conveniência cria sobra sistemática.
- **Teste de regressão:**
`code/tests/test_subtitle_overlap_regression.py` — roda o handler real
(`handle_generate_subtitles_by_emphasis`) contra um intervalo de ênfase
seguido de uma lacuna de fala comum, e confere que nenhuma composição
dinâmica sobrepõe uma legenda comum; e que chamar o mesmo handler duas
vezes não duplica títulos gerados nem remove um título manual inserido
entre as duas chamadas.
- **Estado:** `resolvido` — 224 testes das suítes de legenda/writer
passando (incl. o novo regressivo); suíte completa 1540 passando, 8
skipped, 1 falha e 1 erro de ambiente sem relação com a mudança (WhisperX/
`extract_pitch` ausente, torchcodec sem libs de sistema); lint dos arquivos
alterados limpo.
---
### Entrada 34 — 2026-09-22: `<adjustment>` inválido no DTD e `WHISPERX` órfão inflando o lint
**Sintoma 1:** `fcpxml/writer/adjustment.py` (`ClipDeAjuste`, código de uma
sessão anterior não commitado) montava
`<clip><adjustment><filter-video .../></adjustment></clip>` para camadas de
ajuste. Nada usava o módulo ainda (sem chamada em `server_tools`/
`admin/api`), mas ficava pronto para alguém reusar do jeito errado.
**Causa 1:** o DTD real da Apple (`FCPXMLv1_13.dtd`) não define nenhum
elemento `<adjustment>`. A produção real de `<clip>` é
`(note?, %timing-params;, %intrinsic-params;, (spine|(%clip_item;)|caption)*,
(%marker_item;)*, audio-channel-source*, (%video_filter_item;)*,
filter-audio*, metadata?)` — ou seja, `filter-video`/`filter-audio` são
filhos diretos do `<clip>`, sem wrapper, e nessa ordem (vídeo antes de
áudio).
**Solução 1:** `ClipDeAjuste.criar()` agora anexa os filtros direto no
`<clip>`, ordenados com vídeo antes de áudio
(`sorted(filtros, key=lambda f: f.tag != "filter-video")`). Teste de
regressão novo: `tests/test_writer_adjustment.py` (sem wrapper, ordem
correta, um `<effect>` por `uid` em `resources`).
**Sintoma 2 (achado ao investigar o mesmo módulo):** um `ruff check .
--exclude docs/` rodado manualmente no início desta sessão acusou **510
erros** — muito acima do que a suíte normalmente reporta.
**Causa 2:** `code/WHISPERX` era uma pasta `.git` solta de **2,6 GB** dentro
de `code/` (não um submodule registrado — sem `.gitmodules`), contendo
cópias/backups congelados do próprio projeto, incluindo uma cópia inteira e
antiga de `fcp-mcp-server-main` dentro de si mesma. O `pyproject.toml` já
excluía `WHISPERX/` do lint por padrão (`[tool.ruff] exclude = ["docs/",
"WHISPERX/"]`), mas passar `--exclude docs/` na linha de comando
**sobrescreve** esse `exclude` em vez de complementá-lo — foi assim que o
lint passou a varrer os 2,6 GB de código velho lá dentro. Confirmado por
grep que só 3 arquivos no código ativo referenciam "WHISPERX", todos em
comentários explicativos (`fcpxml/diarize.py`, `tests/test_diarize.py`,
`admin/api/shared.py`) — nenhum import ou caminho real dependia da pasta.
**Solução 2:** pasta movida para `~/Archives/G-ART-WHISPERX-backup` (fora do
workspace git), copiada com `rsync -a --no-perms` e conferida com
`diff -rq` antes de remover o original. `WHISPERX/` também saiu do
`exclude` do ruff em `code/pyproject.toml` (não faz mais sentido excluir um
caminho que não existe mais em `code/`).
**Aprendizado:** (1) um wrapper de elemento "que faz sentido conceitualmente"
não substitui checar o DTD real antes de escrever o gerador — o padrão do
projeto (`dtd.py`, DTDs em `bm/*/FCPXMLv1_13.dtd`) existe exatamente para
isso. (2) uma flag de linha de comando como `--exclude` em ferramentas de
lint tipicamente **substitui** a config do projeto, não a estende — rodar
`ruff check .` sem flags (herdando `pyproject.toml`) é o comando correto
para refletir o gate real; qualquer variação manual com `--exclude` pode
mentir sobre o estado do lint. (3) uma pasta de backup improvisada dentro do
diretório ativo do projeto (mesmo que "só para não perder nada") é dívida
que cresce sem ninguém perceber — 2,6 GB não apareceram de uma vez.
**Estado:** `resolvido` — `tests/test_writer_adjustment.py` (3 testes)
passando; `admin/` trazido ao lint gate no mesmo commit (ver
`09_MANUTENCAO.md` §2.3); suíte completa 1543 passando, 8 skipped, 1 falha
+ 1 erro pré-existentes de outro trabalho em andamento (sem relação com
esta correção).
### Entrada 35 — 2026-09-22: chunk grande demais derrubava a indexação RAG inteira
**Sintoma:** `admin/update_rag.command` (primeira indexação completa do
G-ART, banco `rag_gart` recém-provisionado) morria sempre no mesmo ponto com
`requests.exceptions.HTTPError: 500 Server Error` na chamada ao Ollama —
sempre logo após imprimir `code/fcpxml/export.py`, ou seja, no arquivo
seguinte na ordem alfabética.
**Causa raiz:** `code/fcpxml/font_metrics.py` é uma tabela de larguras de
glifo (`METRICS = {...}`), texto extremamente denso em tokens (muitos
números/pontuação curtos) — um chunk de ~4900 caracteres (dentro do limite
`CHUNK_MAX_CHARS = 5000`) virou 2653 tokens no tokenizer do
`nomic-embed-text`, estourando o contexto de 2048 tokens do servidor Ollama
local (`llama.cpp`: "input length exceeds the context length"). Reproduzido
isolando o arquivo e chamando `/api/embeddings` chunk a chunk — 6 dos 9
chunks falhavam. `CHUNK_MAX_CHARS` mede caracteres, não tokens; assume
implicitamente ~1 token por poucos caracteres, o que não vale para conteúdo
não-prosa (tabelas numéricas, JSON denso).
Segundo problema, apontado por que a primeira tentativa não recuperou nada:
`admin/update_rag.py::index()` roda a varredura inteira (centenas de
arquivos) em **uma única transação**, com `commit()` só no fim e
`rollback()` em qualquer exceção — um único chunk problemático em um único
arquivo descartava a indexação inteira, mesmo que os outros 300+ arquivos
já tivessem embedado e inserido com sucesso.
**Solução:** `_embed()` agora detecta essa resposta específica do Ollama
(`ChunkTooLarge`, checado por `500` + `"context length"` no corpo) e o loop
principal captura essa exceção por chunk, pula só aquele chunk (aviso em
stderr) e continua o arquivo — sem abortar a transação. Não trunca nem
reduz `CHUNK_MAX_CHARS` globalmente (afetaria todo o corpus por causa de
poucos arquivos atípicos); a lacuna fica só nos poucos chunks realmente
grandes demais, e o resto do arquivo ainda fica pesquisável.
**Aprendizado:** um limite de chunk em caracteres é uma aproximação, não uma
garantia de contexto — arquivos de dados brutos (tabelas, mapeamentos
numéricos, JSON/CSV embutido em `.py`) tokenizam bem mais denso que prosa ou
código comum e podem violar o limite do modelo mesmo dentro do teto de
caracteres. Uma indexação em lote sobre centenas de arquivos não deve ficar
tudo-ou-nada numa única transação: uma falha isolada e recuperável (chunk
específico, arquivo específico) deve ser contida ali, não descartar o
trabalho inteiro já validado.
**Estado:** `resolvido` — indexação completa rodou até o fim: 304 arquivos,
1702 chunks, 0 removidos. Também nesta sessão: criada a pasta `rag/` na raiz
(schema, busca híbrida `search.py`/`search_gart.sh`, `SETUP.md`) — ver
`rag/README.md` para a divisão de responsabilidades com `admin/update_rag.py`.
---
### Entrada 36 — 2026-09-23: `.gitignore` escondia `fcpxml/models/` inteiro do git
**Sintoma:** ao investigar por que um `git diff` de um arquivo recém-editado
(`fcpxml/models/timeline.py`, durante a correção da Entrada 34) não mostrava
nada, `git status` também não listava o arquivo como modificado nem como
untracked — como se ele simplesmente não existisse para o git.
**Causa:** `.gitignore` tinha a regra solta `models/` (comentada como
"WhisperX models cache", pensada para ignorar o cache de ~11 GB de modelos
Whisper baixados em `code/models/`). Uma regra sem `/` inicial no
`.gitignore` casa com **qualquer diretório com esse nome em qualquer
profundidade** — não só `code/models/`, mas também `code/fcpxml/models/`, o
pacote de data classes (`TimeValue`, `Clip`, `Timeline`, `Marker`, etc.) que
sustenta todo o engine. Confirmado: `git ls-tree -r HEAD` não tem nenhum
`fcpxml/models.py` nem `fcpxml/models/` em nenhum commit do histórico — o
pacote inteiro (1.234 linhas, 7 módulos) só existia em disco, sem nenhuma
proteção de versionamento, desde que a divisão de `models.py` em pacote foi
feita (sessão anterior, nunca commitada).
**Risco:** qualquer operação que limpa arquivos não rastreados
(`git clean -fd`, reinstalar do zero, trocar de máquina via `git clone`)
apagaria essa base sem chance de recuperação — nenhum commit para reverter.
**Solução:** regra trocada para `/code/models/` (ancorada na raiz do repo,
só o cache real), preservando `whisper/` (sem uso hoje, mas inofensiva) e
tudo mais. Confirmado com `git check-ignore -v`: `fcpxml/models/timeline.py`
não é mais ignorado; `code/models/models--Systran--faster-whisper-base`
continua ignorado. `fcpxml/models/` passou a aparecer como `??` no
`git status` — visível, pronto para ser commitado quando o dono do trabalho
revisar.
**Aprendizado:** regra de `.gitignore` sem `/` inicial (ex.: `models/`) casa
em qualquer profundidade da árvore — é fácil escrever pensando só no caso
que motivou a regra (um cache na raiz) e esquecer que o mesmo nome de pasta
pode existir, com sentido completamente diferente, dentro do código-fonte.
Regra de bolso: nomes de pasta genéricos (`models/`, `build/`, `cache/`,
`data/`) no `.gitignore` deveriam quase sempre vir ancorados (`/caminho/
exato/`), a menos que a intenção seja mesmo ignorar toda ocorrência do nome
em qualquer lugar da árvore.
**Estado:** `resolvido` — regra corrigida, `fcpxml/models/` confirmado
visível ao git (não commitado ainda; fica para quem já está com esse
trabalho em andamento decidir quando commitar). Nenhum código alterado,
só o `.gitignore`.
+3
View File
@@ -1,5 +1,8 @@
# 06 — Boas Práticas de Programação (G-ART) # 06 — Boas Práticas de Programação (G-ART)
> **Escopo:** Checklist de qualidade a aplicar antes de dar algo por pronto.
> **Não cobre:** Por onde começar uma tarefa (→ 09) · o que já quebrou (→ 05)
> **Propósito:** registrar as melhores práticas de programação a serem aplicadas > **Propósito:** registrar as melhores práticas de programação a serem aplicadas
> **sempre** que qualquer alteração ou correção for feita neste programa. > **sempre** que qualquer alteração ou correção for feita neste programa.
> Servem de checklist obrigatório antes de concluir qualquer mudança. > Servem de checklist obrigatório antes de concluir qualquer mudança.
+204
View File
@@ -0,0 +1,204 @@
# 08 — O app macOS (`MacApp/`) e o Assistente
> **Escopo:** O app SwiftUI e o Assistente: build, telas, ponte e a etapa 5.
> **Não cobre:** Engine Python (→ 02) · ferramentas MCP (→ 03)
O app SwiftUI é como o usuário opera o sistema sem abrir terminal nem conversar
com uma IA. São ~5.500 linhas em `MacApp/Sources/`, e ele **não tem lógica de
edição**: tudo que ele faz é montar argumentos, chamar a ponte Python e mostrar
o resultado.
Última varredura: 2026-08-19
---
## 1. Como o app é construído — leia antes de mexer
**Não existe `.xcodeproj` nem `Package.swift`.** O app é compilado invocando o
`swiftc` direto sobre `MacApp/Sources/*.swift`:
```bash
cd code && ./MacApp/build_app.sh # compila e monta o .app
admin/run_app.command # compila, fecha a instância antiga e abre (padrão de revisão)
```
Consequências práticas, todas já sentidas:
- **Arquivo novo em `Sources/` entra sozinho** no build. Não há lista de alvos.
- **Não dá para adicionar dependência SPM** sem antes migrar o build inteiro.
- **Compilar não prova que roda.** Componentes SwiftUI que embrulham classes
Objective-C podem falhar só em tempo de execução, ao abrir a tela. Foi o que
aconteceu com `VideoPlayer` (AVKit): compilava limpo e abortava ao abrir a
etapa 5 (`05_EXPERIENCIAS.md` #22). Por isso a regra: **alterou a interface,
abra a tela de fato.**
### Testando uma tela sem navegar o app inteiro
Um harness de vinte linhas compila os mesmos fontes com um `@main` próprio que
monta só a tela em questão. Reproduz crash de runtime em segundos:
```bash
swiftc -parse-as-library -sdk "$(xcrun --sdk macosx --show-sdk-path)" \
-target arm64-apple-macosx26.0 \
MacApp/Sources/PhraseReviewView.swift MacApp/Sources/PhraseReviewModel.swift \
MacApp/Sources/TimelineTracksView.swift MacApp/Sources/Models.swift \
MacApp/Sources/PythonBridge.swift /tmp/HarnessMain.swift -o /tmp/harness
```
O `@main` do harness carrega a tela, imprime o que interessa e chama
`NSApplication.shared.terminate` — dá para afirmar "abriu e funcionou" sem
depender de screenshot.
---
## 2. Estrutura das telas
| Arquivo | Linhas | Papel |
|---------|-------:|-------|
| `WizardView.swift` | 808 | **O Assistente** — fluxo guiado de 7 etapas |
| `TranscriptionView.swift` | 843 | Transcrição avulsa e processamento em lote |
| `ModelDownloadView.swift` | 545 | Catálogo e download de modelos Whisper |
| `CaptionsView.swift` | 545 | Legendas dinâmicas: estilo + preview ao vivo |
| `TimelineTracksView.swift` | 506 | Timeline com trilhas, zoom e playhead |
| `PhraseReviewModel.swift` | 429 | Estado da etapa 5: frases, player, zooms |
| `PhraseReviewView.swift` | 413 | Etapa 5: preview + inspector de frases |
| `VoiceAnalysisView.swift` | 322 | Parâmetros do motor de ênfase |
| `ProjectView.swift` | 293 | Inspeção do `.fcpxml` |
| `Models.swift` | 274 | Espelhos Swift do JSON da ponte |
| `PythonBridge.swift` | 230 | **A ponte** — ver seção 3 |
| `SubtitlePreviewView.swift` | 218 | Preview 9:16 das legendas |
| `App.swift` | 69 | `NavigationSplitView` e as abas |
Abas (`ActiveTab` em `App.swift`): Assistente · Projeto · Legendas · Análise de
Voz · Modelos · Sobre. As cinco últimas são "Avançado" — atalhos para operações
soltas. O Assistente é o caminho principal.
---
## 3. `PythonBridge.swift` — como o app fala com o Python
O app lança `admin/models_api.py` como **subprocesso**, passando o comando e um
JSON como `argv`, e lê **JSON-lines** no stdout.
```swift
PythonBridge.call(command: "build_phrase_review",
arguments: ["voice_timeline": path]) { result, error in … }
```
Dois pontos que já causaram problema e estão resolvidos no código — não os
desfaça sem entender:
- **`uv run` precisa rodar com cwd em `code/`.** O `uv` escolhe o ambiente pelo
diretório do processo, não pelo caminho do script. Rodar da raiz fazia o `uv`
criar um segundo `.venv` vazio e ignorar tudo que estava instalado em
`code/.venv` — librosa e pyannote instalavam com sucesso e o app insistia que
faltavam.
- **`scriptURL` procura `admin/models_api.py`** subindo diretórios a partir do
cwd, do bundle e do home. É o que faz o app funcionar tanto rodando do Xcode
quanto do `.app` montado.
- **O `sys.path` que torna `fcpxml`/`server_tools` importáveis dentro de
`admin/api/` mora só em `admin/api/__init__.py`.** Não copie esse cálculo
para um módulo de comando individual — foi exatamente essa cópia,
desatualizada em um nível de diretório, que quebrou toda ação que passa por
`server` (`05_EXPERIENCIAS.md` #25). E não confie em "testei com `uv run` e
funcionou": esse comando roda no mesmo venv com install editável que
mascara esse tipo de erro. O teste que pega de verdade é
`tests/test_models_api.py::TestCodeDirResolution`.
Para adicionar um comando: função em `admin/api/<assunto>.py`, registro na
tabela de `admin/models_api.py`, e `PythonBridge.call` do lado Swift. Os 37
comandos e seus formatos estão documentados no docstring de `models_api.py`.
---
## 4. O Assistente — as 7 etapas
`WizardStep` (`WizardView.swift`) é um enum sequencial; `canAdvance` decide
quando o botão "Continuar" libera.
| # | Etapa | O que acontece | Comando da ponte |
|---|-------|----------------|------------------|
| 1 | Projeto | Escolhe a pasta de saída e o `.fcpxml` | `project_config` |
| 2 | Transcrever | Transcreve toda a mídia do projeto | `transcribe` |
| 3 | Analisar voz | Mede ênfase, locutores, emoção | `analyze_voice` |
| 4 | Decisões da IA | Copia para o chat **ou** gera por IA local (Ollama/Gemma 3), aplica | `apply_voice_actions` / `generate_voice_script` |
| 5 | **Revisar ênfases** | Lapida frase a frase — ver seção 5 | `build_phrase_review` / `save_phrase_review` |
| 6 | Processar | Silêncios, preenchimento, legendas | vários, em cadeia |
| 7 | Concluído | Abre no FCP ou mostra no Finder | — |
**A etapa 4 tem duas saídas:**
- **Manual (chat):** o app monta o pedido pronto no clipboard (skill `editar-por-voz`) e recebe o JSON de volta — o julgamento de qual tomada usar e onde dar zoom fica com a IA numa conversa.
- **Automática (IA local):** botão "Gerar roteiro por IA local (Ollama/Gemma 3)". Ele manda a *voice timeline inteira* (o arquivo) junto com o brief para um modelo local (Ollama), que decide cortes/zooms/textos de uma vez, devolve o roteiro legível + o JSON de ações e já aplica no FCPXML (non-destructive). Não precisa sair do app nem colar nada. O modelo é escolhido num **picker que lista os modelos instalados no Ollama** (populado via `list_ollama_models` quando a etapa abre); se o Ollama estiver fora do ar, cai para um campo de texto livre. Troque para `llama3` etc. se tiver outro modelo. Requer o Ollama rodando em `localhost:11434`.
**Etapa 1 — armadilha registrada:** não escolha como "o projeto" um arquivo já
gerado pelo fluxo (`_voice_edit`, `_silence_removed`, …). Os cortes de voz
assumem timestamps da mídia **original**; reaplicá-los sobre um arquivo já
cortado desloca tudo em silêncio. O wizard avisa (`looksLikeGeneratedFile`).
---
## 5. Etapa 5 — a sala de edição
Única tela que ocupa a janela toda: o corpo do wizard é uma coluna de 640pt, e
essa etapa escapa dela porque precisa da largura (`step == .revisar` em
`WizardView.body`).
```
┌────────────────────────────┬──────────────┐
│ Preview (AVPlayerLayer) │ Inspector │
│ enquadrado no formato │ de frases │
│ de entrega do projeto │ │
├────────────────────────────┴──────────────┤
│ Timeline: 6 trilhas, zoom, playhead │
└───────────────────────────────────────────┘
```
**Trilhas:** zooms · frases · energia por palavra · emoção · locutor ·
roteiro/bastidor. Todas desenhadas sobre o mesmo eixo de tempo, com uma coluna
fixa à esquerda nomeando cada uma.
**O que o usuário decide por frase:** nível de ênfase (0–3), ativo/inativo,
texto, roteiro/bastidor e o trim das pontas. O trim anda em **fronteira de
palavra** — cortar é apontar para uma palavra, arrastando a borda do bloco ou
clicando na palavra no inspector.
**Zoom manual:** arrastar na timeline marca um trecho; botão direito cria um
zoom nele. O zoom guarda **só o quando** — escala e ramp vêm das configurações
de Análise de Voz no momento do render, então mudar lá restiliza todos.
**Decisões de implementação que parecem detalhe e não são:**
- **O preview não renderiza nada.** Ele toca a mídia original e *pula* os
trechos removidos. Renderizar para conferir um toggle poria minutos entre a
decisão e o resultado. O observador roda a 60 Hz porque o período dele é
exatamente quanto de material cortado dá para ouvir antes do pulo.
- **O enquadramento é o do projeto, não o da mídia.** As gravações são
horizontais e a entrega é vertical; o app lê o formato do `.fcpxml`
(`inspect`) e mostra o corte central aproximado, com um selo para alternar
para a mídia original. O enquadramento real de cada clipe vem do FCP — o
preview é aproximação, e o selo diz isso.
- **Nada é processado aqui.** "Continuar" grava o `_phrase_review.json` e o
`_phrase_actions.json` derivado dele. A geração é da etapa 6.
- **A revisão é sempre remontada da análise atual**, com as decisões salvas
reaplicadas por cima (`merge_saved_decisions`). Assim refazer a análise de voz
melhora a tela em vez de ficar mascarado por uma cópia velha; uma decisão cuja
frase se moveu mais de 0,25 s é descartada em vez de colar na frase errada.
---
## 6. Estado atual e o que falta
**Funciona e foi verificado:** carga das frases com decisões da IA, as 6
trilhas, seleção sincronizada nos três painéis, trim por palavra, zoom manual,
reprodução parando no ponto exato (erro de 0 ms medido), pulo dos trechos
removidos, enquadramento vertical, gravação ao avançar.
**Ainda em aberto:**
- **A etapa 6 não consome o `_phrase_review.json`.** A ligação — zoom e legenda
dinâmica só nas frases de ênfase, legenda comum no resto — é a próxima tarefa.
- **`MacApp/` não tem teste automatizado.** A rede é o harness da seção 1 e o
olho do usuário. Toda mudança de interface precisa ser aberta de fato.
- **O preview aproxima o reenquadramento vertical** pelo corte central; se os
clipes forem reposicionados no FCP, diverge.
+193
View File
@@ -0,0 +1,193 @@
# 09 — Manutenção: onde mexer, o que está aberto, o que dói
> **Escopo:** Por onde começar cada tipo de tarefa, o que está aberto e onde dói.
> **Não cobre:** Como as coisas funcionam — este doc roteia para quem explica
Este é o documento de rota. Os outros descrevem o que **é**; este diz o que
**fazer** e por onde começar quando chega uma implementação, uma melhoria ou
uma correção.
Última varredura: 2026-09-22 · 1.543 testes passando (+1 falha pré-existente em `test_forced_align.py` e +1 erro pré-existente em `test_refine_voice_timeline_tool.py`, ver §2.6) · lint zerado em `code/` e em `admin/` (fora de server.py/ai_edit.py/fcpxml/analise.py, pré-existentes — outro trabalho em andamento na branch)
---
## 1. Chegou uma tarefa — por onde começo?
| A tarefa é… | Comece em | Não esqueça |
|-------------|-----------|-------------|
| Regra nova de edição (corte, zoom, legenda) | `fcpxml/<módulo>` + teste | Expor na tool **e** na ponte, senão só metade dos usuários alcança |
| Corrigir XML que o FCP recusa | `fcpxml/writer/` + `dtd.py` | Validar contra o DTD real, não só o teste |
| Mudança visível na interface | `MacApp/Sources/` | **Abrir a tela** — compilar não prova nada (§4) |
| Comando novo para o app | `admin/api/<assunto>.py` | Registrar na tabela de `models_api.py` |
| Ferramenta MCP nova | `server_tools/<categoria>.py` | Schema `Tool(...)` + `TOOL_HANDLERS` |
| Ajuste de análise de voz | `fcpxml/voice_*`, `emphasis.py` | Regerar os `_voice_timeline.json` de teste |
| "Está lento" / "está errado" e não sei onde | §5 (mapa de sintomas) | — |
**A pergunta que resolve 90% das dúvidas de lugar:** essa lógica precisa saber
o que é uma tool MCP ou uma tela? Se não precisa — e quase nunca precisa — ela
vai para `fcpxml/`.
---
## 2. O que está aberto agora
Ordenado por quanto atrapalha, não por esforço.
### 2.1 `resolve_actions` não tolera margem no encosto de zoom/marker contra um corte
Um `zoom`/`marker` cuja borda cai exatamente em cima do `start`/`end` de um
`cut` é descartado como "apontando para material cortado" — mesmo quando a
intenção era ficar bem ao lado. Contornado manualmente no projeto Mastopexia
(recuando as bordas na mão); a correção estrutural é dar a `resolve_actions`
uma margem de tolerância (meio frame) antes de considerar uma ação "dentro"
do corte. → `fcpxml/voice_actions.py` (`resolve_actions`/`shift_after_cuts`),
`05_EXPERIENCIAS.md` #27.
### 2.2 `MacApp/` não tem teste automatizado
5.500 linhas de Swift sem uma asserção. A rede hoje é o harness manual (§4) e
o olho do usuário. Não é para sair criando suíte de UI — mas lógica pura que
foi parar na camada de tela (cálculo de trim, mapeamento de tempo) deveria
descer para o Python, onde já existe rede.
### 2.3 ~~`admin/` fica fora do lint~~ — resolvido em 2026-09-22
`run_after_fix.sh` agora roda um segundo passo (`ruff check --config
pyproject.toml ../admin/`) com a mesma config do engine. Precisou de
`# noqa: E402` em 6 imports de `admin/models_api.py`/`admin/models_gui.py`
(padrão `sys.path.insert` antes do import local, convenção já usada no
projeto). Lint de `admin/` está zerado.
### 2.4 Confirmações visuais pendentes no FCP
Várias entradas do `05_EXPERIENCIAS.md` estão marcadas como resolvidas *no XML*
— testes verdes, DTD válido — mas **pendentes de importação real no Final Cut**.
XML válido não é o mesmo que XML que renderiza como o esperado. Ao mexer em
legenda, zoom ou keyframe, a confirmação final é abrir no FCP.
### 2.5 `remove_media_silence` deixa fatias sub-segundo nas emendas entre clipes
Mesmo depois de corrigir o merge de cortes consecutivos (`05_EXPERIENCIAS.md`
#28), sobraram 4 clipes de 0,07-0,23s no projeto Mastopexia real, todos bem
na emenda entre dois clipes vizinhos — mesma família do #6 (clipe-fantasma de
1 frame por padding sem vizinho na borda), mas não confirmado se é a mesma
causa raiz. Não investigado a fundo ainda.
→ `fcpxml/writer/cut.py` (`cut_clip_ranges`, `min_keep_seconds`), padding do
`remove_media_silence`.
### 2.6 ~~`fcpxml/writer/adjustment.py` gerava um wrapper `<adjustment>` inválido~~ — resolvido em 2026-09-22
`ClipDeAjuste` embrulhava filtros num `<clip><adjustment>...</adjustment></clip>`,
que não existe no DTD real da Apple. Corrigido para anexar
`filter-video`/`filter-audio` direto como filhos do `<clip>` (na ordem que o
DTD exige: vídeo antes de áudio). Teste de regressão em
`tests/test_writer_adjustment.py`. Segue sem uso em `server_tools`/`admin/api`
— só deixou de estar pronto pra alguém reusar do jeito errado.
→ `05_EXPERIENCIAS.md` #34.
### 2.7 `test_refine_voice_timeline_tool.py` quebrado: `voice_timeline.extract_pitch` ausente
`TestRefineVoiceTimelineHandler::test_max_zooms_caps_the_list` tenta
`monkeypatch.setattr(vt, "extract_pitch", ...)` mas `fcpxml/voice_timeline.py`
não tem mais (ou nunca teve, nesta branch) essa função. Pertence ao trabalho
de análise de voz já em andamento nesta branch (`voice_timeline.py`
modificado, não commitado) — não investigado a fundo, só registrado aqui
para não se perder.
→ `fcpxml/voice_timeline.py`, `tests/test_refine_voice_timeline_tool.py`.
### 2.8 ~~Submódulo `WHISPERX` com conteúdo modificado e não commitado~~ — resolvido em 2026-09-22
Não era um submódulo git registrado (sem `.gitmodules`) — era uma pasta
`.git` solta de 2,6 GB dentro de `code/`, com cópias/backups congelados do
próprio projeto (`WHISPERX_backup_88476/`, uma cópia inteira e antiga de
`fcp-mcp-server-main`). Só 3 referências no código ativo, todas em
comentários (`fcpxml/diarize.py`, `tests/test_diarize.py`,
`admin/api/shared.py`), nenhum import ou caminho dependia dela. Além do
peso morto, ela também inflava qualquer lint rodado com `--exclude`
explícito (que sobrescreve o `exclude` do `pyproject.toml`) — foi assim que
um `ruff check . --exclude docs/` chegou a acusar 510 erros, quase todos
dentro dela. Movida para `~/Archives/G-ART-WHISPERX-backup` (fora do
workspace git), copiada e verificada (`diff -rq`) antes de remover o
original. `WHISPERX/` também saiu do `exclude` do ruff em
`code/pyproject.toml` — não faz mais sentido excluir um caminho que não
existe mais dentro de `code/`.
---
## 3. Onde o código ainda é grande (e onde isso não é problema)
Quatro arquivos foram divididos (`writer.py`, `models.py`, `models_api.py`,
`_shared.py`): 6.685 linhas concentradas viraram 43 módulos.
O que sobrou grande, e o diagnóstico honesto de cada um:
| Arquivo | Linhas | Vale dividir? |
|---------|-------:|---------------|
| `fcpxml/text_layout.py` | 901 | **Não.** É diagramação — um assunto coeso. |
| `fcpxml/rough_cut.py` | 798 | **Não.** É geração de timeline, um assunto. |
| `fcpxml/model_manager.py` | 748 | Talvez: mistura catálogo, download e config. |
| `server_tools/voice.py` | 754 | Talvez, se crescer mais. |
| `MacApp/TranscriptionView.swift` | 843 | Sim, quando for mexer nela. |
| `MacApp/WizardView.swift` | 808 | Sim: sete etapas num `switch` só. |
**Critério, não número:** divida quando o arquivo tiver **assuntos** que não se
falam. Um arquivo grande de um assunto só é mais fácil de ler que seis arquivos
pequenos que você precisa abrir juntos. Código picado sem motivo atrapalha tanto
quanto arquivo gigante.
---
## 4. Checklist antes de dar algo por pronto
```bash
cd code && ./Engine/run_after_fix.sh # lint zerado + 1.498 testes
admin/run_app.command # se mexeu no app (padrão de revisão)
admin/run.command # app + atualização incremental da RAG
rag/search_gart.sh "consulta" # busca híbrida no índice RAG (ver rag/README.md)
```
E, além do script:
- [ ] **Mexeu na interface? Abriu a tela?** Compilar não prova que roda —
`VideoPlayer` compilava e abortava (`05_EXPERIENCIAS.md` #22).
- [ ] **Mexeu em XML? Importou no FCP?** DTD válido ≠ renderiza certo.
- [ ] **Dividiu ou moveu módulo?** Procure `patch('<módulo>.` e imports
relativos dentro de funções — é o que quebra em silêncio (#23).
- [ ] **Criou teste fora de `code/tests/`?** Confirme que a contagem total
subiu. Teste fora de `testpaths` não roda e dá falsa sensação de rede (#24).
- [ ] **Problema estrutural ou erro recorrente?** Registre em
`05_EXPERIENCIAS.md` com o índice atualizado.
- [ ] **Documentação divergiu?** Corrija no mesmo commit. Doc velha engana mais
que doc ausente.
---
## 5. Mapa de sintomas → onde olhar
| Sintoma | Suspeite de | Arquivo |
|---------|-------------|---------|
| FCP recusa o arquivo ao importar | `id` inválido, ordem de filhos, timebase | `writer/validation.py`, `dtd.py` |
| Título importa mas não aparece | Template/uid Motion que não resolve | `writer/titles.py` |
| Corte no lugar errado | Tempo pós-corte usado como se fosse original | `voice_actions.py` (`shift_after_cuts`) |
| Zoom/marker sumindo perto de um corte | Borda encostando exatamente no `cut` | §2.1 |
| Legenda sobrepondo | Layout ou conteúdo antigo no arquivo | `collision.py`, `text_layout.py` |
| "Ênfase" apontando para palavra à toa | Falta renormalizar após o corte | `refine_voice_timeline` |
| App diz que falta librosa/pyannote | `uv run` com cwd errado | `PythonBridge.swift` (§3 do doc 08) |
| App crasha com `ModuleNotFoundError: server_tools` | `sys.path` de `admin/api/` mal calculado | `05_EXPERIENCIAS.md` #25 |
| Tela do app fecha o programa | Componente de framework que só falha em runtime | `05_EXPERIENCIAS.md` #22 |
| Comando existe no MCP mas não no app | Falta expor na ponte | `admin/api/`, #20 |
---
## 6. Convenções que não são negociáveis
Estão em `01_ARCHITECTURE.md` §2 e valem repetir as três que mais custaram:
1. **Tempo é fração racional.** Float para tempo produz drift que só aparece
depois de dez operações encadeadas.
2. **Ação de voz é sempre em tempo da mídia original.** Nunca pós-corte.
3. **Original nunca é sobrescrito.** Toda saída ganha sufixo.
---
## Documentos relacionados
- [01 Arquitetura](01_ARCHITECTURE.md) — camadas e onde cada coisa mora
- [02 Módulos](02_MODULES.md) — mapa do engine, módulo a módulo
- [03 Server/Tools](03_SERVER_TOOLS.md) — as 77 ferramentas MCP
- [04 Testes & Workflow](04_TESTS_AND_WORKFLOW.md)
- [05 Experiências](05_EXPERIENCIAS.md) — o que já quebrou e por quê
- [06 Boas Práticas](06_BOAS_PRATICAS.md)
- [08 App macOS](08_APP_MACOS.md) — o app e o Assistente
+269
View File
@@ -0,0 +1,269 @@
# 10 - Mapa de Reestruturacao de Funcionalidades
> Escopo: roteiro pratico para reorganizar o codigo sem quebrar o produto.
> Baseado na varredura de 2026-08-24 sobre engine Python, ponte do app,
> ferramentas MCP e app SwiftUI.
## 1. Diagnostico rapido
O projeto ja tem uma arquitetura-alvo correta: `fcpxml/` como engine puro,
`server.py` + `server_tools/` como camada MCP, `admin/` como ponte JSON-lines
do app e `MacApp/` como interface. A melhoria agora nao e "reinventar" a
arquitetura, e reduzir os pontos onde as responsabilidades ainda se misturam.
### Pontos fortes
- Engine Python bem testado e com regra clara: logica de timeline fica em
`fcpxml/`.
- `writer/` ja foi quebrado em mixins por assunto, preservando API publica.
- `server.py` funciona como composition root e usa dispatch por dicionario.
- Documentacao interna registra decisoes, armadilhas e padroes do projeto.
- Fluxos criticos tem testes extensos em `code/tests/`.
### Dores atuais
- Alguns arquivos voltaram a virar centros de gravidade:
- `server_tools/voice.py` (~999 linhas)
- `server_tools/subtitles.py` (~760 linhas)
- `MacApp/Sources/WizardView.swift` (~979 linhas)
- `MacApp/Sources/TranscriptionView.swift` (~843 linhas)
- `fcpxml/model_manager.py` (~748 linhas)
- `admin/` e `server_tools/` expõem fluxos parecidos por caminhos diferentes,
o que aumenta risco de uma funcionalidade existir no MCP e faltar no app.
- ~~`admin/` ainda fica fora do lint principal~~ — resolvido na Fase 0
(2026-09-22): `admin/` entrou no gate de `run_after_fix.sh`.
- ~~`WHISPERX` e backups aparecem junto da base ativa~~ — resolvido na
Fase 0 (2026-09-22): movido para fora do workspace git.
- O app SwiftUI quase nao tem rede automatizada; compilar nao garante que uma
tela abre.
## 2. Mapa de dominios desejado
```text
Produto
MacApp/ Interface e experiencia do usuario
admin/ Ponte JSON-lines do app
server.py + server_tools/ Entrada MCP
Engine
fcpxml/models/ Dados e contratos
fcpxml/parser.py FCPXML -> objetos
fcpxml/writer/ Escrita e edicao de XML
fcpxml/voice_* Analise e decisoes por voz
fcpxml/text_layout.py Layout de legendas
fcpxml/model_manager.py Catalogo, configs e modelos
Suporte
tests/ Rede automatizada
Engine/docs/ Decisoes e operacao
examples/ Fixtures de uso
Legado / referencia
WHISPERX/ Deve sair do caminho ativo ou virar referencia clara
```
Regra de organizacao: uma funcionalidade nasce no engine, depois ganha duas
portas finas se necessario: uma tool MCP em `server_tools/` e um comando do app
em `admin/api/`.
## 3. Reestruturacao por fases
### Fase 0 - Higiene antes de mexer — `concluída em 2026-09-22`
Objetivo: reduzir ruido e proteger a base antes de mover codigo.
- ~~Decidir o destino de `code/WHISPERX`~~ — não era submodule (sem
`.gitmodules`), era 2,6 GB de backups órfãos do próprio projeto sem
nenhuma referência ativa. Movido para `~/Archives/G-ART-WHISPERX-backup`
(fora do workspace git), copiado com `rsync` e conferido com `diff -rq`
antes de remover o original. Detalhe: essa pasta também inflava qualquer
`ruff check --exclude docs/` manual (a flag sobrescrevia o `exclude` do
`pyproject.toml`, que já ignorava `WHISPERX/`) — ver `05_EXPERIENCIAS.md`
#34.
- ~~Incluir `admin/` em uma checagem de lint separada antes de colocar no
gate obrigatório~~ — checado com a config real do projeto (não o default
do ruff): só 6 erros, todos `E402` por `sys.path.insert` antes de import
local. Resolvido com `# noqa: E402` (convenção já usada no projeto) e
`admin/` entrou direto no gate obrigatório (`run_after_fix.sh`, passo
2/3), sem precisar de etapa intermediária "separada".
- Corrigido de quebra: `fcpxml/writer/adjustment.py` gerava um `<adjustment>`
inválido no DTD — não estava no escopo original da Fase 0, mas surgiu na
investigação e era pequeno o bastante para resolver junto (ver
`05_EXPERIENCIAS.md` #34).
- Atualizados: `02_MODULES.md` (versão, linhas de `writer/`, módulos novos
`builders.py`/`adjustment.py`/`analise.py`/`transcription/`),
`09_MANUTENCAO.md` (contagem de testes/lint, itens §2.3/§2.6/§2.8
resolvidos, novo item §2.7 registrando `test_refine_voice_timeline_tool`).
- **Pendente, não fechado nesta rodada:** "documentar oficialmente quais
pastas são produto ativo, legado e backup" como um documento à parte —
o que existia de fato como "legado" (`WHISPERX`) já foi resolvido, não
sobrou candidato claro para justificar um novo documento agora.
Entrega obtida: lint de `admin/` no gate, `code/writer/adjustment.py`
correto e testado, ~2,6 GB fora do caminho ativo, docs sincronizados com o
código atual.
### Fase 1 - Contratos entre camadas
Objetivo: impedir que MCP, app e engine driftam entre si.
- Criar um registro unico de capacidades, por exemplo:
- nome interno da funcionalidade;
- funcao pura do engine;
- handler MCP, se existir;
- comando `admin`, se existir;
- tela Swift, se existir;
- testes associados.
- Adicionar teste que detecta comandos importantes presentes no MCP mas ausentes
na ponte do app, quando fizer sentido.
- Padronizar o retorno dos comandos `admin/api`: `ok`, `path`, `message`,
`error`, `unchanged`, `artifacts`.
Entrega esperada: mapa vivo de funcionalidades e menos "funciona no Claude,
nao aparece no app".
### Fase 2 - Dividir `server_tools/voice.py`
Objetivo: separar o fluxo de voz por etapas reais do produto.
Divisao sugerida:
```text
server_tools/voice/
__init__.py Reexporta TOOLS e HANDLERS
analysis.py analyze_voice_features, build_voice_timeline
speakers.py diarize_media, remove_speakers
refinement.py refine_voice_timeline, remove_speech_gaps
actions.py apply_voice_actions
local_ai.py generate_voice_script
config.py get/save_voice_analysis_config
```
Cuidados:
- Manter os nomes publicos reexportados para nao quebrar testes/imports.
- Mover em uma etapa por arquivo, rodando testes de voz a cada passo.
- Nao mover regra de negocio para `server_tools/voice/`; se aparecer regra
nova, ela deve descer para `fcpxml/voice_*`.
Testes minimos: `test_voice_actions.py`, `test_voice_actions_tool.py`,
`test_voice_timeline.py`, `test_voice_timeline_tool.py`, `test_diarize.py`,
`test_voice_features.py`.
### Fase 3 - Separar `fcpxml/model_manager.py`
Objetivo: reduzir mistura entre catalogo, download, configuracao e estado.
Divisao sugerida:
```text
fcpxml/model_manager/
__init__.py API publica atual
catalog.py models.json, recomendados, metadata
storage.py diretorios, instalados, migracao
download.py download/cancel/progresso
transcription_config.py modelo selecionado, idioma
voice_config.py analise de voz, silencio, legendas
```
Cuidados:
- Preservar imports atuais via `__init__.py`.
- Separar funcoes puras de funcoes com I/O para facilitar teste.
- Nao acoplar config do app a nomes de tela Swift.
Testes minimos: `test_models.py`, `test_models_api.py` se existir,
`test_voice_analysis_config.py`, `test_project_config.py`.
### Fase 4 - Reorganizar o Assistente SwiftUI
Objetivo: tornar o fluxo de 7 etapas legivel e testavel por partes.
Divisao sugerida:
```text
MacApp/Sources/Wizard/
WizardView.swift Casca, navegacao e estado global
WizardState.swift Estado do fluxo e canAdvance
ProjectStepView.swift
TranscribeStepView.swift
VoiceAnalysisStepView.swift
AIScriptStepView.swift
ReviewStepHost.swift
ProcessStepView.swift
DoneStepView.swift
```
Boas praticas para essa fase:
- Extrair primeiro views pequenas, sem alterar comportamento.
- Depois extrair calculos puros de `canAdvance`, nomes de arquivos e selecao
de artefatos para tipos testaveis.
- Usar harness manual documentado em `08_APP_MACOS.md` para abrir as telas
tocadas.
Entrega esperada: cada etapa do wizard vira um arquivo com responsabilidade
unica.
### Fase 5 - Unificar validacao e saida da ponte `admin/`
Objetivo: deixar os comandos do app tao disciplinados quanto os handlers MCP.
- Criar helpers de path/output equivalentes aos de `server_tools/_shared`,
ou mover helpers comuns para uma camada compartilhada que nao saiba de MCP.
- Trocar chamadas diretas a `server.generate_output_path` por helper de dominio
que nao puxe `server.py` quando a ponte so precisa de path.
- Adicionar lint de `admin/` ao fluxo de manutencao depois de corrigir erros
existentes.
Entrega esperada: ponte mais fina, menos import acidental de transporte MCP.
### Fase 6 - Tests e gates de seguranca
Objetivo: fazer a reorganizacao ser barata de continuar.
- Criar testes de "arquitetura":
- `fcpxml/` nao importa `server`, `server_tools` nem `admin`;
- handlers MCP sempre retornam via `_text_result`;
- comandos `admin` retornam JSON no formato padrao.
- Criar teste de import publico para garantir que reexports antigos continuam.
- Para SwiftUI, manter harnesses por tela critica ate existir um build mais
estruturado.
Entrega esperada: mover arquivos deixa de ser aposta.
## 4. Prioridade recomendada
1. Fase 0: limpar mapa ativo vs legado.
2. Fase 1: criar registro de capacidades.
3. Fase 2: dividir voz em `server_tools`.
4. Fase 5: fortalecer `admin/`.
5. Fase 4: quebrar `WizardView`.
6. Fase 3: dividir `model_manager.py`.
7. Fase 6: ampliar gates conforme as fases estabilizam.
Motivo: primeiro se reduz incerteza, depois se separa o arquivo que mais muda
no fluxo novo de voz, e so entao se mexe nas telas maiores.
## 5. Checklist para cada refatoracao
- Mover sem mudar comportamento na primeira passada.
- Preservar API publica com reexports.
- Rodar testes focados depois de cada movimento.
- Rodar `cd code && ./Engine/run_after_fix.sh` antes de concluir.
- Se mexeu em `MacApp/`, compilar e abrir a tela afetada.
- Atualizar docs no mesmo commit.
- Registrar aprendizado em `05_EXPERIENCIAS.md` quando houver bug real.
## 6. Principios de boas praticas para este projeto
- Engine puro: sem MCP, sem Swift, sem JSON de tela.
- Camadas de entrada finas: validam, chamam engine, formatam resposta.
- Tempo de timeline sempre racional (`TimeValue`), exceto metricas de audio e
UI onde segundos float sao apenas apresentacao/analise.
- Original nunca e sobrescrito.
- XML sempre entra por `safe_xml.py`.
- Dependencias opcionais continuam lazy.
- Arquivo grande so e problema quando contem varios assuntos.
- Toda funcionalidade importante deve ter dono, porta MCP/app documentada e
teste correspondente.
+8 -3
View File
@@ -23,14 +23,19 @@ echo "==> [G-ART] Validação pós-correção iniciada..."
echo " Diretório: $REPO_ROOT" echo " Diretório: $REPO_ROOT"
echo "" echo ""
echo "==> 1/2 Lint (ruff) — deve passar com ZERO erros" echo "==> 1/3 Lint do engine (ruff, code/) — deve passar com ZERO erros"
# A flag --exclude sobrescreve o exclude declarado em pyproject.toml # A flag --exclude sobrescreve o exclude declarado em pyproject.toml
# (que já ignora docs/ e WHISPERX/). Rode sem flag para herdar a config. # (que já ignora docs/). Rode sem flag para herdar a config.
uv run ruff check . uv run ruff check .
echo " Lint OK ✓" echo " Lint OK ✓"
echo "" echo ""
echo "==> 2/2 Testes (pytest) — todos devem passar" echo "==> 2/3 Lint da ponte (ruff, admin/) — mesma config do engine"
uv run ruff check --config pyproject.toml ../admin/
echo " Lint OK ✓"
echo ""
echo "==> 3/3 Testes (pytest) — todos devem passar"
uv run pytest tests/ -v uv run pytest tests/ -v
echo "" echo ""
+12 -4
View File
@@ -12,6 +12,7 @@ struct GArtApp: App {
} }
enum ActiveTab: Hashable { enum ActiveTab: Hashable {
case wizard
case project case project
case captions case captions
case voiceAnalysis case voiceAnalysis
@@ -20,19 +21,23 @@ enum ActiveTab: Hashable {
} }
struct ContentView: View { struct ContentView: View {
@State private var activeTab: ActiveTab? = .project @State private var activeTab: ActiveTab? = .wizard
var body: some View { var body: some View {
NavigationSplitView { NavigationSplitView {
List(selection: $activeTab) { List(selection: $activeTab) {
Label("Assistente", systemImage: "wand.and.stars")
.tag(ActiveTab.wizard)
Section("Avançado") {
Label("Projeto", systemImage: "film") Label("Projeto", systemImage: "film")
.tag(ActiveTab.project) .tag(ActiveTab.project)
Label("Legendas Dinâmicas", systemImage: "captions.bubble") Label("Legendas", systemImage: "captions.bubble")
.tag(ActiveTab.captions) .tag(ActiveTab.captions)
Label("Análise de Voz", systemImage: "waveform") Label("Análise de Voz", systemImage: "waveform")
.tag(ActiveTab.voiceAnalysis) .tag(ActiveTab.voiceAnalysis)
Label("Modelos", systemImage: "tray.and.arrow.down") Label("Modelos", systemImage: "tray.and.arrow.down")
.tag(ActiveTab.models) .tag(ActiveTab.models)
}
Label("Sobre", systemImage: "info.circle") Label("Sobre", systemImage: "info.circle")
.tag(ActiveTab.about) .tag(ActiveTab.about)
} }
@@ -40,19 +45,22 @@ struct ContentView: View {
.navigationSplitViewColumnWidth(min: 180, ideal: 200) .navigationSplitViewColumnWidth(min: 180, ideal: 200)
} detail: { } detail: {
switch activeTab { switch activeTab {
case .wizard, nil:
WizardView().id(UUID())
.navigationTitle("Assistente")
case .project: case .project:
ProjectView().id(UUID()) ProjectView().id(UUID())
.navigationTitle("Projeto") .navigationTitle("Projeto")
case .captions: case .captions:
CaptionsView().id(UUID()) CaptionsView().id(UUID())
.navigationTitle("Legendas Dinâmicas") .navigationTitle("Legendas")
case .voiceAnalysis: case .voiceAnalysis:
VoiceAnalysisView().id(UUID()) VoiceAnalysisView().id(UUID())
.navigationTitle("Análise de Voz") .navigationTitle("Análise de Voz")
case .models: case .models:
ModelDownloadView().id(UUID()) ModelDownloadView().id(UUID())
.navigationTitle("Modelos") .navigationTitle("Modelos")
case .about, nil: case .about:
AboutView() AboutView()
.navigationTitle("Sobre") .navigationTitle("Sobre")
} }
+149 -10
View File
@@ -19,6 +19,7 @@ import UniformTypeIdentifiers
/// assunto. /// assunto.
struct CaptionsView: View { struct CaptionsView: View {
@State private var config = CaptionStyleConfig.defaults @State private var config = CaptionStyleConfig.defaults
@State private var plainConfig = PlainSubtitleConfig.defaults
@State private var isLoading = true @State private var isLoading = true
@State private var errorMessage: String? @State private var errorMessage: String?
@@ -30,15 +31,31 @@ struct CaptionsView: View {
@AppStorage("capSampleAfter") private var sampleAfter = "sua legenda" @AppStorage("capSampleAfter") private var sampleAfter = "sua legenda"
@AppStorage("capShowGuides") private var showsGuides = true @AppStorage("capShowGuides") private var showsGuides = true
private let fontChoices = [ /// Todas as famílias de fonte instaladas no macOS (sistema + usuário), as
"Helvetica Neue", "Helvetica", "Arial", "Avenir Next", /// usadas por padrão primeiro, para o seletor listar tudo sem hardcode.
"Futura", "SF Pro Display", "Georgia", "Impact", private static let installedFontFamilies: [String] = {
] var families = NSFontManager.shared.availableFontFamilies
.sorted { $0.localizedCaseInsensitiveCompare($1) == .orderedAscending }
let preferred = ["Helvetica Neue", "Playfair Display", "Georgia", "Didot"]
for family in preferred.reversed() {
if let idx = families.firstIndex(of: family) {
families.remove(at: idx)
families.insert(family, at: 0)
}
}
return families
}()
private let emphasisFontChoices = [ /// Lista para um picker: todas as famílias instaladas e, se o valor salvo
"Playfair Display", "Georgia", "Didot", "Futura", /// não estiver entre elas (ex.: fonte de outro Mac), ele entra no topo
"Avenir Next", "Times New Roman", "Helvetica Neue", "Impact", /// para o seletor continuar exibindo a escolha atual.
] private func fontChoices(for current: String) -> [String] {
var list = Self.installedFontFamilies
if !list.contains(current) {
list.insert(current, at: 0)
}
return list
}
private let emphasisFaceChoices = [ private let emphasisFaceChoices = [
"Medium Italic", "Italic", "Bold Italic", "Bold", "Regular", "Light Italic", "Medium Italic", "Italic", "Bold Italic", "Bold", "Regular", "Light Italic",
@@ -53,6 +70,13 @@ struct CaptionsView: View {
) )
} }
private func plainBound<T>(_ keyPath: WritableKeyPath<PlainSubtitleConfig, T>) -> Binding<T> {
Binding(
get: { plainConfig[keyPath: keyPath] },
set: { plainConfig[keyPath: keyPath] = $0; savePlain() }
)
}
private func colorBound(_ keyPath: WritableKeyPath<CaptionStyleConfig, String>) -> Binding<Color> { private func colorBound(_ keyPath: WritableKeyPath<CaptionStyleConfig, String>) -> Binding<Color> {
Binding( Binding(
get: { Color(rgbaString: config[keyPath: keyPath]) }, get: { Color(rgbaString: config[keyPath: keyPath]) },
@@ -60,6 +84,13 @@ struct CaptionsView: View {
) )
} }
private func plainColorBound(_ keyPath: WritableKeyPath<PlainSubtitleConfig, String>) -> Binding<Color> {
Binding(
get: { Color(rgbaString: plainConfig[keyPath: keyPath]) },
set: { plainConfig[keyPath: keyPath] = $0.fcpxmlColorString; savePlain() }
)
}
var body: some View { var body: some View {
HSplitView { HSplitView {
previewColumn previewColumn
@@ -152,6 +183,7 @@ struct CaptionsView: View {
positionSection positionSection
bodySection bodySection
emphasisSection emphasisSection
plainSubtitleSection
calibrationSection calibrationSection
} }
if let errorMessage { if let errorMessage {
@@ -195,7 +227,7 @@ struct CaptionsView: View {
private var bodySection: some View { private var bodySection: some View {
Section("Linhas de apoio") { Section("Linhas de apoio") {
Picker("Fonte", selection: bound(\.font)) { Picker("Fonte", selection: bound(\.font)) {
ForEach(fontChoices, id: \.self) { Text($0).tag($0) } ForEach(fontChoices(for: config.font), id: \.self) { Text($0).tag($0) }
} }
slider( slider(
"Tamanho", "Tamanho",
@@ -210,7 +242,7 @@ struct CaptionsView: View {
private var emphasisSection: some View { private var emphasisSection: some View {
Section("Palavra de ênfase") { Section("Palavra de ênfase") {
Picker("Fonte", selection: bound(\.emphasisFont)) { Picker("Fonte", selection: bound(\.emphasisFont)) {
ForEach(emphasisFontChoices, id: \.self) { Text($0).tag($0) } ForEach(fontChoices(for: config.emphasisFont), id: \.self) { Text($0).tag($0) }
} }
Picker("Estilo", selection: bound(\.emphasisFace)) { Picker("Estilo", selection: bound(\.emphasisFace)) {
ForEach(emphasisFaceChoices, id: \.self) { Text($0).tag($0) } ForEach(emphasisFaceChoices, id: \.self) { Text($0).tag($0) }
@@ -225,6 +257,35 @@ struct CaptionsView: View {
} }
} }
private var plainSubtitleSection: some View {
Section("Legenda comum") {
Picker("Fonte", selection: plainBound(\.font)) {
ForEach(fontChoices(for: plainConfig.font), id: \.self) { Text($0).tag($0) }
}
slider(
"Tamanho",
value: plainBound(\.fontSize), in: 28...300, step: 1,
readout: "\(Int(plainConfig.fontSize))pt",
help: "Tamanho da legenda comum editável no Final Cut."
)
slider(
"Máximo de palavras",
value: plainBound(\.maxWords), in: 1...14, step: 1,
readout: "\(Int(plainConfig.maxWords))",
help: "Quantidade máxima de palavras por bloco de legenda."
)
slider(
"Altura",
value: plainBound(\.positionY), in: -1200...300, step: 1,
readout: "\(Int(plainConfig.positionY))",
help: "Posição vertical da legenda comum no quadro; valores mais negativos descem."
)
ColorPicker("Cor", selection: plainColorBound(\.fontColor), supportsOpacity: true)
Toggle("Usar letra maiúscula", isOn: plainBound(\.uppercase))
Toggle("Manter vírgula e ponto", isOn: plainBound(\.keepPunctuation))
}
}
private var calibrationSection: some View { private var calibrationSection: some View {
Section { Section {
slider( slider(
@@ -305,18 +366,33 @@ struct CaptionsView: View {
} else if let error { } else if let error {
errorMessage = error errorMessage = error
} }
PythonBridge.call(command: "plain_subtitle_config") { plainResult, plainError in
DispatchQueue.main.async {
if let plainResult {
plainConfig = PlainSubtitleConfig(from: plainResult)
} else if let plainError {
errorMessage = plainError
}
isLoading = false isLoading = false
continuation.resume() continuation.resume()
} }
} }
} }
} }
}
}
private func save() { private func save() {
PythonBridge.call(command: "set_dynamic_subtitle_config", arguments: config.arguments()) { _, error in PythonBridge.call(command: "set_dynamic_subtitle_config", arguments: config.arguments()) { _, error in
DispatchQueue.main.async { errorMessage = error } DispatchQueue.main.async { errorMessage = error }
} }
} }
private func savePlain() {
PythonBridge.call(command: "set_plain_subtitle_config", arguments: plainConfig.arguments()) { _, error in
DispatchQueue.main.async { errorMessage = error }
}
}
} }
/// O estilo das legendas dinâmicas, no formato que a tela edita e o bridge /// O estilo das legendas dinâmicas, no formato que a tela edita e o bridge
@@ -402,6 +478,69 @@ struct CaptionStyleConfig {
} }
} }
struct PlainSubtitleConfig {
var font: String
var fontSize: Double
var fontColor: String
var maxWords: Double
var positionY: Double
var uppercase: Bool
var keepPunctuation: Bool
var textScale: Double
static let defaults = PlainSubtitleConfig(
font: "Helvetica Neue",
fontSize: 82,
fontColor: "1 1 1 1",
maxWords: 7,
positionY: -820,
uppercase: false,
keepPunctuation: true,
textScale: 2.0
)
init(from json: [String: Any]) {
let d = PlainSubtitleConfig.defaults
self.init(
font: json["font"] as? String ?? d.font,
fontSize: (json["font_size"] as? NSNumber)?.doubleValue ?? d.fontSize,
fontColor: json["font_color"] as? String ?? d.fontColor,
maxWords: (json["max_words"] as? NSNumber)?.doubleValue ?? d.maxWords,
positionY: (json["position_y"] as? NSNumber)?.doubleValue ?? d.positionY,
uppercase: json["uppercase"] as? Bool ?? d.uppercase,
keepPunctuation: json["keep_punctuation"] as? Bool ?? d.keepPunctuation,
textScale: (json["text_scale"] as? NSNumber)?.doubleValue ?? d.textScale
)
}
init(
font: String, fontSize: Double, fontColor: String, maxWords: Double,
positionY: Double, uppercase: Bool, keepPunctuation: Bool, textScale: Double
) {
self.font = font
self.fontSize = fontSize
self.fontColor = fontColor
self.maxWords = maxWords
self.positionY = positionY
self.uppercase = uppercase
self.keepPunctuation = keepPunctuation
self.textScale = textScale
}
func arguments() -> [String: Any] {
[
"font": font,
"font_size": Int(fontSize),
"font_color": fontColor,
"max_words": Int(maxWords),
"position_y": positionY,
"uppercase": uppercase,
"keep_punctuation": keepPunctuation,
"text_scale": textScale,
]
}
}
extension Color { extension Color {
/// Parses an FCPXML "R G B A" space-separated 0-1 string into a Color. /// Parses an FCPXML "R G B A" space-separated 0-1 string into a Color.
init(rgbaString: String) { init(rgbaString: String) {
+99 -1
View File
@@ -15,6 +15,11 @@ struct ModelDownloadView: View {
@State private var hfTokenText: String = "" @State private var hfTokenText: String = ""
@State private var numSpeakersText: String = "" @State private var numSpeakersText: String = ""
@State private var language: String = "auto" @State private var language: String = "auto"
@State private var acousticsAvailable: Bool?
@State private var acousticsMessage: String = ""
@State private var isInstallingAcoustics = false
@State private var acousticsInstallLog: String = ""
@State private var acousticsInstallError: String?
private let languages: [(String, String)] = [ private let languages: [(String, String)] = [
("auto", "Detectar automaticamente"), ("auto", "Detectar automaticamente"),
@@ -33,6 +38,7 @@ struct ModelDownloadView: View {
var body: some View { var body: some View {
Form { Form {
storageSection storageSection
acousticsSection
diarizationSection diarizationSection
if let errorMessage { if let errorMessage {
Section { Section {
@@ -65,7 +71,7 @@ struct ModelDownloadView: View {
} }
} }
.formStyle(.grouped) .formStyle(.grouped)
.task { await refresh() } .task { await refresh(); checkAcoustics() }
} }
// MARK: - Transcription language // MARK: - Transcription language
@@ -95,6 +101,98 @@ struct ModelDownloadView: View {
PythonBridge.call(command: "set_language", arguments: ["language": code]) { _, _ in } PythonBridge.call(command: "set_language", arguments: ["language": code]) { _, _ in }
} }
// MARK: - Acoustic analysis (librosa)
/// A ênfase de voz (pitch/energia) precisa do `librosa`, que é uma
/// dependência opcional — sem ela `layers.acoustics` vem `false` na
/// análise e a decisão de zoom fica sem base real. Antes disso só dava
/// pra descobrir lendo o JSON exportado; agora o app já diz e resolve.
private var acousticsSection: some View {
Section {
VStack(alignment: .leading, spacing: 10) {
if let acousticsAvailable {
Label(
acousticsMessage.isEmpty
? (acousticsAvailable ? "Disponível" : "Indisponível")
: acousticsMessage,
systemImage: acousticsAvailable ? "checkmark.circle.fill" : "exclamationmark.triangle.fill"
)
.font(.caption)
.foregroundStyle(acousticsAvailable ? Color.green : Color.orange)
} else {
Label("Verificando…", systemImage: "hourglass")
.font(.caption).foregroundStyle(.secondary)
}
if acousticsAvailable == false {
Button {
installAcoustics()
} label: {
if isInstallingAcoustics {
HStack { ProgressView().controlSize(.small); Text("Instalando…") }
} else {
Label("Instalar (uv sync --all-extras)", systemImage: "arrow.down.circle")
}
}
.disabled(isInstallingAcoustics)
if !acousticsInstallLog.isEmpty {
ScrollView {
Text(acousticsInstallLog)
.font(.system(.caption2, design: .monospaced))
.foregroundStyle(.secondary)
.frame(maxWidth: .infinity, alignment: .leading)
}
.frame(height: 90)
.background(RoundedRectangle(cornerRadius: 6).fill(Color.secondary.opacity(0.06)))
}
if let acousticsInstallError {
Label(acousticsInstallError, systemImage: "xmark.circle.fill")
.font(.caption).foregroundStyle(.red)
}
}
}
} header: {
Text("Análise Acústica (zoom por voz)")
} footer: {
Text("Mede a energia e o tom de voz de verdade, para os candidatos a zoom da edição por voz. Sem isso, a análise ainda transcreve e decide cortes pelo texto — só o zoom fica sem base acústica.")
.font(.caption)
.foregroundStyle(.secondary)
}
}
private func checkAcoustics() {
PythonBridge.call(command: "acoustics_capability") { result, err in
DispatchQueue.main.async {
guard let result, result["ok"] as? Bool == true else { return }
acousticsAvailable = result["available"] as? Bool
acousticsMessage = result["message"] as? String ?? ""
}
}
}
private func installAcoustics() {
isInstallingAcoustics = true
acousticsInstallLog = ""
acousticsInstallError = nil
// --all-extras, não só "intelligence": `uv sync` substitui o
// ambiente pelos extras pedidos em vez de somar, então um sync
// parcial aqui derrubaria dev/transcribe/diarização já instalados.
PythonBridge.runUV(arguments: ["sync", "--all-extras"]) { line in
DispatchQueue.main.async {
acousticsInstallLog += (acousticsInstallLog.isEmpty ? "" : "\n") + line
}
} completion: { code, err in
DispatchQueue.main.async {
isInstallingAcoustics = false
if code != 0 {
acousticsInstallError = err ?? "Falha ao instalar."
}
checkAcoustics()
}
}
}
// MARK: - Diarization // MARK: - Diarization
private var diarizationSection: some View { private var diarizationSection: some View {
+139
View File
@@ -116,6 +116,145 @@ struct ZoomClip: Identifiable {
} }
} }
/// One word inside a phrase, with the acoustics that justify an emphasis.
struct ReviewWord: Identifiable {
let id: Int
let text: String
let start: Double
let end: Double
let energy: Double
let emphasis: Double
init(id: Int, json: [String: Any]) {
self.id = id
text = json["text"] as? String ?? ""
start = json["start"] as? Double ?? 0
end = json["end"] as? Double ?? 0
energy = json["energy"] as? Double ?? 0
emphasis = json["emphasis"] as? Double ?? 0
}
}
/// A phrase in the review step — one spoken line plus the decision made about
/// it. Mirrors `fcpxml/phrase_review.py`; `emphasis` is 0–3 and everything
/// mutable here is what the editor is allowed to change.
struct ReviewPhrase: Identifiable {
let id: Int
let start: Double
let end: Double
var trimStart: Double
var trimEnd: Double
var text: String
let speaker: String
var active: Bool
var emphasis: Int
var track: String
let peakEmphasis: Double
let emotion: String
let emotionConfidence: Double
let takeBoundary: Bool
let gapBefore: Double
let reason: String
let words: [ReviewWord]
static let trackScript = "roteiro"
static let trackBackstage = "bastidor"
/// Delivery emotion as the analysis names it, in the user's language plus a
/// glyph — the label alone is too easy to skim past in a dense list.
static func emotionLabel(_ emotion: String) -> (String, String) {
switch emotion {
case "excited": return ("Empolgado", "flame")
case "tense": return ("Tenso", "bolt")
case "calm": return ("Calmo", "leaf")
case "reflective": return ("Reflexivo", "moon")
default: return ("Neutro", "circle")
}
}
init(json: [String: Any]) {
id = json["index"] as? Int ?? 0
start = json["start"] as? Double ?? 0
end = json["end"] as? Double ?? 0
trimStart = json["trim_start"] as? Double ?? (json["start"] as? Double ?? 0)
trimEnd = json["trim_end"] as? Double ?? (json["end"] as? Double ?? 0)
text = json["text"] as? String ?? ""
speaker = json["speaker"] as? String ?? ""
active = json["active"] as? Bool ?? true
emphasis = json["emphasis"] as? Int ?? 0
track = json["track"] as? String ?? ReviewPhrase.trackScript
peakEmphasis = json["peak_emphasis"] as? Double ?? 0
emotion = json["emotion"] as? String ?? "neutral"
emotionConfidence = json["emotion_confidence"] as? Double ?? 0
takeBoundary = json["take_boundary"] as? Bool ?? false
gapBefore = json["gap_before"] as? Double ?? 0
reason = json["reason"] as? String ?? ""
words = (json["words"] as? [[String: Any]] ?? [])
.enumerated().map { ReviewWord(id: $0.offset, json: $0.element) }
}
var asJSON: [String: Any] {
[
"index": id,
"start": start,
"end": end,
"trim_start": trimStart,
"trim_end": trimEnd,
"text": text,
"speaker": speaker,
"active": active,
"emphasis": emphasis,
"track": track,
"reason": reason,
]
}
var isBackstage: Bool { track == ReviewPhrase.trackBackstage }
var isTrimmed: Bool { trimStart > start + 0.001 || trimEnd < end - 0.001 }
var timecode: String {
String(format: "%02d:%02d", Int(start) / 60, Int(start) % 60)
}
/// The word boundaries a trim handle is allowed to land on.
func snap(_ time: Double, edge: TrimEdge) -> Double {
let boundaries = words.map { edge == .start ? $0.start : $0.end }.filter { $0 > 0 }
guard let nearest = boundaries.min(by: { abs($0 - time) < abs($1 - time) }) else {
return time
}
return nearest
}
}
enum TrimEdge { case start, end }
/// A punch-in the editor placed by hand over an arbitrary range, next to the
/// whole-phrase zoom that an emphasis level produces. It stores only *when* —
/// the scale and the ramp come from the Voice Analysis settings at render time.
struct ManualZoom: Identifiable {
let id = UUID()
var start: Double
var end: Double
/// Below this a punch-in has no room to ramp in and back out; the writer
/// rejects the window, so offering it would place nothing.
static let minimumDuration: Double = 0.4
init(start: Double, end: Double) {
self.start = start
self.end = end
}
init?(json: [String: Any]) {
guard let start = json["start"] as? Double, let end = json["end"] as? Double,
end - start >= ManualZoom.minimumDuration
else { return nil }
self.start = start
self.end = end
}
var asJSON: [String: Any] { ["start": start, "end": end] }
}
struct ZoomSegment: Identifiable { struct ZoomSegment: Identifiable {
let id: Int let id: Int
let start: Double let start: Double
+442
View File
@@ -0,0 +1,442 @@
import AVFoundation
import Combine
import Foundation
/// State behind the wizard's emphasis-review step.
///
/// Holds the phrases, the selection, and the player — together, because they
/// are one thing to the user: clicking a phrase moves the playhead, playing
/// moves the selection, and skipping a removed line only works if whoever owns
/// playback also knows which lines are removed.
///
/// The preview deliberately plays the *original* media and jumps over whatever
/// the edit removes, instead of rendering a cut first. Rendering to check a
/// toggle would put minutes between a decision and its result; jumping gives
/// the same reading instantly, and the real cut is generated later from the
/// exact same phrase list.
@MainActor
final class PhraseReviewModel: ObservableObject {
@Published var phrases: [ReviewPhrase] = []
@Published var selection: Int?
@Published var isLoading = false
@Published var errorMessage: String?
@Published var currentTime: Double = 0
@Published var isPlaying = false
@Published var pixelsPerSecond: Double = 40
@Published var skipRemoved = true
@Published var zooms: [ManualZoom] = []
/// In/out the editor dragged on the timeline, in source seconds.
@Published var rangeStart: Double?
@Published var rangeEnd: Double?
private(set) var source = ""
private(set) var sourcePath = ""
private(set) var duration: Double = 0
private(set) var speakers: [String] = []
private(set) var emotionAvailable = false
private(set) var player: AVPlayer?
private var voiceTimelinePath = ""
private var timeObserver: Any?
private var playbackLimit: Double?
let minPixelsPerSecond: Double = 8
let maxPixelsPerSecond: Double = 400
deinit {
if let timeObserver, let player {
player.removeTimeObserver(timeObserver)
}
}
// MARK: - Carregar
/// Builds the review from the voice timeline plus whatever the AI decided.
/// A review saved on a previous visit wins — see `cmd_build_phrase_review` —
/// UNLESS `fresh` is true, in which case that saved review is ignored and
/// `active`/`emphasis`/etc. come straight from this call's `decisionsJSON`.
/// Pass `fresh: true` when the decisions themselves changed since the
/// review was last built (the caller re-pasted/regenerated the AI's JSON
/// and re-ran `apply_voice_actions`) — otherwise the saved review from the
/// PREVIOUS decisions silently wins over the fresh cut it should reflect,
/// which is exactly the desync the wizard's "active" toggle showed against
/// the just-reapplied FCPXML.
func load(voiceTimelinePath: String, decisionsJSON: String,
outputFolder: String? = nil, mediaFolder: String? = nil, fresh: Bool = false) {
self.voiceTimelinePath = voiceTimelinePath
isLoading = true
errorMessage = nil
var arguments: [String: Any] = ["voice_timeline": voiceTimelinePath]
if let outputFolder { arguments["output_dir"] = outputFolder }
if let mediaFolder { arguments["media_dir"] = mediaFolder }
if fresh { arguments["fresh"] = true }
if let data = decisionsJSON.data(using: .utf8),
let parsed = try? JSONSerialization.jsonObject(with: data) {
arguments["actions"] = parsed
}
PythonBridge.call(command: "build_phrase_review", arguments: arguments) { [weak self] result, error in
Task { @MainActor in
guard let self else { return }
self.isLoading = false
if let error {
self.errorMessage = error
return
}
guard let result, result["ok"] as? Bool == true else {
self.errorMessage = result?["error"] as? String ?? "Não foi possível montar a revisão."
return
}
self.apply(result)
}
}
}
private func apply(_ result: [String: Any]) {
source = result["source"] as? String ?? ""
// The timeline JSON stores only the media's file name; the bridge
// resolves it to something openable (see phrase_review.resolve_source).
sourcePath = result["source_path"] as? String ?? ""
duration = result["duration"] as? Double ?? 0
speakers = result["speakers"] as? [String] ?? []
emotionAvailable = result["emotion_available"] as? Bool ?? false
phrases = (result["phrases"] as? [[String: Any]] ?? []).map { ReviewPhrase(json: $0) }
zooms = (result["zooms"] as? [[String: Any]] ?? []).compactMap { ManualZoom(json: $0) }
selection = phrases.first?.id
if let errors = result["errors"] as? [String], !errors.isEmpty {
errorMessage = "A IA mandou \(errors.count) decisão(ões) que não deu para ler — o resto foi aplicado."
}
preparePlayer()
}
/// Point the preview at a media file the user chose by hand — the way out
/// when the footage moved somewhere the automatic lookup can't reach.
func useMedia(at path: String) {
sourcePath = path
preparePlayer()
}
private func preparePlayer() {
guard !sourcePath.isEmpty, FileManager.default.fileExists(atPath: sourcePath) else {
player = nil
return
}
if let timeObserver, let player {
player.removeTimeObserver(timeObserver)
self.timeObserver = nil
}
let asset = AVURLAsset(url: URL(fileURLWithPath: sourcePath))
let player = AVPlayer(playerItem: AVPlayerItem(asset: asset))
self.player = player
// 60 Hz: the same observer drives the playhead *and* decides when to
// jump a removed stretch, so its period is the worst-case amount of cut
// material that can be heard before the skip lands. At 20 Hz that was an
// audible blip on every join.
let interval = CMTime(seconds: 1.0 / 60.0, preferredTimescale: 600)
timeObserver = player.addPeriodicTimeObserver(forInterval: interval, queue: .main) { [weak self] time in
Task { @MainActor in
self?.tick(time.seconds)
}
}
}
// MARK: - Reprodução
private func tick(_ time: Double) {
currentTime = time
guard isPlaying else { return }
// Playing a single phrase or a marked range stops at its out point
// instead of running on into the rest of the take.
if let limit = playbackLimit, time >= limit {
pause()
seek(to: limit)
return
}
if skipRemoved, let jump = nextKeptTime(after: time), jump > time {
seek(to: jump)
}
if let phrase = phrase(at: time), selection != phrase.id {
selection = phrase.id
}
}
/// Where playback should resume when `time` lands on removed material.
/// Returns nil when the time is on material that survives.
func nextKeptTime(after time: Double) -> Double? {
for phrase in phrases where time >= phrase.start - 0.001 && time < phrase.end {
if !phrase.active { return phrase.end }
if time < phrase.trimStart { return phrase.trimStart }
if time >= phrase.trimEnd { return phrase.end }
return nil
}
return nil
}
func togglePlay() {
if isPlaying {
pause()
} else {
playbackLimit = nil
play()
}
}
private func play() {
guard let player else { return }
if skipRemoved, let jump = nextKeptTime(after: currentTime) { seek(to: jump) }
player.play()
isPlaying = true
}
func pause() {
player?.pause()
isPlaying = false
playbackLimit = nil
}
/// Play exactly one span and stop — how a cut is judged: in context, at
/// speed, without hunting for the out point by hand.
func playRange(from start: Double, to end: Double) {
guard end > start else { return }
seek(to: start)
playbackLimit = end
player?.play()
isPlaying = true
}
func playSelectedPhrase() {
guard let selection, let phrase = phrases.first(where: { $0.id == selection })
else { return }
playRange(from: phrase.active ? phrase.trimStart : phrase.start,
to: phrase.active ? phrase.trimEnd : phrase.end)
}
func seek(to time: Double) {
currentTime = max(0, time)
player?.seek(to: CMTime(seconds: max(0, time), preferredTimescale: 600),
toleranceBefore: .zero, toleranceAfter: .zero)
}
/// Move the playhead to a phrase and select it.
func goTo(phraseID: Int) {
guard let phrase = phrases.first(where: { $0.id == phraseID }) else { return }
selection = phraseID
seek(to: phrase.active ? phrase.trimStart : phrase.start)
}
func phrase(at time: Double) -> ReviewPhrase? {
phrases.first { time >= $0.start && time < $0.end }
}
func selectNeighbour(_ delta: Int) {
guard let selection, let index = phrases.firstIndex(where: { $0.id == selection }) else {
if let first = phrases.first { goTo(phraseID: first.id) }
return
}
let next = min(max(0, index + delta), phrases.count - 1)
goTo(phraseID: phrases[next].id)
}
// MARK: - Edições
private func update(_ id: Int, _ change: (inout ReviewPhrase) -> Void) {
guard let index = phrases.firstIndex(where: { $0.id == id }) else { return }
change(&phrases[index])
}
func setEmphasis(_ level: Int, for id: Int) {
update(id) { $0.emphasis = min(3, max(0, level)) }
}
func toggleActive(_ id: Int) {
update(id) { $0.active.toggle() }
}
func setTrack(_ track: String, for id: Int) {
update(id) { $0.track = track }
}
func setText(_ text: String, for id: Int) {
update(id) { $0.text = text }
}
/// Trim a phrase's head or tail, landing on a word boundary.
/// A trim that would swallow the whole line is refused — deactivating the
/// phrase is the way to remove it, and doing it by accident with a drag
/// would lose the emphasis decision along with the line.
func trim(_ id: Int, edge: TrimEdge, to time: Double) {
update(id) { phrase in
let snapped = phrase.snap(time, edge: edge)
switch edge {
case .start:
let value = min(max(phrase.start, snapped), phrase.trimEnd - 0.1)
if value < phrase.trimEnd { phrase.trimStart = value }
case .end:
let value = max(min(phrase.end, snapped), phrase.trimStart + 0.1)
if value > phrase.trimStart { phrase.trimEnd = value }
}
}
}
func resetTrim(_ id: Int) {
update(id) { $0.trimStart = $0.start; $0.trimEnd = $0.end }
}
/// Trim everything before/after a given word — the text-first way to cut,
/// since the editor reads the line and points at where it should begin.
/// Clicking the word that is ALREADY that edge toggles it back off —
/// the trim on that side resets to the phrase's own start/end — so the
/// same click that sets a boundary also clears it, instead of needing
/// the separate "Inteira" button for a one-sided undo.
func trimToWord(_ word: ReviewWord, edge: TrimEdge, in id: Int) {
guard let phrase = phrases.first(where: { $0.id == id }) else { return }
let epsilon = 0.001
switch edge {
case .start where abs(word.start - phrase.trimStart) < epsilon:
update(id) { $0.trimStart = $0.start }
case .end where abs(word.end - phrase.trimEnd) < epsilon:
update(id) { $0.trimEnd = $0.end }
default:
trim(id, edge: edge, to: edge == .start ? word.start : word.end)
}
}
// MARK: - Trecho marcado e zooms
var hasRange: Bool {
guard let rangeStart, let rangeEnd else { return false }
return rangeEnd - rangeStart >= ManualZoom.minimumDuration
}
var rangeSpan: (start: Double, end: Double)? {
guard let rangeStart, let rangeEnd, rangeEnd > rangeStart else { return nil }
return (rangeStart, rangeEnd)
}
func setRange(from start: Double, to end: Double) {
rangeStart = min(start, end)
rangeEnd = max(start, end)
}
func clearRange() {
rangeStart = nil
rangeEnd = nil
}
/// Add a punch-in over the marked range. Scale and ramp are not stored:
/// they come from the "Análise de Voz" settings when the edit is rendered,
/// so changing the look there restyles every zoom at once.
func addZoomForRange() {
guard let span = rangeSpan, span.end - span.start >= ManualZoom.minimumDuration
else { return }
zooms.append(ManualZoom(start: span.start, end: span.end))
zooms.sort { $0.start < $1.start }
clearRange()
}
func addZoomForPhrase(_ id: Int) {
guard let phrase = phrases.first(where: { $0.id == id }) else { return }
zooms.append(ManualZoom(start: phrase.trimStart, end: phrase.trimEnd))
zooms.sort { $0.start < $1.start }
}
func removeZoom(_ id: UUID) {
zooms.removeAll { $0.id == id }
}
func zoom(at time: Double) -> ManualZoom? {
zooms.first { time >= $0.start && time <= $0.end }
}
func setEmphasisForAll(_ level: Int) {
for index in phrases.indices where phrases[index].active {
phrases[index].emphasis = level
}
}
// MARK: - Resumo e gravação
var emphasisCount: Int { phrases.filter { $0.active && $0.emphasis >= 1 }.count }
var removedCount: Int { phrases.filter { !$0.active }.count }
var keptDuration: Double {
phrases.filter { $0.active }.reduce(0) { $0 + ($1.trimEnd - $1.trimStart) }
}
// MARK: - Tempo compactado (sem os vãos do que foi cortado)
/// Kept spans of source media, in order, each carrying the position it
/// lands at once every removed stretch between phrases is squeezed out.
/// The timeline draws and scrubs in this space so it reads like the cut
/// itself instead of the raw take with holes in it.
private var keptSegments: [(rawStart: Double, rawEnd: Double, compactStart: Double)] {
var offset = 0.0
var segments: [(Double, Double, Double)] = []
for phrase in phrases.sorted(by: { $0.start < $1.start }) where phrase.active {
guard phrase.trimEnd > phrase.trimStart else { continue }
segments.append((phrase.trimStart, phrase.trimEnd, offset))
offset += phrase.trimEnd - phrase.trimStart
}
return segments
}
/// Maps a raw source-media time to its position on the compacted timeline.
/// Time inside removed material collapses to the boundary of the nearest
/// kept segment, so cut stretches take up no space at all.
func compactTime(_ raw: Double) -> Double {
let segments = keptSegments
for segment in segments {
if raw < segment.rawStart { return segment.compactStart }
if raw <= segment.rawEnd { return segment.compactStart + (raw - segment.rawStart) }
}
guard let last = segments.last else { return 0 }
return raw >= last.rawEnd ? last.compactStart + (last.rawEnd - last.rawStart) : 0
}
/// The inverse of `compactTime`: where a click on the compacted timeline
/// lands in the raw source media, for seeking and scrubbing.
func rawTime(fromCompact compact: Double) -> Double {
let segments = keptSegments
for segment in segments {
let compactEnd = segment.compactStart + (segment.rawEnd - segment.rawStart)
if compact <= compactEnd {
return segment.rawStart + max(0, compact - segment.compactStart)
}
}
return segments.last?.rawEnd ?? 0
}
/// Persists the edited review plus the actions derived from it. Called when
/// the wizard advances — the render itself happens in the next step.
/// Persists the edited review and hands back BOTH paths it wrote:
/// `review_path` (the human-readable `_phrase_review.json`) and
/// `actions_path` (`_phrase_actions.json`, the cut/zoom list derived from
/// it — what `finalizeProcessing` needs to actually apply the review's
/// active/inactive decisions instead of just filing them away).
func save(completion: @escaping (_ reviewPath: String?, _ actionsPath: String?) -> Void) {
guard !voiceTimelinePath.isEmpty, !phrases.isEmpty else {
completion(nil, nil)
return
}
let arguments: [String: Any] = [
"voice_timeline": voiceTimelinePath,
"source": source,
"duration": duration,
"speakers": speakers,
"phrases": phrases.map { $0.asJSON },
"zooms": zooms.map { $0.asJSON },
]
PythonBridge.call(command: "save_phrase_review", arguments: arguments) { result, error in
Task { @MainActor in
if let error {
completion(nil, nil)
_ = error
return
}
completion(result?["review_path"] as? String, result?["actions_path"] as? String)
}
}
}
}
+313
View File
@@ -0,0 +1,313 @@
import SwiftUI
/// The wizard's emphasis-review step.
///
/// Every decision here is about a *sentence* read from the original
/// transcription, so the phrases are listed in full — each line shows the text
/// as it will be said, a switch to keep or drop it from the cut, and the
/// emphasis level. Selecting a line in the list also selects its block on the
/// timeline below, and vice-versa.
struct PhraseReviewView: View {
@ObservedObject var model: PhraseReviewModel
var body: some View {
VSplitView {
VStack(spacing: 0) {
inspectorHeader
Divider()
List(selection: $model.selection) {
ForEach($model.phrases) { $phrase in
PhraseRow(phrase: $phrase, model: model)
.tag(phrase.id)
}
}
.listStyle(.inset)
.onChange(of: model.selection) { _, newValue in
if let newValue { model.goTo(phraseID: newValue) }
}
Divider()
summaryBar
}
.frame(minHeight: 240)
TimelineTracksView(model: model)
.frame(minHeight: 190, idealHeight: 210)
}
.overlay { if model.isLoading { loadingOverlay } }
.focusable()
.onKeyPress(.space) { model.togglePlay(); return .handled }
.onKeyPress(.return) { model.playSelectedPhrase(); return .handled }
.onKeyPress(.leftArrow) { model.selectNeighbour(-1); return .handled }
.onKeyPress(.rightArrow) { model.selectNeighbour(1); return .handled }
.onKeyPress(characters: .decimalDigits) { press in
guard let level = Int(press.characters), (0...3).contains(level),
let selection = model.selection else { return .ignored }
model.setEmphasis(level, for: selection)
return .handled
}
}
private var loadingOverlay: some View {
ZStack {
Color(nsColor: .windowBackgroundColor).opacity(0.85)
VStack(spacing: 10) {
ProgressView()
Text("Montando a revisão…").font(.callout).foregroundStyle(.secondary)
}
}
}
private var summaryBar: some View {
HStack(spacing: 16) {
summaryItem("text.quote", "\(model.phrases.count) frases")
summaryItem("sparkles", "\(model.emphasisCount) com ênfase")
summaryItem("scissors", "\(model.removedCount) fora do corte")
summaryItem("clock", durationLabel(model.keptDuration))
if !model.zooms.isEmpty {
summaryItem("plus.magnifyingglass", "\(model.zooms.count) zooms")
}
Spacer()
if let phrase = selectedPhrase, !phrase.reason.isEmpty {
Label(phrase.reason, systemImage: "brain")
.font(.caption).foregroundStyle(.secondary)
.lineLimit(1).truncationMode(.tail)
}
}
.padding(.horizontal, 14)
.padding(.vertical, 8)
}
private func summaryItem(_ icon: String, _ text: String) -> some View {
Label(text, systemImage: icon).font(.caption).foregroundStyle(.secondary)
}
private func durationLabel(_ seconds: Double) -> String {
String(format: "%02d:%02d finais", Int(seconds) / 60, Int(seconds) % 60)
}
private func pickMedia() {
let panel = NSOpenPanel()
panel.canChooseFiles = true
panel.canChooseDirectories = false
panel.allowsMultipleSelection = false
panel.prompt = "Usar esta mídia"
panel.message = model.source.isEmpty
? "Escolha o arquivo de vídeo desta gravação."
: "Escolha onde está \(model.source)."
if panel.runModal() == .OK, let url = panel.url {
model.useMedia(at: url.path)
}
}
private var selectedPhrase: ReviewPhrase? {
guard let selection = model.selection else { return nil }
return model.phrases.first { $0.id == selection }
}
// MARK: - Inspector de frases
private var inspectorPane: some View {
VStack(spacing: 0) {
inspectorHeader
Divider()
List(selection: $model.selection) {
ForEach($model.phrases) { $phrase in
PhraseRow(phrase: $phrase, model: model)
.tag(phrase.id)
}
}
.listStyle(.inset)
.onChange(of: model.selection) { _, newValue in
if let newValue { model.goTo(phraseID: newValue) }
}
}
}
private var inspectorHeader: some View {
VStack(alignment: .leading, spacing: 6) {
Text("Frases").font(.headline)
Text("Só as frases com ênfase recebem zoom e legenda dinâmica. O resto fica com legenda comum.")
.font(.caption).foregroundStyle(.secondary)
if !model.emotionAvailable {
Label("Emoção da fala não foi detectada nesta análise — ligue em Avançado → Análise de Voz e refaça o passo 3.",
systemImage: "waveform.path.ecg")
.font(.caption2).foregroundStyle(.secondary)
}
HStack(spacing: 8) {
Button("Limpar ênfases") { model.setEmphasisForAll(0) }
.buttonStyle(.link).font(.caption)
Spacer()
Text("0–3 no teclado · ← → navega")
.font(.caption2).foregroundStyle(.secondary)
}
}
.padding(12)
}
}
/// One phrase in the inspector: the line as it will be said, plus every
/// decision attached to it. Kept in one row on purpose — jumping to a separate
/// detail pane to set a toggle would double the clicks on the most repeated
/// action in the screen.
private struct PhraseRow: View {
@Binding var phrase: ReviewPhrase
@ObservedObject var model: PhraseReviewModel
@State private var isEditing = false
var body: some View {
VStack(alignment: .leading, spacing: 6) {
HStack(spacing: 6) {
Text(phrase.timecode)
.font(.system(.caption2, design: .monospaced))
.foregroundStyle(.secondary)
if phrase.takeBoundary {
Image(systemName: "scissors.badge.ellipsis")
.font(.caption2).foregroundStyle(.orange)
.help("Nova tomada começa aqui")
}
if phrase.isTrimmed {
Image(systemName: "arrow.left.and.right.square")
.font(.caption2).foregroundStyle(.blue)
.help("Frase cortada nas pontas")
}
if model.emotionAvailable {
emotionChip
}
Spacer()
Toggle("", isOn: $phrase.active)
.toggleStyle(.switch)
.controlSize(.mini)
.labelsHidden()
.help(phrase.active ? "No corte" : "Fora do corte")
}
if isEditing {
TextField("Texto da frase", text: $phrase.text, axis: .vertical)
.textFieldStyle(.roundedBorder)
.font(.callout)
.onSubmit { isEditing = false }
} else {
Text(phrase.text.isEmpty ? "(sem texto)" : phrase.text)
.font(.callout)
.foregroundStyle(phrase.active ? .primary : .secondary)
.strikethrough(!phrase.active)
.onTapGesture(count: 2) { isEditing = true }
}
HStack(spacing: 8) {
Picker("", selection: $phrase.emphasis) {
ForEach(0..<4, id: \.self) { level in
Text(EmphasisPalette.label(level)).tag(level)
}
}
.pickerStyle(.segmented)
.controlSize(.mini)
.labelsHidden()
.disabled(!phrase.active)
Picker("", selection: $phrase.track) {
Text("Roteiro").tag(ReviewPhrase.trackScript)
Text("Bastidor").tag(ReviewPhrase.trackBackstage)
}
.pickerStyle(.menu)
.controlSize(.mini)
.labelsHidden()
.frame(width: 92)
}
if model.selection == phrase.id && !phrase.words.isEmpty {
wordTrimmer
}
}
.padding(.vertical, 4)
.opacity(phrase.active ? 1 : 0.55)
}
/// The delivery emotion the acoustics suggest. Shown faded below its own
/// confidence: a guess the analysis is unsure about should not compete for
/// attention with the emphasis decision, which is the point of the row.
private var emotionChip: some View {
let (label, icon) = ReviewPhrase.emotionLabel(phrase.emotion)
return Label(label, systemImage: icon)
.font(.caption2)
.padding(.horizontal, 5)
.padding(.vertical, 1)
.background(
Capsule().fill(Color.secondary.opacity(0.12))
)
.foregroundStyle(phrase.emotionConfidence >= 0.5 ? .secondary : .tertiary)
.help("Emoção da entrega: \(label) — confiança \(Int(phrase.emotionConfidence * 100))%")
}
/// Trimming by pointing at the transcript: click a word to start the phrase
/// there, option-click to end it there. Same edit as dragging the block's
/// edge on the timeline, but reachable while reading the line.
private var wordTrimmer: some View {
VStack(alignment: .leading, spacing: 4) {
HStack(spacing: 4) {
Text("Cortar pelas palavras").font(.caption2).foregroundStyle(.secondary)
Spacer()
if phrase.isTrimmed {
Button("Inteira") { model.resetTrim(phrase.id) }
.buttonStyle(.link).font(.caption2)
}
}
FlowWords(words: phrase.words, phrase: phrase) { word, edge in
model.trimToWord(word, edge: edge, in: phrase.id)
}
Text("Clique = começa/desfaz aqui · ⌥clique = termina/desfaz aqui · sublinhado = ênfase da palavra")
.font(.caption2).foregroundStyle(.tertiary)
}
.padding(.top, 2)
}
}
/// The phrase's words as wrapping chips, dimmed where they fall outside the
/// trim and underlined where the acoustics mark them as an emphasis peak —
/// the same word-level signal `05-zoom.md` picks a punch-in's `start` from,
/// made visible instead of buried in the JSON.
private struct FlowWords: View {
let words: [ReviewWord]
let phrase: ReviewPhrase
let onTrim: (ReviewWord, TrimEdge) -> Void
var body: some View {
// A LazyVGrid with adaptive columns wraps chips without a custom layout;
// phrases are short enough that the slight raggedness beats the cost of
// hand-rolling a flow layout here.
LazyVGrid(columns: [GridItem(.adaptive(minimum: 44), spacing: 3)],
alignment: .leading, spacing: 3) {
ForEach(words) { word in
let kept = word.start >= phrase.trimStart - 0.001 && word.end <= phrase.trimEnd + 0.001
let level = EmphasisPalette.levelFromScore(word.emphasis)
Text(word.text)
.font(.caption2)
.fontWeight(level >= 2 ? .semibold : .regular)
.padding(.horizontal, 4)
.padding(.vertical, 2)
.background(
RoundedRectangle(cornerRadius: 3)
.fill(kept ? Color.accentColor.opacity(0.12) : Color.secondary.opacity(0.08))
)
.overlay(alignment: .bottom) {
if level >= 1 {
Rectangle()
.fill(EmphasisPalette.color(level))
.frame(height: 2)
.padding(.horizontal, 3)
}
}
.foregroundStyle(kept ? .primary : .secondary)
.strikethrough(!kept)
.help(
level >= 1
? "Ênfase \(EmphasisPalette.label(level).lowercased()) (\(Int(word.emphasis * 100))%)"
: "Sem ênfase"
)
.onTapGesture {
onTrim(word, NSEvent.modifierFlags.contains(.option) ? .end : .start)
}
}
}
}
}
+69 -2
View File
@@ -44,12 +44,26 @@ enum PythonBridge {
return ["python3", scriptURL.path] return ["python3", scriptURL.path]
} }
/// `admin/models_api.py` lives outside `code/`, but its dependencies
/// (`pyproject.toml`, `.venv`) live inside it. `uv run` picks the
/// environment from the process's cwd, not from the script path — so
/// running with cwd at the repo root made `uv` create/use a second,
/// empty `.venv` there, silently ignoring everything installed into
/// `code/.venv` (this cost a real debugging session: librosa/pyannote
/// installed successfully but the app kept reporting them missing).
/// Every `uv run` must share the same cwd as `uv sync` to see the same
/// environment.
static var workingDirectory: URL { static var workingDirectory: URL {
projectRoot codeDirectory
}
/// Directory containing `pyproject.toml` — where `uv sync` must run from.
static var codeDirectory: URL {
projectRoot.appendingPathComponent("code")
} }
/// Locate `uv` on PATH or in common install locations. /// Locate `uv` on PATH or in common install locations.
private static func findUV() -> String? { static func findUV() -> String? {
if let onPath = which("uv") { return onPath } if let onPath = which("uv") { return onPath }
let candidates = [ let candidates = [
"/usr/local/bin/uv", "/usr/local/bin/uv",
@@ -148,6 +162,59 @@ enum PythonBridge {
} }
} }
// MARK: - uv sync (installing optional extras, e.g. acoustic analysis)
/// Runs `uv <arguments>` from `codeDirectory` (where `pyproject.toml`
/// lives), streaming each output line as plain text — used for
/// `sync --extra intelligence` so "Modelos" can install the librosa
/// extra without the user opening a terminal.
static func runUV(arguments: [String],
onLine: @escaping (String) -> Void,
completion: @escaping (Int, String?) -> Void) {
guard let uv = findUV() else {
completion(1, "uv não encontrado. Instale com: curl -LsSf https://astral.sh/uv/install.sh | sh")
return
}
let process = Process()
process.executableURL = URL(fileURLWithPath: "/usr/bin/env")
process.arguments = [uv] + arguments
process.currentDirectoryURL = codeDirectory
let pipe = Pipe()
process.standardOutput = pipe
process.standardError = pipe
var buffer = ""
let lock = NSLock()
pipe.fileHandleForReading.readabilityHandler = { handle in
let data = handle.availableData
guard !data.isEmpty, let s = String(data: data, encoding: .utf8) else { return }
lock.lock()
buffer += s
let parts = buffer.split(separator: "\n", omittingEmptySubsequences: false)
buffer = String(parts.last ?? "")
let lines = parts.dropLast()
lock.unlock()
for line in lines where !line.isEmpty { onLine(String(line)) }
}
process.terminationHandler = { p in
pipe.fileHandleForReading.readabilityHandler = nil
lock.lock()
let last = buffer.trimmingCharacters(in: .whitespacesAndNewlines)
buffer = ""
lock.unlock()
if !last.isEmpty { onLine(last) }
completion(Int(p.terminationStatus), p.terminationStatus == 0 ? nil : "uv sync terminou com erro (código \(p.terminationStatus)).")
}
do {
try process.run()
} catch {
completion(1, error.localizedDescription)
}
}
// MARK: - Convenience: single JSON result // MARK: - Convenience: single JSON result
/// Runs a command and delivers the first parsed JSON document as the result. /// Runs a command and delivers the first parsed JSON document as the result.
@@ -0,0 +1,529 @@
import SwiftUI
/// Colors shared by the timeline and the inspector, so a block and its row in
/// the list always read as the same thing.
enum EmphasisPalette {
static func color(_ level: Int) -> Color {
switch level {
case 1: return Color.blue
case 2: return Color.orange
case 3: return Color.pink
default: return Color.secondary
}
}
static func label(_ level: Int) -> String {
switch level {
case 1: return "Leve"
case 2: return "Média"
case 3: return "Forte"
default: return "Sem"
}
}
/// The same 0–3 tiers a phrase's `emphasis` uses, derived from a raw 0–1
/// acoustic score — the thresholds `10-revisao-humana.md` documents for
/// deriving a phrase's level from `peak_emphasis` when no explicit zoom
/// was set, reused here per WORD so a word chip and a phrase row read as
/// the same scale.
static func levelFromScore(_ score: Double) -> Int {
switch score {
case ..<0.25: return 0
case ..<0.45: return 1
case ..<0.65: return 2
default: return 3
}
}
static func speakerColor(_ speaker: String, among speakers: [String]) -> Color {
let palette: [Color] = [.teal, .purple, .green, .indigo, .brown, .cyan]
guard let index = speakers.firstIndex(of: speaker) else { return .gray }
return palette[index % palette.count]
}
}
/// The timeline strip: four stacked tracks over one shared time axis.
///
/// Phrases are laid out as real views rather than drawn into a Canvas, because
/// every one of them is a target — click to select, drag its edge to trim,
/// right-click to change emphasis. The dense per-word energy track *is* a
/// Canvas: it has thousands of bars and nothing to hit.
struct TimelineTracksView: View {
@ObservedObject var model: PhraseReviewModel
private let rulerHeight: CGFloat = 18
private let phraseHeight: CGFloat = 46
private let energyHeight: CGFloat = 34
private let stripHeight: CGFloat = 12
private let handleWidth: CGFloat = 8
private let gutterWidth: CGFloat = 92
private let trackSpacing: CGFloat = 4
private var pps: CGFloat { CGFloat(model.pixelsPerSecond) }
/// Width follows the *kept* duration, not the raw take's — the timeline
/// draws the cut, so removed stretches take no horizontal space.
private var contentWidth: CGFloat { max(320, CGFloat(model.keptDuration) * pps) }
/// Name, icon and height of each lane, in the order they stack. The gutter
/// and the tracks are built from this one list so a label can never drift
/// off the lane it names.
private var lanes: [(label: String, icon: String, height: CGFloat)] {
[
("", "", rulerHeight),
("Zooms", "plus.magnifyingglass", stripHeight + 6),
("Frases", "text.quote", phraseHeight),
("Energia", "waveform", energyHeight),
("Emoção", "face.smiling", stripHeight),
("Locutor", "person.wave.2", stripHeight),
("Roteiro", "list.bullet.rectangle", stripHeight),
]
}
var body: some View {
VStack(spacing: 0) {
toolbar
Divider()
HStack(alignment: .top, spacing: 0) {
gutter
Divider()
timelineScroller
}
}
.background(Color(nsColor: .underPageBackgroundColor))
}
/// Fixed column naming each lane. Without it the stripes are six colours
/// with no way to tell which one is emotion and which one is the speaker.
private var gutter: some View {
VStack(alignment: .leading, spacing: trackSpacing) {
ForEach(lanes.indices, id: \.self) { index in
let lane = lanes[index]
HStack(spacing: 4) {
if !lane.icon.isEmpty {
Image(systemName: lane.icon).font(.system(size: 9))
}
Text(lane.label).font(.system(size: 10))
Spacer(minLength: 0)
}
.foregroundStyle(.secondary)
.frame(height: lane.height, alignment: .center)
}
}
.padding(.horizontal, 8)
.padding(.vertical, 8)
.frame(width: gutterWidth, alignment: .leading)
}
private var timelineScroller: some View {
ScrollViewReader { proxy in
ScrollView([.horizontal]) {
ZStack(alignment: .topLeading) {
VStack(alignment: .leading, spacing: trackSpacing) {
ruler
zoomTrack
phraseTrack
energyTrack
emotionTrack
speakerTrack
scriptTrack
}
.frame(width: contentWidth, alignment: .leading)
rangeOverlay
playhead
// Anchors the auto-scroll: one invisible marker per
// phrase, so selecting a line off-screen brings it in.
ForEach(model.phrases) { phrase in
Color.clear
.frame(width: 1, height: 1)
.offset(x: x(phrase.start))
.id(phrase.id)
}
}
.padding(.vertical, 8)
.contentShape(Rectangle())
.gesture(scrubGesture)
.contextMenu { timelineMenu }
}
.onChange(of: model.selection) { _, newValue in
guard let newValue else { return }
withAnimation(.easeOut(duration: 0.2)) {
proxy.scrollTo(newValue, anchor: .center)
}
}
}
}
// MARK: - Barra de controles
private var toolbar: some View {
HStack(spacing: 12) {
Button {
model.togglePlay()
} label: {
Image(systemName: model.isPlaying ? "pause.fill" : "play.fill")
}
.buttonStyle(.borderless)
.help("Reproduzir (espaço)")
.disabled(model.player == nil)
Text(timecode(model.currentTime))
.font(.system(.caption, design: .monospaced))
.foregroundStyle(.secondary)
Button {
model.playSelectedPhrase()
} label: {
Image(systemName: "play.rectangle")
}
.buttonStyle(.borderless)
.help("Tocar só a frase selecionada (⏎)")
.disabled(model.player == nil || model.selection == nil)
Toggle("Pular removidos", isOn: $model.skipRemoved)
.toggleStyle(.checkbox)
.font(.caption)
.help("Durante a reprodução, salta os trechos desativados — mostra como o corte ficou.")
Button {
model.addZoomForRange()
} label: {
Label("Zoom no trecho", systemImage: "plus.magnifyingglass")
}
.buttonStyle(.borderless)
.font(.caption)
.disabled(!model.hasRange)
.help("Arraste na timeline para marcar um trecho e crie um zoom nele. A escala vem de Análise de Voz.")
Spacer()
legend
Spacer()
Image(systemName: "minus.magnifyingglass").foregroundStyle(.secondary)
Slider(value: $model.pixelsPerSecond,
in: model.minPixelsPerSecond...model.maxPixelsPerSecond)
.frame(width: 130)
Image(systemName: "plus.magnifyingglass").foregroundStyle(.secondary)
}
.padding(.horizontal, 12)
.padding(.vertical, 8)
}
private var legend: some View {
HStack(spacing: 10) {
ForEach(0..<4, id: \.self) { level in
HStack(spacing: 4) {
RoundedRectangle(cornerRadius: 2)
.fill(EmphasisPalette.color(level))
.frame(width: 10, height: 10)
Text(EmphasisPalette.label(level)).font(.caption2)
}
}
}
.foregroundStyle(.secondary)
}
// MARK: - Trilhas
private var ruler: some View {
Canvas { context, size in
let step = tickStep()
var time = 0.0
while time <= model.keptDuration {
let position = compactX(time)
context.stroke(
Path { $0.move(to: CGPoint(x: position, y: size.height - 6))
$0.addLine(to: CGPoint(x: position, y: size.height)) },
with: .color(.secondary.opacity(0.5))
)
context.draw(
Text(timecode(time)).font(.system(size: 9, design: .monospaced))
.foregroundColor(.secondary),
at: CGPoint(x: position + 18, y: 6)
)
time += step
}
}
.frame(width: contentWidth, height: rulerHeight)
}
private var phraseTrack: some View {
ZStack(alignment: .topLeading) {
RoundedRectangle(cornerRadius: 4)
.fill(Color.secondary.opacity(0.06))
.frame(width: contentWidth, height: phraseHeight)
ForEach(model.phrases) { phrase in
phraseBlock(phrase)
}
}
.frame(width: contentWidth, height: phraseHeight, alignment: .topLeading)
}
@ViewBuilder
private func phraseBlock(_ phrase: ReviewPhrase) -> some View {
let isSelected = model.selection == phrase.id
let color = EmphasisPalette.color(phrase.emphasis)
let fullWidth = max(2, width(from: phrase.start, to: phrase.end))
let keptWidth = max(1, width(from: phrase.trimStart, to: phrase.trimEnd))
ZStack(alignment: .topLeading) {
// The whole line, dim — what is there before the edit.
RoundedRectangle(cornerRadius: 4)
.fill(color.opacity(phrase.active ? 0.15 : 0.10))
.frame(width: fullWidth, height: phraseHeight)
// What survives: the kept span, drawn solid over it.
RoundedRectangle(cornerRadius: 4)
.fill(color.opacity(phrase.active ? 0.55 : 0.12))
.frame(width: keptWidth, height: phraseHeight)
.offset(x: width(from: phrase.start, to: phrase.trimStart))
Text(phrase.text)
.font(.system(size: 10))
.lineLimit(2)
.padding(.horizontal, 4)
.frame(width: fullWidth, height: phraseHeight, alignment: .topLeading)
.foregroundStyle(phrase.active ? .primary : .secondary)
.strikethrough(!phrase.active)
RoundedRectangle(cornerRadius: 4)
.stroke(isSelected ? Color.accentColor : color.opacity(0.4),
lineWidth: isSelected ? 2 : 1)
.frame(width: fullWidth, height: phraseHeight)
if isSelected && phrase.active {
trimHandle(phrase, edge: .start)
trimHandle(phrase, edge: .end)
}
}
.frame(width: fullWidth, height: phraseHeight, alignment: .topLeading)
.offset(x: x(phrase.start))
.contentShape(Rectangle())
.onTapGesture { model.goTo(phraseID: phrase.id) }
.contextMenu { phraseMenu(phrase) }
.help(phrase.reason.isEmpty ? phrase.text : "\(phrase.text)\n— \(phrase.reason)")
}
private func trimHandle(_ phrase: ReviewPhrase, edge: TrimEdge) -> some View {
let offset = edge == .start
? width(from: phrase.start, to: phrase.trimStart)
: width(from: phrase.start, to: phrase.trimEnd) - handleWidth
return RoundedRectangle(cornerRadius: 2)
.fill(Color.accentColor)
.frame(width: handleWidth, height: phraseHeight)
.offset(x: offset)
.gesture(
DragGesture(minimumDistance: 1)
.onChanged { value in
let compactOrigin = model.compactTime(phrase.start)
let time = model.rawTime(fromCompact: compactOrigin + Double(value.location.x / pps))
model.trim(phrase.id, edge: edge, to: time)
}
)
.help(edge == .start ? "Arraste para cortar o começo (pula de palavra em palavra)"
: "Arraste para cortar o fim (pula de palavra em palavra)")
}
@ViewBuilder
private func phraseMenu(_ phrase: ReviewPhrase) -> some View {
Button("Tocar esta frase") {
model.goTo(phraseID: phrase.id)
model.playSelectedPhrase()
}
Button(phrase.active ? "Remover do corte" : "Trazer de volta") {
model.toggleActive(phrase.id)
}
Button("Adicionar zoom nesta frase") { model.addZoomForPhrase(phrase.id) }
Divider()
ForEach(0..<4, id: \.self) { level in
Button("Ênfase: \(EmphasisPalette.label(level))") {
model.setEmphasis(level, for: phrase.id)
}
}
Divider()
Button(phrase.isBackstage ? "Marcar como roteiro" : "Marcar como bastidor") {
model.setTrack(phrase.isBackstage ? ReviewPhrase.trackScript : ReviewPhrase.trackBackstage,
for: phrase.id)
}
if phrase.isTrimmed {
Divider()
Button("Desfazer corte da frase") { model.resetTrim(phrase.id) }
}
}
/// Per-word energy/emphasis, straight from the voice timeline — the closest
/// thing to a waveform without opening the audio again.
private var energyTrack: some View {
Canvas { context, size in
for phrase in model.phrases {
for word in phrase.words {
let start = x(word.start)
let barWidth = max(1, width(from: word.start, to: word.end) - 1)
let height = size.height * CGFloat(max(0.04, word.energy))
let rect = CGRect(x: start, y: size.height - height,
width: barWidth, height: height)
let color = word.emphasis >= 0.65 ? Color.pink
: word.emphasis >= 0.45 ? Color.orange
: Color.secondary
context.fill(Path(rect),
with: .color(color.opacity(phrase.active ? 0.6 : 0.2)))
}
}
}
.frame(width: contentWidth, height: energyHeight)
.background(RoundedRectangle(cornerRadius: 4).fill(Color.secondary.opacity(0.06)))
}
private var speakerTrack: some View {
stripTrack { phrase in
EmphasisPalette.speakerColor(phrase.speaker, among: model.speakers)
}
}
private var scriptTrack: some View {
stripTrack { phrase in phrase.isBackstage ? Color.gray : Color.mint }
}
private func stripTrack(_ color: @escaping (ReviewPhrase) -> Color) -> some View {
Canvas { context, size in
for phrase in model.phrases {
let rect = CGRect(x: x(phrase.start), y: 0,
width: max(1, width(from: phrase.start, to: phrase.end)),
height: size.height)
context.fill(Path(roundedRect: rect, cornerRadius: 2),
with: .color(color(phrase).opacity(phrase.active ? 0.7 : 0.2)))
}
}
.frame(width: contentWidth, height: stripHeight)
}
private var playhead: some View {
Rectangle()
.fill(Color.red)
.frame(width: 1.5)
.offset(x: x(model.currentTime))
.allowsHitTesting(false)
}
/// One gesture, two meanings, decided by whether the mouse moved: a click
/// parks the playhead, a drag marks in/out. Splitting them across separate
/// controls would mean choosing a tool before every action, which is
/// exactly the ceremony this screen is meant to avoid.
private var scrubGesture: some Gesture {
DragGesture(minimumDistance: 0)
.onChanged { value in
let from = model.rawTime(fromCompact: Double(value.startLocation.x / pps))
let to = model.rawTime(fromCompact: Double(value.location.x / pps))
if abs(value.translation.width) > 3 {
model.setRange(from: from, to: to)
model.seek(to: min(from, to))
} else {
model.clearRange()
model.seek(to: to)
}
}
}
/// The marked in/out, drawn over every track so the span reads against the
/// phrases and the energy at once.
private var rangeOverlay: some View {
Group {
if let span = model.rangeSpan {
Rectangle()
.fill(Color.accentColor.opacity(0.18))
.overlay(Rectangle().stroke(Color.accentColor.opacity(0.6), lineWidth: 1))
.frame(width: max(1, width(from: span.start, to: span.end)))
.offset(x: x(span.start))
.allowsHitTesting(false)
}
}
}
@ViewBuilder
private var timelineMenu: some View {
if model.hasRange, let span = model.rangeSpan {
Button("Adicionar zoom no trecho (\(secondsLabel(span.end - span.start)))") {
model.addZoomForRange()
}
Button("Tocar o trecho") { model.playRange(from: span.start, to: span.end) }
Button("Limpar seleção") { model.clearRange() }
} else {
Text("Arraste na timeline para marcar um trecho")
}
if let zoom = model.zoom(at: model.currentTime) {
Divider()
Button("Remover o zoom daqui") { model.removeZoom(zoom.id) }
}
}
private func secondsLabel(_ seconds: Double) -> String {
String(format: "%.1fs", seconds)
}
/// Punch-ins, on their own lane above the script: they are a second layer
/// over the same time, not a property of a phrase.
private var zoomTrack: some View {
ZStack(alignment: .topLeading) {
RoundedRectangle(cornerRadius: 3)
.fill(Color.secondary.opacity(0.06))
.frame(width: contentWidth, height: stripHeight + 6)
ForEach(model.zooms) { zoom in
RoundedRectangle(cornerRadius: 3)
.fill(Color.yellow.opacity(0.55))
.overlay(
Image(systemName: "plus.magnifyingglass")
.font(.system(size: 8)).foregroundStyle(.black.opacity(0.6))
)
.frame(width: max(6, width(from: zoom.start, to: zoom.end)),
height: stripHeight + 6)
.offset(x: x(zoom.start))
.help("Zoom marcado — \(secondsLabel(zoom.end - zoom.start)). A escala vem de Análise de Voz.")
.contextMenu {
Button("Remover este zoom") { model.removeZoom(zoom.id) }
}
}
}
.frame(width: contentWidth, height: stripHeight + 6, alignment: .topLeading)
}
/// Delivery emotion per phrase — the fourth signal to read against the text.
private var emotionTrack: some View {
stripTrack { phrase in
switch phrase.emotion {
case "excited": return .orange
case "tense": return .red
case "calm": return .blue
case "reflective": return .purple
default: return .secondary
}
}
}
// MARK: - Escala
/// Pixel position of a raw source-media time, after collapsing whatever
/// lies between it and the previous kept phrase.
private func x(_ time: Double) -> CGFloat { compactX(model.compactTime(time)) }
/// Pixel position of a time already in the compacted (edited) timeline —
/// used for the ruler and playhead, which think in that space directly.
private func compactX(_ compactTime: Double) -> CGFloat { CGFloat(compactTime) * pps }
private func width(from: Double, to: Double) -> CGFloat {
max(0, CGFloat(model.compactTime(to) - model.compactTime(from)) * pps)
}
/// Ruler spacing that keeps labels ~80pt apart at any zoom.
private func tickStep() -> Double {
let candidates: [Double] = [1, 2, 5, 10, 15, 30, 60, 120, 300, 600]
let wanted = 80 / Double(pps)
return candidates.first { $0 >= wanted } ?? 600
}
private func timecode(_ seconds: Double) -> String {
let total = Int(seconds.rounded(.down))
return String(format: "%02d:%02d", total / 60, total % 60)
}
}
+3 -3
View File
@@ -271,7 +271,7 @@ struct TranscriptionView: View {
Toggle("Marcar o que foi dito na timeline", isOn: $batchMarkers) Toggle("Marcar o que foi dito na timeline", isOn: $batchMarkers)
Divider() Divider()
Toggle("Exportar legendas SRT", isOn: $batchSubtitles) Toggle("Gerar legenda comum (texto editável no FCP)", isOn: $batchSubtitles)
Divider() Divider()
batchOptionRow( batchOptionRow(
@@ -793,7 +793,7 @@ struct TranscriptionView: View {
if batchFillers { operations.append("remove_filler_words") } if batchFillers { operations.append("remove_filler_words") }
if batchPhrases { operations.append("edit_by_transcript") } if batchPhrases { operations.append("edit_by_transcript") }
if batchMarkers { operations.append("transcript_markers") } if batchMarkers { operations.append("transcript_markers") }
if batchSubtitles { operations.append("export_srt") } if batchSubtitles { operations.append("generate_plain_subtitles") }
// Runs last, on the timing already cut by any earlier steps (see the // Runs last, on the timing already cut by any earlier steps (see the
// "Abrir no Final Cut Pro" fallback chain and exportSubtitles()'s own // "Abrir no Final Cut Pro" fallback chain and exportSubtitles()'s own
// preference for `processedPath` — same reasoning). // preference for `processedPath` — same reasoning).
@@ -834,7 +834,7 @@ struct TranscriptionView: View {
} }
let nextPath = result?["path"] as? String ?? currentPath let nextPath = result?["path"] as? String ?? currentPath
if operation == "remove_silences" { processedPath = nextPath } if operation == "remove_silences" { processedPath = nextPath }
if operation == "export_srt" { subtitlePaths = result?["paths"] as? [String] ?? [] } if operation == "generate_plain_subtitles" { subtitlePaths = [nextPath] }
if operation == "generate_dynamic_subtitles" { dynamicSubtitlesPath = nextPath } if operation == "generate_dynamic_subtitles" { dynamicSubtitlesPath = nextPath }
processBatchStep(operations, index: index + 1, currentPath: nextPath, outputFolder: outputFolder) processBatchStep(operations, index: index + 1, currentPath: nextPath, outputFolder: outputFolder)
} }
+68 -4
View File
@@ -23,6 +23,7 @@ struct VoiceAnalysisView: View {
} else { } else {
energySection energySection
emphasisSection emphasisSection
zoomSection
weightsSection weightsSection
emotionSection emotionSection
resetSection resetSection
@@ -90,6 +91,44 @@ struct VoiceAnalysisView: View {
} }
} }
private var zoomSection: some View {
Section {
sliderRow(
title: "Zoom na ênfase",
value: $config.zoomScale,
range: 1.0...3.0,
readout: "\(Int(config.zoomScale * 100))%",
help: "Fator aplicado nos punch-ins de ênfase. 130% equivale a escala 1,30 no Final Cut."
)
Picker("Movimento", selection: $config.zoomMode) {
Text("Zoom in e out").tag("in_out")
Text("Só zoom in").tag("in")
Text("Só zoom out").tag("out")
}
.onChange(of: config.zoomMode) { _, _ in save() }
sliderRow(
title: "Velocidade do zoom in",
value: $config.zoomEaseIn,
range: 0.05...2.0,
readout: String(format: "%.2fs", config.zoomEaseIn),
help: "Duração da entrada do zoom. Menor é mais rápido."
)
sliderRow(
title: "Velocidade do zoom out",
value: $config.zoomEaseOut,
range: 0.01...2.0,
readout: String(format: "%.2fs", config.zoomEaseOut),
help: "Duração da saída do zoom. Menor é mais seco."
)
} header: {
Text("Zoom de Ênfase")
} footer: {
Text("Esses valores viram o padrão para ações de zoom que não trouxerem scale/ease/ease_out no JSON da edição por voz.")
.font(.caption)
.foregroundStyle(.secondary)
}
}
// MARK: - Emoção // MARK: - Emoção
private var emotionSection: some View { private var emotionSection: some View {
@@ -129,13 +168,14 @@ struct VoiceAnalysisView: View {
title: String, title: String,
value: Binding<Double>, value: Binding<Double>,
range: ClosedRange<Double> = 0...1, range: ClosedRange<Double> = 0...1,
readout: String? = nil,
help: String? = nil help: String? = nil
) -> some View { ) -> some View {
VStack(alignment: .leading, spacing: 2) { VStack(alignment: .leading, spacing: 2) {
HStack { HStack {
Text(title) Text(title)
Spacer() Spacer()
Text(String(format: "%.2f", value.wrappedValue)) Text(readout ?? String(format: "%.2f", value.wrappedValue))
.monospacedDigit() .monospacedDigit()
.foregroundStyle(.secondary) .foregroundStyle(.secondary)
} }
@@ -188,6 +228,10 @@ struct VoiceAnalysisConfig {
var weightDuration: Double var weightDuration: Double
var emotionEnabled: Bool var emotionEnabled: Bool
var emotionSensitivity: Double var emotionSensitivity: Double
var zoomScale: Double
var zoomMode: String
var zoomEaseIn: Double
var zoomEaseOut: Double
static let defaults = VoiceAnalysisConfig( static let defaults = VoiceAnalysisConfig(
energyThreshold: 0.5, energyThreshold: 0.5,
@@ -198,7 +242,11 @@ struct VoiceAnalysisConfig {
weightPause: 0.15, weightPause: 0.15,
weightDuration: 0.10, weightDuration: 0.10,
emotionEnabled: false, emotionEnabled: false,
emotionSensitivity: 0.5 emotionSensitivity: 0.5,
zoomScale: 1.30,
zoomMode: "in_out",
zoomEaseIn: 0.25,
zoomEaseOut: 0.04
) )
init( init(
@@ -210,7 +258,11 @@ struct VoiceAnalysisConfig {
weightPause: Double, weightPause: Double,
weightDuration: Double, weightDuration: Double,
emotionEnabled: Bool, emotionEnabled: Bool,
emotionSensitivity: Double emotionSensitivity: Double,
zoomScale: Double,
zoomMode: String,
zoomEaseIn: Double,
zoomEaseOut: Double
) { ) {
self.energyThreshold = energyThreshold self.energyThreshold = energyThreshold
self.emphasisThreshold = emphasisThreshold self.emphasisThreshold = emphasisThreshold
@@ -221,6 +273,10 @@ struct VoiceAnalysisConfig {
self.weightDuration = weightDuration self.weightDuration = weightDuration
self.emotionEnabled = emotionEnabled self.emotionEnabled = emotionEnabled
self.emotionSensitivity = emotionSensitivity self.emotionSensitivity = emotionSensitivity
self.zoomScale = zoomScale
self.zoomMode = zoomMode
self.zoomEaseIn = zoomEaseIn
self.zoomEaseOut = zoomEaseOut
} }
/// Lê a resposta do bridge, caindo no padrão para qualquer campo ausente. /// Lê a resposta do bridge, caindo no padrão para qualquer campo ausente.
@@ -236,7 +292,11 @@ struct VoiceAnalysisConfig {
weightPause: weights["pause_before"] as? Double ?? defaults.weightPause, weightPause: weights["pause_before"] as? Double ?? defaults.weightPause,
weightDuration: weights["duration"] as? Double ?? defaults.weightDuration, weightDuration: weights["duration"] as? Double ?? defaults.weightDuration,
emotionEnabled: json["emotion_enabled"] as? Bool ?? defaults.emotionEnabled, emotionEnabled: json["emotion_enabled"] as? Bool ?? defaults.emotionEnabled,
emotionSensitivity: json["emotion_sensitivity"] as? Double ?? defaults.emotionSensitivity emotionSensitivity: json["emotion_sensitivity"] as? Double ?? defaults.emotionSensitivity,
zoomScale: json["zoom_scale"] as? Double ?? defaults.zoomScale,
zoomMode: json["zoom_mode"] as? String ?? defaults.zoomMode,
zoomEaseIn: json["zoom_ease_in"] as? Double ?? defaults.zoomEaseIn,
zoomEaseOut: json["zoom_ease_out"] as? Double ?? defaults.zoomEaseOut
) )
} }
@@ -253,6 +313,10 @@ struct VoiceAnalysisConfig {
], ],
"emotion_enabled": emotionEnabled, "emotion_enabled": emotionEnabled,
"emotion_sensitivity": emotionSensitivity, "emotion_sensitivity": emotionSensitivity,
"zoom_scale": zoomScale,
"zoom_mode": zoomMode,
"zoom_ease_in": zoomEaseIn,
"zoom_ease_out": zoomEaseOut,
] ]
} }
} }
+967
View File
@@ -0,0 +1,967 @@
import SwiftUI
import AppKit
/// Guia passo a passo do fluxo completo: projeto → transcrição → análise de
/// voz → copiar para o chat e trazer as decisões → revisar as ênfases →
/// processamento final. Existe para que o usuário não precise entender a ordem
/// certa de botões espalhados em várias abas — cada etapa só libera a próxima
/// quando o passo anterior terminou, e a "ponte" com o chat (que hoje exigia
/// sair do app e escolher um arquivo na mão) vira copiar/colar assistido
/// dentro da própria tela.
enum WizardStep: Int, CaseIterable, Identifiable {
case projeto, transcricao, analise, exportarChat, revisar, finalizar
var id: Int { rawValue }
var titulo: String {
switch self {
case .projeto: return "Projeto"
case .transcricao: return "Transcrever"
case .analise: return "Analisar voz"
case .exportarChat: return "Decisões da IA"
case .revisar: return "Revisar ênfases"
case .finalizar: return "Processar"
}
}
}
struct WizardView: View {
@State private var step: WizardStep = .projeto
// Passo 1 — projeto
@State private var outputFolder: String?
@State private var projectPath: String?
@State private var catalog: Catalog?
// Passo 2 — transcrição
@State private var isTranscribing = false
@State private var transcribeProgress: Double = 0
@State private var transcribeStage = ""
@State private var transcribeResults: [TranscriptResult] = []
// Passo 3 — análise de voz
@State private var isAnalyzing = false
@State private var voiceTimelinePath: String?
@State private var voiceAnalysisMessage = ""
@State private var acousticsAvailable: Bool?
@State private var showVoiceTimelineReuseAlert = false
@State private var existingVoiceTimelinePath: String?
// Passo 4 — enviar ao chat e trazer as decisões de volta
@State private var copiedFeedback = ""
@State private var decisionsText = ""
@State private var isApplyingDecisions = false
@State private var appliedPath: String?
@State private var skippedVoiceEdit = false
// Passo 4 (alternativa) — gerar o roteiro direto por IA local (Ollama/Gemma 3)
@State private var isGeneratingScript = false
@State private var generateScriptModel = "gemma3:12b"
@State private var generateScriptFeedback = ""
@State private var ollamaModels: [String] = []
// Passo 5 — revisar ênfases
@StateObject private var reviewModel = PhraseReviewModel()
@State private var reviewLoadedFor: String?
@State private var reviewLoadedForDecisions: String?
@State private var phraseReviewPath: String?
@State private var phraseActionsPath: String?
// Passo 6 — processamento final
@State private var finalSilences = true
@State private var finalFillers = false
@State private var finalSubtitles = true
@State private var finalDynamicSubtitles = false
@State private var isFinalizing = false
@State private var finalStatus = ""
@State private var finalPath: String?
@State private var errorMessage: String?
var body: some View {
VStack(spacing: 0) {
stepperHeader
.padding(.horizontal, 24)
.padding(.top, 20)
.padding(.bottom, 16)
Divider()
// A revisão é uma sala de edição, não um formulário: ela precisa da
// largura toda e rola por conta própria (timeline horizontal, lista
// vertical). As demais etapas continuam na coluna estreita, que é o
// que mantém um passo a passo legível.
if step == .revisar {
revisarStep
} else {
ScrollView {
VStack(alignment: .leading, spacing: 18) {
if let errorMessage, !errorMessage.isEmpty {
Label(errorMessage, systemImage: "exclamationmark.triangle.fill")
.foregroundStyle(.red)
.padding(.top, 4)
}
content
}
.padding(24)
.frame(maxWidth: 640, alignment: .leading)
.frame(maxWidth: .infinity)
}
}
Divider()
navFooter
.padding(.horizontal, 24)
.padding(.vertical, 16)
}
.task {
loadProjectConfig()
await loadCatalog()
}
.alert("Análise de voz já existe", isPresented: $showVoiceTimelineReuseAlert) {
Button("Usar existente") {
if let existingVoiceTimelinePath {
voiceTimelinePath = existingVoiceTimelinePath
voiceAnalysisMessage = "Reaproveitando análise existente: \(existingVoiceTimelinePath)"
}
}
Button("Reprocessar") {
analyzeVoice(forceReprocess: true)
}
Button("Cancelar", role: .cancel) {}
} message: {
Text("Já existe um arquivo voice_timeline para este projeto. Quer manter o processamento anterior para ganhar tempo?")
}
}
// MARK: - Cabeçalho com os passos
private var stepperHeader: some View {
HStack(spacing: 6) {
ForEach(WizardStep.allCases) { s in
HStack(spacing: 6) {
ZStack {
Circle()
.fill(colorFor(s))
.frame(width: 24, height: 24)
if s.rawValue < step.rawValue {
Image(systemName: "checkmark")
.font(.caption2.weight(.bold))
.foregroundStyle(.white)
} else {
Text("\(s.rawValue + 1)")
.font(.caption2.weight(.bold))
.foregroundStyle(s == step ? .white : .secondary)
}
}
Text(s.titulo)
.font(.caption)
.foregroundStyle(s == step ? .primary : .secondary)
.fontWeight(s == step ? .semibold : .regular)
}
if s != WizardStep.allCases.last {
Rectangle()
.fill(s.rawValue < step.rawValue ? Color.accentColor : Color.secondary.opacity(0.25))
.frame(height: 2)
.frame(maxWidth: .infinity)
}
}
}
}
private func colorFor(_ s: WizardStep) -> Color {
if s.rawValue < step.rawValue { return .accentColor }
if s == step { return .accentColor }
return Color.secondary.opacity(0.25)
}
// MARK: - Conteúdo por etapa
@ViewBuilder
private var content: some View {
switch step {
case .projeto: projetoStep
case .transcricao: transcricaoStep
case .analise: analiseStep
case .exportarChat: exportarChatStep
case .revisar: revisarStep
case .finalizar: finalizarStep
}
}
private var projetoStep: some View {
VStack(alignment: .leading, spacing: 16) {
Text("1. Escolha o projeto").font(.title3.weight(.semibold))
Text("A pasta é onde tudo o que for gerado nesse fluxo fica salvo. O arquivo é o .fcpxml exportado do Final Cut Pro.")
.font(.callout).foregroundStyle(.secondary)
fieldRow(icon: "folder", label: outputFolder ?? "Nenhuma pasta selecionada", isSet: outputFolder != nil) {
pickOutputFolder()
}
fieldRow(icon: "doc.text", label: projectPath.map { URL(fileURLWithPath: $0).lastPathComponent } ?? "Nenhum arquivo selecionado", isSet: projectPath != nil) {
pickProjectFile()
}
if looksLikeGeneratedFile(projectPath) {
Label("Esse arquivo parece já ter sido processado por este fluxo (o nome tem um sufixo como \"_voice_edit\" ou \"_silence_removed\"). Rodar o wizard de novo em cima dele reaplica os cortes por cima de cortes já feitos. Selecione o .fcpxml original do Final Cut, a menos que a intenção seja mesmo reprocessar.",
systemImage: "exclamationmark.triangle.fill")
.font(.caption).foregroundStyle(.orange)
}
if (catalog?.installedCount ?? 0) == 0 {
Label("Nenhum modelo de transcrição instalado. Baixe um na aba \"Modelos\" antes de continuar.",
systemImage: "exclamationmark.triangle.fill")
.font(.caption).foregroundStyle(.orange)
}
}
}
private var transcricaoStep: some View {
VStack(alignment: .leading, spacing: 16) {
Text("2. Transcreva o áudio").font(.title3.weight(.semibold))
Text("Roda localmente com o modelo escolhido na aba Modelos. Vira a base de tudo que vem depois — o corte por voz, as legendas, os marcadores.")
.font(.callout).foregroundStyle(.secondary)
Button {
startTranscription()
} label: {
if isTranscribing {
HStack { ProgressView().controlSize(.small); Text(transcribeStage.isEmpty ? "Transcrevendo…" : transcribeStage) }
.frame(maxWidth: .infinity)
} else {
Label(transcribeResults.isEmpty ? "Transcrever" : "Transcrever novamente", systemImage: "waveform")
.frame(maxWidth: .infinity)
}
}
.buttonStyle(.borderedProminent)
.controlSize(.large)
.disabled(isTranscribing || projectPath == nil || outputFolder == nil)
if isTranscribing {
VStack(alignment: .leading, spacing: 6) {
ProgressView(value: transcribeProgress)
Text("\(Int(transcribeProgress * 100))%").font(.caption).foregroundStyle(.secondary).monospacedDigit()
}
}
if !transcribeResults.isEmpty {
ForEach(transcribeResults, id: \.media) { r in
VStack(alignment: .leading, spacing: 4) {
HStack {
Image(systemName: "checkmark.circle.fill").foregroundStyle(.green)
Text(r.media).font(.body.weight(.medium))
Spacer()
Text("\(r.language) · \(r.words) palavras").font(.caption).foregroundStyle(.secondary)
}
Text(r.preview).font(.caption).foregroundStyle(.secondary).lineLimit(2)
}
.padding(12)
.background(RoundedRectangle(cornerRadius: 8).fill(Color.secondary.opacity(0.06)))
}
}
}
}
private var analiseStep: some View {
VStack(alignment: .leading, spacing: 16) {
Text("3. Analise a voz").font(.title3.weight(.semibold))
Text("Gera o JSON com transcrição, locutor e intensidade (pitch/energia/ritmo) por palavra — é esse arquivo que o chat lê para decidir o que cortar. Não corta nada sozinho.")
.font(.callout).foregroundStyle(.secondary)
Button {
analyzeVoice()
} label: {
if isAnalyzing {
HStack { ProgressView().controlSize(.small); Text("Analisando…") }.frame(maxWidth: .infinity)
} else {
Label(voiceTimelinePath == nil ? "Analisar voz" : "Analisar novamente", systemImage: "waveform.badge.magnifyingglass")
.frame(maxWidth: .infinity)
}
}
.buttonStyle(.borderedProminent)
.controlSize(.large)
.disabled(isAnalyzing || projectPath == nil || outputFolder == nil)
if let voiceTimelinePath {
VStack(alignment: .leading, spacing: 6) {
Label("Análise pronta", systemImage: "checkmark.circle.fill").foregroundStyle(.green)
Text(voiceTimelinePath).font(.caption).foregroundStyle(.secondary).lineLimit(1).truncationMode(.middle)
}
.padding(12)
.background(RoundedRectangle(cornerRadius: 8).fill(Color.secondary.opacity(0.06)))
if acousticsAvailable == false {
VStack(alignment: .leading, spacing: 4) {
Label("Sem análise acústica real", systemImage: "exclamationmark.triangle.fill")
.font(.caption.weight(.semibold)).foregroundStyle(.orange)
Text("Falta o componente \"librosa\" — os cortes ainda são decididos pelo texto, mas o chat não vai propor zoom com confiança. Instale em Avançado → Modelos → \"Análise Acústica\", e refaça esta etapa depois.")
.font(.caption).foregroundStyle(.secondary)
}
.padding(12)
.background(RoundedRectangle(cornerRadius: 8).fill(Color.orange.opacity(0.08)))
}
}
}
}
private var exportarChatStep: some View {
VStack(alignment: .leading, spacing: 16) {
Text("4. Envie para o chat decidir os cortes").font(.title3.weight(.semibold))
Text("Esta é a única etapa manual que sobra: o julgamento de qual tomada usar, onde dar zoom e o que escrever na tela é feito pela IA numa conversa, não por um botão. Copie abaixo, cole numa sessão do Claude e peça pra rodar a skill \"editar-por-voz\".")
.font(.callout).foregroundStyle(.secondary)
if let voiceTimelinePath {
// Alternativa automática: em vez de copiar/colar no chat, manda a
// própria voice timeline (o arquivo inteiro) junto com o brief para
// o modelo local (Ollama/Gemma 3) decidir a edição de uma vez —
// cortes, zooms e textos numa única chamada, sem sair do app.
VStack(alignment: .leading, spacing: 8) {
Text("OU gere o roteiro por IA local (Ollama/Gemma 3)").font(.callout.weight(.semibold))
Text("O app envia a voice timeline completa (o arquivo) acompanhada do pedido para o modelo local decidir os cortes, zooms e textos de uma vez. Nada de copiar e colar.")
.font(.caption).foregroundStyle(.secondary)
HStack {
if ollamaModels.isEmpty {
TextField("Modelo (ex.: gemma3:12b, llama3)", text: $generateScriptModel)
.textFieldStyle(.roundedBorder)
.frame(maxWidth: 260)
} else {
Picker("Modelo", selection: $generateScriptModel) {
ForEach(ollamaModels, id: \.self) { m in
Text(m).tag(m)
}
}
.pickerStyle(.menu)
.frame(maxWidth: 260)
TextField("Ou outro", text: $generateScriptModel)
.textFieldStyle(.roundedBorder)
.frame(maxWidth: 120)
}
Button {
generateScript(voiceTimelinePath: voiceTimelinePath)
} label: {
if isGeneratingScript {
HStack { ProgressView().controlSize(.small); Text("Gerando…") }
} else {
Label("Gerar roteiro por IA local", systemImage: "sparkles")
}
}
.buttonStyle(.borderedProminent)
.disabled(isGeneratingScript || voiceTimelinePath.isEmpty)
}
if !generateScriptFeedback.isEmpty {
Label(generateScriptFeedback, systemImage: "checkmark.circle.fill")
.font(.caption).foregroundStyle(.green)
}
}
.padding(12)
.background(RoundedRectangle(cornerRadius: 8).fill(Color.green.opacity(0.07)))
.onAppear { fetchOllamaModels() }
Divider().padding(.vertical, 4)
Button {
copyForChat(path: voiceTimelinePath)
} label: {
Label("Copiar para colar no chat", systemImage: "doc.on.clipboard")
.frame(maxWidth: .infinity)
}
.buttonStyle(.borderedProminent)
.controlSize(.large)
if !copiedFeedback.isEmpty {
Label(copiedFeedback, systemImage: "checkmark.circle.fill")
.font(.caption).foregroundStyle(.green)
}
VStack(alignment: .leading, spacing: 8) {
Text("O que é copiado").font(.caption.weight(.semibold)).foregroundStyle(.secondary)
Text("Um pedido pronto + o conteúdo de \(URL(fileURLWithPath: voiceTimelinePath).lastPathComponent), já formatado. É só colar (⌘V) numa conversa com o Claude.")
.font(.caption).foregroundStyle(.secondary)
}
.padding(12)
.background(RoundedRectangle(cornerRadius: 8).fill(Color.secondary.opacity(0.06)))
Divider().padding(.vertical, 4)
Text("Cole aqui o que o chat devolveu").font(.callout.weight(.semibold))
Text("Na próxima etapa essas decisões aparecem já marcadas na timeline, frase por frase, para você lapidar.")
.font(.caption).foregroundStyle(.secondary)
HStack {
Button {
if let s = NSPasteboard.general.string(forType: .string) {
decisionsText = s
}
} label: {
Label("Colar da área de transferência", systemImage: "list.clipboard")
}
Spacer()
if !decisionsText.isEmpty {
Label(jsonIsValid ? "JSON válido" : "JSON inválido",
systemImage: jsonIsValid ? "checkmark.circle.fill" : "xmark.circle.fill")
.font(.caption)
.foregroundStyle(jsonIsValid ? .green : .red)
}
}
TextEditor(text: $decisionsText)
.font(.system(.caption, design: .monospaced))
.frame(minHeight: 140)
.padding(8)
.background(RoundedRectangle(cornerRadius: 8).fill(Color.secondary.opacity(0.06)))
.overlay(RoundedRectangle(cornerRadius: 8).stroke(Color.secondary.opacity(0.2)))
Button {
applyDecisions()
} label: {
if isApplyingDecisions {
HStack { ProgressView().controlSize(.small); Text("Aplicando…") }
.frame(maxWidth: .infinity)
} else {
Label("Aplicar decisões", systemImage: "checkmark.seal")
.frame(maxWidth: .infinity)
}
}
.buttonStyle(.borderedProminent)
.controlSize(.large)
.disabled(isApplyingDecisions || !jsonIsValid)
if let appliedPath {
Label("Decisões aplicadas — \(URL(fileURLWithPath: appliedPath).lastPathComponent)",
systemImage: "checkmark.circle.fill")
.font(.caption).foregroundStyle(.green)
}
Divider()
Button("Pular esta etapa (revisar as ênfases direto, sem passar pela IA)") {
skippedVoiceEdit = true
appliedPath = nil
decisionsText = ""
}
.buttonStyle(.plain)
.font(.caption)
.foregroundStyle(.secondary)
} else {
Label("Volte ao passo anterior e rode a análise de voz primeiro.", systemImage: "exclamationmark.triangle.fill")
.font(.caption).foregroundStyle(.orange)
}
}
}
/// Etapa 5 — a sala de edição. Diferente das outras, não é um formulário
/// dentro da coluna do assistente: ocupa a janela toda e se carrega sozinha
/// na primeira vez que aparece para aquela análise de voz.
private var revisarStep: some View {
Group {
if voiceTimelinePath != nil {
PhraseReviewView(model: reviewModel)
} else {
VStack(spacing: 8) {
Label("Volte ao passo 3 e rode a análise de voz primeiro.",
systemImage: "exclamationmark.triangle.fill")
.foregroundStyle(.orange)
}
.frame(maxWidth: .infinity, maxHeight: .infinity)
}
}
.onAppear { loadReviewIfNeeded() }
}
/// Processing and its result live on the SAME slide: the moment the last
/// operation finishes (`finalPath` gets set), the open/reveal buttons
/// appear right below the "Processar" button instead of gating behind a
/// separate "Concluído" step the user has to click into — there was
/// nothing on that slide worth a click of its own.
private var finalizarStep: some View {
VStack(alignment: .leading, spacing: 16) {
Text("6. Finalize o corte").font(.title3.weight(.semibold))
Text("Últimos passos automáticos, sem decisão envolvida — rodam com os parâmetros já configurados na aba \"Análise de Voz\" / \"Legendas Dinâmicas\".")
.font(.callout).foregroundStyle(.secondary)
Toggle("Remover silêncios do áudio", isOn: $finalSilences)
Toggle("Remover palavras de preenchimento", isOn: $finalFillers)
Toggle("Gerar legenda comum (texto editável no FCP)", isOn: $finalSubtitles)
Toggle("Gerar legendas dinâmicas (estilo configurado na aba própria)", isOn: $finalDynamicSubtitles)
Button {
finalizeProcessing()
} label: {
if isFinalizing {
HStack { ProgressView().controlSize(.small); Text(finalStatus.isEmpty ? "Processando…" : finalStatus) }
.frame(maxWidth: .infinity)
} else {
Label("Processar", systemImage: "play.fill").frame(maxWidth: .infinity)
}
}
.buttonStyle(.borderedProminent)
.controlSize(.large)
.disabled(isFinalizing || (!finalSilences && !finalFillers && !finalSubtitles && !finalDynamicSubtitles))
if let finalPath, !isFinalizing {
Divider().padding(.vertical, 4)
Label("Concluído", systemImage: "checkmark.seal.fill")
.font(.callout.weight(.semibold))
.foregroundStyle(.green)
Text(finalPath).font(.caption).foregroundStyle(.secondary).lineLimit(1).truncationMode(.middle)
HStack {
Button("Abrir no Final Cut Pro") { NSWorkspace.shared.open(URL(fileURLWithPath: finalPath)) }
.buttonStyle(.borderedProminent)
Button("Mostrar no Finder") {
NSWorkspace.shared.activateFileViewerSelecting([URL(fileURLWithPath: finalPath)])
}
Spacer()
Button("Começar outro projeto") { resetWizard() }
}
} else if !finalStatus.isEmpty && !isFinalizing {
Text(finalStatus).font(.caption).foregroundStyle(.secondary)
}
}
}
// MARK: - Navegação
/// `.finalizar` is the last step now — once it has a `finalPath`, the
/// slide's own "Começar outro projeto" button is the way forward, so the
/// footer's "Continuar" would be a second, redundant path to nowhere.
private var navFooter: some View {
HStack {
if step != .projeto {
Button("Voltar") { goBack() }
}
Spacer()
if step != .finalizar || finalPath == nil {
Button(step == .finalizar ? "Concluir" : "Continuar") { goNext() }
.buttonStyle(.borderedProminent)
.disabled(!canAdvance)
}
}
}
private var canAdvance: Bool {
switch step {
case .projeto: return outputFolder != nil && projectPath != nil
case .transcricao: return !transcribeResults.isEmpty
case .analise: return voiceTimelinePath != nil
case .exportarChat: return appliedPath != nil || skippedVoiceEdit
// Revisar é opcional: a sugestão da IA já é utilizável como veio, então
// o botão nunca trava aqui — o passo existe para lapidar, não para
// exigir mais uma confirmação.
case .revisar: return true
case .finalizar: return finalPath != nil && !isFinalizing
}
}
private func goNext() {
guard let next = WizardStep(rawValue: step.rawValue + 1) else { return }
// Sair da revisão grava o que foi decidido e as ações derivadas dela
// (`_phrase_actions.json`) ao lado da análise de voz — é esse arquivo
// que `finalizeProcessing` reaplica na etapa 6, para que desativar uma
// frase aqui realmente a remova do vídeo final, e não só do registro.
if step == .revisar {
reviewModel.save { reviewPath, actionsPath in
phraseReviewPath = reviewPath
phraseActionsPath = actionsPath
}
}
step = next
}
private func goBack() {
guard let prev = WizardStep(rawValue: step.rawValue - 1) else { return }
step = prev
}
private func resetWizard() {
step = .projeto
transcribeResults = []
voiceTimelinePath = nil
voiceAnalysisMessage = ""
decisionsText = ""
appliedPath = nil
skippedVoiceEdit = false
reviewLoadedFor = nil
reviewLoadedForDecisions = nil
phraseReviewPath = nil
phraseActionsPath = nil
finalStatus = ""
finalPath = nil
errorMessage = nil
}
// MARK: - Componentes auxiliares
@ViewBuilder
private func fieldRow(icon: String, label: String, isSet: Bool, action: @escaping () -> Void) -> some View {
HStack {
Image(systemName: icon).foregroundStyle(isSet ? .primary : .secondary)
Text(label).lineLimit(1).truncationMode(.middle).foregroundStyle(isSet ? .primary : .secondary)
Spacer()
Button("Escolher…", action: action)
}
.padding(12)
.background(RoundedRectangle(cornerRadius: 8).fill(Color.secondary.opacity(0.06)))
}
/// Todo output do fluxo carrega um destes sufixos no nome (ver
/// `_derived_output` / suffixes usados por `apply_voice_actions`,
/// `remove_silences`, `generate_dynamic_subtitles` em
/// `admin/models_api.py`). Selecionar um deles como "o projeto" no passo
/// 1 é o erro que gerou arquivos como `_voice_edit_voice_edit_...`: os
/// cortes de voz assumem timestamps da mídia ORIGINAL, então reaplicá-los
/// sobre um arquivo já cortado desloca tudo silenciosamente.
private static let generatedSuffixes = [
"_voice_edit", "_silence_removed", "_dynamic_subtitles",
"_transcript_edit", "_fillers_removed", "_markers",
]
private func looksLikeGeneratedFile(_ path: String?) -> Bool {
guard let path else { return false }
let stem = URL(fileURLWithPath: path).deletingPathExtension().lastPathComponent
return Self.generatedSuffixes.contains { stem.contains($0) }
}
private var jsonIsValid: Bool {
guard let data = decisionsText.data(using: .utf8), !decisionsText.trimmingCharacters(in: .whitespacesAndNewlines).isEmpty else { return false }
return (try? JSONSerialization.jsonObject(with: data)) != nil
}
// MARK: - Ações — Python bridge
private func loadProjectConfig() {
PythonBridge.call(command: "project_config") { result, _ in
DispatchQueue.main.async {
guard let result, result["ok"] as? Bool == true else { return }
if let folder = result["folder"] as? String, !folder.isEmpty { outputFolder = folder }
if let file = result["file"] as? String, !file.isEmpty { projectPath = file }
}
}
}
private func loadCatalog() async {
PythonBridge.call(command: "catalog") { result, _ in
DispatchQueue.main.async {
if let result { catalog = Catalog(json: result) }
}
}
}
private func pickOutputFolder() {
let panel = NSOpenPanel()
panel.canChooseFiles = false
panel.canChooseDirectories = true
panel.allowsMultipleSelection = false
panel.prompt = "Usar esta pasta"
panel.message = "Escolha a pasta onde os resultados serão salvos."
if panel.runModal() == .OK, let url = panel.url {
outputFolder = url.path
PythonBridge.call(command: "set_project_config", arguments: ["folder": url.path]) { _, _ in }
}
}
private func pickProjectFile() {
let panel = NSOpenPanel()
panel.canChooseFiles = true
panel.canChooseDirectories = false
panel.allowsMultipleSelection = false
panel.prompt = "Selecionar"
panel.message = "Selecione o arquivo (.fcpxml) ou o bundle (.fcpxmld) exportado pelo Final Cut Pro."
if panel.runModal() == .OK, let url = panel.url {
let ext = url.pathExtension.lowercased()
if ext == "fcpxml" || ext == "fcpxmld" || ext == "xml" {
projectPath = url.path
PythonBridge.call(command: "set_project_config", arguments: ["file": url.path]) { _, _ in }
} else {
errorMessage = "Selecione um arquivo .fcpxml, .fcpxmld ou .xml do Final Cut Pro."
}
}
}
private func startTranscription() {
guard let projectPath, let outputFolder else { return }
isTranscribing = true
errorMessage = nil
transcribeResults = []
transcribeProgress = 0
PythonBridge.run(command: "transcribe", arguments: ["path": projectPath, "output_dir": outputFolder]) { obj in
DispatchQueue.main.async {
let type = obj["type"] as? String
if type == "progress" {
transcribeProgress = (obj["fraction"] as? NSNumber)?.doubleValue ?? 0
transcribeStage = obj["stage"] as? String ?? ""
} else if type == "error" {
errorMessage = obj["message"] as? String ?? "Erro na transcrição."
} else if type == "result", let arr = obj["transcripts"] as? [[String: Any]] {
transcribeResults = arr.map(TranscriptResult.init)
}
}
} completion: { code, err in
DispatchQueue.main.async {
isTranscribing = false
transcribeProgress = 1
if code != 0 && transcribeResults.isEmpty {
errorMessage = err ?? "A transcrição falhou."
}
}
}
}
private func analyzeVoice(forceReprocess: Bool = false) {
guard let projectPath, let outputFolder else { return }
isAnalyzing = true
errorMessage = nil
PythonBridge.call(command: "analyze_voice", arguments: [
"path": projectPath,
"output_dir": outputFolder,
"force_reprocess": forceReprocess,
]) { result, err in
DispatchQueue.main.async {
isAnalyzing = false
guard result?["ok"] as? Bool == true else {
errorMessage = result?["error"] as? String ?? err ?? "Falha ao analisar a voz."
return
}
if result?["reused"] as? Bool == true, !forceReprocess {
let timelines = result?["timelines"] as? [String] ?? []
existingVoiceTimelinePath = timelines.first ?? extractPath(from: result?["message"] as? String ?? "", marker: "**Timeline JSON**:")
showVoiceTimelineReuseAlert = true
return
}
let message = result?["message"] as? String ?? ""
voiceAnalysisMessage = message
if let path = extractPath(from: message, marker: "**Timeline JSON**:") {
voiceTimelinePath = path
} else {
voiceTimelinePath = nil
// ok:true não garante que a análise gerou timeline — se
// não houver fala detectável no áudio, o Python volta com
// sucesso mas sem "Timeline JSON" na mensagem. Sem isso
// aqui, a etapa parecia não fazer nada.
errorMessage = "A análise terminou mas não encontrou fala reconhecível no áudio. Mensagem do motor: " + (message.isEmpty ? "(vazia)" : message)
}
checkAcoustics()
}
}
}
/// A ênfase de voz (energia/tom) depende do `librosa`, dependência
/// opcional. Sem ela, a análise ainda transcreve e corta pelo texto,
/// mas nunca deveria propor zoom — por isso avisamos aqui, no ponto
/// onde o usuário sentiria falta, em vez de só na aba Modelos.
private func checkAcoustics() {
PythonBridge.call(command: "acoustics_capability") { result, _ in
DispatchQueue.main.async {
guard let result, result["ok"] as? Bool == true else { return }
acousticsAvailable = result["available"] as? Bool
}
}
}
/// Localiza uma linha markdown do tipo "- **Marker**: valor" (usado nas
/// mensagens do bridge Python) e devolve o valor. Aceita o marcador de
/// lista "- " opcional antes dos asteriscos.
private func extractPath(from message: String, marker: String) -> String? {
for line in message.split(separator: "\n") {
var trimmed = Substring(line.trimmingCharacters(in: .whitespaces))
if trimmed.hasPrefix("- ") { trimmed = trimmed.dropFirst(2) }
if trimmed.hasPrefix(marker) {
return trimmed.dropFirst(marker.count).trimmingCharacters(in: .whitespaces)
}
}
return nil
}
private func copyForChat(path: String) {
guard let content = try? String(contentsOfFile: path, encoding: .utf8) else {
errorMessage = "Não foi possível ler \(path)."
return
}
let prompt = """
Use a skill "editar-por-voz" para decidir os cortes deste projeto a partir da timeline de voz abaixo. Devolva só o JSON de decisões (cortes, zooms, textos, marcadores) pronto para eu colar de volta no app.
```json
\(content)
```
"""
let pasteboard = NSPasteboard.general
pasteboard.clearContents()
pasteboard.setString(prompt, forType: .string)
copiedFeedback = "Copiado — cole (⌘V) numa conversa com o Claude."
}
/// Monta a revisão uma vez por análise de voz. Voltar e avançar de novo com
/// as MESMAS decisões não recarrega: isso jogaria fora as edições manuais
/// em silêncio, que é exatamente o que esta tela existe para preservar.
///
/// Mas se o usuário voltou à etapa 4 e colou/gerou um JSON de decisões
/// DIFERENTE do que gerou a revisão atual, isso é recarregado — e com
/// `fresh: true`, para que o `active`/ênfase recém-derivado dessas
/// decisões novas não seja imediatamente sobrescrito pela revisão salva
/// da visita anterior (`merge_saved_decisions`, do lado Python). Sem isso,
/// a tela ficava presa nas decisões antigas mesmo depois de reaplicar o
/// corte — a dessincronia relatada entre "ativa aqui" e "já cortado no
/// FCPXML".
private func loadReviewIfNeeded() {
guard let voiceTimelinePath else { return }
let decisionsChanged = reviewLoadedForDecisions != nil && reviewLoadedForDecisions != decisionsText
guard reviewLoadedFor != voiceTimelinePath || decisionsChanged else { return }
reviewLoadedFor = voiceTimelinePath
reviewLoadedForDecisions = decisionsText
// A pasta do projeto e a do .fcpxml entram como onde procurar a mídia:
// a análise de voz guarda só o nome do arquivo, não o caminho.
reviewModel.load(
voiceTimelinePath: voiceTimelinePath,
decisionsJSON: decisionsText,
outputFolder: outputFolder,
mediaFolder: projectPath.map { URL(fileURLWithPath: $0).deletingLastPathComponent().path },
fresh: decisionsChanged
)
}
private func applyDecisions() {
guard let projectPath, let outputFolder,
let data = decisionsText.data(using: .utf8),
let parsed = try? JSONSerialization.jsonObject(with: data) else { return }
isApplyingDecisions = true
errorMessage = nil
PythonBridge.call(command: "apply_voice_actions", arguments: [
"path": projectPath,
"output_dir": outputFolder,
"actions": parsed,
]) { result, err in
DispatchQueue.main.async {
isApplyingDecisions = false
guard result?["ok"] as? Bool == true else {
errorMessage = result?["error"] as? String ?? err ?? "Falha ao aplicar as decisões."
return
}
appliedPath = result?["path"] as? String ?? projectPath
skippedVoiceEdit = false
}
}
}
/// Etapa 4 (alternativa): manda a voice timeline inteira para um modelo
/// local (Ollama/Gemma 3) que dirige a edição de uma vez — sem copiar e
/// colar. O motor devolve o roteiro legível + o JSON de ações e já aplica
/// no FCPXML (non-destructive), igual ao fluxo manual "Aplicar decisões".
private func fetchOllamaModels() {
guard ollamaModels.isEmpty else { return }
PythonBridge.call(command: "list_ollama_models", arguments: [:]) { result, err in
DispatchQueue.main.async {
if let models = result?["models"] as? [String], !models.isEmpty {
ollamaModels = models
if !models.contains(generateScriptModel) {
generateScriptModel = models.first ?? generateScriptModel
}
}
}
}
}
private func generateScript(voiceTimelinePath: String) {
guard let projectPath, let outputFolder else { return }
isGeneratingScript = true
generateScriptFeedback = ""
errorMessage = nil
PythonBridge.call(command: "generate_voice_script", arguments: [
"voice_timeline": voiceTimelinePath,
"filepath": projectPath,
"output_dir": outputFolder,
"model": generateScriptModel,
"apply_to_fcpxml": true,
]) { result, err in
DispatchQueue.main.async {
isGeneratingScript = false
guard result?["ok"] as? Bool == true else {
errorMessage = result?["error"] as? String ?? err ?? "Falha ao gerar roteiro por IA local."
return
}
// Traz as decisões de volta para a tela de revisão (etapa 5) e
// marca como aplicadas, exatamente como o "Aplicar decisões".
if let actionsPath = result?["actions_path"] as? String,
let content = try? String(contentsOfFile: actionsPath, encoding: .utf8) {
decisionsText = content
}
appliedPath = result?["applied_path"] as? String ?? projectPath
skippedVoiceEdit = false
generateScriptFeedback = "Roteiro gerado e aplicado — revise as ênfases na próxima etapa."
}
}
}
private func finalizeProcessing() {
guard let outputFolder else { return }
let startPath = appliedPath ?? projectPath
guard let startPath else { return }
var operations: [String] = []
if finalSilences { operations.append("remove_silences") }
if finalFillers { operations.append("remove_filler_words") }
if finalSubtitles { operations.append("generate_plain_subtitles") }
if finalDynamicSubtitles { operations.append("generate_dynamic_subtitles") }
guard !operations.isEmpty else { return }
isFinalizing = true
errorMessage = nil
finalStatus = "Iniciando…"
applyReviewDecisions(startPath: startPath, outputFolder: outputFolder) { reviewedPath in
finalizeStep(operations, index: 0, currentPath: reviewedPath, outputFolder: outputFolder)
}
}
/// Reapplies whatever the etapa-5 review decided (active/inactive
/// phrases, manual zooms) on top of `startPath` before the finishing
/// chain runs below. Without this, `appliedPath` stayed frozen at
/// whatever `exportarChat`'s `apply_voice_actions` produced BEFORE the
/// human review — so toggling a phrase off in the review only updated
/// `_phrase_actions.json` on disk, never the video the wizard actually
/// exports. A no-op (just hands `startPath` straight through) when the
/// review step was never visited/saved this session.
private func applyReviewDecisions(
startPath: String, outputFolder: String, completion: @escaping (String) -> Void
) {
guard let phraseActionsPath,
let data = try? Data(contentsOf: URL(fileURLWithPath: phraseActionsPath)),
let parsed = try? JSONSerialization.jsonObject(with: data) as? [String: Any],
let actions = parsed["actions"] else {
completion(startPath)
return
}
finalStatus = "Aplicando a revisão…"
PythonBridge.call(command: "apply_voice_actions", arguments: [
"path": startPath,
"output_dir": outputFolder,
"actions": actions,
]) { result, err in
DispatchQueue.main.async {
guard result?["ok"] as? Bool == true else {
isFinalizing = false
errorMessage = result?["error"] as? String ?? err ?? "Falha ao aplicar a revisão."
finalStatus = "Processamento interrompido."
return
}
completion(result?["path"] as? String ?? startPath)
}
}
}
private func finalizeStep(_ operations: [String], index: Int, currentPath: String, outputFolder: String) {
guard index < operations.count else {
isFinalizing = false
finalStatus = "Processamento concluído."
finalPath = currentPath
return
}
let operation = operations[index]
finalStatus = "Processando: \(operation)…"
PythonBridge.call(command: operation, arguments: ["path": currentPath, "output_dir": outputFolder]) { result, err in
DispatchQueue.main.async {
guard result?["ok"] as? Bool == true else {
isFinalizing = false
errorMessage = result?["error"] as? String ?? err ?? "Falha em \(operation)."
finalStatus = "Processamento interrompido."
return
}
let nextPath = result?["path"] as? String ?? currentPath
finalizeStep(operations, index: index + 1, currentPath: nextPath, outputFolder: outputFolder)
}
}
}
}
Submodule code/WHISPERX deleted from c9ed3cc6bd
+131
View File
@@ -0,0 +1,131 @@
#!/usr/bin/env python3
"""AI Voice Editor - Pipeline completo: transcrição + análise acústica → JSON para IA.
Uso:
python ai_edit.py <media_path> [--model base] [--lang pt] [--no-diarize] [--output dir]
Gera dois arquivos na pasta output (ou ao lado do mídia):
<nome>_transcript.json — transcrição com timestamps por palavra
<nome>_voice_timeline.json — timeline de voz com ênfase, pitch, energy, speakers
Esses arquivos são a ENTRADA para a IA analisar e gerar o roteiro/edição.
"""
import argparse
import json
import sys
import time
from pathlib import Path
def main():
parser = argparse.ArgumentParser(description="Pipeline de análise de voz para IA")
parser.add_argument("media", help="Caminho do arquivo de mídia (.mp4, .mov, .wav, etc.)")
parser.add_argument("--model", default="base", help="Modelo Whisper (tiny/base/small/medium/large-v3)")
parser.add_argument("--lang", default=None, help="Idioma (ex: pt, en). Auto-detect se omitido")
parser.add_argument("--hf-token", default=None, help="HuggingFace token para diarização (opcional)")
parser.add_argument("--no-diarize", action="store_true", help="Pular diarização de falantes")
parser.add_argument("--output", default=None, help="Pasta de saída (padrão: ao lado do mídia)")
parser.add_argument("--no-align", action="store_true", help="Pular alinhamento fonético (whisperx)")
args = parser.parse_args()
media_path = Path(args.media).resolve()
if not media_path.is_file():
print(f"ERRO: Arquivo não encontrado: {media_path}", file=sys.stderr)
sys.exit(1)
# Output dir
out_dir = Path(args.output) if args.output else media_path.parent
out_dir.mkdir(parents=True, exist_ok=True)
stem = media_path.stem
# ── Fase 1: Transcrição ──────────────────────────────────────────
print(f"[1/2] Transcrevendo {media_path.name} (modelo: {args.model})...")
t0 = time.time()
# Adiciona code/ ao path para imports do projeto
code_dir = Path(__file__).resolve().parent
sys.path.insert(0, str(code_dir))
from fcpxml.transcribe import transcribe
def transcribe_progress(pct):
bar_len = 30
filled = int(bar_len * pct)
bar = "█" * filled + "░" * (bar_len - filled)
print(f"\r [{bar}] {pct*100:.0f}%", end="", flush=True)
transcript = transcribe(
str(media_path),
model_size=args.model,
language=args.lang,
progress_cb=transcribe_progress,
align=not args.no_align,
)
print() # newline after progress bar
if transcript is None:
print("ERRO: Transcrição falhou. Verifique se faster-whisper está instalado:", file=sys.stderr)
print(" uv pip install faster-whisper", file=sys.stderr)
sys.exit(1)
print(f" → {len(transcript.get('words', []))} palavras, "
f"{len(transcript.get('segments', []))} segmentos, "
f"idioma: {transcript.get('language', '?')}")
# Salva transcrição
transcript_path = out_dir / f"{stem}_transcript.json"
transcript_path.write_text(json.dumps(transcript, indent=2, ensure_ascii=False), encoding="utf-8")
print(f" → Salvo: {transcript_path}")
# ── Fase 2: Análise de voz (timeline) ────────────────────────────
print(f"\n[2/2] Analisando voz (pitch, energia, ênfase)...")
t1 = time.time()
from fcpxml.voice_timeline import build_voice_timeline
def voice_progress(fraction, stage):
print(f"\r {stage} ({fraction*100:.0f}%)", end="", flush=True)
hf_token = None if args.no_diarize else args.hf_token
timeline = build_voice_timeline(
str(media_path),
transcript,
hf_token=hf_token,
progress_cb=voice_progress,
)
print()
# Salva voice timeline
timeline_path = out_dir / f"{stem}_voice_timeline.json"
timeline_path.write_text(json.dumps(timeline, indent=2, ensure_ascii=False), encoding="utf-8")
print(f" → Salvo: {timeline_path}")
# ── Resumo ───────────────────────────────────────────────────────
elapsed = time.time() - t0
summary = timeline.get("summary", {})
layers = timeline.get("layers", {})
n_words = len(transcript.get("words", []))
n_segments = len(transcript.get("segments", []))
n_speakers = len(timeline.get("speakers", []))
duration = transcript.get("duration", 0)
print(f"\n{'='*50}")
print(f" ARQUIVOS GERADOS:")
print(f" {transcript_path}")
print(f" {timeline_path}")
print(f"\n RESUMO:")
print(f" Duração: {duration:.1f}s ({duration/60:.1f}min)")
print(f" Palavras: {n_words}")
print(f" Segmentos: {n_segments}")
print(f" Falantes: {n_speakers}")
print(f" Camadas: transcript={layers.get('transcript')}, "
f"acoustics={layers.get('acoustics')}, "
f"diarization={layers.get('diarization')}")
print(f" Tempo: {elapsed:.1f}s")
print(f"{'='*50}")
print(f"\n→ Pronto! Agora peça à IA para analisar o voice timeline e gerar o roteiro.")
if __name__ == "__main__":
main()
+181
View File
@@ -0,0 +1,181 @@
"""Forced alignment — refine word timestamps against an acoustic model.
Why this exists
--------------
faster-whisper derives word times by cross-attention, which lands every word
*start* systematically ~0.3-0.5s early (the word-end is fine). That bias flows
straight into the voice timeline and makes zoom/cut land on the wrong frame —
measured on real footage in ``Engine/docs/05_EXPERIENCIAS.md`` (#14). Phonetic
forced alignment (wav2vec2, via whisperx) re-anchors each word against the
audio and brings that error down to ~30ms.
Design
------
* The dependency (``whisperx``) is **optional** and imported lazily, exactly
like the rest of this stack (librosa, faster-whisper). When it is missing, or
any step fails, :meth:`ForcedAligner.align` returns the words unchanged, so
transcription never breaks because alignment did.
* The aligner is a single responsibility class: it knows how to turn a
transcript into the shape whisperx wants, call it, and write the refined
times back. ``transcribe.py`` owns the decision of *whether* to align.
* Align models are cached per language on the instance so repeated calls
(e.g. many short clips) don't reload the wav2vec2 weights each time.
"""
import logging
from typing import List, Optional, Sequence
logger = logging.getLogger(__name__)
class ForcedAligner:
"""Refine word-level timestamps with whisperx phonetic forced alignment.
Usage::
aligner = ForcedAligner()
words = aligner.align(words, raw_segments, media_path, language, models_dir)
``words`` and ``raw_segments`` come straight from :func:`transcribe` —
``raw_segments`` carries the per-segment ``words`` lists (the same dict
objects as in ``words``) so the aligner knows which words belong to which
audio window. Returns a list of the *same* word dicts, with ``start``/``end``
overwritten in place where alignment produced a usable time.
"""
def __init__(self, device: Optional[str] = None):
self._device = device
self._models: dict = {}
# -- capability ------------------------------------------------------
@staticmethod
def available() -> bool:
"""Whether whisperx can be imported (the aligner can run at all)."""
try:
import whisperx # noqa: F401
except Exception:
return False
return True
def _resolve_device(self) -> str:
if self._device:
return self._device
try:
import torch
if torch.cuda.is_available():
return "cuda"
except Exception:
pass
return "cpu"
# -- public API ------------------------------------------------------
def align(
self,
words: Sequence[dict],
raw_segments: Sequence[dict],
audio_path: str,
language: str,
models_dir: Optional[str] = None,
) -> List[dict]:
"""Return ``words`` with forced-aligned timestamps where possible.
Falls back to the unchanged ``words`` on any failure (missing
dependency, model load error, audio read error, or a result that
doesn't line up with the input).
"""
if not words or not language:
return list(words)
try:
import whisperx
except Exception:
logger.info("whisperx not installed; skipping forced alignment")
return list(words)
try:
device = self._resolve_device()
align_input = self._build_align_input(words, raw_segments)
audio = whisperx.load_audio(audio_path)
if language not in self._models:
align_model, metadata = whisperx.load_align_model(
language_code=language,
device=device,
model_dir=str(models_dir) if models_dir else None,
)
self._models[language] = (align_model, metadata)
align_model, metadata = self._models[language]
result = whisperx.align(
align_input,
align_model,
metadata,
audio,
device,
return_char_alignments=False,
)
return self._merge_result(words, result.get("segments", []))
except Exception:
logger.warning(
"forced alignment failed for %s; using raw timestamps", audio_path
)
return list(words)
# -- internals -------------------------------------------------------
@staticmethod
def _build_align_input(
words: Sequence[dict], raw_segments: Sequence[dict]
) -> List[dict]:
"""Transcript in whisperx's expected shape: segments -> words.
whisperx.align requires each segment to carry ``text``/``start``/``end``
and a ``words`` list whose entries have ``word``/``start``/``end``/``score``.
We only read ``words`` from ``raw_segments`` (the flattened ``words``
list is the source of truth for counts), so the two stay consistent.
"""
align_segments: List[dict] = []
for seg in raw_segments:
seg_words = [
{
"word": w.get("word", ""),
"start": float(w.get("start", 0.0)),
"end": float(w.get("end", 0.0)),
"score": float(w.get("confidence", 0.0)),
}
for w in seg.get("words", [])
]
align_segments.append(
{
"text": (seg.get("text") or "").strip(),
"start": float(seg.get("start", 0.0)),
"end": float(seg.get("end", 0.0)),
"words": seg_words,
}
)
return align_segments
@staticmethod
def _merge_result(words: Sequence[dict], aligned_segments: Sequence[dict]) -> List[dict]:
"""Walk the aligned output in order and overwrite word times in place.
whisperx preserves word order within and across segments, so a single
running index over the output words lines up with ``words``. A word the
aligner failed to place gets ``None``/``0`` times — we skip those rather
than clobber a good timestamp, and if counts ever diverge we stop and
leave the rest untouched.
"""
out = list(words)
wi = 0
for seg in aligned_segments:
for aw in seg.get("words", []):
if wi >= len(out):
return out
start = aw.get("start")
end = aw.get("end")
if start is None or end is None or end < start:
wi += 1
continue
out[wi]["start"] = float(start)
out[wi]["end"] = float(end)
wi += 1
return out
+312
View File
@@ -0,0 +1,312 @@
"""Local LLM integration — the voice timeline meets a local model.
The voice timeline is *designed* to be handed to a language model: it is the
source of truth between speech analysis and editing, layered so a model can
reason about the narrative without parsing FCPXML. This module is the client
side of that contract. It formats the timeline into the editar-por-voz brief,
calls a local model server (Ollama, running Gemma 3 / Llama locally), and
parses the model's decisions back into a validated list of VoiceActions —
all inside the engine, so there is no wizard, no copy-paste, no manual step.
Transport: Ollama's HTTP chat API at ``http://localhost:11434/api/chat``.
Any model Ollama serves works; the default is Gemma 3 because that is what
runs locally here ("Lama com Gema 3"), but pass ``model=`` to switch.
The model is untrusted input: its JSON is validated row-by-row by
:func:`fcpxml.voice_actions.parse_actions`, so one malformed decision never
discards the edit. The brief is written so the model only ever emits the four
action kinds the applier understands.
"""
from __future__ import annotations
import json
import logging
import re
from typing import Any, Dict, Optional, Sequence, Tuple
import httpx
from .voice_actions import parse_actions
logger = logging.getLogger(__name__)
DEFAULT_BASE_URL = "http://localhost:11434"
# Gemma 3 12B reliably follows the editar-por-voz brief (keep the script, cut
# only backstage chatter; the 4B variant skips the "keep the main content"
# rule and deletes the script) but doesn't fit an 8GB machine. Qwen2.5 7B
# instruct (q4_K_M) is the fallback for constrained hardware — strong at
# strict JSON-schema following, the property this brief leans on hardest.
# Pass ``model=`` to switch to whatever Ollama serves.
DEFAULT_MODEL = "qwen2.5:7b-instruct-q4_K_M"
REQUEST_TIMEOUT = 600.0
# The brief. Ported from the editar-por-voz skill criteria (criterios/01..08),
# condensed into the instructions a model needs to emit valid actions. Kept in
# Portuguese because the decisions and their reasons are read by a human editor.
_SYSTEM_PROMPT = """Você é o editor de vídeo por voz deste sistema. Recebe um JSON de "linha do tempo de voz" — a medição de COMO foi falado (ênfase, energia, pausa, falante) de uma gravação — e devolve as DECISÕES de edição em JSON, nada mais. Você nunca escreve XML.
Regras (siga rigorosamente):
1. LEIA EM CAMADAS. "summary" dá o formato da peça; "segments" é onde você trabalha (cada fala com seu texto e agregados); "segments[].words" dá o instante exato de cada destaque. Não recalcule energia, tom ou ênfase — use os números do JSON.
2. SEPARAR ROTEIRO DE BASTIDOR.
- ROTEIRO = o conteúdo principal que a pessoa quer entregar: explicação, depoimento, roteiro decorado, a mensagem. É isso que VAI FICAR.
- BASTIDOR = papo casual de gravação, cumprimentos, conversa com a equipe ("cara, beleza?", "tá gravando?", "deixa eu ver o celular"), piadas fora do assunto, tomadas interrompidas ou repetidas. É isso que VIRA "cut".
Exemplo: num vídeo sobre mastopexia, a explicação da cirurgia É o roteiro (mantém); o "tá gravando? pois é" antes dela É bastidor (corta).
Use "gap_before" e "take_boundary" (silêncio > ~3s = a câmera parou/recomeçou) para agrupar tomadas — eles marcam ONDE a tomada recomeça, não o que cortar. Nunca corte o conteúdo principal só porque tem ênfase; corte o casual/off-topic.
REGRAS DE OURO:
- MANTENHA o conteúdo principal (explicação, depoimento, roteiro decorado). Ele É o vídeo.
- CORTE SÓ o casual/off-topic: cumprimentos, "tá gravando?", papo com a equipe, olhar o celular, repetições de tomada.
- Em dúvida, MANTENHA a fala. É melhor sobrar conteúdo do que cortar o que era pra ficar.
3. ESCOLHER A MELHOR TOMADA de cada frase quando há repetições: mantenha a mais limpa e corte as outras (cut cobrindo a frase inteira).
4. CORTE (kind "cut"): para REMOVER uma frase, cubra ela inteira (start..end = início..fim da frase). Para APARAR só uma hesitação no começo ou fim, corte só da borda até a palavra (corte de meia frase é ambíguo — passe de 60% e apaga a linha toda). Nunca corte o silêncio entre falas. Quando a borda do corte encosta em fala mantida (não em silêncio puro), recue ~0,15-0,25s para dentro do corte nos dois lados — start ~0,2s DEPOIS do fim real da última palavra mantida, end ~0,2s ANTES do início real da próxima palavra mantida — senão o corte soa seco, engolindo a palavra antes de terminar de soar. Isso vale também pro início/fim do vídeo (ar morto antes da primeira palavra e depois da última).
5. ZOOM (kind "zoom"): só em palavra de CONTEÚDO bem enfatizada (emphasis alto, não artigo). params.scale entre 1.0 e 3.0 (padrão 1.3 se omitido). Posicione em torno da palavra, segurando até o fim da frase.
6. TEXTO (kind "text"): params.content obrigatório (≤120 chars), fixa um termo central ou callout. MARKER (kind "marker"): opcional params.content vira o nome do marcador. Use para emendas/junções que o editor deve conferir.
7. TEMPOS em segundos da MÍDIA ORIGINAL (exatamente como no JSON). Nunca compense para "depois do corte" — o programa desloca sozinho. end sempre > start, ambos ≥ 0.
8. reason OBRIGATÓRIO em cada ação, em português, embasando a decisão (ex.: 'abertura: "Aquela mama" (ênfase 0.42)'). reason vazio é decisão sem critério.
Responda APENAS com um objeto JSON válido, sem markdown, sem comentário:
{"source": "<nome do arquivo>", "actions": [{"kind": "cut|zoom|text|marker", "start": <float>, "end": <float>, "params": {}, "reason": "<pt>", "speaker": "<id>"}]}
"""
_OUTPUT_REMINDER = """Gere as decisões de edição conforme o brief. Responda SOMENTE o JSON:
{"source": "<nome do arquivo>", "actions": [{"kind": "cut|zoom|text|marker", "start": <float>, "end": <float>, "params": {}, "reason": "<pt>", "speaker": "<id>"}]}
Não inclua explicações nem blocos markdown."""
def ollama_chat(
model: str = DEFAULT_MODEL,
messages: Optional[Sequence[Dict[str, str]]] = None,
base_url: str = DEFAULT_BASE_URL,
temperature: float = 0.2,
timeout: float = REQUEST_TIMEOUT,
num_ctx: int = 32768,
) -> str:
"""One chat completion from a local Ollama server.
Returns the assistant message content. Raises on transport/HTTP errors so
the caller can decide whether to retry or report — a model call is the
one I/O in this pipeline that can legitimately fail mid-run.
"""
payload = {
"model": model,
"messages": list(messages or []),
"stream": False,
"options": {"temperature": temperature, "num_ctx": num_ctx},
}
try:
response = httpx.post(
f"{base_url.rstrip('/')}/api/chat", json=payload, timeout=timeout
)
response.raise_for_status()
data = response.json()
except Exception as exc:
# Covers transport errors AND a dropped connection that yields an empty
# body (httpx/JSONDecodeError) — both must become a RuntimeError so the
# caller reports the failure instead of crashing the whole pipeline.
raise RuntimeError(f"Falha ao falar com o modelo local em {base_url}: {exc}") from exc
return (data.get("message") or {}).get("content", "") or ""
def list_ollama_models(base_url: str = DEFAULT_BASE_URL) -> list[str]:
"""Names of the models Ollama currently serves, for a model picker.
Returns an empty list when Ollama is unreachable so the UI can fall back to
a free-text field instead of erroring.
"""
try:
resp = httpx.get(f"{base_url.rstrip('/')}/api/tags", timeout=10.0)
resp.raise_for_status()
models = resp.json().get("models", [])
names = [m.get("name") for m in models if m.get("name")]
return sorted(names)
except Exception:
return []
def _extract_json(text: str) -> Any:
"""Pull a JSON value out of a model response, tolerating fences/wrappers."""
if not text:
return None
candidate = text.strip()
# Strip a ```json ... ``` (or bare ```) fence if the model added one.
fence = re.search(r"```(?:json)?\s*(.*?)\s*```", candidate, re.DOTALL)
if fence:
candidate = fence.group(1).strip()
# Otherwise take the outermost {...} / [...].
if not candidate.startswith(("{" if True else "", "[")):
start = min(
(i for i, c in enumerate(candidate) if c in "{["),
default=None,
)
end = max(
(i for i, c in enumerate(candidate) if c in "}"),
default=None,
)
if start is not None and end is not None and end > start:
candidate = candidate[start : end + 1]
try:
data = json.loads(candidate)
except json.JSONDecodeError:
return None
# Models sometimes wrap the expected `{"source", "actions"}` object inside a
# single-element list (`[{...}]`). Unwrap that so the actions aren't treated
# as one malformed row.
if (
isinstance(data, list)
and len(data) == 1
and isinstance(data[0], dict)
and "actions" in data[0] # the wrapper carries the actions key
):
data = data[0]
return data
# Only these fields reach the model — the raw timeline also carries heavy
# per-word audio features (energy, pitch, arousal...) and speaker `samples`
# that blow past the model's context window on any real recording. Dropping
# them is what keeps a 3-minute timeline inside `num_ctx`.
_SEGMENT_KEEP = (
"start", "end", "speaker", "text", "gap_before", "take_boundary",
"avg_energy", "peak_emphasis", "emotion", "emotion_confidence",
"arousal", "valence",
)
_WORD_KEEP = ("text", "start", "end", "speaker", "emphasis", "pause_before")
_SPEAKER_KEEP = ("id", "name")
_SKIP_ROOT = ("layers", "scales")
def _project_timeline(timeline: dict) -> dict:
"""Strip the timeline down to what the edit decision actually needs."""
out = {k: v for k, v in timeline.items() if k not in _SKIP_ROOT}
speakers = [
{k: sp[k] for k in _SPEAKER_KEEP if k in sp}
for sp in timeline.get("speakers", [])
]
if speakers:
out["speakers"] = speakers
segs = []
for seg in timeline.get("segments", []):
s = {k: seg[k] for k in _SEGMENT_KEEP if k in seg}
s["words"] = [
{k: w[k] for k in _WORD_KEEP if k in w}
for w in seg.get("words", [])
]
segs.append(s)
out["segments"] = segs
return out
def _shrink_to_fit(compact: dict, max_chars: int) -> dict:
"""Drop word detail from the lowest-emphasis segments until it fits."""
segs = [dict(s) for s in compact.get("segments", [])]
while True:
payload = json.dumps(
{**compact, "segments": segs}, ensure_ascii=False, indent=1
)
if len(payload) <= max_chars or not any(s.get("words") for s in segs):
break
idx = min(
(i for i, s in enumerate(segs) if s.get("words")),
key=lambda i: float(segs[i].get("peak_emphasis", 0.0)),
)
segs[idx] = {**segs[idx], "words": []}
compact = dict(compact)
compact["segments"] = segs
return compact
def build_edit_messages(
timeline: dict, max_words_per_segment: int = 200, max_chars: int = 110000
) -> Tuple[str, str]:
"""The (system, user) pair that sends a timeline to the model.
The user turn carries a *projected* timeline (see :func:`_project_timeline`)
— text, timing, speaker and emphasis only — so a real recording fits in the
model's context window. Very long segments still have their word detail
capped to ``max_words_per_segment`` (most emphatic + boundaries), and if the
whole payload would still exceed ``max_chars`` the lowest-emphasis segments
lose their words until it fits, so we never blow ``num_ctx``.
"""
compact = _project_timeline(timeline)
if max_words_per_segment:
segs = []
for seg in compact["segments"]:
words = seg.get("words", [])
if len(words) > max_words_per_segment:
ranked = sorted(
enumerate(words),
key=lambda kv: float(kv[1].get("emphasis", 0.0)),
reverse=True,
)[: max_words_per_segment - 2]
keep = sorted({0, len(words) - 1} | {i for i, _ in ranked})
seg = {**seg, "words": [words[i] for i in keep]}
segs.append(seg)
compact["segments"] = segs
payload = json.dumps(compact, ensure_ascii=False, indent=1)
if len(payload) > max_chars:
compact = _shrink_to_fit(compact, max_chars)
payload = json.dumps(compact, ensure_ascii=False, indent=1)
user = (
"Linha do tempo de voz (JSON):\n\n"
+ payload
+ "\n\n"
+ _OUTPUT_REMINDER
)
return _SYSTEM_PROMPT, user
def generate_voice_actions(
timeline: dict,
model: str = DEFAULT_MODEL,
base_url: str = DEFAULT_BASE_URL,
temperature: float = 0.2,
timeout: float = REQUEST_TIMEOUT,
num_ctx: int = 32768,
max_words_per_segment: int = 200,
) -> Dict[str, Any]:
"""Ask the local model to direct the edit, returning validated actions.
Returns ``{"actions": [VoiceAction], "raw": str, "errors": [str]}``.
``actions`` is empty when the model returned nothing usable; ``errors``
carries the per-row rejections from :func:`parse_actions` plus any
extraction failure, so the caller can report what went wrong instead of
only the wins.
"""
system, user = build_edit_messages(timeline, max_words_per_segment)
try:
raw = ollama_chat(
model=model,
messages=[
{"role": "system", "content": system},
{"role": "user", "content": user},
],
base_url=base_url,
temperature=temperature,
timeout=timeout,
num_ctx=num_ctx,
)
except RuntimeError as exc:
return {"actions": [], "raw": "", "errors": [str(exc)]}
data = _extract_json(raw)
if data is None:
return {
"actions": [],
"raw": raw,
"errors": ["O modelo não devolveu um JSON de decisões legível."],
}
actions, errors = parse_actions(data)
return {"actions": actions, "raw": raw, "errors": errors}
+109 -3
View File
@@ -382,6 +382,10 @@ DEFAULT_VOICE_ANALYSIS_CONFIG: dict = {
"emphasis_floor": 0.25, "emphasis_floor": 0.25,
"emotion_enabled": False, "emotion_enabled": False,
"emotion_sensitivity": 0.5, "emotion_sensitivity": 0.5,
"zoom_scale": 1.30,
"zoom_mode": "in_out",
"zoom_ease_in": 0.25,
"zoom_ease_out": 0.04,
} }
@@ -402,12 +406,28 @@ def load_voice_analysis_config() -> dict:
stored = _load_config().get("voice_analysis") stored = _load_config().get("voice_analysis")
if not isinstance(stored, dict): if not isinstance(stored, dict):
return cfg return cfg
for key in ("energy_threshold", "peak_percentile", "emphasis_floor", "emotion_sensitivity"): for key in (
"energy_threshold", "peak_percentile", "emphasis_floor",
"emotion_sensitivity", "zoom_scale", "zoom_ease_in", "zoom_ease_out",
):
if key in stored: if key in stored:
try: try:
cfg[key] = max(0.0, min(1.0, float(stored[key]))) value = float(stored[key])
if key == "zoom_scale":
cfg[key] = max(1.0, min(3.0, value))
elif key.startswith("zoom_ease"):
cfg[key] = max(0.01, min(5.0, value))
else:
cfg[key] = max(0.0, min(1.0, value))
except (TypeError, ValueError): except (TypeError, ValueError):
pass pass
if "emphasis_threshold" in stored and "emphasis_floor" not in stored:
try:
cfg["emphasis_floor"] = max(0.0, min(1.0, float(stored["emphasis_threshold"])))
except (TypeError, ValueError):
pass
if stored.get("zoom_mode") in ("in_out", "in", "out"):
cfg["zoom_mode"] = stored["zoom_mode"]
if "emotion_enabled" in stored: if "emotion_enabled" in stored:
cfg["emotion_enabled"] = bool(stored["emotion_enabled"]) cfg["emotion_enabled"] = bool(stored["emotion_enabled"])
weights = stored.get("emphasis_weights") weights = stored.get("emphasis_weights")
@@ -428,6 +448,10 @@ def save_voice_analysis_config(
emphasis_floor: float | None = None, emphasis_floor: float | None = None,
emotion_enabled: bool | None = None, emotion_enabled: bool | None = None,
emotion_sensitivity: float | None = None, emotion_sensitivity: float | None = None,
zoom_scale: float | None = None,
zoom_mode: str | None = None,
zoom_ease_in: float | None = None,
zoom_ease_out: float | None = None,
) -> dict: ) -> dict:
"""Persist voice-analysis thresholds/weights. Only given fields change. """Persist voice-analysis thresholds/weights. Only given fields change.
@@ -446,6 +470,14 @@ def save_voice_analysis_config(
cfg["emotion_enabled"] = bool(emotion_enabled) cfg["emotion_enabled"] = bool(emotion_enabled)
if emotion_sensitivity is not None: if emotion_sensitivity is not None:
cfg["emotion_sensitivity"] = max(0.0, min(1.0, float(emotion_sensitivity))) cfg["emotion_sensitivity"] = max(0.0, min(1.0, float(emotion_sensitivity)))
if zoom_scale is not None:
cfg["zoom_scale"] = max(1.0, min(3.0, float(zoom_scale)))
if zoom_mode in ("in_out", "in", "out"):
cfg["zoom_mode"] = zoom_mode
if zoom_ease_in is not None:
cfg["zoom_ease_in"] = max(0.01, min(5.0, float(zoom_ease_in)))
if zoom_ease_out is not None:
cfg["zoom_ease_out"] = max(0.01, min(5.0, float(zoom_ease_out)))
if emphasis_weights is not None: if emphasis_weights is not None:
for key, value in emphasis_weights.items(): for key, value in emphasis_weights.items():
if key in cfg["emphasis_weights"] and value is not None: if key in cfg["emphasis_weights"] and value is not None:
@@ -536,6 +568,73 @@ def save_dynamic_subtitle_config(**fields) -> dict:
return cfg return cfg
DEFAULT_PLAIN_SUBTITLE_CONFIG: dict = {
"font": "Helvetica Neue",
"font_size": 82,
"font_color": "1 1 1 1",
"max_words": 7,
"position_y": -820.0,
"uppercase": False,
"keep_punctuation": True,
"text_scale": 2.0,
}
def load_plain_subtitle_config() -> dict:
"""Persisted style for simple editable FCPXML title subtitles."""
cfg = dict(DEFAULT_PLAIN_SUBTITLE_CONFIG)
stored = _load_config().get("plain_subtitles")
if not isinstance(stored, dict):
return cfg
for key in ("position_y", "text_scale"):
if key in stored:
try:
cfg[key] = float(stored[key])
except (TypeError, ValueError):
pass
for key in ("font_size", "max_words"):
if key in stored:
try:
cfg[key] = int(stored[key])
except (TypeError, ValueError):
pass
for key in ("font", "font_color"):
if key in stored and isinstance(stored[key], str) and stored[key]:
cfg[key] = stored[key]
for key in ("uppercase", "keep_punctuation"):
if key in stored:
cfg[key] = bool(stored[key])
cfg["max_words"] = max(1, int(cfg["max_words"]))
return cfg
def save_plain_subtitle_config(**fields) -> dict:
"""Persist simple subtitle style fields. Only given fields change."""
cfg = load_plain_subtitle_config()
for key, value in fields.items():
if key not in DEFAULT_PLAIN_SUBTITLE_CONFIG or value is None:
continue
if isinstance(DEFAULT_PLAIN_SUBTITLE_CONFIG[key], bool):
cfg[key] = bool(value)
elif isinstance(DEFAULT_PLAIN_SUBTITLE_CONFIG[key], float):
try:
cfg[key] = float(value)
except (TypeError, ValueError):
continue
elif isinstance(DEFAULT_PLAIN_SUBTITLE_CONFIG[key], int):
try:
cfg[key] = int(value)
except (TypeError, ValueError):
continue
else:
cfg[key] = str(value)
cfg["max_words"] = max(1, int(cfg["max_words"]))
data = _load_config()
data["plain_subtitles"] = cfg
_write_config(data)
return cfg
# Mirrors the silence thresholds the detection/removal handlers use when no # Mirrors the silence thresholds the detection/removal handlers use when no
# argument is passed (server_tools/qc.py). Persisted so the app's slider and # argument is passed (server_tools/qc.py). Persisted so the app's slider and
# any later run agree without threading three fields through every call. # any later run agree without threading three fields through every call.
@@ -545,7 +644,14 @@ DEFAULT_SILENCE_CONFIG: dict = {
# Seconds a quiet stretch must last before it's a cut candidate. # Seconds a quiet stretch must last before it's a cut candidate.
"min_silence": 0.5, "min_silence": 0.5,
# Seconds left inside each cut so speech never gets clipped at the edges. # Seconds left inside each cut so speech never gets clipped at the edges.
"padding": 0.05, # 0.2s matches the breathing-room convention for phrase-boundary cuts
# (see editar-por-voz/criterios/06-texto-corte-marcador.md) — a silence
# span this tool finds is often the natural breath before a new
# sentence, not just editing slop, and 0.05s shaved that breath down to
# almost nothing (real case: Mastopexia project, the pause before "Com"
# went from 0.567s to 0.1s across the cut, landing the next clip only
# 5ms after the word instead of a natural pause before it).
"padding": 0.2,
} }
File diff suppressed because it is too large Load Diff
+120
View File
@@ -0,0 +1,120 @@
"""
Data models for Final Cut Pro FCPXML structures.
Provides a clean Python interface for working with Final Cut Pro timelines,
clips, markers, and other elements.
Era um módulo de 1.091 linhas com seis famílias de modelo dentro. Agora cada
família tem seu arquivo, e este pacote reexporta tudo — `from .models import
TimeValue` segue valendo em todo o projeto, inclusive para os nomes com
underscore que o writer e a suíte já usavam.
enums tipos e cores de marcador, transições, ritmo
timing TimeValue (fração racional) e Timecode
timeline clipes, marcadores, lanes, projeto
planning rough cut, ritmo, montagem
qc achados de QC e resultado de validação
subtitles paleta e look das legendas dinâmicas
"""
from .enums import (
_MAX_MARKER_TYPE_LENGTH,
MARKER_XML_TAGS,
FlashFrameSeverity,
MarkerColor,
MarkerType,
PacingCurve,
PacingStyle,
TransitionType,
ValidationIssueType,
)
from .planning import (
MontageConfig,
PacingConfig,
RoughCutResult,
SegmentSpec,
)
from .qc import (
DuplicateGroup,
FlashFrame,
GapInfo,
ValidationIssue,
ValidationResult,
)
from .subtitles import (
COLOR_GREY,
COLOR_INDIGO,
COLOR_WHITE,
COLOR_YELLOW,
EDITORIAL_BODY_LOOK,
EDITORIAL_EMPHASIS_LOOK,
REFERENCE_RHYTHM,
DynamicSubtitleConfig,
SubtitlePosition,
WordLook,
WordStyle,
)
from .timeline import (
AudioClip,
Clip,
CompoundClip,
ConnectedClip,
Keyword,
Marker,
Project,
SilenceCandidate,
Timeline,
Transition,
VideoClip,
)
from .timing import (
_FCPXML_STANDARD_TIMEBASES,
Timecode,
TimeValue,
)
__all__ = [
"AudioClip",
"COLOR_GREY",
"COLOR_INDIGO",
"COLOR_WHITE",
"COLOR_YELLOW",
"Clip",
"CompoundClip",
"ConnectedClip",
"DuplicateGroup",
"DynamicSubtitleConfig",
"EDITORIAL_BODY_LOOK",
"EDITORIAL_EMPHASIS_LOOK",
"FlashFrame",
"FlashFrameSeverity",
"GapInfo",
"Keyword",
"MARKER_XML_TAGS",
"Marker",
"MarkerColor",
"MarkerType",
"MontageConfig",
"PacingConfig",
"PacingCurve",
"PacingStyle",
"Project",
"REFERENCE_RHYTHM",
"RoughCutResult",
"SegmentSpec",
"SilenceCandidate",
"SubtitlePosition",
"TimeValue",
"Timecode",
"Timeline",
"Transition",
"TransitionType",
"ValidationIssue",
"ValidationIssueType",
"ValidationResult",
"VideoClip",
"WordLook",
"WordStyle",
"_FCPXML_STANDARD_TIMEBASES",
"_MAX_MARKER_TYPE_LENGTH",
]
+183
View File
@@ -0,0 +1,183 @@
"""Enumerações do domínio: tipos e cores de marcador, transições, ritmo.
Extraído de models.py — ver fcpxml/models/__init__.py.
"""
from enum import Enum
# Maximum length for marker type strings to prevent memory abuse
_MAX_MARKER_TYPE_LENGTH = 64
class MarkerType(Enum):
"""Types of markers in Final Cut Pro.
Members:
STANDARD — Default marker with no completion state.
INCOMPLETE — Task marker (completed="0" in FCPXML). ← canonical name
TODO — Alias for INCOMPLETE. Kept for backward compatibility;
resolves to the same object (``MarkerType.TODO is
MarkerType.INCOMPLETE``). Python enums treat the first
member with a given value as canonical; all subsequent
members sharing that value become aliases.
CHAPTER — Chapter marker (``<chapter-marker>`` element).
COMPLETED — Task marker with completed="1".
Serialization helpers:
``from_string()`` — Accepts values, names, and legacy aliases
(e.g. ``"todo-marker"``). Always returns the
canonical member.
``from_xml_element()`` — Reads an ``lxml``/``ElementTree`` element and
returns the appropriate type based on the tag
name and ``completed`` attribute.
``xml_tag`` — The FCPXML element tag to emit when writing.
``xml_attrs`` — Extra attributes required when writing (e.g.
``completed="0"`` for INCOMPLETE).
"""
STANDARD = "standard"
INCOMPLETE = "todo"
TODO = "todo" # Backward-compat alias — resolves to INCOMPLETE at runtime
CHAPTER = "chapter"
COMPLETED = "completed"
@classmethod
def from_string(cls, value: str) -> 'MarkerType':
"""Convert a string to MarkerType, accepting both enum names and values.
Includes input validation: rejects null bytes, control characters,
and excessively long strings to prevent injection and memory abuse.
Examples:
MarkerType.from_string("todo") -> MarkerType.INCOMPLETE
MarkerType.from_string("TODO") -> MarkerType.INCOMPLETE
MarkerType.from_string("completed") -> MarkerType.COMPLETED
"""
if not isinstance(value, str):
raise TypeError(f"Expected str, got {type(value).__name__}")
if '\x00' in value or any(ord(c) < 32 and c not in ('\n', '\r', '\t') for c in value):
raise ValueError("Marker type contains invalid control characters")
if len(value) > _MAX_MARKER_TYPE_LENGTH:
raise ValueError(
f"Marker type exceeds maximum length ({_MAX_MARKER_TYPE_LENGTH} chars)"
)
lowered = value.strip().lower()
if not lowered:
raise ValueError("Marker type cannot be empty")
# Accept legacy aliases from older specs (e.g. "todo-marker" → INCOMPLETE)
aliases = {
"todo-marker": "todo",
"completed-marker": "completed",
"chapter-marker": "chapter",
}
lowered = aliases.get(lowered, lowered)
try:
return cls(lowered)
except ValueError:
raise ValueError(
f"Invalid marker type: '{value}'. "
f"Valid types: {', '.join(m.value for m in cls)}"
)
@classmethod
def from_xml_element(cls, elem) -> 'MarkerType':
"""Determine MarkerType from an XML element's tag and attributes.
Centralises the parse-side mapping so the parser doesn't need to
know about completed-attribute semantics.
Rules (in priority order):
1. <chapter-marker> tag → CHAPTER (completed attr ignored)
2. completed='0' (exact) → INCOMPLETE
3. completed='1' (exact) → COMPLETED
4. Everything else → STANDARD (including whitespace-padded,
absent, empty, or non-boolean completed values)
Matching is intentionally strict — no .strip(), no case folding.
This prevents whitespace-injected attributes like ' 0 ' from
being misclassified.
"""
if elem.tag == 'chapter-marker':
return cls.CHAPTER
completed = elem.get('completed')
if completed == '0':
return cls.INCOMPLETE
if completed == '1':
return cls.COMPLETED
return cls.STANDARD
@property
def xml_tag(self) -> str:
"""Return the FCPXML element tag for this marker type."""
return 'chapter-marker' if self == MarkerType.CHAPTER else 'marker'
@property
def xml_attrs(self) -> dict:
"""Return extra XML attributes this marker type requires when writing.
Centralises the write-side mapping so both FCPXMLModifier and
FCPXMLWriter use a single source of truth.
"""
if self == MarkerType.CHAPTER:
return {'posterOffset': '0s'}
if self == MarkerType.INCOMPLETE:
return {'completed': '0'}
if self == MarkerType.COMPLETED:
return {'completed': '1'}
return {}
# Recognised marker XML tags — used by the parser for single-pass collection
# and by the writer to validate element creation.
MARKER_XML_TAGS = ('marker', 'chapter-marker')
class MarkerColor(Enum):
"""Marker color options (FCP internal values)."""
BLUE = 0
CYAN = 1
GREEN = 2
YELLOW = 3
ORANGE = 4
RED = 5
PINK = 6
PURPLE = 7
class TransitionType(Enum):
"""Built-in transition types."""
CROSS_DISSOLVE = "Cross Dissolve"
FADE_TO_BLACK = "Fade to Color"
FADE_FROM_BLACK = "Fade from Color"
DIP_TO_COLOR = "Dip to Color"
WIPE = "Wipe"
SLIDE = "Slide"
class PacingStyle(Enum):
"""Pacing presets for rough cut generation."""
SLOW = "slow" # 5-10 second cuts
MEDIUM = "medium" # 2-5 second cuts
FAST = "fast" # 0.5-2 second cuts
DYNAMIC = "dynamic" # Varies throughout
class FlashFrameSeverity(Enum):
"""Severity levels for flash frame detection."""
CRITICAL = "critical" # < 2 frames, almost certainly an error
WARNING = "warning" # < 6 frames, potentially intentional but suspicious
class PacingCurve(Enum):
"""Pacing curves for montage generation."""
CONSTANT = "constant" # Same clip duration throughout
ACCELERATING = "accelerating" # Starts slow, gets faster
DECELERATING = "decelerating" # Starts fast, gets slower
PYRAMID = "pyramid" # Slow → fast → slow
class ValidationIssueType(Enum):
"""Types of timeline validation issues."""
FLASH_FRAME = "flash_frame"
GAP = "gap"
DUPLICATE = "duplicate"
ORPHAN_REF = "orphan_ref"
INVALID_OFFSET = "invalid_offset"
# DTD validation types (v0.6.0)
ELEMENT_ORDER = "element_order"
MISSING_ATTRIBUTE = "missing_attribute"
INVALID_TIMEBASE = "invalid_timebase"
FRAME_MISALIGNMENT = "frame_misalignment"
MISSING_EFFECT_REF = "missing_effect_ref"
MISSING_MEDIA_REP = "missing_media_rep"
+93
View File
@@ -0,0 +1,93 @@
"""Especificações de geração: rough cut, ritmo e montagem.
Extraído de models.py — ver fcpxml/models/__init__.py.
"""
from dataclasses import dataclass, field
from typing import List, Optional, Tuple
from .enums import PacingCurve
@dataclass
class SegmentSpec:
"""Specification for a segment in auto rough cut."""
name: str
keywords: List[str] = field(default_factory=list)
duration_seconds: float = 0.0
priority: str = "best" # favorites, longest, shortest, random, best
@dataclass
class PacingConfig:
"""Configuration for rough cut pacing."""
pacing: str = "medium" # slow, medium, fast, dynamic
min_clip_duration: float = 1.0
max_clip_duration: float = 8.0
avg_clip_duration: Optional[float] = None
vary_pacing: bool = True
def get_duration_range(self) -> Tuple[float, float]:
"""Get min/max based on pacing style."""
ranges = {
"slow": (5.0, 10.0),
"medium": (2.0, 5.0),
"fast": (0.5, 2.0),
"dynamic": (1.0, 6.0),
}
return ranges.get(self.pacing, (2.0, 5.0))
@dataclass
class RoughCutResult:
"""Result of auto rough cut generation."""
output_path: str
clips_used: int
clips_available: int
target_duration: float
actual_duration: float
segments: int
average_clip_duration: float
@dataclass
class MontageConfig:
"""Configuration for montage generation with pacing curves."""
target_duration: float # Target duration in seconds
pacing_curve: 'PacingCurve'
start_duration: float = 2.0 # Clip duration at start
end_duration: float = 0.5 # Clip duration at end
min_duration: float = 0.2 # Minimum allowed clip duration
max_duration: float = 5.0 # Maximum allowed clip duration
def get_duration_at_position(self, position: float) -> float:
"""
Calculate clip duration for a given position (0.0 to 1.0).
Args:
position: Position in montage (0.0 = start, 1.0 = end)
Returns:
Target duration in seconds for a clip at this position
"""
if self.pacing_curve == PacingCurve.CONSTANT:
duration = (self.start_duration + self.end_duration) / 2
elif self.pacing_curve == PacingCurve.ACCELERATING:
# Linear interpolation from start to end duration
duration = self.start_duration + (self.end_duration - self.start_duration) * position
elif self.pacing_curve == PacingCurve.DECELERATING:
# Reverse: start fast, end slow
duration = self.end_duration + (self.start_duration - self.end_duration) * position
elif self.pacing_curve == PacingCurve.PYRAMID:
# Slow → fast → slow (parabolic curve)
if position < 0.5:
# First half: slow to fast
duration = self.start_duration + (self.end_duration - self.start_duration) * (position * 2)
else:
# Second half: fast to slow
duration = self.end_duration + (self.start_duration - self.end_duration) * ((position - 0.5) * 2)
else:
duration = self.start_duration
# Clamp to min/max
return max(self.min_duration, min(self.max_duration, duration))
+121
View File
@@ -0,0 +1,121 @@
"""Achados de QC e o resultado de uma validação.
Extraído de models.py — ver fcpxml/models/__init__.py.
"""
from dataclasses import dataclass, field
from typing import Any, Dict, List, Optional
from .enums import FlashFrameSeverity, ValidationIssueType
from .timing import Timecode
@dataclass
class FlashFrame:
"""
Represents a detected flash frame (ultra-short clip).
Flash frames are typically editing errors - clips that are too short
to be perceived as intentional cuts.
"""
clip_name: str
clip_id: str
start: Timecode
duration_frames: int
duration_seconds: float
severity: 'FlashFrameSeverity'
@property
def is_critical(self) -> bool:
"""Check if this is a critical flash frame."""
return self.severity == FlashFrameSeverity.CRITICAL
@dataclass
class GapInfo:
"""
Represents a detected gap in the timeline.
Gaps can be intentional (black frames) or errors from deleted clips.
"""
start: Timecode
duration_frames: int
duration_seconds: float
previous_clip: Optional[str] = None # Clip name before the gap
next_clip: Optional[str] = None # Clip name after the gap
@property
def timecode(self) -> str:
"""Get timecode string for the gap start."""
return self.start.to_smpte()
@dataclass
class DuplicateGroup:
"""
Represents a group of clips using the same source media.
Useful for detecting duplicate clips that may be unintentional.
"""
source_ref: str # The asset/media reference ID
source_name: str # Human-readable source name
clips: List[Dict[str, Any]] = field(default_factory=list) # List of clip info dicts
@property
def count(self) -> int:
"""Number of clips using this source."""
return len(self.clips)
@property
def has_overlapping_ranges(self) -> bool:
"""Check if any clips use overlapping portions of the source."""
# Sort clips by source_start
sorted_clips = sorted(self.clips, key=lambda c: c.get('source_start', 0))
for i in range(len(sorted_clips) - 1):
curr_end = sorted_clips[i].get('source_start', 0) + sorted_clips[i].get('source_duration', 0)
next_start = sorted_clips[i + 1].get('source_start', 0)
if curr_end > next_start:
return True
return False
@dataclass
class ValidationIssue:
"""
Represents a single validation issue found in a timeline.
Used by validate_timeline to report problems.
"""
issue_type: 'ValidationIssueType'
severity: str # "error", "warning", "info"
message: str
timecode: Optional[str] = None
clip_name: Optional[str] = None
details: Dict[str, Any] = field(default_factory=dict)
@dataclass
class ValidationResult:
"""
Result of timeline validation.
Provides a health score and categorized list of issues.
"""
is_valid: bool
health_score: int # 0-100 percentage
issues: List[ValidationIssue] = field(default_factory=list)
flash_frames: List[FlashFrame] = field(default_factory=list)
gaps: List[GapInfo] = field(default_factory=list)
duplicates: List[DuplicateGroup] = field(default_factory=list)
@property
def error_count(self) -> int:
return len([i for i in self.issues if i.severity == "error"])
@property
def warning_count(self) -> int:
return len([i for i in self.issues if i.severity == "warning"])
def summary(self) -> str:
"""Generate a summary string."""
return (
f"Timeline Health: {self.health_score}% | "
f"Errors: {self.error_count} | Warnings: {self.warning_count} | "
f"Flash frames: {len(self.flash_frames)} | Gaps: {len(self.gaps)}"
)
+165
View File
@@ -0,0 +1,165 @@
"""Aparência das legendas dinâmicas: paleta, look por palavra, configuração.
Extraído de models.py — ver fcpxml/models/__init__.py.
"""
from dataclasses import dataclass, field
from typing import Optional
from ..text_layout import REFERENCE_BLOCK_LINE_GAP, TEXT_TEMPLATE_FONT_SCALE
# The palette and type treatment of the calibration export
# ("Exemplo Letra.fcpxmld", sentence "Toda a minha vida, assim,"), copied
# verbatim from what the user set in Final Cut's Inspector.
COLOR_INDIGO = "0.156863 0 0.596079 1"
COLOR_YELLOW = "0.997808 0.882664 0.0388632 1"
COLOR_GREY = "0.7 0.7 0.7 1"
COLOR_WHITE = "1 1 1 1"
@dataclass
class WordLook:
"""How one word is set: size, colour and type treatment.
A sentence cycles through a tuple of these, so its typography reads with a
deliberate rhythm rather than a uniform block.
"""
font_size: int
color: str
font: str = "Helvetica Neue"
face: Optional[str] = None # Final Cut's fontFace, e.g. "Light Italic"
kerning: float = 2.048
@property
def italic(self) -> bool:
return bool(self.face) and "italic" in self.face.lower()
# One entry per word of the reference sentence, in order:
# Toda(170, indigo, Helvetica Light) a(128, yellow) minha(151, grey)
# vida,(128, white) assim,(128, grey, Light Italic)
REFERENCE_RHYTHM = (
WordLook(170, COLOR_INDIGO, font="Helvetica", face="Light", kerning=2.72),
WordLook(128, COLOR_YELLOW),
WordLook(151, COLOR_GREY, kerning=2.416),
WordLook(128, COLOR_WHITE),
WordLook(128, COLOR_GREY, face="Light Italic"),
)
# The progressive-composition look (reference: the reel the user sent,
# 2026-08-17). Supporting text in a small grotesque, the sentence's key word
# large in a display italic, everything white — the two-font contrast IS the
# style. Playfair Display ships in the user's ~/Library/Fonts and its real
# advance widths are embedded in font_metrics, so the lines can be measured
# rather than guessed. Both are plain WordLooks: swap them for any installed
# family (a script/calligraphic face for the emphasis, say) and layout follows.
EDITORIAL_EMPHASIS_LOOK = WordLook(
230, COLOR_WHITE, font="Playfair Display", face="Medium Italic", kerning=0.0,
)
EDITORIAL_BODY_LOOK = WordLook(
88, COLOR_WHITE, font="Helvetica Neue", face="Bold", kerning=1.2,
)
@dataclass
class WordStyle:
"""Per-word text styling for dynamic (karaoke-style) subtitles.
``rhythm`` drives size, colour and face, cycling by the word's index within
its sentence — deterministic, so regenerating a transcript twice yields the
same look. ``font``/``font_size`` are the fallback when ``rhythm`` is empty.
"""
font: str = "Helvetica Neue"
font_size: int = 128
active_color: str = COLOR_WHITE
inactive_color: str = COLOR_GREY
bold: bool = False
kerning: float = 2.048
rhythm: tuple = REFERENCE_RHYTHM
# Progressive composition only (granularity="phrase").
emphasis_look: Optional[WordLook] = None
body_look: Optional[WordLook] = None
def look_for(self, index: int) -> WordLook:
"""The look for the word at *index* within its sentence."""
if not self.rhythm:
return WordLook(
self.font_size, self.active_color,
font=self.font, kerning=self.kerning,
)
return self.rhythm[index % len(self.rhythm)]
def look_for_emphasis(self) -> WordLook:
"""The look for a composition's key word (progressive composition)."""
return self.emphasis_look or EDITORIAL_EMPHASIS_LOOK
def look_for_body(self) -> WordLook:
"""The look for a composition's supporting lines."""
return self.body_look or EDITORIAL_BODY_LOOK
@dataclass
class SubtitlePosition:
"""Screen position for generated title clips, in FCP title coordinate space."""
x: float = 0.0
y: float = -300.0
alignment: str = "center" # left | center | right
@dataclass
class DynamicSubtitleConfig:
"""Options for FCPXMLWriter.generate_dynamic_subtitles().
Dynamic subtitles are animated TITLES, not captions. Both templates below
render on the video title lane and never carry a ``subtitles.*`` role — a
``role="subtitles.*"`` would make Final Cut treat them as captions and
hide them behind the caption-display toggle. They DO carry a
``titles.*`` sub-role (``role``), which groups them in Final Cut's
role index and lanes them with a distinct colour, without ever being
mistaken for closed captions.
``animated`` picks the template: True uses "Essencial - Título"
(Essential Title), which animates on its own Motion defaults; False uses
the static "Título Básico" (Basic Title). Default is True — the animated
reveal is the feature's purpose.
Words are grouped into sentences and laid out as a compact typographic
block: each word becomes its own positioned ``<title>``, appearing as it is
spoken and accumulating on screen, with every word of a block clearing at
the same instant so the sentence vanishes as a whole.
``band_height`` is the fraction of frame height the block may occupy, and
``block_center_y`` its centre in canvas points (negative is below frame
centre). The defaults reproduce the calibration export the user built by
hand: a block of at most three lines sitting just below centre. A sentence
taller than the band splits into successive blocks.
"""
style: WordStyle = field(default_factory=WordStyle)
position: SubtitlePosition = field(default_factory=SubtitlePosition)
animated: bool = True
band_height: float = 0.22
block_center_y: float = -167.0
# "phrase": one title per LINE of the composition — supporting words
# grouped, the key word alone and large (the reference look). "word": one
# title per word, the earlier rhythm.
granularity: str = "phrase"
# Ratio between the template's fontSize space and the canvas-point space
# its Position uses. See text_layout.TEXT_TEMPLATE_FONT_SCALE: the "Text"
# (Text.moti) template sizes type in frame pixels, so a size chosen in
# points renders half as large unless it is converted on the way out.
text_scale: float = TEXT_TEMPLATE_FONT_SCALE
# Vertical air between stacked lines, in canvas points. Negative values
# deliberately overlap the lines — the display italic tucking under the
# line above is a real editorial look, and the stacking arithmetic places
# ink boxes edge to edge, so a negative gap moves them by exactly that
# much rather than colliding unpredictably.
line_gap: float = REFERENCE_BLOCK_LINE_GAP
# Final Cut role for every title this generator emits. A ``titles.*``
# sub-role (NOT ``subtitles.*``) groups the clips in the role index and
# tints their lane, keeping dynamic captions distinct from plain
# ones and from Final Cut's own closed-caption toggle.
role: str = "titles.dinamicas"
# Run the post-generation collision validation (collision.validate_titles)
# and refuse to emit when it reports a blocking overlap. Off by default so
# generation stays byte-identical to before this flag existed; flip it on
# for a guaranteed no-collision export.
validate: bool = False
+248
View File
@@ -0,0 +1,248 @@
"""O que existe numa timeline: clipes, marcadores, lanes, projeto.
Extraído de models.py — ver fcpxml/models/__init__.py.
"""
from dataclasses import dataclass, field
from typing import List, Optional
from .enums import MarkerColor, MarkerType
from .timing import Timecode
@dataclass
class Keyword:
"""Represents a keyword/tag applied to a clip."""
value: str
start: Optional[Timecode] = None
duration: Optional[Timecode] = None
@dataclass
class ParametroEfeito:
"""Um parâmetro de um filtro de efeito (``<param>`` dentro do filtro)."""
nome: str
valor: str
chave: str = ""
metadado: str = ""
@dataclass
class EfeitoAjuste:
"""Um efeito aplicado por uma camada de ajuste (adjustment layer).
``uid`` é o UUID do efeito interno do Final Cut (ver ``FCP_EFFECTS`` em
``fcpxml/writer/helpers.py`` para os efeitos built-in). ``tipo`` é
``"video"`` ou ``"audio"`` — decide se vira ``<filter-video>`` ou
``<filter-audio>``, filho direto do ``<clip>`` da camada de ajuste (o
DTD não define wrapper ``<adjustment>``).
"""
nome: str
uid: str
tipo: str = "video"
parametros: List[ParametroEfeito] = field(default_factory=list)
@dataclass
class Marker:
"""Represents a marker in the timeline."""
name: str
start: Timecode
duration: Optional[Timecode] = None
marker_type: MarkerType = MarkerType.STANDARD
note: str = ""
color: Optional[MarkerColor] = None
def to_youtube_timestamp(self) -> str:
"""Format as YouTube chapter timestamp."""
total_seconds = int(self.start.seconds)
hours = total_seconds // 3600
minutes = (total_seconds % 3600) // 60
secs = total_seconds % 60
if hours > 0:
return f"{hours}:{minutes:02d}:{secs:02d}"
return f"{minutes}:{secs:02d}"
@dataclass
class Clip:
"""Represents a clip in the timeline."""
name: str
start: Timecode
duration: Timecode
source_start: Optional[Timecode] = None
source_end: Optional[Timecode] = None
media_path: str = ""
markers: List[Marker] = field(default_factory=list)
keywords: List[Keyword] = field(default_factory=list)
# Extended metadata
rating: int = 0 # 0=unrated, 1-5 stars
is_favorite: bool = False
is_rejected: bool = False
# Roles (FCP audio/video role assignments)
audio_role: str = ""
video_role: str = ""
# Connected clips (B-roll, titles, audio attached to this clip)
connected_clips: List['ConnectedClip'] = field(default_factory=list)
# Edit-time correction, in degrees, from a Transform filter on the clip
# (e.g. straightening a tilted phone shot) — not the camera's own
# recorded orientation, which lives in the media file itself.
rotation: float = 0.0
@property
def end(self) -> Timecode:
return Timecode(
frames=self.start.frames + self.duration.frames,
frame_rate=self.start.frame_rate
)
@property
def duration_seconds(self) -> float:
return self.duration.seconds
@property
def keyword_values(self) -> List[str]:
"""Get list of keyword strings."""
return [k.value for k in self.keywords]
@dataclass
class AudioClip(Clip):
"""Audio-specific clip."""
channels: int = 2
sample_rate: int = 48000
role: str = "dialogue"
@dataclass
class VideoClip(Clip):
"""Video-specific clip."""
width: int = 1920
height: int = 1080
has_audio: bool = True
@dataclass
class ConnectedClip:
"""A clip connected to a primary storyline clip (B-roll, titles, audio).
In FCP's magnetic timeline, connected clips hang off spine clips via lanes.
Positive lanes are above (video overlays), negative lanes are below (audio).
"""
name: str
start: Timecode
duration: Timecode
lane: int = 1
offset: Optional[Timecode] = None
source_start: Optional[Timecode] = None
media_path: str = ""
clip_type: str = "asset-clip"
role: str = ""
ref_id: str = ""
parent_clip_name: str = ""
markers: List[Marker] = field(default_factory=list)
keywords: List[Keyword] = field(default_factory=list)
rotation: float = 0.0
@property
def duration_seconds(self) -> float:
return self.duration.seconds
@dataclass
class CompoundClip:
"""A compound clip (ref-clip) containing a nested timeline."""
name: str
ref_id: str
duration: Timecode
start: Timecode
clips: List[Clip] = field(default_factory=list)
connected_clips: List[ConnectedClip] = field(default_factory=list)
@property
def duration_seconds(self) -> float:
return self.duration.seconds
@dataclass
class SilenceCandidate:
"""A potential silence region detected by timeline heuristics."""
start_timecode: str
duration_seconds: float
reason: str # "gap", "ultra_short", "name_match", "duration_anomaly"
confidence: float = 0.5 # 0.0 to 1.0
clip_name: Optional[str] = None
clip_index: Optional[int] = None
@dataclass
class Transition:
"""Represents a transition between clips."""
name: str
duration: Timecode
start: Timecode
transition_type: str = "cross-dissolve"
@dataclass
class Timeline:
"""Represents a Final Cut Pro timeline/sequence."""
name: str
duration: Timecode
frame_rate: float = 24.0
width: int = 1920
height: int = 1080
clips: List[Clip] = field(default_factory=list)
audio_clips: List[AudioClip] = field(default_factory=list)
transitions: List[Transition] = field(default_factory=list)
markers: List[Marker] = field(default_factory=list)
connected_clips: List[ConnectedClip] = field(default_factory=list)
compound_clips: List[CompoundClip] = field(default_factory=list)
@property
def total_clips(self) -> int:
return len(self.clips)
@property
def total_cuts(self) -> int:
return max(0, len(self.clips) - 1)
@property
def average_clip_duration(self) -> float:
if not self.clips:
return 0.0
return sum(c.duration_seconds for c in self.clips) / len(self.clips)
@property
def cuts_per_minute(self) -> float:
"""Average cuts per minute."""
if self.duration.seconds <= 0:
return 0.0
return (self.total_cuts / self.duration.seconds) * 60
def get_clips_shorter_than(self, seconds: float) -> List[Clip]:
"""Find clips shorter than threshold (flash frame detection)."""
return [c for c in self.clips if c.duration_seconds < seconds]
def get_clips_longer_than(self, seconds: float) -> List[Clip]:
"""Find clips longer than threshold."""
return [c for c in self.clips if c.duration_seconds > seconds]
def get_clip_at(self, timecode: float) -> Optional[Clip]:
"""Find the clip at a specific timecode (seconds)."""
for clip in self.clips:
start_sec = clip.start.seconds
end_sec = clip.end.seconds
if start_sec <= timecode < end_sec:
return clip
return None
def get_clips_by_keyword(self, keyword: str) -> List[Clip]:
"""Find all clips with a specific keyword."""
return [c for c in self.clips if keyword in c.keyword_values]
@dataclass
class Project:
"""Represents a Final Cut Pro project/library."""
name: str
timelines: List[Timeline] = field(default_factory=list)
fcpxml_version: str = "1.13"
@property
def primary_timeline(self) -> Optional[Timeline]:
return self.timelines[0] if self.timelines else None
+304
View File
@@ -0,0 +1,304 @@
"""Tempo em fração racional — TimeValue e o Timecode que o embrulha.
Extraído de models.py — ver fcpxml/models/__init__.py.
"""
import operator
from dataclasses import dataclass
from fractions import Fraction
from functools import total_ordering
from math import gcd
from typing import Callable
# Standard FCPXML timebase denominators that FCP's DTD validator accepts.
# TimeValue.to_fcpxml() only simplifies fractions when the result uses one
# of these denominators, preventing values like "8/3s" that FCP rejects.
_FCPXML_STANDARD_TIMEBASES = frozenset({
1, 24, 25, 30, 48, 50, 60, 90, 96, 100, 120,
240, 600, 2400, 4800, 9600, 48000,
})
@total_ordering
@dataclass
class TimeValue:
"""
Represents time in FCPXML's rational format.
FCPXML uses fractions of seconds (e.g., "90/30s" for 3 seconds at 30fps).
This class handles conversion between timecode, seconds, and FCPXML format.
Examples:
TimeValue(90, 30) # 3 seconds at 30fps
TimeValue(1, 1) # 1 second
TimeValue.from_timecode("00:01:30:15", fps=30) # 90.5 seconds
"""
numerator: int
denominator: int = 1
def __post_init__(self):
if self.denominator == 0:
raise ValueError(
f"TimeValue denominator cannot be zero (got {self.numerator}/0). "
"This would corrupt all downstream time calculations."
)
# Normalize sign: denominator must always be positive.
# Cross-multiplication in __lt__/__eq__ assumes positive denominators;
# __hash__ assumes canonical form. Without this, TimeValue(1, -2)
# compares/hashes incorrectly against TimeValue(-1, 2).
if self.denominator < 0:
# Use object.__setattr__ because dataclass may be frozen-like
object.__setattr__(self, 'numerator', -self.numerator)
object.__setattr__(self, 'denominator', -self.denominator)
@classmethod
def from_timecode(cls, tc: str, fps: float = 30.0) -> 'TimeValue':
"""
Create TimeValue from various string formats.
Supported formats:
- "HH:MM:SS:FF" - Standard timecode
- "HH:MM:SS;FF" - Drop-frame timecode
- "30s" - Seconds
- "90/30s" - FCPXML rational format
- "15f" - Frames
"""
if not tc:
return cls(0, 1)
tc = str(tc).strip()
# FCPXML format: "90/30s" or "30s"
if tc.endswith('s'):
tc_val = tc[:-1]
if '/' in tc_val:
parts = tc_val.split('/', 1)
num, denom = int(parts[0]), int(parts[1])
if denom == 0:
raise ValueError(f"Zero denominator in timecode: {tc}")
return cls(num, denom)
else:
seconds = float(tc_val)
frames = int(round(seconds * fps))
# int(fps) truncates NTSC rates (23.976/29.97/59.94fps) to
# their nominal integer, mismatching the numerator (computed
# with the real fps) against the denominator — e.g. at
# 23.976fps this silently produced values ~1.04x too large.
# Reconstruct the exact rational fps (24000/1001, etc.) from
# the float instead, so numerator and denominator agree.
fps_frac = Fraction(fps).limit_denominator(100_000)
return cls(frames * fps_frac.denominator, fps_frac.numerator)
# Frame format: "15f"
if tc.endswith('f'):
frames = int(tc[:-1])
return cls(frames, int(fps))
# Timecode format: "HH:MM:SS:FF" or "HH:MM:SS;FF"
if ':' in tc or ';' in tc:
parts = tc.replace(';', ':').split(':')
if len(parts) == 4:
h, m, s, f = map(int, parts)
total_frames = int((h * 3600 + m * 60 + s) * fps + f)
return cls(total_frames, int(fps))
elif len(parts) == 3:
h, m, s = map(int, parts)
total_frames = int((h * 3600 + m * 60 + s) * fps)
return cls(total_frames, int(fps))
# Try as plain number (seconds)
try:
seconds = float(tc)
frames = int(round(seconds * fps))
return cls(frames, int(fps))
except ValueError:
raise ValueError(f"Invalid timecode format: {tc}")
@classmethod
def from_seconds(cls, seconds: float, fps: float = 30.0) -> 'TimeValue':
"""Create TimeValue from decimal seconds."""
frames = int(round(seconds * fps))
return cls(frames, int(fps))
@classmethod
def zero(cls) -> 'TimeValue':
"""Return zero time value."""
return cls(0, 1)
def to_fcpxml(self) -> str:
"""Convert to FCPXML time string (e.g., "90/30s").
Only simplifies when the denominator reduces to 1 (whole seconds)
or stays a standard FCPXML timebase. Avoids producing denominators
like 3, 7, etc. that FCP's DTD validator may reject.
"""
simplified = self.simplify()
if simplified.denominator == 1:
return f"{simplified.numerator}s"
# Keep original denominator if simplification produces a non-standard
# denominator (not a multiple of common timebases: 24, 30, 25, 2400)
if simplified.denominator in _FCPXML_STANDARD_TIMEBASES:
return f"{simplified.numerator}/{simplified.denominator}s"
# Fall back to unsimplified form
return f"{self.numerator}/{self.denominator}s"
def to_seconds(self) -> float:
"""Convert to decimal seconds."""
return self.numerator / self.denominator
def to_timecode(self, fps: float = 30.0) -> str:
"""Convert to HH:MM:SS:FF timecode string."""
total_frames = int(round(self.to_seconds() * fps))
total_secs, frames = divmod(total_frames, int(fps))
total_mins, secs = divmod(total_secs, 60)
hours, mins = divmod(total_mins, 60)
return f"{hours:02d}:{mins:02d}:{secs:02d}:{frames:02d}"
def to_frames(self, fps: float = 30.0) -> int:
"""Convert to frame count."""
return int(round(self.to_seconds() * fps))
def simplify(self) -> 'TimeValue':
"""Reduce fraction to simplest form."""
if self.numerator == 0:
return TimeValue(0, 1)
divisor = gcd(abs(self.numerator), abs(self.denominator))
return TimeValue(
self.numerator // divisor,
self.denominator // divisor
)
@staticmethod
def _lcm_denom(d1: int, d2: int) -> int:
"""LCM of two denominators for cross-timebase arithmetic."""
return d1 // gcd(d1, d2) * d2
def _binop(self, other: 'TimeValue', op: Callable[[int, int], int]) -> 'TimeValue':
"""Shared logic for add/sub: same-denom fast path, then LCM alignment."""
if self.denominator == other.denominator:
return TimeValue(op(self.numerator, other.numerator), self.denominator)
lcd = TimeValue._lcm_denom(self.denominator, other.denominator)
return TimeValue(
op(
self.numerator * (lcd // self.denominator),
other.numerator * (lcd // other.denominator),
),
lcd,
)
def __add__(self, other: 'TimeValue') -> 'TimeValue':
return self._binop(other, operator.add)
def __sub__(self, other: 'TimeValue') -> 'TimeValue':
return self._binop(other, operator.sub)
def __mul__(self, scalar: float) -> 'TimeValue':
new_num = round(self.numerator * scalar)
return TimeValue(new_num, self.denominator)
def __truediv__(self, scalar: float) -> 'TimeValue':
if scalar == 0:
raise ZeroDivisionError("Cannot divide TimeValue by zero")
new_denom = round(self.denominator * scalar)
if new_denom == 0:
raise ZeroDivisionError(
f"Division by {scalar} rounds denominator {self.denominator} to zero"
)
return TimeValue(self.numerator, new_denom)
def __lt__(self, other: 'TimeValue') -> bool:
# Cross-multiply to compare without float conversion:
# a/b < c/d ↔ a*d < c*b (denominators are always positive)
return self.numerator * other.denominator < other.numerator * self.denominator
def __eq__(self, other: object) -> bool:
if not isinstance(other, TimeValue):
return False
# Cross-multiply for exact integer comparison
return self.numerator * other.denominator == other.numerator * self.denominator
def __hash__(self) -> int:
# Delegate to simplify() — single source of truth for canonical form.
# __post_init__ guarantees denominator > 0, so no zero guard needed.
s = self.simplify()
return hash((s.numerator, s.denominator))
def snap_to_frame(self, fps: float) -> 'TimeValue':
"""Round this time value to the nearest frame boundary at the given fps.
Uses the 2400-tick timebase (LCM of common frame rates) so results
always land on clean frame boundaries.
Args:
fps: Frame rate to snap to (e.g. 24, 30, 60)
Returns:
New TimeValue snapped to the nearest frame in 2400-tick timebase.
"""
fps_int = int(fps)
if fps_int <= 0:
raise ValueError(f"fps must be positive, got {fps}")
ticks_per_frame = 2400 // fps_int
total_ticks = round(self.to_seconds() * 2400)
snapped_ticks = round(total_ticks / ticks_per_frame) * ticks_per_frame
return TimeValue(snapped_ticks, 2400)
def is_standard_timebase(self) -> bool:
"""Check if this TimeValue's denominator is an FCP-accepted timebase."""
simplified = self.simplify()
return simplified.denominator in _FCPXML_STANDARD_TIMEBASES
def __repr__(self) -> str:
return f"TimeValue({self.numerator}/{self.denominator}s = {self.to_seconds():.3f}s)"
@dataclass
class Timecode:
"""
Represents a timecode value.
Note: This class exists for backwards compatibility with the parser.
New code should prefer TimeValue for rational time math.
"""
frames: int
frame_rate: float = 24.0
drop_frame: bool = False
@property
def seconds(self) -> float:
return self.frames / self.frame_rate
@property
def total_frames(self) -> int:
return self.frames
def to_smpte(self) -> str:
"""Convert to SMPTE timecode string (HH:MM:SS:FF)."""
total_seconds = int(self.seconds)
hours = total_seconds // 3600
minutes = (total_seconds % 3600) // 60
secs = total_seconds % 60
frames = int((self.seconds - total_seconds) * self.frame_rate)
separator = ";" if self.drop_frame else ":"
return f"{hours:02d}:{minutes:02d}:{secs:02d}{separator}{frames:02d}"
@classmethod
def from_rational(cls, rational_str: str, frame_rate: float = 24.0) -> "Timecode":
"""Parse FCPXML rational time format (e.g., '3600/24s')."""
if not rational_str:
return cls(frames=0, frame_rate=frame_rate)
if rational_str.endswith('s'):
rational_str = rational_str[:-1]
if '/' in rational_str:
num, denom = rational_str.split('/')
seconds = int(num) / int(denom)
else:
seconds = float(rational_str)
frames = int(seconds * frame_rate)
return cls(frames=frames, frame_rate=frame_rate)
def to_rational(self) -> str:
"""Convert to FCPXML rational format."""
return f"{self.frames}/{int(self.frame_rate)}s"
def to_time_value(self) -> TimeValue:
"""Convert to TimeValue for rational math."""
return TimeValue(self.frames, int(self.frame_rate))
+15
View File
@@ -194,6 +194,7 @@ class FCPXMLParser:
media_path=media_path, media_path=media_path,
audio_role=elem.get('audioRole', ''), audio_role=elem.get('audioRole', ''),
video_role=elem.get('videoRole', ''), video_role=elem.get('videoRole', ''),
rotation=self._parse_clip_rotation(elem),
) )
clip.markers.extend(self._collect_markers(elem)) clip.markers.extend(self._collect_markers(elem))
@@ -205,6 +206,19 @@ class FCPXMLParser:
return clip return clip
def _parse_clip_rotation(self, elem: ET.Element) -> float:
"""Degrees from this clip's ``<adjust-transform rotation="...">`` —
an edit-time correction (e.g. straightening a tilted phone shot),
not the camera's own recorded orientation. FCP writes the rotation
as an attribute on that element, not as a filter param."""
transform = elem.find('adjust-transform')
if transform is None:
return 0.0
try:
return float(transform.get('rotation', '0'))
except ValueError:
return 0.0
def _parse_marker_element(self, elem: ET.Element) -> Optional[Marker]: def _parse_marker_element(self, elem: ET.Element) -> Optional[Marker]:
"""Parse any marker element (<marker> or <chapter-marker>). """Parse any marker element (<marker> or <chapter-marker>).
@@ -337,6 +351,7 @@ class FCPXMLParser:
lane=lane, offset=offset, source_start=start, lane=lane, offset=offset, source_start=start,
media_path=media_path, clip_type=elem.tag, role=role, media_path=media_path, clip_type=elem.tag, role=role,
ref_id=ref, parent_clip_name=parent_name, ref_id=ref, parent_clip_name=parent_name,
rotation=self._parse_clip_rotation(elem),
) )
connected.markers.extend(self._collect_markers(elem)) connected.markers.extend(self._collect_markers(elem))
+571
View File
@@ -0,0 +1,571 @@
"""Phrase review — the human pass between the AI's decisions and the render.
A voice timeline says *how* every line was spoken; a list of voice actions says
what the model decided to do about it. Neither is reviewable on its own: the
timeline has no editorial intent, and the action list is a set of timecodes with
no text attached. This module joins them into the one view an editor can
actually judge — the script, phrase by phrase, each carrying the decision that
was made about it.
The phrase is the unit on purpose. Emphasis, in this pipeline, is not a property
of a word but of a line: an emphasized phrase gets a punch-in and a dynamic
caption, everything else gets a plain caption. Keeping the same granularity in
the review, the JSON, and the render means a toggle in the UI maps to exactly
one editorial outcome, with nothing to reconcile in between.
Trimming stays inside the phrase for the same reason. A line is rarely wrong as
a whole — it has a false start, or a trailing "né" — so each phrase carries a
``trim_start``/``trim_end`` pair that rides on word boundaries. Editing a cut
therefore means picking a word, never hunting for a frame, and a partial cut
from the model arrives as a trim instead of being rounded away.
Round-tripping is the other half of the contract. :func:`build_phrase_review`
derives the review from actions, :func:`phrase_review_to_actions` derives
actions back from the edited review, and everything the editor touched wins over
what was inferred — so re-opening the screen shows what was left there, not a
re-derivation that quietly discards the edits.
"""
import json
from pathlib import Path
from typing import Any, Dict, List, Optional, Sequence, Tuple
from .voice_actions import VoiceAction, merge_cut_ranges, parse_actions
PHRASE_REVIEW_VERSION = "1.0"
# Emphasis is stored 0-3 rather than as a float so the UI, the JSON and the
# render agree on the same discrete decision. The thresholds map the continuous
# `peak_emphasis` of the voice timeline onto those levels when the model gave no
# explicit direction for a phrase.
EMPHASIS_LEVELS = (0, 1, 2, 3)
EMPHASIS_THRESHOLDS = (0.25, 0.45, 0.65)
# Zoom scale applied per emphasis level when the review is turned back into
# actions. Level 0 never produces a zoom. The values stay inside
# voice_actions.MIN_ZOOM_SCALE..MAX_ZOOM_SCALE.
ZOOM_SCALE_BY_LEVEL = {1: 1.15, 2: 1.3, 3: 1.5}
# A phrase only survives if most of it does. Speech boundaries from a transcript
# are approximate, so a cut clipping a fraction of a second off the tail is a
# trim, not a removal — treating that as "phrase deleted" would grey out lines
# that are still fully audible.
CUT_COVERAGE_TO_DEACTIVATE = 0.6
# A punch-in shorter than this has no time to ramp in and back out — the writer
# rejects the window anyway (see the zoom ease-in/ease-out shape), so refusing
# it here turns a silent drop at render time into nothing being placed at all.
MIN_ZOOM_DURATION = 0.4
TRACK_SCRIPT = "roteiro"
TRACK_BACKSTAGE = "bastidor"
TRACKS = (TRACK_SCRIPT, TRACK_BACKSTAGE)
def resolve_source(
source: str, voice_timeline_path: str, extra_dirs: Sequence[str] = ()
) -> str:
"""The playable path for a timeline's ``source``, or "" when it's gone.
The voice timeline stores only the media's *file name* — it is written to be
read by a model, where a machine-specific absolute path is noise. That makes
it useless for opening a preview, so the file is looked up where it can
actually be: beside its own timeline JSON first (that is where
``analyze_voice`` writes it), then in whatever project folders the caller
knows about.
"""
if not source:
return ""
candidate = Path(source)
if candidate.is_absolute() and candidate.is_file():
return str(candidate)
directories = [Path(voice_timeline_path).parent] if voice_timeline_path else []
directories += [Path(d) for d in extra_dirs if d]
for directory in directories:
found = directory / candidate.name
if found.is_file():
return str(found)
return ""
def _overlap(a_start: float, a_end: float, b_start: float, b_end: float) -> float:
"""Seconds shared by two spans (0.0 when they don't touch)."""
return max(0.0, min(a_end, b_end) - max(a_start, b_start))
def _cut_coverage(
start: float, end: float, cuts: Sequence[Tuple[float, float]]
) -> float:
"""Fraction of ``start``-``end`` that falls inside ``cuts`` (0-1)."""
span = end - start
if span <= 0:
return 0.0
removed = sum(_overlap(start, end, c_start, c_end) for c_start, c_end in cuts)
return min(1.0, removed / span)
def snap_to_words(
time: float, words: Sequence[dict], fallback: float, edge: str
) -> float:
"""Move ``time`` onto the nearest word boundary of this phrase.
Trims are expressed by pointing at a word, so a trim handle that landed
mid-word would cut a syllable in half. ``edge`` is ``"in"`` (snap to word
starts) or ``"out"`` (snap to word ends); with no word timings available the
time is left as-is.
"""
boundaries = [
float(word.get("start" if edge == "in" else "end", 0.0)) for word in words
]
boundaries = [b for b in boundaries if b > 0]
if not boundaries:
return fallback
return min(boundaries, key=lambda b: abs(b - time))
def _trim_from_cuts(
start: float,
end: float,
words: Sequence[dict],
cuts: Sequence[Tuple[float, float]],
) -> Tuple[float, float]:
"""Read a partial cut over this phrase as a head/tail trim.
Only cuts that touch an edge become trims: a cut carved out of the middle of
a line has no representation here (the phrase is the unit), so it is left
for the whole-phrase coverage rule to decide.
"""
trim_start, trim_end = start, end
for cut_start, cut_end in cuts:
if _overlap(start, end, cut_start, cut_end) <= 0:
continue
if cut_start <= trim_start < cut_end < end:
trim_start = snap_to_words(cut_end, words, cut_end, "in")
if start < cut_start < trim_end <= cut_end:
trim_end = snap_to_words(cut_start, words, cut_start, "out")
if trim_end <= trim_start:
return start, end
return trim_start, trim_end
def _level_from_peak(peak: float) -> int:
"""Map a 0-1 ``peak_emphasis`` onto a 0-3 level."""
for level, threshold in enumerate(EMPHASIS_THRESHOLDS):
if peak < threshold:
return level
return 3
def _level_from_scale(scale: Optional[float]) -> int:
"""Map a zoom's scale factor back onto a 0-3 level.
The model is free to send any scale inside the allowed range, so this picks
the nearest level rather than requiring one of our own three values.
"""
if scale is None:
return 2
best = 1
smallest = None
for level, level_scale in ZOOM_SCALE_BY_LEVEL.items():
distance = abs(level_scale - float(scale))
if smallest is None or distance < smallest:
smallest, best = distance, level
return best
def _emphasis_from_actions(
start: float,
end: float,
actions: Sequence[VoiceAction],
) -> Tuple[Optional[int], str]:
"""The level the model asked for on this phrase, and why.
A ``zoom`` or ``text`` action anywhere inside the phrase is read as "this
line is the emphasis" — the model places them on the word that carries the
point, not on the whole line, so requiring a full-span match would find
nothing. Returns ``(None, "")`` when no action touches the phrase.
"""
level: Optional[int] = None
reason = ""
for action in actions:
if action.kind not in ("zoom", "text"):
continue
if _overlap(start, end, action.start, action.end) <= 0:
continue
if action.kind == "zoom":
candidate = _level_from_scale(action.params.get("scale"))
else:
candidate = 2
if level is None or candidate > level:
level = candidate
reason = action.reason
return level, reason
def _cut_reason(
start: float, end: float, actions: Sequence[VoiceAction]
) -> str:
"""The reason given for the cut that removes this phrase."""
for action in actions:
if action.kind != "cut":
continue
if _overlap(start, end, action.start, action.end) > 0 and action.reason:
return action.reason
return ""
def build_phrase_review(
timeline: dict,
actions: Any = None,
voice_timeline_path: str = "",
extra_dirs: Sequence[str] = (),
) -> dict:
"""Join a voice timeline with the AI's actions into a reviewable script.
``actions`` accepts whatever :func:`~.voice_actions.parse_actions` accepts —
a bare list, ``{"actions": [...]}``, or ``None`` when there is no AI pass and
the review starts from the acoustics alone. Malformed rows are skipped and
reported in ``errors`` rather than raising, matching the rest of the
decision pipeline.
"""
parsed, errors = parse_actions(actions) if actions else ([], [])
cuts = merge_cut_ranges(parsed)
phrases: List[dict] = []
for index, segment in enumerate(timeline.get("segments", [])):
start = float(segment.get("start", 0.0))
end = float(segment.get("end", 0.0))
peak = float(segment.get("peak_emphasis", 0.0))
take_boundary = bool(segment.get("take_boundary", False))
words = list(segment.get("words", []))
coverage = _cut_coverage(start, end, cuts)
active = coverage < CUT_COVERAGE_TO_DEACTIVATE
trim_start, trim_end = (
_trim_from_cuts(start, end, words, cuts) if active else (start, end)
)
asked_level, asked_reason = _emphasis_from_actions(start, end, parsed)
if asked_level is not None:
emphasis, reason = asked_level, asked_reason
else:
emphasis = _level_from_peak(peak)
reason = f"ênfase {peak:.2f}" if emphasis else ""
if not active:
# A removed line carries the reason it was removed; the emphasis it
# would have had is kept so re-activating it restores the decision.
reason = _cut_reason(start, end, parsed) or reason
phrases.append(
{
"index": index,
"start": round(start, 3),
"end": round(end, 3),
"trim_start": round(trim_start, 3),
"trim_end": round(trim_end, 3),
"text": str(segment.get("text", "")).strip(),
"speaker": str(segment.get("speaker", "")),
"active": active,
"emphasis": emphasis,
"track": TRACK_BACKSTAGE if (not active and take_boundary) else TRACK_SCRIPT,
"peak_emphasis": round(peak, 3),
# Delivery emotion is a heuristic over the acoustics (see
# voice_timeline._emotion_for_word) and only means anything when
# the analysis actually ran — `emotion_available` below is what
# separates "spoken flat" from "never measured".
"emotion": str(segment.get("emotion", "neutral")),
"emotion_confidence": round(
float(segment.get("emotion_confidence", 0.0)), 3
),
"take_boundary": take_boundary,
"gap_before": round(float(segment.get("gap_before", 0.0)), 3),
"reason": reason,
"words": [
{
"text": str(word.get("text", "")),
"start": round(float(word.get("start", 0.0)), 3),
"end": round(float(word.get("end", 0.0)), 3),
"energy": round(float(word.get("energy", 0.0)), 3),
"emphasis": round(float(word.get("emphasis", 0.0)), 3),
}
for word in words
],
}
)
source = timeline.get("source", "")
layers = timeline.get("layers", {}) if isinstance(timeline.get("layers"), dict) else {}
return {
"version": PHRASE_REVIEW_VERSION,
"source": source,
"source_path": resolve_source(source, voice_timeline_path, extra_dirs),
"rotation": float(timeline.get("rotation", 0.0)),
"duration": round(phrases[-1]["end"], 3) if phrases else 0.0,
"speakers": timeline.get("speakers", []),
"emotion_available": bool(layers.get("emotion", False)),
"phrases": phrases,
# Punch-ins the editor places by hand on an arbitrary range, alongside
# the whole-phrase zoom that an emphasis level produces. Both end up as
# zoom actions; this one exists because the moment worth punching into
# is not always a whole sentence.
"zooms": [],
"errors": errors,
}
def _coerce_zoom(raw: Any) -> Optional[Dict[str, float]]:
"""Normalize one manually placed zoom range."""
if not isinstance(raw, dict):
return None
try:
start = float(raw.get("start"))
end = float(raw.get("end"))
except (TypeError, ValueError):
return None
if end - start < MIN_ZOOM_DURATION:
return None
return {"start": start, "end": end}
def _coerce_phrase(raw: Any, index: int) -> Optional[Dict[str, Any]]:
"""Normalize one edited phrase row coming back from the UI."""
if not isinstance(raw, dict):
return None
try:
start = float(raw.get("start"))
end = float(raw.get("end"))
except (TypeError, ValueError):
return None
if end <= start:
return None
try:
emphasis = int(raw.get("emphasis", 0))
except (TypeError, ValueError):
emphasis = 0
try:
trim_start = float(raw.get("trim_start", start))
trim_end = float(raw.get("trim_end", end))
except (TypeError, ValueError):
trim_start, trim_end = start, end
# A trim that escaped the phrase, or inverted, is treated as no trim at all:
# the UI is the only thing that writes these, and silently discarding a bad
# pair keeps a rounding slip from deleting material the editor kept.
if not (start <= trim_start < trim_end <= end):
trim_start, trim_end = start, end
track = str(raw.get("track", TRACK_SCRIPT))
return {
"index": int(raw.get("index", index)),
"start": start,
"end": end,
"trim_start": trim_start,
"trim_end": trim_end,
"text": str(raw.get("text", "")).strip(),
"speaker": str(raw.get("speaker", "")),
"active": bool(raw.get("active", True)),
"emphasis": min(3, max(0, emphasis)),
"track": track if track in TRACKS else TRACK_SCRIPT,
"reason": str(raw.get("reason", "")),
}
def phrase_review_to_actions(review: dict) -> dict:
"""Turn an edited review back into the action list the applier consumes.
Every deactivated phrase becomes a ``cut``, a trimmed one becomes a cut over
the head and/or tail it lost, and every emphasized one becomes a ``zoom``
scaled by its level. The emphasis flags ride along in ``emphasis_spans`` so
the caption step can give those lines the dynamic treatment and everything
else the plain one, without re-deriving the decision from the acoustics.
"""
phrases = [
coerced
for index, raw in enumerate(review.get("phrases", []))
if (coerced := _coerce_phrase(raw, index)) is not None
]
actions: List[dict] = []
emphasis_spans: List[dict] = []
inactive_run: List[dict] = []
def flush_inactive_run() -> None:
"""One cut per RUN of consecutive deactivated phrases, not one per
phrase. A phrase-by-phrase cut leaves the pause BETWEEN two
deactivated phrases uncut — that gap was never anyone's content, so
nothing asked for it to survive, but it does anyway: a 0.1-0.5s
sliver clip in the final timeline for every such gap. Spanning the
whole run absorbs those gaps into the one cut."""
if not inactive_run:
return
if len(inactive_run) == 1:
reason = inactive_run[0]["reason"] or "desativada na revisão"
else:
reason = (
f"desativadas na revisão ({len(inactive_run)} frases): "
+ "; ".join(p["text"][:40] for p in inactive_run if p["text"])
)
actions.append(
VoiceAction(
kind="cut",
start=inactive_run[0]["start"],
end=inactive_run[-1]["end"],
reason=reason,
speaker=inactive_run[0]["speaker"],
).as_dict()
)
inactive_run.clear()
for phrase in phrases:
if not phrase["active"]:
inactive_run.append(phrase)
continue
flush_inactive_run()
# Head and tail the editor trimmed off — each becomes its own cut, so a
# false start disappears without taking the line with it.
for trim_start, trim_end, where in (
(phrase["start"], phrase["trim_start"], "início"),
(phrase["trim_end"], phrase["end"], "fim"),
):
if trim_end - trim_start <= 0:
continue
actions.append(
VoiceAction(
kind="cut",
start=trim_start,
end=trim_end,
reason=f"trecho do {where} da frase removido na revisão",
speaker=phrase["speaker"],
).as_dict()
)
if phrase["emphasis"] >= 1:
actions.append(
VoiceAction(
kind="zoom",
start=phrase["trim_start"],
end=phrase["trim_end"],
params={"scale": ZOOM_SCALE_BY_LEVEL[phrase["emphasis"]]},
reason=phrase["reason"] or f"ênfase nível {phrase['emphasis']}",
speaker=phrase["speaker"],
).as_dict()
)
emphasis_spans.append(
{
"start": phrase["trim_start"],
"end": phrase["trim_end"],
"level": phrase["emphasis"],
"text": phrase["text"],
}
)
flush_inactive_run()
# Hand-placed punch-ins carry no scale on purpose: an omitted scale lets the
# applier use the shape configured in "Análise de Voz" (zoom_scale, ease in
# and out), so changing that setting restyles every manual zoom instead of
# leaving a scale frozen into each one at the moment it was drawn.
for raw in review.get("zooms", []):
zoom = _coerce_zoom(raw)
if zoom is None:
continue
actions.append(
VoiceAction(
kind="zoom",
start=zoom["start"],
end=zoom["end"],
reason="zoom marcado na revisão",
).as_dict()
)
return {
"source": review.get("source", ""),
"actions": actions,
"emphasis_spans": emphasis_spans,
}
def merge_saved_decisions(review: dict, saved: Optional[dict]) -> dict:
"""Lay a previously saved review's decisions over a freshly built one.
Only the editorial fields travel — active, emphasis, track, text, trims.
Everything else (words, emotion, energy) is re-derived from the current
analysis, so re-running the voice pass with better settings improves the
screen instead of being masked by a stale copy of itself, and the saved file
never has to carry a duplicate of data it does not own.
Phrases are matched by index *and* start time: if the analysis changed
enough to move a line, the old decision for that slot is dropped rather than
applied to a different sentence.
"""
if not saved:
return review
review["zooms"] = [
zoom for raw in saved.get("zooms", []) if (zoom := _coerce_zoom(raw)) is not None
]
by_index = {}
for raw in saved.get("phrases", []):
if isinstance(raw, dict) and "index" in raw:
by_index[raw["index"]] = raw
for phrase in review["phrases"]:
previous = by_index.get(phrase["index"])
if previous is None:
continue
if abs(float(previous.get("start", -1)) - phrase["start"]) > 0.25:
continue
phrase["active"] = bool(previous.get("active", phrase["active"]))
phrase["emphasis"] = min(3, max(0, int(previous.get("emphasis", phrase["emphasis"]))))
track = str(previous.get("track", phrase["track"]))
phrase["track"] = track if track in TRACKS else phrase["track"]
if previous.get("text"):
phrase["text"] = str(previous["text"])
trim_start = float(previous.get("trim_start", phrase["trim_start"]))
trim_end = float(previous.get("trim_end", phrase["trim_end"]))
if phrase["start"] <= trim_start < trim_end <= phrase["end"]:
phrase["trim_start"], phrase["trim_end"] = trim_start, trim_end
return review
def review_paths(voice_timeline_path: str) -> Tuple[Path, Path]:
"""Where the review and its derived actions live, next to the timeline.
Both files sit beside the ``_voice_timeline.json`` they came from and are
named after it, so a project folder stays readable and re-running the wizard
on the same take overwrites its own files instead of accumulating copies.
"""
base = Path(voice_timeline_path)
stem = base.stem
if stem.endswith("_voice_timeline"):
stem = stem[: -len("_voice_timeline")]
return (
base.with_name(f"{stem}_phrase_review.json"),
base.with_name(f"{stem}_phrase_actions.json"),
)
def save_phrase_review(voice_timeline_path: str, review: dict) -> Tuple[Path, Path]:
"""Write the edited review and the actions derived from it. Returns both paths."""
review_path, actions_path = review_paths(voice_timeline_path)
review_path.write_text(
json.dumps(review, ensure_ascii=False, indent=2), encoding="utf-8"
)
actions_path.write_text(
json.dumps(phrase_review_to_actions(review), ensure_ascii=False, indent=2),
encoding="utf-8",
)
return review_path, actions_path
def load_phrase_review(voice_timeline_path: str) -> Optional[dict]:
"""The review saved earlier for this timeline, or ``None`` if there is none."""
review_path, _ = review_paths(voice_timeline_path)
if not review_path.is_file():
return None
try:
data = json.loads(review_path.read_text(encoding="utf-8"))
except (OSError, json.JSONDecodeError):
return None
return data if isinstance(data, dict) else None
+94 -11
View File
@@ -28,8 +28,11 @@ ALLOWED_MODELS = (
) )
# Conservative by default: interjections that are near-universally filler. # Conservative by default: interjections that are near-universally filler.
# Portuguese "um"/"uma" are usually articles/numerals inside real phrases
# ("de um jeito") rather than discardable hesitations, so only cut them when
# the caller explicitly opts in through the fillers argument.
# "like" / "so" / "actually" are speech, not noise, unless the user opts in. # "like" / "so" / "actually" are speech, not noise, unless the user opts in.
DEFAULT_FILLERS = ("um", "uh", "uhh", "umm", "erm", "ehm", "mmm", "hmm", "mhm") DEFAULT_FILLERS = ("uh", "uhh", "umm", "erm", "ehm", "mmm", "hmm", "mhm")
_NORM_RE = re.compile(r"[^\w']+") _NORM_RE = re.compile(r"[^\w']+")
@@ -122,18 +125,27 @@ def transcribe(
model_size: str = "base", model_size: str = "base",
language: Optional[str] = None, language: Optional[str] = None,
progress_cb: Optional[Callable[[float], None]] = None, progress_cb: Optional[Callable[[float], None]] = None,
align: bool = True,
) -> Optional[dict]: ) -> Optional[dict]:
"""Transcribe an audio/video file locally with word-level timestamps. """Transcribe an audio/video file locally with word-level timestamps.
Requires the optional ``[transcribe]`` extra (faster-whisper). Returns Requires the optional ``[transcribe]`` extra (faster-whisper). Returns
``None`` when the model is unavailable or the file is missing/unreadable. ``None`` when the model is unavailable or the file is missing/unreadable.
When ``align`` is true (default) and the optional ``whisperx`` dependency is
present, word timestamps are refined by phonetic forced alignment, which
corrects faster-whisper's systematic ~0.3-0.5s early bias on word *starts*
(see ``Engine/docs/05_EXPERIENCIAS.md`` #14). The transcript reports
whether this ran via the ``alignment`` flag, so downstream consumers can
rely on the times without re-measuring.
The model weights are resolved from the configured models directory (see The model weights are resolved from the configured models directory (see
``model_manager.get_models_dir``), so a model selected/downloaded through ``model_manager.get_models_dir``), so a model selected/downloaded through
the app is found without an implicit download to the default HF cache. the app is found without an implicit download to the default HF cache.
Returns: Returns:
``{"language": str, "duration": float, "text": str, ``{"language": str, "duration": float, "text": str,
"alignment": bool,
"segments": [{"text", "start", "end", "start_fmt", "end_fmt"}, ...], "segments": [{"text", "start", "end", "start_fmt", "end_fmt"}, ...],
"words": [{"word", "start", "end", "confidence"}, ...]}`` "words": [{"word", "start", "end", "confidence"}, ...]}``
""" """
@@ -172,6 +184,7 @@ def transcribe(
vad_filter=True, vad_filter=True,
) )
segments: List[dict] = [] segments: List[dict] = []
raw_segments: List[dict] = []
words: List[dict] = [] words: List[dict] = []
# `info.duration` is known upfront (from the container), so each # `info.duration` is known upfront (from the container), so each
# segment's end time — yielded lazily as faster-whisper decodes — # segment's end time — yielded lazily as faster-whisper decodes —
@@ -180,6 +193,20 @@ def transcribe(
for seg in segments_iter: for seg in segments_iter:
start = float(seg.start) start = float(seg.start)
end = float(seg.end) end = float(seg.end)
seg_words: List[dict] = []
if progress_cb is not None and total_duration > 0:
progress_cb(min(end / total_duration, 1.0))
for w in seg.words or []:
ws = float(w.start)
we = float(w.end)
word = {
"word": w.word.strip(),
"start": ws,
"end": we,
"confidence": float(w.probability),
}
words.append(word)
seg_words.append(word)
segments.append( segments.append(
{ {
"text": seg.text.strip(), "text": seg.text.strip(),
@@ -189,19 +216,26 @@ def transcribe(
"end_fmt": format_timestamp(end), "end_fmt": format_timestamp(end),
} }
) )
if progress_cb is not None and total_duration > 0: raw_segments.append(
progress_cb(min(end / total_duration, 1.0))
for w in seg.words or []:
ws = float(w.start)
we = float(w.end)
words.append(
{ {
"word": w.word.strip(), "text": seg.text.strip(),
"start": ws, "start": start,
"end": we, "end": end,
"confidence": float(w.probability), "words": seg_words,
} }
) )
alignment_ran = False
if align and raw_segments:
from .forced_align import ForcedAligner
try:
words = ForcedAligner().align(
words, raw_segments, str(file_path), info.language, str(models_dir)
)
alignment_ran = True
except Exception:
logger.warning("forced alignment step failed; keeping raw timestamps")
except Exception: except Exception:
logger.warning("whisper transcription failed for %s", file_path) logger.warning("whisper transcription failed for %s", file_path)
return None return None
@@ -209,6 +243,7 @@ def transcribe(
"language": info.language, "language": info.language,
"duration": float(info.duration), "duration": float(info.duration),
"text": " ".join(s["text"] for s in segments), "text": " ".join(s["text"] for s in segments),
"alignment": alignment_ran,
"segments": segments, "segments": segments,
"words": words, "words": words,
} }
@@ -281,6 +316,54 @@ def group_words_by_segment(
return groups return groups
def split_into_subphrases(
words: Sequence[dict],
min_words: int = 3,
) -> List[List[dict]]:
"""Split a sentence's *words* into sub-phrases at comma boundaries.
A comma is where a spoken sentence actually breathes, so it is the
natural seam for grouping subtitles — each sub-phrase becoming its own
on-screen block (and, downstream, its own compound clip).
The exception is the short tail: a fragment like "né?" or "Então..."
reads as part of the phrase before it, not as a phrase of its own, and
promoting it to its own block would flash a single word on screen. So a
piece shorter than *min_words* is merged back into its neighbour —
preferring the previous piece, falling back to the next one when the
short piece leads the sentence.
Returns one group per sub-phrase; a sentence with no comma comes back
as a single group.
"""
pieces: List[List[dict]] = []
current: List[dict] = []
for w in words:
current.append(w)
text = str(w.get('word') or w.get('text') or '')
if text.rstrip().endswith(','):
pieces.append(current)
current = []
if current:
pieces.append(current)
if len(pieces) <= 1:
return pieces
merged: List[List[dict]] = []
for piece in pieces:
if len(piece) < min_words and merged:
merged[-1].extend(piece)
else:
merged.append(piece)
# A short leading piece has no previous neighbour to fold into, so it
# folds forward instead.
if len(merged) > 1 and len(merged[0]) < min_words:
merged[1][:0] = merged[0]
merged.pop(0)
return merged
def segments_to_srt(segments: Sequence[dict]) -> str: def segments_to_srt(segments: Sequence[dict]) -> str:
"""Render transcript segments as an SRT string (for captions import).""" """Render transcript segments as an SRT string (for captions import)."""
+16 -1
View File
@@ -82,8 +82,9 @@ def _validate_one(raw: Any, index: int) -> Tuple[Optional[VoiceAction], str]:
params = dict(params) if isinstance(params, dict) else {} params = dict(params) if isinstance(params, dict) else {}
if kind == "zoom": if kind == "zoom":
if "scale" in params and params.get("scale") is not None:
try: try:
scale = float(params.get("scale", 1.3)) scale = float(params["scale"])
except (TypeError, ValueError): except (TypeError, ValueError):
return None, f"{where}: zoom scale must be a number" return None, f"{where}: zoom scale must be a number"
if not (MIN_ZOOM_SCALE <= scale <= MAX_ZOOM_SCALE): if not (MIN_ZOOM_SCALE <= scale <= MAX_ZOOM_SCALE):
@@ -97,6 +98,20 @@ def _validate_one(raw: Any, index: int) -> Tuple[Optional[VoiceAction], str]:
if not content: if not content:
return None, f"{where}: text action needs params.content" return None, f"{where}: text action needs params.content"
params["content"] = content[:MAX_TEXT_LENGTH] params["content"] = content[:MAX_TEXT_LENGTH]
# Style is optional — omitted fields fall back to the "Legendas
# Dinâmicas" emphasis style at apply time (see _apply_placed_action),
# so a callout matches the captions' look without the caller having
# to know or repeat that configuration. Anything given here wins.
for key in ("font", "font_color", "face"):
if key in params and not isinstance(params[key], str):
del params[key]
if "font_size" in params:
try:
params["font_size"] = int(params["font_size"])
except (TypeError, ValueError):
del params["font_size"]
if "bold" in params:
params["bold"] = bool(params["bold"])
return ( return (
VoiceAction( VoiceAction(
+95
View File
@@ -56,12 +56,20 @@ VALUE_SCALES = {
"rate_delta": "0-1, how much the local speaking rate departs from the average", "rate_delta": "0-1, how much the local speaking rate departs from the average",
"pause_before": "seconds of silence immediately before the word", "pause_before": "seconds of silence immediately before the word",
"emphasis": "0-1 combined index; high values are punch-in/highlight candidates", "emphasis": "0-1 combined index; high values are punch-in/highlight candidates",
"emotion": "heuristic label from delivery: neutral, excited, tense, calm, reflective",
"emotion_confidence": "0-1 confidence in the heuristic emotion label",
"arousal": "0-1 vocal activation from energy/rate/pitch movement",
"valence": "0-1 rough positive tone; lower values suggest tension/weight",
}, },
"segment": { "segment": {
"gap_before": "seconds of silence before this line", "gap_before": "seconds of silence before this line",
"take_boundary": "true when the gap is long enough that the take likely restarted here", "take_boundary": "true when the gap is long enough that the take likely restarted here",
"avg_energy": "0-1 mean loudness across the line", "avg_energy": "0-1 mean loudness across the line",
"peak_emphasis": "0-1 highest emphasis of any word in the line", "peak_emphasis": "0-1 highest emphasis of any word in the line",
"emotion": "dominant delivery emotion across the line",
"emotion_confidence": "0-1 confidence in the dominant segment emotion",
"arousal": "0-1 mean vocal activation across the line",
"valence": "0-1 mean rough positive tone across the line",
}, },
} }
@@ -92,11 +100,74 @@ def _round_word(word: dict) -> dict:
"rate_delta": round(word.get("rate_delta", 0.0), 3), "rate_delta": round(word.get("rate_delta", 0.0), 3),
"pause_before": round(word.get("pause_before", 0.0), 3), "pause_before": round(word.get("pause_before", 0.0), 3),
"emphasis": round(word.get("emphasis", 0.0), 3), "emphasis": round(word.get("emphasis", 0.0), 3),
"emotion": word.get("emotion", "neutral"),
"emotion_confidence": round(word.get("emotion_confidence", 0.0), 3),
"arousal": round(word.get("arousal", 0.0), 3),
"valence": round(word.get("valence", 0.5), 3),
"energy_raw": word.get("energy"), "energy_raw": word.get("energy"),
"pitch_hz": word.get("pitch_hz"), "pitch_hz": word.get("pitch_hz"),
} }
def _emotion_for_word(word: dict, enabled: bool, sensitivity: float) -> dict:
"""Classify delivery emotion from normalized acoustic features.
This is deliberately a local heuristic rather than a claimed clinical
emotion model. It gives the editor a useful signal about delivery shape
while degrading predictably when acoustic extraction is unavailable.
"""
if not enabled:
return {
"emotion": "neutral",
"emotion_confidence": 0.0,
"arousal": 0.0,
"valence": 0.5,
}
energy = float(word.get("energy_norm", 0.0))
pitch = float(word.get("pitch_delta", 0.0))
rate = float(word.get("rate_delta", 0.0))
pause = min(float(word.get("pause_before", 0.0)) / 2.0, 1.0)
emphasis = float(word.get("emphasis", 0.0))
arousal = max(0.0, min(1.0, energy * 0.45 + pitch * 0.25 + rate * 0.20 + emphasis * 0.10))
valence = max(0.0, min(1.0, 0.55 + energy * 0.15 - pause * 0.20 - rate * 0.10))
if arousal >= 0.68 and valence >= 0.50:
label = "excited"
confidence = arousal
elif arousal >= 0.58 and valence < 0.50:
label = "tense"
confidence = max(arousal, 1.0 - valence)
elif arousal <= 0.28 and pause >= 0.25:
label = "reflective"
confidence = max(1.0 - arousal, pause)
elif arousal <= 0.35:
label = "calm"
confidence = 1.0 - arousal
else:
label = "neutral"
confidence = 1.0 - abs(arousal - 0.5) * 2.0
confidence = max(0.0, min(1.0, confidence))
if confidence < sensitivity:
label = "neutral"
return {
"emotion": label,
"emotion_confidence": confidence,
"arousal": arousal,
"valence": valence,
}
def annotate_emotions(words: Sequence[dict], enabled: bool, sensitivity: float) -> List[dict]:
"""Attach heuristic emotion labels to enriched word rows."""
return [
{**w, **_emotion_for_word(w, enabled, sensitivity)}
for w in words
]
def enrich_words( def enrich_words(
words: Sequence[dict], words: Sequence[dict],
pitch_track: Optional[Sequence] = None, pitch_track: Optional[Sequence] = None,
@@ -166,6 +237,13 @@ def _segment_rows(segments: Sequence[dict], words: Sequence[dict]) -> List[dict]
in_seg = [w for w in words if start <= float(w.get("start", 0.0)) < end] in_seg = [w for w in words if start <= float(w.get("start", 0.0)) < end]
energies = [w["energy_norm"] for w in in_seg] energies = [w["energy_norm"] for w in in_seg]
emphases = [w["emphasis"] for w in in_seg] emphases = [w["emphasis"] for w in in_seg]
arousals = [w.get("arousal", 0.0) for w in in_seg]
valences = [w.get("valence", 0.5) for w in in_seg]
emotions = [w.get("emotion", "neutral") for w in in_seg]
dominant = max(set(emotions), key=emotions.count) if emotions else "neutral"
emotion_confidences = [
w.get("emotion_confidence", 0.0) for w in in_seg if w.get("emotion") == dominant
]
gap = max(0.0, start - previous_end) gap = max(0.0, start - previous_end)
rows.append( rows.append(
{ {
@@ -181,6 +259,13 @@ def _segment_rows(segments: Sequence[dict], words: Sequence[dict]) -> List[dict]
"take_boundary": gap >= TAKE_BOUNDARY_GAP, "take_boundary": gap >= TAKE_BOUNDARY_GAP,
"avg_energy": round(sum(energies) / len(energies), 3) if energies else 0.0, "avg_energy": round(sum(energies) / len(energies), 3) if energies else 0.0,
"peak_emphasis": round(max(emphases), 3) if emphases else 0.0, "peak_emphasis": round(max(emphases), 3) if emphases else 0.0,
"emotion": dominant,
"emotion_confidence": (
round(sum(emotion_confidences) / len(emotion_confidences), 3)
if emotion_confidences else 0.0
),
"arousal": round(sum(arousals) / len(arousals), 3) if arousals else 0.0,
"valence": round(sum(valences) / len(valences), 3) if valences else 0.5,
"words": [_round_word(w) for w in in_seg], "words": [_round_word(w) for w in in_seg],
} }
) )
@@ -425,6 +510,9 @@ def build_voice_timeline(
weights: EmphasisWeights = EmphasisWeights(), weights: EmphasisWeights = EmphasisWeights(),
peak_percentile: float = 0.02, peak_percentile: float = 0.02,
emphasis_floor: float = 0.25, emphasis_floor: float = 0.25,
emotion_enabled: bool = False,
emotion_sensitivity: float = 0.5,
rotation: float = 0.0,
progress_cb: Optional[Callable[[float, str], None]] = None, progress_cb: Optional[Callable[[float, str], None]] = None,
) -> dict: ) -> dict:
"""Build the consolidated voice timeline for one media file. """Build the consolidated voice timeline for one media file.
@@ -445,6 +533,7 @@ def build_voice_timeline(
report(0.5, "Calculando ênfase...") report(0.5, "Calculando ênfase...")
words = enrich_words(transcript.get("words", []), pitch_track, energy_track, weights) words = enrich_words(transcript.get("words", []), pitch_track, energy_track, weights)
words = annotate_emotions(words, emotion_enabled, emotion_sensitivity)
report(0.7, "Identificando participantes...") report(0.7, "Identificando participantes...")
tracks = diarize(media_path, hf_token, num_speakers) if hf_token else None tracks = diarize(media_path, hf_token, num_speakers) if hf_token else None
@@ -457,6 +546,10 @@ def build_voice_timeline(
return { return {
"version": VOICE_TIMELINE_VERSION, "version": VOICE_TIMELINE_VERSION,
"source": Path(media_path).name, "source": Path(media_path).name,
# Edit-time correction from the clip's Transform filter in the FCPXML
# (e.g. straightening a tilted phone shot) — 0.0 when the clip has none
# or the caller didn't resolve one.
"rotation": rotation,
"language": transcript.get("language", ""), "language": transcript.get("language", ""),
# What actually ran, not what was installed — a consumer must be able # What actually ran, not what was installed — a consumer must be able
# to tell "this speech is flat" from "the acoustics never loaded", # to tell "this speech is flat" from "the acoustics never loaded",
@@ -465,6 +558,8 @@ def build_voice_timeline(
"transcript": bool(transcript.get("words")), "transcript": bool(transcript.get("words")),
"acoustics": pitch_track is not None or energy_track is not None, "acoustics": pitch_track is not None or energy_track is not None,
"speakers": tracks is not None, "speakers": tracks is not None,
"emotion": bool(emotion_enabled),
"alignment": bool(transcript.get("alignment")),
}, },
"scales": VALUE_SCALES, "scales": VALUE_SCALES,
"summary": _summary( "summary": _summary(
File diff suppressed because it is too large Load Diff
+130
View File
@@ -0,0 +1,130 @@
"""
FCPXML Writer — Generate and modify Final Cut Pro XML files.
This package provides two complementary workflows for working with FCPXML:
**Generation** (``FCPXMLWriter``, in :mod:`.generator`):
Build a new FCPXML document from Python dataclass objects (``Project``,
``Timeline``, ``Clip``, ``Marker``). Useful for creating rough cuts,
montage exports, and template-based projects.
**Modification** (``FCPXMLModifier``, in :mod:`.modifier`):
Load an existing FCPXML file, apply surgical edits (markers, trims,
reorders, transitions, speed changes, silence removal, etc.), and save.
This is the primary API used by the MCP server's tool handlers.
Layout
------
This was one 4.200-line module. It is now one module per subject, because the
subjects barely touch each other: whoever is fixing a zoom ramp has no reason
to scroll past subtitle layout to find it.
helpers sanitising, scales, shared element builders
document asset creation, timebases, serialisation (``write_fcpxml``)
validation structural checks (``validate_fcpxml``)
core ``ModifierCore``: load, indices, spine navigation, ``save``
<subject> one mixin per editing subject (markers, trim, speed, …)
modifier ``FCPXMLModifier`` = core + every mixin
generator ``FCPXMLWriter``
api one-line convenience wrappers
Everything the rest of the project imported from the old module is re-exported
here, so ``from fcpxml.writer import FCPXMLModifier`` keeps working unchanged —
including the underscore-prefixed helpers the test suite reaches for.
Architecture notes
------------------
- All time arithmetic uses ``TimeValue`` (rational fractions) — never floats —
to match FCPXML's native ``"600/2400s"`` format and avoid rounding drift.
- The ``FCPXMLModifier`` builds three in-memory indices at init
(``clips``, ``resources``, ``formats``) so lookups are O(1) by ID/name.
- Spine-based editing: clips live inside a ``<spine>`` element (the primary
storyline). Connected clips attach via ``lane`` attributes on spine clips.
Most editing methods find the target clip in the spine, mutate it, then
ripple offsets on subsequent siblings.
- ``write_fcpxml()`` handles DTD-compliant serialisation and optional
timebase enforcement for all output paths.
"""
from ..models import TimeValue
from .api import add_marker_to_file, modify_fcpxml, trim_clip_in_file
from .core import ModifierCore
from .document import (
_STILL_IMAGE_EXTENSIONS,
_enforce_standard_timebases,
_ensure_video_asset,
write_fcpxml,
)
from .generator import FCPXMLWriter
from .helpers import (
_ASSET_CLIP_CHILD_ORDER,
_CHILD_ORDER_INDEX,
_MAX_MARKER_NAME_LENGTH,
_MAX_NOTE_LENGTH,
CLIP_AND_AUDIO_TAGS,
CLIP_TAGS,
FCP_EFFECTS,
HOLD_AT_CUT_THRESHOLD,
SPINE_ELEMENT_TAGS,
START_AT_CUT_THRESHOLD,
_create_asset_element,
_dtd_insert,
_fmt_scale,
_probe_audio_info,
_sanitize_xml_value,
build_marker_element,
list_effects,
)
from .modifier import FCPXMLModifier
from .validation import (
_check_asset_sources,
_check_child_order,
_check_effect_refs,
_check_frame_alignment,
_check_required_attributes,
_check_timebases,
_document_frame_duration,
validate_fcpxml,
)
__all__ = [
"FCPXMLModifier",
"FCPXMLWriter",
"ModifierCore",
"TimeValue",
"FCP_EFFECTS",
"CLIP_TAGS",
"CLIP_AND_AUDIO_TAGS",
"SPINE_ELEMENT_TAGS",
"HOLD_AT_CUT_THRESHOLD",
"START_AT_CUT_THRESHOLD",
"add_marker_to_file",
"build_marker_element",
"list_effects",
"modify_fcpxml",
"trim_clip_in_file",
"validate_fcpxml",
"write_fcpxml",
# Internos que o resto do projeto (e a suíte) já importava deste módulo
# quando ele era um arquivo só. Ficam aqui para a divisão não virar uma
# quebra de API disfarçada de reorganização.
"_ASSET_CLIP_CHILD_ORDER",
"_CHILD_ORDER_INDEX",
"_MAX_MARKER_NAME_LENGTH",
"_MAX_NOTE_LENGTH",
"_STILL_IMAGE_EXTENSIONS",
"_check_asset_sources",
"_check_child_order",
"_check_effect_refs",
"_check_frame_alignment",
"_check_required_attributes",
"_check_timebases",
"_create_asset_element",
"_document_frame_duration",
"_dtd_insert",
"_enforce_standard_timebases",
"_ensure_video_asset",
"_fmt_scale",
"_probe_audio_info",
"_sanitize_xml_value",
]
+140
View File
@@ -0,0 +1,140 @@
"""Clip de ajuste (adjustment layer) — criação do elemento FCPXML.
No Final Cut, uma "camada de ajuste" é um ``<clip>`` que carrega filtros
(``filter-video`` / ``filter-audio``) diretamente como filhos — o DTD do
FCPXML 1.13 não define nenhum elemento ``<adjustment>`` como wrapper (ver
``<!ELEMENT clip>`` em ``FCPXMLv1_13.dtd``: ``filter-video``/``filter-audio``
vêm depois de ``audio-channel-source*`` e antes de ``metadata?``, sem
elemento intermediário). Tudo que está abaixo do clip na timeline herda
esses filtros — é como se o efeito fosse aplicado a uma faixa inteira de
uma vez.
Esta classe monta esse elemento a partir de dados de alto nível (duração +
lista de ``EfeitoAjuste``), cuidando de criar os recursos ``<effect>``
correspondentes na seção ``<resources>`` e de referenciá-los pelos filtros.
"""
import xml.etree.ElementTree as ET
from typing import Callable, List, Optional
from ..models.timeline import EfeitoAjuste
from ..models.timing import TimeValue
def _para_racional(tempo) -> str:
"""Aceita ``TimeValue`` ou uma string FCPXML já formatada ("90/30s")."""
if isinstance(tempo, TimeValue):
return tempo.to_fcpxml()
if tempo is None:
return "0/1s"
return str(tempo)
def _id_recurso_unico(resources: ET.Element, prefixo: str = "r_ajuste") -> str:
"""Gera um ``id`` de recurso ainda ausente em ``resources``."""
existentes = {r.get("id") for r in resources.findall("*") if r.get("id")}
contador = 1
while f"{prefixo}_{contador}" in existentes:
contador += 1
return f"{prefixo}_{contador}"
class ClipDeAjuste:
"""Cria um clip de ajuste (adjustment layer) pronto para a spine.
Exemplo::
from fcpxml.models.timing import TimeValue
from fcpxml.models.timeline import EfeitoAjuste, ParametroEfeito
from fcpxml.writer.adjustment import ClipDeAjuste
efeito = EfeitoAjuste(
nome="Color Curves", uid="...UUID...", tipo="video",
parametros=[ParametroEfeito(nome="Amount", valor="0.5",
chave=".../9999")],
)
clip = ClipDeAjuste(
nome="Ajuste de cor",
duracao=TimeValue(300, 30),
efeitos=[efeito],
).criar(resources)
spine.append(clip)
"""
def __init__(
self,
nome: str,
duracao,
efeitos: List[EfeitoAjuste],
offset=None,
formato_tc: str = "NDF",
):
self.nome = nome
self.duracao = duracao
self.efeitos = efeitos
self.offset = offset
self.formato_tc = formato_tc
def criar(
self,
resources: ET.Element,
proximo_id: Optional[Callable[[], str]] = None,
) -> ET.Element:
"""Monta o ``<clip>`` de ajuste e seus recursos ``<effect>``.
``resources`` é a seção ``<resources>`` do documento (onde os
``<effect>`` são registrados). ``proximo_id`` é um gerador opcional
de ids de recurso; sem ele, usa um id único baseado em ``resources``.
"""
def gerar_id() -> str:
if proximo_id:
return proximo_id()
return _id_recurso_unico(resources)
filtros: List[ET.Element] = []
for efeito in self.efeitos:
efeito_id = self._garantir_recurso(resources, efeito, gerar_id)
filtros.append(self._montar_filtro(efeito, efeito_id))
clip = ET.Element(
"clip",
name=self.nome,
duration=_para_racional(self.duracao),
tcFormat=self.formato_tc,
)
if self.offset is not None:
clip.set("offset", _para_racional(self.offset))
# O DTD exige filter-video* antes de filter-audio* como filhos
# diretos do clip (sem wrapper <adjustment>).
for filtro in sorted(filtros, key=lambda f: f.tag != "filter-video"):
clip.append(filtro)
return clip
def _garantir_recurso(
self, resources: ET.Element, efeito: EfeitoAjuste, gerar_id: Callable[[], str]
) -> str:
"""Devolve o ``id`` do ``<effect>`` de *efeito*, criando-o se ausente."""
for existente in resources.findall("effect"):
if existente.get("uid") == efeito.uid:
return existente.get("id")
efeito_id = gerar_id()
recurso = ET.SubElement(resources, "effect")
recurso.set("id", efeito_id)
recurso.set("name", efeito.nome)
recurso.set("uid", efeito.uid)
return efeito_id
def _montar_filtro(self, efeito: EfeitoAjuste, efeito_id: str) -> ET.Element:
"""Monta o ``<filter-video>``/``<filter-audio>`` de um efeito."""
tag = "filter-video" if efeito.tipo == "video" else "filter-audio"
filtro = ET.Element(tag, ref=efeito_id, name=efeito.nome)
for parametro in efeito.parametros:
param = ET.SubElement(filtro, "param")
param.set("name", parametro.nome)
if parametro.chave:
param.set("key", parametro.chave)
param.set("value", parametro.valor)
if parametro.metadado:
param.set("metadata", parametro.metadado)
return filtro
+55
View File
@@ -0,0 +1,55 @@
"""Atalhos de uma linha para as operações mais comuns.
Extraído de writer.py — ver fcpxml/writer/__init__.py para o conjunto.
"""
from typing import Optional
from ..models import (
MarkerType,
)
from .modifier import FCPXMLModifier
# ============================================================================
# CONVENIENCE FUNCTIONS
# ============================================================================
def modify_fcpxml(filepath: str) -> FCPXMLModifier:
"""
Open an FCPXML file for modification.
Usage:
modifier = modify_fcpxml("project.fcpxml")
modifier.add_marker(...)
modifier.save("output.fcpxml")
"""
return FCPXMLModifier(filepath)
def add_marker_to_file(
filepath: str,
timecode: str,
name: str,
marker_type: str = "standard",
output_path: Optional[str] = None
) -> str:
"""Convenience function to add a marker to an FCPXML file."""
modifier = FCPXMLModifier(filepath)
modifier.add_marker_at_timeline(
timecode, name,
MarkerType.from_string(marker_type)
)
return modifier.save(output_path)
def trim_clip_in_file(
filepath: str,
clip_id: str,
trim_start: Optional[str] = None,
trim_end: Optional[str] = None,
output_path: Optional[str] = None
) -> str:
"""Convenience function to trim a clip in an FCPXML file."""
modifier = FCPXMLModifier(filepath)
modifier.trim_clip(clip_id, trim_start, trim_end)
return modifier.save(output_path)
+162
View File
@@ -0,0 +1,162 @@
"""Clipes de áudio e cama musical.
Extraído de writer.py — ver fcpxml/writer/__init__.py para o conjunto.
"""
import xml.etree.ElementTree as ET
from pathlib import Path
from typing import Optional
from ..models import (
TimeValue,
)
from .helpers import _create_asset_element, _dtd_insert, _probe_audio_info, _sanitize_xml_value
class AudioMixin:
"""Clipes de áudio e cama musical."""
# AUDIO CLIP OPERATIONS (v0.6.0)
# ========================================================================
def add_audio_clip(
self,
parent_clip_id: str,
asset_id: Optional[str] = None,
offset: str = "0s",
duration: Optional[str] = None,
role: str = "dialogue",
lane: int = -1,
src: Optional[str] = None,
) -> ET.Element:
"""Add an audio clip connected to an existing timeline clip.
Creates an <asset-clip> at a negative lane with audioRole attribute.
Supports hierarchical roles like "dialogue.boom", "music.score",
"effects.foley".
Args:
parent_clip_id: Name/ID of the clip to attach audio to.
asset_id: Existing asset reference ID. If None and src provided,
creates a new asset.
offset: Position relative to parent clip start.
duration: Duration of audio clip.
role: Audio role (e.g. "dialogue", "music.score", "effects.foley").
lane: Lane number (negative = below primary, default -1).
src: Path to audio file. Used to create a new asset if asset_id
is not provided.
Returns:
The created audio clip element.
"""
parent = self._require_clip(parent_clip_id)
# Resolve or create asset
if asset_id and asset_id in self.resources:
asset = self.resources[asset_id]
elif src:
# Create new asset in resources
resources = self.root.find('.//resources')
if resources is None:
raise ValueError("No <resources> element found in FCPXML")
asset_id = self._unique_resource_id(resources, 'r_audio1')
# The asset duration must reflect the real media length, not the
# requested clip duration — FCP flags assets that claim more
# media than the file contains.
probed = _probe_audio_info(src)
if probed:
rate = probed['sample_rate']
asset_duration = f"{round(probed['duration'] * rate)}/{rate}s"
else:
asset_duration = duration or "0s"
asset_elem = _create_asset_element(
resources, asset_id, Path(src).stem, src,
duration=asset_duration,
has_video="0", has_audio="1",
)
if probed:
asset_elem.set('audioSources', '1')
asset_elem.set('audioChannels', str(probed['channels']))
asset_elem.set('audioRate', str(probed['sample_rate']))
asset = {
'id': asset_id,
'name': Path(src).stem,
'duration': asset_duration,
'element': asset_elem,
}
self.resources[asset_id] = asset
else:
raise ValueError("Must provide either asset_id or src for audio clip")
clip_duration, source_start = self._resolve_clip_duration(asset, duration)
# Clamp so the clip never claims more media than the asset contains
asset_duration_tv = self._parse_time(asset.get('duration', '0s'))
if asset_duration_tv > TimeValue.zero():
available = asset_duration_tv - source_start
if available < TimeValue.zero():
raise ValueError(
f"Source start {source_start.to_fcpxml()} is beyond the end "
f"of audio asset '{asset.get('name')}' "
f"({asset_duration_tv.to_fcpxml()})"
)
if clip_duration > available:
clip_duration = available
new_clip = self._make_asset_clip(
asset_id, asset.get('name', 'Audio'),
self._parse_time(offset), source_start, clip_duration,
lane=str(lane),
audioRole=_sanitize_xml_value(role, 256),
)
_dtd_insert(parent, new_clip)
return new_clip
def add_music_bed(
self,
asset_id: Optional[str] = None,
duration: Optional[str] = None,
role: str = "music",
src: Optional[str] = None,
) -> ET.Element:
"""Add a music bed spanning the full timeline at lane -1.
Convenience method: attaches to the first spine clip and spans
the full timeline duration.
Args:
asset_id: Existing asset reference ID.
duration: Override duration (default: full timeline).
role: Audio role (default "music").
src: Path to audio file (creates asset if asset_id not given).
Returns:
The created music bed clip element.
"""
spine = self._get_spine()
first_clip = None
first_clip_id = None
for clip_id, clip in self.clips.items():
if clip in list(spine):
first_clip = clip
first_clip_id = clip_id
break
if first_clip is None:
raise ValueError("No clips in spine to attach music bed to")
# Calculate full timeline duration if not specified
if not duration:
duration = self._timeline_duration().to_fcpxml()
return self.add_audio_clip(
parent_clip_id=first_clip_id,
asset_id=asset_id,
offset="0s",
duration=duration,
role=role,
lane=-1,
src=src,
)
# ========================================================================
+290
View File
@@ -0,0 +1,290 @@
"""Compound clips: criar e achatar.
Extraído de writer.py — ver fcpxml/writer/__init__.py para o conjunto.
"""
import copy
import uuid
import xml.etree.ElementTree as ET
from typing import List
from ..models import (
TimeValue,
)
from .helpers import (
_dtd_insert,
_sanitize_xml_value,
)
class CompoundMixin:
"""Compound clips: criar e achatar."""
# COMPOUND CLIP OPERATIONS (v0.6.0)
# ========================================================================
def create_compound_clip(
self,
clip_ids: List[str],
name: str = "Compound Clip",
) -> ET.Element:
"""Group spine clips into a compound clip.
Creates a <media> resource with a nested <sequence><spine> containing
the specified clips, then replaces the originals in the main spine
with a single <ref-clip>.
Args:
clip_ids: IDs of clips in the spine to group.
name: Name for the compound clip.
Returns:
The created <ref-clip> element.
"""
spine = self._get_spine()
resources = self.root.find('.//resources')
if resources is None:
raise ValueError("No <resources> element found in FCPXML")
# Collect clips and validate they're in spine
spine_children = list(spine)
clips_to_group = []
for cid in clip_ids:
clip = self._require_clip(cid)
if clip not in spine_children:
raise ValueError(f"Clip not in spine: {cid}")
clips_to_group.append((cid, clip))
if not clips_to_group:
raise ValueError("No valid clips to group")
# Sort by offset so the compound maintains order
clips_to_group.sort(
key=lambda c: self._parse_time(c[1].get('offset', '0s'))
)
# Calculate compound duration and starting offset
first_offset = self._parse_time(clips_to_group[0][1].get('offset', '0s'))
total_duration = TimeValue.zero()
for _, clip in clips_to_group:
total_duration = total_duration + self._parse_time(clip.get('duration', '0s'))
# Get format ref
format_id = None
for fmt_id in self.formats:
format_id = fmt_id
break
# Create media resource with nested sequence
media_id = self._unique_resource_id(resources, 'r_compound1')
media = ET.SubElement(resources, 'media')
media.set('id', media_id)
media.set('name', _sanitize_xml_value(name, 512))
media.set('uid', str(uuid.uuid4()).upper())
seq = ET.SubElement(media, 'sequence')
seq.set('format', format_id or 'r1')
seq.set('duration', total_duration.to_fcpxml())
seq.set('tcStart', '0s')
seq.set('tcFormat', 'NDF')
inner_spine = ET.SubElement(seq, 'spine')
# Move clips into the compound's inner spine
inner_offset = TimeValue.zero()
for _, clip in clips_to_group:
new_clip = copy.deepcopy(clip)
new_clip.set('offset', inner_offset.to_fcpxml())
inner_spine.append(new_clip)
inner_offset = inner_offset + self._parse_time(clip.get('duration', '0s'))
# Get the insert position (where first clip was)
spine_children = list(spine)
insert_idx = spine_children.index(clips_to_group[0][1])
# Remove originals from spine
for cid, clip in clips_to_group:
spine.remove(clip)
if cid in self.clips:
del self.clips[cid]
# Create ref-clip in main spine
ref_clip = ET.Element('ref-clip')
ref_clip.set('ref', media_id)
ref_clip.set('offset', first_offset.to_fcpxml())
ref_clip.set('name', _sanitize_xml_value(name, 512))
ref_clip.set('duration', total_duration.to_fcpxml())
spine.insert(insert_idx, ref_clip)
# Index the new ref-clip
compound_id = f"compound_{name}"
self.clips[compound_id] = ref_clip
return ref_clip
def wrap_titles_in_compound(
self,
parent_clip: ET.Element,
titles: List[ET.Element],
name: str = "Legenda",
) -> ET.Element:
"""Pack lane-nested *titles* of *parent_clip* into one compound clip.
A dynamic-subtitle sub-phrase is a dozen overlapping ``<title>``
elements stacked across as many lanes — legible on screen, unreadable
in the timeline. Collapsing each sub-phrase into a single compound
gives one bar per phrase to drag, mute or retime as a unit.
Mirrors the structure Final Cut itself produces for "New Compound
Clip" over stacked titles: the earliest title becomes the compound's
spine anchor at offset 0, the rest hang off it as lane children, and
a ``<ref-clip>`` takes their place in *parent_clip* on the anchor's
original lane.
Child offsets are rebased from *parent_clip*'s source-time space onto
the anchor's, since a lane child is anchored at its parent's
``start`` — leaving them untouched would shift every word of the
phrase by the gap between the two starts.
Args:
parent_clip: The spine clip the titles currently hang off.
titles: The ``<title>`` elements to pack; must all be direct
children of *parent_clip*.
name: Name for the resulting compound clip.
Returns:
The created ``<ref-clip>`` element, now in *parent_clip*.
"""
if not titles:
raise ValueError("No titles to wrap")
resources = self.root.find('.//resources')
if resources is None:
raise ValueError("No <resources> element found in FCPXML")
ordered = sorted(
titles, key=lambda t: self._parse_time(t.get('offset', '0s'))
)
anchor = ordered[0]
anchor_offset = self._parse_time(anchor.get('offset', '0s'))
anchor_start = self._parse_time(anchor.get('start', '0s'))
anchor_lane = anchor.get('lane')
total = TimeValue.zero()
for title in ordered:
rel = self._parse_time(title.get('offset', '0s')) - anchor_offset
end = rel + self._parse_time(title.get('duration', '0s'))
if end > total:
total = end
format_id = next(iter(self.formats), None) or 'r1'
media_id = self._unique_resource_id(resources, 'r_compound1')
media = ET.SubElement(resources, 'media')
media.set('id', media_id)
media.set('name', _sanitize_xml_value(name, 512))
media.set('uid', str(uuid.uuid4()).upper())
seq = ET.SubElement(media, 'sequence')
seq.set('format', format_id)
seq.set('duration', total.to_fcpxml())
seq.set('tcStart', '0s')
seq.set('tcFormat', 'NDF')
inner_spine = ET.SubElement(seq, 'spine')
for title in ordered:
parent_clip.remove(title)
anchor.set('offset', '0s')
if anchor_lane is not None:
del anchor.attrib['lane']
inner_spine.append(anchor)
for lane, title in enumerate(ordered[1:], start=1):
rel = self._parse_time(title.get('offset', '0s')) - anchor_offset
title.set('offset', (anchor_start + rel).to_fcpxml())
title.set('lane', str(lane))
anchor.append(title)
ref_clip = ET.Element('ref-clip')
ref_clip.set('ref', media_id)
if anchor_lane is not None:
ref_clip.set('lane', anchor_lane)
ref_clip.set('offset', anchor_offset.to_fcpxml())
ref_clip.set('name', _sanitize_xml_value(name, 512))
ref_clip.set('duration', total.to_fcpxml())
_dtd_insert(parent_clip, ref_clip)
return ref_clip
def flatten_compound_clip(
self,
ref_clip_id: str,
) -> List[ET.Element]:
"""Flatten a compound clip back into individual spine clips.
Extracts clips from the compound's inner sequence and places them
back in the main spine at the ref-clip's position.
Args:
ref_clip_id: ID of the ref-clip to flatten.
Returns:
List of extracted clip elements now in the main spine.
"""
spine = self._get_spine()
ref_clip = self._require_clip(ref_clip_id)
if ref_clip.tag != 'ref-clip':
raise ValueError(f"Element is not a ref-clip: {ref_clip_id}")
media_ref = ref_clip.get('ref', '')
ref_offset = self._parse_time(ref_clip.get('offset', '0s'))
# Find the media resource
resources = self.root.find('.//resources')
media_elem = None
if resources is not None:
for m in resources.findall('media'):
if m.get('id') == media_ref:
media_elem = m
break
if media_elem is None:
raise ValueError(f"Media resource not found for ref: {media_ref}")
inner_spine = media_elem.find('.//spine')
if inner_spine is None:
raise ValueError("No spine found in compound clip media")
# Get insert position
spine_children = list(spine)
insert_idx = spine_children.index(ref_clip)
# Remove ref-clip from spine
spine.remove(ref_clip)
if ref_clip_id in self.clips:
del self.clips[ref_clip_id]
# Extract clips from inner spine into main spine
extracted = []
current_offset = ref_offset
for child in list(inner_spine):
new_clip = copy.deepcopy(child)
new_clip.set('offset', current_offset.to_fcpxml())
spine.insert(insert_idx, new_clip)
insert_idx += 1
extracted.append(new_clip)
current_offset = current_offset + self._parse_time(
child.get('duration', '0s')
)
# Index the extracted clip
clip_name = new_clip.get('name') or new_clip.get('id') or f"flat_{len(self.clips)}"
self.clips[clip_name] = new_clip
# Clean up media resource
if resources is not None:
resources.remove(media_elem)
return extracted
# ========================================================================
+49
View File
@@ -0,0 +1,49 @@
"""Clipes conectados (lanes acima/abaixo da spine).
Extraído de writer.py — ver fcpxml/writer/__init__.py para o conjunto.
"""
import xml.etree.ElementTree as ET
from typing import Optional
class ConnectedMixin:
"""Clipes conectados (lanes acima/abaixo da spine)."""
# CONNECTED CLIP OPERATIONS (v0.5.0)
# ========================================================================
def add_connected_clip(
self,
parent_clip_id: str,
asset_id: Optional[str] = None,
asset_name: Optional[str] = None,
offset: str = "0s",
duration: Optional[str] = None,
lane: int = 1,
) -> ET.Element:
"""Add a connected clip (B-roll, title, audio) to an existing timeline clip.
Args:
parent_clip_id: Name/ID of the clip to attach to
asset_id: Asset reference ID
asset_name: Asset name (alternative to asset_id)
offset: Position relative to parent clip start
duration: Duration of connected clip (default: full asset)
lane: Lane number (positive=above, negative=below)
Returns:
The created connected clip element
"""
parent = self._require_clip(parent_clip_id)
asset, asset_id = self._resolve_asset(asset_id, asset_name)
clip_duration, source_start = self._resolve_clip_duration(asset, duration)
new_clip = self._make_asset_clip(
asset_id, asset.get('name', 'Untitled'),
self._parse_time(offset), source_start, clip_duration,
parent=parent, lane=str(lane),
)
return new_clip
# ========================================================================
+725
View File
@@ -0,0 +1,725 @@
"""Núcleo do FCPXMLModifier: carga, índices, navegação na spine e save.
Extraído de writer.py — ver fcpxml/writer/__init__.py para o conjunto.
"""
import xml.etree.ElementTree as ET
from fractions import Fraction
from pathlib import Path
from typing import Any, Dict, Optional, Tuple
from ..models import (
TimeValue,
)
from .document import write_fcpxml
from .helpers import CLIP_TAGS
class ModifierCore:
"""Load an existing FCPXML file, apply edits, and save.
This is the primary editing interface used by every MCP server write-tool
handler. It wraps an ElementTree parsed from disk and maintains three
in-memory indices so that clip/asset lookups are fast.
Index design
------------
``clips`` : ``Dict[str, ET.Element]``
Every ``<clip>``, ``<asset-clip>``, and ``<video>`` element keyed by
its ``id`` attribute, falling back to ``name``, then a generated key.
**Gotcha**: duplicate clip names (e.g. multiple "Interview_A") mean
only the *last* element indexed under that name is accessible. Use
unique ``id`` attributes when possible.
``resources`` : ``Dict[str, Dict[str, Any]]``
Every ``<asset>`` element keyed by ``id``, with pre-extracted ``name``,
``src``, ``start``, ``duration``, and a reference to the raw element.
``formats`` : ``Dict[str, Dict[str, Any]]``
Every ``<format>`` element keyed by ``id``.
Editing model
-------------
1. Look up the target clip via ``_require_clip`` / ``_require_spine_clip``.
2. Mutate the clip's XML attributes (``start``, ``duration``, ``offset``).
3. If the edit changes duration, ripple subsequent spine siblings via
``_ripple_from_index`` so downstream offsets stay contiguous.
4. Call ``save()`` to serialise the modified tree back to disk.
Example::
modifier = FCPXMLModifier("project.fcpxml")
modifier.add_marker("clip_0", "00:00:10:00", "Review", MarkerType.INCOMPLETE)
modifier.trim_clip("clip_1", trim_end="-2s")
modifier.save("project_modified.fcpxml")
Attributes:
path (Path): Filesystem path to the source FCPXML file.
tree (ET.ElementTree): Parsed XML tree (mutated in-place by edits).
root (ET.Element): Root ``<fcpxml>`` element.
fps (float): Detected frame rate from the first ``<format>`` resource.
clips (Dict[str, ET.Element]): Clip index — see *Index design* above.
resources (Dict[str, Dict]): Asset index.
formats (Dict[str, Dict]): Format index.
"""
def __init__(self, fcpxml_path: str):
"""Load *fcpxml_path*, parse its XML, and build lookup indices.
The constructor eagerly builds all three indices (clips, resources,
formats) and detects the project frame rate. After construction the
modifier is ready for any editing operation.
Args:
fcpxml_path: Absolute or relative path to an ``.fcpxml`` file or
an ``.fcpxmld`` bundle (a directory wrapping ``Info.fcpxml``
plus sidecar data files for object tracking / Cinematic mode).
Raises:
FileNotFoundError: If *fcpxml_path* does not exist.
ET.ParseError: If the file is not valid XML.
ValueError: If no ``<spine>`` is found (checked lazily on first edit).
"""
path = Path(fcpxml_path)
self.bundle_dir: Optional[Path] = None
if path.suffix.lower() == '.fcpxmld':
self.bundle_dir = path
inner = path / 'Info.fcpxml'
if not inner.exists():
raise FileNotFoundError(
f"Info.fcpxml not found in bundle: {fcpxml_path}"
)
fcpxml_path = str(inner)
self.path = Path(fcpxml_path)
from ..safe_xml import safe_parse
self.tree = safe_parse(fcpxml_path)
self.root = self.tree.getroot()
self.fps = self._detect_fps()
# Lazily filled on the first generated title; see _unique_text_style_id.
self._text_style_ids: Optional[set] = None
# Lazily filled on the first clip split/cut; see _unique_tracking_shape_id.
self._tracking_shape_ids: Optional[set] = None
self._build_resource_index()
self._build_clip_index()
def _detect_fps(self) -> float:
"""Extract frame rate from format resource."""
for fmt in self.root.findall('.//format'):
frame_dur = fmt.get('frameDuration', '1/30s')
if '/' in frame_dur:
parts = frame_dur.replace('s', '').split('/', 1)
num, denom = int(parts[0]), int(parts[1])
if num <= 0:
return 30.0
return denom / num
return 30.0
def frame_duration_fraction(self):
"""Exact ``frameDuration`` as a Fraction (e.g. 1001/24000 at 23.976fps).
Unlike ``_detect_fps()`` (a float, lossy for NTSC rates), this is
exact — use it wherever a cut boundary is snapped to the frame grid,
so 23.976/29.97/59.94 timebases don't drift off-grid the way a
hardcoded tick base like 2400 does.
"""
for fmt in self.root.findall('.//format'):
raw = fmt.get('frameDuration', '')
if raw.endswith('s') and '/' in raw:
n, d = raw[:-1].split('/', 1)
fd = Fraction(int(n), int(d))
if fd > 0:
return fd
return Fraction(1, 30)
def frame_size(self) -> 'Tuple[float, float]':
"""The sequence's frame size in pixels, as ``(width, height)``.
Reads the sequence's own ``<format>`` when it references one, since a
document may carry several (an asset's source format need not match
the timeline's). Falls back to the first format that declares a size,
then to 1920x1080.
"""
formats = {f.get('id'): f for f in self.root.findall('.//format')}
candidates = []
seq = self.root.find('.//sequence')
if seq is not None and formats.get(seq.get('format')) is not None:
candidates.append(formats[seq.get('format')])
candidates.extend(formats.values())
for fmt in candidates:
try:
width = float(fmt.get('width') or 0)
height = float(fmt.get('height') or 0)
except (TypeError, ValueError):
continue
if width > 0 and height > 0:
return width, height
return 1920.0, 1080.0
def frame_width(self) -> float:
"""The sequence's frame width in pixels."""
return self.frame_size()[0]
def frame_height(self) -> float:
"""The sequence's frame height in pixels."""
return self.frame_size()[1]
def snap_seconds_to_frame(self, seconds: float) -> 'TimeValue':
"""Round *seconds* to the nearest exact frame boundary as a TimeValue."""
fd = self.frame_duration_fraction()
frames = round(seconds / float(fd))
snapped = fd * frames
return TimeValue(snapped.numerator, snapped.denominator)
def snap_spine_times_to_frames(self) -> None:
"""Snap primary-storyline offsets and durations to sequence frames.
Final Cut rejects otherwise valid XML when ripple edits leave a clip
boundary between frames. Use the exact ``frameDuration`` fraction,
rather than a float FPS, to preserve 23.976/29.97 timebases.
"""
frame_duration = None
for fmt in self.root.findall('.//format'):
raw = fmt.get('frameDuration', '')
if raw.endswith('s') and '/' in raw:
n, d = raw[:-1].split('/', 1)
frame_duration = Fraction(int(n), int(d))
break
if frame_duration is None or frame_duration <= 0:
return
for element in self.root.findall('.//spine/*'):
for attr in ('offset', 'duration'):
raw = element.get(attr)
if not raw or not raw.endswith('s'):
continue
value = raw[:-1]
if '/' in value:
n, d = value.split('/', 1)
seconds = Fraction(int(n), int(d))
else:
seconds = Fraction(value)
frames = int(round(float(seconds / frame_duration)))
snapped = frame_duration * frames
element.set(attr, f'{snapped.numerator}/{snapped.denominator}s')
def _build_resource_index(self) -> None:
"""Build ``self.resources`` and ``self.formats`` from ``<asset>``/``<format>`` elements.
Called once during ``__init__``. Each asset entry stores the raw
element plus pre-extracted metadata so callers don't need to
re-parse attributes on every access.
"""
self.resources: Dict[str, Dict[str, Any]] = {}
self.formats: Dict[str, Dict[str, Any]] = {}
for asset in self.root.findall('.//asset'):
asset_id = asset.get('id', '')
self.resources[asset_id] = {
'id': asset_id,
'name': asset.get('name', ''),
'src': asset.get('src', '') or (asset.find('media-rep').get('src', '') if asset.find('media-rep') is not None else ''),
'start': asset.get('start', '0s'),
'duration': asset.get('duration', '0s'),
'element': asset
}
for fmt in self.root.findall('.//format'):
fmt_id = fmt.get('id', '')
self.formats[fmt_id] = {
'id': fmt_id,
'name': fmt.get('name', ''),
'element': fmt
}
def _index_elements(self, tag: str, fallback_prefix: str) -> None:
"""Index XML elements of *tag* into ``self.clips`` by id/name.
Each element is keyed by its ``id`` attribute, falling back to
``name``, then a generated ``{fallback_prefix}_{i}`` key. This
replaces three near-identical loops that only differed in the tag
name and fallback prefix.
"""
for i, elem in enumerate(self.root.findall(f'.//{tag}')):
key = elem.get('id') or elem.get('name') or f"{fallback_prefix}_{i}"
self.clips[key] = elem
def _build_clip_index(self) -> None:
"""Build ``self.clips`` index from all clip-type elements.
Indexes ``<clip>``, ``<asset-clip>``, and ``<video>`` tags. Keys are
resolved by ``_index_elements`` (``id`` → ``name`` → generated).
.. warning::
Duplicate names cause last-one-wins overwrites. If your project
has multiple clips named "Interview_A", only the last one parsed
will be reachable by name. Prefer unique ``id`` attributes.
"""
self.clips: Dict[str, ET.Element] = {}
for tag, prefix in (('clip', 'clip'), ('asset-clip', 'asset_clip'), ('video', 'video')):
self._index_elements(tag, prefix)
def _get_spine(self) -> ET.Element:
"""Get the primary storyline spine.
Finds the spine inside the project/sequence hierarchy, NOT inside
compound clip media resources.
"""
# Prefer the main timeline spine (under project/sequence)
spine = self.root.find('.//project/sequence/spine')
if spine is None:
# Fall back to any spine (for simple FCPXML without project wrapper)
spine = self.root.find('.//spine')
if spine is None:
raise ValueError("No spine found in FCPXML")
return spine
def _iter_spine_clips(self) -> list[tuple[int, ET.Element]]:
"""Return an indexed list of clip-type elements in the primary spine.
Filters out gaps, transitions, and other non-clip elements, returning
only ``(index_in_spine, element)`` pairs where the tag is in
``CLIP_TAGS``. The index is the element's position among *all* spine
children (not just clips), so it stays valid for insertion/removal.
"""
spine = self._get_spine()
return [
(i, child)
for i, child in enumerate(spine.findall('*'))
if child.tag in CLIP_TAGS
]
def _find_spine_clip_at_seconds(self, target_seconds: float) -> tuple[ET.Element, float]:
"""Find the spine clip containing *target_seconds* and return it with the relative offset.
Returns:
``(clip_element, relative_seconds)`` — the clip and the time
within that clip corresponding to *target_seconds*.
Raises:
ValueError: If no clip spans the requested position.
"""
spine = self._get_spine()
for child in spine.findall('*'):
if child.tag not in CLIP_TAGS:
continue
offset = self._parse_time(child.get('offset', '0s')).to_seconds()
dur = self._parse_time(child.get('duration', '0s')).to_seconds()
if offset <= target_seconds < offset + dur:
return child, target_seconds - offset
raise ValueError(f"No spine clip at position {target_seconds:.3f}s")
def _parse_time(self, tc: str) -> TimeValue:
"""Parse a timecode string to TimeValue."""
return TimeValue.from_timecode(tc, self.fps)
def _get_clip_times(
self, clip: ET.Element
) -> tuple:
"""Return (start, duration, offset) TimeValues for a clip element."""
return (
self._parse_time(clip.get('start', '0s')),
self._parse_time(clip.get('duration', '0s')),
self._parse_time(clip.get('offset', '0s')),
)
def source_file_start(self, clip: ET.Element) -> 'TimeValue':
"""Return a clip's in-point measured from the head of its media file.
FCPXML ``start`` on an asset-clip is a source *timecode*, and the
asset's own ``start`` is the timecode of the source media's first
frame. Media analysis (ffmpeg silencedetect, Whisper) reports
file-relative time, so subtract the asset's start timecode to land
both on the same origin. When the asset starts at 0s (the common
case, and every test fixture) this is a no-op.
"""
ref = clip.get('ref', '')
asset = self.resources.get(ref, {})
asset_start = self._parse_time(asset.get('start', '0s'))
clip_start = self._parse_time(clip.get('start', '0s'))
return clip_start - asset_start
def _resolve_clip_duration(
self,
asset: dict,
duration: Optional[str] = None,
in_point: Optional[str] = None,
out_point: Optional[str] = None,
) -> tuple['TimeValue', 'TimeValue']:
"""Compute clip duration and source start from optional overrides.
Centralises the three-way fallback logic shared by insert_clip,
add_connected_clip, and add_audio_clip:
1. If *in_point* and *out_point* are given → subclip range.
2. Else if *duration* is given → explicit duration, source start = 0.
3. Else → full asset duration, source start = 0.
Returns:
``(clip_duration, source_start)`` TimeValue pair.
"""
if in_point and out_point:
in_time = self._parse_time(in_point)
out_time = self._parse_time(out_point)
return out_time - in_time, in_time
if duration:
return self._parse_time(duration), TimeValue.zero()
return self._parse_time(asset.get('duration', '0s')), TimeValue.zero()
def _make_asset_clip(
self,
asset_id: str,
name: str,
offset: 'TimeValue',
start: 'TimeValue',
duration: 'TimeValue',
*,
parent: Optional[ET.Element] = None,
**extra_attrs: str,
) -> ET.Element:
"""Build an ``<asset-clip>`` element with standard attributes.
Centralises the repeated element creation shared by insert_clip,
add_connected_clip, and add_audio_clip. Each caller can pass
additional attributes (``lane``, ``audioRole``, ``format``) via
*extra_attrs*.
Args:
asset_id: Resource reference (e.g. ``'r3'``).
name: Human-readable clip name.
offset: Timeline offset (or offset within parent for connected clips).
start: Source media start point.
duration: Clip duration.
parent: If given, create the element as a SubElement of *parent*;
otherwise create a detached Element.
**extra_attrs: Additional XML attributes (``lane``, ``audioRole``).
Returns:
The new ``<asset-clip>`` Element.
"""
if parent is not None:
elem = ET.SubElement(parent, 'asset-clip')
else:
elem = ET.Element('asset-clip')
elem.set('ref', asset_id)
elem.set('offset', offset.to_fcpxml())
elem.set('name', name)
elem.set('start', start.to_fcpxml())
elem.set('duration', duration.to_fcpxml())
for attr, val in extra_attrs.items():
elem.set(attr, val)
return elem
def _require_clip(self, clip_id: 'str | ET.Element') -> ET.Element:
"""Look up a clip by ID/name, raising if not found.
Centralises the get-or-raise pattern used by every clip-mutating
method so the error message stays consistent and future
enhancements (fuzzy matching, suggestions) only need one site.
An Element is returned as-is. That matters after ``split_clip`` or
``cut_clip_ranges``: the resulting pieces all carry the *same* name,
so a name lookup would always resolve to the first one and silently
put the edit on the wrong piece. Callers holding the exact element
pass it directly.
"""
if isinstance(clip_id, ET.Element):
return clip_id
clip = self.clips.get(clip_id)
if clip is None:
raise ValueError(f"Clip not found: {clip_id}")
return clip
def _require_spine_clip(self, clip_id: str) -> tuple[ET.Element, ET.Element, int]:
"""Look up a clip and verify it lives in the primary spine.
Returns:
``(spine, clip, index_in_spine)`` tuple.
Raises:
ValueError: If the clip doesn't exist or isn't in the spine.
"""
clip = self._require_clip(clip_id)
spine = self._get_spine()
clip_index = self._find_clip_index(spine, clip)
if clip_index is None:
raise ValueError(f"Clip not in spine: {clip_id}")
return spine, clip, clip_index
def _find_clip_index(self, spine: ET.Element, clip: ET.Element) -> int | None:
"""Find the index of a clip in the spine. Returns None if not found."""
for i, child in enumerate(spine):
if child == clip:
return i
return None
@staticmethod
def _find_neighbor_clip(
spine_list: list, index: int, direction: str
) -> Optional[ET.Element]:
"""Find the nearest non-gap clip before or after *index* in *spine_list*.
Args:
spine_list: Materialised list of spine children.
index: Position to search from (exclusive).
direction: ``'prev'`` to search backward, ``'next'`` to search forward.
Returns:
The first clip-type element found, or ``None``.
"""
if direction == 'prev':
for j in range(index - 1, -1, -1):
if spine_list[j].tag in CLIP_TAGS:
return spine_list[j]
else:
for j in range(index + 1, len(spine_list)):
if spine_list[j].tag in CLIP_TAGS:
return spine_list[j]
return None
def _resolve_asset(
self, asset_id: Optional[str], asset_name: Optional[str]
) -> tuple:
"""Look up an asset by ID or name from ``self.resources``.
Returns:
``(asset_dict, resolved_asset_id)`` tuple.
Raises:
ValueError: If neither ID nor name matches a known asset.
"""
if asset_id and asset_id in self.resources:
return self.resources[asset_id], asset_id
if asset_name:
for res_id, res_data in self.resources.items():
if res_data.get('name') == asset_name:
return res_data, res_id
raise ValueError(f"Asset not found: {asset_id or asset_name}")
@staticmethod
def _unique_resource_id(resources: ET.Element, prefix: str) -> str:
"""Generate a unique resource ID with the given *prefix*.
Starts with ``prefix`` (e.g. ``'r_audio1'``), appending an
incrementing counter until no collision exists in *resources*.
"""
existing_ids = {el.get('id', '') for el in resources}
candidate = prefix
counter = 2
while candidate in existing_ids:
# Strip trailing digits from prefix for the counter suffix
base = prefix.rstrip('0123456789')
candidate = f'{base}{counter}'
counter += 1
return candidate
def _find_spine_element_at_timecode(
self, spine: ET.Element, target_tc: str, *, require_clip: bool = False
) -> Optional[ET.Element]:
"""Find the first spine child whose offset matches *target_tc*.
Normalises both sides through ``TimeValue`` round-trip so format
differences (e.g. ``"3600/2400s"`` vs ``"1800/1200s"``) don't
cause false negatives.
Args:
spine: The ``<spine>`` element to search.
target_tc: Timecode string to match against each child's offset.
require_clip: If True, skip non-clip elements (gaps, etc.).
"""
for child in spine:
offset_str = child.get('offset', '0s')
tc = TimeValue.from_timecode(offset_str, self.fps).to_timecode(self.fps)
if tc == target_tc:
if require_clip and child.tag not in CLIP_TAGS:
continue
return child
return None
def _absorb_into_neighbor(
self,
spine: ET.Element,
element: ET.Element,
direction: str,
) -> Optional[ET.Element]:
"""Extend a neighbor clip to absorb *element*'s duration, then remove *element*.
Shared by ``fix_flash_frames`` (absorbing flash-frame clips) and
``fill_gaps`` (absorbing gap elements). Both operations find the
nearest clip in *direction*, grow it by the absorbed element's
duration, and remove the absorbed element from the spine.
When extending backward (``direction='next'``), the neighbor's
source in-point is also pulled earlier so the extra frames come
from before the original cut, not after.
Does **not** call ``_recalculate_offsets`` — callers decide when to
recalculate (per-iteration vs. once at the end).
Args:
spine: The primary storyline ``<spine>`` element.
element: The clip or gap to absorb (will be removed).
direction: ``'prev'`` to extend the previous clip forward,
``'next'`` to extend the next clip backward.
Returns:
The neighbor clip that absorbed the duration, or ``None`` if
no suitable neighbor exists.
"""
spine_list = list(spine)
element_index = spine_list.index(element)
neighbor = self._find_neighbor_clip(spine_list, element_index, direction)
if neighbor is None:
return None
absorbed_dur = self._parse_time(element.get('duration', '0s'))
neighbor_dur = self._parse_time(neighbor.get('duration', '0s'))
if direction == 'next':
neighbor_start = self._parse_time(neighbor.get('start', '0s'))
new_start = neighbor_start - absorbed_dur
if new_start >= TimeValue.zero():
neighbor.set('start', new_start.to_fcpxml())
neighbor.set('duration', (neighbor_dur + absorbed_dur).to_fcpxml())
else:
# Can't shift start negative — only extend by what's available
available = neighbor_start
neighbor.set('start', TimeValue(0, 1).to_fcpxml())
neighbor.set('duration', (neighbor_dur + available).to_fcpxml())
else:
neighbor.set('duration', (neighbor_dur + absorbed_dur).to_fcpxml())
spine.remove(element)
return neighbor
def _resolve_insert_position(
self, position: str, spine_children: list
) -> tuple:
"""Translate a human-friendly position spec into (target_offset, insert_index).
Supported formats:
``'start'`` — beginning of spine
``'end'`` — after last element
``'after:clip_id'`` — after the named clip
``'before:clip_id'``— before the named clip
*timecode* — absolute timeline position
Returns:
``(TimeValue, int)`` — the offset and child-index for spine insertion.
"""
if position == 'start':
return TimeValue.zero(), 0
if position == 'end':
if spine_children:
last = spine_children[-1]
last_offset = self._parse_time(last.get('offset', '0s'))
last_dur = self._parse_time(last.get('duration', '0s'))
return last_offset + last_dur, len(spine_children)
return TimeValue.zero(), len(spine_children)
if position.startswith('after:') or position.startswith('before:'):
is_after = position.startswith('after:')
ref_id = position.split(':', 1)[1]
ref_clip = self.clips.get(ref_id)
if ref_clip is None or ref_clip not in spine_children:
raise ValueError(f"Reference clip not found: {ref_id}")
idx = spine_children.index(ref_clip)
ref_offset = self._parse_time(ref_clip.get('offset', '0s'))
if is_after:
ref_dur = self._parse_time(ref_clip.get('duration', '0s'))
return ref_offset + ref_dur, idx + 1
return ref_offset, idx
# Assume timecode
target_offset = self._parse_time(position)
insert_index = 0
for i, child in enumerate(spine_children):
child_offset = self._parse_time(child.get('offset', '0s'))
if child_offset >= target_offset:
insert_index = i
break
insert_index = i + 1
return target_offset, insert_index
def _make_transition_element(
self,
effect_name: str,
trans_offset: 'TimeValue',
trans_duration: 'TimeValue',
effect_ref_id: str | None,
) -> ET.Element:
"""Build a <transition> element with optional filter-video child."""
transition = ET.Element('transition')
transition.set('name', effect_name)
transition.set('offset', trans_offset.to_fcpxml())
transition.set('duration', trans_duration.to_fcpxml())
if effect_ref_id:
fv = ET.SubElement(transition, 'filter-video')
fv.set('ref', effect_ref_id)
fv.set('name', effect_name)
return transition
def save(self, output_path: Optional[str] = None) -> str:
"""Serialise the modified XML tree to disk.
When the destination ends in ``.fcpxmld`` a bundle directory is
created and the XML lands in ``Info.fcpxml`` inside it. If the
source was also a bundle, every sidecar file (object-tracking /
Cinematic-mode ``dataLocator`` payloads — anything that isn't
``Info.fcpxml``) is copied across so the round-trip is lossless.
Writing a bundle source to a flat ``.fcpxml`` destination drops
those sidecars by definition.
Args:
output_path: Destination ``.fcpxml`` file or ``.fcpxmld``
bundle path. Defaults to overwriting the original
file/bundle loaded in ``__init__``.
Returns:
The absolute path written to (the bundle path when writing
a bundle, not the inner ``Info.fcpxml``).
"""
if output_path is None:
out = self.bundle_dir if self.bundle_dir is not None else self.path
else:
out = Path(output_path)
# Every write path goes through here, so snapping here (rather than
# in each handler) guarantees ripple edits never leave a spine clip
# off the frame grid — see snap_spine_times_to_frames() docstring.
# No-op (each value already equals its own snapped form) on content
# that was already frame-aligned.
self.snap_spine_times_to_frames()
if out.suffix.lower() == '.fcpxmld':
out.mkdir(exist_ok=True)
if (
self.bundle_dir is not None
and self.bundle_dir.resolve() != out.resolve()
):
self._copy_bundle_sidecars(self.bundle_dir, out)
write_fcpxml(self.root, str(out / 'Info.fcpxml'), fps=self.fps)
return str(out)
return write_fcpxml(self.root, str(out), fps=self.fps)
@staticmethod
def _copy_bundle_sidecars(src_bundle: Path, dst_bundle: Path) -> None:
"""Copy every sidecar entry of *src_bundle* into *dst_bundle*.
Sidecars are all bundle members except ``Info.fcpxml`` itself —
e.g. the external data files that ``locator``/``dataLocator``
elements reference for object tracking and Cinematic mode.
"""
import shutil
for entry in src_bundle.iterdir():
if entry.name == 'Info.fcpxml':
continue
target = dst_bundle / entry.name
if entry.is_dir():
shutil.copytree(entry, target, dirs_exist_ok=True)
else:
shutil.copy2(entry, target)
+367
View File
@@ -0,0 +1,367 @@
"""Dividir, cortar faixas e apagar clipes.
Extraído de writer.py — ver fcpxml/writer/__init__.py para o conjunto.
"""
import copy
import xml.etree.ElementTree as ET
from typing import List, Tuple
from ..models import (
TimeValue,
)
class CutMixin:
"""Dividir, cortar faixas e apagar clipes."""
# SPLIT & DELETE OPERATIONS
# ========================================================================
@staticmethod
def _filter_children_for_segment(
clip: ET.Element,
seg_start: 'TimeValue',
seg_duration: 'TimeValue',
) -> None:
"""Remove markers/keywords/titles from *clip* that fall outside the segment range.
After ``split_clip`` deepcopy's the original clip into each segment, every
segment inherits all child elements. Markers whose ``start`` falls outside
``[seg_start, seg_start + seg_duration)`` are phantom duplicates and must be
removed. Keywords that partially overlap get their ``start``/``duration``
clamped to the segment boundaries.
A lane-nested ``<title>`` (a "text" voice action's on-screen callout,
or a caption from an earlier `generate_dynamic_subtitles` pass) is
the same kind of phantom duplicate, just keyed on ``offset`` instead
of ``start`` — its offset lives in the same source-media coordinate
space as a marker's ``start`` (see ``add_text_title``/``add_marker``,
both anchored at ``parent.start``). Left unfiltered, every further
cut (silence removal, filler removal) duplicates it into every
resulting piece, so the same word shows up several times across the
edited timeline instead of once where it was placed.
A lane-nested ``<video>`` zoom (the "Clipe de Ajuste" adjustment
layer ``add_zoom`` creates, ``role`` starting with ``"adjustments."``)
is the exact same phantom-duplicate case, keyed on ``offset``+
``duration`` like a keyword. Left unfiltered, every further cut
duplicates the zoom into every resulting piece with its original
offset untouched — each copy then draws at the same absolute
position, so two "Clipe de Ajuste" bars appear stacked on top of
each other in the timeline instead of the one real zoom window.
"""
seg_end = seg_start + seg_duration
to_remove = []
for child in clip:
tag = child.tag
if tag in ('marker', 'chapter-marker'):
child_start = TimeValue.from_timecode(child.get('start', '0s'))
if child_start < seg_start or child_start >= seg_end:
to_remove.append(child)
elif tag == 'title':
title_offset = TimeValue.from_timecode(child.get('offset', '0s'))
if title_offset < seg_start or title_offset >= seg_end:
to_remove.append(child)
elif tag == 'video' and (child.get('role') or '').startswith('adjustments.'):
v_offset = TimeValue.from_timecode(child.get('offset', '0s'))
v_dur = TimeValue.from_timecode(child.get('duration', '0s'))
v_end = v_offset + v_dur
if v_end <= seg_start or v_offset >= seg_end:
to_remove.append(child)
elif tag == 'keyword':
kw_start = TimeValue.from_timecode(child.get('start', '0s'))
kw_dur = TimeValue.from_timecode(child.get('duration', '0s'))
kw_end = kw_start + kw_dur
# Completely outside segment → remove
if kw_end <= seg_start or kw_start >= seg_end:
to_remove.append(child)
else:
# Clamp keyword range to segment boundaries
clamped_start = max(kw_start, seg_start)
clamped_end = min(kw_end, seg_end)
child.set('start', clamped_start.to_fcpxml())
child.set('duration', (clamped_end - clamped_start).to_fcpxml())
for child in to_remove:
clip.remove(child)
def split_clip(
self,
clip_id: str,
split_points: List[str]
) -> List[ET.Element]:
"""
Split a clip at specified timecodes.
Args:
clip_id: Clip to split
split_points: Timecodes within the clip to split at
Returns:
List of resulting clip elements
"""
spine, clip, clip_index = self._require_spine_clip(clip_id)
# Get clip properties
clip_start, clip_duration, clip_offset = self._get_clip_times(clip)
clip_name = clip.get('name', 'Clip')
# Sort split points
split_times = sorted([self._parse_time(sp) for sp in split_points])
# Remove original clip
spine.remove(clip)
# Create new clips
new_clips = []
current_offset = clip_offset
current_start = clip_start
all_points = split_times + [clip_duration]
for i, split_time in enumerate(all_points):
if i == 0:
segment_duration = split_time
else:
segment_duration = split_time - split_times[i - 1]
if segment_duration <= TimeValue.zero():
continue
# Create new clip
new_clip = copy.deepcopy(clip)
new_clip.set('name', clip_name)
new_clip.set('offset', current_offset.to_fcpxml())
new_clip.set('start', current_start.to_fcpxml())
new_clip.set('duration', segment_duration.to_fcpxml())
# Remove markers/keywords that belong to other segments
self._filter_children_for_segment(
new_clip, current_start, segment_duration
)
self._reassign_text_style_ids(new_clip)
self._reassign_tracking_shape_ids(new_clip)
spine.insert(clip_index + len(new_clips), new_clip)
new_clips.append(new_clip)
# Update for next iteration
current_offset = current_offset + segment_duration
current_start = current_start + segment_duration
# Update clip index: remove stale original entry, add split entries
self.clips.pop(clip_id, None)
for i, new_clip in enumerate(new_clips):
new_id = f"{clip_id}_split_{i}"
self.clips[new_id] = new_clip
return new_clips
def cut_clip_ranges(
self,
clip: ET.Element,
cut_ranges: List[Tuple['TimeValue', 'TimeValue']],
) -> 'TimeValue':
"""Remove clip-relative time ranges from a spine clip, rippling after.
Element-based on purpose: callers that walk the spine (e.g. media
silence removal) pass the exact element, so duplicate-named clips are
never ambiguous the way name-keyed operations are.
Args:
clip: The spine clip element to cut (must be a direct spine child).
cut_ranges: (start, end) TimeValue pairs measured from the clip's
own head. Overlapping/unsorted ranges are merged; portions
outside [0, clip duration] are clamped. A cut covering the
whole clip removes it entirely.
Returns:
Total removed duration (zero if no effective ranges).
"""
spine = self._get_spine()
clip_start, clip_duration, clip_offset = self._get_clip_times(clip)
clip_index = list(spine).index(clip)
zero = TimeValue.zero()
# Clamp, sort, merge.
clamped = []
for start, end in cut_ranges:
start = start if start > zero else zero
end = end if end < clip_duration else clip_duration
if end > start:
clamped.append((start, end))
clamped.sort(key=lambda r: r[0])
merged: List[Tuple[TimeValue, TimeValue]] = []
for start, end in clamped:
if merged and start <= merged[-1][1]:
if end > merged[-1][1]:
merged[-1] = (merged[-1][0], end)
else:
merged.append((start, end))
if not merged:
return zero
# Keep ranges = complement of the merged cuts.
keeps: List[Tuple[TimeValue, TimeValue]] = []
cursor = zero
for start, end in merged:
if start > cursor:
keeps.append((cursor, start))
cursor = end
if cursor < clip_duration:
keeps.append((cursor, clip_duration))
# A keep segment shorter than MIN_KEEP_SECONDS is leftover between
# two cuts, not a real clip — at the very start/end of the clip it's
# cut padding with no kept audio on the outer side; in the interior
# it's the pause BETWEEN two things that were both cut (e.g. two
# consecutive deactivated phrases in the voice-editing flow), which
# belongs to neither side by construction. At the edges we fold it
# into the one neighboring KEEP segment there is, which simply starts
# earlier / ends later to absorb it. In the interior both neighbors
# are CUT, not keep, so there is nothing to fold into — it is just
# dropped, extending the surrounding cut across it instead of
# surviving as a third near-invisible micro-clip.
#
# The threshold is bigger than one frame on purpose: measured on a
# real voice-edit (0.07-0.23s residues), a single frame did not catch
# them — this is pause/padding leftover, not intentional short
# content, so treating anything under a third of a second this way
# is safe for this cut path.
min_keep_seconds = max(6 * float(self.frame_duration_fraction()), 0.3)
i = 0
while len(keeps) > 1 and i < len(keeps):
start, end = keeps[i]
if (end - start).to_seconds() >= min_keep_seconds:
i += 1
continue
if i == 0:
keeps[1] = (start, keeps[1][1])
keeps.pop(0)
elif i == len(keeps) - 1:
keeps[i - 1] = (keeps[i - 1][0], end)
keeps.pop(i)
else:
keeps.pop(i)
# Re-check the same index: the segment now there might itself be
# short enough to absorb again (two short keeps in a row).
spine.remove(clip)
new_clips: List[ET.Element] = []
current_offset = clip_offset
kept_total = zero
for keep_start, keep_end in keeps:
seg_duration = keep_end - keep_start
seg_start = clip_start + keep_start
new_clip = copy.deepcopy(clip)
new_clip.set('offset', current_offset.to_fcpxml())
new_clip.set('start', seg_start.to_fcpxml())
new_clip.set('duration', seg_duration.to_fcpxml())
self._filter_children_for_segment(new_clip, seg_start, seg_duration)
self._reassign_text_style_ids(new_clip)
self._reassign_tracking_shape_ids(new_clip)
spine.insert(clip_index + len(new_clips), new_clip)
new_clips.append(new_clip)
current_offset = current_offset + seg_duration
kept_total = kept_total + seg_duration
removed = clip_duration - kept_total
self._ripple_from_index(spine, clip_index + len(new_clips), zero - removed)
self._update_sequence_duration()
# Keep the name index coherent, mirroring delete_clip/split_clip.
name = clip.get('id') or clip.get('name') or ''
if name and self.clips.get(name) is clip:
if new_clips:
self.clips[name] = new_clips[0]
else:
remaining = [
sc for _, sc in self._iter_spine_clips()
if (sc.get('id') or sc.get('name') or '') == name
]
if remaining:
self.clips[name] = remaining[0]
else:
self.clips.pop(name, None)
return removed
def remove_trailing_gaps(self) -> None:
"""Remove empty ``<gap>`` elements at the end of the timeline.
Silence removal (and FCP round-trips) can leave a trailing gap holding
the timeline open past the last real clip. This removes only *trailing*
gaps — a gap in the middle is left untouched — and re-syncs the sequence
duration so the exported file ends where the content ends.
"""
spine = self._get_spine()
children = list(spine)
if not children:
return
last = children[-1]
if last.tag != 'gap':
return
spine.remove(last)
self._update_sequence_duration()
def delete_clip(
self,
clip_ids: List[str],
ripple: bool = True
) -> None:
"""
Delete clips from timeline.
Uses spine iteration instead of the name-indexed dict so that
duplicate-named clips (e.g. four ``Interview_A``) are resolved
correctly — always targeting the *first* spine match rather than
the last-indexed entry.
Args:
clip_ids: Clips to delete
ripple: If True, shift subsequent clips. If False, leave gaps.
"""
spine = self._get_spine()
for clip_id in clip_ids:
# Walk spine directly to find the first clip matching this name,
# avoiding the last-one-wins problem in self.clips.
target = None
for _spine_idx, spine_clip in self._iter_spine_clips():
name = spine_clip.get('id') or spine_clip.get('name') or ''
if name == clip_id:
target = spine_clip
break
if target is None:
continue
_, clip_duration, clip_offset = self._get_clip_times(target)
clip_index = list(spine).index(target)
if ripple:
spine.remove(target)
self._ripple_from_index(
spine, clip_index, TimeValue.zero() - clip_duration
)
else:
# Replace with gap
gap = ET.Element('gap')
gap.set('name', 'Gap')
gap.set('offset', clip_offset.to_fcpxml())
gap.set('duration', clip_duration.to_fcpxml())
spine.remove(target)
spine.insert(clip_index, gap)
# Re-index: if other spine clips share this name, point the
# dict entry at the next one; otherwise remove entirely.
remaining = [
sc for _, sc in self._iter_spine_clips()
if (sc.get('id') or sc.get('name') or '') == clip_id
]
if remaining:
self.clips[clip_id] = remaining[0]
else:
self.clips.pop(clip_id, None)
# ========================================================================
+170
View File
@@ -0,0 +1,170 @@
"""Escrita do documento FCPXML: assets de vídeo, timebases e serialização.
Extraído de writer.py — ver fcpxml/writer/__init__.py para o conjunto.
"""
import logging
import subprocess
import xml.etree.ElementTree as ET
from pathlib import Path
from typing import Optional
from ..models import (
TimeValue,
)
from .validation import validate_fcpxml
_log = logging.getLogger(__name__)
# ============================================================================
# STILL IMAGE AUTO-CONVERSION (v0.6.0)
# ============================================================================
_STILL_IMAGE_EXTENSIONS = {'.png', '.jpg', '.jpeg', '.tiff', '.tif', '.bmp'}
def _ensure_video_asset(
src_path: str,
duration: float = 10.0,
fps: int = 24,
width: int = 1920,
height: int = 1080,
) -> str:
"""Convert a still image to a video file if needed.
Detects still images by extension and converts them to MOV using ffmpeg.
Video files are returned as-is.
Args:
src_path: Path to the source media file.
duration: Duration in seconds for the still-to-video conversion.
fps: Frame rate for the output video.
width: Output width (even number).
height: Output height (even number).
Returns:
Path to the video file (original path if already video, new .mov path
if converted from still).
Raises:
FileNotFoundError: If ffmpeg is not installed.
"""
# Validate numeric parameters to prevent ffmpeg abuse / resource exhaustion.
if not isinstance(duration, (int, float)) or duration <= 0 or duration > 3600:
raise ValueError(f"duration must be 0 < d <= 3600, got {duration!r}")
if not isinstance(fps, int) or fps < 1 or fps > 240:
raise ValueError(f"fps must be 1–240, got {fps!r}")
if not isinstance(width, int) or width < 2 or width > 7680 or width % 2:
raise ValueError(f"width must be even, 2–7680, got {width!r}")
if not isinstance(height, int) or height < 2 or height > 4320 or height % 2:
raise ValueError(f"height must be even, 2–4320, got {height!r}")
path = Path(src_path)
if path.suffix.lower() not in _STILL_IMAGE_EXTENSIONS:
return src_path
output_path = path.with_suffix('.mov')
if output_path.exists():
return str(output_path)
# Build ffmpeg command: still image → video with specified duration
cmd = [
'ffmpeg', '-y',
'-loop', '1',
'-i', str(path),
'-c:v', 'prores_ks',
'-profile:v', '0',
'-t', str(duration),
'-r', str(fps),
'-vf', f'scale={width}:{height}:force_original_aspect_ratio=decrease,'
f'pad={width}:{height}:(ow-iw)/2:(oh-ih)/2',
'-pix_fmt', 'yuva444p10le',
str(output_path),
]
try:
subprocess.run(cmd, check=True, capture_output=True, timeout=120)
except FileNotFoundError:
raise FileNotFoundError(
"ffmpeg not found. Install ffmpeg to use still image auto-conversion: "
"brew install ffmpeg"
)
except subprocess.TimeoutExpired:
raise RuntimeError(
f"Image conversion timed out after 120s: {path}"
)
except subprocess.CalledProcessError as e:
stderr_msg = e.stderr.decode(errors='replace') if e.stderr else str(e)
raise RuntimeError(f"ffmpeg conversion failed: {stderr_msg}")
return str(output_path)
def _enforce_standard_timebases(root: ET.Element) -> None:
"""Walk all elements and snap time attributes to standard FCPXML timebases.
Targets offset, start, duration, and tcStart attributes. Values that
already use a standard denominator are left untouched.
"""
time_attrs = ('offset', 'start', 'duration', 'tcStart')
for elem in root.iter():
for attr in time_attrs:
val = elem.get(attr)
if val and val.endswith('s') and '/' in val:
try:
tv = TimeValue.from_timecode(val)
if not tv.is_standard_timebase():
# Snap to nearest frame at 2400 ticks/sec
snapped = tv.snap_to_frame(24)
elem.set(attr, snapped.to_fcpxml())
except (ValueError, ZeroDivisionError):
pass # Skip unparseable values
def write_fcpxml(
root: ET.Element,
filepath: str,
enforce_timebases: bool = False,
strict: bool = False,
fps: Optional[float] = None,
) -> str:
"""Format an ElementTree root as pretty-printed FCPXML and write to disk.
Handles XML declaration, DOCTYPE insertion, and blank-line cleanup
consistently across all FCPXML output paths (modifier, writer, rough cut).
Args:
root: The <fcpxml> root Element to serialize.
filepath: Destination file path.
enforce_timebases: If True, snap all time values to standard FCPXML
timebases before writing. Default False for backward compat.
strict: If True, raise ValueError on validation errors.
If False (default), log warnings.
fps: Frame rate for the frame-alignment validation check. Defaults
to 24 when omitted — pass the sequence's real (float) rate so
NTSC projects (23.976/29.97/59.94fps) don't get spurious
"not frame-aligned at 24fps" warnings for values that are
exactly aligned at their own true rate.
Returns:
The filepath written to.
"""
if enforce_timebases:
_enforce_standard_timebases(root)
# Auto-validate before writing
issues = validate_fcpxml(root, fps=fps if fps is not None else 24.0)
if issues:
errors = [i for i in issues if i.severity == "error"]
warnings = [i for i in issues if i.severity == "warning"]
for w in warnings:
_log.warning("FCPXML validation: %s", w.message)
if errors and strict:
msg = "; ".join(e.message for e in errors)
raise ValueError(f"FCPXML validation failed: {msg}")
for e in errors:
_log.error("FCPXML validation: %s", e.message)
from ..safe_xml import serialize_xml
return serialize_xml(root, filepath, doctype='<!DOCTYPE fcpxml>')
+147
View File
@@ -0,0 +1,147 @@
"""FCPXMLWriter: gera um documento novo a partir de objetos Python.
Extraído de writer.py — ver fcpxml/writer/__init__.py para o conjunto.
"""
import uuid
import xml.etree.ElementTree as ET
from datetime import datetime
from ..models import (
Marker,
Project,
Timecode,
)
from .document import write_fcpxml
from .helpers import build_marker_element
# ============================================================================
# FCPXML GENERATOR - Create from Python objects
# ============================================================================
class FCPXMLWriter:
"""Generate a new FCPXML document from Python dataclass objects.
Converts a ``Project`` (containing ``Timeline`` → ``Clip`` → ``Marker``
hierarchies) into a spec-compliant FCPXML v1.11 element tree and writes
it to disk. Used by ``RoughCutGenerator`` and the ``generate_*`` MCP
tools to create fresh timelines from scratch.
Unlike ``FCPXMLModifier`` (which mutates existing XML), this class
*creates* XML from structured Python objects.
Example::
from fcpxml.models import Project, Timeline, Clip, Timecode
project = Project(name="My Edit", timelines=[...])
writer = FCPXMLWriter()
writer.write_project(project, "output.fcpxml")
"""
def __init__(self, version: str = "1.13"):
"""Initialize writer targeting the given FCPXML version."""
self.version = version
self.resource_counter = 1
def _next_resource_id(self) -> str:
"""Return an auto-incrementing resource ID (r1, r2, ...)."""
rid = f"r{self.resource_counter}"
self.resource_counter += 1
return rid
def _generate_uid(self) -> str:
"""Generate a unique identifier for FCPXML elements."""
return str(uuid.uuid4()).upper()
def _tc_to_rational(self, tc: Timecode) -> str:
"""Convert a Timecode to FCPXML rational time string (e.g. '48/24s')."""
return f"{tc.frames}/{int(tc.frame_rate)}s"
def write_project(self, project: Project, filepath: str):
"""Write a project to an FCPXML file."""
root = self._build_fcpxml(project)
write_fcpxml(root, filepath)
def _build_fcpxml(self, project: Project) -> ET.Element:
"""Build the full FCPXML element tree: fcpxml > resources + library > event > project."""
root = ET.Element('fcpxml', version=self.version)
resources = ET.SubElement(root, 'resources')
resource_map = {}
if project.timelines:
timeline = project.timelines[0]
format_id = self._next_resource_id()
ET.SubElement(resources, 'format',
id=format_id,
name=f"FFVideoFormat{timeline.height}p{int(timeline.frame_rate)}",
frameDuration=f"1/{int(timeline.frame_rate)}s",
width=str(timeline.width), height=str(timeline.height))
resource_map['_format'] = format_id
library = ET.SubElement(root, 'library',
location=f"file:///Users/editor/Movies/{project.name}.fcpbundle/")
event = ET.SubElement(library, 'event', name=project.name, uid=self._generate_uid())
for timeline in project.timelines:
self._add_timeline(event, timeline, resources, resource_map)
return root
def _add_timeline(self, event, timeline, resources, resource_map):
"""Add a timeline as a project > sequence > spine structure under the event."""
project_elem = ET.SubElement(event, 'project',
name=timeline.name, uid=self._generate_uid(),
modDate=datetime.now().strftime("%Y-%m-%d %H:%M:%S -0500"))
format_id = resource_map.get('_format', 'r1')
sequence = ET.SubElement(project_elem, 'sequence',
format=format_id, duration=self._tc_to_rational(timeline.duration),
tcStart="0s", tcFormat="NDF", audioLayout="stereo", audioRate="48k")
spine = ET.SubElement(sequence, 'spine')
for clip in timeline.clips:
self._add_clip(spine, clip, resources, resource_map)
for marker in timeline.markers:
self._add_marker(sequence, marker)
def _add_clip(self, spine, clip, resources, resource_map):
"""Add a clip as an asset-clip element, creating its asset resource if needed."""
if clip.media_path and clip.media_path not in resource_map:
asset_id = self._next_resource_id()
ET.SubElement(resources, 'asset', id=asset_id, name=clip.name,
uid=self._generate_uid(), src=clip.media_path, start="0s",
duration=self._tc_to_rational(clip.duration), hasVideo="1", hasAudio="1")
resource_map[clip.media_path] = asset_id
asset_id = resource_map.get(clip.media_path, 'r1')
format_id = resource_map.get('_format', 'r1')
clip_elem = ET.SubElement(spine, 'asset-clip',
ref=asset_id, offset=self._tc_to_rational(clip.start), name=clip.name,
start=self._tc_to_rational(clip.source_start) if clip.source_start else "0s",
duration=self._tc_to_rational(clip.duration), format=format_id, tcFormat="NDF")
for marker in clip.markers:
self._add_marker(clip_elem, marker)
for keyword in clip.keywords:
self._add_keyword(clip_elem, keyword)
def _add_marker(self, parent: ET.Element, marker: Marker):
"""Add a marker or chapter-marker element to a parent clip or sequence."""
build_marker_element(
parent=parent,
marker_type=marker.marker_type,
start=self._tc_to_rational(marker.start),
duration=self._tc_to_rational(marker.duration) if marker.duration else "1/24s",
name=marker.name,
note=marker.note or None,
)
def _add_keyword(self, parent, keyword):
"""Add a keyword element with optional start/duration range to a parent clip."""
attrs = {'value': keyword.value}
if keyword.start:
attrs['start'] = self._tc_to_rational(keyword.start)
if keyword.duration:
attrs['duration'] = self._tc_to_rational(keyword.duration)
ET.SubElement(parent, 'keyword', **attrs)
+279
View File
@@ -0,0 +1,279 @@
"""Ajudantes de nível de módulo do writer: sanitização, escalas, elementos base.
Extraído de writer.py — ver fcpxml/writer/__init__.py para o conjunto.
"""
import subprocess
import xml.etree.ElementTree as ET
from pathlib import Path
from typing import Any, Dict, List, Optional
from ..models import (
MarkerType,
)
# Maximum lengths for XML attribute values to prevent memory abuse
_MAX_MARKER_NAME_LENGTH = 1024
_MAX_NOTE_LENGTH = 4096
# ============================================================================
# EFFECT RESOURCE REGISTRY (v0.6.0)
# ============================================================================
# FCP built-in transition/filter effect UUIDs extracted from Filters.bundle.
# Maps slug → (display_name, uuid).
FCP_EFFECTS: Dict[str, tuple] = {
# Dissolves
'cross-dissolve': ('Cross Dissolve', '4731E73A-8DAC-4113-9A30-AE85B1761265'),
'fade': ('Fade', '8154D0DA-C99B-4EF8-8FF8-006FE5ED57F1'),
'dip-to-color': ('Dip to Color', 'F779C565-486D-4633-8035-0374B4DB8F5C'),
'noise-dissolve': ('Noise Dissolve', 'ABFED81E-35D9-429C-AB47-438C1FB5D9DE'),
# Wipes
'edge-wipe': ('Edge Wipe', '857E2FBA-98DB-411B-A88C-CE6ABC1F65D8'),
'slide': ('Slide', '6AAB0D54-FCD8-4EBD-A62D-D352A5ED1648'),
'band-wipe': ('Band Wipe', 'A4E0B8E4-E916-474B-A14C-E3A9E0B1A3C1'),
'center-wipe': ('Center Wipe', 'B3F2D4A1-7C8E-4B9D-A5F6-D1E2C3B4A5D6'),
'checker-wipe': ('Checker Wipe', 'C4D3E2F1-8A7B-4C6D-B5E4-F2A1D3C4B5E6'),
'clock-wipe': ('Clock Wipe', 'D5E4F3A2-9B8C-4D7E-C6F5-A3B2E4D5C6F7'),
'gradient-wipe': ('Gradient Wipe', 'E6F5A4B3-AC9D-4E8F-D7A6-B4C3F5E6D7A8'),
'inset-wipe': ('Inset Wipe', 'F7A6B5C4-BD0E-4F9A-E8B7-C5D4A6F7E8B9'),
'star-wipe': ('Star Wipe', 'A8B7C6D5-CE1F-4A0B-F9C8-D6E5B7A8F9C0'),
# Legacy aliases — map common shorthand to canonical slugs
'fade-to-black': ('Fade', '8154D0DA-C99B-4EF8-8FF8-006FE5ED57F1'),
'fade-from-black': ('Fade', '8154D0DA-C99B-4EF8-8FF8-006FE5ED57F1'),
'wipe': ('Edge Wipe', '857E2FBA-98DB-411B-A88C-CE6ABC1F65D8'),
'dissolve': ('Cross Dissolve', '4731E73A-8DAC-4113-9A30-AE85B1761265'),
}
def list_effects() -> List[Dict[str, str]]:
"""Return a list of all available FCP transition effects.
Each entry contains slug, display_name, and uuid.
Legacy aliases are excluded to avoid duplicates.
"""
seen_uuids: set = set()
effects = []
for slug, (name, uid) in FCP_EFFECTS.items():
if uid in seen_uuids:
continue
seen_uuids.add(uid)
effects.append({'slug': slug, 'name': name, 'uuid': uid})
return effects
# Named constants for clip-tag sets used across operations.
# Using named tuples prevents inconsistent ad-hoc tag lists and ensures
# new clip types only need adding in one place.
CLIP_TAGS = ('clip', 'asset-clip', 'video', 'ref-clip')
CLIP_AND_AUDIO_TAGS = ('clip', 'asset-clip', 'video', 'audio', 'ref-clip')
SPINE_ELEMENT_TAGS = ('clip', 'asset-clip', 'video', 'audio', 'gap', 'transition', 'ref-clip')
def _sanitize_xml_value(value: str, max_length: int = _MAX_MARKER_NAME_LENGTH) -> str:
"""Sanitize a string value before writing it into an XML attribute.
Strips null bytes, control characters (except tab/newline/CR), and
enforces a length limit to prevent memory abuse or malformed XML.
"""
if not isinstance(value, str):
return str(value)
# Remove null bytes and non-printable control characters
cleaned = ''.join(
c for c in value
if c in ('\t', '\n', '\r') or ord(c) >= 32
)
if len(cleaned) > max_length:
cleaned = cleaned[:max_length]
return cleaned
# FCPXML DTD child element ordering for asset-clip / clip elements.
# Elements MUST appear in this order for DTD validation.
# See: https://developer.apple.com/documentation/professional-video-applications/fcpxml-reference
_ASSET_CLIP_CHILD_ORDER = [
'note',
'conform-rate', 'timeMap',
'adjust-crop', 'adjust-corners', 'adjust-conform', 'adjust-transform',
'adjust-blend', 'adjust-stabilization', 'adjust-rollingShutter',
'adjust-360-transform', 'adjust-reorient', 'adjust-orientation',
'adjust-volume', 'adjust-panner',
# anchor items (connected clips, titles, etc.)
'audio', 'video', 'clip', 'title', 'caption',
'mc-clip', 'ref-clip', 'sync-clip', 'asset-clip', 'audition', 'spine',
# marker items
'marker', 'chapter-marker', 'rating', 'keyword', 'analysis-marker',
# trailing
'audio-channel-source',
'filter-video', 'filter-video-mask',
'filter-audio',
'metadata',
]
# Build a priority lookup: tag → index for fast comparison
_CHILD_ORDER_INDEX = {tag: i for i, tag in enumerate(_ASSET_CLIP_CHILD_ORDER)}
# How close to the end of a clip a zoom must finish for the return to be
# skipped. Within this margin the cut arrives before the eye registers the
# move back, so the return reads as a twitch rather than a resolution.
HOLD_AT_CUT_THRESHOLD = 1.0
# How close to the start of a clip a zoom must begin for the ramp-in to be
# skipped and the shot to simply open already zoomed. Tighter than the end
# margin on purpose: at the end the cut hides an unfinished return, but at
# the start a ramp is visible from frame one and reads as the shot settling.
START_AT_CUT_THRESHOLD = 0.5
def _fmt_scale(value: float) -> str:
"""Format a scale factor without trailing float noise (1.0 -> "1")."""
return f"{value:.6f}".rstrip("0").rstrip(".") or "0"
def _dtd_insert(parent: ET.Element, child: ET.Element) -> ET.Element:
"""Insert a child element into parent at the correct DTD-ordered position.
Instead of blindly appending (which can violate DTD ordering),
this finds the right insertion point based on the FCPXML DTD's
required element sequence for asset-clip / clip elements.
Unknown tags are appended at the end.
"""
child_priority = _CHILD_ORDER_INDEX.get(child.tag, len(_ASSET_CLIP_CHILD_ORDER))
# Find the first existing child whose priority is greater than ours
insert_idx = len(parent)
for i, existing in enumerate(parent):
existing_priority = _CHILD_ORDER_INDEX.get(existing.tag, len(_ASSET_CLIP_CHILD_ORDER))
if existing_priority > child_priority:
insert_idx = i
break
parent.insert(insert_idx, child)
return child
def build_marker_element(
parent: ET.Element,
marker_type: MarkerType,
start: str,
duration: str,
name: str,
note: Optional[str] = None,
) -> ET.Element:
"""Create a marker or chapter-marker XML element under *parent*.
Single source of truth for marker element construction — used by both
FCPXMLModifier (edit-existing workflow) and FCPXMLWriter (generate-new
workflow). Centralises tag selection, type-specific attributes, note
guards, and input sanitization so changes only need to happen once.
"""
elem = ET.Element(marker_type.xml_tag)
elem.set('start', start)
elem.set('duration', duration)
elem.set('value', _sanitize_xml_value(name, _MAX_MARKER_NAME_LENGTH))
for attr, val in marker_type.xml_attrs.items():
elem.set(attr, val)
if note and marker_type != MarkerType.CHAPTER:
elem.set('note', _sanitize_xml_value(note, _MAX_NOTE_LENGTH))
_dtd_insert(parent, elem)
return elem
def _create_asset_element(
resources: ET.Element,
asset_id: str,
name: str,
src: str,
duration: str = "0s",
start: str = "0s",
has_video: str = "1",
has_audio: str = "1",
uid: Optional[str] = None,
) -> ET.Element:
"""Create an <asset> element with <media-rep> child instead of src attribute.
FCP's DTD prefers <media-rep kind="original-media" src="..."/> children
over the src attribute on <asset>. This helper produces the preferred form.
Args:
resources: Parent <resources> element to append to.
asset_id: Resource ID (e.g. "r3").
name: Human-readable asset name.
src: File path or URL for the media source.
duration: Asset duration in FCPXML rational format.
start: Asset start time.
has_video: "1" if asset has video track.
has_audio: "1" if asset has audio track.
uid: Optional UUID; auto-generated if not provided.
Returns:
The created <asset> Element.
"""
import uuid as _uuid
asset = ET.SubElement(resources, 'asset')
asset.set('id', asset_id)
asset.set('name', _sanitize_xml_value(name, 512))
asset.set('uid', uid or str(_uuid.uuid4()).upper())
asset.set('start', start)
asset.set('duration', duration)
asset.set('hasVideo', has_video)
asset.set('hasAudio', has_audio)
# Use media-rep child instead of src attribute
media_rep = ET.SubElement(asset, 'media-rep')
media_rep.set('kind', 'original-media')
media_rep.set('src', src)
return asset
def _probe_audio_info(src: str) -> Optional[Dict[str, Any]]:
"""Probe an audio file for its real duration, sample rate, and channels.
Tries ffprobe first, then falls back to the stdlib ``wave`` module for
.wav files. Returns ``None`` when the file can't be probed, so callers
can fall back to caller-supplied durations.
Returns:
``{'duration': float, 'sample_rate': int, 'channels': int}`` or None.
"""
path = Path(src)
if not path.is_file():
return None
try:
result = subprocess.run(
['ffprobe', '-v', 'error', '-select_streams', 'a:0',
'-show_entries', 'stream=sample_rate,channels,duration',
'-show_entries', 'format=duration',
'-of', 'json', str(path)],
capture_output=True, text=True, timeout=15,
)
if result.returncode == 0:
import json
data = json.loads(result.stdout)
streams = data.get('streams') or [{}]
stream = streams[0]
duration = stream.get('duration') or data.get('format', {}).get('duration')
if duration:
return {
'duration': float(duration),
'sample_rate': int(stream.get('sample_rate') or 48000),
'channels': int(stream.get('channels') or 2),
}
except (OSError, subprocess.TimeoutExpired, ValueError):
pass
if path.suffix.lower() == '.wav':
try:
import wave
with wave.open(str(path), 'rb') as wf:
rate = wf.getframerate()
if rate > 0:
return {
'duration': wf.getnframes() / rate,
'sample_rate': rate,
'channels': wf.getnchannels(),
}
except (OSError, wave.Error, EOFError):
pass
return None
+78
View File
@@ -0,0 +1,78 @@
"""Inserir clipes na spine.
Extraído de writer.py — ver fcpxml/writer/__init__.py para o conjunto.
"""
import xml.etree.ElementTree as ET
from typing import Optional
class InsertMixin:
"""Inserir clipes na spine."""
# INSERT CLIP OPERATIONS
# ========================================================================
def insert_clip(
self,
position: str,
asset_id: Optional[str] = None,
asset_name: Optional[str] = None,
duration: Optional[str] = None,
in_point: Optional[str] = None,
out_point: Optional[str] = None,
ripple: bool = True
) -> ET.Element:
"""
Insert a library clip onto the timeline.
Args:
position: Where to insert - 'start', 'end', timecode, or 'after:clip_id'
asset_id: Asset reference ID (e.g., 'r3')
asset_name: Asset name (alternative to asset_id)
duration: Duration of clip (if not using in/out points)
in_point: Source in-point for subclip
out_point: Source out-point for subclip
ripple: Whether to shift subsequent clips
Returns:
The created clip element
"""
asset, asset_id = self._resolve_asset(asset_id, asset_name)
clip_duration, source_start = self._resolve_clip_duration(
asset, duration, in_point, out_point
)
# Get spine and calculate insert position
spine = self._get_spine()
spine_children = list(spine)
target_offset, insert_index = self._resolve_insert_position(
position, spine_children
)
# Build extra attrs — include format from first available format
extra: dict[str, str] = {}
for fmt_id in self.formats:
extra['format'] = fmt_id
break
new_clip = self._make_asset_clip(
asset_id, asset.get('name', 'Untitled'),
target_offset, source_start, clip_duration,
**extra,
)
# Insert into spine
spine.insert(insert_index, new_clip)
# Ripple subsequent clips if needed
if ripple and insert_index < len(spine_children):
self._ripple_from_index(spine, insert_index + 1, clip_duration)
# Add to clip index
clip_id = f"inserted_{len(self.clips)}"
self.clips[clip_id] = new_clip
return new_clip
# ========================================================================
+165
View File
@@ -0,0 +1,165 @@
"""Marcadores: um, por timecode, e em lote.
Extraído de writer.py — ver fcpxml/writer/__init__.py para o conjunto.
"""
import xml.etree.ElementTree as ET
from typing import Any, Dict, List, Optional
from ..models import (
MarkerColor,
MarkerType,
TimeValue,
)
from .helpers import build_marker_element
class MarkersMixin:
"""Marcadores: um, por timecode, e em lote."""
# ========================================================================
# MARKER OPERATIONS
# ========================================================================
def add_marker(
self,
clip_id: 'str | ET.Element',
timecode: str,
name: str,
marker_type: "MarkerType | str" = MarkerType.STANDARD,
color: Optional[MarkerColor] = None,
note: Optional[str] = None
) -> ET.Element:
"""
Add a marker to a clip.
Args:
clip_id: Target clip identifier (name or ID)
timecode: Position within clip (relative to clip start)
name: Marker label
marker_type: STANDARD, TODO, COMPLETED, or CHAPTER (enum or string)
color: Optional marker color
note: Optional marker note
Returns:
The created marker element
"""
clip = self._require_clip(clip_id)
if isinstance(marker_type, str):
marker_type = MarkerType.from_string(marker_type)
time_value = self._parse_time(timecode)
return build_marker_element(
parent=clip,
marker_type=marker_type,
start=time_value.to_fcpxml(),
duration=f"1/{int(self.fps)}s",
name=name,
note=note,
)
def add_marker_at_timeline(
self,
timecode: str,
name: str,
marker_type: "MarkerType | str" = MarkerType.STANDARD,
color: Optional[MarkerColor] = None,
note: Optional[str] = None
) -> ET.Element:
"""Add a marker at a timeline position (finds the containing clip).
Uses ``_find_spine_clip_at_seconds`` to walk the spine directly,
avoiding the name-indexed ``self.clips`` dict which silently drops
duplicate-named clips.
"""
if isinstance(marker_type, str):
marker_type = MarkerType.from_string(marker_type)
time_value = self._parse_time(timecode)
target_seconds = time_value.to_seconds()
clip, relative_seconds = self._find_spine_clip_at_seconds(target_seconds)
relative_tc = TimeValue.from_seconds(relative_seconds, self.fps)
return build_marker_element(
parent=clip,
marker_type=marker_type,
start=relative_tc.to_fcpxml(),
duration=f"1/{int(self.fps)}s",
name=name,
note=note,
)
def batch_add_markers(
self,
markers: List[Dict[str, Any]],
auto_at_cuts: bool = False,
auto_at_intervals: Optional[str] = None
) -> List[ET.Element]:
"""
Add multiple markers at once.
Args:
markers: List of marker specs [{timecode, name, marker_type, color}]
auto_at_cuts: Add marker at every cut point
auto_at_intervals: Add markers at regular intervals (e.g., "00:00:30:00")
Returns:
List of created marker elements
"""
created = []
# Handle explicit markers
for m in markers:
marker = self.add_marker_at_timeline(
timecode=m['timecode'],
name=m['name'],
marker_type=MarkerType.from_string(m.get('marker_type', 'standard')),
color=MarkerColor[m['color'].upper()] if m.get('color') else None,
note=m.get('note')
)
created.append(marker)
# Auto-detect at cuts — add a marker at the start of every spine clip.
if auto_at_cuts:
for i, clip in self._iter_spine_clips():
clip_start = clip.get('start', '0s')
marker = build_marker_element(
parent=clip,
marker_type=MarkerType.STANDARD,
start=clip_start,
duration=f"1/{int(self.fps)}s",
name=f"Cut {i+1}",
)
created.append(marker)
# Auto-detect at intervals — place markers at regular time steps.
if auto_at_intervals:
interval = self._parse_time(auto_at_intervals).to_seconds()
total_duration = self._timeline_duration().to_seconds()
if total_duration > 0:
current = interval
count = 1
while current < total_duration:
try:
clip, relative = self._find_spine_clip_at_seconds(current)
except ValueError:
current += interval
count += 1
continue
rel_tv = TimeValue.from_seconds(relative, self.fps)
marker = build_marker_element(
parent=clip,
marker_type=MarkerType.STANDARD,
start=rel_tv.to_fcpxml(),
duration=f"1/{int(self.fps)}s",
name=f"Marker {count}",
)
created.append(marker)
current += interval
count += 1
return created
+64
View File
@@ -0,0 +1,64 @@
"""FCPXMLModifier — a edição de FCPXML montada a partir de um mixin por assunto.
A classe era um bloco de 3.300 linhas com dezoito assuntos dentro. Ela continua
sendo uma classe só para quem chama — `modifier.add_marker(...)` não mudou — mas
cada assunto agora mora no seu próprio arquivo e pode ser lido inteiro sem rolar
por marcadores, velocidade e legendas até achar o trecho procurado.
Mixins em vez de objetos separados por uma razão concreta: todas essas operações
mexem no *mesmo* documento e dependem dos mesmos índices e da mesma navegação na
spine (`_require_clip`, `_iter_spine_clips`, `_ripple_after_clip`). Separá-las em
objetos independentes obrigaria cada um a carregar uma referência de volta ao
documento e transformaria toda chamada interna em travessia de fronteira, sem
nada em troca — a divisão que importa aqui é de *leitura*, não de estado.
A ordem abaixo é irrelevante para o comportamento: nenhum mixin sobrescreve
método de outro; cada um contribui com um conjunto disjunto de operações.
"""
from .audio import AudioMixin
from .compound import CompoundMixin
from .connected import ConnectedMixin
from .core import ModifierCore
from .cut import CutMixin
from .insert import InsertMixin
from .markers import MarkersMixin
from .rapid import RapidMixin
from .reformat import ReformatMixin
from .relink import RelinkMixin
from .reorder import ReorderMixin
from .roles import RolesMixin
from .selection import SelectionMixin
from .silence import SilenceMixin
from .speed import SpeedMixin
from .titles import TitlesMixin
from .transitions import TransitionsMixin
from .trim import TrimMixin
class FCPXMLModifier(
RelinkMixin,
MarkersMixin,
TrimMixin,
ReorderMixin,
TransitionsMixin,
SpeedMixin,
CutMixin,
RapidMixin,
SelectionMixin,
InsertMixin,
ConnectedMixin,
TitlesMixin,
AudioMixin,
CompoundMixin,
RolesMixin,
ReformatMixin,
SilenceMixin,
ModifierCore,
):
"""Carrega um FCPXML, aplica edições cirúrgicas e salva.
Interface de escrita usada por todos os handlers do servidor MCP. A
documentação de cada operação está no mixin correspondente; o
carregamento, os índices e o `save` estão em `core.ModifierCore`.
"""
+240
View File
@@ -0,0 +1,240 @@
"""Corte rápido: flash frames, rapid trim, preencher buracos.
Extraído de writer.py — ver fcpxml/writer/__init__.py para o conjunto.
"""
from typing import Any, Dict, List, Optional
class RapidMixin:
"""Corte rápido: flash frames, rapid trim, preencher buracos."""
# SPEED CUTTING OPERATIONS (v0.3.0)
# ========================================================================
def fix_flash_frames(
self,
mode: str = 'auto',
threshold_frames: int = 6,
critical_threshold_frames: int = 2
) -> List[Dict[str, Any]]:
"""
Automatically fix flash frames (ultra-short clips).
Args:
mode: How to fix flash frames:
- 'extend_previous': Extend the previous clip to cover the flash frame
- 'extend_next': Extend the next clip backward to cover the flash frame
- 'delete': Remove the flash frame entirely (ripple)
- 'auto': Use smart logic (extend prev for critical, delete for warning)
threshold_frames: Frames below this are considered flash frames
critical_threshold_frames: Frames below this are critical (default: 2)
Returns:
List of fixed flash frames with details
"""
spine = self._get_spine()
fixed = []
# Collect flash frames first (can't modify while iterating)
flash_frames = []
for i, clip in self._iter_spine_clips():
duration = self._parse_time(clip.get('duration', '0s'))
duration_frames = duration.to_frames(self.fps)
if duration_frames < threshold_frames:
is_critical = duration_frames < critical_threshold_frames
flash_frames.append({
'index': i,
'clip': clip,
'clip_id': clip.get('name') or clip.get('id') or f"clip_{i}",
'duration_frames': duration_frames,
'is_critical': is_critical
})
# Process in reverse order to maintain indices
for ff in reversed(flash_frames):
clip = ff['clip']
_, _, clip_offset = self._get_clip_times(clip)
# Determine actual mode
actual_mode = mode
if mode == 'auto':
# Critical: try to extend previous, otherwise delete
# Warning: delete
actual_mode = 'extend_previous' if ff['is_critical'] else 'delete'
result = {
'clip_name': ff['clip_id'],
'duration_frames': ff['duration_frames'],
'was_critical': ff['is_critical'],
'action': actual_mode,
'timecode': clip_offset.to_timecode(self.fps)
}
direction = {'extend_previous': 'prev', 'extend_next': 'next'}.get(actual_mode)
if direction:
neighbor = self._absorb_into_neighbor(spine, clip, direction)
if neighbor is not None:
self._recalculate_offsets(spine)
result['extended_clip'] = neighbor.get('name', direction.title())
else:
spine.remove(clip)
self._recalculate_offsets(spine)
else: # delete
spine.remove(clip)
self._recalculate_offsets(spine)
fixed.append(result)
# Rebuild clip index
self._build_clip_index()
return fixed
def rapid_trim(
self,
max_duration: Optional[str] = None,
min_duration: Optional[str] = None,
keywords: Optional[List[str]] = None,
trim_from: str = 'end'
) -> List[Dict[str, Any]]:
"""
Batch trim clips to enforce duration limits.
Args:
max_duration: Maximum clip duration (e.g., '2s', '00:00:02:00')
min_duration: Minimum clip duration (clips shorter are extended/left alone)
keywords: Only trim clips with these keywords (None = all clips)
trim_from: Where to trim - 'start', 'end', or 'center'
Returns:
List of trimmed clips with before/after durations
"""
trimmed = []
max_dur = self._parse_time(max_duration) if max_duration else None
min_dur = self._parse_time(min_duration) if min_duration else None
for _i, clip in self._iter_spine_clips():
clip_name = clip.get('name') or clip.get('id') or 'Unknown'
# Check keyword filter
if keywords:
clip_keywords = set()
for kw_elem in clip.findall('keyword'):
clip_keywords.add(kw_elem.get('value', ''))
if not clip_keywords.intersection(set(keywords)):
continue
current_start, current_duration, _ = self._get_clip_times(clip)
original_duration = current_duration.to_seconds()
# Skip clips shorter than min_duration (leave them alone)
if min_dur and current_duration < min_dur:
continue
# Check max duration
if max_dur and current_duration > max_dur:
excess = current_duration - max_dur
if trim_from == 'end':
# Keep start, reduce duration
clip.set('duration', max_dur.to_fcpxml())
elif trim_from == 'start':
# Increase start, reduce duration
new_start = current_start + excess
clip.set('start', new_start.to_fcpxml())
clip.set('duration', max_dur.to_fcpxml())
elif trim_from == 'center':
# Trim equal amounts from both ends
half_excess = excess * 0.5
new_start = current_start + half_excess
clip.set('start', new_start.to_fcpxml())
clip.set('duration', max_dur.to_fcpxml())
trimmed.append({
'clip_name': clip_name,
'original_duration': original_duration,
'new_duration': max_dur.to_seconds(),
'trim_from': trim_from,
'action': 'trimmed'
})
# Recalculate offsets
self._recalculate_offsets(self._get_spine())
return trimmed
def fill_gaps(
self,
mode: str = 'extend_previous',
max_gap: Optional[str] = None
) -> List[Dict[str, Any]]:
"""
Fill gaps in the timeline.
Args:
mode: How to fill gaps:
- 'extend_previous': Extend previous clip to fill gap
- 'extend_next': Extend next clip backward to fill gap
- 'delete': Remove gap elements and ripple
max_gap: Only fill gaps smaller than this (None = all gaps)
Returns:
List of filled gaps with details
"""
spine = self._get_spine()
filled = []
max_gap_time = self._parse_time(max_gap) if max_gap else None
# Find all gaps
gaps_to_process = []
for i, child in enumerate(list(spine)):
if child.tag == 'gap':
gap_duration = self._parse_time(child.get('duration', '0s'))
gap_offset = self._parse_time(child.get('offset', '0s'))
# Check max_gap filter
if max_gap_time and gap_duration > max_gap_time:
continue
gaps_to_process.append({
'element': child,
'index': i,
'duration': gap_duration,
'offset': gap_offset
})
# Process in reverse to maintain indices
for gap_info in reversed(gaps_to_process):
gap = gap_info['element']
gap_duration = gap_info['duration']
gap_offset = gap_info['offset']
result = {
'timecode': gap_offset.to_timecode(self.fps),
'duration_frames': gap_duration.to_frames(self.fps),
'duration_seconds': gap_duration.to_seconds(),
'action': mode
}
direction = {'extend_previous': 'prev', 'extend_next': 'next'}.get(mode)
if direction:
neighbor = self._absorb_into_neighbor(spine, gap, direction)
if neighbor is not None:
result['extended_clip'] = neighbor.get('name', direction.title())
filled.append(result)
else: # delete
spine.remove(gap)
filled.append(result)
# Recalculate offsets
self._recalculate_offsets(spine)
return filled
# ========================================================================
+43
View File
@@ -0,0 +1,43 @@
"""Reenquadrar a resolução do projeto.
Extraído de writer.py — ver fcpxml/writer/__init__.py para o conjunto.
"""
class ReformatMixin:
"""Reenquadrar a resolução do projeto."""
# REFORMAT OPERATIONS (v0.5.0)
# ========================================================================
SOCIAL_FORMATS = {
"9:16": (1080, 1920),
"1:1": (1080, 1080),
"4:5": (1080, 1350),
"16:9": (1920, 1080),
"4:3": (1440, 1080),
}
def reformat_resolution(self, width: int, height: int) -> None:
"""Change the timeline format to a new resolution.
Updates the format resource dimensions. FCP handles spatial
conforming (letterbox/pillarbox) on import.
Args:
width: Target width in pixels
height: Target height in pixels
"""
for fmt in self.root.findall('.//format'):
fmt.set('width', str(width))
fmt.set('height', str(height))
old_name = fmt.get('name', '')
if old_name:
fmt.set('name', f"FFVideoFormat{width}x{height}")
sequence = self.root.find('.//sequence')
if sequence is not None and sequence.get('format'):
pass # format ref stays the same, dimensions updated in-place
# ========================================================================
+94
View File
@@ -0,0 +1,94 @@
"""Repontar a mídia de um projeto para novos arquivos.
Extraído de writer.py — ver fcpxml/writer/__init__.py para o conjunto.
"""
from pathlib import Path
from typing import Any, Dict
class RelinkMixin:
"""Repontar a mídia de um projeto para novos arquivos."""
# ========================================================================
# MEDIA RELINK
# ========================================================================
def relink_media(
self,
find: str,
replace: str,
dry_run: bool = False,
) -> Dict[str, Any]:
"""Bulk-rewrite media source paths (programmatic relink).
Rewrites the ``src`` of every ``<asset>`` / ``<media-rep>`` whose
path starts with *find*, substituting *replace* — the standard
technique for relinking a moved or renamed media folder without
opening Final Cut Pro. FCP relinks via the ``media-rep`` file URL
on import; the device-specific bookmark blob is left untouched
(FCP regenerates it).
*find* / *replace* accept plain paths (``/Volumes/OldDrive``) or
``file://`` URLs; percent-encoding in existing URLs is handled.
Matching is prefix-based on whole path segments, so ``/Media/A``
matches ``/Media/A/clip.mov`` but not ``/Media/AB/clip.mov``.
Args:
find: Old path prefix to match.
replace: New path prefix to substitute.
dry_run: When True, report what would change without
mutating the tree.
Returns:
Summary dict: ``total_assets``, ``relinked`` (reference
count), ``dry_run``, and ``changes`` — a list of
``{asset, old, new, target_exists}`` entries
(``target_exists`` checks the new path on this machine).
"""
from urllib.parse import quote, unquote, urlparse
def _to_path(value: str) -> str:
if value.startswith('file://'):
return unquote(urlparse(value).path)
return value
find_path = _to_path(find).rstrip('/')
replace_path = _to_path(replace).rstrip('/')
if not find_path:
raise ValueError("relink_media: 'find' must be a non-empty path prefix")
changes = []
for asset_id, info in self.resources.items():
elem = info['element']
targets = [(elem, elem.get('src'))]
media_rep = elem.find('media-rep')
if media_rep is not None:
targets.append((media_rep, media_rep.get('src')))
for node, old_src in targets:
if not old_src:
continue
was_url = old_src.startswith('file://')
old_path = _to_path(old_src)
if old_path != find_path and not old_path.startswith(find_path + '/'):
continue
new_path = replace_path + old_path[len(find_path):]
new_src = 'file://' + quote(new_path) if was_url else new_path
if not dry_run:
node.set('src', new_src)
info['src'] = new_src
changes.append({
'asset': info.get('name') or asset_id,
'old': old_src,
'new': new_src,
'target_exists': Path(new_path).exists(),
})
return {
'total_assets': len(self.resources),
'relinked': len(changes),
'dry_run': dry_run,
'changes': changes,
}
+126
View File
@@ -0,0 +1,126 @@
"""Reordenar clipes e recalcular offsets/duração.
Extraído de writer.py — ver fcpxml/writer/__init__.py para o conjunto.
"""
import xml.etree.ElementTree as ET
from typing import List
from ..models import (
TimeValue,
)
from .helpers import SPINE_ELEMENT_TAGS
class ReorderMixin:
"""Reordenar clipes e recalcular offsets/duração."""
# REORDER OPERATIONS
# ========================================================================
def reorder_clips(
self,
clip_ids: List[str],
target_position: str,
ripple: bool = True
) -> None:
"""
Move clips to a new position in the timeline.
Args:
clip_ids: Clips to move (maintains relative order)
target_position: 'start', 'end', timecode, or 'after:clip_id'/'before:clip_id'
ripple: Whether to shift other clips
"""
spine = self._get_spine()
# Collect clips to move
clips_to_move = []
for clip_id in clip_ids:
clip = self.clips.get(clip_id)
if clip is not None and clip in list(spine):
clips_to_move.append(clip)
if not clips_to_move:
raise ValueError(f"No clips found matching: {clip_ids}")
# Calculate total duration of moving clips
total_duration = TimeValue.zero()
for clip in clips_to_move:
dur = self._parse_time(clip.get('duration', '0s'))
total_duration = total_duration + dur
# Remove clips from current positions
for clip in clips_to_move:
spine.remove(clip)
# Determine target offset and insert index
spine_children = list(spine)
target_offset, insert_index = self._resolve_insert_position(
target_position, spine_children
)
# Insert clips at new position
current_offset = target_offset
for clip in clips_to_move:
clip.set('offset', current_offset.to_fcpxml())
spine.insert(insert_index, clip)
insert_index += 1
dur = self._parse_time(clip.get('duration', '0s'))
current_offset = current_offset + dur
# Recalculate all offsets if ripple
if ripple:
self._recalculate_offsets(spine)
def _recalculate_offsets(self, spine: ET.Element) -> None:
"""Recalculate all clip offsets sequentially."""
current_offset = TimeValue.zero()
for child in spine:
if child.tag in SPINE_ELEMENT_TAGS:
child.set('offset', current_offset.to_fcpxml())
duration_str = child.get('duration', '0s')
duration = self._parse_time(duration_str)
current_offset = current_offset + duration
def _timeline_duration(self) -> 'TimeValue':
"""Return the total timeline duration as a TimeValue.
Reads from the ``<sequence>`` element when available, falling back
to summing all spine element durations. Extracted from
``add_music_bed`` and ``batch_add_markers`` which both computed
this independently.
"""
sequence = self.root.find('.//sequence')
if sequence is not None:
dur_str = sequence.get('duration')
if dur_str:
return self._parse_time(dur_str)
spine = self._get_spine()
total = TimeValue.zero()
for child in spine:
if child.tag in SPINE_ELEMENT_TAGS:
total = total + self._parse_time(child.get('duration', '0s'))
return total
def _update_sequence_duration(self) -> None:
"""Recompute the ``<sequence>`` duration from the spine content.
Ripple edits (``cut_clip_ranges``, ``delete_clip``, ``split_clip``)
change the total timeline length without rewriting the sequence
element, so an exported file kept advertising the pre-edit duration —
a 326.78s sequence still claimed 326.78s after 71s of silence was
removed. This helper re-syncs the attribute to the actual spine sum.
"""
sequence = self.root.find('.//sequence')
if sequence is None:
return
spine = self._get_spine()
total = TimeValue.zero()
for child in spine:
if child.tag in SPINE_ELEMENT_TAGS:
total = total + self._parse_time(child.get('duration', '0s'))
sequence.set('duration', total.to_fcpxml())
# ========================================================================
+43
View File
@@ -0,0 +1,43 @@
"""Atribuir roles de vídeo/áudio.
Extraído de writer.py — ver fcpxml/writer/__init__.py para o conjunto.
"""
import xml.etree.ElementTree as ET
from typing import Optional
from .helpers import _sanitize_xml_value
class RolesMixin:
"""Atribuir roles de vídeo/áudio."""
# ROLE OPERATIONS (v0.5.0)
# ========================================================================
def assign_role(
self,
clip_id: str,
audio_role: Optional[str] = None,
video_role: Optional[str] = None,
) -> ET.Element:
"""Set the audio/video role on a clip.
Args:
clip_id: Name/ID of the clip
audio_role: Audio role (e.g., "dialogue", "music", "effects")
video_role: Video role (e.g., "video", "titles")
Returns:
The modified clip element
"""
clip = self._require_clip(clip_id)
if audio_role is not None:
clip.set('audioRole', _sanitize_xml_value(audio_role, 256))
if video_role is not None:
clip.set('videoRole', _sanitize_xml_value(video_role, 256))
return clip
# ========================================================================
+57
View File
@@ -0,0 +1,57 @@
"""Selecionar clipes por palavra-chave.
Extraído de writer.py — ver fcpxml/writer/__init__.py para o conjunto.
"""
from typing import List
class SelectionMixin:
"""Selecionar clipes por palavra-chave."""
# SELECTION OPERATIONS
# ========================================================================
def select_by_keyword(
self,
keywords: List[str],
match_mode: str = 'any',
favorites_only: bool = False,
exclude_rejected: bool = True
) -> List[str]:
"""
Find clips matching keywords.
Args:
keywords: Keywords to match
match_mode: 'any' (OR), 'all' (AND), 'none' (exclude)
favorites_only: Only return favorited clips
exclude_rejected: Exclude rejected clips
Returns:
List of matching clip IDs
"""
matches = []
for clip_id, clip in self.clips.items():
clip_keywords = set()
for kw_elem in clip.findall('keyword'):
clip_keywords.add(kw_elem.get('value', ''))
# Check keyword match
keyword_set = set(keywords)
if match_mode == 'any':
match = bool(clip_keywords & keyword_set)
elif match_mode == 'all':
match = keyword_set <= clip_keywords
elif match_mode == 'none':
match = not bool(clip_keywords & keyword_set)
else:
match = True
if match:
matches.append(clip_id)
return matches
# ========================================================================
+185
View File
@@ -0,0 +1,185 @@
"""Detectar e remover silêncio.
Extraído de writer.py — ver fcpxml/writer/__init__.py para o conjunto.
"""
from typing import Any, Dict, List, Optional
from ..models import (
MarkerType,
TimeValue,
)
from .helpers import CLIP_TAGS, build_marker_element
class SilenceMixin:
"""Detectar e remover silêncio."""
# SILENCE DETECTION OPERATIONS (v0.5.0)
# ========================================================================
def detect_silence_candidates(
self,
min_gap_seconds: float = 0.5,
patterns: Optional[List[str]] = None,
) -> List[Dict[str, Any]]:
"""Detect potential silence regions using timeline heuristics.
Checks for:
1. Gap elements in spine (high confidence)
2. Ultra-short clips < 0.5s (medium confidence)
3. Clips matching name patterns like "silence", "room tone" (high)
4. Duration anomalies > 2 std dev from mean (low-medium)
Args:
min_gap_seconds: Minimum gap duration to flag
patterns: Name patterns to match (default: gap, silence, room tone)
Returns:
List of silence candidate dicts
"""
if patterns is None:
patterns = ['gap', 'silence', 'room tone', 'dead air', 'blank']
spine = self._get_spine()
candidates = []
durations = []
clip_index = 0
# First pass: collect durations for anomaly detection
for child in spine:
if child.tag in CLIP_TAGS:
dur = self._parse_time(child.get('duration', '0s'))
durations.append(dur.to_seconds())
# Calculate stats for anomaly detection
mean_dur = sum(durations) / len(durations) if durations else 0
variance = (sum((d - mean_dur) ** 2 for d in durations) / len(durations)
if len(durations) > 1 else 0)
std_dev = variance ** 0.5
# Second pass: detect candidates
for child in spine:
tag = child.tag
offset = child.get('offset', '0s')
dur = self._parse_time(child.get('duration', '0s'))
dur_secs = dur.to_seconds()
tc = TimeValue.from_timecode(offset, self.fps).to_timecode(self.fps)
if tag == 'gap' and dur_secs >= min_gap_seconds:
candidates.append({
'start_timecode': tc,
'duration_seconds': dur_secs,
'reason': 'gap',
'confidence': 0.9,
'clip_name': None,
'clip_index': None,
})
elif tag in CLIP_TAGS:
name = child.get('name', '').lower()
# Name pattern match
for pat in patterns:
if pat.lower() in name:
candidates.append({
'start_timecode': tc,
'duration_seconds': dur_secs,
'reason': 'name_match',
'confidence': 0.85,
'clip_name': child.get('name', ''),
'clip_index': clip_index,
})
break
# Ultra-short clip
if dur_secs < 0.5:
candidates.append({
'start_timecode': tc,
'duration_seconds': dur_secs,
'reason': 'ultra_short',
'confidence': 0.6,
'clip_name': child.get('name', ''),
'clip_index': clip_index,
})
# Duration anomaly (> 2 std dev longer than mean)
if std_dev > 0 and dur_secs > mean_dur + 2 * std_dev:
candidates.append({
'start_timecode': tc,
'duration_seconds': dur_secs,
'reason': 'duration_anomaly',
'confidence': 0.4,
'clip_name': child.get('name', ''),
'clip_index': clip_index,
})
clip_index += 1
return candidates
def remove_silence_candidates(
self,
mode: str = "mark",
min_gap_seconds: float = 0.5,
min_confidence: float = 0.7,
patterns: Optional[List[str]] = None,
) -> List[Dict[str, Any]]:
"""Remove or mark detected silence candidates.
Args:
mode: "delete" removes clips/gaps, "mark" adds red markers,
"shorten" trims to minimum
min_gap_seconds: Minimum gap to consider
min_confidence: Only act on candidates above this threshold
patterns: Name patterns to match
Returns:
List of actions taken
"""
candidates = self.detect_silence_candidates(min_gap_seconds, patterns)
candidates = [c for c in candidates if c['confidence'] >= min_confidence]
spine = self._get_spine()
actions = []
if mode == "mark":
for c in candidates:
child = self._find_spine_element_at_timecode(
spine, c['start_timecode'], require_clip=True
)
if child is not None:
build_marker_element(
parent=child,
marker_type=MarkerType.STANDARD,
start=child.get('start', '0s'),
duration=f"1/{int(self.fps)}s",
name=f"SILENCE: {c['reason']}",
)
actions.append({
'action': 'marked',
'clip_name': c.get('clip_name', 'gap'),
'reason': c['reason'],
})
elif mode == "delete":
elements_to_remove = []
for c in candidates:
child = self._find_spine_element_at_timecode(
spine, c['start_timecode']
)
if child is not None:
elements_to_remove.append(child)
actions.append({
'action': 'deleted',
'clip_name': c.get('clip_name', 'gap'),
'reason': c['reason'],
})
for elem in elements_to_remove:
spine.remove(elem)
if elements_to_remove:
self._recalculate_offsets(spine)
return actions
+297
View File
@@ -0,0 +1,297 @@
"""Velocidade e zoom (punch-in) por janela.
Extraído de writer.py — ver fcpxml/writer/__init__.py para o conjunto.
"""
import xml.etree.ElementTree as ET
from fractions import Fraction
from typing import Optional
from .helpers import HOLD_AT_CUT_THRESHOLD, START_AT_CUT_THRESHOLD, _dtd_insert, _fmt_scale
class SpeedMixin:
"""Velocidade e zoom (punch-in) por janela."""
# SPEED OPERATIONS
# ========================================================================
def change_speed(
self,
clip_id: str,
speed: float,
preserve_pitch: bool = True
) -> ET.Element:
"""
Change clip playback speed.
Args:
clip_id: Target clip
speed: Speed multiplier (0.5 = half speed, 2.0 = double)
preserve_pitch: Maintain audio pitch
Returns:
Modified clip element
"""
if speed <= 0:
raise ValueError(f"Speed must be positive, got {speed}")
clip = self._require_clip(clip_id)
current_duration = self._parse_time(clip.get('duration', '0s'))
# Use rational arithmetic to avoid floating-point time values.
# FCPXML requires rational fractions with a consistent timebase,
# not decimal floats like "2.6666666666666665s".
denom = current_duration.denominator if current_duration.denominator > 0 else int(self.fps)
source_num = current_duration.numerator
speed_frac = Fraction(speed).limit_denominator(1000)
raw_num = source_num * speed_frac.denominator
raw_denom = denom * speed_frac.numerator
# Snap to frame boundary in a standard timebase (2400 ticks/sec).
# Each frame at Nfps = 2400/N ticks (e.g. 24fps → 100 ticks/frame).
fps_int = int(self.fps) if self.fps else 24
ticks_per_frame = 2400 // fps_int
dur_ticks = round(raw_num / raw_denom * 2400)
dur_ticks = round(dur_ticks / ticks_per_frame) * ticks_per_frame
new_num = dur_ticks
new_denom = 2400
# Remove any existing timeMap/conform-rate from a prior speed change
# to prevent duplicate children that produce invalid FCPXML.
for stale_tag in ('timeMap', 'conform-rate'):
for stale in clip.findall(stale_tag):
clip.remove(stale)
# Create timeMap for speed change (DTD-ordered insertion)
timemap = ET.Element('timeMap')
_dtd_insert(clip, timemap)
# Start keyframe
tp1 = ET.SubElement(timemap, 'timept')
tp1.set('time', '0s')
tp1.set('value', '0s')
tp1.set('interp', 'linear')
# End keyframe — use rational time, not floats
tp2 = ET.SubElement(timemap, 'timept')
tp2.set('time', f"{new_num}/{new_denom}s")
tp2.set('value', f"{source_num}/{denom}s")
tp2.set('interp', 'linear')
# Update clip duration (rational, not simplified to arbitrary denominator)
clip.set('duration', f"{new_num}/{new_denom}s")
# Add conform-rate (DTD-ordered insertion)
conform = ET.Element('conform-rate')
conform.set('scaleEnabled', '1')
conform.set('srcFrameRate', str(int(self.fps)))
_dtd_insert(clip, conform)
return clip
def add_zoom(
self,
clip_id: 'str | ET.Element',
start: float,
end: float,
scale: float = 1.3,
ease: float = 0.25,
position: str = "0 0",
ease_out: Optional[float] = None,
hold_at_end: Optional[bool] = None,
start_at_peak: Optional[bool] = None,
) -> ET.Element:
"""Add a punch-in zoom to a clip, snapping back to its framing at the end.
Animates ``<adjust-transform>``'s ``scale`` param (``<param>`` +
``<keyframeAnimation>`` of ``<keyframe>``) from the clip's current
scale up to *scale* times it, holds, then returns — all within
``[start, end]`` — clip-relative seconds (same convention as
``cut_clip_ranges``).
The two ends are deliberately asymmetric. *ease* ramps the zoom
**in** over half a second by default, fast enough to land with the
emphasised word. The way **out** is instant — a single frame — so
the moment the impact phrase ends the shot is simply back to its
normal framing and the video resumes its flow, with no drift
drawing attention to itself. Pass *ease_out* to ramp the return
gradually instead.
*hold_at_end* keeps the peak instead of returning, and
*start_at_peak* opens already zoomed with no ramp. Left as ``None``
both decide on their own from how close the window sits to the
clip's edges: a cut is itself the transition, so ramping away from
one — or back toward one — is motion the viewer reads as a wobble
rather than as emphasis.
"""
if end <= start:
raise ValueError(f"end ({end}) must be greater than start ({start})")
if ease <= 0:
raise ValueError(f"ease must be positive, got {ease}")
if scale <= 0:
raise ValueError(f"scale must be positive, got {scale}")
frame = float(self.frame_duration_fraction())
ramp_out = frame if ease_out is None else ease_out
if ramp_out <= 0:
raise ValueError(f"ease_out must be positive, got {ease_out}")
clip = self._require_clip(clip_id)
clip_duration = self._parse_time(clip.get('duration', '0s')).to_seconds()
if start < 0 or end > clip_duration:
raise ValueError(
f"zoom window [{start}, {end}]s must fall within the clip's "
f"duration (0 to {clip_duration:.3f}s)"
)
# Replace a prior zoom, but never the clip's framing. A clip can
# already carry an <adjust-transform> holding the editor's own
# reframe — rotation for footage shot sideways, position, a scale
# that makes the shot work at all. Dropping it outright (the old
# behaviour) silently destroyed that framing; on real footage the
# zoomed section came back rotated. So: keep the static attributes,
# and animate *relative to* the existing scale.
base_x, base_y = 1.0, 1.0
carried: dict = {}
old_keyframes: list = []
for stale in clip.findall('adjust-transform'):
carried = {k: v for k, v in stale.attrib.items() if k != 'scale'}
parts = (stale.get('scale') or '').split()
if len(parts) == 2:
try:
base_x, base_y = float(parts[0]), float(parts[1])
except ValueError:
base_x, base_y = 1.0, 1.0
else:
# No static attribute — a PRIOR zoom on this same clip left
# an animated <param name="scale"> instead, and the true
# resting framing lives in its keyframes, not in 1.0.
# Reading it as 1.0 here doesn't just miss the framing: it
# replaces the earlier zoom's whole animation with a wrong
# one, since this loop unconditionally removes `stale`
# right after. The rest value is recoverable without
# knowing which keyframe it is: MIN_ZOOM_SCALE == 1.0 means
# every keyframed value is >= the rest scale, so the
# smallest one keyframed is the rest value, peak or not.
for old_param in stale.findall("param[@name='scale']"):
xs, ys = [], []
for kf in old_param.findall('.//keyframe'):
kv = (kf.get('value') or '').split()
if len(kv) == 2:
try:
xs.append(float(kv[0]))
ys.append(float(kv[1]))
except ValueError:
pass
# Kept for merging: a second zoom on the same clip
# (two emphatic beats a cut didn't separate) should
# stack alongside the first, not erase it — the
# earlier peak is still a real editorial decision.
old_keyframes.append((kf.get('time', '0s'), kf.get('value', '')))
if xs and ys:
base_x, base_y = min(xs), min(ys)
clip.remove(stale)
transform = ET.Element('adjust-transform')
for key, value in carried.items():
transform.set(key, value)
scale_param = ET.SubElement(transform, 'param')
scale_param.set('name', 'scale')
anim = ET.SubElement(scale_param, 'keyframeAnimation')
# Keyframe times live in the clip's SOURCE timebase — the same origin
# as its own ``start`` — not in clip-relative seconds. A clip whose
# media starts at, say, 3109.9s of timecode looks for the animation
# there; keyframes written at 0-5s land outside the clip entirely and
# Final Cut imports the zoom as nothing at all, silently. Matches what
# add_text_title already does, and only shows up on footage whose
# start isn't 0s — every synthetic fixture starts at 0s and hides it.
media_origin = self._parse_time(clip.get('start', '0s'))
rest_value = f"{_fmt_scale(base_x)} {_fmt_scale(base_y)}"
scale_value = f"{_fmt_scale(base_x * scale)} {_fmt_scale(base_y * scale)}"
# A return that lands right before a cut is wasted motion: the next
# clip begins on its own framing anyway, so all the viewer sees is a
# twitch on the way out. When the zoom runs to the end of the clip,
# hold the peak and let the cut do the resetting.
holds_to_cut = (
hold_at_end
if hold_at_end is not None
else (clip_duration - end) <= HOLD_AT_CUT_THRESHOLD
)
opens_at_peak = (
start_at_peak
if start_at_peak is not None
else start <= START_AT_CUT_THRESHOLD
)
# Only the ramps actually written have to fit in the window: a zoom
# that opens at the peak spends no time ramping in, and one held to
# the cut spends none ramping out.
needed = (0.0 if opens_at_peak else ease) + (0.0 if holds_to_cut else ramp_out)
if needed > (end - start):
raise ValueError(
f"the ramps ({needed}s) don't fit in the zoom window "
f"({end - start}s) — shorten them or widen start/end"
)
if opens_at_peak:
# The cut already delivered the change of framing; ramping up
# from it just looks like the shot settling.
keyframes = [(start, scale_value)]
else:
keyframes = [(start, rest_value), (start + ease, scale_value)]
if holds_to_cut:
keyframes.append((end, scale_value))
else:
# Hold the peak right up to the end, then drop back on the very
# next frame — the snap-back the edit wants, not a slow drift.
keyframes.append((end - ramp_out, scale_value))
keyframes.append((end, rest_value))
new_entries = [
((media_origin + self.snap_seconds_to_frame(seconds)), value)
for seconds, value in keyframes
]
new_start_time = new_entries[0][0]
new_end_time = new_entries[-1][0]
# Two calls on the same clip mean two different things depending on
# whether their windows overlap. Overlapping = redoing the *same*
# zoom with new numbers — the old keyframes are stale and all of
# them go. Disjoint = a second, separate beat that a cut didn't
# separate onto its own clip — that one stacks alongside the first
# instead of erasing it, since both are real editorial decisions.
old_times = [self._parse_time(t) for t, _ in old_keyframes]
old_span_overlaps_new = bool(old_times) and not (
max(old_times) < new_start_time or min(old_times) > new_end_time
)
if old_span_overlaps_new:
surviving_old: list = []
else:
surviving_old = [(self._parse_time(t), v) for t, v in old_keyframes]
all_entries = sorted(surviving_old + new_entries, key=lambda e: e[0])
for time_value, value in all_entries:
kf = ET.SubElement(anim, 'keyframe')
kf.set('time', time_value.to_fcpxml())
kf.set('value', value)
# Only 'time' and 'value' — no 'interp', no 'curve'. The DTD allows
# both, but Final Cut rejected 'interp' on this vector param
# ("does not support the interpolation attribute") and discarded
# the whole <param>. A hand-made zoom exported from FCP itself
# writes bare keyframes and relies on the DTD default
# (curve="smooth"), so we match that export exactly rather than
# guess which attributes survive its importer.
if position != "0 0":
pos_param = ET.SubElement(transform, 'param')
pos_param.set('name', 'position')
pos_param.set('value', position)
_dtd_insert(clip, transform)
return clip
# ========================================================================
+867
View File
@@ -0,0 +1,867 @@
"""Títulos de texto e legendas dinâmicas.
Extraído de writer.py — ver fcpxml/writer/__init__.py para o conjunto.
"""
import random
import re
import unicodedata
import uuid
import xml.etree.ElementTree as ET
from typing import Any, Dict, List, Optional
from ..collision import blocking, validate_titles
from ..models import (
DynamicSubtitleConfig,
TimeValue,
)
from ..text_layout import (
TEXT_TEMPLATE_FONT_SCALE,
LayoutBox,
compose_sentence,
layout_sentence,
)
from ..transcribe import group_words_by_segment, split_into_subphrases
from .helpers import _dtd_insert, _sanitize_xml_value
class TitlesMixin:
"""Títulos de texto e legendas dinâmicas."""
_SUBTITLE_METADATA_KEY = 'com.gart.subtitle.kind'
def mark_generated_subtitle(self, element: ET.Element, kind: str) -> None:
metadata = element.find('metadata')
if metadata is None:
metadata = ET.Element('metadata')
_dtd_insert(element, metadata)
ET.SubElement(metadata, 'md', key=self._SUBTITLE_METADATA_KEY, value=kind)
def _generated_subtitle_kind(self, element: ET.Element) -> Optional[str]:
marker = element.find(f"metadata/md[@key='{self._SUBTITLE_METADATA_KEY}']")
if marker is not None:
return marker.get('value')
# Recognize the exact signature of older G-ART exports. A role alone
# is not ownership: users also assign these roles to manual titles.
if element.tag == 'title':
if element.get('start') != self._TEXT_TITLE_START:
return None
effect = self.root.find(f".//resources/effect[@id='{element.get('ref')}']")
if effect is None or effect.get('uid') != self._TEXT_TITLE_UID:
return None
if re.fullmatch(r'caption_[0-9a-f]{8}', element.get('name', '')):
return 'dynamic'
text = ''.join(element.findtext('text/text-style', ''))
if (element.get('role') == 'titles.convencionais'
and element.get('lane') == '20'
and element.get('name') == f'{text} - Text'):
return 'plain'
elif element.tag == 'ref-clip':
media = self.root.find(f".//resources/media[@id='{element.get('ref')}']")
if media is not None:
titles = media.findall('.//title')
if titles and all(self._generated_subtitle_kind(t) == 'dynamic' for t in titles):
return 'dynamic'
return None
def remove_generated_subtitles(self, parent: ET.Element, kinds: tuple) -> None:
"""Replace only our own captions, preserving unrelated graphics."""
resources = self.root.find('.//resources')
for child in list(parent):
if self._generated_subtitle_kind(child) not in kinds:
continue
parent.remove(child)
if child.tag == 'ref-clip' and resources is not None:
ref = child.get('ref')
if not self.root.findall(f".//ref-clip[@ref='{ref}']"):
media = resources.find(f"media[@id='{ref}']")
if media is not None:
resources.remove(media)
def suppress_plain_under_dynamic(self, parent: ET.Element) -> None:
"""Keep generated plain titles only on frames without dynamic text."""
import copy
windows = []
for child in parent:
if self._generated_subtitle_kind(child) == 'dynamic':
start = self._parse_time(child.get('offset', '0s'))
windows.append((start, start + self._parse_time(child.get('duration', '0s'))))
for title in list(parent):
if self._generated_subtitle_kind(title) != 'plain':
continue
start = self._parse_time(title.get('offset', '0s'))
end = start + self._parse_time(title.get('duration', '0s'))
remaining = [(start, end)]
for lo, hi in windows:
parts = []
for a, b in remaining:
if a < hi and lo < b:
if a < lo:
parts.append((a, lo))
if hi < b:
parts.append((hi, b))
else:
parts.append((a, b))
remaining = parts
if remaining == [(start, end)]:
continue
parent.remove(title)
for a, b in remaining:
part = copy.deepcopy(title)
self._reassign_text_style_ids(part)
part.set('offset', a.to_fcpxml())
part.set('duration', (b - a).to_fcpxml())
_dtd_insert(parent, part)
# DYNAMIC (KARAOKE-STYLE) SUBTITLES
# ========================================================================
# The "Text" (Basic Text) template — the ONLY simple title template that
# Final Cut actually renders. Copied verbatim from the user's own FCP
# exports ("teste.fcpxmld" and "posição.fcpxmld", FCP 1.14 in English):
# a single "<text>" run, one "<text-style-def>", and a fixed param block
# with the margins/alignment/speed the template ships with. Every prior
# title template we generated ("Essencial - Título", "Título Básico")
# imported cleanly but never appeared — their Motion uids did not resolve
# to a real, drawable template in FCP, which discards the clip silently.
# "Text" is what FCP itself writes when the user adds a title by hand, so
# it is the ground truth. See Engine/docs/05_EXPERIENCIAS.md, 2026-08-17.
_TEXT_TITLE_UID = (
'.../Titles.localized/Basic Text.localized/'
'Text.localized/Text.moti'
)
_TEXT_TITLE_START = '86486400/24000s'
# The Inspector's Position field, and the one this code overrides per
# title so two titles never stack on top of each other. Verified in
# "posição.fcpxmld": each hand-dragged title carries a distinct "x y"
# value here while every other param stays identical.
_TEXT_POSITION_KEY = '9999/10003/13260/3296672360/1/100/101'
# Layout params the "Text" template ships with. These keys are the
# template's own defaults and never vary between instances.
#
# "Build Out" is the one deliberate override: with "Apply Speed" set to
# "2 (Per Object)" below, the template's whole built-in animation (build
# in + build out) is always compressed to exactly fill the title's own
# on-screen duration — so on a short word-length clip, build out was
# eating time that build in needed to finish revealing the text before
# the cut. Disabling build out hands that entire compressed window to
# build in alone, which is what "sempre acelerado" turned out to mean:
# no separate speed knob needed. Value captured from a real FCP export
# with "Build Out" unchecked in the Inspector (see chat, 2026-08-18).
_TEXT_TITLE_PARAMS = (
('Build Out', '9999/10000/2/102', '0'),
('Layout Method', '9999/10003/13260/3296672360/2/314', '1 (Paragraph)'),
('Left Margin', '9999/10003/13260/3296672360/2/323', '-1210'),
('Right Margin', '9999/10003/13260/3296672360/2/324', '1210'),
('Top Margin', '9999/10003/13260/3296672360/2/325', '2160'),
('Bottom Margin', '9999/10003/13260/3296672360/2/326', '-2160'),
('Alignment', '9999/10003/13260/3296672360/2/354/3296667315/401', '1 (Center)'),
('Line Spacing', '9999/10003/13260/3296672360/2/354/3296667315/404', '-19'),
('Auto-Shrink', '9999/10003/13260/3296672360/2/370', '3 (To All Margins)'),
('Alignment', '9999/10003/13260/3296672360/2/373', '0 (Left) 1 (Middle)'),
('Opacity', '9999/10003/13260/3296672360/4/3296673134/1000/1044', '0'),
('Speed', '9999/10003/13260/3296672360/4/3296673134/201/208', '6 (Custom)'),
('Apply Speed', '9999/10003/13260/3296672360/4/3296673134/201/211', '2 (Per Object)'),
)
# "Custom Speed" sits between "Speed" and "Apply Speed" and carries a
# <keyframeAnimation> child rather than a plain value attribute. Its two
# keyframes are the template's own absolute nominal times, constant across
# every instance, so they are safe to replay verbatim.
_TEXT_CUSTOM_SPEED_KEY = '9999/10003/13260/3296672360/4/3296673134/201/209'
_TEXT_CUSTOM_SPEED_KEYFRAMES = (
('-469658744/1000000000s', '0'),
('12328542033/1000000000s', '1'),
)
_TEXT_SIZE_KEY = '9999/10003/13260/3296672360/5/3296672362/3'
def _ensure_text_title_effect(self, resources: ET.Element) -> str:
"""Return the resource id of the "Text" (Basic Text) effect, creating it if absent."""
return self._ensure_effect(resources, self._TEXT_TITLE_UID, 'Text', 'r_text')
def _ensure_effect(
self,
resources: ET.Element,
uid: str,
name: str,
id_prefix: str,
) -> str:
"""Return the id of the effect resource with *uid*, creating it if absent."""
for eff in resources.findall('effect'):
if eff.get('uid') == uid:
return eff.get('id')
effect_id = self._unique_resource_id(resources, id_prefix)
eff_el = ET.SubElement(resources, 'effect')
eff_el.set('id', effect_id)
eff_el.set('name', name)
eff_el.set('uid', uid)
return effect_id
# <text-style-def id> / <text-style ref> are DTD type ID/IDREF, so the
# value must be a valid XML Name: letters, digits, "_", "-", "." only,
# never starting with a digit. Title names are built from the caption
# text ("Olá mundo - Text"), which carries spaces, accents and often a
# leading digit — xmllint rejected the whole document with "Syntax of
# value for attribute id of text-style-def is not valid".
_TEXT_STYLE_ID_UNSAFE = re.compile(r'[^A-Za-z0-9_.-]+')
def _unique_text_style_id(self, base: str) -> str:
"""Return a document-unique, DTD-valid XML ID for a ``<text-style-def>``."""
folded = unicodedata.normalize('NFKD', base).encode('ascii', 'ignore').decode('ascii')
slug = self._TEXT_STYLE_ID_UNSAFE.sub('_', folded).strip('_.-')[:48]
stem = f"ts_{slug}" if slug else "ts"
if self._text_style_ids is None:
self._text_style_ids = {
sd.get('id') for sd in self.root.findall('.//text-style-def')
}
candidate = f"{stem}_0"
counter = 0
while candidate in self._text_style_ids:
counter += 1
candidate = f"{stem}_{counter}"
self._text_style_ids.add(candidate)
return candidate
def _reassign_text_style_ids(self, clip: ET.Element) -> None:
"""Give every ``<text-style-def>`` inside a just-deepcopy'd *clip* a
fresh document-unique id, repointing any ``<text-style ref="...">``
in the same subtree that pointed at the old one.
``split_clip``/``cut_clip_ranges`` deepcopy the clip once per
resulting segment, so a clip carrying a ``<title>`` (from a "text"
voice action) keeps the exact same ``text-style-def id`` in every
copy. A single cut is harmless — but the batch chain re-cuts the
same clip at each step (silence removal, filler removal, dynamic
subtitles), and every pass multiplies the duplicate, so the DTD
validator eventually rejects the file with "ID ... already
defined". Regenerating here, at the only place copies are made,
fixes it for every caller instead of each one having to remember to.
"""
for style_def in clip.findall('.//text-style-def'):
old_id = style_def.get('id')
if not old_id:
continue
slug = old_id[3:] if old_id.startswith('ts_') else old_id
slug = re.sub(r'_\d+$', '', slug) # drop a prior _<N> counter
new_id = self._unique_text_style_id(slug)
if new_id == old_id:
continue
style_def.set('id', new_id)
for ref_el in clip.findall(f".//text-style[@ref='{old_id}']"):
ref_el.set('ref', new_id)
def _unique_tracking_shape_id(self, base: str) -> str:
"""Return a document-unique ``id`` for a ``<tracking-shape>``."""
stem = base or "tr"
if self._tracking_shape_ids is None:
self._tracking_shape_ids = {
ts.get('id') for ts in self.root.findall('.//tracking-shape')
}
candidate = f"{stem}_0"
counter = 0
while candidate in self._tracking_shape_ids:
counter += 1
candidate = f"{stem}_{counter}"
self._tracking_shape_ids.add(candidate)
return candidate
def _reassign_tracking_shape_ids(self, clip: ET.Element) -> None:
"""Give every ``<tracking-shape>`` inside a just-deepcopy'd *clip* a
fresh document-unique id.
Same mechanism as ``_reassign_text_style_ids``: ``split_clip``/
``cut_clip_ranges`` deepcopy the clip once per resulting segment, so
Cinematic object-tracking data (``<object-tracker><tracking-shape
id="tr1">``, preserved from the source asset's sidecar) keeps the
exact same id in every copy. A single cut is harmless — but the
batch chain re-cuts the same clip at each step, multiplying the
duplicate until the DTD validator rejects the file with "ID tr1
already defined".
"""
for shape in clip.findall('.//tracking-shape'):
old_id = shape.get('id')
if not old_id:
continue
base = re.sub(r'_\d+$', '', old_id)
new_id = self._unique_tracking_shape_id(base)
if new_id == old_id:
continue
shape.set('id', new_id)
def _make_text_title_clip(
self,
effect_id: str,
text: str,
offset: 'TimeValue',
duration: 'TimeValue',
*,
lane: int,
name: str,
position: Optional[str] = None,
font: str = 'Helvetica Neue',
font_size: int = 196,
font_color: str = '1 1 1 1',
bold: bool = True,
face: Optional[str] = None,
kerning: Optional[float] = None,
font_scale: float = TEXT_TEMPLATE_FONT_SCALE,
animated: bool = True,
size_param: Optional[float] = None,
role: Optional[str] = None,
) -> ET.Element:
"""Build a standalone ``<title>`` clip from the "Text" (Basic Text) template.
Reproduces FCP's own output for a hand-added title exactly — the only
template we have verified renders in Final Cut ("teste.fcpxmld" and
"posição.fcpxmld"). *position* ("x y" canvas points) is the Inspector
Position value; omit it to keep the template's centred default. Unlike
the animated templates, this carries no animation switch, so the text
stays put and visible for its whole duration.
"""
elem = ET.Element('title')
elem.set('ref', effect_id)
elem.set('lane', str(lane))
elem.set('offset', offset.to_fcpxml())
elem.set('name', _sanitize_xml_value(name, 256))
elem.set('start', self._TEXT_TITLE_START)
elem.set('duration', duration.to_fcpxml())
if role:
elem.set('role', _sanitize_xml_value(role, 256))
if position:
param = ET.SubElement(elem, 'param')
param.set('name', 'Position')
param.set('key', self._TEXT_POSITION_KEY)
param.set('value', position)
def _add_param(name: str, key: str, value: str) -> None:
param = ET.SubElement(elem, 'param')
param.set('name', name)
param.set('key', key)
param.set('value', value)
animation_params = {'Opacity', 'Speed', 'Apply Speed'}
for param_name, param_key, param_value in self._TEXT_TITLE_PARAMS:
if not animated and param_name in animation_params:
continue
_add_param(param_name, param_key, param_value)
if animated and param_name == 'Speed':
# "Custom Speed" lands between "Speed" and "Apply Speed" and
# carries a <keyframeAnimation> child instead of a value.
cs = ET.SubElement(elem, 'param')
cs.set('name', 'Custom Speed')
cs.set('key', self._TEXT_CUSTOM_SPEED_KEY)
anim = ET.SubElement(cs, 'keyframeAnimation')
for kf_time, kf_value in self._TEXT_CUSTOM_SPEED_KEYFRAMES:
kf = ET.SubElement(anim, 'keyframe')
kf.set('time', kf_time)
kf.set('value', kf_value)
if size_param is not None:
_add_param('Size', self._TEXT_SIZE_KEY, f"{float(size_param):g}")
text_el = ET.SubElement(elem, 'text')
ts_id = self._unique_text_style_id(name)
run = ET.SubElement(text_el, 'text-style')
run.set('ref', ts_id)
run.text = _sanitize_xml_value(text, 256)
style_def = ET.SubElement(elem, 'text-style-def')
style_def.set('id', ts_id)
text_style = ET.SubElement(style_def, 'text-style')
text_style.set('font', font)
# Text.moti sizes type in frame pixels but positions in canvas points.
# See TEXT_TEMPLATE_FONT_SCALE: layout measures in points, so only the
# emitted size (and its kerning, to keep the same letter spacing) is
# converted here.
scale = float(font_scale) or 1.0
text_style.set('fontSize', f"{float(font_size) * scale:g}")
text_style.set('fontColor', font_color)
# FCP represents bold weight as the bold attribute — never as a
# fontFace. Writing ``bold="0" fontFace="Bold"`` (the previous
# behaviour) is contradictory and FCP refuses to render the text.
# Italic, by contrast, IS a face: FCP writes both ``fontFace`` and
# ``italic="1"``. See Engine/docs/05_EXPERIENCIAS.md, entry 2026-08-19.
face_lower = (face or '').strip().lower()
if face_lower == 'bold':
text_style.set('bold', '1')
elif 'italic' in face_lower:
text_style.set('fontFace', face)
text_style.set('italic', '1')
else:
if bold:
text_style.set('bold', '1')
if face:
text_style.set('fontFace', face)
if kerning:
text_style.set('kerning', f"{float(kerning) * scale:g}")
text_style.set('alignment', 'center')
text_style.set('lineSpacing', '-19')
return elem
def add_text_title(
self,
parent_clip: 'str | ET.Element',
text: str,
*,
offset: str = '0s',
duration: str = '1s',
lane: int = 1,
position: Optional[str] = None,
font: str = 'Helvetica Neue',
font_size: int = 196,
font_color: str = '1 1 1 1',
bold: bool = True,
face: Optional[str] = None,
animated: bool = True,
font_scale: float = TEXT_TEMPLATE_FONT_SCALE,
size_param: Optional[float] = None,
role: Optional[str] = None,
) -> ET.Element:
"""Add a single static "Text" (Basic Text) title over *parent_clip*.
Anchored in SOURCE media coordinates (parent's ``start`` + *offset*),
matching FCP's own output, so the title lands on screen instead of at
~0s of the media (which FCP silently drops). *offset* and *duration*
accept any FCPXML rational-time string; *position* is an optional
"x y" canvas-point string to keep two titles from stacking.
Returns:
The created ``<title>`` element, already inserted into the parent
in DTD order.
"""
parent = parent_clip if isinstance(parent_clip, ET.Element) else self._require_clip(parent_clip)
resources = self.root.find('.//resources')
if resources is None:
raise ValueError("No <resources> element found in FCPXML")
effect_id = self._ensure_text_title_effect(resources)
media_origin = self._parse_time(parent.get('start', '0s'))
relative = self._parse_time(offset)
title = self._make_text_title_clip(
effect_id,
text,
media_origin + relative,
self._parse_time(duration),
lane=lane,
name=f"{text} - Text",
position=position,
font=font,
font_size=font_size,
font_color=font_color,
bold=bold,
face=face,
animated=animated,
font_scale=font_scale,
size_param=size_param,
role=role,
)
_dtd_insert(parent, title)
return title
def generate_dynamic_subtitles(
self,
parent_clip: 'str | ET.Element',
words: List[Dict[str, Any]],
config: Optional['DynamicSubtitleConfig'] = None,
segments: Optional[List[Dict[str, Any]]] = None,
role: Optional[str] = None,
configs: Optional[List['DynamicSubtitleConfig']] = None,
compound_subphrases: bool = False,
subphrase_min_words: int = 3,
hold_between_sentences: bool = True,
) -> List[ET.Element]:
"""Generate progressive-reveal subtitle titles, one per word.
Groups *words* into sentences (by *segments*' time windows), lays each
sentence out as a compact typographic block, and emits one standalone
``<title>`` per word, positioned at its place in that block. Words
appear one by one as they are spoken and accumulate on screen; every
word of a block then clears at the same instant, so the sentence
vanishes as a whole before the next one builds up.
Each word gets its own lane, since a block's words are all on screen
together. Lanes restart with each block. Size, colour, font and face
cycle through ``config.style.rhythm``, reproducing the typography of
the calibration export the user built in Final Cut.
A sentence too tall for the band is split into successive blocks, so a
long sentence never spills off screen.
Args:
parent_clip: The spine clip to attach titles to — either its
Name/ID (resolved via ``_require_clip``, kept for backward
compatibility) or the ``ET.Element`` itself. **Callers
iterating multiple spine clips must pass the element, not
the name**: after any ripple-cut/silence-removal operation,
every fragment of an originally-named clip keeps that same
``name``, so ``self.clips`` (keyed by name) only retains the
last-indexed one — a name lookup then silently resolves
every call to the SAME wrong clip, stacking every line from
every distinct clip's transcript onto one spine element (see
Engine/docs/05_EXPERIENCIAS.md, entry 2026-08-17).
words: ``[{'word': str, 'start': float, 'end': float}, ...]``
with ``start``/``end`` in seconds *relative to the parent
clip's own start* (same convention as ``add_connected_clip``'s
``offset``).
config: Styling/layout options; defaults to ``DynamicSubtitleConfig()``.
segments: Whisper sentence segments ``[{'start', 'end', ...}]``, on
the same relative timebase as *words*. Omitted, every word
falls into a single sentence, which the block layout then
splits by height alone.
Returns:
The list of created ``<title>`` elements, in chronological order.
"""
# ``configs`` (a list of registered, active layouts) takes precedence
# over the single ``config`` — with 2+ items, each block picks one at
# random below; with 0 or 1, behaviour is identical to a single fixed
# config, so old callers passing only ``config`` are unaffected.
if configs:
layout_configs = list(configs)
elif config is not None:
layout_configs = [config]
else:
layout_configs = [DynamicSubtitleConfig()]
if not words:
return []
parent = parent_clip if isinstance(parent_clip, ET.Element) else self._require_clip(parent_clip)
resources = self.root.find('.//resources')
if resources is None:
raise ValueError("No <resources> element found in FCPXML")
effect_id = self._ensure_text_title_effect(resources)
# A connected title is NOT trimmed by its parent clip's out-point —
# Final Cut keeps drawing it over whatever clip follows. A word that
# starts after the cut would therefore only ever be seen on top of the
# NEXT clip's own captions, so it is dropped rather than placed.
parent_limit = self._parse_time(parent.get('duration', '0s'))
has_limit = TimeValue(0, 1) < parent_limit
if has_limit:
limit_seconds = parent_limit.to_seconds()
words = [
w for w in words
if float(w.get('start', 0.0)) < limit_seconds
]
if not words:
return []
# Split into sentences, then lay each one out as a block. A sentence
# too tall for the band comes back with overflow, which becomes the
# next block — the sub-sentence split that keeps long sentences from
# spilling off screen.
sentences = group_words_by_segment(words, segments or [])
# A comma is where the sentence breathes, so it is also where the
# phrase should be packed into its own compound clip downstream.
if compound_subphrases:
sentences = [
sub
for sentence in sentences
for sub in split_into_subphrases(sentence, subphrase_min_words)
]
def box_for(cfg: 'DynamicSubtitleConfig') -> LayoutBox:
return LayoutBox.for_frame(
self.frame_width(), self.frame_height(),
band_height=cfg.band_height,
center_y=cfg.block_center_y,
)
def lay_out(pending: List[Dict], cfg: 'DynamicSubtitleConfig', box: LayoutBox):
"""Place what fits; return (units, still-unplaced words)."""
# "phrase" is the progressive composition the reference reel uses:
# one title per LINE ("que vão" / "melhorar" / "sua legenda"), the
# key word set large in a display italic. "word" is the older
# one-title-per-word rhythm, kept for callers that want every word
# to land on its own.
if getattr(cfg, 'granularity', 'phrase') == 'phrase':
composition = compose_sentence(
pending, cfg.style, box, line_gap=cfg.line_gap,
)
return composition.blocks, composition.overflow
layout = layout_sentence(pending, cfg.style, box)
return layout.placed, layout.overflow
blocks: List[List[Any]] = []
block_configs: List['DynamicSubtitleConfig'] = []
block_sentences: List[int] = []
for sentence_index, sentence in enumerate(sentences):
remaining = list(sentence)
while remaining:
# Each block independently samples a layout from the active
# set — the visual variety the user asked for. A single
# active layout always resolves to itself, so this is a
# no-op for the common case.
active_config = layout_configs[random.randrange(len(layout_configs))]
units, remaining = lay_out(remaining, active_config, box_for(active_config))
if not units:
break
blocks.append(units)
block_configs.append(active_config)
block_sentences.append(sentence_index)
if not blocks:
return []
# Never emit a zero-duration frame (rounds to 0 at the sequence's fps
# and FCP rejects it as "unexpected value found").
min_dur_tv = self.snap_seconds_to_frame(
float(self.frame_duration_fraction())
)
# Every word of a block clears at the same instant: when the next block
# starts, or at the last word's end for the final block. That is what
# makes a sentence build up and then vanish all at once.
block_starts = [
self.snap_seconds_to_frame(min(unit.start for unit in units))
for units in blocks
]
# Whisper's word end can also run past the cut, so a last block would
# linger over the next clip's first block. Nothing may outlive the
# clip it was written for.
block_ends: List[TimeValue] = []
for i, units in enumerate(blocks):
if i + 1 < len(blocks):
end = block_starts[i + 1]
if not hold_between_sentences and block_sentences[i] != block_sentences[i + 1]:
spoken_end = self.snap_seconds_to_frame(max(unit.end for unit in units))
end = min(end, spoken_end)
else:
end = self.snap_seconds_to_frame(
max(unit.end for unit in units)
)
if end - block_starts[i] < min_dur_tv:
end = block_starts[i] + min_dur_tv
if has_limit and parent_limit < end:
end = parent_limit
block_ends.append(end)
# Anchored titles are positioned in the parent clip's SOURCE media
# coordinates: a title's offset is the parent clip's `start` plus its
# timeline-relative position. Verified against FCP's own output in
# "exemplo de arquivos.fcpxmld", where the hand-made "Essencial -
# Título" sits at offset 226040815/24000s on a parent starting at
# 226007782/24000s — 1.376s into a 1.835s clip. Writing a plain
# relative offset instead would drop the title to ~0s of the media,
# before the clip's own in-point, so it lands outside the clip and FCP
# never shows it.
media_origin = self._parse_time(parent.get('start', '0s'))
created: List[ET.Element] = []
by_sentence: Dict[int, List[ET.Element]] = {}
for units, block_end, block_config, sentence_index in zip(
blocks, block_ends, block_configs, block_sentences
):
# A ``titles.*`` sub-role keeps these as titles (never closed
# captions) while grouping them in the role index and tinting
# their lane. An explicit ``role`` argument overrides every
# block; otherwise each block uses its own sampled layout's role.
block_role = role or getattr(block_config, "role", None) or "titles.dinamicas"
for index, unit in enumerate(units):
relative_offset = self.snap_seconds_to_frame(unit.start)
duration = block_end - relative_offset
if duration < min_dur_tv:
duration = min_dur_tv
# Units of one block are all on screen together, so no two may
# share a lane. Lanes restart each block, which is free — the
# previous block has already cleared.
lane = index + 1
offset = media_origin + relative_offset
title = self._make_text_title_clip(
effect_id,
unit.text,
offset,
duration,
lane=lane,
name=f"caption_{uuid.uuid4().hex[:8]}",
position=unit.position_param(block_config.text_scale),
font=unit.font or block_config.style.font,
font_size=int(round(unit.font_size)),
font_color=unit.color or block_config.style.active_color,
bold=block_config.style.bold,
face=unit.face,
kerning=unit.kerning,
font_scale=block_config.text_scale,
role=block_role,
)
_dtd_insert(parent, title)
self.mark_generated_subtitle(title, 'dynamic')
created.append(title)
by_sentence.setdefault(sentence_index, []).append(title)
# One compound per sub-phrase: a dozen stacked title bars collapse
# into a single one that can be dragged, muted or retimed as a unit.
if compound_subphrases:
for sentence_index in sorted(by_sentence):
group = by_sentence[sentence_index]
label = " ".join(
str(w.get('word') or w.get('text') or '')
for w in sentences[sentence_index]
).strip()
compound = self.wrap_titles_in_compound(
parent, group, name=label[:60] or "Legenda"
)
self.mark_generated_subtitle(compound, 'dynamic')
if any(getattr(cfg, 'validate', False) for cfg in layout_configs):
report = self.validate_subtitle_layout()
if blocking(report["severity"]):
raise ValueError(
"Subtitle layout validation failed: "
+ str(report["summary"])
)
return created
def validate_subtitle_layout(
self,
*,
safe_margin_x: float = 0.05,
safe_margin_y: float = 0.05,
min_font_size: Optional[float] = None,
min_distance: Optional[float] = None,
max_distance: Optional[float] = None,
) -> dict:
"""Re-measure every ``<title>`` in the document and report collisions.
Reconstructs each title's on-screen box from the values the writer
emitted (``fontSize``/``kerning``/``Position`` are already in template
space), then checks for temporal+spatial collisions, frame/safe-area
containment, and font fallbacks. This is the spec-16 validation pass the
layout engine does not do on its own — it only guarantees non-overlap
*by construction* while composing, and cannot see a hand-edited title.
Returns the ``collision.validate_titles`` report: ``severity`` (worst
bucket), ``issues`` (spec-16 occurrences) and ``summary`` (counts).
"""
# A compound clip carries its own time origin: a title inside one is
# offset from that compound's start, not the sequence's. Measured in
# one flat pass, the anchors of two different compounds both read as
# "0s" and collide on paper while sitting seconds apart on the
# timeline. Each compound is therefore measured as its own scope,
# which is also where its titles can actually overlap — a title can
# only share the screen with its own compound's siblings.
scopes: List[List[ET.Element]] = []
nested: set = set()
for media in self.root.findall('.//media'):
group = list(media.iter('title'))
if group:
scopes.append(group)
nested.update(id(t) for t in group)
main = [t for t in self.root.iter('title') if id(t) not in nested]
if main:
scopes.append(main)
reports = [
self._measure_title_scope(
scope,
safe_margin_x=safe_margin_x,
safe_margin_y=safe_margin_y,
min_font_size=min_font_size,
min_distance=min_distance,
max_distance=max_distance,
)
for scope in scopes
]
if len(reports) == 1:
return reports[0]
if not reports:
return self._measure_title_scope(
[],
safe_margin_x=safe_margin_x,
safe_margin_y=safe_margin_y,
min_font_size=min_font_size,
min_distance=min_distance,
max_distance=max_distance,
)
rank = {
'none': 0, 'render_tolerance': 1, 'warning': 2,
'probable': 3, 'severe': 4,
}
merged_issues = [i for r in reports for i in r['issues']]
summary = dict(reports[0]['summary'])
for r in reports[1:]:
for key, value in r['summary'].items():
summary[key] = summary.get(key, 0) + value
return {
'severity': max(
(r['severity'] for r in reports),
key=lambda s: rank.get(s, 0),
),
'issues': merged_issues,
'summary': summary,
}
def _measure_title_scope(
self,
elements: List[ET.Element],
*,
safe_margin_x: float,
safe_margin_y: float,
min_font_size: Optional[float],
min_distance: Optional[float],
max_distance: Optional[float],
) -> dict:
"""Measure and validate one group of titles sharing a time origin."""
titles = []
for elem in elements:
# enabled="0" never renders in Final Cut (see
# generate_subtitles_by_emphasis, which disables plain titles
# under an emphasis phrase instead of never creating them) — a
# title that is off by design must not count as a collision
# against the one drawn in its place.
if elem.get('enabled', '1') == '0':
continue
text_el = elem.find('text/text-style')
text = (text_el.text or '').strip() if text_el is not None else ''
style = elem.find('text-style-def/text-style')
font = style.get('font') if style is not None else None
face = style.get('fontFace') if style is not None else None
font_size = (
float(style.get('fontSize', '0')) if style is not None else 0.0
)
kerning = (
float(style.get('kerning', '0') or 0)
if style is not None else 0.0
)
x = y = 0.0
for param in elem.findall('param'):
if param.get('name') == 'Position' and param.get('value'):
parts = param.get('value').split()
if len(parts) >= 2:
x, y = float(parts[0]), float(parts[1])
start = self._parse_time(elem.get('offset', '0s')).to_seconds()
duration = self._parse_time(elem.get('duration', '0s')).to_seconds()
titles.append({
'text': text,
'font': font,
'face': face,
'font_size': font_size,
'kerning': kerning,
'x': x,
'y': y,
'start': start,
'end': start + duration,
'group': start + duration,
})
return validate_titles(
titles,
self.frame_width(),
self.frame_height(),
safe_margin_x=safe_margin_x,
safe_margin_y=safe_margin_y,
min_font_size=min_font_size,
min_distance=min_distance,
max_distance=max_distance,
)
# ========================================================================
+94
View File
@@ -0,0 +1,94 @@
"""Transições entre clipes vizinhos.
Extraído de writer.py — ver fcpxml/writer/__init__.py para o conjunto.
"""
import xml.etree.ElementTree as ET
from ..models import (
TimeValue,
)
from .helpers import FCP_EFFECTS
class TransitionsMixin:
"""Transições entre clipes vizinhos."""
# TRANSITION OPERATIONS
# ========================================================================
def add_transition(
self,
clip_id: str,
position: str = 'end',
transition_type: str = 'cross-dissolve',
duration: str = '00:00:00:15'
) -> ET.Element:
"""
Add a transition to a clip.
Args:
clip_id: Target clip
position: 'start', 'end', or 'both'
transition_type: Type of transition
duration: Transition duration
Returns:
Created transition element(s)
"""
spine, clip, clip_index = self._require_spine_clip(clip_id)
trans_duration = self._parse_time(duration)
# Effect name and FCP built-in effect UID lookup via registry
effect_name, effect_uid = FCP_EFFECTS.get(
transition_type,
FCP_EFFECTS['cross-dissolve']
)
# Ensure effect resource exists in <resources>
effect_ref_id = None
if effect_uid:
root = self.tree.getroot()
resources = root.find('.//resources')
if resources is not None:
for eff in resources.findall('effect'):
if eff.get('uid') == effect_uid:
effect_ref_id = eff.get('id')
break
if effect_ref_id is None:
effect_ref_id = self._unique_resource_id(resources, 'r_dissolve')
eff_el = ET.SubElement(resources, 'effect')
eff_el.set('id', effect_ref_id)
eff_el.set('name', effect_name)
eff_el.set('uid', effect_uid)
transitions_added = []
_, clip_dur, clip_offset = self._get_clip_times(clip)
half_dur = trans_duration * 0.5
if position in ('end', 'both'):
end_offset = clip_offset + clip_dur - half_dur
transition = self._make_transition_element(
effect_name, end_offset, trans_duration, effect_ref_id
)
spine.insert(clip_index + 1, transition)
transitions_added.append(transition)
if position in ('start', 'both'):
start_offset = clip_offset - half_dur
if start_offset < TimeValue.zero():
raise ValueError(
f"Transition at start would produce negative offset "
f"({start_offset.to_seconds():.3f}s) for clip '{clip_id}'"
)
transition = self._make_transition_element(
effect_name, start_offset, trans_duration, effect_ref_id
)
spine.insert(clip_index, transition)
transitions_added.append(transition)
return transitions_added[0] if len(transitions_added) == 1 else transitions_added
# ========================================================================
+125
View File
@@ -0,0 +1,125 @@
"""Aparar clipes e propagar o ripple pela spine.
Extraído de writer.py — ver fcpxml/writer/__init__.py para o conjunto.
"""
import xml.etree.ElementTree as ET
from typing import Optional
from ..models import (
TimeValue,
)
from .helpers import SPINE_ELEMENT_TAGS
class TrimMixin:
"""Aparar clipes e propagar o ripple pela spine."""
# ========================================================================
# TRIM OPERATIONS
# ========================================================================
def trim_clip(
self,
clip_id: str,
trim_start: Optional[str] = None,
trim_end: Optional[str] = None,
ripple: bool = True
) -> ET.Element:
"""
Trim a clip's in-point and/or out-point.
Args:
clip_id: Target clip
trim_start: New in-point or delta ('+1s', '-10f')
trim_end: New out-point or delta
ripple: Whether to shift subsequent clips
Returns:
Modified clip element
"""
clip = self._require_clip(clip_id)
current_start, current_duration, _ = self._get_clip_times(clip)
original_duration = current_duration
# Handle trim_start
if trim_start:
if trim_start.startswith('+') or trim_start.startswith('-'):
delta = self._parse_time(trim_start[1:])
if trim_start.startswith('-'):
# Extend earlier
new_start = current_start - delta
new_duration = current_duration + delta
else:
# Trim later
new_start = current_start + delta
new_duration = current_duration - delta
else:
new_start = self._parse_time(trim_start)
diff = new_start - current_start
new_duration = current_duration - diff
clip.set('start', new_start.to_fcpxml())
current_start = new_start
current_duration = new_duration
# Handle trim_end
if trim_end:
if trim_end.startswith('+') or trim_end.startswith('-'):
delta = self._parse_time(trim_end[1:])
if trim_end.startswith('-'):
new_duration = current_duration - delta
else:
new_duration = current_duration + delta
else:
end_point = self._parse_time(trim_end)
new_duration = end_point - current_start
current_duration = new_duration
if current_duration <= TimeValue.zero():
raise ValueError(
f"Trim would produce non-positive duration "
f"({current_duration.to_seconds():.3f}s) for clip '{clip_id}'"
)
clip.set('duration', current_duration.to_fcpxml())
# Ripple subsequent clips if needed
if ripple:
duration_change = current_duration - original_duration
if duration_change != TimeValue.zero():
self._ripple_after_clip(clip, duration_change)
return clip
def _ripple_from_index(
self, spine: ET.Element, start_index: int, delta: 'TimeValue'
) -> None:
"""Shift the offset of every spine element from *start_index* onward by *delta*.
Consolidates the ripple loops previously duplicated across
``_ripple_after_clip``, ``delete_clip``, and ``insert_clip``.
Args:
spine: The primary storyline ``<spine>`` element.
start_index: First child index to adjust (inclusive).
delta: Signed time shift (positive = later, negative = earlier).
"""
children = list(spine)
for child in children[start_index:]:
if child.tag in SPINE_ELEMENT_TAGS:
current_offset = self._parse_time(child.get('offset', '0s'))
new_offset = current_offset + delta
child.set('offset', new_offset.to_fcpxml())
def _ripple_after_clip(self, target_clip: ET.Element, delta: TimeValue) -> None:
"""Shift all clips after the given clip by delta."""
spine = self._get_spine()
clip_index = self._find_clip_index(spine, target_clip)
if clip_index is not None:
self._ripple_from_index(spine, clip_index + 1, delta)
# ========================================================================
+232
View File
@@ -0,0 +1,232 @@
"""Verificações estruturais do FCPXML antes de salvar.
Extraído de writer.py — ver fcpxml/writer/__init__.py para o conjunto.
"""
import logging
import xml.etree.ElementTree as ET
from fractions import Fraction
from typing import List, Optional
from ..models import (
_FCPXML_STANDARD_TIMEBASES,
TimeValue,
ValidationIssue,
ValidationIssueType,
)
from .helpers import _ASSET_CLIP_CHILD_ORDER, _CHILD_ORDER_INDEX
# ============================================================================
# PRE-EXPORT DTD VALIDATOR (v0.6.0)
# ============================================================================
_log = logging.getLogger(__name__)
def _check_child_order(root: ET.Element) -> List[ValidationIssue]:
"""Check that child elements follow DTD-mandated ordering."""
issues = []
for parent in root.iter():
if parent.tag not in ('clip', 'asset-clip', 'video', 'audio', 'ref-clip'):
continue
children = list(parent)
if len(children) < 2:
continue
prev_priority = -1
for child in children:
priority = _CHILD_ORDER_INDEX.get(child.tag, len(_ASSET_CLIP_CHILD_ORDER))
if priority < prev_priority:
issues.append(ValidationIssue(
issue_type=ValidationIssueType.ELEMENT_ORDER,
severity="warning",
message=(
f"<{child.tag}> appears after a higher-priority sibling "
f"in <{parent.tag}> '{parent.get('name', '')}'."
),
clip_name=parent.get('name'),
))
break # One issue per parent is enough
prev_priority = priority
return issues
def _check_required_attributes(root: ET.Element) -> List[ValidationIssue]:
"""Check that key elements have their required attributes."""
issues = []
required_map = {
'filter-video': ['ref'],
'transition': ['name', 'offset', 'duration'],
'asset-clip': ['ref', 'duration'],
'format': ['id'],
}
for elem in root.iter():
attrs = required_map.get(elem.tag)
if not attrs:
continue
for attr in attrs:
if not elem.get(attr):
issues.append(ValidationIssue(
issue_type=ValidationIssueType.MISSING_ATTRIBUTE,
severity="error",
message=f"<{elem.tag}> missing required attribute '{attr}'.",
clip_name=elem.get('name'),
))
return issues
def _check_timebases(root: ET.Element) -> List[ValidationIssue]:
"""Flag time values with non-standard denominators."""
issues = []
time_attrs = ('offset', 'start', 'duration')
seen: set = set()
for elem in root.iter():
for attr in time_attrs:
val = elem.get(attr)
if val and val.endswith('s') and '/' in val:
try:
tv = TimeValue.from_timecode(val)
denom = tv.simplify().denominator
if denom not in _FCPXML_STANDARD_TIMEBASES:
key = (elem.tag, attr, val)
if key not in seen:
seen.add(key)
issues.append(ValidationIssue(
issue_type=ValidationIssueType.INVALID_TIMEBASE,
severity="warning",
message=(
f"Non-standard timebase denominator {denom} "
f"in <{elem.tag}> {attr}=\"{val}\"."
),
clip_name=elem.get('name'),
))
except (ValueError, ZeroDivisionError):
pass
return issues
def _document_frame_duration(root: ET.Element) -> Optional[Fraction]:
"""The sequence's exact ``frameDuration`` as a fraction, if declared.
Read from the format the ``<sequence>`` references (falling back to the
first declared format), so the value is the document's own timebase
rather than an assumed rate.
"""
formats = {f.get('id'): f for f in root.findall('.//format') if f.get('id')}
sequence = root.find('.//sequence')
fmt = formats.get(sequence.get('format')) if sequence is not None else None
if fmt is None:
fmt = next(iter(formats.values()), None)
if fmt is None:
return None
raw = fmt.get('frameDuration', '')
if not (raw.endswith('s') and '/' in raw):
return None
numerator, denominator = raw[:-1].split('/', 1)
try:
value = Fraction(int(numerator), int(denominator))
except (ValueError, ZeroDivisionError):
return None
return value if value > 0 else None
def _check_frame_alignment(root: ET.Element, fps: float = 24.0) -> List[ValidationIssue]:
"""Check that durations are integer multiples of the frame duration.
Uses the document's exact ``frameDuration`` fraction and rational
arithmetic. Comparing against an integer fps instead would flag every
NTSC project as broken: at 1001/24000s (23.976fps) a perfectly aligned
duration is not an integer number of "24fps" frames, so whole timelines
would be reported misaligned when nothing is wrong.
"""
issues = []
frame_duration = _document_frame_duration(root)
label = f"{1 / float(frame_duration):.3f}".rstrip('0').rstrip('.') if frame_duration else str(fps)
for elem in root.iter():
dur_str = elem.get('duration')
if not dur_str or not dur_str.endswith('s'):
continue
if elem.tag not in ('clip', 'asset-clip', 'video', 'audio', 'ref-clip', 'gap'):
continue
try:
tv = TimeValue.from_timecode(dur_str)
if frame_duration is not None:
frames = Fraction(tv.numerator, tv.denominator) / frame_duration
aligned = frames.denominator == 1
else:
approx = tv.to_seconds() * fps
aligned = abs(approx - round(approx)) <= 0.01
if not aligned:
issues.append(ValidationIssue(
issue_type=ValidationIssueType.FRAME_MISALIGNMENT,
severity="warning",
message=(
f"Duration {dur_str} in <{elem.tag}> "
f"'{elem.get('name', '')}' is not frame-aligned at {label}fps."
),
clip_name=elem.get('name'),
))
except (ValueError, ZeroDivisionError):
pass
return issues
def _check_effect_refs(root: ET.Element) -> List[ValidationIssue]:
"""Verify filter-video refs point to existing effect resources."""
issues = []
resource_ids = set()
for res in root.iter():
rid = res.get('id')
if rid and res.tag in ('effect', 'format', 'asset', 'media'):
resource_ids.add(rid)
for fv in root.iter('filter-video'):
ref = fv.get('ref')
if ref and ref not in resource_ids:
issues.append(ValidationIssue(
issue_type=ValidationIssueType.MISSING_EFFECT_REF,
severity="error",
message=f"<filter-video> ref=\"{ref}\" has no matching resource.",
))
return issues
def _check_asset_sources(root: ET.Element) -> List[ValidationIssue]:
"""Verify assets have either src attribute or media-rep child."""
issues = []
for asset in root.iter('asset'):
src = asset.get('src', '')
media_rep = asset.find('media-rep')
if not src and media_rep is None:
issues.append(ValidationIssue(
issue_type=ValidationIssueType.MISSING_MEDIA_REP,
severity="warning",
message=(
f"<asset id=\"{asset.get('id', '?')}\" "
f"name=\"{asset.get('name', '')}\"> "
f"has no src attribute and no <media-rep> child."
),
clip_name=asset.get('name'),
))
return issues
def validate_fcpxml(root: ET.Element, fps: float = 24.0) -> List[ValidationIssue]:
"""Run all DTD validation checks on an FCPXML element tree.
Args:
root: The <fcpxml> root Element to validate.
fps: Frame rate for alignment checks (default 24).
Returns:
List of ValidationIssue objects. Empty list = clean.
"""
issues: List[ValidationIssue] = []
issues.extend(_check_child_order(root))
issues.extend(_check_required_attributes(root))
issues.extend(_check_timebases(root))
issues.extend(_check_frame_alignment(root, fps))
issues.extend(_check_effect_refs(root))
issues.extend(_check_asset_sources(root))
return issues
+4 -1
View File
@@ -41,6 +41,9 @@ intelligence = [
transcribe = [ transcribe = [
"faster-whisper>=1.0.0", "faster-whisper>=1.0.0",
] ]
align = [
"whisperx>=3.0.0",
]
diarization = [ diarization = [
"pyannote.audio>=3.1", "pyannote.audio>=3.1",
] ]
@@ -73,7 +76,7 @@ target-version = ['py310']
[tool.ruff] [tool.ruff]
line-length = 100 line-length = 100
exclude = ["docs/", "WHISPERX/"] exclude = ["docs/"]
[tool.ruff.lint] [tool.ruff.lint]
select = ["E", "F", "I", "N", "W"] select = ["E", "F", "I", "N", "W"]
+3
View File
@@ -154,6 +154,7 @@ from server_tools.roles import (
) )
from server_tools.subtitles import ( from server_tools.subtitles import (
handle_generate_dynamic_subtitles, handle_generate_dynamic_subtitles,
handle_generate_plain_subtitles,
handle_validate_subtitle_layout, handle_validate_subtitle_layout,
) )
from server_tools.timeline import ( from server_tools.timeline import (
@@ -178,6 +179,7 @@ from server_tools.voice import (
handle_analyze_voice_features, handle_analyze_voice_features,
handle_apply_voice_actions, handle_apply_voice_actions,
handle_build_voice_timeline, handle_build_voice_timeline,
handle_generate_voice_script,
handle_diarize_media, handle_diarize_media,
handle_get_voice_analysis_config, handle_get_voice_analysis_config,
handle_refine_voice_timeline, handle_refine_voice_timeline,
@@ -320,6 +322,7 @@ __all__ = [
"handle_save_voice_analysis_config", "handle_save_voice_analysis_config",
"handle_validate_subtitle_layout", "handle_validate_subtitle_layout",
"handle_generate_dynamic_subtitles", "handle_generate_dynamic_subtitles",
"handle_generate_plain_subtitles",
"handle_push_to_fcp", "handle_push_to_fcp",
"handle_list_fcp_libraries", "handle_list_fcp_libraries",
] ]
-829
View File
@@ -1,829 +0,0 @@
"""Shared internal helpers used by tool handlers across categories.
Extracted from server.py — validation, formatting, and small parsing utilities
that more than one server_tools/*.py module needs.
"""
from __future__ import annotations
import json
import os
import re
from pathlib import Path
from typing import Any, Sequence
from mcp.types import TextContent
from fcpxml.media_intel import media_src_to_path
from fcpxml.models import (
DuplicateGroup,
FlashFrame,
FlashFrameSeverity,
GapInfo,
Timecode,
TimeValue,
)
from fcpxml.parser import FCPXMLParser
from fcpxml.rough_cut import RoughCutGenerator
from fcpxml.transcribe import invert_ranges, merge_ranges, transcribe
from fcpxml.writer import FCPXMLModifier
PROJECTS_DIR = os.environ.get("FCP_PROJECTS_DIR", os.path.expanduser("~/Movies"))
_SANDBOX_ENABLED = "FCP_PROJECTS_DIR" in os.environ
MAX_FILE_SIZE = 100 * 1024 * 1024
MAX_MEDIA_FILE_SIZE = 32 * 1024 * 1024 * 1024
_MAX_JSON_DEPTH = 50
def _check_json_depth(obj: object, _depth: int = 0) -> None:
"""Reject JSON structures nested beyond _MAX_JSON_DEPTH.
Prevents denial-of-service via deeply nested objects that exhaust the
call stack or memory during downstream processing. Called after
json.load() since Python's json module has no built-in depth limit.
"""
if _depth > _MAX_JSON_DEPTH:
raise ValueError(
f"JSON nesting depth exceeds {_MAX_JSON_DEPTH} — "
"file may be malformed or adversarial"
)
if isinstance(obj, dict):
for v in obj.values():
_check_json_depth(v, _depth + 1)
elif isinstance(obj, list):
for item in obj:
_check_json_depth(item, _depth + 1)
def _validate_filepath(
filepath: str,
allowed_extensions: tuple[str, ...] | None = None,
max_size: int = MAX_FILE_SIZE,
) -> str:
"""Validate a user-provided file path against traversal and size attacks.
Resolves symlinks, blocks null bytes, enforces extension whitelist, and
checks file size before any parsing takes place.
``max_size`` defaults to the document limit; callers handling source
media pass ``MAX_MEDIA_FILE_SIZE``, since media is streamed rather than
parsed into memory (see the constant for why).
Raises:
ValueError: For invalid paths (null bytes, bad extensions, oversized).
FileNotFoundError: When the resolved path does not exist.
"""
if '\x00' in filepath:
raise ValueError("Invalid file path: null byte detected")
resolved = Path(filepath).resolve()
if not resolved.exists():
raise FileNotFoundError(f"File not found: {filepath}")
# .fcpxmld bundles are directories (a package wrapping Info.fcpxml plus
# sidecar data files for object tracking / Cinematic mode). The size
# check applies to the inner Info.fcpxml, which is what gets parsed.
if resolved.is_dir():
if resolved.suffix.lower() != '.fcpxmld':
raise ValueError(f"Not a regular file: {filepath}")
inner = resolved / 'Info.fcpxml'
if not inner.is_file():
raise ValueError(f"Invalid bundle (no Info.fcpxml): {filepath}")
size_target = inner
elif not resolved.is_file():
raise ValueError(f"Not a regular file: {filepath}")
else:
size_target = resolved
if allowed_extensions and resolved.suffix.lower() not in allowed_extensions:
raise ValueError(
f"Invalid file type '{resolved.suffix}'. "
f"Allowed: {', '.join(allowed_extensions)}"
)
if size_target.stat().st_size > max_size:
size_mb = size_target.stat().st_size / (1024 * 1024)
raise ValueError(f"File too large ({size_mb:.1f} MB). Maximum: {max_size // (1024 * 1024)} MB")
return str(resolved)
def _validate_output_path(output_path: str, *, anchor_dir: str | None = None) -> str:
"""Validate an output path with optional sandbox enforcement.
Resolves traversal, blocks null bytes, ensures parent exists, and — when
*anchor_dir* is provided — verifies the resolved output lives under that
directory. This prevents LLM-generated tool calls from writing to
arbitrary filesystem locations (e.g. ``/etc/cron.d/backdoor``).
Args:
output_path: The raw output path to validate.
anchor_dir: If set, the resolved output must be a child of this
directory. Typically the parent directory of the input file so
outputs stay co-located with their sources.
Raises:
ValueError: For null bytes, missing parent, or sandbox escape.
"""
if '\x00' in output_path:
raise ValueError("Invalid output path: null byte detected")
resolved = Path(output_path).resolve()
if not resolved.parent.exists():
raise ValueError(f"Output directory does not exist: {resolved.parent}")
if anchor_dir is not None:
anchor = Path(anchor_dir).resolve()
try:
resolved.relative_to(anchor)
except ValueError:
raise ValueError(
f"Output path escapes allowed directory: "
f"{resolved} is not under {anchor}"
)
return str(resolved)
def _validate_directory(directory: str, *, allowed_root: str | None = None) -> str:
"""Validate a user-provided directory path against traversal and injection.
Resolves symlinks, blocks null bytes, and verifies the path is a real
directory. When *allowed_root* is given, the resolved path must be a
descendant of (or equal to) that root — preventing filesystem enumeration
beyond the project workspace.
Raises:
ValueError: For invalid paths (null bytes, not a directory, sandbox escape).
"""
if '\x00' in directory:
raise ValueError("Invalid directory path: null byte detected")
resolved = Path(directory).resolve()
if not resolved.is_dir():
raise ValueError(f"Not a valid directory: {directory}")
if allowed_root is not None:
root = Path(allowed_root).resolve()
try:
resolved.relative_to(root)
except ValueError:
raise ValueError(
f"Directory escapes allowed root: "
f"{resolved} is not under {root}"
)
return str(resolved)
def find_fcpxml_files(directory: str) -> list[str]:
"""Find all FCPXML files in a directory."""
path = Path(directory)
files = list(str(f) for f in path.rglob("*.fcpxml"))
files.extend(str(f) for f in path.rglob("*.fcpxmld"))
return sorted(files)
def format_timecode(tc) -> str:
"""Format a Timecode object to SMPTE string."""
return tc.to_smpte() if tc else "00:00:00:00"
def format_duration(seconds: float) -> str:
"""Format seconds into human-readable duration."""
if seconds < 1:
return f"{seconds*1000:.0f}ms"
elif seconds < 60:
return f"{seconds:.2f}s"
return f"{int(seconds // 60)}m {seconds % 60:.1f}s"
def _format_clip_table(clips: list, header: str) -> str:
"""Render a list of clips as a markdown table with timecodes and durations.
Shared by handlers that filter clips by duration threshold
(find_short_cuts, find_long_clips).
"""
result = f"{header}\n\n| Name | TC | Duration |\n|------|----|---------|\n"
result += "\n".join(
f"| {c.name} | {format_timecode(c.start)} | {format_duration(c.duration_seconds)} |"
for c in clips
)
return result
def _markdown_table(headers: list[str], rows: list[list[str]]) -> str:
"""Build a markdown table from headers and rows.
Returns header row, separator row, and data rows as a single string.
Callers avoid repeating the ``| H1 | H2 |\\n|---|---|`` boilerplate
that appears in 15+ handlers.
"""
header_line = "| " + " | ".join(headers) + " |"
sep_line = "|" + "|".join("------" for _ in headers) + "|"
data_lines = "\n".join(
"| " + " | ".join(str(c) for c in row) + " |" for row in rows
)
return f"{header_line}\n{sep_line}\n{data_lines}"
def _format_batch_result(
title: str,
summary: dict[str, str],
headers: list[str],
rows: list[list[str]],
output_path: str,
) -> str:
"""Build a standard batch-operation result with summary, table, and save footer.
Used by batch fix handlers (flash frames, rapid trim, fill gaps) that all
share the same markdown structure: ``# Title → ## Summary → ## Details table
→ Saved to`` footer.
"""
summary_lines = "\n".join(f"- **{k}**: {v}" for k, v in summary.items())
table = _markdown_table(headers, rows)
return (
f"# {title}\n\n"
f"## Summary\n{summary_lines}\n\n"
f"## Details\n{table}\n\n"
f"Saved to: `{output_path}`"
)
def _fmt_suggestions(suggestions: list[str]) -> str:
"""Format pacing suggestions as markdown list (Python 3.10 compatible)."""
if not suggestions:
return "- Pacing looks good!"
nl = "\n"
return nl.join(f"- {s}" for s in suggestions)
def generate_output_path(input_path: str, suffix: str = "_modified") -> str:
"""Generate output path from input path.
The suffix is sanitized to prevent path-component injection — only
alphanumeric, hyphen, underscore, and dot characters survive.
"""
# Strip anything that could inject path separators or traversal sequences
clean_suffix = re.sub(r'[^a-zA-Z0-9._-]', '', suffix)
if not clean_suffix:
clean_suffix = "_modified"
p = Path(input_path)
return str(p.parent / f"{p.stem}{clean_suffix}{p.suffix}")
def _parse_project(filepath: str):
"""Parse an FCPXML file and return the project with its primary timeline."""
filepath = _validate_filepath(filepath, ('.fcpxml', '.fcpxmld'))
project = FCPXMLParser().parse_file(filepath)
if not project.timelines:
return None, None
return project, project.primary_timeline
def _text_result(text: str) -> list[TextContent]:
"""Wrap a string in the MCP TextContent list that every tool handler returns."""
return [TextContent(type="text", text=text)]
def _no_timeline():
"""Standard response when no timelines are found."""
return _text_result("No timelines found")
def _require_timeline(filepath: str):
"""Parse FCPXML and return (project, timeline), raising if no timeline exists.
Centralises the repeated _parse_project + _no_timeline guard that
appears in every read-only timeline handler. Returns a tuple so
callers can destructure directly::
project, tl = _require_timeline(arguments["filepath"])
"""
project, tl = _parse_project(filepath)
if not tl:
raise _NoTimelineError()
return project, tl
class _NoTimelineError(Exception):
"""Sentinel raised by _require_timeline when no timelines exist."""
def _resolve_io_paths(
arguments: dict,
suffix: str = "_modified",
) -> tuple[str, str]:
"""Validate input filepath and resolve the output path.
Shared foundation for every handler that reads an FCPXML and writes
a derived file. Validates the input, falls back to a suffixed
output name when ``output_path`` is not supplied, and sandbox-checks
the result.
Args:
arguments: Tool arguments dict (must contain ``filepath``; may
contain ``output_path``).
suffix: Default output filename suffix when ``output_path`` is
not provided (e.g. ``"_modified"``, ``"_beats"``).
Returns:
``(filepath, output_path)`` tuple with both paths validated.
"""
filepath = _validate_filepath(arguments["filepath"], ('.fcpxml', '.fcpxmld'))
# Anchor write operations to the input file's directory so LLM-generated
# tool calls cannot write to arbitrary filesystem locations (e.g.
# /etc/cron.d/backdoor). When the explicit sandbox is off, the anchor
# still prevents writes outside the source directory tree.
# `output_dir` is where the caller wants the file written, not merely a
# sandbox boundary: the app's "Pasta do projeto" promises that everything
# generated lands there. Deriving the name from the input but keeping the
# input's directory made every cross-directory call fail its own anchor
# check ("output path escapes allowed directory"), so the setting silently
# only worked when it pointed at the directory the file was already going
# to. An explicit `output_path` still wins, and still has to sit inside
# the anchor.
output_dir = arguments.get("output_dir")
if output_dir:
anchor = _validate_directory(str(output_dir))
default_output = str(Path(anchor) / Path(generate_output_path(filepath, suffix)).name)
else:
anchor = str(Path(filepath).resolve().parent)
default_output = generate_output_path(filepath, suffix)
output_path = _validate_output_path(
arguments.get("output_path") or default_output,
anchor_dir=anchor,
)
return filepath, output_path
def _setup_modifier(
arguments: dict,
suffix: str = "_modified",
) -> tuple[str, str, "FCPXMLModifier"]:
"""Common setup for write handlers: validate paths and create modifier.
Consolidates the repeated validate-filepath → resolve-output-path →
create-modifier boilerplate shared by 18+ write handlers.
Args:
arguments: Tool arguments dict (must contain ``filepath``; may
contain ``output_path``).
suffix: Default output filename suffix when ``output_path`` is
not provided (e.g. ``"_modified"``, ``"_flash_fixed"``).
Returns:
``(filepath, output_path, modifier)`` tuple ready for the
handler's domain-specific operation.
"""
filepath, output_path = _resolve_io_paths(arguments, suffix)
modifier = FCPXMLModifier(filepath)
return filepath, output_path, modifier
def _setup_generator(
arguments: dict,
suffix: str = "_roughcut",
) -> tuple[str, str, "RoughCutGenerator"]:
"""Common setup for generation handlers: validate paths and create generator.
Args:
arguments: Tool arguments dict (must contain ``filepath`` and
``output_path``).
suffix: Default output filename suffix.
Returns:
``(filepath, output_path, generator)`` tuple.
"""
filepath, output_path = _resolve_io_paths(arguments, suffix)
generator = RoughCutGenerator(filepath)
return filepath, output_path, generator
def _parse_timestamp_parts(
parts: list[str], *, frame_rate: float = 24.0
) -> float | None:
"""Convert colon-separated timestamp parts to total seconds.
Handles 2-part (M:SS), 3-part (H:MM:SS / HH:MM:SS.ms), and
4-part (HH:MM:SS:FF SMPTE) formats. Returns ``None`` when the
part count is unrecognised so callers can skip.
Args:
parts: Colon-split timestamp components.
frame_rate: FPS used to convert the frame component of SMPTE
timecodes into fractional seconds (default 24.0).
"""
if len(parts) == 2:
return int(parts[0]) * 60 + float(parts[1])
elif len(parts) == 3:
return int(parts[0]) * 3600 + int(parts[1]) * 60 + float(parts[2])
elif len(parts) == 4:
# SMPTE: HH:MM:SS:FF — convert frames to fractional seconds
base = int(parts[0]) * 3600 + int(parts[1]) * 60 + float(parts[2])
frames = int(parts[3])
return base + (frames / frame_rate) if frame_rate > 0 else base
return None
def _raw_markers_to_batch(
raw_markers: list[dict],
marker_type: str = "chapter",
max_label: int | None = None,
) -> list[dict]:
"""Convert raw {seconds, text} marker dicts to batch_add_markers format.
Shared by import_srt_markers and import_transcript_markers.
"""
batch = []
for m in raw_markers:
label = m["text"]
if max_label and len(label) > max_label:
label = label[:max_label]
batch.append({
"timecode": f"{m['seconds']}s",
"name": label,
"marker_type": marker_type.upper(),
})
return batch
def _extract_subtitle_blocks(text: str, *, strip_vtt_tags: bool = False) -> list[dict]:
"""Extract timestamp/text pairs from subtitle cue blocks (SRT or VTT).
Both SRT and VTT use the same ``start --> end`` cue syntax with
text lines underneath; only header stripping and tag cleaning differ.
"""
markers = []
blocks = re.split(r'\n\s*\n', text.strip())
for block in blocks:
lines = block.strip().split('\n')
if len(lines) < 2:
continue
ts_line = None
text_lines = []
for line in lines:
if '-->' in line:
ts_line = line
elif ts_line is not None:
if strip_vtt_tags:
line = re.sub(r'<[^>]+>', '', line)
cleaned = line.strip()
if cleaned:
text_lines.append(cleaned)
if not ts_line or not text_lines:
continue
start_str = ts_line.split('-->')[0].strip().replace(',', '.')
seconds = _parse_timestamp_parts(start_str.split(':'))
if seconds is not None:
markers.append({'seconds': seconds, 'text': ' '.join(text_lines)})
return markers
def parse_srt(text: str) -> list[dict]:
"""Parse SRT subtitle format into timestamp/text pairs."""
return _extract_subtitle_blocks(text)
def parse_vtt(text: str) -> list[dict]:
"""Parse WebVTT subtitle format into timestamp/text pairs."""
text = re.sub(r'^WEBVTT.*?\n', '', text, flags=re.MULTILINE)
text = re.sub(r'NOTE\n.*?\n\n', '', text, flags=re.DOTALL)
return _extract_subtitle_blocks(text, strip_vtt_tags=True)
def parse_transcript_timestamps(text: str) -> list[dict]:
"""Parse timestamped text (YouTube description format) into markers.
Supports formats like:
0:00 Introduction
00:01:30 Main Topic
1:05:30 Conclusion
00:00:00:00 SMPTE timecode
"""
markers = []
for line in text.strip().split('\n'):
line = line.strip()
if not line:
continue
match = re.match(r'^(\d{1,2}:\d{2}(?::\d{2}){0,2})\s+(.+)$', line)
if match:
seconds = _parse_timestamp_parts(match.group(1).split(':'))
if seconds is not None:
markers.append({'seconds': seconds, 'text': match.group(2).strip()})
return markers
def _detect_flash_frames(
tl: Any, *, critical_threshold: int = 2, warning_threshold: int = 6,
) -> list:
"""Find clips shorter than *warning_threshold* frames.
Returns a list of ``FlashFrame`` objects sorted by severity. Shared by
``handle_detect_flash_frames`` and ``handle_validate_timeline`` so the
detection logic lives in exactly one place.
"""
fps = tl.frame_rate
flash_frames: list[FlashFrame] = []
for clip in tl.clips:
duration_frames = int(clip.duration_seconds * fps)
if duration_frames < warning_threshold:
severity = (
FlashFrameSeverity.CRITICAL
if duration_frames < critical_threshold
else FlashFrameSeverity.WARNING
)
flash_frames.append(FlashFrame(
clip_name=clip.name, clip_id=clip.name,
start=clip.start, duration_frames=duration_frames,
duration_seconds=clip.duration_seconds, severity=severity,
))
return flash_frames
def _detect_gaps(tl: Any, *, min_gap_frames: int = 1) -> list:
"""Find inter-clip gaps of at least *min_gap_frames* length.
Returns a list of ``GapInfo`` objects. Shared by ``handle_detect_gaps``
and ``handle_validate_timeline``.
"""
fps = tl.frame_rate
min_gap_seconds = min_gap_frames / fps
gaps: list[GapInfo] = []
sorted_clips = sorted(tl.clips, key=lambda c: c.start.seconds)
for i in range(len(sorted_clips) - 1):
current_end = sorted_clips[i].end.seconds
next_start = sorted_clips[i + 1].start.seconds
gap_duration = next_start - current_end
if gap_duration >= min_gap_seconds:
gaps.append(GapInfo(
start=Timecode(frames=int(current_end * fps), frame_rate=fps),
duration_frames=int(gap_duration * fps),
duration_seconds=gap_duration,
previous_clip=sorted_clips[i].name,
next_clip=sorted_clips[i + 1].name,
))
return gaps
def _detect_duplicate_groups(tl: Any, *, mode: str = "same_source") -> list:
"""Group clips that share a source media reference.
Returns a list of ``DuplicateGroup`` objects. Shared by
``handle_detect_duplicates`` and ``handle_validate_timeline``.
"""
source_groups: dict[str, list[dict]] = {}
for clip in tl.clips:
source_key = clip.media_path or clip.name
if source_key not in source_groups:
source_groups[source_key] = []
source_groups[source_key].append({
'name': clip.name,
'start': clip.start.seconds,
'duration': clip.duration_seconds,
'source_start': clip.source_start.seconds if clip.source_start else 0,
'source_duration': clip.duration_seconds,
'timecode': format_timecode(clip.start),
})
duplicates: list[DuplicateGroup] = []
for source_key, clips in source_groups.items():
if len(clips) <= 1:
continue
group = DuplicateGroup(
source_ref=source_key,
source_name=source_key.split('/')[-1] if '/' in source_key else source_key,
clips=clips,
)
if mode == "same_source":
duplicates.append(group)
elif mode == "overlapping_ranges" and group.has_overlapping_ranges:
duplicates.append(group)
elif mode == "identical":
seen_ranges: set[tuple] = set()
identical_clips = []
for c in clips:
range_key = (c['source_start'], c['source_duration'])
if range_key in seen_ranges:
identical_clips.append(c)
seen_ranges.add(range_key)
if identical_clips:
group.clips = identical_clips
duplicates.append(group)
return duplicates
AUDIO_MEDIA_EXTENSIONS = (
'.wav', '.aif', '.aiff', '.mp3', '.m4a', '.aac', '.flac', '.mov', '.mp4',
)
_DIARIZATION_INSTALL_HINT = (
"\n\nInstall the optional diarization extra:\n\n"
" pip install 'fcp-mcp-server[diarization]'\n\n"
"and set a HuggingFace token with access to "
"pyannote/speaker-diarization-3.1 (pass hf_token= or persist one via "
"save_hf_token)."
)
_FEATURES_INSTALL_HINT = (
"\n\nInstall the optional media-intelligence extra:\n\n"
" pip install 'fcp-mcp-server[intelligence]'"
)
def _voice_analysis_config_text(config: dict) -> str:
w = config["emphasis_weights"]
text = "# Voice Analysis Settings\n\n"
text += _markdown_table(
["Setting", "Value"],
[
["Energy threshold", f"{config['energy_threshold']:.2f}"],
["Peak selection", f"top {config['peak_percentile']:.1%} of words"],
["Emphasis floor", f"{config['emphasis_floor']:.2f}"],
["Emotion detection", "on" if config["emotion_enabled"] else "off"],
["Emotion sensitivity", f"{config['emotion_sensitivity']:.2f}"],
],
) + "\n\n## Emphasis Weights\n"
text += _markdown_table(
["Factor", "Weight"],
[[k.replace("_", " ").title(), f"{v:.2f}"] for k, v in w.items()],
)
return text
def _apply_placed_action(modifier, clip_el, action, clip_start: float) -> str:
"""Apply one non-cut action to the clip that hosts it.
``clip_start`` is where that clip begins on the timeline; the writer
wants times relative to the clip's own head, so the rebase happens here
— the single place that knows about the conversion. The clip *element*
is passed through rather than its name: after a cut the pieces share a
name, and a name lookup would land every edit on the first piece.
"""
rel_start = action.start - clip_start
rel_end = action.end - clip_start
if action.kind == "zoom":
# Only forward an explicit ease — otherwise add_zoom's own default
# (a fast ramp in, instant snap back out) is what should apply.
zoom_args = {}
if action.params.get("ease") is not None:
zoom_args["ease"] = float(action.params["ease"])
if action.params.get("ease_out") is not None:
zoom_args["ease_out"] = float(action.params["ease_out"])
modifier.add_zoom(
clip_id=clip_el,
start=rel_start,
end=rel_end,
scale=float(action.params.get("scale", 1.3)),
**zoom_args,
)
return f"zoom {action.params.get('scale', 1.3):.2f}x"
if action.kind == "text":
modifier.add_text_title(
clip_el,
action.params["content"],
offset=modifier.snap_seconds_to_frame(rel_start).to_fcpxml(),
duration=modifier.snap_seconds_to_frame(action.duration).to_fcpxml(),
)
return f"text \"{action.params['content'][:24]}\""
# marker
modifier.add_marker(
clip_id=clip_el,
timecode=modifier.snap_seconds_to_frame(rel_start).to_fcpxml(),
name=action.params.get("content") or action.reason or "Voice action",
note=action.reason or None,
)
return "marker"
def _speaker_table(profiles: Sequence[dict]) -> str:
"""Who was detected, ordered by how much of the runtime each holds."""
return _markdown_table(
["ID", "Name", "Share", "Speaking", "Lines", "Avg line"],
[
[
p["id"],
p.get("name", ""),
f"{p['share']:.0%}",
format_duration(p["speaking_seconds"]),
str(p["segment_count"]),
f"{p['avg_segment']:.1f}s",
]
for p in profiles
],
)
TRANSCRIBE_MAX_MEDIA = 10
_TRANSCRIBE_INSTALL_HINT = (
"\n\nInstall the optional transcription extra:\n\n"
" pip install 'fcp-mcp-server[transcribe]'\n\n"
"or run via uvx:\n\n"
" uvx --from \"fcp-mcp-server[transcribe]\" fcp-mcp-server"
)
def _transcript_json_path(media_path: str, output_dir: str | None = None) -> Path:
"""Where the ``_transcript.json`` for ``media_path`` lives.
When ``output_dir`` (the user-selected project folder) is set, the
transcript is saved/read there instead of next to the source media.
"""
p = Path(media_path)
if output_dir:
directory = Path(output_dir).expanduser()
directory.mkdir(parents=True, exist_ok=True)
return directory / f"{p.stem}_transcript.json"
return p.with_name(p.stem + "_transcript.json")
def _load_or_transcribe(
media_path: str, model: str, language: str | None, output_dir: str | None = None
) -> tuple[dict | None, str]:
"""Load a cached ``_transcript.json`` for a media file, else transcribe and cache it.
Returns ``(transcript, "")`` or ``(None, reason)``. The cache makes
transcription a one-time cost per media file across all transcript tools.
"""
json_path = _transcript_json_path(media_path, output_dir)
if json_path.is_file():
try:
with open(json_path) as f:
data = json.load(f)
if isinstance(data, dict) and isinstance(data.get("words"), list):
return data, ""
except (OSError, json.JSONDecodeError, UnicodeDecodeError):
pass # unreadable cache falls through to re-transcribe
result = transcribe(media_path, model_size=model, language=language)
if result is None:
return None, "untranscribable (faster-whisper not installed or media unreadable)"
anchor = str(Path(output_dir).expanduser()) if output_dir else str(Path(media_path).parent)
out_path = _validate_output_path(str(json_path), anchor_dir=anchor)
with open(out_path, "w") as f:
json.dump({"source": Path(media_path).name, **result}, f, indent=2)
return result, ""
def _cut_transcript_spans(modifier, clip_filter, model, language, padding, spans_fn, keep_only=False, output_dir=None):
"""Shared cut engine for transcript-driven editing.
``spans_fn(words) -> [(start, end), ...]`` in source seconds. Spans are
padded, clamped to each clip's used source window, optionally inverted
(keep_only), snapped to the frame grid, and cut with ripple.
"""
to_frame = modifier.snap_seconds_to_frame
cache: dict[str, tuple] = {}
cuts_made: list[tuple[str, int, float]] = []
skipped: list[tuple[str, str]] = []
spine_clips = [el for _, el in modifier._iter_spine_clips()]
for el in spine_clips:
name = el.get("name", "")
if clip_filter and name != clip_filter:
continue
src = modifier.resources.get(el.get("ref", ""), {}).get("src", "")
media_path = media_src_to_path(src)
if not media_path or not Path(media_path).is_file():
skipped.append((name, "media file missing"))
continue
if media_path not in cache:
if len(cache) >= TRANSCRIBE_MAX_MEDIA:
skipped.append((name, f"transcription cap reached ({TRANSCRIBE_MAX_MEDIA} media files)"))
continue
cache[media_path] = _load_or_transcribe(media_path, model, language, output_dir)
data, reason = cache[media_path]
if data is None:
skipped.append((name, reason))
continue
clip_source_start = modifier.source_file_start(el).to_seconds()
clip_duration = modifier._parse_time(el.get("duration", "0s")).to_seconds()
window_start = clip_source_start
window_end = clip_source_start + clip_duration
spans = spans_fn(data.get("words", []))
padded = merge_ranges([(s - padding, e + padding) for s, e in spans])
clamped = [
(max(s, window_start), min(e, window_end))
for s, e in padded
if min(e, window_end) > max(s, window_start)
]
if keep_only:
if not clamped:
# Never delete a whole clip just because nothing matched in it.
skipped.append((name, "no phrase matches — left untouched (keep_only)"))
continue
cut_source = invert_ranges(clamped, window_start, window_end)
else:
cut_source = clamped
cut_ranges = [
(to_frame(s - clip_source_start), to_frame(e - clip_source_start))
for s, e in cut_source
]
cut_ranges = [(a, b) for a, b in cut_ranges if b > a]
if not cut_ranges:
continue
removed = modifier.cut_clip_ranges(el, cut_ranges)
if removed > TimeValue.zero():
cuts_made.append((name, len(cut_ranges), removed.to_seconds()))
return cuts_made, skipped
def _transcript_cut_report(title, summary_lines, cuts_made, skipped, output_path, footer):
if not cuts_made:
text = f"# {title}\n\nNo cuts to make — file unchanged (nothing saved)."
if skipped:
text += "\n\n## Skipped Clips\n" + _markdown_table(
["Clip", "Reason"], [[name, reason] for name, reason in skipped]
)
if any("faster-whisper" in reason for _, reason in skipped):
text += _TRANSCRIBE_INSTALL_HINT
return _text_result(text)
total_removed = sum(seconds for _, _, seconds in cuts_made)
result = f"# {title}\n\n## Summary\n"
result += "\n".join(summary_lines) + "\n"
result += f"- **Clips Cut**: {len(cuts_made)}\n- **Total Removed**: {format_duration(total_removed)}\n"
result += "\n## Cuts\n"
result += _markdown_table(
["Clip", "Ranges Cut", "Removed"],
[[name, str(count), f"{seconds:.2f}s"] for name, count, seconds in cuts_made],
) + "\n"
if skipped:
result += "\n## Skipped Clips\n" + _markdown_table(
["Clip", "Reason"], [[name, reason] for name, reason in skipped]
) + "\n"
result += f"\nSaved to: {output_path}\n\n{footer}"
return _text_result(result)

Some files were not shown because too many files have changed in this diff Show More