Compare commits
19
Commits
main
...
e9a17c1b62
@@ -28,6 +28,22 @@ minutos.
|
||||
`apply_voice_actions` aplica direto, mas é para **teste**. O produto do seu
|
||||
trabalho é a lista de decisões.
|
||||
|
||||
**Caminho automatizado (sem wizard, sem copiar-e-colar):** a tool
|
||||
`generate_voice_script` (MCP) / comando `generate_voice_script` (ponte do app)
|
||||
corre o fluxo fechado: transcreve → `build_voice_timeline` → entrega a timeline
|
||||
a um **modelo local Ollama (Gemma 3 / Llama)** que age exatamente como este
|
||||
skill descreve (separa roteiro de bastidor, escolhe tomadas, decide zoom/corte)
|
||||
→ devolve o roteiro legível **e** o JSON de ações, e opcionalmente aplica no
|
||||
FCPXML. O cliente fica em `code/fcpxml/llm_local.py`; o prompt que embute este
|
||||
contrato está em `_SYSTEM_PROMPT`. Use essa tool quando o usuário pedir para
|
||||
"rodar tudo internamente" ou "gerar o roteiro por IA local".
|
||||
|
||||
**Para onde ela vai (modo manual):** o usuário cola o seu JSON no app, e ele
|
||||
abre na etapa 5 do Assistente — uma tela onde cada frase do roteiro aparece com
|
||||
a sua decisão já marcada, para ser revisada antes de gerar. Você é o **ponto de
|
||||
partida** da edição, não a palavra final; escreva decisões defensáveis e motivos
|
||||
legíveis. Como o app traduz cada ação sua: `criterios/10-revisao-humana.md`.
|
||||
|
||||
## Ordem de trabalho
|
||||
|
||||
Siga nesta ordem. Pular a Fase 2 ou a 3 leva a decisões erradas.
|
||||
@@ -44,6 +60,9 @@ Siga nesta ordem. Pular a Fase 2 ou a 3 leva a decisões erradas.
|
||||
| **7** | Cortar a lista pelo ritmo | `criterios/07-ritmo.md` |
|
||||
| **8** | Montar o JSON de saída | `criterios/08-formato-de-saida.md` |
|
||||
|
||||
**Antes da Fase 5, leia `criterios/10-revisao-humana.md`.** Ele descreve o que
|
||||
o app faz com o seu JSON — e muda *como* escrever cortes e zooms, não só quais.
|
||||
|
||||
## As três armadilhas
|
||||
|
||||
Cada uma já causou erro silencioso em material real:
|
||||
@@ -69,10 +88,14 @@ disso e o efeito cai no frame errado — sem erro visível.
|
||||
|
||||
```
|
||||
build_voice_timeline → [você decide] → refine_voice_timeline → [você corta
|
||||
pelo ritmo] → apply_voice_actions → remove_media_silence →
|
||||
generate_dynamic_subtitles
|
||||
pelo ritmo] → [revisão humana na etapa 5 do app] → apply_voice_actions →
|
||||
remove_media_silence → generate_dynamic_subtitles
|
||||
```
|
||||
|
||||
A revisão humana entra entre a sua decisão e a aplicação. É por isso que o
|
||||
`reason` importa tanto: ele é lido ali, na hora de decidir se a sua escolha
|
||||
fica.
|
||||
|
||||
Vícios de linguagem e lacunas longas entram na **sua** lista, num ripple só
|
||||
(`06-texto-corte-marcador.md`). Silêncio fino e legendas vêm depois.
|
||||
|
||||
|
||||
@@ -1,5 +1,8 @@
|
||||
# 01 — Leitura do JSON
|
||||
|
||||
> **Escopo:** Como ler o voice_timeline em camadas, sem recalcular o que já foi medido.
|
||||
> **Quando:** Fase 1 — ver a ordem de trabalho em `../SKILL.md`.
|
||||
|
||||
O arquivo `<mídia>_voice_timeline.json` é a entrada de todo o trabalho.
|
||||
Leia em camadas, de cima para baixo, e só desça quando precisar.
|
||||
|
||||
@@ -50,14 +53,19 @@ use para decidir; existem para permitir a reanálise da Fase 2.
|
||||
## O timestamp por palavra tem um viés conhecido
|
||||
|
||||
O início de cada palavra vem sistematicamente **adiantado em ~0,3-0,5s** em
|
||||
relação ao ataque real da fala — medido em material real com ffmpeg
|
||||
(`astats`), consistente em 6 pontos do mesmo vídeo. O fim da palavra não
|
||||
tem esse problema (erro de poucos centésimos). Causa: `word_timestamps` do
|
||||
faster-whisper deriva por atenção cruzada, sem alinhamento forçado — ver
|
||||
`05_EXPERIENCIAS.md`, entrada de 2026-08-19.
|
||||
relação ao ataque real da fala — medido em material real com ffmpeg (`astats`),
|
||||
consistente em 6 pontos do mesmo vídeo. O fim da palavra não tem esse problema
|
||||
(erro de poucos centésimos). Causa: `word_timestamps` do faster-whisper deriva
|
||||
por atenção cruzada, sem alinhamento forçado — ver `05_EXPERIENCIAS.md`, entrada
|
||||
de 2026-08-19.
|
||||
|
||||
**Quando o pipeline já corrigiu isso:** se `layers.alignment` for `true`
|
||||
(transcript gerado com alinhamento forçado fonético via whisperx, implementado
|
||||
depois desse aviso), o viés foi removido na origem — **não aplique o offset
|
||||
manual** abaixo. O aviso vale só para transcripts antigos sem `layers.alignment`.
|
||||
|
||||
Isso não é "reestimar no olho" — é um bug de medição na fonte, não um
|
||||
julgamento seu. Na prática:
|
||||
julgamento seu. Na prática (somente sem `layers.alignment`):
|
||||
|
||||
- Ao posicionar um `zoom` cujo `start` precisa cair exatamente na palavra
|
||||
(não uma frase inteira), some **+0,3 a +0,4s** ao timestamp do JSON antes
|
||||
@@ -66,6 +74,3 @@ julgamento seu. Na prática:
|
||||
- **Não aplique essa correção a `gap_before` para decidir corte** — a régua
|
||||
de silêncio (`06-texto-corte-marcador.md`) já é conservadora o bastante
|
||||
para absorver esse erro; corrigir os dois ao mesmo tempo é redundante.
|
||||
- Se um dia o pipeline ganhar alinhamento forçado (WhisperX), este aviso
|
||||
perde a razão de existir — confira se `layers` ou a versão do documento
|
||||
já indicam isso antes de aplicar o offset manualmente.
|
||||
|
||||
@@ -1,5 +1,8 @@
|
||||
# 02 — Triagem: roteiro vs. conversa de bastidor
|
||||
|
||||
> **Escopo:** Separar o texto do roteiro da conversa de bastidor — tarefa de texto, nunca de limiar.
|
||||
> **Quando:** Fase 2 — ver a ordem de trabalho em `../SKILL.md`.
|
||||
|
||||
**Primeira coisa a fazer, antes de qualquer decisão de efeito.**
|
||||
|
||||
Material bruto de gravação quase nunca é uma tomada só. A pessoa lê o
|
||||
|
||||
@@ -1,5 +1,8 @@
|
||||
# 03 — Escolha da melhor tomada
|
||||
|
||||
> **Escopo:** Qual tomada de cada frase sobrevive, e o que fazer em caso de empate.
|
||||
> **Quando:** Fase 3 — ver a ordem de trabalho em `../SKILL.md`.
|
||||
|
||||
A mesma frase costuma aparecer 2, 3, 4 vezes. Seu trabalho é ficar com
|
||||
**uma**.
|
||||
|
||||
|
||||
@@ -1,5 +1,8 @@
|
||||
# 04 — Reanálise do material que sobrou
|
||||
|
||||
> **Escopo:** Renormalizar a ênfase sobre o que sobrou, antes de escolher zooms.
|
||||
> **Quando:** Fase 4 — ver a ordem de trabalho em `../SKILL.md`.
|
||||
|
||||
**Não escolha zooms com os números da análise bruta.**
|
||||
|
||||
## O problema
|
||||
|
||||
@@ -1,5 +1,8 @@
|
||||
# 05 — Zoom (punch-in)
|
||||
|
||||
> **Escopo:** Onde dar punch-in, qual janela e qual escala — e o que a escala significa além do zoom.
|
||||
> **Quando:** Fase 5 — ver a ordem de trabalho em `../SKILL.md`.
|
||||
|
||||
## Quando usar
|
||||
|
||||
No momento em que o argumento vira. Um pico acústico só merece zoom se for
|
||||
@@ -46,14 +49,24 @@ automático acerta na quase totalidade dos casos.
|
||||
|
||||
## Escala
|
||||
|
||||
| Valor | Uso |
|
||||
|---|---|
|
||||
| 1,15 | sutil |
|
||||
| 1,18 – 1,3 | padrão |
|
||||
| 1,5 | forte |
|
||||
| Valor | Uso | Vira, na tela de revisão |
|
||||
|---|---|---|
|
||||
| 1,15 | sutil | ênfase **1 — Leve** |
|
||||
| 1,18 – 1,3 | padrão | ênfase **2 — Média** |
|
||||
| 1,5 | forte | ênfase **3 — Forte** |
|
||||
|
||||
Em vídeo institucional, fique na faixa baixa. Acima de 3,0 é rejeitado.
|
||||
|
||||
**A escala tem um segundo efeito, e ele é maior que o zoom.** A frase que
|
||||
recebe um zoom é marcada como **ênfase** na etapa 5, e frase de ênfase recebe
|
||||
**legenda dinâmica**; as demais ficam com legenda comum. Ou seja: escolher onde
|
||||
dar zoom é também escolher onde o texto ganha tratamento tipográfico.
|
||||
|
||||
Consequência prática: **não espalhe zoom "por segurança"**. Cada um promove uma
|
||||
frase a destaque em duas dimensões ao mesmo tempo. Na dúvida, deixe sem — o
|
||||
editor promove numa tecla, e despromover custa mais que promover.
|
||||
Detalhe: `10-revisao-humana.md`.
|
||||
|
||||
O zoom é **relativo ao enquadramento existente**: se o clipe já tem escala
|
||||
1,77 (material gravado de lado e reenquadrado), um zoom 1,18 anima de 1,77
|
||||
para 2,09 e preserva rotação e posição.
|
||||
|
||||
@@ -1,5 +1,8 @@
|
||||
# 06 — Texto, corte e marcador
|
||||
|
||||
> **Escopo:** Texto na tela, o que cortar (inclui muletas e lacunas) e quando marcar.
|
||||
> **Quando:** Fase 6 — ver a ordem de trabalho em `../SKILL.md`.
|
||||
|
||||
## Texto
|
||||
|
||||
Para fixar um **conceito, número ou nome** que o espectador precisa reter.
|
||||
@@ -50,6 +53,34 @@ Acima de 3s a pausa deixa de contar como ênfase por construção — medido em
|
||||
material real, lacunas de 6–9s rankeavam como os momentos mais enfáticos da
|
||||
gravação só porque a escala saturava.
|
||||
|
||||
### Nunca corte rente à palavra — deixe uma folga
|
||||
|
||||
Um `cut` cujo `start`/`end` cai exatamente no timestamp da palavra (fim da
|
||||
última palavra mantida = início do corte) produz um corte seco: a palavra é
|
||||
engolida antes de terminar de soar, e a fala seguinte começa sem nenhum ar.
|
||||
Isso é diferente de cortar a pausa curta (que seria apagar a própria ênfase,
|
||||
proibido acima) — aqui a pausa **já existe** entre o fim de um bloco mantido
|
||||
e o início do próximo, e o corte está comendo justamente essa margem.
|
||||
|
||||
Ao escrever a borda de um `cut` que encosta em fala mantida (não em silêncio
|
||||
puro), recue **~0,15–0,25s** para dentro do próprio corte, nos dois lados:
|
||||
|
||||
- o `start` do corte fica ~0,2s **depois** do fim real da última palavra
|
||||
mantida;
|
||||
- o `end` do corte fica ~0,2s **antes** do início real da próxima palavra
|
||||
mantida.
|
||||
|
||||
Caso real (projeto Mastopexia): um corte escrito rente (`10.77 → 95.50`,
|
||||
exatamente nos timestamps de palavra) soava abrupto nas duas emendas.
|
||||
Recuado para `10.97 → 95.30`, cada lado ganhou ~0,2s de respiro sem alterar
|
||||
o que é dito — e não empurra o próximo zoom/marcador contra a borda do corte
|
||||
(ver `05-zoom.md` sobre janelas encostadas em corte).
|
||||
|
||||
Isso vale também para o **início e o fim do vídeo**: ar morto antes da
|
||||
primeira palavra e depois da última também leva `cut`, com a mesma folga —
|
||||
não é "silêncio dentro da fala" (isso é `remove_media_silence`), é o mesmo
|
||||
corte de tomada/bastidor que você já está decidindo.
|
||||
|
||||
### O que continua NÃO sendo seu trabalho
|
||||
|
||||
| Tarefa | Ferramenta | Por quê |
|
||||
|
||||
@@ -1,5 +1,8 @@
|
||||
# 07 — Ritmo
|
||||
|
||||
> **Escopo:** Quantos efeitos cabem: os tetos e como escolher o que fica.
|
||||
> **Quando:** Fase 7 — ver a ordem de trabalho em `../SKILL.md`.
|
||||
|
||||
**O erro mais comum é efeito demais.** Cansa mais que efeito de menos, e
|
||||
denuncia edição automática.
|
||||
|
||||
|
||||
@@ -1,5 +1,8 @@
|
||||
# 08 — Formato de saída
|
||||
|
||||
> **Escopo:** O JSON de entrega: estrutura, regras e como o programa trata erros.
|
||||
> **Quando:** Fase 8 — ver a ordem de trabalho em `../SKILL.md`.
|
||||
|
||||
O produto do seu trabalho é **este JSON**. É ele que vai para o programa
|
||||
gerar o FCPXML. Você nunca escreve XML.
|
||||
|
||||
@@ -53,6 +56,19 @@ uma. Um `reason` vazio é sinal de decisão sem critério.
|
||||
Inclua o dado que embasou: *"abertura: 'Aquela mama' (ênfase 0.42)"* é útil;
|
||||
*"zoom"* não é.
|
||||
|
||||
Não é campo de log: o texto é **exibido na tela de revisão**, ao lado da frase,
|
||||
e é o que o editor lê antes de manter ou desfazer o que você decidiu.
|
||||
|
||||
### 6. Corte: alinhe à intenção
|
||||
A tela lê cada `cut` contra as frases da transcrição:
|
||||
|
||||
- cobre **≥ 60%** de uma frase → aquela frase é **removida**;
|
||||
- toca só o **começo** ou só o **fim** → vira **trim** (a frase fica, aparada).
|
||||
|
||||
Então corte a frase **inteira** quando quiser removê-la, e corte **só da borda
|
||||
até a palavra** quando quiser aparar uma hesitação. Um corte de meia frase é
|
||||
ambíguo — passa de 60% e apaga a linha toda. Detalhe: `10-revisao-humana.md`.
|
||||
|
||||
## Como o programa trata erros
|
||||
|
||||
- **Ação inválida** → rejeitada e reportada **individualmente**. Uma linha
|
||||
|
||||
@@ -1,5 +1,8 @@
|
||||
# 09 — Quando a análise veio incompleta
|
||||
|
||||
> **Escopo:** O que fazer quando uma camada da análise não rodou.
|
||||
> **Quando:** Fase 0 — ver a ordem de trabalho em `../SKILL.md`.
|
||||
|
||||
O bloco `layers` no topo do JSON diz **o que de fato rodou**. Leia antes de
|
||||
qualquer outra coisa.
|
||||
|
||||
|
||||
@@ -0,0 +1,121 @@
|
||||
# 10 — A revisão humana: o que acontece com o seu JSON
|
||||
|
||||
> **Escopo:** O que o app faz com o seu JSON na etapa 5 — muda como escrever as ações.
|
||||
> **Quando:** ler antes da Fase 5 — ver a ordem de trabalho em `../SKILL.md`.
|
||||
|
||||
> Leia antes de decidir cortes e zooms. Muda **como** escrever as ações, não
|
||||
> apenas quais.
|
||||
|
||||
Seu JSON não vai direto para o FCPXML. Ele é colado no app e abre na **etapa 5
|
||||
do Assistente**, uma tela onde o editor vê cada frase do roteiro com a sua
|
||||
decisão já aplicada e lapida antes de gerar.
|
||||
|
||||
Isso tem duas consequências práticas:
|
||||
|
||||
1. **Suas decisões são lidas por uma pessoa, frase a frase.** Uma decisão sem
|
||||
motivo explícito parece arbitrária — e será desfeita.
|
||||
2. **A tela traduz suas ações para o vocabulário dela.** Se você não escrever
|
||||
as ações do jeito que essa tradução espera, a intenção se perde no caminho.
|
||||
|
||||
---
|
||||
|
||||
## Como cada ação sua é lida
|
||||
|
||||
O app quebra a gravação em **frases** (os segmentos do voice timeline) e
|
||||
projeta suas ações sobre elas.
|
||||
|
||||
### `cut`
|
||||
|
||||
| O corte cobre… | Vira | Na tela |
|
||||
|---|---|---|
|
||||
| **≥ 60%** da frase | frase **desativada** | apagada, riscada, reativável num clique |
|
||||
| só o **começo** ou só o **fim** | **trim** da frase | a frase fica, aparada nas pontas |
|
||||
| um pedaço no **meio** | nada em si | só conta para a regra dos 60% |
|
||||
|
||||
O trim é **encaixado na fronteira de palavra** mais próxima. Você não precisa
|
||||
acertar o frame: mire na palavra onde a frase deve começar ou terminar.
|
||||
|
||||
**O que isso pede de você:** decida se está removendo *a linha* ou *aparando*
|
||||
uma ponta, e escreva o corte de acordo.
|
||||
|
||||
- Removendo a linha → corte a frase inteira, de ponta a ponta.
|
||||
- Aparando um falso começo → corte só da borda até a palavra onde a fala
|
||||
engata. Um corte que cobre meia frase é ambíguo: passa de 60% e apaga a linha
|
||||
toda, quando você só queria tirar a hesitação.
|
||||
|
||||
### `zoom` e `text`
|
||||
|
||||
Qualquer `zoom` ou `text` que toque uma frase marca aquela frase como
|
||||
**ênfase** — e ênfase, nesta tela, significa **duas coisas**:
|
||||
|
||||
> **A frase de ênfase recebe zoom E legenda dinâmica. As demais recebem
|
||||
> legenda comum.**
|
||||
|
||||
O nível vem da sua `scale`:
|
||||
|
||||
| `scale` | Nível na tela | |
|
||||
|---|---|---|
|
||||
| 1,15 | 1 — Leve | |
|
||||
| 1,3 | 2 — Média | |
|
||||
| 1,5 | 3 — Forte | |
|
||||
| omitida, ou uma ação `text` | 2 — Média | padrão |
|
||||
|
||||
Sem nenhuma ação sua, a tela deriva o nível do `peak_emphasis` da frase
|
||||
(< 0,25 → sem ênfase; < 0,45 → leve; < 0,65 → média; acima → forte). **A sua
|
||||
decisão sempre ganha da derivação automática.**
|
||||
|
||||
**O que isso pede de você:** escolher a escala com intenção. Ela não é só
|
||||
"quanto amplia" — é o peso que aquela frase terá no vídeo inteiro, incluindo o
|
||||
tratamento da legenda. Um zoom leve numa frase de apoio não é neutro: promove
|
||||
aquela frase a destaque tipográfico também.
|
||||
|
||||
### `marker`
|
||||
|
||||
Não altera a frase. Continua sendo o seu recado para o editor conferir uma
|
||||
emenda — e é a ferramenta certa quando você está em dúvida (ver
|
||||
`03-escolha-da-melhor-tomada.md`).
|
||||
|
||||
---
|
||||
|
||||
## `reason` aparece na tela
|
||||
|
||||
Não é campo de log. O texto que você escreve em `reason` é exibido para o
|
||||
editor ao lado da frase selecionada, e é o que ele lê antes de manter ou
|
||||
desfazer a sua decisão.
|
||||
|
||||
Escreva para quem está com pressa e vai decidir na hora:
|
||||
|
||||
- **Bom:** `"fecho, pico em 'devolver' (ênfase 0.34) — escala mais forte por ser o fechamento da peça"`
|
||||
- **Ruim:** `"zoom"` · `"corte necessário"` · `"melhor tomada"`
|
||||
|
||||
A regra prática: se o `reason` não contém **o dado** que embasou (a palavra, o
|
||||
número, a comparação entre tomadas), você provavelmente não tinha critério —
|
||||
tinha impressão.
|
||||
|
||||
---
|
||||
|
||||
## O que a tela NÃO desfaz por você
|
||||
|
||||
- **Tempo errado continua errado.** A tela mostra suas ações no eixo da mídia
|
||||
original; se você compensou para pós-corte, tudo aparece no lugar errado e o
|
||||
editor não tem como adivinhar o que você quis dizer.
|
||||
- **Excesso de zoom continua excesso.** A tela não impõe o teto de 2–4 por
|
||||
minuto (`07-ritmo.md`) — ela mostra o que você mandou. Efeito demais chega
|
||||
ao editor como trabalho de limpeza.
|
||||
- **Frase promovida a ênfase sem querer.** Como zoom e legenda dinâmica andam
|
||||
juntos, espalhar zooms "de segurança" enche o vídeo de legenda dinâmica. Na
|
||||
dúvida, deixe sem — o editor promove; é mais barato que despromover.
|
||||
|
||||
---
|
||||
|
||||
## Depois da revisão
|
||||
|
||||
O editor pode, na tela: mudar o nível de ênfase (0–3), desativar ou reativar
|
||||
frases, corrigir o texto, aparar as pontas por palavra, reclassificar entre
|
||||
roteiro e bastidor e acrescentar zooms manuais em trechos arbitrários.
|
||||
|
||||
O resultado vira um `_phrase_review.json` e o `_phrase_actions.json` derivado —
|
||||
e é esse que a geração usa. **Seu JSON é o ponto de partida da conversa, não a
|
||||
palavra final.** Trabalhe para ser um bom ponto de partida: decisões
|
||||
defensáveis, motivos legíveis e nenhuma escolha que o editor precise desfazer
|
||||
antes de começar.
|
||||
+6
-2
@@ -30,6 +30,7 @@ Thumbs.db
|
||||
# Env files (NUNCA commitar — contêm segredos)
|
||||
*.env
|
||||
.env
|
||||
admin/gart-rag.env
|
||||
|
||||
# Graphify output (gerado, não rastrear)
|
||||
graphify-out/
|
||||
@@ -37,6 +38,9 @@ graphify-out/
|
||||
# FCPXML bundles de exemplo (podem ser grandes)
|
||||
*.fcpxmld/
|
||||
|
||||
# WhisperX models cache
|
||||
models/
|
||||
# Cache de modelos Whisper baixados (código/models, ~11 GB, HuggingFace hub
|
||||
# format). Âncora em /code/models/ — NUNCA "models/" solto: isso também
|
||||
# ignorava fcpxml/models/, o pacote de dados do engine (ver
|
||||
# Engine/docs/05_EXPERIENCIAS.md #36).
|
||||
/code/models/
|
||||
whisper/
|
||||
|
||||
@@ -9,25 +9,91 @@ normalmente; a regra é sobre a comunicação com o usuário.
|
||||
|
||||
## What This Is
|
||||
|
||||
MCP server that reads/writes Final Cut Pro XML (FCPXML) files. 73 tools for timeline analysis, batch editing, QC, generation, multi-track support, media relink, NLE export, transcript-based editing (local Whisper), and LIVE FCP control (push_to_fcp / list_fcp_libraries via Apple events). Reads FCPXML 1.8–1.14 (incl. `.fcpxmld` bundles with sidecar preservation), writes 1.13 by default. Dual-mode (XML + Live) direction: `code/docs/CAPABILITY-AUDIT-2026-06.md`.
|
||||
MCP server that reads/writes Final Cut Pro XML (FCPXML) files. 77 tools for timeline analysis, batch editing, QC, generation, multi-track support, media relink, NLE export, transcript-based editing (local Whisper), LIVE FCP control (push_to_fcp / list_fcp_libraries via Apple events), and local-LLM voice scripting (editar-por-voz against Ollama/Gemma 3). Reads FCPXML 1.8–1.14 (incl. `.fcpxmld` bundles with sidecar preservation), writes 1.13 by default. Dual-mode (XML + Live) direction: `code/docs/CAPABILITY-AUDIT-2026-06.md`.
|
||||
|
||||
## Architecture
|
||||
|
||||
Toda a estrutura do projeto fica em `code/`. A pasta `admin/` fica na raiz (fora de `code/`).
|
||||
|
||||
```
|
||||
code/server.py — MCP server entry point. All 62 tool definitions, handlers, resources, prompts.
|
||||
Dispatch dict pattern: TOOL_HANDLERS maps tool names → async handler functions.
|
||||
Há **duas portas de entrada** para o mesmo engine: o MCP (Claude decide a
|
||||
edição) e a ponte JSON (o app macOS opera). Nenhuma das duas tem lógica de
|
||||
timeline — as duas delegam a `fcpxml/`.
|
||||
|
||||
code/fcpxml/parser.py — Reads FCPXML → Python objects (Timeline, Clip, ConnectedClip, Marker, etc.)
|
||||
code/fcpxml/writer.py — Writes modifications back to FCPXML. Handles markers, trimming, gaps, transitions.
|
||||
code/fcpxml/rough_cut.py — Generates new timelines from source clips (rough cuts, montages, A/B rolls).
|
||||
code/fcpxml/diff.py — Timeline comparison engine. Detects added/removed/moved/trimmed clips & markers.
|
||||
code/fcpxml/export.py — DaVinci Resolve FCPXML v1.9 export + FCP7 XMEML v5 export for cross-NLE workflows.
|
||||
code/fcpxml/models.py — Data classes: TimeValue, Timecode, Clip, ConnectedClip, CompoundClip, Timeline, etc.
|
||||
code/fcpxml/media_intel.py — Real media analysis. Audio silence detection + beat detection.
|
||||
code/fcpxml/dtd.py — Validates output against Apple's official DTDs.
|
||||
```
|
||||
code/server.py — MCP entry point (592 linhas). Só dispatch: TOOL_HANDLERS.
|
||||
code/server_tools/ — Os handlers das 77 tools, um módulo por categoria.
|
||||
code/server_tools/_shared/ — Helpers compartilhados (paths, project, formatting,
|
||||
captions, detection, media).
|
||||
|
||||
code/fcpxml/parser.py — Reads FCPXML → Python objects (Timeline, Clip, Marker…)
|
||||
code/fcpxml/writer/ — PACOTE. Edição/escrita de FCPXML. FCPXMLModifier é
|
||||
montado por mixins, um por assunto (markers, trim,
|
||||
speed, titles, cut, silence…). Ver writer/modifier.py.
|
||||
code/fcpxml/models/ — PACOTE. Data classes por família: timing, timeline,
|
||||
enums, subtitles, qc, planning.
|
||||
code/fcpxml/rough_cut.py — Generates new timelines (rough cuts, montages, A/B rolls).
|
||||
code/fcpxml/diff.py — Timeline comparison engine.
|
||||
code/fcpxml/export.py — DaVinci Resolve v1.9 + FCP7 XMEML v5 export.
|
||||
code/fcpxml/media_intel.py — Silence detection + beat detection.
|
||||
code/fcpxml/dtd.py — Validates output against Apple's official DTDs.
|
||||
code/fcpxml/voice_*.py — Pipeline de voz: features → emphasis → voice_timeline
|
||||
→ voice_actions → phrase_review. Ver Engine/docs/02.
|
||||
|
||||
admin/models_api.py — Ponte JSON com o app: docstring de comandos + dispatch.
|
||||
admin/api/ — Os 37 comandos, um módulo por assunto.
|
||||
code/MacApp/Sources/ — App SwiftUI. Compilado por swiftc (sem Xcode/SPM).
|
||||
```
|
||||
|
||||
Os dois `__init__.py` de pacote (`writer/`, `models/`) reexportam tudo, então
|
||||
`from .writer import FCPXMLModifier` e `from .models import TimeValue` seguem
|
||||
valendo em todo o projeto.
|
||||
|
||||
## Documentação (MANDATORY)
|
||||
|
||||
A documentação viva fica em `code/Engine/docs/`. Cada arquivo tem **uma função
|
||||
específica** — leia só o que a tarefa exige, não o conjunto. Carregar
|
||||
documentação que não é do assunto custa tempo e processamento sem entregar nada.
|
||||
|
||||
### Qual arquivo abrir
|
||||
|
||||
| Sua tarefa | Abra | Não precisa de |
|
||||
|-----------|------|----------------|
|
||||
| Entender como o sistema é dividido | `01_ARCHITECTURE.md` | o resto |
|
||||
| Achar onde mora uma função do engine | `02_MODULES.md` | 01, 03 |
|
||||
| Criar/alterar uma ferramenta MCP | `03_SERVER_TOOLS.md` | 08 |
|
||||
| Entender ou rodar os testes | `04_TESTS_AND_WORKFLOW.md` | — |
|
||||
| "Isso já quebrou antes?" | `05_EXPERIENCIAS.md` — **só o índice no topo** | as entradas que não são a sua |
|
||||
| Checklist antes de fechar | `06_BOAS_PRATICAS.md` | — |
|
||||
| Mexer no app / no Assistente | `08_APP_MACOS.md` | 02, 03 |
|
||||
| Escolher o que fazer, ver o que está aberto | `09_MANUTENCAO.md` | — |
|
||||
|
||||
Quando não souber por onde começar: `09_MANUTENCAO.md`. Ele roteia para o resto.
|
||||
|
||||
### Regra de atualização (obrigatória)
|
||||
|
||||
**Toda alteração de código atualiza a documentação no mesmo commit.** Doc velha
|
||||
engana mais do que doc ausente — quem lê confia nela e erra com confiança.
|
||||
|
||||
| Você alterou | Atualize |
|
||||
|--------------|----------|
|
||||
| Estrutura de pastas, camadas ou dependências | `01_ARCHITECTURE.md` |
|
||||
| Criou/moveu/dividiu módulo em `fcpxml/` | `02_MODULES.md` (tabela + linhas) |
|
||||
| Criou/removeu ferramenta MCP | `03_SERVER_TOOLS.md` + contagem no `CLAUDE.md` |
|
||||
| Comando da ponte | docstring de `admin/models_api.py` + `08_APP_MACOS.md` |
|
||||
| Tela ou fluxo do app | `08_APP_MACOS.md` |
|
||||
| Resolveu ou abriu uma dívida | `09_MANUTENCAO.md` §2 |
|
||||
| Bateu num problema estrutural ou erro recorrente | `05_EXPERIENCIAS.md` + **índice no topo** |
|
||||
|
||||
Se um número (tools, testes, linhas) mudou, corrija onde ele aparece. Se um
|
||||
documento divergir do código, **o código está certo** — conserte o documento.
|
||||
|
||||
### Ao escrever documentação
|
||||
|
||||
- **Um assunto por arquivo.** Se um doc começar a cobrir dois, divida.
|
||||
- **Diga o que não está ali** e para onde ir — economiza a leitura seguinte.
|
||||
- **Fatos verificados**, não suposições: rode o comando e use o número real.
|
||||
- **Registre o porquê**, não só o quê. O "o quê" está no código; o "por quê"
|
||||
se perde, e é o que evita alguém desfazer uma decisão por engano.
|
||||
|
||||
## Key Patterns
|
||||
|
||||
@@ -69,12 +135,19 @@ não passar. Equivalente a rodar manualmente os dois comandos abaixo.
|
||||
|
||||
Sempre que uma alteração for feita no app (MacApp/) durante o período de
|
||||
implementação, **compile e rode o programa localmente no computador** para
|
||||
validar visualmente a alteração, além de rodar os testes:
|
||||
validar visualmente a alteração, além de rodar os testes. O comando padrão
|
||||
para isso — que fecha a instância anterior, recompila e abre o app para
|
||||
conferência — é:
|
||||
|
||||
```bash
|
||||
cd code && ./MacApp/build_app.sh --run # compila e abre o app localmente
|
||||
admin/run_app.command # compila e abre o app localmente (padrão de revisão)
|
||||
```
|
||||
|
||||
Equivalente a `cd code && ./MacApp/build_app.sh --run`, mas desacoplado do
|
||||
Terminal. **Toda vez que uma alteração for concluída, rode este arquivo
|
||||
automaticamente** para já conseguirmos revisar o que foi feito antes de
|
||||
fechar a tarefa.
|
||||
|
||||
Regra geral: após qualquer alteração, o app deve ser executado localmente
|
||||
antes de concluir a tarefa. Se houver erro de compilação, corrija antes de
|
||||
seguir.
|
||||
@@ -88,7 +161,7 @@ CI runs both on every push to main. If either fails, the commit gets an X on Git
|
||||
|
||||
## Testing
|
||||
|
||||
1342 tests across 34 files. `test_models.py` covers TimeValue arithmetic, Timecode parsing/formatting, Clip properties, validation models, and Timeline helpers. `test_writer.py` covers insert_clip, add_marker (all types), trim_clip, delete_clip, split_clip, and change_speed operations. `test_server.py` covers MCP tool handlers, parsers, and dispatch. `test_rough_cut.py` covers RoughCutGenerator. `test_features_v05.py` covers connected clips, roles, timeline diff, reformat, silence detection, export, and backward compatibility. `test_marker_pipeline.py` covers build_marker_element shared builder, batch auto-modes, clip index duplicate-name behavior, and write_fcpxml output format. `test_refactored_helpers.py` covers _index_elements, _iter_spine_clips, _find_spine_clip_at_seconds, _resolve_clip_duration, _make_asset_clip, _format_batch_result, and serialize_xml edge cases. `test_transcribe.py` covers phrase/filler span matching, range merge/invert algebra, whisper graceful degradation, and transcript-driven handler cuts against cached transcripts. `test_media_intel.py` covers silencedetect stderr parsing, source-to-timeline mapping, parameter bounds, and real-WAV ffmpeg integration (skips without ffmpeg; CI installs it). Tests use `examples/sample.fcpxml` as fixture data and inline XML fixtures. Tests create temp files and clean up after.
|
||||
1498 tests across 43 files, all under `code/tests/`. Um teste fora dessa pasta não roda (`testpaths = ["tests"]`) — se você criar um em outro lugar, confirme que a contagem total subiu. Cobertura por área: `test_models.py` (TimeValue/Timecode/Clip/Timeline), `test_writer.py` (insert/marker/trim/delete/split/speed), `test_server.py` (handlers e dispatch), `test_rough_cut.py`, `test_features_v05.py` (connected clips, roles, diff, reformat, silêncio, export), `test_marker_pipeline.py`, `test_refactored_helpers.py`, `test_transcribe.py`, `test_media_intel.py` (pula sem ffmpeg; o CI instala), `test_phrase_review.py` (revisão de frases da etapa 5) e `test_models_api.py` (comandos da ponte). Fixtures: `examples/sample.fcpxml` e XML inline. Os testes criam temporários e limpam depois.
|
||||
|
||||
## FCPXML Gotchas
|
||||
|
||||
|
||||
@@ -0,0 +1,20 @@
|
||||
"""Comandos da ponte JSON usada pelo app, agrupados por assunto.
|
||||
|
||||
O setup de sys.path mora aqui, e só aqui, porque o pacote é importado antes de
|
||||
qualquer um dos seus módulos (`from admin.api import models, voice, ...`
|
||||
dispara este arquivo primeiro). Cada módulo de comando importa `fcpxml.*`
|
||||
antes de importar `.shared` — sem o path já pronto neste ponto, o primeiro
|
||||
desses imports falha com `ModuleNotFoundError`. Repetir o cálculo em cada
|
||||
módulo (como era antes) é frágil por ordem: o app roda `admin/models_api.py`
|
||||
por caminho absoluto, então `__file__` está sempre correto, mas cada arquivo
|
||||
que refizesse essa conta um nível de diretório errado — como aconteceu quando
|
||||
`_shared.py` virou este pacote e `admin/code` (inexistente) saiu no lugar de
|
||||
`code/` — quebrava em silêncio até alguém rodar o comando de verdade.
|
||||
"""
|
||||
|
||||
import sys
|
||||
from pathlib import Path
|
||||
|
||||
_CODE_DIR = str(Path(__file__).resolve().parent.parent.parent / "code")
|
||||
if _CODE_DIR not in sys.path:
|
||||
sys.path.insert(0, _CODE_DIR)
|
||||
@@ -0,0 +1,122 @@
|
||||
"""Edições no projeto: silêncio, corte por texto, preenchimento, marcadores.
|
||||
|
||||
Extraído de models_api.py — a tabela de comandos segue lá.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import asyncio
|
||||
from pathlib import Path
|
||||
|
||||
from fcpxml.model_manager import (
|
||||
load_silence_config,
|
||||
save_silence_config,
|
||||
)
|
||||
|
||||
from . import shared
|
||||
from .shared import (
|
||||
_derived_output,
|
||||
_emit_no_change_or_error,
|
||||
)
|
||||
|
||||
|
||||
def cmd_remove_silences(args: dict) -> int:
|
||||
"""Run the canonical server silence remover into a suffixed copy."""
|
||||
path = str(args.get("path", ""))
|
||||
if not path or not Path(path).exists():
|
||||
shared.emit({"ok": False, "error": "Arquivo de projeto não encontrado."})
|
||||
return 1
|
||||
try:
|
||||
from server import handle_remove_media_silence
|
||||
|
||||
output = _derived_output(path, "_silence_removed", args)
|
||||
contents = asyncio.run(handle_remove_media_silence({**args, "filepath": path, "output_path": output}))
|
||||
message = "\n".join(getattr(content, "text", str(content)) for content in contents)
|
||||
if not Path(output).exists():
|
||||
return _emit_no_change_or_error(path, message)
|
||||
shared.emit({"ok": True, "path": output, "message": message})
|
||||
return 0
|
||||
except Exception as exc:
|
||||
shared.emit({"ok": False, "error": str(exc)})
|
||||
return 1
|
||||
|
||||
def cmd_edit_by_transcript(args: dict) -> int:
|
||||
"""Cut (or keep only) spoken phrases, using each media's cached transcript."""
|
||||
path = str(args.get("path", ""))
|
||||
phrases = args.get("phrases") or []
|
||||
if not path or not Path(path).exists():
|
||||
shared.emit({"ok": False, "error": "Arquivo de projeto não encontrado."})
|
||||
return 1
|
||||
if not isinstance(phrases, list) or not [p for p in phrases if str(p).strip()]:
|
||||
shared.emit({"ok": False, "error": "Informe ao menos uma frase para cortar."})
|
||||
return 1
|
||||
try:
|
||||
from server import handle_edit_by_transcript
|
||||
|
||||
output = _derived_output(path, "_transcript_edit", args)
|
||||
contents = asyncio.run(handle_edit_by_transcript({**args, "filepath": path, "output_path": output}))
|
||||
message = "\n".join(getattr(content, "text", str(content)) for content in contents)
|
||||
if not Path(output).exists():
|
||||
shared.emit({"ok": False, "error": message})
|
||||
return 1
|
||||
shared.emit({"ok": True, "path": output, "message": message})
|
||||
return 0
|
||||
except Exception as exc:
|
||||
shared.emit({"ok": False, "error": str(exc)})
|
||||
return 1
|
||||
|
||||
def cmd_remove_filler_words(args: dict) -> int:
|
||||
"""Cut filler words (um, uh, ...) out, using each media's cached transcript."""
|
||||
path = str(args.get("path", ""))
|
||||
if not path or not Path(path).exists():
|
||||
shared.emit({"ok": False, "error": "Arquivo de projeto não encontrado."})
|
||||
return 1
|
||||
try:
|
||||
from server import handle_remove_filler_words
|
||||
|
||||
output = _derived_output(path, "_defillered", args)
|
||||
contents = asyncio.run(handle_remove_filler_words({**args, "filepath": path, "output_path": output}))
|
||||
message = "\n".join(getattr(content, "text", str(content)) for content in contents)
|
||||
if not Path(output).exists():
|
||||
return _emit_no_change_or_error(path, message)
|
||||
shared.emit({"ok": True, "path": output, "message": message})
|
||||
return 0
|
||||
except Exception as exc:
|
||||
shared.emit({"ok": False, "error": str(exc)})
|
||||
return 1
|
||||
|
||||
def cmd_transcript_markers(args: dict) -> int:
|
||||
"""Add a marker per transcribed segment, using each media's cached transcript."""
|
||||
path = str(args.get("path", ""))
|
||||
if not path or not Path(path).exists():
|
||||
shared.emit({"ok": False, "error": "Arquivo de projeto não encontrado."})
|
||||
return 1
|
||||
try:
|
||||
from server import handle_transcript_markers
|
||||
|
||||
output = _derived_output(path, "_transcript_markers", args)
|
||||
contents = asyncio.run(handle_transcript_markers({**args, "filepath": path, "output_path": output}))
|
||||
message = "\n".join(getattr(content, "text", str(content)) for content in contents)
|
||||
if not Path(output).exists():
|
||||
shared.emit({"ok": False, "error": message})
|
||||
return 1
|
||||
shared.emit({"ok": True, "path": output, "message": message})
|
||||
return 0
|
||||
except Exception as exc:
|
||||
shared.emit({"ok": False, "error": str(exc)})
|
||||
return 1
|
||||
|
||||
def cmd_silence_config(args: dict) -> int:
|
||||
"""Read the persisted silence thresholds (noise floor, duration, padding)."""
|
||||
shared.emit({"ok": True, **load_silence_config()})
|
||||
return 0
|
||||
|
||||
def cmd_set_silence_config(args: dict) -> int:
|
||||
"""Persist silence thresholds. Only the given fields change."""
|
||||
config = save_silence_config(
|
||||
noise_db=args.get("noise_db"),
|
||||
min_silence=args.get("min_silence"),
|
||||
padding=args.get("padding"),
|
||||
)
|
||||
shared.emit({"ok": True, **config})
|
||||
return 0
|
||||
@@ -0,0 +1,137 @@
|
||||
"""Catálogo de modelos: listar, baixar, escolher, apagar.
|
||||
|
||||
Extraído de models_api.py — a tabela de comandos segue lá.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import shutil
|
||||
import subprocess
|
||||
import threading
|
||||
|
||||
from fcpxml.diarize import (
|
||||
diarization_capability,
|
||||
)
|
||||
from fcpxml.model_manager import (
|
||||
download_model,
|
||||
get_models_dir,
|
||||
is_model_downloaded,
|
||||
list_installed_models,
|
||||
load_catalog,
|
||||
load_hf_token,
|
||||
load_num_speakers,
|
||||
load_selected_model,
|
||||
load_transcript_language,
|
||||
model_cache_dir,
|
||||
save_models_dir,
|
||||
save_selected_model,
|
||||
save_transcript_language,
|
||||
)
|
||||
|
||||
from . import shared
|
||||
from .shared import (
|
||||
RECOMMENDED,
|
||||
)
|
||||
|
||||
# Downloads em andamento, para o comando `cancel` conseguir interrompê-los.
|
||||
# Mora aqui, e não no shared, porque só `download` e `cancel` o tocam — e o
|
||||
# lock é próprio: ele protege este dicionário, não a saída em stdout.
|
||||
_CANCEL: dict[str, threading.Event] = {}
|
||||
_CANCEL_LOCK = threading.Lock()
|
||||
|
||||
|
||||
def cmd_catalog() -> None:
|
||||
catalog = load_catalog()
|
||||
installed = list_installed_models()
|
||||
diar_ok, diar_msg = diarization_capability(load_hf_token())
|
||||
shared.emit(
|
||||
{
|
||||
"models": catalog,
|
||||
"installed": installed,
|
||||
"selected": load_selected_model(),
|
||||
"language": load_transcript_language(),
|
||||
"models_dir": str(get_models_dir()),
|
||||
"installed_count": len(installed),
|
||||
"recommended": list(RECOMMENDED),
|
||||
"diarization": diar_ok,
|
||||
"diarization_message": diar_msg,
|
||||
"hf_token_set": bool(load_hf_token()),
|
||||
"num_speakers": load_num_speakers(),
|
||||
}
|
||||
)
|
||||
|
||||
def cmd_download(args: dict) -> int:
|
||||
model = str(args.get("model", ""))
|
||||
if model not in _model_names():
|
||||
shared.emit({"type": "error", "message": f"Modelo desconhecido: {model}"})
|
||||
return 1
|
||||
ev = threading.Event()
|
||||
with _CANCEL_LOCK:
|
||||
_CANCEL[model] = ev
|
||||
try:
|
||||
download_model(model, progress_cb=lambda f: shared.emit({"type": "progress", "fraction": f}), cancel_event=ev)
|
||||
installed = is_model_downloaded(model)
|
||||
shared.emit({"type": "done", "installed": installed})
|
||||
if installed:
|
||||
save_selected_model(model)
|
||||
return 0 if installed else 1
|
||||
except Exception as exc:
|
||||
shared.emit({"type": "error", "message": str(exc)})
|
||||
return 1
|
||||
finally:
|
||||
with _CANCEL_LOCK:
|
||||
_CANCEL.pop(model, None)
|
||||
|
||||
def cmd_cancel(args: dict) -> None:
|
||||
model = str(args.get("model", ""))
|
||||
ev = _CANCEL.get(model)
|
||||
if ev is not None:
|
||||
ev.set()
|
||||
shared.emit({"ok": True})
|
||||
|
||||
def cmd_select(args: dict) -> None:
|
||||
model = str(args.get("model", ""))
|
||||
if not is_model_downloaded(model):
|
||||
shared.emit({"ok": False, "error": "Modelo não está instalado."})
|
||||
return
|
||||
save_selected_model(model)
|
||||
shared.emit({"ok": True, "selected": load_selected_model()})
|
||||
|
||||
def cmd_set_language(args: dict) -> int:
|
||||
"""Persist the transcription language (the default for every transcription)."""
|
||||
lang = str(args.get("language", "auto"))
|
||||
try:
|
||||
saved = save_transcript_language(lang)
|
||||
except ValueError as exc:
|
||||
shared.emit({"ok": False, "error": str(exc)})
|
||||
return 1
|
||||
shared.emit({"ok": True, "language": saved})
|
||||
return 0
|
||||
|
||||
def cmd_delete(args: dict) -> None:
|
||||
model = str(args.get("model", ""))
|
||||
try:
|
||||
shutil.rmtree(model_cache_dir(model), ignore_errors=True)
|
||||
except Exception:
|
||||
pass
|
||||
shared.emit({"ok": True})
|
||||
|
||||
def cmd_open_finder(args: dict) -> None:
|
||||
target = str(args.get("path") or model_cache_dir(str(args.get("model", ""))))
|
||||
try:
|
||||
subprocess.Popen(["open", target])
|
||||
except OSError:
|
||||
pass
|
||||
shared.emit({"ok": True})
|
||||
|
||||
def cmd_set_models_dir(args: dict) -> int:
|
||||
try:
|
||||
d = save_models_dir(str(args.get("dir", "")))
|
||||
shared.emit({"ok": True, "models_dir": d})
|
||||
return 0
|
||||
except ValueError as exc:
|
||||
shared.emit({"ok": False, "error": str(exc)})
|
||||
return 1
|
||||
|
||||
def _model_names() -> list[str]:
|
||||
return [m["internal_name"] for m in load_catalog()]
|
||||
@@ -0,0 +1,69 @@
|
||||
"""Projeto: inspecionar o .fcpxml e lembrar a pasta/arquivo em uso.
|
||||
|
||||
Extraído de models_api.py — a tabela de comandos segue lá.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
from pathlib import Path
|
||||
|
||||
from fcpxml.model_manager import (
|
||||
load_project_config,
|
||||
save_project_config,
|
||||
)
|
||||
from fcpxml.parser import parse_fcpxml
|
||||
|
||||
from . import shared
|
||||
|
||||
|
||||
def cmd_inspect(args: dict) -> int:
|
||||
"""Validate an FCPXML file and return a summary of its projects/timelines."""
|
||||
path = str(args.get("path", ""))
|
||||
if not path:
|
||||
shared.emit({"ok": False, "error": "Nenhum arquivo informado."})
|
||||
return 1
|
||||
if not Path(path).exists():
|
||||
shared.emit({"ok": False, "error": "Arquivo não encontrado."})
|
||||
return 1
|
||||
try:
|
||||
proj = parse_fcpxml(path)
|
||||
except Exception as exc:
|
||||
shared.emit({"ok": False, "error": f"Erro ao ler o projeto: {exc}"})
|
||||
return 1
|
||||
|
||||
timelines = []
|
||||
for tl in proj.timelines:
|
||||
timelines.append(
|
||||
{
|
||||
"name": tl.name,
|
||||
"duration_seconds": round(tl.duration.seconds, 3),
|
||||
"frame_rate": round(tl.frame_rate, 3),
|
||||
"width": tl.width,
|
||||
"height": tl.height,
|
||||
"clips": tl.total_clips,
|
||||
"cuts": tl.total_cuts,
|
||||
"connected": len(tl.connected_clips),
|
||||
"markers": len(tl.markers),
|
||||
}
|
||||
)
|
||||
shared.emit(
|
||||
{
|
||||
"ok": True,
|
||||
"path": path,
|
||||
"name": proj.name,
|
||||
"fcpxml_version": proj.fcpxml_version,
|
||||
"timelines": timelines,
|
||||
}
|
||||
)
|
||||
return 0
|
||||
|
||||
def cmd_project_config(args: dict) -> int:
|
||||
"""Read the last project folder/file the app was working on."""
|
||||
shared.emit({"ok": True, **load_project_config()})
|
||||
return 0
|
||||
|
||||
def cmd_set_project_config(args: dict) -> int:
|
||||
"""Persist the last project folder/file. Only the given fields change."""
|
||||
config = save_project_config(folder=args.get("folder"), file=args.get("file"))
|
||||
shared.emit({"ok": True, **config})
|
||||
return 0
|
||||
@@ -0,0 +1,90 @@
|
||||
"""Revisão de frases: montar a tela de ênfases e salvar o que foi decidido.
|
||||
|
||||
Extraído de models_api.py — a tabela de comandos segue lá.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import json
|
||||
from pathlib import Path
|
||||
|
||||
from . import shared
|
||||
|
||||
|
||||
def cmd_build_phrase_review(args: dict) -> int:
|
||||
"""Build the reviewable script (phrases + the AI's decisions) for the wizard.
|
||||
|
||||
`voice_timeline` points at the _voice_timeline.json; `actions` carries the
|
||||
decision list the model returned (inline, in any of the shapes the skill
|
||||
emits). The review is always rebuilt from the current analysis, then the
|
||||
decisions saved on a previous visit are laid back over it — reopening the
|
||||
step must show the edits the user left there without freezing the acoustics
|
||||
as they were when they left.
|
||||
"""
|
||||
from fcpxml.phrase_review import (
|
||||
build_phrase_review,
|
||||
load_phrase_review,
|
||||
merge_saved_decisions,
|
||||
)
|
||||
|
||||
timeline_path = str(args.get("voice_timeline", ""))
|
||||
if not timeline_path or not Path(timeline_path).exists():
|
||||
shared.emit({"ok": False, "error": "Análise de voz (voice_timeline.json) não encontrada."})
|
||||
return 1
|
||||
|
||||
try:
|
||||
with open(timeline_path, encoding="utf-8") as fh:
|
||||
timeline = json.load(fh)
|
||||
except (OSError, ValueError) as exc:
|
||||
shared.emit({"ok": False, "error": f"Erro ao ler a análise de voz: {exc}"})
|
||||
return 1
|
||||
|
||||
extra = [d for d in (args.get("output_dir"), args.get("media_dir")) if d]
|
||||
review = build_phrase_review(
|
||||
timeline,
|
||||
args.get("actions"),
|
||||
voice_timeline_path=timeline_path,
|
||||
extra_dirs=extra,
|
||||
)
|
||||
|
||||
saved = None if args.get("fresh") else load_phrase_review(timeline_path)
|
||||
review = merge_saved_decisions(review, saved)
|
||||
shared.emit({"ok": True, "reused": saved is not None, **review})
|
||||
return 0
|
||||
|
||||
def cmd_save_phrase_review(args: dict) -> int:
|
||||
"""Persist the edited review and the actions derived from it."""
|
||||
from fcpxml.phrase_review import save_phrase_review
|
||||
|
||||
timeline_path = str(args.get("voice_timeline", ""))
|
||||
if not timeline_path:
|
||||
shared.emit({"ok": False, "error": "Caminho da análise de voz não informado."})
|
||||
return 1
|
||||
|
||||
phrases = args.get("phrases")
|
||||
if not isinstance(phrases, list):
|
||||
shared.emit({"ok": False, "error": "Nenhuma frase para salvar."})
|
||||
return 1
|
||||
|
||||
review = {
|
||||
"version": args.get("version", "1.0"),
|
||||
"source": args.get("source", ""),
|
||||
"duration": args.get("duration", 0.0),
|
||||
"speakers": args.get("speakers", []),
|
||||
"phrases": phrases,
|
||||
"zooms": args.get("zooms", []),
|
||||
}
|
||||
try:
|
||||
review_path, actions_path = save_phrase_review(timeline_path, review)
|
||||
except OSError as exc:
|
||||
shared.emit({"ok": False, "error": f"Erro ao salvar a revisão: {exc}"})
|
||||
return 1
|
||||
|
||||
shared.emit({
|
||||
"ok": True,
|
||||
"review_path": str(review_path),
|
||||
"actions_path": str(actions_path),
|
||||
"emphasis_count": sum(1 for p in phrases if int(p.get("emphasis", 0) or 0) >= 1),
|
||||
"removed_count": sum(1 for p in phrases if not p.get("active", True)),
|
||||
})
|
||||
return 0
|
||||
@@ -0,0 +1,156 @@
|
||||
"""Base comum dos comandos da ponte: saída JSON, caminhos derivados e cache.
|
||||
|
||||
A saída passa toda por `emit`. Os módulos de comando chamam `shared.emit(...)`
|
||||
pelo módulo, e não pelo nome importado, de propósito: assim trocar `emit` num
|
||||
lugar só — como a suíte faz para capturar a saída — continua alcançando todos
|
||||
os comandos, o que deixaria de valer se cada um tivesse ligado o nome no seu
|
||||
próprio import.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import json
|
||||
import os
|
||||
import sys
|
||||
import threading
|
||||
from pathlib import Path
|
||||
from typing import Any
|
||||
|
||||
from fcpxml.diarize import build_speakers
|
||||
from fcpxml.media_intel import media_src_to_path
|
||||
from fcpxml.parser import parse_fcpxml
|
||||
|
||||
RECOMMENDED = ("large-v3", "distil-large-v3", "small", "base")
|
||||
|
||||
def _derived_output(path: str, suffix: str, args: dict) -> str:
|
||||
"""Resolve a derived XML path, optionally inside the chosen output folder."""
|
||||
output_dir = str(args.get("output_dir", "")).strip()
|
||||
if output_dir:
|
||||
directory = Path(output_dir).expanduser()
|
||||
directory.mkdir(parents=True, exist_ok=True)
|
||||
source = Path(path)
|
||||
extension = ".fcpxmld" if source.is_dir() else source.suffix
|
||||
return str(directory / f"{source.stem}{suffix}{extension}")
|
||||
from server import generate_output_path
|
||||
return generate_output_path(path, suffix)
|
||||
|
||||
def _is_no_change_message(message: str) -> bool:
|
||||
"""Whether a tool completed cleanly without needing to save a new file."""
|
||||
text = message.lower()
|
||||
return any(
|
||||
token in text
|
||||
for token in (
|
||||
"no cuts to make",
|
||||
"no silence",
|
||||
"file unchanged",
|
||||
"nothing saved",
|
||||
)
|
||||
)
|
||||
|
||||
def _emit_no_change_or_error(path: str, message: str) -> int:
|
||||
if _is_no_change_message(message):
|
||||
emit({"ok": True, "path": path, "unchanged": True, "message": message})
|
||||
return 0
|
||||
emit({"ok": False, "error": message})
|
||||
return 1
|
||||
|
||||
|
||||
# Serializa a escrita em stdout. A ponte é JSON-lines: dois comandos
|
||||
# escrevendo ao mesmo tempo entrelaçariam documentos e o app leria lixo.
|
||||
_OUT_LOCK = threading.Lock()
|
||||
|
||||
def emit(obj: Any) -> None:
|
||||
sys.stdout.write(json.dumps(obj, ensure_ascii=False) + "\n")
|
||||
sys.stdout.flush()
|
||||
|
||||
def _transcript_json_path(media_path: str, output_dir: str = "") -> Path:
|
||||
"""Where the ``_transcript.json`` for ``media_path`` lives.
|
||||
|
||||
When ``output_dir`` (the user-selected project folder) is set, the
|
||||
transcript is saved/read there — never next to the source media, which
|
||||
may sit on a read-only volume or a Final Cut Library the user never
|
||||
browses. Falls back to the media's own folder only when no project
|
||||
folder has been chosen (legacy/MCP callers).
|
||||
"""
|
||||
p = Path(media_path)
|
||||
if output_dir:
|
||||
directory = Path(output_dir).expanduser()
|
||||
directory.mkdir(parents=True, exist_ok=True)
|
||||
return directory / f"{p.stem}_transcript.json"
|
||||
return p.with_name(p.stem + "_transcript.json")
|
||||
|
||||
def _save_json_atomic(path: Path, data: Any) -> None:
|
||||
"""Write ``data`` to ``path`` atomically and validate the result on disk.
|
||||
|
||||
Mirrors the reference WHISPERX save path: write a ``.tmp``, ``os.replace``
|
||||
into place, then confirm the file exists, is non-empty, and parses as JSON.
|
||||
"""
|
||||
tmp_path = str(path) + ".tmp"
|
||||
with open(tmp_path, "w", encoding="utf-8") as fh:
|
||||
json.dump(data, fh, ensure_ascii=False, indent=2)
|
||||
os.replace(tmp_path, path)
|
||||
if not path.exists() or os.path.getsize(path) == 0:
|
||||
raise RuntimeError("O arquivo salvo está vazio ou não foi encontrado.")
|
||||
with open(path, encoding="utf-8") as fh:
|
||||
json.load(fh)
|
||||
|
||||
def _project_media_paths(path: str) -> list[str]:
|
||||
proj = parse_fcpxml(path)
|
||||
tl = proj.primary_timeline or (proj.timelines[0] if proj.timelines else None)
|
||||
media_paths: list[str] = []
|
||||
if tl is not None:
|
||||
for clip in getattr(tl, "clips", []):
|
||||
mp = media_src_to_path(clip.media_path or "")
|
||||
if mp and Path(mp).is_file() and mp not in media_paths:
|
||||
media_paths.append(mp)
|
||||
return media_paths
|
||||
|
||||
def _project_media_rotations(path: str) -> dict[str, float]:
|
||||
"""Degrees each source media was rotated by via a Transform filter on its
|
||||
clip in the FCPXML — keyed by the same resolved media path
|
||||
``_project_media_paths`` returns, so the two can be joined by media_path."""
|
||||
proj = parse_fcpxml(path)
|
||||
tl = proj.primary_timeline or (proj.timelines[0] if proj.timelines else None)
|
||||
rotations: dict[str, float] = {}
|
||||
if tl is not None:
|
||||
for clip in getattr(tl, "clips", []):
|
||||
mp = media_src_to_path(clip.media_path or "")
|
||||
if mp and clip.rotation:
|
||||
rotations[mp] = clip.rotation
|
||||
return rotations
|
||||
|
||||
def _voice_timeline_json_path(media_path: str, output_dir: str = "") -> Path:
|
||||
p = Path(media_path)
|
||||
if output_dir:
|
||||
directory = Path(output_dir).expanduser()
|
||||
directory.mkdir(parents=True, exist_ok=True)
|
||||
return directory / f"{p.stem}_voice_timeline.json"
|
||||
return p.with_name(p.stem + "_voice_timeline.json")
|
||||
|
||||
def _load_cached_voice_timeline(json_path: Path, media_path: str) -> dict | None:
|
||||
try:
|
||||
with open(json_path, encoding="utf-8") as fh:
|
||||
data = json.load(fh)
|
||||
except (OSError, json.JSONDecodeError, UnicodeDecodeError):
|
||||
return None
|
||||
if not isinstance(data, dict):
|
||||
return None
|
||||
if data.get("source") != Path(media_path).name:
|
||||
return None
|
||||
if not isinstance(data.get("segments"), list):
|
||||
return None
|
||||
return data
|
||||
|
||||
def _load_cached_transcript(json_path: Path) -> dict | None:
|
||||
"""Return a valid cached transcript dict, or ``None`` if absent/unreadable."""
|
||||
if not json_path.is_file():
|
||||
return None
|
||||
try:
|
||||
data = json.loads(json_path.read_text(encoding="utf-8"))
|
||||
except (OSError, ValueError):
|
||||
return None
|
||||
if isinstance(data, dict) and isinstance(data.get("words"), list):
|
||||
if "speakers" not in data:
|
||||
data["speakers"] = build_speakers(data.get("segments", []))
|
||||
return data
|
||||
return None
|
||||
@@ -0,0 +1,234 @@
|
||||
"""Legendas: dinâmicas, comuns, SRT e as configurações de estilo.
|
||||
|
||||
Extraído de models_api.py — a tabela de comandos segue lá.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import asyncio
|
||||
from pathlib import Path
|
||||
|
||||
from fcpxml.media_intel import media_src_to_path
|
||||
from fcpxml.model_manager import (
|
||||
load_dynamic_subtitle_config,
|
||||
load_plain_subtitle_config,
|
||||
save_dynamic_subtitle_config,
|
||||
save_plain_subtitle_config,
|
||||
)
|
||||
from fcpxml.writer import FCPXMLModifier
|
||||
|
||||
from . import shared
|
||||
from .shared import (
|
||||
_derived_output,
|
||||
_emit_no_change_or_error,
|
||||
_load_cached_transcript,
|
||||
_transcript_json_path,
|
||||
)
|
||||
|
||||
|
||||
def cmd_generate_dynamic_subtitles(args: dict) -> int:
|
||||
"""Generate word-by-word ("karaoke") caption compound clips, one per line,
|
||||
using each media's cached transcript."""
|
||||
path = str(args.get("path", ""))
|
||||
if not path or not Path(path).exists():
|
||||
shared.emit({"ok": False, "error": "Arquivo de projeto não encontrado."})
|
||||
return 1
|
||||
try:
|
||||
from server import handle_generate_dynamic_subtitles
|
||||
|
||||
output = _derived_output(path, "_dynamic_subtitles", args)
|
||||
contents = asyncio.run(
|
||||
handle_generate_dynamic_subtitles({**args, "filepath": path, "output_path": output})
|
||||
)
|
||||
message = "\n".join(getattr(content, "text", str(content)) for content in contents)
|
||||
if not Path(output).exists():
|
||||
shared.emit({"ok": False, "error": message})
|
||||
return 1
|
||||
shared.emit({"ok": True, "path": output, "message": message})
|
||||
return 0
|
||||
except Exception as exc:
|
||||
shared.emit({"ok": False, "error": str(exc)})
|
||||
return 1
|
||||
|
||||
def cmd_generate_plain_subtitles(args: dict) -> int:
|
||||
"""Generate simple static editable subtitle title clips."""
|
||||
path = str(args.get("path", ""))
|
||||
if not path or not Path(path).exists():
|
||||
shared.emit({"ok": False, "error": "Arquivo de projeto não encontrado."})
|
||||
return 1
|
||||
try:
|
||||
from server import handle_generate_plain_subtitles
|
||||
|
||||
output = _derived_output(path, "_plain_subtitles", args)
|
||||
contents = asyncio.run(
|
||||
handle_generate_plain_subtitles({**args, "filepath": path, "output_path": output})
|
||||
)
|
||||
message = "\n".join(getattr(content, "text", str(content)) for content in contents)
|
||||
if not Path(output).exists():
|
||||
return _emit_no_change_or_error(path, message)
|
||||
shared.emit({"ok": True, "path": output, "message": message})
|
||||
return 0
|
||||
except Exception as exc:
|
||||
shared.emit({"ok": False, "error": str(exc)})
|
||||
return 1
|
||||
|
||||
def cmd_export_srt(args: dict) -> int:
|
||||
"""Write a captions .srt synced to the edited timeline.
|
||||
|
||||
Each transcribed segment is mapped from its SOURCE-media timestamp to its
|
||||
real TIMELINE position (``clip_offset + (seg_start - clip_source_start)``),
|
||||
so captions only cover the frames that remain after cuts/silence removal —
|
||||
not the whole source file. One .srt is produced per media, in timeline order.
|
||||
"""
|
||||
path = str(args.get("path", ""))
|
||||
output_dir = str(args.get("output_dir", "")).strip()
|
||||
if not path or not Path(path).exists():
|
||||
shared.emit({"ok": False, "error": "Arquivo de projeto não encontrado."})
|
||||
return 1
|
||||
try:
|
||||
modifier = FCPXMLModifier(path)
|
||||
except Exception as exc:
|
||||
shared.emit({"ok": False, "error": f"Erro ao ler o projeto: {exc}"})
|
||||
return 1
|
||||
|
||||
# Group spine clips by media so each transcript is loaded once.
|
||||
by_media: dict[str, list] = {}
|
||||
for _, el in modifier._iter_spine_clips():
|
||||
src = modifier.resources.get(el.get("ref", ""), {}).get("src", "")
|
||||
mp = media_src_to_path(src)
|
||||
if not mp or not Path(mp).is_file():
|
||||
continue
|
||||
by_media.setdefault(mp, []).append(el)
|
||||
|
||||
# Never emit a caption past the end of the project — Final Cut rejects an
|
||||
# SRT whose last cue overruns the timeline ("subtitle extends beyond project
|
||||
# duration"). Clamp every mapped cue end to this ceiling.
|
||||
timeline_total = modifier._timeline_duration().to_seconds()
|
||||
|
||||
srt_paths: list[str] = []
|
||||
for mp, clips in by_media.items():
|
||||
cached = _load_cached_transcript(_transcript_json_path(mp, output_dir))
|
||||
if cached is None:
|
||||
continue
|
||||
segments = cached.get("segments") or []
|
||||
if not segments:
|
||||
continue
|
||||
|
||||
rows: list[tuple[float, float, str, int]] = []
|
||||
for el in clips:
|
||||
clip_source_start = modifier.source_file_start(el).to_seconds()
|
||||
clip_duration = modifier._parse_time(el.get("duration", "0s")).to_seconds()
|
||||
clip_offset = modifier._parse_time(el.get("offset", "0s")).to_seconds()
|
||||
window_end = clip_source_start + clip_duration
|
||||
for seg_index, seg in enumerate(segments):
|
||||
seg_start = float(seg.get("start", 0.0))
|
||||
seg_end = float(seg.get("end", seg_start))
|
||||
text = seg.get("text", "").strip()
|
||||
if not text or seg_end <= seg_start:
|
||||
continue
|
||||
# Intersect the complete source segment with this kept clip.
|
||||
# Testing only seg_start loses speech whose first words fall in
|
||||
# a removed range; interval intersection preserves the part
|
||||
# that remains and avoids duplicating a segment wholesale.
|
||||
source_start = max(seg_start, clip_source_start)
|
||||
source_end = min(seg_end, window_end)
|
||||
if source_end <= source_start:
|
||||
continue
|
||||
tl_start = clip_offset + (source_start - clip_source_start)
|
||||
tl_end = clip_offset + (source_end - clip_source_start)
|
||||
tl_start = max(0.0, min(tl_start, timeline_total))
|
||||
tl_end = max(0.0, min(tl_end, timeline_total))
|
||||
if tl_end > tl_start:
|
||||
rows.append((tl_start, tl_end, text, seg_index))
|
||||
|
||||
if not rows:
|
||||
continue
|
||||
rows.sort(key=lambda r: (r[0], r[1], r[3]))
|
||||
# Merge only pieces from the same original Whisper segment when their
|
||||
# mapped intervals touch. Never merge unrelated speech or invent time.
|
||||
merged: list[tuple[float, float, str, int]] = []
|
||||
for row in rows:
|
||||
if merged and row[3] == merged[-1][3] and row[0] <= merged[-1][1] + 0.001:
|
||||
prev = merged[-1]
|
||||
merged[-1] = (prev[0], max(prev[1], row[1]), prev[2], prev[3])
|
||||
else:
|
||||
merged.append(row)
|
||||
|
||||
blocks = []
|
||||
for index, (s, e, text, _) in enumerate(merged, 1):
|
||||
start_stamp = srt_stamp(s)
|
||||
end_stamp = srt_stamp(e)
|
||||
# Millisecond SRT precision can collapse a sub-millisecond span;
|
||||
# omit it rather than emit an invalid zero-duration cue.
|
||||
if start_stamp == end_stamp:
|
||||
continue
|
||||
blocks.append(f"{index}\n{start_stamp} --> {end_stamp}\n{text}\n")
|
||||
if not blocks:
|
||||
continue
|
||||
|
||||
out = (
|
||||
Path(output_dir).expanduser() / f"{Path(mp).stem}_captions.srt"
|
||||
if output_dir
|
||||
else Path(mp).with_name(Path(mp).stem + "_captions.srt")
|
||||
)
|
||||
if output_dir:
|
||||
out.parent.mkdir(parents=True, exist_ok=True)
|
||||
try:
|
||||
out.write_text("\n".join(blocks), encoding="utf-8")
|
||||
except OSError as exc:
|
||||
shared.emit({"ok": False, "error": f"Não foi possível salvar a legenda: {exc}"})
|
||||
return 1
|
||||
srt_paths.append(str(out))
|
||||
|
||||
if not srt_paths:
|
||||
shared.emit({"ok": False, "error": "Nenhuma transcrição encontrada. Transcreva o projeto primeiro."})
|
||||
return 1
|
||||
|
||||
shared.emit({"ok": True, "paths": srt_paths, "message": f"{len(srt_paths)} legenda(s) .srt sincronizada(s) com o corte."})
|
||||
return 0
|
||||
|
||||
def srt_stamp(seconds: float) -> str:
|
||||
"""Format float seconds as ``HH:MM:SS,mmm`` (SRT uses a comma).
|
||||
|
||||
Uses ``floor`` (not ``round``) so a timestamp never rounds up past a frame
|
||||
boundary — an SRT cue ending on the last frame must not overrun the
|
||||
project duration, or Final Cut flags it as extending beyond the project.
|
||||
"""
|
||||
ms = int((seconds if seconds > 0 else 0.0) * 1000)
|
||||
h, rem = divmod(ms, 3600000)
|
||||
m, rem = divmod(rem, 60000)
|
||||
s, ms = divmod(rem, 1000)
|
||||
return f"{h:02d}:{m:02d}:{s:02d},{ms:03d}"
|
||||
|
||||
def cmd_dynamic_subtitle_config(args: dict) -> int:
|
||||
"""Read the persisted dynamic-subtitle style (font, size, color, layout)."""
|
||||
shared.emit({"ok": True, **load_dynamic_subtitle_config()})
|
||||
return 0
|
||||
|
||||
def cmd_set_dynamic_subtitle_config(args: dict) -> int:
|
||||
"""Persist dynamic-subtitle style fields. Only the given fields change."""
|
||||
config = save_dynamic_subtitle_config(**{
|
||||
k: args.get(k) for k in (
|
||||
"band_height", "block_center_y", "line_gap", "font", "font_size",
|
||||
"emphasis_font", "emphasis_face", "emphasis_size",
|
||||
"active_color", "emphasis_color", "text_scale",
|
||||
)
|
||||
})
|
||||
shared.emit({"ok": True, **config})
|
||||
return 0
|
||||
|
||||
def cmd_plain_subtitle_config(args: dict) -> int:
|
||||
"""Read the persisted simple subtitle style."""
|
||||
shared.emit({"ok": True, **load_plain_subtitle_config()})
|
||||
return 0
|
||||
|
||||
def cmd_set_plain_subtitle_config(args: dict) -> int:
|
||||
"""Persist simple subtitle style fields. Only the given fields change."""
|
||||
config = save_plain_subtitle_config(**{
|
||||
k: args.get(k) for k in (
|
||||
"font", "font_size", "font_color", "max_words",
|
||||
"position_y", "uppercase", "keep_punctuation", "text_scale",
|
||||
)
|
||||
})
|
||||
shared.emit({"ok": True, **config})
|
||||
return 0
|
||||
@@ -0,0 +1,185 @@
|
||||
"""Transcrição e locutores.
|
||||
|
||||
Extraído de models_api.py — a tabela de comandos segue lá.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import json
|
||||
from pathlib import Path
|
||||
|
||||
from fcpxml.diarize import (
|
||||
assign_speakers,
|
||||
build_speakers,
|
||||
diarization_capability,
|
||||
diarize,
|
||||
)
|
||||
from fcpxml.media_intel import media_src_to_path
|
||||
from fcpxml.model_manager import (
|
||||
is_model_downloaded,
|
||||
load_hf_token,
|
||||
load_num_speakers,
|
||||
load_selected_model,
|
||||
load_transcript_language,
|
||||
save_hf_token,
|
||||
save_num_speakers,
|
||||
)
|
||||
from fcpxml.parser import parse_fcpxml
|
||||
from fcpxml.transcribe import transcribe
|
||||
|
||||
from . import shared
|
||||
from .shared import (
|
||||
_load_cached_transcript,
|
||||
_save_json_atomic,
|
||||
_transcript_json_path,
|
||||
)
|
||||
|
||||
|
||||
def cmd_transcribe(args: dict) -> int:
|
||||
proj_path = str(args.get("path", ""))
|
||||
output_dir = str(args.get("output_dir", "")).strip()
|
||||
# Honra o modelo selecionado no programa quando nenhum é passado.
|
||||
model = str(args.get("model", "") or load_selected_model() or "")
|
||||
language = args.get("language")
|
||||
if language is None:
|
||||
language = load_transcript_language()
|
||||
if language == "auto":
|
||||
language = None
|
||||
if not proj_path:
|
||||
shared.emit({"type": "error", "message": "Nenhum projeto selecionado."})
|
||||
return 1
|
||||
if not output_dir:
|
||||
shared.emit({"type": "error", "message": "Selecione a pasta do projeto antes de transcrever."})
|
||||
return 1
|
||||
if not model or not is_model_downloaded(model):
|
||||
shared.emit(
|
||||
{
|
||||
"type": "error",
|
||||
"message": "Nenhum modelo de transcrição instalado. Baixe e selecione um modelo na aba Modelos.",
|
||||
}
|
||||
)
|
||||
return 1
|
||||
|
||||
token = str(args.get("hf_token") or load_hf_token() or "")
|
||||
if args.get("num_speakers") is not None:
|
||||
num_speakers = str(args.get("num_speakers"))
|
||||
else:
|
||||
num_speakers = load_num_speakers()
|
||||
|
||||
# Load project.
|
||||
try:
|
||||
proj = parse_fcpxml(proj_path)
|
||||
except Exception as exc:
|
||||
shared.emit({"type": "error", "message": f"Erro ao ler o projeto: {exc}"})
|
||||
return 1
|
||||
tl = proj.primary_timeline or (proj.timelines[0] if proj.timelines else None)
|
||||
media_paths: list[str] = []
|
||||
if tl is not None:
|
||||
for clip in getattr(tl, "clips", []):
|
||||
mp = media_src_to_path(clip.media_path or "")
|
||||
if mp and Path(mp).is_file() and mp not in media_paths:
|
||||
media_paths.append(mp)
|
||||
if not media_paths:
|
||||
shared.emit({"type": "error", "message": "Nenhum arquivo de mídia acessível encontrado."})
|
||||
return 1
|
||||
|
||||
total = len(media_paths)
|
||||
results: list[dict] = []
|
||||
for i, mp in enumerate(media_paths, 1):
|
||||
stage = f"Transcrevendo {Path(mp).name} ({i}/{total})…"
|
||||
shared.emit({"type": "progress", "fraction": (i - 1) / total, "stage": stage})
|
||||
json_path = _transcript_json_path(mp, output_dir)
|
||||
cached = _load_cached_transcript(json_path)
|
||||
if cached is not None:
|
||||
shared.emit({"type": "progress", "fraction": i / total, "stage": stage})
|
||||
results.append(_result_row(mp, cached))
|
||||
continue
|
||||
|
||||
def _on_progress(file_fraction: float, _i: int = i, _stage: str = stage) -> None:
|
||||
# Blend this file's own progress into the overall fraction so a
|
||||
# single-media project doesn't jump straight to 100% before the
|
||||
# actual (slow) decoding work has even started.
|
||||
overall = (_i - 1 + file_fraction) / total
|
||||
shared.emit({"type": "progress", "fraction": overall, "stage": _stage})
|
||||
|
||||
data = transcribe(mp, model_size=model, language=language, progress_cb=_on_progress)
|
||||
if data is None:
|
||||
shared.emit({"type": "error", "message": f"Não foi possível transcrever: {Path(mp).name}"})
|
||||
return 1
|
||||
|
||||
# Diarização opcional (necessita token HF): assina speaker por segmento/palavra.
|
||||
if token:
|
||||
tracks = diarize(mp, token, num_speakers)
|
||||
segments, words = assign_speakers(
|
||||
data.get("segments", []), data.get("words", []), tracks
|
||||
)
|
||||
data = {**data, "segments": segments, "words": words}
|
||||
data["speakers"] = build_speakers(data.get("segments", []))
|
||||
|
||||
payload = {
|
||||
"schema_version": "1.0",
|
||||
"source": Path(mp).name,
|
||||
"model": model,
|
||||
**data,
|
||||
}
|
||||
try:
|
||||
_save_json_atomic(json_path, payload)
|
||||
except (OSError, RuntimeError, ValueError) as exc:
|
||||
shared.emit({"type": "error", "message": f"Não foi possível salvar o JSON: {exc}"})
|
||||
return 1
|
||||
results.append(_result_row(mp, data))
|
||||
|
||||
shared.emit({"type": "result", "transcripts": results})
|
||||
return 0
|
||||
|
||||
|
||||
def cmd_rename_speakers(args: dict) -> int:
|
||||
"""Apply real names to speakers already saved in a transcript JSON."""
|
||||
json_path = Path(str(args.get("path", "")))
|
||||
names = args.get("speakers") or {}
|
||||
if not json_path.is_file():
|
||||
shared.emit({"type": "error", "message": "Transcrição não encontrada."})
|
||||
return 1
|
||||
try:
|
||||
data = json.loads(json_path.read_text(encoding="utf-8"))
|
||||
except (OSError, ValueError) as exc:
|
||||
shared.emit({"type": "error", "message": f"Não foi possível ler o JSON: {exc}"})
|
||||
return 1
|
||||
mapping = {str(sid): str(name).strip() for sid, name in (names or {}).items()}
|
||||
for sp in data.get("speakers", []):
|
||||
sid = str(sp.get("id", ""))
|
||||
if mapping.get(sid):
|
||||
sp["name"] = mapping[sid]
|
||||
try:
|
||||
_save_json_atomic(json_path, data)
|
||||
except (OSError, RuntimeError, ValueError) as exc:
|
||||
shared.emit({"type": "error", "message": f"Não foi possível salvar: {exc}"})
|
||||
return 1
|
||||
shared.emit({"ok": True, "speakers": data.get("speakers", [])})
|
||||
return 0
|
||||
|
||||
def cmd_set_diarization(args: dict) -> int:
|
||||
"""Persist the HuggingFace token and expected speaker count for diarization."""
|
||||
token = args.get("token")
|
||||
num = args.get("num_speakers")
|
||||
if token is not None:
|
||||
save_hf_token(str(token))
|
||||
if num is not None:
|
||||
save_num_speakers(str(num))
|
||||
ok, msg = diarization_capability(load_hf_token())
|
||||
shared.emit({"ok": True, "diarization": ok, "diarization_message": msg, "num_speakers": load_num_speakers()})
|
||||
return 0
|
||||
|
||||
def _result_row(mp: str, data: dict) -> dict:
|
||||
words = data.get("words", [])
|
||||
preview = (data.get("text", "") or "")[:160]
|
||||
speakers = data.get("speakers") or []
|
||||
return {
|
||||
"media": Path(mp).name,
|
||||
"language": data.get("language", "?"),
|
||||
"words": len(words),
|
||||
"duration": float(data.get("duration", 0.0)),
|
||||
"preview": preview,
|
||||
"saved": str(_transcript_json_path(mp)),
|
||||
"speakers": [s.get("name", s.get("id", "")) for s in speakers],
|
||||
}
|
||||
@@ -0,0 +1,300 @@
|
||||
"""Análise de voz e aplicação das decisões de edição.
|
||||
|
||||
Extraído de models_api.py — a tabela de comandos segue lá.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import asyncio
|
||||
import json
|
||||
import re
|
||||
from pathlib import Path
|
||||
|
||||
from fcpxml.model_manager import (
|
||||
load_hf_token,
|
||||
load_num_speakers,
|
||||
load_selected_model,
|
||||
load_transcript_language,
|
||||
load_voice_analysis_config,
|
||||
save_voice_analysis_config,
|
||||
)
|
||||
|
||||
from . import shared
|
||||
from .shared import (
|
||||
_load_cached_transcript,
|
||||
_load_cached_voice_timeline,
|
||||
_project_media_paths,
|
||||
_project_media_rotations,
|
||||
_transcript_json_path,
|
||||
_voice_timeline_json_path,
|
||||
)
|
||||
|
||||
|
||||
def cmd_analyze_voice(args: dict) -> int:
|
||||
"""Build the voice timeline (transcript+diarization+acoustics -> emphasis)
|
||||
for every unique source media in the project, so `refine_voice_timeline`
|
||||
and friends have something to read without ever reopening the audio.
|
||||
|
||||
Analysis only — writes _voice_timeline.json next to each media, doesn't
|
||||
touch the project XML. `path` passes through unchanged so it composes
|
||||
with the other batch steps (silence removal, captions) regardless of
|
||||
where in the list it runs.
|
||||
"""
|
||||
path = str(args.get("path", ""))
|
||||
if not path or not Path(path).exists():
|
||||
shared.emit({"ok": False, "error": "Arquivo de projeto não encontrado."})
|
||||
return 1
|
||||
|
||||
model = str(args.get("model", "") or load_selected_model() or "")
|
||||
language = args.get("language")
|
||||
if language is None:
|
||||
language = load_transcript_language()
|
||||
if language == "auto":
|
||||
language = None
|
||||
token = str(args.get("hf_token") or load_hf_token() or "")
|
||||
num_speakers = str(args.get("num_speakers") or load_num_speakers() or "")
|
||||
|
||||
try:
|
||||
media_paths = _project_media_paths(path)
|
||||
rotations = _project_media_rotations(path)
|
||||
except Exception as exc:
|
||||
shared.emit({"ok": False, "error": f"Erro ao ler o projeto: {exc}"})
|
||||
return 1
|
||||
if not media_paths:
|
||||
shared.emit({"ok": False, "error": "Nenhum arquivo de mídia acessível encontrado."})
|
||||
return 1
|
||||
|
||||
from server import handle_build_voice_timeline
|
||||
|
||||
messages: list[str] = []
|
||||
output_dir = str(args.get("output_dir") or "").strip()
|
||||
existing: list[Path] = []
|
||||
for mp in media_paths:
|
||||
timeline_path = _voice_timeline_json_path(mp, output_dir)
|
||||
if _load_cached_voice_timeline(timeline_path, mp) is not None:
|
||||
existing.append(timeline_path)
|
||||
if existing and len(existing) == len(media_paths) and not bool(args.get("force_reprocess", False)):
|
||||
message = "# Voice Timeline Cache\n\n"
|
||||
message += "Reaproveitando análise de voz existente. Nada foi reprocessado.\n\n"
|
||||
for timeline_path in existing:
|
||||
message += f"- **Timeline JSON**: {timeline_path}\n"
|
||||
shared.emit({
|
||||
"ok": True,
|
||||
"path": path,
|
||||
"reused": True,
|
||||
"timelines": [str(p) for p in existing],
|
||||
"message": message,
|
||||
})
|
||||
return 0
|
||||
|
||||
for mp in media_paths:
|
||||
transcript_path = _transcript_json_path(mp, output_dir)
|
||||
reused_prefix = ""
|
||||
if _load_cached_transcript(transcript_path) is not None:
|
||||
reused_prefix = f"# Cache\n\nReaproveitando transcrição existente: `{transcript_path}`\n\n"
|
||||
try:
|
||||
contents = asyncio.run(handle_build_voice_timeline({
|
||||
"media_path": mp, "model": model, "language": language,
|
||||
"hf_token": token, "num_speakers": num_speakers,
|
||||
"output_dir": output_dir, "rotation": rotations.get(mp, 0.0),
|
||||
}))
|
||||
except Exception as exc:
|
||||
shared.emit({"ok": False, "error": f"Falha analisando {Path(mp).name}: {exc}"})
|
||||
return 1
|
||||
messages.append(reused_prefix + "\n".join(getattr(c, "text", str(c)) for c in contents))
|
||||
|
||||
shared.emit({"ok": True, "path": path, "message": "\n\n---\n\n".join(messages)})
|
||||
return 0
|
||||
|
||||
def cmd_acoustics_capability(args: dict) -> int:
|
||||
"""Whether librosa (pitch/energy extraction) is installed in this venv.
|
||||
|
||||
Surfaces `features_capability()` — previously computed but never
|
||||
exposed to the app, so `layers.acoustics: false` in a voice timeline
|
||||
had no explanation the user could act on.
|
||||
"""
|
||||
from fcpxml.voice_features import features_capability
|
||||
ok, msg = features_capability()
|
||||
shared.emit({"ok": True, "available": ok, "message": msg})
|
||||
return 0
|
||||
|
||||
def cmd_voice_analysis(args: dict) -> int:
|
||||
"""Read the persisted voice-analysis settings (energy/emphasis/emotion)."""
|
||||
config = load_voice_analysis_config()
|
||||
shared.emit({"ok": True, **config, "emphasis_threshold": config["emphasis_floor"]})
|
||||
return 0
|
||||
|
||||
def cmd_set_voice_analysis(args: dict) -> int:
|
||||
"""Persist voice-analysis settings. Only the given fields change."""
|
||||
weights = args.get("emphasis_weights")
|
||||
config = save_voice_analysis_config(
|
||||
energy_threshold=args.get("energy_threshold"),
|
||||
emphasis_weights=weights if isinstance(weights, dict) else None,
|
||||
emphasis_floor=args.get("emphasis_threshold"),
|
||||
emotion_enabled=args.get("emotion_enabled"),
|
||||
emotion_sensitivity=args.get("emotion_sensitivity"),
|
||||
zoom_scale=args.get("zoom_scale"),
|
||||
zoom_mode=args.get("zoom_mode"),
|
||||
zoom_ease_in=args.get("zoom_ease_in"),
|
||||
zoom_ease_out=args.get("zoom_ease_out"),
|
||||
)
|
||||
shared.emit({"ok": True, **config})
|
||||
return 0
|
||||
|
||||
def cmd_apply_voice_actions(args: dict) -> int:
|
||||
"""Apply a decision list (cuts/zooms/texts/markers) to the project XML.
|
||||
|
||||
The list is produced by a model reading the _voice_timeline.json — this
|
||||
is the step that turns those decisions into an edit, and the one the
|
||||
batch chain was missing: without it the app could measure the voice and
|
||||
caption the result, but never cut by it.
|
||||
|
||||
`actions_path` points at the JSON; either a bare list or the
|
||||
``{"actions": [...]}`` wrapper the skill emits is accepted. Times stay in
|
||||
ORIGINAL source seconds — the handler resolves cuts first and shifts
|
||||
everything else itself.
|
||||
"""
|
||||
path = str(args.get("path", ""))
|
||||
if not path or not Path(path).exists():
|
||||
shared.emit({"ok": False, "error": "Arquivo de projeto não encontrado."})
|
||||
return 1
|
||||
|
||||
actions = args.get("actions")
|
||||
if actions is None:
|
||||
actions_path = str(args.get("actions_path", ""))
|
||||
if not actions_path or not Path(actions_path).exists():
|
||||
shared.emit({"ok": False, "error": "Arquivo de decisões (JSON) não encontrado."})
|
||||
return 1
|
||||
try:
|
||||
with open(actions_path, encoding="utf-8") as fh:
|
||||
loaded = json.load(fh)
|
||||
except (OSError, ValueError) as exc:
|
||||
shared.emit({"ok": False, "error": f"Erro ao ler as decisões: {exc}"})
|
||||
return 1
|
||||
actions = loaded.get("actions") if isinstance(loaded, dict) else loaded
|
||||
|
||||
# The documented output format is {"source": ..., "actions": [...]} —
|
||||
# callers passing that whole object inline (e.g. the wizard pasting the
|
||||
# skill's JSON verbatim) need the same unwrap the actions_path branch
|
||||
# above already does, or a well-formed payload gets rejected as
|
||||
# "malformed" for having one extra layer of nesting.
|
||||
if isinstance(actions, dict):
|
||||
actions = actions.get("actions")
|
||||
|
||||
if not isinstance(actions, list) or not actions:
|
||||
shared.emit({"ok": False, "error": "A lista de decisões está vazia ou malformada."})
|
||||
return 1
|
||||
|
||||
from server import handle_apply_voice_actions
|
||||
|
||||
try:
|
||||
contents = asyncio.run(handle_apply_voice_actions({
|
||||
"filepath": path,
|
||||
"actions": actions,
|
||||
"output_dir": args.get("output_dir"),
|
||||
}))
|
||||
except Exception as exc:
|
||||
shared.emit({"ok": False, "error": f"Falha ao aplicar as decisões: {exc}"})
|
||||
return 1
|
||||
|
||||
message = "\n".join(getattr(c, "text", str(c)) for c in contents)
|
||||
# The handler reports dropped/rejected actions individually; hand the
|
||||
# whole report back so the app can surface them instead of only the count.
|
||||
out_path = path
|
||||
for line in message.splitlines():
|
||||
if line.startswith("- **Saved to**:"):
|
||||
out_path = line.split("`")[1] if "`" in line else path
|
||||
break
|
||||
shared.emit({"ok": True, "path": out_path, "message": message})
|
||||
return 0
|
||||
|
||||
|
||||
def cmd_generate_voice_script(args: dict) -> int:
|
||||
"""Run the ENTIRE voice-edit pass against a LOCAL model, inside the engine.
|
||||
|
||||
Transcribe (cached) -> build the voice timeline -> hand it to a local
|
||||
Ollama model (Gemma 3 / Llama) that directs the edit -> return the readable
|
||||
script (roteiro) and the action JSON, and optionally apply to a FCPXML. No
|
||||
wizard, no copy-paste: the model's decisions are validated and applied by
|
||||
the same pipeline the rules engine uses.
|
||||
|
||||
Args (all optional except one of ``media_path`` / ``voice_timeline``):
|
||||
media_path audio/video to analyze and direct (required when there is
|
||||
no voice_timeline yet)
|
||||
voice_timeline path to an existing _voice_timeline.json; when given the
|
||||
analysis is reused and media_path is not required
|
||||
filepath optional FCPXML to apply the decisions to (non-destructive)
|
||||
model local model Ollama serves (default gemma3:12b)
|
||||
base_url Ollama base URL (default http://localhost:11434)
|
||||
model_size whisper size if transcription is needed
|
||||
language ISO language hint for transcription
|
||||
hf_token HuggingFace token for diarization
|
||||
num_speakers known speaker count, if any
|
||||
output_dir folder for the timeline/review/actions JSON
|
||||
apply_to_fcpxml apply to filepath when given (default true)
|
||||
-> {"ok": true, "message": "...", "roteiro_path", "actions_path",
|
||||
"applied_path"} or {"ok": false, "error": "..."}
|
||||
"""
|
||||
media_path = str(args.get("media_path", ""))
|
||||
voice_timeline = str(args.get("voice_timeline", ""))
|
||||
if not voice_timeline and (not media_path or not Path(media_path).exists()):
|
||||
shared.emit({"ok": False, "error": "Arquivo de mídia não encontrado (informe media_path ou voice_timeline)."})
|
||||
return 1
|
||||
|
||||
from server import handle_generate_voice_script
|
||||
|
||||
try:
|
||||
contents = asyncio.run(handle_generate_voice_script({
|
||||
"media_path": media_path,
|
||||
"voice_timeline": args.get("voice_timeline"),
|
||||
"filepath": args.get("filepath"),
|
||||
"model": args.get("model"),
|
||||
"base_url": args.get("base_url"),
|
||||
"model_size": args.get("model_size"),
|
||||
"language": args.get("language"),
|
||||
"hf_token": args.get("hf_token"),
|
||||
"num_speakers": args.get("num_speakers"),
|
||||
"output_dir": args.get("output_dir"),
|
||||
"apply_to_fcpxml": args.get("apply_to_fcpxml", True),
|
||||
}))
|
||||
except Exception as exc:
|
||||
shared.emit({"ok": False, "error": f"Falha ao gerar roteiro por IA local: {exc}"})
|
||||
return 1
|
||||
|
||||
message = "\n".join(getattr(c, "text", str(c)) for c in contents)
|
||||
|
||||
def _path_after(label: str) -> str:
|
||||
m = re.search(rf"\*\*{label}\*\*: (.+)", message)
|
||||
return m.group(1).strip() if m else ""
|
||||
|
||||
roteiro_path = _path_after(r"Roteiro \(legível\)")
|
||||
actions_path = _path_after("Ações JSON")
|
||||
applied_path = ""
|
||||
for line in message.splitlines():
|
||||
if line.startswith("- **Saved to**:"):
|
||||
applied_path = line.split("`")[1] if "`" in line else ""
|
||||
break
|
||||
shared.emit({
|
||||
"ok": True,
|
||||
"message": message,
|
||||
"roteiro_path": roteiro_path,
|
||||
"actions_path": actions_path,
|
||||
"applied_path": applied_path,
|
||||
})
|
||||
return 0
|
||||
|
||||
|
||||
def cmd_list_ollama_models(args: dict) -> int:
|
||||
"""List the models Ollama currently serves, for the app's model picker.
|
||||
|
||||
Args:
|
||||
base_url Ollama base URL (default http://localhost:11434)
|
||||
-> {"ok": true, "models": ["gemma3:12b", ...]} (empty list if Ollama
|
||||
is unreachable, so the UI can fall back to a text field)
|
||||
"""
|
||||
from fcpxml.llm_local import list_ollama_models
|
||||
|
||||
base_url = str(args.get("base_url") or "http://localhost:11434")
|
||||
models = list_ollama_models(base_url=base_url)
|
||||
shared.emit({"ok": True, "models": models})
|
||||
return 0
|
||||
@@ -0,0 +1,107 @@
|
||||
"""Zoom (punch-in): por janela, por clipe e por trecho da transcrição.
|
||||
|
||||
Extraído de models_api.py — a tabela de comandos segue lá.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import asyncio
|
||||
from pathlib import Path
|
||||
|
||||
from . import shared
|
||||
from .shared import (
|
||||
_derived_output,
|
||||
_load_cached_transcript,
|
||||
_transcript_json_path,
|
||||
)
|
||||
|
||||
|
||||
def cmd_add_zoom(args: dict) -> int:
|
||||
"""Add an ease-in/ease-out punch-in zoom to one clip."""
|
||||
path = str(args.get("path", ""))
|
||||
clip_id = str(args.get("clip_id", "")).strip()
|
||||
if not path or not Path(path).exists():
|
||||
shared.emit({"ok": False, "error": "Arquivo de projeto não encontrado."})
|
||||
return 1
|
||||
if not clip_id:
|
||||
shared.emit({"ok": False, "error": "Informe o nome do clipe."})
|
||||
return 1
|
||||
try:
|
||||
from server import handle_add_zoom
|
||||
|
||||
output = _derived_output(path, "_zoom", args)
|
||||
contents = asyncio.run(handle_add_zoom({**args, "filepath": path, "output_path": output}))
|
||||
message = "\n".join(getattr(content, "text", str(content)) for content in contents)
|
||||
if not Path(output).exists():
|
||||
shared.emit({"ok": False, "error": message})
|
||||
return 1
|
||||
shared.emit({"ok": True, "path": output, "message": message})
|
||||
return 0
|
||||
except Exception as exc:
|
||||
shared.emit({"ok": False, "error": str(exc)})
|
||||
return 1
|
||||
|
||||
def cmd_zoom_clips(args: dict) -> int:
|
||||
"""Return timeline clips with enough identity for the zoom picker."""
|
||||
path = Path(str(args.get("path", "")))
|
||||
output_dir = str(args.get("output_dir", "")).strip()
|
||||
if not path.exists():
|
||||
shared.emit({"ok": False, "error": "Arquivo de projeto não encontrado."})
|
||||
return 1
|
||||
try:
|
||||
from server import _require_timeline
|
||||
|
||||
_, timeline = _require_timeline(str(path))
|
||||
clips = []
|
||||
for index, clip in enumerate(timeline.clips):
|
||||
media = clip.media_path or ""
|
||||
cached = _load_cached_transcript(_transcript_json_path(media, output_dir)) if media else None
|
||||
clips.append({
|
||||
"id": f"{index}:{clip.start.seconds:.6f}",
|
||||
"index": index,
|
||||
"name": clip.name,
|
||||
"start": clip.start.seconds,
|
||||
"duration": clip.duration_seconds,
|
||||
"media": Path(media).name if media else "",
|
||||
"preview": ((cached or {}).get("text", "") or "")[:180],
|
||||
"has_transcript": cached is not None,
|
||||
})
|
||||
shared.emit({"ok": True, "clips": clips})
|
||||
return 0
|
||||
except Exception as exc:
|
||||
shared.emit({"ok": False, "error": str(exc)})
|
||||
return 1
|
||||
|
||||
def cmd_zoom_segments(args: dict) -> int:
|
||||
"""Return sentence/word ranges for one timeline clip."""
|
||||
path = Path(str(args.get("path", "")))
|
||||
output_dir = str(args.get("output_dir", "")).strip()
|
||||
try:
|
||||
from server import _require_timeline
|
||||
|
||||
_, timeline = _require_timeline(str(path))
|
||||
index = int(args.get("index", -1))
|
||||
if index < 0 or index >= len(timeline.clips):
|
||||
raise ValueError("Clipe selecionado não existe.")
|
||||
clip = timeline.clips[index]
|
||||
if not clip.media_path:
|
||||
raise ValueError("Este clipe não possui mídia associada.")
|
||||
data = _load_cached_transcript(_transcript_json_path(clip.media_path, output_dir))
|
||||
if data is None:
|
||||
shared.emit({"ok": True, "segments": [], "message": "Transcreva este clipe primeiro."})
|
||||
return 0
|
||||
segments = []
|
||||
for number, segment in enumerate(data.get("segments", [])):
|
||||
text = str(segment.get("text", "")).strip()
|
||||
if text:
|
||||
segments.append({
|
||||
"id": number,
|
||||
"start": float(segment.get("start", 0)),
|
||||
"end": float(segment.get("end", 0)),
|
||||
"text": text,
|
||||
})
|
||||
shared.emit({"ok": True, "segments": segments})
|
||||
return 0
|
||||
except Exception as exc:
|
||||
shared.emit({"ok": False, "error": str(exc)})
|
||||
return 1
|
||||
@@ -0,0 +1,12 @@
|
||||
# Copie para admin/gart-rag.env e preencha a senha. Este arquivo é apenas um
|
||||
# modelo; admin/gart-rag.env é ignorado pelo git.
|
||||
RAG_DB_HOST=127.0.0.1
|
||||
RAG_DB_PORT=55435
|
||||
RAG_DB_NAME=rag_gart
|
||||
RAG_DB_SCHEMA=gart
|
||||
RAG_DB_USER=gart_rag_indexer
|
||||
RAG_DB_PASSWORD=
|
||||
|
||||
# Ollama que fornece nomic-embed-text.
|
||||
OLLAMA_URL=http://127.0.0.1:11434
|
||||
RAG_EMBED_MODEL=nomic-embed-text
|
||||
+174
-980
File diff suppressed because it is too large
Load Diff
+3
-4
@@ -20,7 +20,6 @@ import subprocess
|
||||
import sys
|
||||
import threading
|
||||
from pathlib import Path
|
||||
from typing import Optional
|
||||
|
||||
import flet as ft
|
||||
|
||||
@@ -101,9 +100,9 @@ class ModelManagerApp:
|
||||
def __init__(self, page: ft.Page) -> None:
|
||||
self.page = page
|
||||
self.selected = load_selected_model()
|
||||
self.downloading: Optional[str] = None
|
||||
self.downloading: str | None = None
|
||||
self._cancel_events: dict[str, threading.Event] = {}
|
||||
self._picker: Optional[ft.FilePicker] = None
|
||||
self._picker: ft.FilePicker | None = None
|
||||
|
||||
# ── helpers ────────────────────────────────────────────────────────────
|
||||
|
||||
@@ -121,7 +120,7 @@ class ModelManagerApp:
|
||||
self._picker = ft.FilePicker()
|
||||
self._picker.on_result = self._on_file_picked
|
||||
self.page.overlay.append(self._picker)
|
||||
self._pending_target: Optional[dict] = None
|
||||
self._pending_target: dict | None = None
|
||||
|
||||
def _on_file_picked(self, e) -> None:
|
||||
if self._pending_target == "project":
|
||||
|
||||
@@ -1,243 +0,0 @@
|
||||
"""Tests for admin/models_api.py — the SwiftUI JSON bridge commands.
|
||||
|
||||
Focused on the transcription-flow changes: atomic save, speaker renaming, and
|
||||
the "use the selected model" default plus model-availability guard.
|
||||
"""
|
||||
|
||||
import json
|
||||
|
||||
import admin.models_api as api
|
||||
|
||||
|
||||
def _capture(monkeypatch):
|
||||
captured: list[dict] = []
|
||||
|
||||
def _emit(obj):
|
||||
captured.append(obj)
|
||||
|
||||
monkeypatch.setattr(api, "_emit", _emit)
|
||||
return captured
|
||||
|
||||
|
||||
def test_save_json_atomic(tmp_path):
|
||||
p = tmp_path / "t.json"
|
||||
api._save_json_atomic(p, {"a": [1, 2], "text": "olá"})
|
||||
assert p.exists()
|
||||
assert not (tmp_path / "t.json.tmp").exists()
|
||||
assert json.loads(p.read_text(encoding="utf-8"))["text"] == "olá"
|
||||
|
||||
|
||||
def test_rename_speakers(tmp_path, monkeypatch):
|
||||
captured = _capture(monkeypatch)
|
||||
p = tmp_path / "t.json"
|
||||
p.write_text(
|
||||
json.dumps(
|
||||
{
|
||||
"speakers": [
|
||||
{"id": "SPEAKER_00", "name": "Speaker 1"},
|
||||
{"id": "SPEAKER_01", "name": "Speaker 2"},
|
||||
]
|
||||
}
|
||||
),
|
||||
encoding="utf-8",
|
||||
)
|
||||
assert api.cmd_rename_speakers({"path": str(p), "speakers": {"SPEAKER_01": "Erika"}}) == 0
|
||||
assert captured[0]["ok"] is True
|
||||
saved = json.loads(p.read_text(encoding="utf-8"))
|
||||
assert saved["speakers"][0]["name"] == "Speaker 1"
|
||||
assert saved["speakers"][1]["name"] == "Erika"
|
||||
|
||||
|
||||
def test_rename_speakers_missing_file(monkeypatch):
|
||||
captured = _capture(monkeypatch)
|
||||
assert api.cmd_rename_speakers({"path": "/nonexistent/x.json"}) == 1
|
||||
assert captured[0]["type"] == "error"
|
||||
|
||||
|
||||
def test_transcribe_requires_output_dir(monkeypatch):
|
||||
captured = _capture(monkeypatch)
|
||||
monkeypatch.setattr(api, "load_selected_model", lambda: "small")
|
||||
monkeypatch.setattr(api, "is_model_downloaded", lambda m: True)
|
||||
assert api.cmd_transcribe({"path": "/some/project.fcpxml"}) == 1
|
||||
assert captured[0]["type"] == "error"
|
||||
assert "pasta do projeto" in captured[0]["message"]
|
||||
|
||||
|
||||
def test_transcribe_requires_installed_model(monkeypatch, tmp_path):
|
||||
captured = _capture(monkeypatch)
|
||||
monkeypatch.setattr(api, "load_selected_model", lambda: "")
|
||||
monkeypatch.setattr(api, "is_model_downloaded", lambda m: False)
|
||||
assert api.cmd_transcribe({"path": "/some/project.fcpxml", "output_dir": str(tmp_path)}) == 1
|
||||
assert captured[0]["type"] == "error"
|
||||
assert "instalado" in captured[0]["message"]
|
||||
|
||||
|
||||
def test_transcribe_defaults_to_selected_model(monkeypatch, tmp_path):
|
||||
captured = _capture(monkeypatch)
|
||||
monkeypatch.setattr(api, "load_selected_model", lambda: "small")
|
||||
monkeypatch.setattr(api, "is_model_downloaded", lambda m: m == "small")
|
||||
|
||||
class FakeTL:
|
||||
clips = []
|
||||
|
||||
class FakeProject:
|
||||
primary_timeline = None
|
||||
timelines = [FakeTL()]
|
||||
|
||||
monkeypatch.setattr(api, "parse_fcpxml", lambda p: FakeProject())
|
||||
# No media accessible -> reaches the media-path check (past model validation).
|
||||
assert api.cmd_transcribe({"path": "/some/project.fcpxml", "output_dir": str(tmp_path)}) == 1
|
||||
assert captured[0]["type"] == "error"
|
||||
assert "mídia" in captured[0]["message"]
|
||||
|
||||
|
||||
def test_set_language_persists(monkeypatch):
|
||||
captured = _capture(monkeypatch)
|
||||
assert api.cmd_set_language({"language": "pt"}) == 0
|
||||
assert captured[0]["ok"] is True
|
||||
assert captured[0]["language"] == "pt"
|
||||
assert api.load_transcript_language() == "pt"
|
||||
|
||||
|
||||
def test_set_language_rejects_unknown(monkeypatch):
|
||||
captured = _capture(monkeypatch)
|
||||
assert api.cmd_set_language({"language": "xx"}) == 1
|
||||
assert captured[0]["ok"] is False
|
||||
assert "language" in captured[0]["error"]
|
||||
|
||||
|
||||
def test_transcribe_defaults_language_to_persisted(monkeypatch, tmp_path):
|
||||
monkeypatch.setattr(api, "load_selected_model", lambda: "small")
|
||||
monkeypatch.setattr(api, "is_model_downloaded", lambda m: m == "small")
|
||||
monkeypatch.setattr(api, "load_transcript_language", lambda: "pt")
|
||||
|
||||
media = tmp_path / "clip.mov"
|
||||
media.write_bytes(b"fake")
|
||||
|
||||
class FakeClip:
|
||||
media_path = ""
|
||||
|
||||
class FakeTL:
|
||||
clips = [FakeClip()]
|
||||
|
||||
class FakeProject:
|
||||
primary_timeline = None
|
||||
timelines = [FakeTL()]
|
||||
|
||||
monkeypatch.setattr(api, "parse_fcpxml", lambda p: FakeProject())
|
||||
monkeypatch.setattr(api, "media_src_to_path", lambda mp: str(media))
|
||||
called = {}
|
||||
monkeypatch.setattr(
|
||||
api, "transcribe", lambda mp, model_size, language, **kw: called.update(lang=language)
|
||||
)
|
||||
assert api.cmd_transcribe({"path": "/some/project.fcpxml", "output_dir": str(tmp_path / "out")}) == 1
|
||||
assert called["lang"] == "pt"
|
||||
|
||||
|
||||
def test_srt_stamp_format():
|
||||
assert api.srt_stamp(0.0) == "00:00:00,000"
|
||||
assert api.srt_stamp(1.5) == "00:00:01,500"
|
||||
assert api.srt_stamp(3661.234) == "01:01:01,234"
|
||||
|
||||
|
||||
_FCPXML_SAMPLE = """<?xml version="1.0" encoding="UTF-8"?>
|
||||
<fcpxml version="1.13">
|
||||
<resources>
|
||||
<asset id="r1" name="clip" uid="u1" start="0s" duration="100s"
|
||||
hasVideo="1" format="f1" hasAudio="1">
|
||||
<media-rep kind="original-media" src="file:///tmp/clip.mp4"/>
|
||||
</asset>
|
||||
<format id="f1" name="FFVideoFormat1080p25" frameDuration="1/25s" width="1920" height="1080"/>
|
||||
</resources>
|
||||
<library>
|
||||
<event name="Event">
|
||||
<project name="P">
|
||||
<sequence format="f1">
|
||||
<spine>
|
||||
<asset-clip ref="r1" offset="0s" start="10s" duration="10s" name="clip"/>
|
||||
<gap name="Espaço" offset="10s" duration="90s" start="10s"/>
|
||||
</spine>
|
||||
</sequence>
|
||||
</project>
|
||||
</event>
|
||||
</library>
|
||||
</fcpxml>
|
||||
"""
|
||||
|
||||
|
||||
def test_cmd_export_srt_maps_to_edited_timeline(tmp_path, monkeypatch):
|
||||
"""Captions must reflect the EDITED timeline, not the whole source file."""
|
||||
captured = _capture(monkeypatch)
|
||||
project = tmp_path / "proj.fcpxml"
|
||||
project.write_text(_FCPXML_SAMPLE, encoding="utf-8")
|
||||
media = tmp_path / "clip.mp4"
|
||||
media.write_bytes(b"fake")
|
||||
# Transcript covers 0..100s; the clip only USES source 10..20s -> timeline 0..10s.
|
||||
transcript = {
|
||||
"words": [],
|
||||
"segments": [
|
||||
{"start": 5.0, "end": 6.0, "text": "antes do corte"},
|
||||
{"start": 12.0, "end": 14.0, "text": "dentro do corte"},
|
||||
{"start": 50.0, "end": 51.0, "text": "depois do corte"},
|
||||
]
|
||||
}
|
||||
tj = api._transcript_json_path(media)
|
||||
tj.parent.mkdir(parents=True, exist_ok=True)
|
||||
api._save_json_atomic(tj, transcript)
|
||||
monkeypatch.setattr(api, "media_src_to_path", lambda src: str(media))
|
||||
|
||||
assert api.cmd_export_srt({"path": str(project)}) == 0
|
||||
assert captured[0]["ok"] is True
|
||||
srt = tmp_path / "clip_captions.srt"
|
||||
assert srt.exists()
|
||||
text = srt.read_text(encoding="utf-8")
|
||||
# Only the segment inside the used source window (12s) survives.
|
||||
assert "dentro do corte" in text
|
||||
assert "antes do corte" not in text
|
||||
assert "depois do corte" not in text
|
||||
# Mapped to timeline 0..10s -> the 12s source segment lands at 2s.
|
||||
assert "00:00:02,000 --> 00:00:04,000" in text
|
||||
|
||||
|
||||
def test_cmd_export_srt_no_transcript(tmp_path, monkeypatch):
|
||||
captured = _capture(monkeypatch)
|
||||
project = tmp_path / "proj.fcpxml"
|
||||
project.write_text(_FCPXML_SAMPLE, encoding="utf-8")
|
||||
media = tmp_path / "clip.mp4"
|
||||
media.write_bytes(b"fake")
|
||||
monkeypatch.setattr(api, "media_src_to_path", lambda src: str(media))
|
||||
assert api.cmd_export_srt({"path": str(project)}) == 1
|
||||
assert captured[0]["ok"] is False
|
||||
|
||||
|
||||
def test_cmd_export_srt_clamps_past_project_duration(tmp_path, monkeypatch):
|
||||
"""A segment ending after the last clip must be clamped to the project end.
|
||||
|
||||
Final Cut rejects an SRT whose final cue overruns the timeline
|
||||
("subtitle extends beyond project duration").
|
||||
"""
|
||||
captured = _capture(monkeypatch)
|
||||
project = tmp_path / "proj.fcpxml"
|
||||
project.write_text(_FCPXML_SAMPLE, encoding="utf-8")
|
||||
media = tmp_path / "clip.mp4"
|
||||
media.write_bytes(b"fake")
|
||||
# Clip uses source 10..20s -> timeline 0..10s. A segment 12..30s maps to
|
||||
# timeline 2..20s, but the project only lasts 10s: must clamp end to 10s.
|
||||
transcript = {
|
||||
"words": [],
|
||||
"segments": [
|
||||
{"start": 12.0, "end": 30.0, "text": "longa fala"},
|
||||
]
|
||||
}
|
||||
tj = api._transcript_json_path(media)
|
||||
tj.parent.mkdir(parents=True, exist_ok=True)
|
||||
api._save_json_atomic(tj, transcript)
|
||||
monkeypatch.setattr(api, "media_src_to_path", lambda src: str(media))
|
||||
|
||||
assert api.cmd_export_srt({"path": str(project)}) == 0
|
||||
assert captured[0]["ok"] is True
|
||||
srt = tmp_path / "clip_captions.srt"
|
||||
text = srt.read_text(encoding="utf-8")
|
||||
# Timeline is 10s; the cue must not end past it.
|
||||
assert "00:00:02,000 --> 00:00:10,000" in text
|
||||
assert "00:00:20,000" not in text
|
||||
Executable
+68
@@ -0,0 +1,68 @@
|
||||
#!/bin/zsh
|
||||
# Atualiza incrementalmente a RAG do G-ART usando o banco compartilhado.
|
||||
#
|
||||
# Credenciais: defina RAG_DB_PASSWORD no ambiente ou crie
|
||||
# admin/gart-rag.env (ignorado pelo git). O arquivo pode conter também
|
||||
# RAG_DB_USER, RAG_DB_PORT, OLLAMA_URL e RAG_EMBED_MODEL.
|
||||
set -euo pipefail
|
||||
|
||||
ROOT="$(cd "$(dirname "$0")/.." && pwd -P)"
|
||||
ENV_FILE="$ROOT/admin/gart-rag.env"
|
||||
if [[ -f "$ENV_FILE" ]]; then
|
||||
set -a
|
||||
source "$ENV_FILE"
|
||||
set +a
|
||||
fi
|
||||
|
||||
PYTHON="${RAG_PYTHON:-}"
|
||||
if [[ -z "$PYTHON" ]]; then
|
||||
for candidate in "$ROOT/admin/.venv/bin/python3" "$ROOT/rag/.venv/bin/python3"; do
|
||||
if [[ -x "$candidate" ]]; then PYTHON="$candidate"; break; fi
|
||||
done
|
||||
fi
|
||||
PYTHON="${PYTHON:-$(command -v python3)}"
|
||||
|
||||
if ! "$PYTHON" -c 'import psycopg2, requests' >/dev/null 2>&1; then
|
||||
echo "ERRO: o Python da RAG precisa dos pacotes psycopg2 e requests." >&2
|
||||
echo "Instale-os no ambiente indicado por RAG_PYTHON e tente novamente." >&2
|
||||
exit 1
|
||||
fi
|
||||
|
||||
if [[ -z "${RAG_DB_PASSWORD:-}" ]]; then
|
||||
echo "ERRO: defina RAG_DB_PASSWORD ou configure $ENV_FILE" >&2
|
||||
exit 1
|
||||
fi
|
||||
|
||||
HOST="${RAG_VPS_HOST:-179.197.228.240}"
|
||||
LOCAL_PORT="${RAG_DB_PORT:-55435}"
|
||||
REMOTE_PORT="${RAG_REMOTE_PORT:-55435}"
|
||||
TUNNEL_PID=""
|
||||
cleanup() {
|
||||
if [[ -n "$TUNNEL_PID" ]] && kill -0 "$TUNNEL_PID" 2>/dev/null; then
|
||||
kill "$TUNNEL_PID" 2>/dev/null || true
|
||||
wait "$TUNNEL_PID" 2>/dev/null || true
|
||||
fi
|
||||
}
|
||||
trap cleanup EXIT
|
||||
|
||||
if ! nc -z 127.0.0.1 "$LOCAL_PORT" 2>/dev/null; then
|
||||
echo "==> Abrindo túnel RAG (127.0.0.1:$LOCAL_PORT)..."
|
||||
ssh -N -o ExitOnForwardFailure=yes -o ServerAliveInterval=30 \
|
||||
-o ServerAliveCountMax=3 -L "127.0.0.1:$LOCAL_PORT:127.0.0.1:$REMOTE_PORT" \
|
||||
"${RAG_VPS_USER:-root}@$HOST" &
|
||||
TUNNEL_PID=$!
|
||||
for _ in {1..20}; do
|
||||
nc -z 127.0.0.1 "$LOCAL_PORT" 2>/dev/null && break
|
||||
kill -0 "$TUNNEL_PID" 2>/dev/null || break
|
||||
sleep 0.25
|
||||
done
|
||||
fi
|
||||
|
||||
if ! nc -z 127.0.0.1 "$LOCAL_PORT" 2>/dev/null; then
|
||||
echo "ERRO: não foi possível abrir o túnel RAG." >&2
|
||||
exit 1
|
||||
fi
|
||||
|
||||
echo "==> Atualizando RAG do G-ART (incremental)..."
|
||||
cd "$ROOT"
|
||||
exec "$PYTHON" "$ROOT/admin/update_rag.py"
|
||||
Executable
+216
@@ -0,0 +1,216 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Atualiza incrementalmente o índice RAG do G-ART.
|
||||
|
||||
As credenciais são fornecidas pelo ambiente; este arquivo nunca deve conter
|
||||
senha. O indexador usa o banco ``rag_gart`` e o schema ``gart`` por padrão.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import hashlib
|
||||
import os
|
||||
import re
|
||||
import sys
|
||||
from pathlib import Path
|
||||
|
||||
import psycopg2
|
||||
import requests
|
||||
|
||||
ROOT = Path(__file__).resolve().parents[1]
|
||||
DB_NAME = os.environ.get("RAG_DB_NAME", "rag_gart")
|
||||
DB_SCHEMA = os.environ.get("RAG_DB_SCHEMA", "gart")
|
||||
OLLAMA_URL = os.environ.get("OLLAMA_URL", "http://127.0.0.1:11434")
|
||||
EMBED_MODEL = os.environ.get("RAG_EMBED_MODEL", "nomic-embed-text")
|
||||
EMBED_DIM = int(os.environ.get("RAG_EMBED_DIM", "768"))
|
||||
|
||||
INCLUDE_EXTENSIONS = {
|
||||
".command", ".md", ".py", ".sh", ".sql", ".swift", ".txt", ".yml", ".yaml",
|
||||
}
|
||||
EXCLUDE_DIRS = {
|
||||
".git", ".venv", ".pytest_cache", ".ruff_cache", "__pycache__", "build",
|
||||
"dist", "node_modules", "graphify-out", "bm", "models", "whisper",
|
||||
}
|
||||
EXCLUDE_FILES = {".env", "admin/genial-crm.env", "admin/genial-crm.local.env"}
|
||||
CHUNK_LINES = 60
|
||||
CHUNK_OVERLAP = 10
|
||||
CHUNK_MAX_CHARS = 5000
|
||||
|
||||
|
||||
def _sql_id(value: str) -> str:
|
||||
return '"' + value.replace('"', '""') + '"'
|
||||
|
||||
|
||||
def _connect():
|
||||
password = os.environ.get("RAG_DB_PASSWORD")
|
||||
if not password:
|
||||
raise RuntimeError("RAG_DB_PASSWORD não foi definida")
|
||||
return psycopg2.connect(
|
||||
host=os.environ.get("RAG_DB_HOST", "127.0.0.1"),
|
||||
port=os.environ.get("RAG_DB_PORT", "55435"),
|
||||
dbname=DB_NAME,
|
||||
user=os.environ.get("RAG_DB_USER", "gart_rag_indexer"),
|
||||
password=password,
|
||||
connect_timeout=5,
|
||||
)
|
||||
|
||||
|
||||
def _iter_files():
|
||||
for path in ROOT.rglob("*"):
|
||||
if not path.is_file() or path.suffix.lower() not in INCLUDE_EXTENSIONS:
|
||||
continue
|
||||
rel = path.relative_to(ROOT).as_posix()
|
||||
parts = set(path.relative_to(ROOT).parts)
|
||||
if parts & EXCLUDE_DIRS or rel in EXCLUDE_FILES or path.name in EXCLUDE_FILES:
|
||||
continue
|
||||
if any(part.startswith(".") for part in path.relative_to(ROOT).parts[:-1]):
|
||||
continue
|
||||
yield path, rel
|
||||
|
||||
|
||||
def _chunks(text: str):
|
||||
lines = text.splitlines()
|
||||
if not lines:
|
||||
return []
|
||||
step = max(1, CHUNK_LINES - CHUNK_OVERLAP)
|
||||
result = []
|
||||
for start in range(0, len(lines), step):
|
||||
window_start = start
|
||||
buffer = []
|
||||
size = 0
|
||||
for offset, line in enumerate(lines[start:start + CHUNK_LINES]):
|
||||
if buffer and size + len(line) + 1 > CHUNK_MAX_CHARS:
|
||||
result.append((window_start + 1, window_start + len(buffer), "\n".join(buffer).strip()))
|
||||
buffer = []
|
||||
window_start = start + offset
|
||||
size = 0
|
||||
buffer.append(line)
|
||||
size += len(line) + 1
|
||||
if buffer:
|
||||
result.append((window_start + 1, window_start + len(buffer), "\n".join(buffer).strip()))
|
||||
if start + CHUNK_LINES >= len(lines):
|
||||
break
|
||||
return [(start, end, content) for start, end, content in result if content]
|
||||
|
||||
|
||||
def _facts(text: str, rel_path: str):
|
||||
lines = text.splitlines()
|
||||
summary = next(
|
||||
(line.strip().lstrip("#! ").strip() for line in lines[:30] if line.strip()),
|
||||
None,
|
||||
)
|
||||
symbols = re.findall(
|
||||
r"^\s*(?:class|def|async\s+def|func|struct|enum|protocol|actor|interface)\s+([A-Za-z_]\w*)",
|
||||
text,
|
||||
re.MULTILINE,
|
||||
)
|
||||
parts = Path(rel_path).parts
|
||||
module = parts[0] if len(parts) > 1 else None
|
||||
return module, Path(rel_path).stem, summary, sorted(set(symbols)), len(lines)
|
||||
|
||||
|
||||
class ChunkTooLargeError(Exception):
|
||||
"""Chunk excede o contexto do modelo de embedding (ver EXCLUDE_FILES/CHUNK_MAX_CHARS)."""
|
||||
|
||||
|
||||
def _embed(text: str):
|
||||
response = requests.post(
|
||||
f"{OLLAMA_URL.rstrip('/')}/api/embeddings",
|
||||
json={"model": EMBED_MODEL, "prompt": f"search_document: {text}"},
|
||||
timeout=60,
|
||||
)
|
||||
if response.status_code == 500 and "context length" in response.text.lower():
|
||||
raise ChunkTooLargeError(response.text)
|
||||
response.raise_for_status()
|
||||
vector = response.json()["embedding"]
|
||||
if len(vector) != EMBED_DIM:
|
||||
raise ValueError(f"embedding com {len(vector)} dimensões; esperado {EMBED_DIM}")
|
||||
return vector
|
||||
|
||||
|
||||
def _hash(text: str) -> str:
|
||||
return hashlib.md5(text.encode("utf-8")).hexdigest()
|
||||
|
||||
|
||||
def index():
|
||||
schema = _sql_id(DB_SCHEMA)
|
||||
conn = _connect()
|
||||
conn.autocommit = False
|
||||
indexed = skipped = deleted = chunks_written = 0
|
||||
seen = set()
|
||||
try:
|
||||
with conn.cursor() as cur:
|
||||
for path, rel_path in sorted(_iter_files(), key=lambda item: item[1]):
|
||||
try:
|
||||
text = path.read_text(encoding="utf-8", errors="ignore")
|
||||
except OSError as exc:
|
||||
print(f"[RAG] ignorado {rel_path}: {exc}", file=sys.stderr)
|
||||
continue
|
||||
seen.add(rel_path)
|
||||
digest = _hash(text)
|
||||
cur.execute(f"SELECT content_hash FROM {schema}.indexed_files WHERE file_path = %s", (rel_path,))
|
||||
row = cur.fetchone()
|
||||
if row and row[0] == digest:
|
||||
skipped += 1
|
||||
continue
|
||||
|
||||
module, main_type, summary, symbols, n_lines = _facts(text, rel_path)
|
||||
cur.execute(f"DELETE FROM {schema}.code_chunks WHERE file_path = %s", (rel_path,))
|
||||
for index_number, (start, end, content) in enumerate(_chunks(text)):
|
||||
try:
|
||||
vector = _embed(content)
|
||||
except ChunkTooLargeError:
|
||||
# Chunks densos em tokens (ex: tabelas de dados numéricas
|
||||
# como font_metrics.py) podem passar de CHUNK_MAX_CHARS em
|
||||
# caracteres mas estourar o contexto do modelo em tokens.
|
||||
# Pular o chunk em vez de abortar a transação inteira.
|
||||
print(f"[RAG] chunk grande demais, pulado: {rel_path}:{start}-{end}", file=sys.stderr)
|
||||
continue
|
||||
cur.execute(
|
||||
f"""INSERT INTO {schema}.code_chunks
|
||||
(file_path, content, chunk_index, embedding, content_hash,
|
||||
file_mtime, start_line, end_line, symbols, module, kind)
|
||||
VALUES (%s, %s, %s, %s, %s, %s, %s, %s, %s, %s, %s)""",
|
||||
(rel_path, content, index_number, vector, digest,
|
||||
path.stat().st_mtime, start, end, ", ".join(symbols), module, "window"),
|
||||
)
|
||||
chunks_written += 1
|
||||
cur.execute(
|
||||
f"""INSERT INTO {schema}.file_index
|
||||
(file_path, module, main_type, public_symbols, summary, n_lines, content_hash)
|
||||
VALUES (%s, %s, %s, %s, %s, %s, %s)
|
||||
ON CONFLICT (file_path) DO UPDATE SET
|
||||
module = EXCLUDED.module, main_type = EXCLUDED.main_type,
|
||||
public_symbols = EXCLUDED.public_symbols, summary = EXCLUDED.summary,
|
||||
n_lines = EXCLUDED.n_lines, content_hash = EXCLUDED.content_hash,
|
||||
updated_at = CURRENT_TIMESTAMP""",
|
||||
(rel_path, module, main_type, symbols, summary, n_lines, digest),
|
||||
)
|
||||
cur.execute(
|
||||
f"""INSERT INTO {schema}.indexed_files (file_path, content_hash)
|
||||
VALUES (%s, %s)
|
||||
ON CONFLICT (file_path) DO UPDATE SET
|
||||
content_hash = EXCLUDED.content_hash, updated_at = CURRENT_TIMESTAMP""",
|
||||
(rel_path, digest),
|
||||
)
|
||||
indexed += 1
|
||||
print(f"[RAG] {rel_path}")
|
||||
|
||||
cur.execute(f"SELECT file_path FROM {schema}.indexed_files")
|
||||
for (rel_path,) in cur.fetchall():
|
||||
if rel_path not in seen:
|
||||
cur.execute(f"DELETE FROM {schema}.code_chunks WHERE file_path = %s", (rel_path,))
|
||||
cur.execute(f"DELETE FROM {schema}.file_index WHERE file_path = %s", (rel_path,))
|
||||
cur.execute(f"DELETE FROM {schema}.indexed_files WHERE file_path = %s", (rel_path,))
|
||||
deleted += 1
|
||||
print(f"[RAG] removido {rel_path}")
|
||||
conn.commit()
|
||||
except Exception:
|
||||
conn.rollback()
|
||||
raise
|
||||
finally:
|
||||
conn.close()
|
||||
print(f"[RAG] concluído: {indexed} atualizado(s), {skipped} sem mudança, {deleted} removido(s), {chunks_written} chunk(s).")
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
index()
|
||||
@@ -0,0 +1 @@
|
||||
analysis/
|
||||
+43
-32
@@ -10,12 +10,20 @@ opera **fora** do Final Cut Pro: você exporta o XML, o servidor processa o
|
||||
documento como dados estruturados e devolve um XML modificado para importação.
|
||||
Nada é patcheado, nenhuma API privada é usada.
|
||||
|
||||
Toda a análise foi feita a partir do código-fonte em `server.py` e `fcpxml/`.
|
||||
Toda a análise foi feita a partir do código-fonte. Este README é a visão
|
||||
geral; o detalhe módulo a módulo mora em `docs/02_MODULES.md`, que é o
|
||||
documento a manter atualizado quando a estrutura mudar.
|
||||
|
||||
> **Guia rápido:** [01 Arquitetura](docs/01_ARCHITECTURE.md) ·
|
||||
> **Começando agora?** Leia [01 Arquitetura](docs/01_ARCHITECTURE.md) e depois
|
||||
> [09 Manutenção](docs/09_MANUTENCAO.md) — o primeiro diz como o sistema é
|
||||
> dividido, o segundo diz por onde começar a mexer e o que está em aberto.
|
||||
>
|
||||
> **Guia completo:** [01 Arquitetura](docs/01_ARCHITECTURE.md) ·
|
||||
> [02 Módulos](docs/02_MODULES.md) · [03 Server/Tools](docs/03_SERVER_TOOLS.md) ·
|
||||
> [04 Testes & Workflow](docs/04_TESTS_AND_WORKFLOW.md) ·
|
||||
> [05 Experiências](docs/05_EXPERIENCIAS.md) · [06 Boas Práticas](docs/06_BOAS_PRATICAS.md)
|
||||
> [05 Experiências](docs/05_EXPERIENCIAS.md) · [06 Boas Práticas](docs/06_BOAS_PRATICAS.md) ·
|
||||
> [07 Projeto Ativo no FCP](docs/07_ESTUDO_PROJETO_ATIVO_FCP.md) ·
|
||||
> [08 App macOS](docs/08_APP_MACOS.md) · [09 Manutenção](docs/09_MANUTENCAO.md)
|
||||
|
||||
---
|
||||
|
||||
@@ -26,7 +34,7 @@ Toda a análise foi feita a partir do código-fonte em `server.py` e `fcpxml/`.
|
||||
Python, e reescreve de volta sem perda de sidecars (object tracking,
|
||||
Cinematic).
|
||||
|
||||
2. **Uma camada MCP de 62 ferramentas** — expõe análise, edição em lote, QC,
|
||||
2. **Uma camada MCP de 74 ferramentas** — expõe análise, edição em lote, QC,
|
||||
geração, exportação cross-NLE, inteligência de mídia (silêncio/beats) e
|
||||
edição baseada em transcrição, tudo acessível por um cliente MCP (Claude).
|
||||
|
||||
@@ -40,7 +48,7 @@ Toda a análise foi feita a partir do código-fonte em `server.py` e `fcpxml/`.
|
||||
|
||||
| Camada | Tecnologia |
|
||||
|--------|-----------|
|
||||
| Linguagem | **Python 3.10+** (~7.1k linhas em `server.py` + `fcpxml/`) |
|
||||
| Linguagem | **Python 3.10+** (~13k linhas em `server.py`, `server_tools/` e `fcpxml/`) |
|
||||
| Protocolo MCP | **mcp** (`mcp` SDK), servidor por stdio |
|
||||
| Parsing XML | **defusedxml** em todos os 4 entry points + `lxml`/`ElementTree` |
|
||||
| Tempo racional | frações `numerador/denominador` no formato `"600/2400s"` |
|
||||
@@ -56,26 +64,29 @@ Toda a análise foi feita a partir do código-fonte em `server.py` e `fcpxml/`.
|
||||
|
||||
```
|
||||
G-ART/
|
||||
├── server.py # MCP server — 62 tools, prompts, resources, dispatch
|
||||
├── fcpxml/ # "Engine" — biblioteca Python de núcleo
|
||||
│ ├── models.py # TimeValue, Timecode, Clip, Timeline, enums, QC models
|
||||
│ ├── parser.py # FCPXML → objetos Python (spine, connected clips, roles)
|
||||
│ ├── writer.py # Modifica e grava FCPXML (markers, trim, gaps, speed)
|
||||
│ ├── rough_cut.py # Gera timelines novas (rough cuts, montages, A/B)
|
||||
│ ├── diff.py # Motor de comparação de timelines
|
||||
│ ├── export.py # Export DaVinci Resolve v1.9 + FCP7 XMEML v5
|
||||
│ ├── media_intel.py # Detecção real de silêncio (ffmpeg) e beats (librosa)
|
||||
│ ├── transcribe.py # Transcrição Whisper local + edição por transcrição
|
||||
│ ├── templates.py # Templates de timeline (intro/outro, lower thirds)
|
||||
│ ├── live.py # Modo Live — push_to_fcp / list_fcp_libraries
|
||||
│ ├── safe_xml.py # Wrappers defusedxml + serialize_xml()
|
||||
│ └── dtd.py # Validação contra DTDs oficiais da Apple
|
||||
├── Engine/ # Esta documentação da arquitetura
|
||||
├── admin/ # Scripts de manutenção (graphify.sh, graphify.md)
|
||||
├── docs/ # WORKFLOWS, CAPABILITY-AUDIT, specs
|
||||
├── examples/ # Fixture de teste (sample.fcpxml)
|
||||
├── tests/ # 1032 testes / 24 suítes
|
||||
└── tools/ # Pacote Python (__init__)
|
||||
├── CLAUDE.md # Regras do projeto para o agente
|
||||
├── admin/ # Ponte com o app (fora de code/)
|
||||
│ ├── models_api.py # Entry point: docstring dos comandos + dispatch
|
||||
│ └── api/ # Os 37 comandos, um módulo por assunto
|
||||
└── code/
|
||||
├── server.py # MCP entry point — só dispatch
|
||||
├── server_tools/ # Handlers das 74 tools + _shared/
|
||||
├── fcpxml/ # "Engine" — biblioteca Python de núcleo
|
||||
│ ├── writer/ # PACOTE: edição/escrita (mixins por assunto)
|
||||
│ ├── models/ # PACOTE: dados por família (timing, timeline…)
|
||||
│ ├── parser.py # FCPXML → objetos Python
|
||||
│ ├── rough_cut.py # Gera timelines novas
|
||||
│ ├── voice_*.py # Pipeline de voz (features → timeline → actions)
|
||||
│ ├── phrase_review.py # Revisão de frases da etapa 5
|
||||
│ ├── text_layout.py # Diagramação das legendas
|
||||
│ ├── live.py # Modo Live — push_to_fcp
|
||||
│ ├── safe_xml.py # defusedxml + serialize_xml()
|
||||
│ └── dtd.py # Validação contra DTDs da Apple
|
||||
├── MacApp/Sources/ # App SwiftUI (compilado por swiftc)
|
||||
├── Engine/ # Esta documentação
|
||||
├── docs/ # WORKFLOWS, CAPABILITY-AUDIT, specs
|
||||
├── examples/ # Fixture de teste (sample.fcpxml)
|
||||
└── tests/ # 1.466 testes / 42 suítes
|
||||
```
|
||||
|
||||
---
|
||||
@@ -102,7 +113,7 @@ TimeValue(600, 2400) # "600/2400s" == 0.25s
|
||||
- Soma/subtração compartilham um único caminho `_binop()` (fast-path de mesmo
|
||||
denominador + alinhamento por LCM).
|
||||
|
||||
### 4.2 Modelos principais — `models.py`
|
||||
### 4.2 Modelos principais — `models/`
|
||||
|
||||
| Classe | Função |
|
||||
|--------|--------|
|
||||
@@ -126,8 +137,8 @@ escrita. `from_xml_element` faz match estrito do atributo `completed`
|
||||
| Subsistema | Módulo | Função |
|
||||
|-----------|--------|--------|
|
||||
| Parser | `parser.py` | FCPXML → objetos Python: espinha, connected clips, secondary storylines, roles |
|
||||
| Modifier | `writer.FCPXMLModifier` | Edição index-based (clips/resources/formats dicts) do documento existente |
|
||||
| Writer | `writer.FCPXMLWriter` | Gera FCPXML novo a partir de objetos Python |
|
||||
| Modifier | `writer/` (`FCPXMLModifier`) | Edição index-based (clips/resources/formats dicts) do documento existente |
|
||||
| Writer | `writer/generator.py` | Gera FCPXML novo a partir de objetos Python |
|
||||
| Rough cut | `rough_cut.py` | Gera timelines (rough cuts, montages, A/B roll) |
|
||||
| Diff | `diff.py` | Compara timelines — detecta added/removed/moved/trimmed |
|
||||
| Export | `export.py` | DaVinci Resolve v1.9 + FCP7 XMEML v5 |
|
||||
@@ -151,7 +162,7 @@ assíncrono:
|
||||
TOOL_HANDLERS = {
|
||||
"analyze_timeline": handle_analyze_timeline,
|
||||
"list_clips": handle_list_clips,
|
||||
# ... 62 tools
|
||||
# ... 74 tools, todos em server_tools/
|
||||
}
|
||||
```
|
||||
|
||||
@@ -249,7 +260,7 @@ correção, sem depender de lembrar dos dois comandos no pre-commit.
|
||||
- [docs/02_MODULES.md](docs/02_MODULES.md) — guia módulo a módulo do `fcpxml/`
|
||||
(responsabilidade, tamanho, APIs públicas).
|
||||
- [docs/03_SERVER_TOOLS.md](docs/03_SERVER_TOOLS.md) — a camada MCP `server.py`,
|
||||
62 ferramentas, helpers e o padrão de handler.
|
||||
74 ferramentas, helpers e o padrão de handler.
|
||||
- [docs/04_TESTS_AND_WORKFLOW.md](docs/04_TESTS_AND_WORKFLOW.md) — suíte de testes,
|
||||
fluxo de trabalho (lint + pytest), execução e estado atual do sistema.
|
||||
- [docs/05_EXPERIENCIAS.md](docs/05_EXPERIENCIAS.md) — **memória de projeto**:
|
||||
@@ -259,10 +270,10 @@ correção, sem depender de lembrar dos dois comandos no pre-commit.
|
||||
programação** a aplicar em toda alteração/correção; inclui checklist final.
|
||||
|
||||
### Outros documentos
|
||||
- [../CLAUDE.md](../CLAUDE.md) — visão geral, key patterns, execução e pre-commit.
|
||||
- [../CLAUDE.md](../../CLAUDE.md) — visão geral, key patterns, execução e pre-commit.
|
||||
- [../docs/CAPABILITY-AUDIT-2026-06.md](../docs/CAPABILITY-AUDIT-2026-06.md) —
|
||||
auditoria do ecossistema e roadmap dual-mode (XML + Live).
|
||||
- [../docs/WORKFLOWS.md](../docs/WORKFLOWS.md) — 8 receitas de workflow de produção.
|
||||
- [../docs/specs/](../docs/specs/) — schemas de tools, estrutura FCPXML, pseudocódigo
|
||||
do writer, algoritmo de rough cut, implementação do server, roadmap, modelos.
|
||||
- [../admin/graphify.md](../admin/graphify.md) — pipeline de graphify do código.
|
||||
- [../admin/graphify.md](../../admin/graphify.md) — pipeline de graphify do código.
|
||||
|
||||
@@ -1,109 +1,177 @@
|
||||
# 01 — Arquitetura do Sistema (G-ART / fcp-mcp-server)
|
||||
|
||||
> Referência canônica de como o sistema está dividido e implementado. Leia este
|
||||
> documento antes de qualquer mudança de código.
|
||||
> **Escopo:** Como o sistema é dividido em camadas e onde cada responsabilidade mora.
|
||||
> **Não cobre:** Detalhe módulo a módulo (→ 02) · ferramentas MCP (→ 03) · app (→ 08)
|
||||
|
||||
> Referência canônica de como o sistema está dividido. Leia antes de qualquer
|
||||
> mudança de código. Se algo aqui divergir do código, **o código está certo e
|
||||
> este documento está velho** — corrija-o no mesmo commit.
|
||||
|
||||
Última varredura: 2026-08-19 · 77 ferramentas MCP · 1.498 testes · versão `0.6.35`
|
||||
|
||||
---
|
||||
|
||||
## 1. Visão de cima (camadas)
|
||||
|
||||
O sistema é um **servidor MCP em Python** que lê/analisa/reescreve arquivos
|
||||
**FCPXML** do Final Cut Pro. Há **três camadas** bem separadas:
|
||||
O sistema lê, analisa e reescreve **FCPXML** do Final Cut Pro. Ele opera *fora*
|
||||
do FCP: você exporta o XML, o programa processa como dados estruturados e
|
||||
devolve um XML para importar. Nada é patcheado, nenhuma API privada é usada.
|
||||
|
||||
São **quatro camadas**, e o ponto importante é que existem **duas portas de
|
||||
entrada diferentes** para o mesmo motor:
|
||||
|
||||
```
|
||||
┌─────────────────────────────────────────────────────────────┐
|
||||
│ admin/ — Aplicações complementares (fora do MCP) │
|
||||
│ models_api.py API (FastAPI) p/ gerenciar modelos │
|
||||
│ models_gui.py UI desktop (Flet) p/ gerenciar modelos │
|
||||
│ graphify.sh/.md Pipeline de graphify do código │
|
||||
├─────────────────────────────────────────────────────────────┤
|
||||
│ server.py — CAMADA MCP / TRANSPORTE (NÃO tem lógica) │
|
||||
│ 73 tools, handlers, prompts, resources, dispatch │
|
||||
│ Só valida entrada/saída e traduz JSON-RPC → chamadas │
|
||||
├─────────────────────────────────────────────────────────────┤
|
||||
│ fcpxml/ — "ENGINE" = NÚCLEO PURO Python (desacoplado) │
|
||||
│ Não conhece MCP nem argumentos de tool. │
|
||||
│ Trabalha com objetos Python e XML. │
|
||||
│ É o foco / onde quase tudo mora. │
|
||||
└─────────────────────────────────────────────────────────────┘
|
||||
┌──────────────────────────┐ ┌──────────────────────────────┐
|
||||
│ MacApp/ (SwiftUI) │ │ Cliente MCP (Claude) │
|
||||
│ O app que o usuário usa │ │ Conversa, decide a edição │
|
||||
└───────────┬──────────────┘ └───────────────┬──────────────┘
|
||||
│ subprocesso + JSON-lines │ JSON-RPC (stdio)
|
||||
▼ ▼
|
||||
┌──────────────────────────┐ ┌──────────────────────────────┐
|
||||
│ admin/models_api.py │ │ server.py + server_tools/ │
|
||||
│ + admin/api/ │ │ 77 tools, dispatch, schemas │
|
||||
│ 37 comandos da ponte │ │ NÃO tem lógica de timeline │
|
||||
└───────────┬──────────────┘ └───────────────┬──────────────┘
|
||||
└───────────────┬────────────────────┘
|
||||
▼
|
||||
┌───────────────────────────────────┐
|
||||
│ fcpxml/ — O ENGINE │
|
||||
│ Núcleo puro Python, desacoplado. │
|
||||
│ Não conhece MCP nem o app. │
|
||||
│ É onde quase tudo mora. │
|
||||
└───────────────────────────────────┘
|
||||
```
|
||||
|
||||
**Regra de arquitetura:** `server.py` NUNCA implementa lógica de timeline —
|
||||
ele delega ao `fcpxml/`. Tudo em `fcpxml/` é testável isoladamente (1032 testes).
|
||||
**A regra que sustenta tudo:** nem `server.py` nem `admin/api/` implementam
|
||||
lógica de timeline. Os dois validam entrada, chamam o engine e formatam a
|
||||
saída. Toda regra de negócio é testável sem MCP e sem app.
|
||||
|
||||
## 2. Regras transversais (convenções em todo o código)
|
||||
**Por que duas portas.** O MCP existe para o julgamento editorial — qual tomada
|
||||
usar, onde dar zoom — que é conversa com uma IA. A ponte existe para o que o
|
||||
usuário faz sozinho no app — transcrever, configurar, processar. As duas caem
|
||||
no mesmo engine, então uma correção ali vale para as duas.
|
||||
|
||||
---
|
||||
|
||||
## 2. Regras transversais (valem em todo o código)
|
||||
|
||||
| Conceito | Regra |
|
||||
|----------|-------|
|
||||
| **Tempo** | `TimeValue` fração racional `"600/2400s"`. Nunca use float p/ tempo. |
|
||||
| **I/O paths** | Sempre via helpers `_validate_filepath` / `_validate_output_path` (sandbox). |
|
||||
| **Nome de saída** | Nunca sobrescrever original: `output_<suffix>.fcpxml`. |
|
||||
| **Segurança XML** | Sempre `defusedxml` (via `safe_xml.py`). Nunca `xml.etree` direto. |
|
||||
| **Tempo** | `TimeValue`, fração racional `"600/2400s"`. **Nunca float para tempo.** |
|
||||
| **Tempo de decisão** | Ações de voz usam sempre segundos da **mídia original**, nunca pós-corte. |
|
||||
| **I/O paths** | Sempre via `_validate_filepath` / `_validate_output_path` (sandbox). |
|
||||
| **Nome de saída** | Nunca sobrescrever o original: `generate_output_path()` gera `_suffix`. |
|
||||
| **Segurança XML** | Sempre `defusedxml` via `safe_xml.py`. Nunca `xml.etree` direto para ler. |
|
||||
| **Deps opcionais** | `librosa`/`ffmpeg`/`huggingface_hub` importados **lazy**, degradam com `None`. |
|
||||
| **Lint** | `ruff check . --exclude docs/` — zero erros. |
|
||||
| **Validação pós-correção** | `./Engine/run_after_fix.sh` SEMPRE após cada correção. |
|
||||
| **Idioma** | Comunicação com o usuário em português. Código e comentários em inglês. |
|
||||
| **Validação** | `./Engine/run_after_fix.sh` **sempre** após cada correção. |
|
||||
| **App** | Alterou `MacApp/`? Compile e rode: `admin/run_app.command` (padrão de revisão; equivale a `./MacApp/build_app.sh --run`). |
|
||||
|
||||
## 3. Fluxo de um request (round-trip)
|
||||
---
|
||||
|
||||
## 3. Fluxo de um request
|
||||
|
||||
### Pela porta MCP (Claude decidindo a edição)
|
||||
|
||||
```
|
||||
Cliente MCP (Claude)
|
||||
│ JSON-RPC (stdio)
|
||||
▼
|
||||
server.py ── dispatcher (TOOL_HANDLERS)
|
||||
│ valida path, parseia projeto, chama engine
|
||||
▼
|
||||
fcpxml/parser.py XML → objetos
|
||||
fcpxml/writer.py edita / grava
|
||||
fcpxml/rough_cut.py gera novas timelines
|
||||
fcpxml/export.py cross-NLE
|
||||
▼
|
||||
output_<suffix>.fcpxml (original intocado)
|
||||
▼
|
||||
Final Cut Pro: File → Import → XML (ou push_to_fcp, sem cliques)
|
||||
Cliente MCP ──JSON-RPC──► server.py
|
||||
│ TOOL_HANDLERS[nome]
|
||||
▼
|
||||
server_tools/<categoria>.py
|
||||
│ _shared/: valida path, parseia projeto
|
||||
▼
|
||||
fcpxml/ (parser → writer → safe_xml)
|
||||
▼
|
||||
projeto_<suffix>.fcpxml (original intocado)
|
||||
```
|
||||
|
||||
### Pela porta do app (usuário operando)
|
||||
|
||||
```
|
||||
MacApp ──Process + argv JSON──► admin/models_api.py
|
||||
│ handlers[comando]
|
||||
▼
|
||||
admin/api/<assunto>.py
|
||||
│ shared.emit() devolve JSON-lines
|
||||
▼
|
||||
fcpxml/ (ou chama um handler do server)
|
||||
▼
|
||||
arquivo gerado + caminho de volta ao app
|
||||
```
|
||||
|
||||
A saída da ponte é **JSON-lines**: um documento JSON por linha, para que
|
||||
comandos longos transmitam progresso enquanto rodam. Toda escrita passa por
|
||||
`admin/api/shared.py::emit`, que serializa o acesso a stdout — dois comandos
|
||||
escrevendo ao mesmo tempo entrelaçariam documentos.
|
||||
|
||||
---
|
||||
|
||||
## 4. Dual-mode: XML + Live
|
||||
|
||||
O sistema opera em **dois modos complementares**:
|
||||
|
||||
- **Modo XML (principal):** exporta FCPXML, processa como dados, reimporta.
|
||||
Roda fora do FCP. Nenhuma API privada.
|
||||
- **Modo Live (`fcpxml/live.py`):** *push* do FCPXML direto p/ o FCP em
|
||||
execução via Apple events oficiais (`Open Document`), com `import-options`.
|
||||
Leitura de bibliotecas via AppleScript read-only.
|
||||
- **Modo Live (`fcpxml/live.py`):** *push* do FCPXML direto para o FCP em
|
||||
execução via Apple events oficiais (`Open Document`). Leitura de bibliotecas
|
||||
via AppleScript read-only.
|
||||
|
||||
**Assimetria estrutural:** import é scriptable, mas a Apple não oferece export
|
||||
programático — round-trips voltam pelas ferramentas XML.
|
||||
programático. Round-trips sempre voltam pelas ferramentas XML.
|
||||
|
||||
---
|
||||
|
||||
## 5. Onde está cada responsabilidade
|
||||
|
||||
| Responsabilidade | Fica em |
|
||||
|------------------|---------|
|
||||
| Modelos de dados (tempo, clips, markers) | `fcpxml/models.py` |
|
||||
| Modelos de dados (tempo, clips, markers, QC, legendas) | `fcpxml/models/` |
|
||||
| Parse FCPXML → objetos | `fcpxml/parser.py` |
|
||||
| Editing/escrita (modifier + writer) | `fcpxml/writer.py` |
|
||||
| Edição e escrita de FCPXML | `fcpxml/writer/` |
|
||||
| Geração de timeline nova | `fcpxml/rough_cut.py` |
|
||||
| Comparação de timelines | `fcpxml/diff.py` |
|
||||
| Export cross-NLE (Resolve, FCP7) | `fcpxml/export.py` |
|
||||
| Inteligência de mídia (silêncio/beats) | `fcpxml/media_intel.py` |
|
||||
| Transcrição Whisper local | `fcpxml/transcribe.py` |
|
||||
| Silêncio e beats | `fcpxml/media_intel.py` |
|
||||
| Transcrição Whisper | `fcpxml/transcribe.py` |
|
||||
| Diarização (quem falou) | `fcpxml/diarize.py` |
|
||||
| Ênfase acústica | `fcpxml/emphasis.py`, `fcpxml/voice_features.py` |
|
||||
| Timeline de voz (o JSON que a IA lê) | `fcpxml/voice_timeline.py` |
|
||||
| Decisões de edição (cut/zoom/text/marker) | `fcpxml/voice_actions.py` |
|
||||
| Revisão de frases da etapa 5 | `fcpxml/phrase_review.py` |
|
||||
| Layout de legendas e métricas de fonte | `fcpxml/text_layout.py`, `font_metrics.py`, `collision.py` |
|
||||
| Gestão de modelos Whisper | `fcpxml/model_manager.py` |
|
||||
| Templates de timeline | `fcpxml/templates.py` |
|
||||
| Controle Live do FCP | `fcpxml/live.py` |
|
||||
| Segurança XML (`defusedxml`, `serialize_xml`) | `fcpxml/safe_xml.py` |
|
||||
| Segurança XML | `fcpxml/safe_xml.py` |
|
||||
| Validação contra DTDs da Apple | `fcpxml/dtd.py` |
|
||||
| Transporte MCP (73 tools) | `server.py` |
|
||||
| Transporte MCP (77 tools) | `server.py` + `server_tools/` |
|
||||
| Ponte com o app (37 comandos) | `admin/models_api.py` + `admin/api/` |
|
||||
| Interface do usuário | `MacApp/Sources/` |
|
||||
|
||||
## 6. Mapa de dependências (você está aqui se for mexer no X → quem tocar)
|
||||
---
|
||||
|
||||
## 6. Mapa de dependências
|
||||
|
||||
```
|
||||
server.py ──► fcpxml/parser, writer, rough_cut, export, diff,
|
||||
media_intel, transcribe, templates, live, dtd
|
||||
admin/models_gui.py ──► fcpxml/media_intel, model_manager,
|
||||
parser, transcribe
|
||||
admin/models_api.py ──► fcpxml/model_manager
|
||||
fcpxml/writer.py ──► fcpxml/models, safe_xml, dtd
|
||||
fcpxml/__init__.py ──► reexporta a API pública
|
||||
MacApp/ ──► admin/models_api.py (subprocesso, por caminho)
|
||||
admin/api/ ──► fcpxml/* e, para algumas operações, server.py
|
||||
server.py ──► server_tools/*
|
||||
server_tools/* ──► server_tools/_shared/ ──► fcpxml/*
|
||||
fcpxml/writer/ ──► fcpxml/models/, safe_xml, dtd, text_layout, collision
|
||||
fcpxml/models/ ──► fcpxml/text_layout (só o pacote subtitles)
|
||||
fcpxml/__init__.py ──► reexporta a API pública
|
||||
```
|
||||
|
||||
> Se você cria uma **nova ferramenta MCP**, o trabalho principal é em `fcpxml/`
|
||||
> (função pura + testes). O handler em `server.py` fica fino: validação de
|
||||
> caminho → `_parse_project` → chama a função → `_text_result`.
|
||||
**A seta que não existe, e não deve existir:** `fcpxml/` nunca importa de
|
||||
`server_tools/`, de `admin/` ou de qualquer coisa que saiba o que é uma tool.
|
||||
Se você precisar disso, a lógica está no lugar errado.
|
||||
|
||||
---
|
||||
|
||||
## 7. Criando algo novo — por onde começar
|
||||
|
||||
| Você quer… | Comece por |
|
||||
|-----------|-----------|
|
||||
| Uma **ferramenta MCP** nova | Função pura em `fcpxml/` + teste. O handler em `server_tools/` fica fino. |
|
||||
| Um **comando do app** novo | Mesmo caminho, e exponha em `admin/api/<assunto>.py` + tabela em `models_api.py`. |
|
||||
| Uma **tela** nova | `MacApp/Sources/`, consumindo comandos que já existem na ponte. |
|
||||
| Uma **regra de edição** nova | `fcpxml/` sempre. Se você está escrevendo `if` sobre timeline fora de `fcpxml/`, pare. |
|
||||
|
||||
O trabalho principal é **sempre** no engine. As camadas de cima são finas de
|
||||
propósito: é o que permite testar 1.498 casos sem abrir o app nem subir o MCP.
|
||||
|
||||
+232
-96
@@ -1,114 +1,250 @@
|
||||
# 02 — Módulos do Engine (`fcpxml/`)
|
||||
|
||||
Guia módulo a módulo do núcleo Python. Tamanho em linhas, responsabilidade e as
|
||||
funções/classes públicas de cada um. APIs públicas são reexportadas em
|
||||
`fcpxml/__init__.py` (fonte da verdade para o `__all__`).
|
||||
> **Escopo:** Mapa do engine `fcpxml/`: qual módulo faz o quê e onde mexer.
|
||||
> **Não cobre:** Camadas e regras gerais (→ 01) · handlers MCP (→ 03) · o que está aberto (→ 09)
|
||||
|
||||
## Versão atual
|
||||
`__version__ = "0.6.35"` — ver `fcpxml/__init__.py`.
|
||||
Mapa módulo a módulo do núcleo Python: onde cada coisa mora e o que ela faz.
|
||||
A API pública é reexportada em `fcpxml/__init__.py` — essa é a fonte da verdade
|
||||
do `__all__`.
|
||||
|
||||
Versão: `0.13.1` · Última varredura: 2026-09-22
|
||||
|
||||
> **Por que existem pacotes aqui.** `writer.py` tinha 4.199 linhas e `models.py`
|
||||
> 1.091, cada um com muitos assuntos dentro. Viraram pacotes com um módulo por
|
||||
> assunto. Do lado de fora **nada mudou**: `from .writer import FCPXMLModifier`
|
||||
> e `from .models import TimeValue` seguem valendo, porque os `__init__.py`
|
||||
> reexportam tudo — inclusive os nomes com underscore que a suíte usa.
|
||||
|
||||
---
|
||||
|
||||
| Módulo | Linhas | Papel |
|
||||
|--------|-------:|-------|
|
||||
| `models.py` | 930 | Data classes e enums (tempo, clips, markers, QC) |
|
||||
| `parser.py` | 367 | FCPXML → objetos Python |
|
||||
| `writer.py` | 3154 | Edição e escrita de FCPXML (o maior) |
|
||||
## Visão geral
|
||||
|
||||
| Módulo / pacote | Linhas | Papel |
|
||||
|-----------------|-------:|-------|
|
||||
| `writer/` | 5.377 | **Edição e escrita de FCPXML** — o coração |
|
||||
| `models/` | 1.195 | Data classes e enums |
|
||||
| `text_layout.py` | 901 | Diagramação das legendas dinâmicas |
|
||||
| `rough_cut.py` | 798 | Geração de timelines novas |
|
||||
| `dtd.py` | 112 | Validação contra DTDs oficiais |
|
||||
| `safe_xml.py` | 113 | Wrappers `defusedxml` + `serialize_xml()` |
|
||||
| `media_intel.py` | 173 | Silêncio (ffmpeg) e beats (librosa) |
|
||||
| `transcribe.py` | 184 | Transcrição Whisper + edição por transcrição |
|
||||
| `model_manager.py` | 298 | Gestão de modelos Whisper (cache/catálogo) |
|
||||
| `export.py` | 226 | Export DaVinci Resolve v1.9 + FCP7 XMEML v5 |
|
||||
| `diff.py` | 269 | Comparação de timelines |
|
||||
| `live.py` | 273 | Modo Live — push_to_fcp / list_fcp_libraries |
|
||||
| `model_manager.py` | 748 | Modelos Whisper: catálogo, download, config |
|
||||
| `voice_timeline.py` | 600 | O JSON de voz que a IA lê |
|
||||
| `analise.py` | 218 | `AnalisadorDeArquivo` — orquestra transcrição/diarização/ênfase/emoção e monta o `_voice_timeline.json`; usado por `voice_timeline.py` |
|
||||
| `phrase_review.py` | 547 | Revisão de frases (etapa 5 do assistente) |
|
||||
| `speaker_review.py` | 208 | Revisão de falantes (etapa 3 do assistente) |
|
||||
| `transcription/` | 325 | Pacote: `engine.py` (adapter faster-whisper), `segments.py`/`text.py`/`timestamps.py` (operações puras sobre transcript); `transcribe.py` é a fachada de compatibilidade |
|
||||
| `collision.py` | 472 | Colisão entre títulos na tela |
|
||||
| `font_metrics.py` | 445 | Largura real de glifos por fonte |
|
||||
| `templates.py` | 387 | Templates de timeline |
|
||||
| `__init__.py` | 139 | Reexporta API pública |
|
||||
| `parser.py` | 367 | FCPXML → objetos Python |
|
||||
| `transcribe.py` | 332 | Transcrição Whisper e corte por texto |
|
||||
| `forced_align.py` | 181 | Alinhamento forçado opcional (whisperx/wav2vec2) que corrige o viés de ~0,4s no início das palavras |
|
||||
| `live.py` | 273 | Modo Live (push_to_fcp) |
|
||||
| `diff.py` | 269 | Comparação de timelines |
|
||||
| `voice_actions.py` | 319 | Decisões de edição (cut/zoom/text/marker) |
|
||||
| `export.py` | 226 | Export Resolve v1.9 + FCP7 XMEML v5 |
|
||||
| `voice_features.py` | 220 | Pitch, energia, ritmo, pausas |
|
||||
| `diarize.py` | 180 | Quem falou (pyannote) |
|
||||
| `media_intel.py` | 177 | Silêncio (ffmpeg) e beats (librosa) |
|
||||
| `emphasis.py` | 133 | Índice de ênfase por palavra |
|
||||
| `safe_xml.py` | 113 | `defusedxml` + `serialize_xml()` |
|
||||
| `dtd.py` | 112 | Validação contra os DTDs da Apple |
|
||||
|
||||
---
|
||||
|
||||
## `models.py` — modelos e enums
|
||||
Single source of truth para estrutura de dados. NUNCA mexa aqui sem rodar
|
||||
`test_models.py`.
|
||||
## `writer/` — edição e escrita
|
||||
|
||||
- **Tempo:** `TimeValue` (fração racional), `Timecode`.
|
||||
- **Clips:** `Clip`, `VideoClip`, `AudioClip`, `ConnectedClip` (lane),
|
||||
`CompoundClip`, `Transition`.
|
||||
- **Contêineres:** `Timeline`, `Project`, `Keyword`.
|
||||
- **Markers:** `Marker`, `MarkerType`, `MarkerColor`, `MARKER_XML_TAGS`.
|
||||
`MarkerType` é o dono da serialização (`from_string`/`from_xml_element`/`xml_attrs`).
|
||||
Match estrito do atributo `completed` (`'0'`/`'1'`, sem padding).
|
||||
- **QC:** `SilenceCandidate`, `FlashFrame`, `GapInfo`, `DuplicateGroup`,
|
||||
`ValidationIssue`, `ValidationResult`.
|
||||
- **Geração:** `SegmentSpec`, `PacingConfig`, `PacingStyle`, `RoughCutResult`.
|
||||
O `FCPXMLModifier` é montado por **composição de mixins**: um mixin por assunto
|
||||
editorial, todos operando sobre o mesmo documento e os mesmos índices.
|
||||
|
||||
## `parser.py` — leitura
|
||||
- `parse_fcpxml(path)` → `Project`.
|
||||
- `FCPXMLParser` — lê spine, connected clips (lanes), secondary storylines, roles.
|
||||
| Módulo | Linhas | Conteúdo |
|
||||
|--------|-------:|----------|
|
||||
| `core.py` | 723 | `ModifierCore`: carga, índices, navegação na spine, `save` |
|
||||
| `titles.py` | 867 | Títulos de texto e legendas dinâmicas |
|
||||
| `cut.py` | 333 | Dividir, cortar faixas, apagar |
|
||||
| `speed.py` | 94 | Velocidade de reprodução |
|
||||
| `zoom.py` | 204 | Zoom (punch-in) via clipe de ajuste conectado |
|
||||
| `helpers.py` | 279 | Sanitização, escalas, construtores de elemento |
|
||||
| `rapid.py` | 240 | Flash frames, rapid trim, preencher buracos |
|
||||
| `validation.py` | 232 | Verificações estruturais antes de salvar |
|
||||
| `compound.py` | 196 | Compound clips: criar e achatar |
|
||||
| `silence.py` | 185 | Detectar e remover silêncio |
|
||||
| `document.py` | 170 | Assets de vídeo, timebases, `write_fcpxml` |
|
||||
| `markers.py` | 165 | Marcadores: um, por timecode, em lote |
|
||||
| `audio.py` | 162 | Clipes de áudio e cama musical |
|
||||
| `generator.py` | 94 | `FCPXMLWriter` — orquestra a criação do zero (estado + delegação) |
|
||||
| `builders.py` | 174 | Um builder por tipo de elemento: `FormatBuilder`, `AssetBuilder`, `MarkerBuilder`, `KeywordBuilder`, `ClipBuilder`, `SequenceBuilder`, `LibraryBuilder` |
|
||||
| `adjustment.py` | 140 | `ClipDeAjuste` — camada de ajuste (filtros `filter-video`/`filter-audio` direto no `<clip>`, sem uso ainda em `server_tools`/`admin/api`) |
|
||||
| `reorder.py` | 126 | Reordenar e recalcular offsets |
|
||||
| `trim.py` | 125 | Aparar e propagar o ripple |
|
||||
| `transitions.py` | 94 | Transições entre vizinhos |
|
||||
| `relink.py` | 94 | Repontar mídia |
|
||||
| `insert.py` | 78 | Inserir clipes na spine |
|
||||
| `modifier.py` | 64 | Monta a classe a partir dos mixins |
|
||||
| `selection.py` | 57 | Selecionar por palavra-chave |
|
||||
| `api.py` | 55 | Atalhos de uma linha |
|
||||
| `connected.py` | 49 | Clipes conectados (lanes) |
|
||||
| `roles.py` | 43 | Atribuir roles |
|
||||
| `reformat.py` | 43 | Reenquadrar resolução |
|
||||
|
||||
## `writer.py` — o coração (3154 linhas)
|
||||
Duas classes principais:
|
||||
**Onde mexer:** ache o assunto na tabela e abra só aquele arquivo. Se a sua
|
||||
mudança precisa de dois mixins ao mesmo tempo, provavelmente o que você quer
|
||||
é um método novo no `core.py` que os dois chamem.
|
||||
|
||||
- **`FCPXMLModifier`** — edita documento existente de forma index-based
|
||||
(dicts de `clips`/`resources`/`formats`), imune a ambiguidade de nomes duplicados.
|
||||
Métodos: `insert_clip`, `add_marker`, `trim_clip`, `delete_clip`, `split_clip`,
|
||||
`change_speed`, `cut_clip_ranges` (usado pela remoção de silêncio), etc.
|
||||
- **`FCPXMLWriter`** — gera FCPXML novo a partir de objetos Python.
|
||||
|
||||
Helpers de nível de arquivo: `modify_fcpxml`, `add_marker_to_file`,
|
||||
`trim_clip_in_file`, `build_marker_element`, `write_fcpxml`, `validate_fcpxml`,
|
||||
`list_effects`, `FCP_EFFECTS`.
|
||||
|
||||
## `rough_cut.py` — geração
|
||||
- `RoughCutGenerator`, `generate_rough_cut`, `generate_segmented_rough_cut`.
|
||||
|
||||
## `media_intel.py` — inteligência de mídia (v0.10)
|
||||
- Silêncio via `ffmpeg silencedetect` (subprocess limitado), `remove_silence_candidates`,
|
||||
mapeamento source→timeline.
|
||||
- Beats via `librosa` (import lazy, extra `[intelligence]`).
|
||||
- Degrada para `None` quando `ffmpeg` ausente.
|
||||
|
||||
## `transcribe.py` — Whisper local
|
||||
- `transcribe(media_path, model_size, language)` → dict com `words` (spans).
|
||||
- `ALLOWED_MODELS` — allowlist de nomes de modelo (também usado por `model_manager`).
|
||||
- Edição por transcrição: remove filler words, aparar por transcrição.
|
||||
|
||||
## `model_manager.py` — gestão de modelos
|
||||
Catálogo `models.json` + cache no HF hub. Config em `~/.fcp-mcp-server/config.json`.
|
||||
Funções: `get/save_models_dir`, `list_installed_models`, `download_model`,
|
||||
`delete_model`, `get/load_selected_model`, `save_selected_model`, `load_catalog`.
|
||||
Permite cancelamento de download via `threading.Event`. Segue convenções:
|
||||
allowlist, lazy imports, degradação graciosa.
|
||||
|
||||
## `export.py` — cross-NLE
|
||||
- `DaVinciExporter` — FCPXML v1.9 p/ DaVinci Resolve.
|
||||
- Export FCP7 XMEML v5.
|
||||
|
||||
## `diff.py` — comparação
|
||||
- `compare_timelines`, `TimelineDiff`, `ClipDiff`, `MarkerDiff`.
|
||||
- Detecta added/removed/moved/trimmed clips & markers.
|
||||
|
||||
## `live.py` — FCP ao vivo (macOS)
|
||||
- `push_to_fcp(path, library, options)` — Apple event *Open Document* + `<import-options>`.
|
||||
Requer `.fcpbundle` p/ zero-click real.
|
||||
- `list_fcp_libraries()` — AppleScript read-only.
|
||||
|
||||
## `templates.py`
|
||||
- `Template`, `TemplateSlot`, `ClipSpec`, `BUILTIN_TEMPLATES`, `apply_template`,
|
||||
`list_templates`. Estruturas prontas: intro/outro, lower thirds, music video.
|
||||
|
||||
## `safe_xml.py`
|
||||
Wrappers `defusedxml` centralizados + `serialize_xml()`. Todo parse/escrita passa aqui.
|
||||
|
||||
## `dtd.py`
|
||||
Valida output contra DTDs oficiais no bundle do FCP (via `xmllint`; exige o caminho
|
||||
do DTD percent-encoded por causa dos espaços em "Final Cut Pro.app").
|
||||
**Cuidado:** os mixins compartilham `self`. Um método novo que colida de nome
|
||||
com outro mixin sobrescreve em silêncio — a ordem em `modifier.py` decide quem
|
||||
ganha. Hoje nenhum colide; mantenha assim.
|
||||
|
||||
---
|
||||
|
||||
## Como adicionar um módulo novo
|
||||
1. Criar `fcpxml/<seu_modulo>.py` — função pura, sem conhecer MCP.
|
||||
2. Reexportar classes/funções em `fcpxml/__init__.py` (`__all__`).
|
||||
3. Cobrir em `tests/test_<seu_modulo>.py`.
|
||||
4. Rodar `./Engine/run_after_fix.sh`.
|
||||
## `models/` — dados e enums
|
||||
|
||||
Fonte única da estrutura de dados. **Nunca mexa aqui sem rodar `test_models.py`.**
|
||||
|
||||
| Módulo | Linhas | Conteúdo |
|
||||
|--------|-------:|----------|
|
||||
| `timing.py` | 304 | `TimeValue` (fração racional), `Timecode` |
|
||||
| `timeline.py` | 217 | `Clip`, `ConnectedClip`, `CompoundClip`, `Timeline`, `Project`, `Marker` |
|
||||
| `enums.py` | 183 | `MarkerType`, `MarkerColor`, `TransitionType`, `PacingStyle`… |
|
||||
| `subtitles.py` | 157 | `WordLook`, `WordStyle`, `DynamicSubtitleConfig`, paleta |
|
||||
| `qc.py` | 121 | `FlashFrame`, `GapInfo`, `DuplicateGroup`, `ValidationIssue` |
|
||||
| `planning.py` | 93 | `SegmentSpec`, `PacingConfig`, `RoughCutResult`, `MontageConfig` |
|
||||
|
||||
`MarkerType` é o dono da serialização de marcador (`from_string`,
|
||||
`from_xml_element`, `xml_attrs`) — não reimplemente isso em outro lugar.
|
||||
|
||||
---
|
||||
|
||||
## O caminho da voz (do áudio à decisão)
|
||||
|
||||
Estes seis módulos formam um pipeline. É o fluxo mais novo e o menos óbvio do
|
||||
projeto, então vale ler nesta ordem:
|
||||
|
||||
```
|
||||
transcribe.py áudio → palavras com tempo
|
||||
+
|
||||
diarize.py quem falou cada trecho
|
||||
+
|
||||
voice_features.py pitch, energia, ritmo, pausas
|
||||
▼
|
||||
emphasis.py combina tudo num índice 0–1 por palavra
|
||||
▼
|
||||
voice_timeline.py monta o _voice_timeline.json ◄── a análise crua
|
||||
▼
|
||||
speaker_review.py (opcional) filtra falante mutado + linha riscada
|
||||
→ _voice_timeline_clean.json ◄── é isto que a IA prefere
|
||||
▼
|
||||
[decisão: skill "editar-por-voz", ou a mão do usuário]
|
||||
▼
|
||||
voice_actions.py valida a lista de cut/zoom/text/marker
|
||||
▼
|
||||
phrase_review.py funde tudo em frases revisáveis (etapa 5 do app)
|
||||
▼
|
||||
writer/ aplica no FCPXML
|
||||
```
|
||||
|
||||
**Regra de ouro do pipeline:** toda ação carrega tempo da **mídia original**,
|
||||
nunca pós-corte. Cortes deslocam tudo depois deles; resolver o deslocamento só
|
||||
na hora de aplicar (`shift_after_cuts`) elimina uma classe inteira de bug.
|
||||
|
||||
### `voice_timeline.py` — o contrato com a IA
|
||||
|
||||
Saída em camadas, para um modelo raciocinar do topo e descer só onde importa:
|
||||
|
||||
```
|
||||
{version, source, language,
|
||||
layers: {transcript, acoustics, speakers, emotion} ← o que rodou de verdade
|
||||
scales: {…} ← como ler cada número
|
||||
summary: {…}
|
||||
speakers: [...]
|
||||
segments: [{start, end, speaker, text, gap_before, take_boundary,
|
||||
avg_energy, peak_emphasis, emotion, emotion_confidence,
|
||||
words: [{text, start, end, energy, pitch_delta, rate_delta,
|
||||
pause_before, emphasis}]}]}
|
||||
```
|
||||
|
||||
`layers` existe para separar *"a fala é monótona"* de *"a análise acústica nunca
|
||||
carregou"* — os dois deixam os mesmos zeros nos dados.
|
||||
|
||||
### `speaker_review.py` — a triagem antes da IA
|
||||
|
||||
Roda logo após `analyze_voice` (etapa 3 do assistente, tela `SpeakerReviewView`
|
||||
no app): lista quem foi detectado (`speaker_profiles`, com % de fala e falas
|
||||
de amostra) e a transcrição segmento a segmento, para o usuário nomear cada
|
||||
falante, mutar quem não interessa (ex.: o entrevistador) e riscar linhas soltas
|
||||
antes de qualquer IA ver o arquivo. `build_speaker_review` nunca toca a
|
||||
timeline crua; `apply_speaker_review`/`write_clean_voice_timeline` produzem
|
||||
uma cópia separada, `_voice_timeline_clean.json`, reaproveitando
|
||||
`enrich_words`/`_segment_rows`/`_summary` de `voice_timeline.py` para
|
||||
recalcular a ênfase só sobre quem sobrou — mesma lógica de `restrict_to_kept`,
|
||||
por falante/segmento em vez de por intervalo de tempo. O merge de decisões
|
||||
salvas segue o padrão de `phrase_review.merge_saved_decisions`: sempre
|
||||
reconstrói da análise atual, só as escolhas humanas persistem.
|
||||
|
||||
A skill "editar-por-voz", `generate_voice_script` e `cmd_build_phrase_review`
|
||||
(etapa 4/5, `admin/api/review.py`) preferem o `_clean` quando ele existe; sem
|
||||
revisão salva, seguem lendo o `_voice_timeline.json` normal — a etapa 3 é
|
||||
sempre opcional. Os três pontos de leitura precisam concordar nessa
|
||||
preferência: se um deles voltar a ler o arquivo cru direto, a revisão de
|
||||
falantes vira letra morta sem nenhum erro visível (ver `05_EXPERIENCIAS.md`).
|
||||
|
||||
Cada linha em `build_speaker_review` carrega suas `words` originais (ênfase
|
||||
por palavra), para a tela desenhar os mesmos chips da etapa 5 sem esperar um
|
||||
recálculo. `apply_speaker_review` é reaproveitada por dois caminhos: gravar
|
||||
(`write_clean_voice_timeline`, via `save_speaker_review`) e só **prever**
|
||||
(`cmd_recalc_speaker_review`, sem tocar disco) — o botão "Recalcular" da tela
|
||||
usa o segundo caminho para atualizar a ênfase só sobre quem sobreviveu ao
|
||||
corte, sem reprocessar áudio.
|
||||
|
||||
### `phrase_review.py` — a revisão humana
|
||||
|
||||
Junta o timeline de voz com as ações da IA numa lista de frases editáveis, e
|
||||
converte de volta. Frase inativa vira `cut`; ênfase ≥ 1 vira `zoom` mais um
|
||||
`emphasis_spans` que a etapa de legendas usa. O trim de cada frase anda em
|
||||
**fronteira de palavra** — cortar é apontar para uma palavra, nunca caçar frame.
|
||||
|
||||
---
|
||||
|
||||
## Legendas dinâmicas (três módulos que andam juntos)
|
||||
|
||||
| Módulo | Papel |
|
||||
|--------|-------|
|
||||
| `text_layout.py` | Quebra a frase em linhas e posiciona cada palavra |
|
||||
| `font_metrics.py` | Largura real de cada glifo na fonte escolhida |
|
||||
| `collision.py` | Detecta título saindo do quadro ou colidindo com outro |
|
||||
|
||||
Estes três não estão divididos porque **cada um já é um assunto só**. O
|
||||
`text_layout.py` tem 901 linhas de um problema coeso: diagramação.
|
||||
|
||||
### Separação de role entre legendas dinâmicas e convencionais
|
||||
|
||||
As duas categorias são ambas `<title>` conectados, mas recebem **roles
|
||||
diferentes** para ficarem didáticas na timeline do FCP (cada role ganha cor
|
||||
própria no índice). O atributo usado em `<title>` é `role` (CDATA) — **nunca**
|
||||
`videoRole`, que é DTD-inválido para títulos (ver `05_EXPERIENCIAS.md`,
|
||||
entrada 32).
|
||||
|
||||
| Categoria | `role` | De onde vem |
|
||||
|-----------|--------|-------------|
|
||||
| Legendas dinâmicas | `titles.dinamicas` | `DynamicSubtitleConfig.role` / `load_dynamic_subtitle_config()["role"]` |
|
||||
| Legendas convencionais | `titles.convencionais` | `load_plain_subtitle_config()["role"]` |
|
||||
|
||||
A cor do texto em si continua nos configs de fonte (abas do app), não no
|
||||
role. Os geradores `generate_dynamic_subtitles` (writer/titles.py),
|
||||
`handle_generate_plain_subtitles` e `handle_generate_subtitles_by_emphasis`
|
||||
(server_tools/subtitles.py) aplicam o role em cada `<title>` criado; o
|
||||
parâmetro `role` das ferramentas MCP sobrescreve o default.
|
||||
|
||||
---
|
||||
|
||||
## Armadilhas do FCPXML (custaram sessões de depuração)
|
||||
|
||||
- Tempo é fração: `"3600/2400s"` = 1,5 s.
|
||||
- `offset` é posição na timeline; `start` é o in-point da mídia.
|
||||
- `<asset-clip>` (biblioteca) é diferente de `<clip>` (timeline).
|
||||
- Marcadores são **filhos** do clipe, não irmãos.
|
||||
- `.fcpxmld` é um **diretório** — sidecars precisam ser copiados no save, ou
|
||||
dados de object tracking e Cinematic são destruídos.
|
||||
- Negrito no FCP é `bold="1"` (atributo); itálico é `fontFace` + `italic="1"`.
|
||||
- `id` de `<text-style-def>` precisa ser XML Name válido — acento, espaço ou
|
||||
dígito inicial fazem o FCP recusar o arquivo inteiro.
|
||||
- `code/examples/sample.fcpxml` **não** é DTD-conformante. Não use como fixture
|
||||
de validade.
|
||||
|
||||
@@ -1,26 +1,49 @@
|
||||
# 03 — Camada MCP (`server.py`) — 73 ferramentas
|
||||
# 03 — Camada MCP (`server.py` + `server_tools/`) — 78 ferramentas
|
||||
|
||||
`server.py` (3824 linhas) é a camada de transporte. Não tem lógica de timeline —
|
||||
mapeia nome → handler e delega ao Engine. O dispatch é um dicionário
|
||||
`TOOL_HANDLERS` (padrão de despacho, sem cadeias gigantes de if/elif).
|
||||
> **Escopo:** As 78 ferramentas MCP: helpers, categorias e como criar uma nova.
|
||||
> **Não cobre:** Lógica de edição, que mora no engine (→ 02) · comandos do app (→ 08)
|
||||
|
||||
`server.py` (592 linhas) é só o transporte: dispatch por dicionário
|
||||
`TOOL_HANDLERS`, sem cadeia de if/elif e **sem lógica de timeline**. Os handlers
|
||||
moram em `server_tools/`, um módulo por categoria, e os helpers que todos usam
|
||||
em `server_tools/_shared/`.
|
||||
|
||||
```
|
||||
server_tools/
|
||||
editing.py (649) qc.py (696) voice.py (754) timeline.py (400)
|
||||
subtitles.py markers_import generation.py transcript.py
|
||||
export.py roles.py live.py
|
||||
_shared/ ← helpers compartilhados, ver abaixo
|
||||
```
|
||||
|
||||
## Helpers centrais (use-os, não reinvente)
|
||||
|
||||
| Helper | Linha | Função |
|
||||
|--------|------:|--------|
|
||||
| `_check_json_depth()` | 83 | Rejeita payloads além de 50 níveis |
|
||||
| `_validate_filepath()` | 103 | Sandbox de entrada |
|
||||
| `_validate_output_path()` | 149 | Sandbox de saída |
|
||||
| `_format_clip_table()` | 245 | Renderização de tabela |
|
||||
| `_markdown_table()` | 259 | Renderização de tabela markdown |
|
||||
| `_parse_project()` | 319 | Parseia FCPXML → `(tree, timeline, project)`; quase todos os handlers começam aqui |
|
||||
| `_resolve_io_paths()` | 357 | Validação de caminho de entrada/saída |
|
||||
| `_setup_modifier()` | 390 | Prepara modifier com validação |
|
||||
| `_setup_generator()` | 414 | Prepara generator com validação |
|
||||
| `_parse_timestamp_parts()` | 433 | Parse de timestamps (min:seg, H:MM:SS, SMPTE) |
|
||||
| `_detect_flash_frames/gaps/duplicate_groups()` | 1667+ | Detectores de QC |
|
||||
Todos reexportados por `server_tools/_shared`, então `from ._shared import X`
|
||||
continua funcionando. A coluna diz o módulo real, para quando você precisar
|
||||
**editar** o helper — ou apontar um `monkeypatch` para ele.
|
||||
|
||||
## As 73 ferramentas por categoria
|
||||
| Helper | Mora em | Função |
|
||||
|--------|---------|--------|
|
||||
| `_validate_filepath()` | `_shared/paths.py` | Sandbox de entrada |
|
||||
| `_validate_output_path()` | `_shared/paths.py` | Sandbox de saída |
|
||||
| `_check_json_depth()` | `_shared/paths.py` | Rejeita payloads além de 50 níveis |
|
||||
| `generate_output_path()` | `_shared/paths.py` | Nome derivado, sem tocar no original |
|
||||
| `_resolve_io_paths()` | `_shared/paths.py` | Entrada + saída de uma vez |
|
||||
| `_parse_project()` | `_shared/project.py` | FCPXML → `(tree, timeline, project)`; quase todo handler começa aqui |
|
||||
| `_setup_modifier()` | `_shared/project.py` | Prepara modifier já validado |
|
||||
| `_setup_generator()` | `_shared/project.py` | Prepara generator já validado |
|
||||
| `_text_result()` | `_shared/project.py` | Envolve o texto em `TextContent` MCP |
|
||||
| `_markdown_table()` | `_shared/formatting.py` | Tabela markdown |
|
||||
| `_format_clip_table()` | `_shared/formatting.py` | Tabela de clipes |
|
||||
| `_format_batch_result()` | `_shared/formatting.py` | Relatório de operação em lote |
|
||||
| `_parse_timestamp_parts()` | `_shared/captions.py` | min:seg, H:MM:SS, SMPTE |
|
||||
| `parse_srt()` / `parse_vtt()` | `_shared/captions.py` | Legendas coladas |
|
||||
| `_detect_flash_frames/gaps/duplicate_groups()` | `_shared/detection.py` | Detectores de QC |
|
||||
| `_load_or_transcribe()` | `_shared/media.py` | Transcrição com cache em disco |
|
||||
| `_cut_transcript_spans()` | `_shared/media.py` | Corte por trecho falado |
|
||||
| `_apply_placed_action()` | `_shared/media.py` | Aplica zoom/text/marker já posicionado |
|
||||
|
||||
## As 77 ferramentas por categoria
|
||||
|
||||
### Timeline & análise (Projeto)
|
||||
`list_projects`, `analyze_timeline`, `list_clips`, `list_markers`, `list_connected_clips`,
|
||||
@@ -59,18 +82,44 @@ mapeia nome → handler e delega ao Engine. O dispatch é um dicionário
|
||||
|
||||
### Voz (análise → decisão → aplicação)
|
||||
`analyze_voice_features`, `build_voice_timeline`, `refine_voice_timeline`,
|
||||
`remove_speakers`, `apply_voice_actions`, `get_voice_analysis_config`,
|
||||
`remove_speakers`, `remove_speech_gaps`, `apply_voice_actions`,
|
||||
`generate_voice_script`, `get_voice_analysis_config`,
|
||||
`save_voice_analysis_config`.
|
||||
|
||||
O fluxo é sempre o mesmo: `build_voice_timeline` mede (caro, roda uma vez) →
|
||||
`remove_speech_gaps` corta pelo que a **transcrição** já sabe que não tem
|
||||
fala — lê `words[].start/end` do `_voice_timeline.json` (função
|
||||
`speech_gap_cut_actions`, em `fcpxml/voice_actions.py`) em vez de medir
|
||||
volume. É o complemento correto para o caso que `remove_media_silence`
|
||||
(silêncio por dB, ver seção "Silêncio e beats") não cobre: um trecho sem
|
||||
fala mas com som real acima do limiar (respiração, ruído de roupa, batida) —
|
||||
`remove_media_silence` nunca vai cortar isso porque tecnicamente não é
|
||||
silêncio. Não corta a lacuna antes da primeiríssima palavra (pode ser quase
|
||||
o arquivo inteiro, antes da tomada realmente começar) — isso continua
|
||||
decisão manual na Fase 6 do `apply_voice_actions`
|
||||
(`.claude/skills/editar-por-voz/criterios/06-texto-corte-marcador.md`).
|
||||
|
||||
O fluxo manual é: `build_voice_timeline` mede (caro, roda uma vez) →
|
||||
o modelo decide os cortes → **`refine_voice_timeline` renormaliza sobre o que
|
||||
sobrou** (barato, sem reabrir áudio) e propõe as janelas de zoom → o modelo
|
||||
corta a lista pelo ritmo → `apply_voice_actions` aplica. Pular a renormalização
|
||||
faz o ranking de ênfase apontar para as palavras erradas (ver
|
||||
`05_EXPERIENCIAS.md`).
|
||||
|
||||
`generate_voice_script` é o fluxo **automático e fechado** (sem wizard, sem
|
||||
copiar-e-colar): transcreve (cache) → `build_voice_timeline` → entrega a
|
||||
timeline a um **modelo local Ollama** que dirige a edição → devolve o roteiro
|
||||
legível (markdown) **e** o JSON de ações, e opcionalmente aplica num FCPXML.
|
||||
O cliente fica em `fcpxml/llm_local.py`; o modelo é tratado como entrada não
|
||||
confiável e cada ação é validada por `parse_actions`. Padrão:
|
||||
`qwen2.5:7b-instruct-q4_K_M` (troca de `gemma3:12b` — não cabia em máquina de
|
||||
8GB de RAM; Gemma 3 4B foi testado antes e falhou por apagar o roteiro
|
||||
principal em vez de só cortar bastidor). Passe `model=` para usar outro
|
||||
servido pelo Ollama.
|
||||
|
||||
### Legendas dinâmicas (geração → validação → aplicação)
|
||||
`generate_dynamic_subtitles`, `validate_subtitle_layout`, `transcript_markers`.
|
||||
`generate_dynamic_subtitles`, `generate_plain_subtitles`,
|
||||
`generate_subtitles_by_emphasis`, `validate_subtitle_layout`,
|
||||
`transcript_markers`.
|
||||
|
||||
**Sempre gere e depois valide — nunca dê a geração como pronta sem
|
||||
`validate_subtitle_layout`.** A composição garante "sem sobreposição" só
|
||||
@@ -88,6 +137,49 @@ severidade probable/severe → investigar CADA colisão pela fração exata do
|
||||
XML antes de mudar código (ver checklist abaixo)
|
||||
```
|
||||
|
||||
**`generate_subtitles_by_emphasis`** gera as duas legendas numa passada só,
|
||||
dividindo por palavra: a dinâmica cobre as frases marcadas como ênfase na
|
||||
etapa 5 (zoom aplicado, nível ≥ 1); a comum cobre **todo o resto** — um bloco
|
||||
comum simplesmente não é criado onde a dinâmica já cobre. A primeira versão
|
||||
gerava a comum inteira e desativava (`enabled="0"`) o que ficava sob a
|
||||
dinâmica, mas um título desativado continua aparecendo como clipe riscado na
|
||||
timeline do Final Cut mesmo sem renderizar — um corte com bastante ênfase
|
||||
enchia a trilha de clipes mortos. Trocado por não gerar ali: o preço é que,
|
||||
se a ênfase for desativada à mão depois, a legenda comum daquele trecho
|
||||
precisa ser regenerada, não só reativada. É a tradução de `10-revisao-humana.md`
|
||||
(skill `editar-por-voz`): "a frase de ênfase recebe zoom E legenda dinâmica;
|
||||
as demais recebem legenda comum". A decisão vem de
|
||||
`<mídia>_phrase_actions.json["emphasis_spans"]`, escrito por
|
||||
`save_phrase_review` quando o editor termina a etapa 5 — sem esse arquivo (ou
|
||||
sem `zoom`/`text` marcados na revisão), a tool gera só a comum, tudo ligado,
|
||||
e avisa no relatório ("Sem revisão de ênfase"). Não expõe overrides de estilo
|
||||
por chamada — usa a config salva ("Legendas Dinâmicas"/plain); para estilo
|
||||
pontual, use `generate_dynamic_subtitles`/`generate_plain_subtitles` direto.
|
||||
|
||||
**Separação por role (didática na timeline):** `generate_dynamic_subtitles` e
|
||||
`generate_plain_subtitles` (e a metade dinâmica/comum do `by_emphasis`) aplicam
|
||||
`role="titles.dinamicas"` e `role="titles.convencionais"` em cada `<title>`
|
||||
criado — sub-roles de `titles`, **nunca** `subtitles.*` (que esconderia o título
|
||||
atrás de Code). O parâmetro `role` de cada ferramenta MCP sobrescreve o default
|
||||
(vindo de `load_dynamic_subtitle_config()["role"]` /
|
||||
`load_plain_subtitle_config()["role"]`). Ver `02_MODULES.md` (seção "Separação
|
||||
de role") e `05_EXPERIENCIAS.md` entrada 32 (DTD: `<title>` leva `role`, não
|
||||
`videoRole`).
|
||||
|
||||
**Compound clip por sub-frase (padrão em `generate_dynamic_subtitles` e na
|
||||
metade dinâmica do `by_emphasis`):** `compound_subphrases=True` divide cada
|
||||
frase em sub-frases pela vírgula (`transcribe.split_into_subphrases`) e
|
||||
empacota os `<title>` de cada uma num `<ref-clip>` — a dúzia de títulos
|
||||
empilhados por lane que uma frase gera vira uma barra só, arrastável/mutável
|
||||
como unidade. Exceção: um trecho curto depois da vírgula ("né?", "Então...",
|
||||
< 3 palavras) funde de volta na sub-frase anterior em vez de virar compound
|
||||
próprio — soa como parte da mesma respiração, não uma frase nova. A estrutura
|
||||
replica o que o próprio Final Cut gera em "New Compound Clip": o título mais
|
||||
cedo vira âncora do spine interno em offset 0, os demais penduram nele por
|
||||
lane. `validate_subtitle_layout` mede cada compound no seu próprio espaço de
|
||||
tempo — sem isso, âncoras de compounds diferentes leem "0s" e colidem no
|
||||
papel mesmo estando segundos distantes na timeline real.
|
||||
|
||||
**Antes de atribuir uma colisão ao gerador, confirme que é o gerador.**
|
||||
Um `<title>` de nome estranho (`ref` diferente, params tipo `Auto-Shrink`/
|
||||
`Left Margin` que `_make_text_title_clip` nunca escreve) é conteúdo humano
|
||||
@@ -144,7 +236,18 @@ Regras:
|
||||
- Sempre retornam via `_text_result(text)` (envolve o texto em `TextContent` MCP).
|
||||
|
||||
## Para adicionar uma ferramenta nova
|
||||
1. Escrever a função no módulo do Engine (`fcpxml/…`) + testes.
|
||||
2. Criar `handle_<nome>` em `server.py` seguindo o padrão acima.
|
||||
3. Registrar no dicionário `TOOL_HANDLERS`.
|
||||
|
||||
1. **Escrever a função no Engine** (`fcpxml/…`) com testes. É aqui que mora o
|
||||
trabalho de verdade; o resto é encanamento.
|
||||
2. **Criar `handle_<nome>`** em `server_tools/<categoria>.py`, seguindo o padrão
|
||||
acima. Escolha a categoria pelo assunto, não pelo tamanho do arquivo.
|
||||
3. **Declarar o schema** (`Tool(...)`) no mesmo módulo.
|
||||
4. **Registrar** no `TOOL_HANDLERS` de `server.py`.
|
||||
5. Rodar `./Engine/run_after_fix.sh`.
|
||||
|
||||
Se a ferramenta também deve aparecer no app, exponha um comando equivalente em
|
||||
`admin/api/<assunto>.py` e registre na tabela de `admin/models_api.py` — ver
|
||||
`08_APP_MACOS.md`. Uma capacidade que só existe como tool MCP **não existe para
|
||||
quem usa o app** (foi exatamente o que aconteceu com `apply_voice_actions`,
|
||||
`05_EXPERIENCIAS.md` #20).
|
||||
4. Rodar `./Engine/run_after_fix.sh`.
|
||||
@@ -1,5 +1,8 @@
|
||||
# 04 — Testes, Fluxo de Trabalho e Estado Atual
|
||||
|
||||
> **Escopo:** Como rodar e escrever testes, e o gate antes de commitar.
|
||||
> **Não cobre:** O que testar em cada módulo (→ 02) · checklist de qualidade (→ 06)
|
||||
|
||||
## 1. Suíte de testes
|
||||
|
||||
**1032 testes em 24 arquivos** em `tests/`. Rode com `uv run pytest tests/ -v`.
|
||||
|
||||
@@ -11,6 +11,56 @@ houver uma correção ou trabalho em torno dele, **adicione um registro aqui**
|
||||
antes de prosseguir. Um problema que se repete em várias tentativas é sinal de
|
||||
que merece entrada.
|
||||
|
||||
|
||||
> **Como usar:** o índice abaixo é o ponto de entrada. Procure o sintoma
|
||||
> aqui primeiro; só abra a entrada completa (mais abaixo) se ela for a sua.
|
||||
> As entradas ficam em ordem cronológica depois do índice.
|
||||
|
||||
## Resumo rápido (índice)
|
||||
|
||||
| # | Data | Problema | Estado |
|
||||
|---|------|----------|--------|
|
||||
| 1 | 2026-08-14 | Início do registro de experiências | `resolvido` |
|
||||
| 4 | 2026-08-14 | XML fora da grade de frame em NTSC (23.976/29.97fps), confirmado no FCP | `resolvido` |
|
||||
| 5 | 2026-08-14 | `TimeValue.from_timecode` corrompia segundos decimais em NTSC (3º ponto do bug) | `resolvido` |
|
||||
| 6 | 2026-08-17 | Clipe-fantasma de 1 frame no início/fim após remoção de silêncio (padding sem vizinho na borda) | `resolvido` |
|
||||
| 7 | 2026-08-17 | Legendas dinâmicas sobrepondo entre clipes (título conectado não é aparado pelo out-point do pai) | `resolvido` |
|
||||
| 8 | 2026-08-17 | Importação recusada: `id` de `<text-style-def>` derivado do texto (acentos/espaços/dígito inicial) não é XML Name válido | `resolvido` |
|
||||
| 9 | 2026-08-17 | Legendas palavra a palavra centradas em vez da composição progressiva diagramada (bloco por trecho, palavra-chave em display italic) | `resolvido` |
|
||||
| 10 | 2026-08-17 | Cedilha/acentos da display italic invadindo a linha vizinha: empilhamento passou a usar a tinta real por classe de glifo | `resolvido` |
|
||||
| 11 | 2026-08-18 | Preview das legendas dinâmicas desproporcional ao render do FCP (stagger/gap/canvas divergentes) e `inactive_color` exposto sem efeito | `resolvido` |
|
||||
| 12 | 2026-08-18 | Espaço de coordenadas do modelo "Text": `fontSize`, `kerning` e `Position` no espaço do quadro — converter só o tamanho descolou o espaçamento | `resolvido` |
|
||||
| 13 | 2026-08-19 | Reanálise de ênfase implementada no Engine mas sem ferramenta MCP — Fase 4 da skill era inexecutável | `resolvido` |
|
||||
| 14 | 2026-08-19 | Offset sistemático de ~0,4s no timing por palavra (faster-whisper sem alinhamento forçado) — agora corrigido em pipeline por alinhamento forçado opcional | `resolvido` |
|
||||
| 15 | 2026-08-19 | `add_zoom` perdia o enquadramento real (voltava a 100%) quando dois zooms caiam no mesmo clipe pós-corte; agora empilha ou substitui conforme as janelas se sobrepõem | `resolvido` |
|
||||
| 16 | 2026-08-19 | `validate_subtitle_layout` acusava colisão severa em títulos que só se tocam na borda, por não-associatividade de float; 7 de 8 colisões reportadas no teste real eram falso positivo | `resolvido` |
|
||||
| 17 | 2026-08-19 | Linha de ênfase das legendas dinâmicas sem limite de largura — palavra longa/maiúscula estourava o frame inteiro; auto-fit encolhe até caber, nunca abaixo do corpo | `resolvido` |
|
||||
| 18 | 2026-08-19 | Legendas dinâmicas geradas com `bold="0" fontFace="Bold"` não renderizam no FCP — negrito deve ser `bold="1"` (atributo) e itálico `fontFace`+`italic="1"` | `resolvido` |
|
||||
| 19 | 2026-08-19 | `output_dir` usado só como cerca de validação e nunca como destino — toda chamada entre pastas falhava acusando o caminho que ela mesma gerou | `resolvido` |
|
||||
| 20 | 2026-08-19 | `apply_voice_actions` ausente da ponte e do encadeamento do app — dava para analisar e legendar, não para cortar | `resolvido` |
|
||||
| 21 | 2026-08-19 | Teste ainda afirmava o default `zoom scale=1.3` removido do parser (agora vem do `zoom_scale` do usuário) | `resolvido` |
|
||||
| 22 | 2026-08-19 | `VideoPlayer` (AVKit) aborta em runtime no app compilado por `swiftc` — etapa 5 fechava o app; trocado por `AVPlayerLayer` | `resolvido` |
|
||||
| 23 | 2026-08-19 | Dividir `writer.py` em pacote quebrou `@patch('fcpxml.writer.subprocess')` — a suíte protege comportamento, não localização | `resolvido` |
|
||||
| 24 | 2026-08-19 | `admin/test_models_api.py` existia mas estava fora de `testpaths` — 13 testes que nunca rodaram | `resolvido` |
|
||||
| 25 | 2026-08-20 | `admin/api/shared.py` apontava para `admin/code` (inexistente) após a divisão — install editável mascarou o bug em toda validação anterior | `resolvido` |
|
||||
| 26 | 2026-08-21 | `generate_voice_script` (IA local/Ollama) caía com "Falha ao gerar roteiro por IA local" — prompt embutia a timeline inteira (47k tokens) e estourava `num_ctx`; e `response.json()` de conexão caída escapava como `JSONDecodeError` | `resolvido` |
|
||||
| 27 | 2026-08-21 | Cortes escritos rente ao timestamp da palavra soam secos — critério da skill e prompt do modelo local não instruíam folga na borda | `resolvido` |
|
||||
| 28 | 2026-08-21 | Frases desativadas em sequência deixavam fatias de 0,1-0,5s sobrando entre clipes | `resolvido` |
|
||||
| 29 | 2026-08-21 | `remove_media_silence` (dB) não corta lacuna sem fala mas com som real — trecho sobrevivia intacto na timeline final | `resolvido` |
|
||||
| 30 | 2026-08-24 | `add_zoom` animava `<adjust-transform>` direto no clipe, diferente de como o FCP realmente exporta zoom (clipe de ajuste conectado) | `resolvido` |
|
||||
| 31 | 2026-08-24 | Revisão de falantes ("Quem fica na edição") salvava certo, mas etapa 4 (Revisão de frases) lia a timeline crua, ignorando falantes mutados/linhas riscadas | `resolvido` |
|
||||
| 32 | 2026-08-24 | Separar legendas dinâmicas de convencionais por role: `<title>` aceita `role` (CDATA), NÃO `videoRole` — este último é DTD-inválido para títulos e quebra a validação | `resolvido` |
|
||||
| 33 | 2026-09-22 | Legenda comum sobreposta à composição dinâmica em `generate_subtitles_by_emphasis`; regenerar acumulava títulos em vez de substituir | `resolvido` |
|
||||
| 34 | 2026-09-22 | `ClipDeAjuste` gerava wrapper `<adjustment>` inválido no DTD; `code/WHISPERX` era 2,6 GB de backup órfão que inflava o lint quando rodado com `--exclude` explícito | `resolvido` |
|
||||
| 35 | 2026-09-22 | Indexação RAG (`admin/update_rag.py`) abortava a transação inteira ao achar um chunk que estoura o contexto do modelo de embedding | `resolvido` |
|
||||
| 36 | 2026-09-23 | `.gitignore` com regra `models/` solta escondia do git o pacote inteiro `fcpxml/models/` (dados do engine), não só o cache do Whisper | `resolvido` |
|
||||
|
||||
> Mantenha o índice acima sempre sincronizado com as entradas mais recentes.
|
||||
|
||||
---
|
||||
|
||||
## Entradas (ordem cronológica)
|
||||
|
||||
---
|
||||
|
||||
## Como registrar (template de entrada)
|
||||
@@ -79,9 +129,9 @@ Use o bloco abaixo como modelo. Uma entrada = um problema resolvido/reconhecido.
|
||||
- **Por que isso importa mais do que parece:** o erro contamina toda decisão temporal a jusante — zoom disparava ~0,4s antes da palavra-alvo, `gap_before` subestimava pausas reais na mesma medida (o que afeta diretamente a régua de silêncio recém-adotada), e as folgas de corte saíam erradas nas emendas.
|
||||
- **Decisão tomada:** não rodei `remove_media_silence` bruto sobre o corte. A detecção (ffmpeg, limiar -30dB/0,5s) não distingue "batida entre frases dentro da régua de 1,5s" de "ar morto de emenda" — cortar ambos teria apertado frases fluidas. Corrigi os tempos manualmente medindo o ataque real nos pontos críticos (cabeça, 2 emendas, cauda, 3 zooms) e refiz o corte numa passada só.
|
||||
- **Solução adotada (paliativa, aplicada manualmente neste teste):** medir o RMS real com `ffmpeg -af astats=metadata=1:reset=1:length=0.05,ametadata=print` em janelas curtas ao redor de cada ponto crítico antes de fixar um corte ou zoom que dependa de precisão de frame. Não é o padrão do sistema — é o que cobre a lacuna até o alinhamento forçado existir.
|
||||
- **Solução estrutural ainda pendente:** ligar o WhisperX (ou alinhamento forçado equivalente) em `transcribe.py`, o que levaria o erro de ~400ms para ~30ms e corrigiria zoom, corte e `gap_before` de uma vez, sem paliativo por projeto. Não implementado ainda — é mudança de pipeline, exige regerar todos os `_transcript.json`/`_voice_timeline.json` existentes.
|
||||
- **Solução estrutural implementada:** `transcribe.py` agora roda alinhamento forçado fonético (wav2vec2 via whisperx) como passo opcional pós-transcrição, em `fcpxml/forced_align.py` (classe `ForcedAligner`). O erro cai de ~400ms para ~30ms e corrige zoom, corte e `gap_before` de uma vez. É **dependência opcional** (`[align]` extra / pacote `whisperx` do PyPI) — quando ausente ou em qualquer falha, degrada e devolve os tempos brutos sem quebrar a transcrição. O `transcript` traz `"alignment": true/false` e o `voice_timeline` expõe `layers.alignment`, para quem lê o JSON saber se o offset manual ainda é necessário. Não reaproveitamos código da pasta `WHISPERX/` local (problemas conhecidos) — só a ideia documentada aqui. Exige regerar os `_transcript.json`/`_voice_timeline.json` existentes para aplicar nos caches antigos.
|
||||
- **Aprendizado:** "não reestime tempos no olho" (critério 01) continua certo para decisão *editorial* — mas não cobre erro sistemático de *medição* na fonte dos tempos. Um offset constante e na mesma direção, em vários pontos do material, é sinal de bug no pipeline de transcrição, não de julgamento errado sobre o material. Vale conferir com uma amostra de áudio real antes de confiar cegamente em timestamp de word-level de qualquer fonte nova.
|
||||
- **Estado:** `parcialmente resolvido` — paliativo documentado e aplicado neste teste; correção estrutural (WhisperX) pendente de implementação.
|
||||
- **Estado:** `resolvido` — alinhamento forçado implementado em `transcribe.py`/`fcpxml/forced_align.py`; paliativo de medição manual mantido apenas para transcripts antigos sem `layers.alignment=true`.
|
||||
|
||||
---
|
||||
|
||||
@@ -1183,27 +1233,663 @@ o outro; percentil entrega um punhado útil nos dois casos.
|
||||
|
||||
---
|
||||
|
||||
## Resumo rápido (índice)
|
||||
## 21 — 2026-08-19 — Teste travado no default antigo de `zoom scale`
|
||||
|
||||
| # | Data | Problema | Estado |
|
||||
|---|------|----------|--------|
|
||||
| 1 | 2026-08-14 | Início do registro de experiências | `resolvido` |
|
||||
| 4 | 2026-08-14 | XML fora da grade de frame em NTSC (23.976/29.97fps), confirmado no FCP | `resolvido` |
|
||||
| 5 | 2026-08-14 | `TimeValue.from_timecode` corrompia segundos decimais em NTSC (3º ponto do bug) | `resolvido` |
|
||||
| 6 | 2026-08-17 | Clipe-fantasma de 1 frame no início/fim após remoção de silêncio (padding sem vizinho na borda) | `resolvido` |
|
||||
| 7 | 2026-08-17 | Legendas dinâmicas sobrepondo entre clipes (título conectado não é aparado pelo out-point do pai) | `resolvido` |
|
||||
| 8 | 2026-08-17 | Importação recusada: `id` de `<text-style-def>` derivado do texto (acentos/espaços/dígito inicial) não é XML Name válido | `resolvido` |
|
||||
| 9 | 2026-08-17 | Legendas palavra a palavra centradas em vez da composição progressiva diagramada (bloco por trecho, palavra-chave em display italic) | `resolvido` |
|
||||
| 10 | 2026-08-17 | Cedilha/acentos da display italic invadindo a linha vizinha: empilhamento passou a usar a tinta real por classe de glifo | `resolvido` |
|
||||
| 11 | 2026-08-18 | Preview das legendas dinâmicas desproporcional ao render do FCP (stagger/gap/canvas divergentes) e `inactive_color` exposto sem efeito | `resolvido` |
|
||||
| 12 | 2026-08-18 | Espaço de coordenadas do modelo "Text": `fontSize`, `kerning` e `Position` no espaço do quadro — converter só o tamanho descolou o espaçamento | `resolvido` |
|
||||
| 13 | 2026-08-19 | Reanálise de ênfase implementada no Engine mas sem ferramenta MCP — Fase 4 da skill era inexecutável | `resolvido` |
|
||||
| 14 | 2026-08-19 | Offset sistemático de ~0,4s no timing por palavra (faster-whisper sem alinhamento forçado) — corrigido manualmente no teste, WhisperX pendente | `parcialmente resolvido` |
|
||||
| 15 | 2026-08-19 | `add_zoom` perdia o enquadramento real (voltava a 100%) quando dois zooms caiam no mesmo clipe pós-corte; agora empilha ou substitui conforme as janelas se sobrepõem | `resolvido` |
|
||||
| 16 | 2026-08-19 | `validate_subtitle_layout` acusava colisão severa em títulos que só se tocam na borda, por não-associatividade de float; 7 de 8 colisões reportadas no teste real eram falso positivo | `resolvido` |
|
||||
| 17 | 2026-08-19 | Linha de ênfase das legendas dinâmicas sem limite de largura — palavra longa/maiúscula estourava o frame inteiro; auto-fit encolhe até caber, nunca abaixo do corpo | `resolvido` |
|
||||
| 18 | 2026-08-19 | Legendas dinâmicas geradas com `bold="0" fontFace="Bold"` não renderizam no FCP — negrito deve ser `bold="1"` (atributo) e itálico `fontFace`+`italic="1"` | `resolvido` |
|
||||
| 19 | 2026-08-19 | `output_dir` usado só como cerca de validação e nunca como destino — toda chamada entre pastas falhava acusando o caminho que ela mesma gerou | `resolvido` |
|
||||
| 20 | 2026-08-19 | `apply_voice_actions` ausente da ponte e do encadeamento do app — dava para analisar e legendar, não para cortar | `resolvido` |
|
||||
- **Sintoma:** `tests/test_voice_actions.py::test_default_scale_when_absent`
|
||||
quebrando com `KeyError: 'scale'`, sem relação com a alteração em curso.
|
||||
- **Causa raiz:** `parse_actions` deixou de carimbar `scale=1.3` quando o
|
||||
parâmetro vem ausente, justamente para que
|
||||
`server_tools/_shared.py` use o `zoom_scale` configurado pelo usuário. O
|
||||
teste continuou afirmando o default antigo, então passou a acusar como erro
|
||||
exatamente o comportamento desejado.
|
||||
- **Solução adotada:** teste reescrito para o contrato novo — um `scale`
|
||||
omitido tem que chegar ausente ao aplicador (`test_absent_scale_is_left_absent`).
|
||||
- **Aprendizado:** quando um default sai do parser e vira configuração, o teste
|
||||
que afirmava o valor antigo passa a defender o bug. Ao remover um default,
|
||||
procure o teste que o fixava no mesmo commit — senão ele fica dizendo o
|
||||
contrário do código, e a próxima pessoa perde tempo achando que quebrou algo.
|
||||
- **Estado:** `resolvido`
|
||||
|
||||
> Mantenha o índice acima sempre sincronizado com as entradas mais recentes.
|
||||
---
|
||||
|
||||
## 22 — 2026-08-19 — `VideoPlayer` (AVKit) derruba o app compilado por `swiftc`
|
||||
|
||||
- **Sintoma:** "G-ART encerrou inesperadamente" (SIGABRT) toda vez que o
|
||||
assistente entrava na etapa 5. Nada aparecia na tela antes do crash.
|
||||
- **Causa raiz:** o app é montado invocando `swiftc` direto
|
||||
(`MacApp/build_app.sh`), não pelo Xcode. Nesse modo o runtime não consegue
|
||||
resolver a superclasse Objective-C de `VideoPlayer`:
|
||||
`failed to demangle superclass of VideoPlayerView from mangled name
|
||||
'So12AVPlayerViewC'` → `getSuperclassMetadata` chama `fatalError`. É erro de
|
||||
runtime, então a compilação passa limpa e o problema só aparece ao abrir a
|
||||
view.
|
||||
- **Solução adotada:** trocar `VideoPlayer` por um `AVPlayerLayer` dentro de um
|
||||
`NSViewRepresentable` (`PlayerSurface`/`PlayerLayerView` em
|
||||
`PhraseReviewView.swift`). Só depende de AVFoundation, que linka normalmente.
|
||||
Os controles de transporte já viviam na barra da timeline, então não se perde
|
||||
nada com a chrome do AVKit.
|
||||
- **Aprendizado:** compilar limpo não prova que um componente de framework
|
||||
existe em runtime neste build. Ao usar uma view SwiftUI que embrulha uma
|
||||
classe AppKit/ObjC (AVKit, WebKit, MapKit), abra a tela de fato antes de
|
||||
concluir. Um harness pequeno (`swiftc` com os mesmos fontes + um `@main` que
|
||||
monta só aquela view e sai) reproduz o crash em segundos, sem precisar
|
||||
navegar o app inteiro até lá.
|
||||
- **Estado:** `resolvido`
|
||||
|
||||
---
|
||||
|
||||
## 23 — 2026-08-19 — Dividir um módulo em pacote quebra quem faz `patch` nele
|
||||
|
||||
- **Sintoma:** ao transformar `fcpxml/writer.py` (4.199 linhas) no pacote
|
||||
`fcpxml/writer/`, quatro testes passaram a falhar com
|
||||
`AttributeError: module 'fcpxml.writer' has no attribute 'subprocess'` —
|
||||
embora nenhuma linha de lógica tivesse mudado.
|
||||
- **Causa raiz:** os testes usavam `@patch('fcpxml.writer.subprocess.run')`.
|
||||
Isso não depende da API pública, e sim de *onde o import mora*: com o
|
||||
módulo dividido, `subprocess` passou a ser importado por
|
||||
`fcpxml/writer/document.py`, então o alvo do patch deixou de existir.
|
||||
Re-exportar no `__init__` não resolveria — substituir
|
||||
`fcpxml.writer.subprocess` não afeta a referência que `document` já tem.
|
||||
- **Solução adotada:** apontar o patch para o módulo real
|
||||
(`fcpxml.writer.document.subprocess.run`). Duas armadilhas do tipo foram
|
||||
evitadas antes: imports relativos precisam de um ponto a mais ao descer um
|
||||
nível (`from .models` → `from ..models`), inclusive os que ficam *dentro*
|
||||
de funções, e o `__all__` precisa listar os nomes com underscore que o
|
||||
resto do projeto já importava, senão a divisão vira quebra de API.
|
||||
- **Aprendizado:** a suíte protege comportamento, não localização. Antes de
|
||||
dividir um módulo, procure por `patch('<modulo>.` e por imports relativos
|
||||
escondidos dentro de funções — são as duas coisas que uma refatoração
|
||||
puramente mecânica quebra em silêncio, e as únicas que os testes pegam
|
||||
tarde.
|
||||
- **Estado:** `resolvido`
|
||||
|
||||
---
|
||||
|
||||
## 24 — 2026-08-19 — Teste existia, mas estava fora da suíte
|
||||
|
||||
- **Sintoma:** `admin/test_models_api.py` (13 testes) nunca rodava. Não
|
||||
falhava — simplesmente não era coletado, então `models_api.py` figurava
|
||||
como "coberto" sem que uma única asserção fosse executada em nenhum
|
||||
commit.
|
||||
- **Causa raiz:** `testpaths = ["tests"]` no `pyproject.toml`, com o pytest
|
||||
rodando de `code/`. O arquivo morava em `admin/`, fora do alcance. Rodá-lo
|
||||
à mão também falhava (`ModuleNotFoundError: admin`), porque a raiz do
|
||||
repositório não entra no `sys.path` — ou seja, o único jeito de executá-lo
|
||||
exigia saber de antemão que ele existia e como.
|
||||
- **Solução adotada:** movido para `code/tests/test_models_api.py`, com o
|
||||
insert da raiz do repositório no `sys.path` ao lado do import que precisa
|
||||
dele. Passou a rodar no gate: 1441 → 1454 testes.
|
||||
- **Aprendizado:** um teste fora de `testpaths` é pior que teste nenhum — ele
|
||||
dá a sensação de rede sem ser rede. Ao mover ou criar teste fora da pasta
|
||||
padrão, confirme que a contagem total subiu; se não subiu, ele não está
|
||||
rodando. Vale também para o lint: `admin/` ainda não é coberto pelo
|
||||
`run_after_fix.sh`, que roda só dentro de `code/`.
|
||||
- **Estado:** `resolvido`
|
||||
|
||||
---
|
||||
|
||||
## 25 — 2026-08-20 — `admin/api/shared.py` apontava para `admin/code` (inexistente)
|
||||
|
||||
- **Sintoma:** app do usuário crashava em toda ação que passa por `server`
|
||||
(ex: "Analisar voz"), com `ModuleNotFoundError: No module named
|
||||
'server_tools'`. Sobreviveu a **duas rodadas de validação minha** na sessão
|
||||
anterior — lint zero, 1454 testes verdes, comando testado manualmente pela
|
||||
ponte — sem nenhuma delas pegar o bug.
|
||||
- **Causa raiz:** ao dividir `admin/_shared.py` (#25 da sessão de refatoração,
|
||||
commit `ffaebb3`) em `admin/api/*.py`, o cálculo
|
||||
`Path(__file__).resolve().parent.parent / "code"` foi copiado sem ajuste.
|
||||
No arquivo original (`admin/models_api.py`, direto em `admin/`), dois
|
||||
`.parent` chegam na raiz do repo. Em `admin/api/shared.py`, um nível mais
|
||||
fundo, dois `.parent` param em `admin/` — e `admin/code` nunca existiu.
|
||||
`sys.path` nunca recebia `code/`, então `import server_tools` (que só
|
||||
funciona com `code/` no path) falhava assim que qualquer handler tentava
|
||||
`from server import ...`.
|
||||
- **Por que passou pela validação anterior:** todo teste que exercitava esse
|
||||
caminho importava `admin.api.*` **dentro do processo do pytest**, que já
|
||||
roda com `cwd=code/` sob um venv com **install editável**
|
||||
(`__editable__.fcp_mcp_server*.pth`) — isso já deixa `fcpxml`/`server_tools`
|
||||
importáveis por conta própria, mascarando qualquer erro no cálculo manual
|
||||
de `sys.path`. O teste manual pela ponte (`uv run python
|
||||
admin/models_api.py analyze_voice ...`) tem o mesmo problema: `uv run`
|
||||
ativa o mesmo venv com o mesmo install editável. **Só o app real, chamando
|
||||
o fallback `python3` sem `uv` ou um venv sem o install editável, expõe o
|
||||
bug** — que é exatamente a diferença entre o ambiente de teste e o do
|
||||
usuário.
|
||||
- **Solução adotada:** o cálculo de `sys.path` saiu de cada módulo de
|
||||
comando e passou a existir **uma única vez**, em `admin/api/__init__.py`
|
||||
— que roda antes de qualquer submódulo do pacote, então nenhum deles
|
||||
precisa da própria cópia. `.parent.parent.parent` (três níveis: `api/` →
|
||||
`admin/` → raiz → `code/`).
|
||||
- **Como o teste de regressão foi validado (e por que precisou de duas
|
||||
tentativas):** a primeira versão do teste também passava com o bug
|
||||
presente, pelo mesmo motivo do parágrafo acima — rodava em processo com o
|
||||
install editável ativo. Só ficou confiável rodando um `subprocess` limpo
|
||||
que remove manualmente qualquer entrada `site-packages` de `sys.path`
|
||||
antes de importar, isolando o mecanismo real que o `__init__.py` precisa
|
||||
fornecer. Confirmado nos dois sentidos: falha com o bug reintroduzido,
|
||||
passa com a correção (`tests/test_models_api.py::TestCodeDirResolution`).
|
||||
- **Aprendizado:** um install editável no venv de teste é uma segunda fonte
|
||||
de verdade que mascara bugs de `sys.path` — o mesmo defeito de "a suíte
|
||||
passa mas o comportamento real não bate" da entrada #23, só que desta vez
|
||||
nem *rodar o comando manualmente* pegou, porque o `uv run` usado para
|
||||
testar caía no mesmo venv "de sorte" que o app não usa. Ao validar correção
|
||||
de caminho/import, rodar num ambiente que não tenha as dependências
|
||||
instaladas por fora do mecanismo sendo testado — ou o teste prova que o
|
||||
ambiente de teste está bem configurado, não que o código está certo.
|
||||
- **Estado:** `resolvido`
|
||||
|
||||
---
|
||||
|
||||
## Entrada #26 — Prompt da IA local estoura o contexto do Ollama (e erro de parse escapa)
|
||||
|
||||
- **Sintoma:** botão "Gerar roteiro por IA local" (etapa 4 do assistente)
|
||||
devolvia "Falha ao gerar roteiro por IA local". Rodando a ponte direto, o
|
||||
erro real aparecia como *"Server disconnected without sending a response"*
|
||||
ou *"Connection refused"* do Ollama, e 0 decisões ("Decisões do modelo: 0").
|
||||
- **Causa raiz (dupla):**
|
||||
1. `build_edit_messages` embutia o JSON da voice timeline **inteiro** no
|
||||
prompt. Uma gravação de 3min vira ~188KB / **~47k tokens** (cada palavra
|
||||
carrega energia, pitch, arousal, valence, `samples`…). Como `num_ctx`
|
||||
estava em 32768, o prompt estourava a janela e o Ollama **dropava a
|
||||
conexão** sem resposta.
|
||||
2. Quando a conexão cai sem resposta, `httpx` entrega um body vazio e
|
||||
`response.json()` lançava `JSONDecodeError` — que **não** é
|
||||
`httpx.HTTPError`, então escapava do `try/except` de `ollama_chat` e
|
||||
virava a exceção genérica que o `cmd_generate_voice_script` transforma
|
||||
em `ok:false` com a mensagem "Falha ao gerar roteiro por IA local: …".
|
||||
- **Correção (em `fcpxml/llm_local.py` + `server_tools/voice.py`):**
|
||||
- `build_edit_messages` agora projeta a timeline (**`_project_timeline`**):
|
||||
mantém só `text`/`start`/`end`/`speaker`/`emphasis`/`pause_before` das
|
||||
palavras e `id`/`name` dos locutores; descarta `layers`, `scales`,
|
||||
`samples` e os floats de áudio. Caiu de ~47k para **~17k tokens** (69KB).
|
||||
- Salvaguarda `_shrink_to_fit`: se ainda passar de `max_chars` (110k),
|
||||
remove os `words` dos segmentos de menor `peak_emphasis` até caber.
|
||||
- `ollama_chat` envolve `post`+`raise_for_status`+`json()` num único
|
||||
`except Exception` que relança como `RuntimeError` claro — fim do
|
||||
`JSONDecodeError` escapando.
|
||||
- `_extract_json` agora desembrulha a lista de 1 elemento `[{source,
|
||||
actions}]` que alguns modelos devolvem, senão o `parse_actions` tratava o
|
||||
objeto-wrapper como uma ação sem `kind` e rejeitava tudo (0 decisões).
|
||||
- `handle_generate_voice_script` levanta `RuntimeError` com a causa quando o
|
||||
modelo não devolve nenhuma decisão utilizável, então o app mostra a
|
||||
mensagem real ("O modelo local não devolveu decisões utilizáveis: …")
|
||||
em vez do genérico.
|
||||
- **Validação:** `tests/test_llm_local.py` ganhou `test_build_edit_messages_is_compact`
|
||||
(prompt < raw, sem `samples`/`energy_raw`/`pitch_hz`) e
|
||||
`test_ollama_chat_wraps_empty_response`. Ponte testada com Ollama mockado
|
||||
nos dois sentidos (sucesso aplica; falha → `ok:false` com msg clara).
|
||||
- **Estado:** `resolvido`
|
||||
|
||||
> **Aprendizado:** modelo local tem contexto finito — nunca embutir o objeto
|
||||
> de análise cru no prompt; projetar só o que a decisão usa. E qualquer parse
|
||||
> de resposta de servidor local deve tratar body vazio/quebrado como erro de
|
||||
> transporte, não como sucesso mudo.
|
||||
|
||||
---
|
||||
|
||||
### 2026-08-21 — Cortes escritos rente ao timestamp da palavra soam secos
|
||||
|
||||
- **Sintoma:** usuário revisou o corte final (projeto Mastopexia) e reportou
|
||||
"os cortes estão muito secos, principalmente no final de frase — falta um
|
||||
tempinho a mais pra concluir as palavras". Também notou que o ar morto
|
||||
antes da primeira fala do vídeo não tinha sido cortado.
|
||||
- **Causa:** o critério `06-texto-corte-marcador.md` (e o prompt embutido do
|
||||
modelo local em `fcpxml/llm_local.py`) instruíam cobrir a frase inteira
|
||||
(`start..end = início..fim da frase`) ao escrever um `cut`, sem nenhuma
|
||||
orientação sobre a borda que encosta em fala **mantida** (não em silêncio
|
||||
puro). Um `cut` com `start` exatamente no fim da última palavra mantida
|
||||
engole essa palavra antes dela terminar de soar; um `cut` com `end` no
|
||||
início exato da próxima engole o ataque da fala seguinte. É um problema
|
||||
diferente de cortar a pausa curta (proibido, é a própria ênfase) — aqui a
|
||||
pausa natural entre os blocos já existe, e o corte estava comendo essa
|
||||
margem sozinho.
|
||||
- **Correção:**
|
||||
- `06-texto-corte-marcador.md` ganhou a seção "Nunca corte rente à
|
||||
palavra — deixe uma folga": recuar `start`/`end` do corte em ~0,15–0,25s
|
||||
para dentro do próprio corte nas bordas que tocam fala mantida (não em
|
||||
silêncio puro), incluindo o início/fim do vídeo.
|
||||
- `fcpxml/llm_local.py::_SYSTEM_PROMPT` (item 4) recebeu a mesma
|
||||
instrução, para o modelo local gerar decisões já com a folga.
|
||||
- **Validação manual:** reaplicado no projeto Mastopexia real —
|
||||
`10.77 → 95.50` (rente) virou `10.97 → 95.30` (folga de ~0,2s nas duas
|
||||
pontas), e as 4 emendas seguintes receberam o mesmo tratamento; zoom/texto/
|
||||
marcador continuaram longe o suficiente da nova borda do corte — a folga
|
||||
também evita o problema relacionado (não corrigido em código, só
|
||||
contornado manualmente nesta sessão): um `zoom`/`marker` cuja borda cai
|
||||
exatamente em cima do início/fim de um `cut` é descartado por
|
||||
`resolve_actions` como "apontando para material cortado", mesmo quando a
|
||||
intenção era ficar bem ao lado. Vale registrar como dívida: `resolve_actions`
|
||||
poderia tolerar uma margem de meio-frame antes de considerar a ação "dentro"
|
||||
do corte.
|
||||
- **Estado:** `resolvido`
|
||||
|
||||
> **Aprendizado:** "cobrir a frase inteira" não é a instrução completa para
|
||||
> um corte — a frase que **sobra** ao lado do corte também precisa de uma
|
||||
> borda que respire. Regra prática: só cortar rente ao timestamp quando a
|
||||
> borda encosta em silêncio real (`gap_before` grande) ou em conteúdo que
|
||||
> também será descartado; encostando em fala mantida, sempre recuar.
|
||||
|
||||
---
|
||||
|
||||
### 2026-08-21 — Frases desativadas em sequência deixavam fatias de 0,1-0,5s sobrando
|
||||
|
||||
- **Sintoma:** usuário viu, no Final Cut, um clipe minúsculo sobrando entre
|
||||
dois clipes normais na timeline (projeto Mastopexia, confirmado por
|
||||
screenshot). Investigação achou 29 `cut`s individuais no
|
||||
`_phrase_actions.json` gerado pela etapa 5, e a timeline final saiu com
|
||||
mais de uma dezena de fatias de 0,1-0,5s entre clipes.
|
||||
- **Causa:** `phrase_review_to_actions()` (`fcpxml/phrase_review.py`) gerava
|
||||
**um `cut` por frase desativada**, cobrindo só `[phrase.start, phrase.end]`.
|
||||
Quando duas ou mais frases seguidas estão desativadas, a pausa **entre**
|
||||
elas nunca pertence a nenhuma frase — não é coberta por nenhum `cut` — e
|
||||
sobrevive como um clipe próprio, minúsculo, que ninguém pediu para manter.
|
||||
- **Correção:** `phrase_review_to_actions()` agora agrupa frases desativadas
|
||||
**consecutivas** (`flush_inactive_run()`) e emite um único `cut` cobrindo do
|
||||
início da primeira ao fim da última do grupo, absorvendo as pausas entre
|
||||
elas. Uma frase ativa no meio ainda quebra o grupo — cuts continuam
|
||||
separados quando há conteúdo mantido entre eles.
|
||||
- **Validação:** `tests/test_phrase_review.py` ganhou
|
||||
`test_consecutive_inactive_phrases_merge_into_one_cut`,
|
||||
`test_inactive_run_at_the_end_still_flushes` e
|
||||
`test_isolated_inactive_phrases_stay_separate_cuts`. No projeto Mastopexia
|
||||
real, 29 cuts individuais viraram 3 cuts mescladas; a contagem de fatias
|
||||
sub-segundo na timeline final caiu de mais de uma dezena para 4 (resíduo
|
||||
menor, provavelmente do padding do `remove_media_silence` na emenda entre
|
||||
clipes — não investigado a fundo nesta sessão, ver `09_MANUTENCAO.md`).
|
||||
- **Estado:** `resolvido` (a causa principal); a sobra residual do
|
||||
`remove_media_silence` continua como dívida separada.
|
||||
|
||||
> **Aprendizado:** "cortar cada frase desativada" não é a mesma coisa que
|
||||
> "cortar o trecho desativado" quando frases se sucedem sem conteúdo mantido
|
||||
> entre elas — a pausa entre duas coisas descartadas também precisa ser
|
||||
> descartada, e ninguém a cobre por definição se o corte for por frase.
|
||||
|
||||
---
|
||||
|
||||
### 2026-08-21 — `remove_media_silence` (dB) não pega lacuna sem fala com som real
|
||||
|
||||
- **Sintoma:** usuário viu, no projeto Mastopexia real, um trecho de ~1,9s
|
||||
sem fala (imagem parada antes da tomada começar) que sobreviveu intacto
|
||||
na timeline final — depois de `apply_voice_actions`, `remove_media_silence`
|
||||
e `generate_dynamic_subtitles` já terem rodado. Achou que era bug de ordem
|
||||
no encadeamento das etapas ("corta e depois volta").
|
||||
- **Investigação:** não era ordem. Extraído o áudio real do trecho
|
||||
(`ffmpeg -af volumedetect`): `mean_volume -21.4dB`, `max_volume 0.0dB` —
|
||||
longe do limiar padrão de silêncio (-30dB). Rodado `detect_silence` nos
|
||||
mesmos limiares do sistema (-30/-25/-20/-16dB): nenhum sinaliza o trecho.
|
||||
O trecho tem som real (roupa, respiração, ambiente) mas nenhuma palavra —
|
||||
exatamente o caso que `06-texto-corte-marcador.md` já descrevia
|
||||
("ausência de fala não é ausência de som"), só que sem ferramenta para
|
||||
agir sobre ele: `remove_media_silence` só enxerga volume, nunca vai
|
||||
cortar algo que soa alto mas não tem fala.
|
||||
- **Correção:** nova função pura `speech_gap_cut_actions()` em
|
||||
`fcpxml/voice_actions.py` — gera `cut`s a partir dos gaps entre
|
||||
`words[].start/end` do `_voice_timeline.json` (tempo de fonte, como todo
|
||||
`VoiceAction`), com a mesma folga por dentro (`padding`) que
|
||||
`speaker_cut_actions()` já usava. Nova tool MCP `remove_speech_gaps`
|
||||
(`server_tools/voice.py`, mesmo molde de `remove_speakers`): resolve o
|
||||
`media_path`, lê a timeline, gera as ações e reaplica via
|
||||
`handle_apply_voice_actions` — não duplica a lógica de corte no FCPXML.
|
||||
Deliberadamente não corta a lacuna antes da primeiríssima palavra (pode
|
||||
ser quase o arquivo inteiro, antes da tomada começar de verdade).
|
||||
- **Ordem revista:** `apply_voice_actions → remove_speech_gaps →
|
||||
remove_media_silence → generate_dynamic_subtitles` — a lacuna "sem fala"
|
||||
some primeiro (cobertura ampla, por transcrição), o que sobra de silêncio
|
||||
técnico *dentro* da fala é apertado depois.
|
||||
- **Validação:** `tests/test_voice_actions.py::TestSpeechGapCutActions`
|
||||
(gap acima/abaixo do limiar, lacuna antes da 1ª palavra nunca cortada,
|
||||
segmentos com palavras sobrepostas não quebram, timeline vazia). Suíte
|
||||
completa (1508 testes) roda limpa.
|
||||
- **Estado:** `resolvido`
|
||||
|
||||
> **Aprendizado:** um detector de silêncio por dB nunca vai cobrir "sem fala
|
||||
> com som" — são categorias diferentes, não uma questão de calibrar o
|
||||
> limiar. Quando já existe transcrição confiável, ela é a fonte melhor para
|
||||
> "onde não tem fala": não depende de threshold nenhum, só da própria
|
||||
> palavra existir ou não naquele instante.
|
||||
|
||||
---
|
||||
|
||||
### 2026-08-24 — Zoom era `<adjust-transform>` no próprio clipe; FCP exporta como clipe de ajuste
|
||||
|
||||
- **Sintoma:** usuário pediu para o zoom parar de mexer diretamente no
|
||||
clipe da timeline e passar a usar um "adjustment clip" com crop
|
||||
animado — o jeito como ele já fazia zoom manualmente no FCP.
|
||||
- **Investigação:** não havia amostra real no projeto para confirmar a
|
||||
forma exata do XML (`adjust-crop`? um `<clip>` com `<adjustment>` como
|
||||
`fcpxml/writer/adjustment.py` já fazia para filtros?). O usuário enviou
|
||||
um `.fcpxmld` exportado pelo próprio FCP com um zoom manual
|
||||
(`exemplo zoom.fcpxmld`), que revelou a forma real: um `<video ref="...">`
|
||||
referenciando o efeito nativo `FFAdjustmentEffect` ("Clipe de Ajuste"),
|
||||
anexado numa lane acima do clipe, com seu **próprio** `<adjust-transform>`
|
||||
animando `scale` de `1 1` até o pico — não `adjust-crop`, e não o wrapper
|
||||
`<adjustment>` que `adjustment.py` usa (que, conferido contra o DTD real
|
||||
da Apple, **não existe** — aquele módulo gera XML inválido; ver dívida
|
||||
em `09_MANUTENCAO.md`). Cruzado com o DTD oficial (`FCPXMLv1_13.dtd`, uma
|
||||
cópia local encontrada fora do projeto): `<video>` é `%anchor_item;`
|
||||
válido sem precisar de asset, e `adjust-transform` é filho direto seu.
|
||||
- **Correção:** `add_zoom` (extraído para `fcpxml/writer/zoom.py`, deixou
|
||||
de compartilhar módulo com `change_speed`) agora cria um `<video>`
|
||||
conectado em vez de animar o clipe base. Isso **simplificou** a lógica
|
||||
antiga: como o clipe de ajuste composita por cima da imagem já
|
||||
reenquadrada, não precisa mais ler/preservar rotação, posição ou escala
|
||||
do clipe original (a classe de teste inteira sobre "preservar
|
||||
enquadramento" — e o bug histórico #15 que ela cobria — deixou de fazer
|
||||
sentido); e dois zooms disjuntos no mesmo clipe agora são dois `<video>`
|
||||
irmãos, não um merge de keyframes num `<adjust-transform>` só.
|
||||
- **Validação:** os 22 testes de zoom em `test_writer.py` reescritos contra
|
||||
a nova forma (`clip.find('video').find('adjust-transform')...`), mais
|
||||
`test_voice_actions_tool.py`. Offset/duration da timeline gerada
|
||||
conferidos byte a byte contra os números reais do `.fcpxmld` de exemplo
|
||||
(bateram exatamente). Suíte completa roda limpa.
|
||||
- **Estado:** `resolvido`
|
||||
|
||||
> **Aprendizado:** para decisões de forma exata de XML, um exemplo real
|
||||
> exportado pelo próprio FCP vale mais que qualquer inferência — a diferença
|
||||
> entre `adjust-crop`, o wrapper inválido de `adjustment.py` e a forma real
|
||||
> (`<video ref="FFAdjustmentEffect">`) não dava para cravar sem um dos dois
|
||||
> (amostra real ou o DTD oficial da Apple, que também foi cruzado aqui).
|
||||
> Peça o exemplo antes de implementar às cegas.
|
||||
|
||||
---
|
||||
|
||||
### 2026-08-24 — Revisão de falantes salvava certo, mas a etapa 4 nunca lia o resultado
|
||||
|
||||
- **Sintoma:** usuário desmarcou falas de bastidor na tela "Quem fica na edição"
|
||||
(etapa 3, `SpeakerReviewView`) e clicou "Salvar seleção", mas as falas
|
||||
desmarcadas continuavam voltando na revisão de frases (etapa 4) e no roteiro
|
||||
final gerado a partir dela.
|
||||
- **Causa raiz:** `save_speaker_review` (`fcpxml/speaker_review.py`) e o
|
||||
`_voice_timeline_clean.json` que ela grava estavam **corretos** — conferido
|
||||
num projeto real: 37 segmentos na timeline crua, 15 marcados `excluded` na
|
||||
revisão salva, 22 sobrando no `_clean.json` (37-15=22, bate exato). O bug
|
||||
estava um passo adiante: `cmd_build_phrase_review`
|
||||
(`admin/api/review.py`), que monta a etapa 4, abria
|
||||
`args.get("voice_timeline")` — o arquivo **cru** — direto, sem nunca checar
|
||||
se existia um `_voice_timeline_clean.json` ao lado. `generate_voice_script`
|
||||
(`server_tools/voice.py`) e `copyForChat` (`WizardView.swift`) já faziam
|
||||
essa checagem corretamente; só a etapa 4 ficou de fora.
|
||||
- **Onde:** `admin/api/review.py::cmd_build_phrase_review`.
|
||||
- **Por que passou despercebido:** a tela de revisão de falantes em si
|
||||
funcionava e mostrava "Salvo" — o problema só aparecia num passo seguinte
|
||||
e sem nenhum erro, então parecia que "a seleção não estava sendo salva"
|
||||
quando na verdade ela salvava certo e era ignorada mais adiante.
|
||||
- **Solução adotada:** `cmd_build_phrase_review` agora resolve
|
||||
`speaker_review.clean_voice_timeline_path(timeline_path)` primeiro e lê
|
||||
esse arquivo quando ele existe, caindo para o cru só na ausência dele —
|
||||
mesma checagem que os outros dois pontos já faziam.
|
||||
- **Aprendizado:** quando existem **múltiplos pontos de leitura** de um
|
||||
mesmo artefato derivado (aqui: três lugares que podem preferir
|
||||
`_voice_timeline_clean.json` sobre o cru), adicionar a checagem em um novo
|
||||
ponto de leitura não é opcional — ela precisa ser replicada em todos, ou o
|
||||
comportamento diverge silenciosamente conforme o caminho que o app tomar.
|
||||
Vale grepar por todo lugar que abre o arquivo "canônico" sempre que um
|
||||
arquivo "_clean"/derivado for introduzido.
|
||||
- **Estado:** `resolvido` — corrigido em `admin/api/review.py`, suíte
|
||||
completa (1506 de 1508 testes; as 2 falhas restantes são de ambiente —
|
||||
WhisperX/torchcodec sem libs de sistema, sem relação com a mudança) e
|
||||
lint do arquivo alterado limpos.
|
||||
|
||||
---
|
||||
|
||||
### Entrada 32 — 2026-08-24: `<title>` leva `role`, nunca `videoRole`
|
||||
|
||||
**Sintoma:** ao atribuir role de vídeo a legendas geradas (para separar
|
||||
legendas dinâmicas de convencionais na timeline), a validação contra o DTD
|
||||
FCPXML v1.13 quebrou com `No declaration for attribute videoRole of element
|
||||
title`.
|
||||
|
||||
**Causa:** no DTD da Apple, `<title>` (`<!ATTLIST title %clip_attrs;>` +
|
||||
`<!ATTLIST title role CDATA #IMPLIED>`) **não** declara `videoRole`. Esse
|
||||
atributo existe em `<video>`, `<asset-clip>`, `<clip>` etc., mas não em
|
||||
títulos. `<title>` usa o atributo genérico `role` (CDATA). Confirmado no
|
||||
`FCPXMLv1_13.dtd` linhas 566–569.
|
||||
|
||||
**Decisão:** legendas dinâmicas e convencionais recebem `role="titles.dinamicas"`
|
||||
e `role="titles.convencionais"` (sub-roles de `titles`, NUNCA `subtitles.*` —
|
||||
ver entrada sobre roteamento de captions). O campo de config e o parâmetro dos
|
||||
geradores chama-se `role` (não `video_role`). `assign_role` (mixin `RolesMixin`)
|
||||
continua correto para clips/vídeos, pois seta `videoRole` neles — não confundir
|
||||
os dois caminhos.
|
||||
|
||||
**Lição:** antes de setar `videoRole` num elemento qualquer, conferir o DTD:
|
||||
títulos usam `role`. Teste de regressão em `tests/test_dynamic_subtitles.py`
|
||||
(`test_titles_carry_title_subrole`) garante `titles.*` e bloqueia `subtitles.*`.
|
||||
|
||||
> **Nota de reconciliação:** entradas antigas deste arquivo (2026-08-17)
|
||||
> afirmavam "nenhum título gerado carrega `role`" e tinham o teste
|
||||
> `test_titles_carry_no_caption_role`. Aquilo referia-se **especificamente**
|
||||
> a `role="subtitles.*"` (que roteia o título para a pista de captions e o
|
||||
> esconde). A regra continua válida: proibido `subtitles.*`. O que mudou é que
|
||||
> agora aplicamos `role="titles.*"` (sub-role de título, válido no DTD e útil
|
||||
> para separar dinâmicas de convencionais na timeline). O teste foi renomeado
|
||||
> para `test_titles_carry_title_subrole` e passa a exigir `titles.*` + bloquear
|
||||
> `subtitles.*`.
|
||||
|
||||
---
|
||||
|
||||
### Entrada 33 — 2026-09-22: legenda comum sob a composição dinâmica; regenerar acumulava títulos
|
||||
|
||||
**Sintoma:** num corte real (Mastopexia), aos 11s a legenda comum "mamas
|
||||
também mudam. É" aparecia simultaneamente com a composição dinâmica de
|
||||
ênfase, poluindo o quadro com texto duplicado. Gerar novamente as legendas
|
||||
(dinâmica ou convencional) sobre um clipe já legendado empilhava um segundo
|
||||
conjunto de títulos por cima do anterior em vez de substituí-lo.
|
||||
|
||||
**Causa raiz — duas falhas distintas:**
|
||||
1. **Sem marcação de autoria.** Os três handlers de legenda
|
||||
(`handle_generate_dynamic_subtitles`, `handle_generate_plain_subtitles`,
|
||||
`handle_generate_subtitles_by_emphasis`) só *adicionavam* títulos —
|
||||
nenhum removia o que uma chamada anterior tinha gerado. Sem uma forma de
|
||||
distinguir "título que este programa gerou" de "título que o editor
|
||||
inseriu manualmente no FCP", uma regeneração não tinha como saber o que é
|
||||
seguro apagar.
|
||||
2. **Janela da legenda de ênfase maior que a fala.** Em
|
||||
`handle_generate_subtitles_by_emphasis`, o cálculo de fim de bloco usava
|
||||
os segmentos brutos do Whisper (`data["segments"]`) para decidir até onde
|
||||
a composição dinâmica se estende — não os spans de ênfase revisados
|
||||
(`spans`). Um segmento do Whisper cobre a frase inteira; a ênfase cobre só
|
||||
o trecho grifado. A dinâmica então ficava "seguindo" além do próprio
|
||||
áudio que a originou, invadindo o intervalo onde a legenda comum já
|
||||
deveria estar sozinha.
|
||||
- **Onde:** `code/fcpxml/writer/titles.py` (`TitlesMixin`) e
|
||||
`code/server_tools/subtitles.py` (os três handlers de geração).
|
||||
- **Solução adotada:**
|
||||
- Todo título/composição gerado por este programa carrega uma marca em
|
||||
`<metadata><md key="com.gart.subtitle.kind" value="dynamic|plain">`
|
||||
(`mark_generated_subtitle`). Um heurístico de compatibilidade
|
||||
(`_generated_subtitle_kind`) reconhece a assinatura exata de exports
|
||||
antigos sem a marca (efeito/uid/start de texto do G-ART + padrão de nome),
|
||||
para não tratar título manual do editor como "nosso" por engano.
|
||||
- Cada handler chama `remove_generated_subtitles(el, kinds)` no início,
|
||||
apagando só os títulos com a marca do próprio tipo que está sendo
|
||||
regerado — títulos manuais e do outro tipo ficam intactos.
|
||||
- `generate_dynamic_subtitles` ganhou o parâmetro `hold_between_sentences`
|
||||
(default `True`, preserva o comportamento anterior nas chamadas normais).
|
||||
`handle_generate_subtitles_by_emphasis` passa `hold_between_sentences=False`
|
||||
e usa os `spans` de ênfase revisados como `emphasis_segments` (em vez dos
|
||||
segmentos brutos do Whisper) — a composição dinâmica agora encerra no fim
|
||||
real da palavra falada quando o próximo bloco pertence a outra frase, e
|
||||
nunca ultrapassa a janela de ênfase que a gerou.
|
||||
- `suppress_plain_under_dynamic` recorta (fatiando o clipe do título, sem
|
||||
duplicar `text-style`) qualquer legenda comum gerada cujo intervalo caia
|
||||
dentro de uma composição dinâmica ainda ativa — mesmo que o cálculo de
|
||||
janela de algum outro caminho volte a divergir no futuro, isso funciona
|
||||
como rede de segurança contra sobreposição visível.
|
||||
- **Aprendizado:** um gerador que pode ser chamado de novo sobre a mesma
|
||||
timeline **precisa** de uma forma de reconhecer sua própria saída anterior
|
||||
antes de decidir "substituir" — sem isso, "regerar" e "empilhar" são
|
||||
indistinguíveis. E ao derivar o fim de uma janela temporal a partir de uma
|
||||
fonte (segmentos do Whisper, spans de ênfase, etc.), confirme que a fonte
|
||||
escolhida tem a granularidade do fenômeno que está sendo delimitado — usar
|
||||
a fonte "mais larga disponível" por conveniência cria sobra sistemática.
|
||||
- **Teste de regressão:**
|
||||
`code/tests/test_subtitle_overlap_regression.py` — roda o handler real
|
||||
(`handle_generate_subtitles_by_emphasis`) contra um intervalo de ênfase
|
||||
seguido de uma lacuna de fala comum, e confere que nenhuma composição
|
||||
dinâmica sobrepõe uma legenda comum; e que chamar o mesmo handler duas
|
||||
vezes não duplica títulos gerados nem remove um título manual inserido
|
||||
entre as duas chamadas.
|
||||
- **Estado:** `resolvido` — 224 testes das suítes de legenda/writer
|
||||
passando (incl. o novo regressivo); suíte completa 1540 passando, 8
|
||||
skipped, 1 falha e 1 erro de ambiente sem relação com a mudança (WhisperX/
|
||||
`extract_pitch` ausente, torchcodec sem libs de sistema); lint dos arquivos
|
||||
alterados limpo.
|
||||
|
||||
---
|
||||
|
||||
### Entrada 34 — 2026-09-22: `<adjustment>` inválido no DTD e `WHISPERX` órfão inflando o lint
|
||||
|
||||
**Sintoma 1:** `fcpxml/writer/adjustment.py` (`ClipDeAjuste`, código de uma
|
||||
sessão anterior não commitado) montava
|
||||
`<clip><adjustment><filter-video .../></adjustment></clip>` para camadas de
|
||||
ajuste. Nada usava o módulo ainda (sem chamada em `server_tools`/
|
||||
`admin/api`), mas ficava pronto para alguém reusar do jeito errado.
|
||||
|
||||
**Causa 1:** o DTD real da Apple (`FCPXMLv1_13.dtd`) não define nenhum
|
||||
elemento `<adjustment>`. A produção real de `<clip>` é
|
||||
`(note?, %timing-params;, %intrinsic-params;, (spine|(%clip_item;)|caption)*,
|
||||
(%marker_item;)*, audio-channel-source*, (%video_filter_item;)*,
|
||||
filter-audio*, metadata?)` — ou seja, `filter-video`/`filter-audio` são
|
||||
filhos diretos do `<clip>`, sem wrapper, e nessa ordem (vídeo antes de
|
||||
áudio).
|
||||
|
||||
**Solução 1:** `ClipDeAjuste.criar()` agora anexa os filtros direto no
|
||||
`<clip>`, ordenados com vídeo antes de áudio
|
||||
(`sorted(filtros, key=lambda f: f.tag != "filter-video")`). Teste de
|
||||
regressão novo: `tests/test_writer_adjustment.py` (sem wrapper, ordem
|
||||
correta, um `<effect>` por `uid` em `resources`).
|
||||
|
||||
**Sintoma 2 (achado ao investigar o mesmo módulo):** um `ruff check .
|
||||
--exclude docs/` rodado manualmente no início desta sessão acusou **510
|
||||
erros** — muito acima do que a suíte normalmente reporta.
|
||||
|
||||
**Causa 2:** `code/WHISPERX` era uma pasta `.git` solta de **2,6 GB** dentro
|
||||
de `code/` (não um submodule registrado — sem `.gitmodules`), contendo
|
||||
cópias/backups congelados do próprio projeto, incluindo uma cópia inteira e
|
||||
antiga de `fcp-mcp-server-main` dentro de si mesma. O `pyproject.toml` já
|
||||
excluía `WHISPERX/` do lint por padrão (`[tool.ruff] exclude = ["docs/",
|
||||
"WHISPERX/"]`), mas passar `--exclude docs/` na linha de comando
|
||||
**sobrescreve** esse `exclude` em vez de complementá-lo — foi assim que o
|
||||
lint passou a varrer os 2,6 GB de código velho lá dentro. Confirmado por
|
||||
grep que só 3 arquivos no código ativo referenciam "WHISPERX", todos em
|
||||
comentários explicativos (`fcpxml/diarize.py`, `tests/test_diarize.py`,
|
||||
`admin/api/shared.py`) — nenhum import ou caminho real dependia da pasta.
|
||||
|
||||
**Solução 2:** pasta movida para `~/Archives/G-ART-WHISPERX-backup` (fora do
|
||||
workspace git), copiada com `rsync -a --no-perms` e conferida com
|
||||
`diff -rq` antes de remover o original. `WHISPERX/` também saiu do
|
||||
`exclude` do ruff em `code/pyproject.toml` (não faz mais sentido excluir um
|
||||
caminho que não existe mais em `code/`).
|
||||
|
||||
**Aprendizado:** (1) um wrapper de elemento "que faz sentido conceitualmente"
|
||||
não substitui checar o DTD real antes de escrever o gerador — o padrão do
|
||||
projeto (`dtd.py`, DTDs em `bm/*/FCPXMLv1_13.dtd`) existe exatamente para
|
||||
isso. (2) uma flag de linha de comando como `--exclude` em ferramentas de
|
||||
lint tipicamente **substitui** a config do projeto, não a estende — rodar
|
||||
`ruff check .` sem flags (herdando `pyproject.toml`) é o comando correto
|
||||
para refletir o gate real; qualquer variação manual com `--exclude` pode
|
||||
mentir sobre o estado do lint. (3) uma pasta de backup improvisada dentro do
|
||||
diretório ativo do projeto (mesmo que "só para não perder nada") é dívida
|
||||
que cresce sem ninguém perceber — 2,6 GB não apareceram de uma vez.
|
||||
|
||||
**Estado:** `resolvido` — `tests/test_writer_adjustment.py` (3 testes)
|
||||
passando; `admin/` trazido ao lint gate no mesmo commit (ver
|
||||
`09_MANUTENCAO.md` §2.3); suíte completa 1543 passando, 8 skipped, 1 falha
|
||||
+ 1 erro pré-existentes de outro trabalho em andamento (sem relação com
|
||||
esta correção).
|
||||
|
||||
### Entrada 35 — 2026-09-22: chunk grande demais derrubava a indexação RAG inteira
|
||||
|
||||
**Sintoma:** `admin/update_rag.command` (primeira indexação completa do
|
||||
G-ART, banco `rag_gart` recém-provisionado) morria sempre no mesmo ponto com
|
||||
`requests.exceptions.HTTPError: 500 Server Error` na chamada ao Ollama —
|
||||
sempre logo após imprimir `code/fcpxml/export.py`, ou seja, no arquivo
|
||||
seguinte na ordem alfabética.
|
||||
|
||||
**Causa raiz:** `code/fcpxml/font_metrics.py` é uma tabela de larguras de
|
||||
glifo (`METRICS = {...}`), texto extremamente denso em tokens (muitos
|
||||
números/pontuação curtos) — um chunk de ~4900 caracteres (dentro do limite
|
||||
`CHUNK_MAX_CHARS = 5000`) virou 2653 tokens no tokenizer do
|
||||
`nomic-embed-text`, estourando o contexto de 2048 tokens do servidor Ollama
|
||||
local (`llama.cpp`: "input length exceeds the context length"). Reproduzido
|
||||
isolando o arquivo e chamando `/api/embeddings` chunk a chunk — 6 dos 9
|
||||
chunks falhavam. `CHUNK_MAX_CHARS` mede caracteres, não tokens; assume
|
||||
implicitamente ~1 token por poucos caracteres, o que não vale para conteúdo
|
||||
não-prosa (tabelas numéricas, JSON denso).
|
||||
|
||||
Segundo problema, apontado por que a primeira tentativa não recuperou nada:
|
||||
`admin/update_rag.py::index()` roda a varredura inteira (centenas de
|
||||
arquivos) em **uma única transação**, com `commit()` só no fim e
|
||||
`rollback()` em qualquer exceção — um único chunk problemático em um único
|
||||
arquivo descartava a indexação inteira, mesmo que os outros 300+ arquivos
|
||||
já tivessem embedado e inserido com sucesso.
|
||||
|
||||
**Solução:** `_embed()` agora detecta essa resposta específica do Ollama
|
||||
(`ChunkTooLarge`, checado por `500` + `"context length"` no corpo) e o loop
|
||||
principal captura essa exceção por chunk, pula só aquele chunk (aviso em
|
||||
stderr) e continua o arquivo — sem abortar a transação. Não trunca nem
|
||||
reduz `CHUNK_MAX_CHARS` globalmente (afetaria todo o corpus por causa de
|
||||
poucos arquivos atípicos); a lacuna fica só nos poucos chunks realmente
|
||||
grandes demais, e o resto do arquivo ainda fica pesquisável.
|
||||
|
||||
**Aprendizado:** um limite de chunk em caracteres é uma aproximação, não uma
|
||||
garantia de contexto — arquivos de dados brutos (tabelas, mapeamentos
|
||||
numéricos, JSON/CSV embutido em `.py`) tokenizam bem mais denso que prosa ou
|
||||
código comum e podem violar o limite do modelo mesmo dentro do teto de
|
||||
caracteres. Uma indexação em lote sobre centenas de arquivos não deve ficar
|
||||
tudo-ou-nada numa única transação: uma falha isolada e recuperável (chunk
|
||||
específico, arquivo específico) deve ser contida ali, não descartar o
|
||||
trabalho inteiro já validado.
|
||||
|
||||
**Estado:** `resolvido` — indexação completa rodou até o fim: 304 arquivos,
|
||||
1702 chunks, 0 removidos. Também nesta sessão: criada a pasta `rag/` na raiz
|
||||
(schema, busca híbrida `search.py`/`search_gart.sh`, `SETUP.md`) — ver
|
||||
`rag/README.md` para a divisão de responsabilidades com `admin/update_rag.py`.
|
||||
|
||||
---
|
||||
|
||||
### Entrada 36 — 2026-09-23: `.gitignore` escondia `fcpxml/models/` inteiro do git
|
||||
|
||||
**Sintoma:** ao investigar por que um `git diff` de um arquivo recém-editado
|
||||
(`fcpxml/models/timeline.py`, durante a correção da Entrada 34) não mostrava
|
||||
nada, `git status` também não listava o arquivo como modificado nem como
|
||||
untracked — como se ele simplesmente não existisse para o git.
|
||||
|
||||
**Causa:** `.gitignore` tinha a regra solta `models/` (comentada como
|
||||
"WhisperX models cache", pensada para ignorar o cache de ~11 GB de modelos
|
||||
Whisper baixados em `code/models/`). Uma regra sem `/` inicial no
|
||||
`.gitignore` casa com **qualquer diretório com esse nome em qualquer
|
||||
profundidade** — não só `code/models/`, mas também `code/fcpxml/models/`, o
|
||||
pacote de data classes (`TimeValue`, `Clip`, `Timeline`, `Marker`, etc.) que
|
||||
sustenta todo o engine. Confirmado: `git ls-tree -r HEAD` não tem nenhum
|
||||
`fcpxml/models.py` nem `fcpxml/models/` em nenhum commit do histórico — o
|
||||
pacote inteiro (1.234 linhas, 7 módulos) só existia em disco, sem nenhuma
|
||||
proteção de versionamento, desde que a divisão de `models.py` em pacote foi
|
||||
feita (sessão anterior, nunca commitada).
|
||||
|
||||
**Risco:** qualquer operação que limpa arquivos não rastreados
|
||||
(`git clean -fd`, reinstalar do zero, trocar de máquina via `git clone`)
|
||||
apagaria essa base sem chance de recuperação — nenhum commit para reverter.
|
||||
|
||||
**Solução:** regra trocada para `/code/models/` (ancorada na raiz do repo,
|
||||
só o cache real), preservando `whisper/` (sem uso hoje, mas inofensiva) e
|
||||
tudo mais. Confirmado com `git check-ignore -v`: `fcpxml/models/timeline.py`
|
||||
não é mais ignorado; `code/models/models--Systran--faster-whisper-base`
|
||||
continua ignorado. `fcpxml/models/` passou a aparecer como `??` no
|
||||
`git status` — visível, pronto para ser commitado quando o dono do trabalho
|
||||
revisar.
|
||||
|
||||
**Aprendizado:** regra de `.gitignore` sem `/` inicial (ex.: `models/`) casa
|
||||
em qualquer profundidade da árvore — é fácil escrever pensando só no caso
|
||||
que motivou a regra (um cache na raiz) e esquecer que o mesmo nome de pasta
|
||||
pode existir, com sentido completamente diferente, dentro do código-fonte.
|
||||
Regra de bolso: nomes de pasta genéricos (`models/`, `build/`, `cache/`,
|
||||
`data/`) no `.gitignore` deveriam quase sempre vir ancorados (`/caminho/
|
||||
exato/`), a menos que a intenção seja mesmo ignorar toda ocorrência do nome
|
||||
em qualquer lugar da árvore.
|
||||
|
||||
**Estado:** `resolvido` — regra corrigida, `fcpxml/models/` confirmado
|
||||
visível ao git (não commitado ainda; fica para quem já está com esse
|
||||
trabalho em andamento decidir quando commitar). Nenhum código alterado,
|
||||
só o `.gitignore`.
|
||||
|
||||
@@ -1,5 +1,8 @@
|
||||
# 06 — Boas Práticas de Programação (G-ART)
|
||||
|
||||
> **Escopo:** Checklist de qualidade a aplicar antes de dar algo por pronto.
|
||||
> **Não cobre:** Por onde começar uma tarefa (→ 09) · o que já quebrou (→ 05)
|
||||
|
||||
> **Propósito:** registrar as melhores práticas de programação a serem aplicadas
|
||||
> **sempre** que qualquer alteração ou correção for feita neste programa.
|
||||
> Servem de checklist obrigatório antes de concluir qualquer mudança.
|
||||
|
||||
@@ -0,0 +1,204 @@
|
||||
# 08 — O app macOS (`MacApp/`) e o Assistente
|
||||
|
||||
> **Escopo:** O app SwiftUI e o Assistente: build, telas, ponte e a etapa 5.
|
||||
> **Não cobre:** Engine Python (→ 02) · ferramentas MCP (→ 03)
|
||||
|
||||
O app SwiftUI é como o usuário opera o sistema sem abrir terminal nem conversar
|
||||
com uma IA. São ~5.500 linhas em `MacApp/Sources/`, e ele **não tem lógica de
|
||||
edição**: tudo que ele faz é montar argumentos, chamar a ponte Python e mostrar
|
||||
o resultado.
|
||||
|
||||
Última varredura: 2026-08-19
|
||||
|
||||
---
|
||||
|
||||
## 1. Como o app é construído — leia antes de mexer
|
||||
|
||||
**Não existe `.xcodeproj` nem `Package.swift`.** O app é compilado invocando o
|
||||
`swiftc` direto sobre `MacApp/Sources/*.swift`:
|
||||
|
||||
```bash
|
||||
cd code && ./MacApp/build_app.sh # compila e monta o .app
|
||||
admin/run_app.command # compila, fecha a instância antiga e abre (padrão de revisão)
|
||||
```
|
||||
|
||||
Consequências práticas, todas já sentidas:
|
||||
|
||||
- **Arquivo novo em `Sources/` entra sozinho** no build. Não há lista de alvos.
|
||||
- **Não dá para adicionar dependência SPM** sem antes migrar o build inteiro.
|
||||
- **Compilar não prova que roda.** Componentes SwiftUI que embrulham classes
|
||||
Objective-C podem falhar só em tempo de execução, ao abrir a tela. Foi o que
|
||||
aconteceu com `VideoPlayer` (AVKit): compilava limpo e abortava ao abrir a
|
||||
etapa 5 (`05_EXPERIENCIAS.md` #22). Por isso a regra: **alterou a interface,
|
||||
abra a tela de fato.**
|
||||
|
||||
### Testando uma tela sem navegar o app inteiro
|
||||
|
||||
Um harness de vinte linhas compila os mesmos fontes com um `@main` próprio que
|
||||
monta só a tela em questão. Reproduz crash de runtime em segundos:
|
||||
|
||||
```bash
|
||||
swiftc -parse-as-library -sdk "$(xcrun --sdk macosx --show-sdk-path)" \
|
||||
-target arm64-apple-macosx26.0 \
|
||||
MacApp/Sources/PhraseReviewView.swift MacApp/Sources/PhraseReviewModel.swift \
|
||||
MacApp/Sources/TimelineTracksView.swift MacApp/Sources/Models.swift \
|
||||
MacApp/Sources/PythonBridge.swift /tmp/HarnessMain.swift -o /tmp/harness
|
||||
```
|
||||
|
||||
O `@main` do harness carrega a tela, imprime o que interessa e chama
|
||||
`NSApplication.shared.terminate` — dá para afirmar "abriu e funcionou" sem
|
||||
depender de screenshot.
|
||||
|
||||
---
|
||||
|
||||
## 2. Estrutura das telas
|
||||
|
||||
| Arquivo | Linhas | Papel |
|
||||
|---------|-------:|-------|
|
||||
| `WizardView.swift` | 808 | **O Assistente** — fluxo guiado de 7 etapas |
|
||||
| `TranscriptionView.swift` | 843 | Transcrição avulsa e processamento em lote |
|
||||
| `ModelDownloadView.swift` | 545 | Catálogo e download de modelos Whisper |
|
||||
| `CaptionsView.swift` | 545 | Legendas dinâmicas: estilo + preview ao vivo |
|
||||
| `TimelineTracksView.swift` | 506 | Timeline com trilhas, zoom e playhead |
|
||||
| `PhraseReviewModel.swift` | 429 | Estado da etapa 5: frases, player, zooms |
|
||||
| `PhraseReviewView.swift` | 413 | Etapa 5: preview + inspector de frases |
|
||||
| `VoiceAnalysisView.swift` | 322 | Parâmetros do motor de ênfase |
|
||||
| `ProjectView.swift` | 293 | Inspeção do `.fcpxml` |
|
||||
| `Models.swift` | 274 | Espelhos Swift do JSON da ponte |
|
||||
| `PythonBridge.swift` | 230 | **A ponte** — ver seção 3 |
|
||||
| `SubtitlePreviewView.swift` | 218 | Preview 9:16 das legendas |
|
||||
| `App.swift` | 69 | `NavigationSplitView` e as abas |
|
||||
|
||||
Abas (`ActiveTab` em `App.swift`): Assistente · Projeto · Legendas · Análise de
|
||||
Voz · Modelos · Sobre. As cinco últimas são "Avançado" — atalhos para operações
|
||||
soltas. O Assistente é o caminho principal.
|
||||
|
||||
---
|
||||
|
||||
## 3. `PythonBridge.swift` — como o app fala com o Python
|
||||
|
||||
O app lança `admin/models_api.py` como **subprocesso**, passando o comando e um
|
||||
JSON como `argv`, e lê **JSON-lines** no stdout.
|
||||
|
||||
```swift
|
||||
PythonBridge.call(command: "build_phrase_review",
|
||||
arguments: ["voice_timeline": path]) { result, error in … }
|
||||
```
|
||||
|
||||
Dois pontos que já causaram problema e estão resolvidos no código — não os
|
||||
desfaça sem entender:
|
||||
|
||||
- **`uv run` precisa rodar com cwd em `code/`.** O `uv` escolhe o ambiente pelo
|
||||
diretório do processo, não pelo caminho do script. Rodar da raiz fazia o `uv`
|
||||
criar um segundo `.venv` vazio e ignorar tudo que estava instalado em
|
||||
`code/.venv` — librosa e pyannote instalavam com sucesso e o app insistia que
|
||||
faltavam.
|
||||
- **`scriptURL` procura `admin/models_api.py`** subindo diretórios a partir do
|
||||
cwd, do bundle e do home. É o que faz o app funcionar tanto rodando do Xcode
|
||||
quanto do `.app` montado.
|
||||
- **O `sys.path` que torna `fcpxml`/`server_tools` importáveis dentro de
|
||||
`admin/api/` mora só em `admin/api/__init__.py`.** Não copie esse cálculo
|
||||
para um módulo de comando individual — foi exatamente essa cópia,
|
||||
desatualizada em um nível de diretório, que quebrou toda ação que passa por
|
||||
`server` (`05_EXPERIENCIAS.md` #25). E não confie em "testei com `uv run` e
|
||||
funcionou": esse comando roda no mesmo venv com install editável que
|
||||
mascara esse tipo de erro. O teste que pega de verdade é
|
||||
`tests/test_models_api.py::TestCodeDirResolution`.
|
||||
|
||||
Para adicionar um comando: função em `admin/api/<assunto>.py`, registro na
|
||||
tabela de `admin/models_api.py`, e `PythonBridge.call` do lado Swift. Os 37
|
||||
comandos e seus formatos estão documentados no docstring de `models_api.py`.
|
||||
|
||||
---
|
||||
|
||||
## 4. O Assistente — as 7 etapas
|
||||
|
||||
`WizardStep` (`WizardView.swift`) é um enum sequencial; `canAdvance` decide
|
||||
quando o botão "Continuar" libera.
|
||||
|
||||
| # | Etapa | O que acontece | Comando da ponte |
|
||||
|---|-------|----------------|------------------|
|
||||
| 1 | Projeto | Escolhe a pasta de saída e o `.fcpxml` | `project_config` |
|
||||
| 2 | Transcrever | Transcreve toda a mídia do projeto | `transcribe` |
|
||||
| 3 | Analisar voz | Mede ênfase, locutores, emoção | `analyze_voice` |
|
||||
| 4 | Decisões da IA | Copia para o chat **ou** gera por IA local (Ollama/Gemma 3), aplica | `apply_voice_actions` / `generate_voice_script` |
|
||||
| 5 | **Revisar ênfases** | Lapida frase a frase — ver seção 5 | `build_phrase_review` / `save_phrase_review` |
|
||||
| 6 | Processar | Silêncios, preenchimento, legendas | vários, em cadeia |
|
||||
| 7 | Concluído | Abre no FCP ou mostra no Finder | — |
|
||||
|
||||
**A etapa 4 tem duas saídas:**
|
||||
|
||||
- **Manual (chat):** o app monta o pedido pronto no clipboard (skill `editar-por-voz`) e recebe o JSON de volta — o julgamento de qual tomada usar e onde dar zoom fica com a IA numa conversa.
|
||||
- **Automática (IA local):** botão "Gerar roteiro por IA local (Ollama/Gemma 3)". Ele manda a *voice timeline inteira* (o arquivo) junto com o brief para um modelo local (Ollama), que decide cortes/zooms/textos de uma vez, devolve o roteiro legível + o JSON de ações e já aplica no FCPXML (non-destructive). Não precisa sair do app nem colar nada. O modelo é escolhido num **picker que lista os modelos instalados no Ollama** (populado via `list_ollama_models` quando a etapa abre); se o Ollama estiver fora do ar, cai para um campo de texto livre. Troque para `llama3` etc. se tiver outro modelo. Requer o Ollama rodando em `localhost:11434`.
|
||||
|
||||
**Etapa 1 — armadilha registrada:** não escolha como "o projeto" um arquivo já
|
||||
gerado pelo fluxo (`_voice_edit`, `_silence_removed`, …). Os cortes de voz
|
||||
assumem timestamps da mídia **original**; reaplicá-los sobre um arquivo já
|
||||
cortado desloca tudo em silêncio. O wizard avisa (`looksLikeGeneratedFile`).
|
||||
|
||||
---
|
||||
|
||||
## 5. Etapa 5 — a sala de edição
|
||||
|
||||
Única tela que ocupa a janela toda: o corpo do wizard é uma coluna de 640pt, e
|
||||
essa etapa escapa dela porque precisa da largura (`step == .revisar` em
|
||||
`WizardView.body`).
|
||||
|
||||
```
|
||||
┌────────────────────────────┬──────────────┐
|
||||
│ Preview (AVPlayerLayer) │ Inspector │
|
||||
│ enquadrado no formato │ de frases │
|
||||
│ de entrega do projeto │ │
|
||||
├────────────────────────────┴──────────────┤
|
||||
│ Timeline: 6 trilhas, zoom, playhead │
|
||||
└───────────────────────────────────────────┘
|
||||
```
|
||||
|
||||
**Trilhas:** zooms · frases · energia por palavra · emoção · locutor ·
|
||||
roteiro/bastidor. Todas desenhadas sobre o mesmo eixo de tempo, com uma coluna
|
||||
fixa à esquerda nomeando cada uma.
|
||||
|
||||
**O que o usuário decide por frase:** nível de ênfase (0–3), ativo/inativo,
|
||||
texto, roteiro/bastidor e o trim das pontas. O trim anda em **fronteira de
|
||||
palavra** — cortar é apontar para uma palavra, arrastando a borda do bloco ou
|
||||
clicando na palavra no inspector.
|
||||
|
||||
**Zoom manual:** arrastar na timeline marca um trecho; botão direito cria um
|
||||
zoom nele. O zoom guarda **só o quando** — escala e ramp vêm das configurações
|
||||
de Análise de Voz no momento do render, então mudar lá restiliza todos.
|
||||
|
||||
**Decisões de implementação que parecem detalhe e não são:**
|
||||
|
||||
- **O preview não renderiza nada.** Ele toca a mídia original e *pula* os
|
||||
trechos removidos. Renderizar para conferir um toggle poria minutos entre a
|
||||
decisão e o resultado. O observador roda a 60 Hz porque o período dele é
|
||||
exatamente quanto de material cortado dá para ouvir antes do pulo.
|
||||
- **O enquadramento é o do projeto, não o da mídia.** As gravações são
|
||||
horizontais e a entrega é vertical; o app lê o formato do `.fcpxml`
|
||||
(`inspect`) e mostra o corte central aproximado, com um selo para alternar
|
||||
para a mídia original. O enquadramento real de cada clipe vem do FCP — o
|
||||
preview é aproximação, e o selo diz isso.
|
||||
- **Nada é processado aqui.** "Continuar" grava o `_phrase_review.json` e o
|
||||
`_phrase_actions.json` derivado dele. A geração é da etapa 6.
|
||||
- **A revisão é sempre remontada da análise atual**, com as decisões salvas
|
||||
reaplicadas por cima (`merge_saved_decisions`). Assim refazer a análise de voz
|
||||
melhora a tela em vez de ficar mascarado por uma cópia velha; uma decisão cuja
|
||||
frase se moveu mais de 0,25 s é descartada em vez de colar na frase errada.
|
||||
|
||||
---
|
||||
|
||||
## 6. Estado atual e o que falta
|
||||
|
||||
**Funciona e foi verificado:** carga das frases com decisões da IA, as 6
|
||||
trilhas, seleção sincronizada nos três painéis, trim por palavra, zoom manual,
|
||||
reprodução parando no ponto exato (erro de 0 ms medido), pulo dos trechos
|
||||
removidos, enquadramento vertical, gravação ao avançar.
|
||||
|
||||
**Ainda em aberto:**
|
||||
|
||||
- **A etapa 6 não consome o `_phrase_review.json`.** A ligação — zoom e legenda
|
||||
dinâmica só nas frases de ênfase, legenda comum no resto — é a próxima tarefa.
|
||||
- **`MacApp/` não tem teste automatizado.** A rede é o harness da seção 1 e o
|
||||
olho do usuário. Toda mudança de interface precisa ser aberta de fato.
|
||||
- **O preview aproxima o reenquadramento vertical** pelo corte central; se os
|
||||
clipes forem reposicionados no FCP, diverge.
|
||||
@@ -0,0 +1,193 @@
|
||||
# 09 — Manutenção: onde mexer, o que está aberto, o que dói
|
||||
|
||||
> **Escopo:** Por onde começar cada tipo de tarefa, o que está aberto e onde dói.
|
||||
> **Não cobre:** Como as coisas funcionam — este doc roteia para quem explica
|
||||
|
||||
Este é o documento de rota. Os outros descrevem o que **é**; este diz o que
|
||||
**fazer** e por onde começar quando chega uma implementação, uma melhoria ou
|
||||
uma correção.
|
||||
|
||||
Última varredura: 2026-09-22 · 1.543 testes passando (+1 falha pré-existente em `test_forced_align.py` e +1 erro pré-existente em `test_refine_voice_timeline_tool.py`, ver §2.6) · lint zerado em `code/` e em `admin/` (fora de server.py/ai_edit.py/fcpxml/analise.py, pré-existentes — outro trabalho em andamento na branch)
|
||||
|
||||
---
|
||||
|
||||
## 1. Chegou uma tarefa — por onde começo?
|
||||
|
||||
| A tarefa é… | Comece em | Não esqueça |
|
||||
|-------------|-----------|-------------|
|
||||
| Regra nova de edição (corte, zoom, legenda) | `fcpxml/<módulo>` + teste | Expor na tool **e** na ponte, senão só metade dos usuários alcança |
|
||||
| Corrigir XML que o FCP recusa | `fcpxml/writer/` + `dtd.py` | Validar contra o DTD real, não só o teste |
|
||||
| Mudança visível na interface | `MacApp/Sources/` | **Abrir a tela** — compilar não prova nada (§4) |
|
||||
| Comando novo para o app | `admin/api/<assunto>.py` | Registrar na tabela de `models_api.py` |
|
||||
| Ferramenta MCP nova | `server_tools/<categoria>.py` | Schema `Tool(...)` + `TOOL_HANDLERS` |
|
||||
| Ajuste de análise de voz | `fcpxml/voice_*`, `emphasis.py` | Regerar os `_voice_timeline.json` de teste |
|
||||
| "Está lento" / "está errado" e não sei onde | §5 (mapa de sintomas) | — |
|
||||
|
||||
**A pergunta que resolve 90% das dúvidas de lugar:** essa lógica precisa saber
|
||||
o que é uma tool MCP ou uma tela? Se não precisa — e quase nunca precisa — ela
|
||||
vai para `fcpxml/`.
|
||||
|
||||
---
|
||||
|
||||
## 2. O que está aberto agora
|
||||
|
||||
Ordenado por quanto atrapalha, não por esforço.
|
||||
|
||||
### 2.1 `resolve_actions` não tolera margem no encosto de zoom/marker contra um corte
|
||||
Um `zoom`/`marker` cuja borda cai exatamente em cima do `start`/`end` de um
|
||||
`cut` é descartado como "apontando para material cortado" — mesmo quando a
|
||||
intenção era ficar bem ao lado. Contornado manualmente no projeto Mastopexia
|
||||
(recuando as bordas na mão); a correção estrutural é dar a `resolve_actions`
|
||||
uma margem de tolerância (meio frame) antes de considerar uma ação "dentro"
|
||||
do corte. → `fcpxml/voice_actions.py` (`resolve_actions`/`shift_after_cuts`),
|
||||
`05_EXPERIENCIAS.md` #27.
|
||||
|
||||
### 2.2 `MacApp/` não tem teste automatizado
|
||||
5.500 linhas de Swift sem uma asserção. A rede hoje é o harness manual (§4) e
|
||||
o olho do usuário. Não é para sair criando suíte de UI — mas lógica pura que
|
||||
foi parar na camada de tela (cálculo de trim, mapeamento de tempo) deveria
|
||||
descer para o Python, onde já existe rede.
|
||||
|
||||
### 2.3 ~~`admin/` fica fora do lint~~ — resolvido em 2026-09-22
|
||||
`run_after_fix.sh` agora roda um segundo passo (`ruff check --config
|
||||
pyproject.toml ../admin/`) com a mesma config do engine. Precisou de
|
||||
`# noqa: E402` em 6 imports de `admin/models_api.py`/`admin/models_gui.py`
|
||||
(padrão `sys.path.insert` antes do import local, convenção já usada no
|
||||
projeto). Lint de `admin/` está zerado.
|
||||
|
||||
### 2.4 Confirmações visuais pendentes no FCP
|
||||
Várias entradas do `05_EXPERIENCIAS.md` estão marcadas como resolvidas *no XML*
|
||||
— testes verdes, DTD válido — mas **pendentes de importação real no Final Cut**.
|
||||
XML válido não é o mesmo que XML que renderiza como o esperado. Ao mexer em
|
||||
legenda, zoom ou keyframe, a confirmação final é abrir no FCP.
|
||||
|
||||
### 2.5 `remove_media_silence` deixa fatias sub-segundo nas emendas entre clipes
|
||||
Mesmo depois de corrigir o merge de cortes consecutivos (`05_EXPERIENCIAS.md`
|
||||
#28), sobraram 4 clipes de 0,07-0,23s no projeto Mastopexia real, todos bem
|
||||
na emenda entre dois clipes vizinhos — mesma família do #6 (clipe-fantasma de
|
||||
1 frame por padding sem vizinho na borda), mas não confirmado se é a mesma
|
||||
causa raiz. Não investigado a fundo ainda.
|
||||
→ `fcpxml/writer/cut.py` (`cut_clip_ranges`, `min_keep_seconds`), padding do
|
||||
`remove_media_silence`.
|
||||
|
||||
### 2.6 ~~`fcpxml/writer/adjustment.py` gerava um wrapper `<adjustment>` inválido~~ — resolvido em 2026-09-22
|
||||
`ClipDeAjuste` embrulhava filtros num `<clip><adjustment>...</adjustment></clip>`,
|
||||
que não existe no DTD real da Apple. Corrigido para anexar
|
||||
`filter-video`/`filter-audio` direto como filhos do `<clip>` (na ordem que o
|
||||
DTD exige: vídeo antes de áudio). Teste de regressão em
|
||||
`tests/test_writer_adjustment.py`. Segue sem uso em `server_tools`/`admin/api`
|
||||
— só deixou de estar pronto pra alguém reusar do jeito errado.
|
||||
→ `05_EXPERIENCIAS.md` #34.
|
||||
|
||||
### 2.7 `test_refine_voice_timeline_tool.py` quebrado: `voice_timeline.extract_pitch` ausente
|
||||
`TestRefineVoiceTimelineHandler::test_max_zooms_caps_the_list` tenta
|
||||
`monkeypatch.setattr(vt, "extract_pitch", ...)` mas `fcpxml/voice_timeline.py`
|
||||
não tem mais (ou nunca teve, nesta branch) essa função. Pertence ao trabalho
|
||||
de análise de voz já em andamento nesta branch (`voice_timeline.py`
|
||||
modificado, não commitado) — não investigado a fundo, só registrado aqui
|
||||
para não se perder.
|
||||
→ `fcpxml/voice_timeline.py`, `tests/test_refine_voice_timeline_tool.py`.
|
||||
|
||||
### 2.8 ~~Submódulo `WHISPERX` com conteúdo modificado e não commitado~~ — resolvido em 2026-09-22
|
||||
Não era um submódulo git registrado (sem `.gitmodules`) — era uma pasta
|
||||
`.git` solta de 2,6 GB dentro de `code/`, com cópias/backups congelados do
|
||||
próprio projeto (`WHISPERX_backup_88476/`, uma cópia inteira e antiga de
|
||||
`fcp-mcp-server-main`). Só 3 referências no código ativo, todas em
|
||||
comentários (`fcpxml/diarize.py`, `tests/test_diarize.py`,
|
||||
`admin/api/shared.py`), nenhum import ou caminho dependia dela. Além do
|
||||
peso morto, ela também inflava qualquer lint rodado com `--exclude`
|
||||
explícito (que sobrescreve o `exclude` do `pyproject.toml`) — foi assim que
|
||||
um `ruff check . --exclude docs/` chegou a acusar 510 erros, quase todos
|
||||
dentro dela. Movida para `~/Archives/G-ART-WHISPERX-backup` (fora do
|
||||
workspace git), copiada e verificada (`diff -rq`) antes de remover o
|
||||
original. `WHISPERX/` também saiu do `exclude` do ruff em
|
||||
`code/pyproject.toml` — não faz mais sentido excluir um caminho que não
|
||||
existe mais dentro de `code/`.
|
||||
|
||||
---
|
||||
|
||||
## 3. Onde o código ainda é grande (e onde isso não é problema)
|
||||
|
||||
Quatro arquivos foram divididos (`writer.py`, `models.py`, `models_api.py`,
|
||||
`_shared.py`): 6.685 linhas concentradas viraram 43 módulos.
|
||||
|
||||
O que sobrou grande, e o diagnóstico honesto de cada um:
|
||||
|
||||
| Arquivo | Linhas | Vale dividir? |
|
||||
|---------|-------:|---------------|
|
||||
| `fcpxml/text_layout.py` | 901 | **Não.** É diagramação — um assunto coeso. |
|
||||
| `fcpxml/rough_cut.py` | 798 | **Não.** É geração de timeline, um assunto. |
|
||||
| `fcpxml/model_manager.py` | 748 | Talvez: mistura catálogo, download e config. |
|
||||
| `server_tools/voice.py` | 754 | Talvez, se crescer mais. |
|
||||
| `MacApp/TranscriptionView.swift` | 843 | Sim, quando for mexer nela. |
|
||||
| `MacApp/WizardView.swift` | 808 | Sim: sete etapas num `switch` só. |
|
||||
|
||||
**Critério, não número:** divida quando o arquivo tiver **assuntos** que não se
|
||||
falam. Um arquivo grande de um assunto só é mais fácil de ler que seis arquivos
|
||||
pequenos que você precisa abrir juntos. Código picado sem motivo atrapalha tanto
|
||||
quanto arquivo gigante.
|
||||
|
||||
---
|
||||
|
||||
## 4. Checklist antes de dar algo por pronto
|
||||
|
||||
```bash
|
||||
cd code && ./Engine/run_after_fix.sh # lint zerado + 1.498 testes
|
||||
admin/run_app.command # se mexeu no app (padrão de revisão)
|
||||
admin/run.command # app + atualização incremental da RAG
|
||||
rag/search_gart.sh "consulta" # busca híbrida no índice RAG (ver rag/README.md)
|
||||
```
|
||||
|
||||
E, além do script:
|
||||
|
||||
- [ ] **Mexeu na interface? Abriu a tela?** Compilar não prova que roda —
|
||||
`VideoPlayer` compilava e abortava (`05_EXPERIENCIAS.md` #22).
|
||||
- [ ] **Mexeu em XML? Importou no FCP?** DTD válido ≠ renderiza certo.
|
||||
- [ ] **Dividiu ou moveu módulo?** Procure `patch('<módulo>.` e imports
|
||||
relativos dentro de funções — é o que quebra em silêncio (#23).
|
||||
- [ ] **Criou teste fora de `code/tests/`?** Confirme que a contagem total
|
||||
subiu. Teste fora de `testpaths` não roda e dá falsa sensação de rede (#24).
|
||||
- [ ] **Problema estrutural ou erro recorrente?** Registre em
|
||||
`05_EXPERIENCIAS.md` com o índice atualizado.
|
||||
- [ ] **Documentação divergiu?** Corrija no mesmo commit. Doc velha engana mais
|
||||
que doc ausente.
|
||||
|
||||
---
|
||||
|
||||
## 5. Mapa de sintomas → onde olhar
|
||||
|
||||
| Sintoma | Suspeite de | Arquivo |
|
||||
|---------|-------------|---------|
|
||||
| FCP recusa o arquivo ao importar | `id` inválido, ordem de filhos, timebase | `writer/validation.py`, `dtd.py` |
|
||||
| Título importa mas não aparece | Template/uid Motion que não resolve | `writer/titles.py` |
|
||||
| Corte no lugar errado | Tempo pós-corte usado como se fosse original | `voice_actions.py` (`shift_after_cuts`) |
|
||||
| Zoom/marker sumindo perto de um corte | Borda encostando exatamente no `cut` | §2.1 |
|
||||
| Legenda sobrepondo | Layout ou conteúdo antigo no arquivo | `collision.py`, `text_layout.py` |
|
||||
| "Ênfase" apontando para palavra à toa | Falta renormalizar após o corte | `refine_voice_timeline` |
|
||||
| App diz que falta librosa/pyannote | `uv run` com cwd errado | `PythonBridge.swift` (§3 do doc 08) |
|
||||
| App crasha com `ModuleNotFoundError: server_tools` | `sys.path` de `admin/api/` mal calculado | `05_EXPERIENCIAS.md` #25 |
|
||||
| Tela do app fecha o programa | Componente de framework que só falha em runtime | `05_EXPERIENCIAS.md` #22 |
|
||||
| Comando existe no MCP mas não no app | Falta expor na ponte | `admin/api/`, #20 |
|
||||
|
||||
---
|
||||
|
||||
## 6. Convenções que não são negociáveis
|
||||
|
||||
Estão em `01_ARCHITECTURE.md` §2 e valem repetir as três que mais custaram:
|
||||
|
||||
1. **Tempo é fração racional.** Float para tempo produz drift que só aparece
|
||||
depois de dez operações encadeadas.
|
||||
2. **Ação de voz é sempre em tempo da mídia original.** Nunca pós-corte.
|
||||
3. **Original nunca é sobrescrito.** Toda saída ganha sufixo.
|
||||
|
||||
---
|
||||
|
||||
## Documentos relacionados
|
||||
|
||||
- [01 Arquitetura](01_ARCHITECTURE.md) — camadas e onde cada coisa mora
|
||||
- [02 Módulos](02_MODULES.md) — mapa do engine, módulo a módulo
|
||||
- [03 Server/Tools](03_SERVER_TOOLS.md) — as 77 ferramentas MCP
|
||||
- [04 Testes & Workflow](04_TESTS_AND_WORKFLOW.md)
|
||||
- [05 Experiências](05_EXPERIENCIAS.md) — o que já quebrou e por quê
|
||||
- [06 Boas Práticas](06_BOAS_PRATICAS.md)
|
||||
- [08 App macOS](08_APP_MACOS.md) — o app e o Assistente
|
||||
@@ -0,0 +1,269 @@
|
||||
# 10 - Mapa de Reestruturacao de Funcionalidades
|
||||
|
||||
> Escopo: roteiro pratico para reorganizar o codigo sem quebrar o produto.
|
||||
> Baseado na varredura de 2026-08-24 sobre engine Python, ponte do app,
|
||||
> ferramentas MCP e app SwiftUI.
|
||||
|
||||
## 1. Diagnostico rapido
|
||||
|
||||
O projeto ja tem uma arquitetura-alvo correta: `fcpxml/` como engine puro,
|
||||
`server.py` + `server_tools/` como camada MCP, `admin/` como ponte JSON-lines
|
||||
do app e `MacApp/` como interface. A melhoria agora nao e "reinventar" a
|
||||
arquitetura, e reduzir os pontos onde as responsabilidades ainda se misturam.
|
||||
|
||||
### Pontos fortes
|
||||
|
||||
- Engine Python bem testado e com regra clara: logica de timeline fica em
|
||||
`fcpxml/`.
|
||||
- `writer/` ja foi quebrado em mixins por assunto, preservando API publica.
|
||||
- `server.py` funciona como composition root e usa dispatch por dicionario.
|
||||
- Documentacao interna registra decisoes, armadilhas e padroes do projeto.
|
||||
- Fluxos criticos tem testes extensos em `code/tests/`.
|
||||
|
||||
### Dores atuais
|
||||
|
||||
- Alguns arquivos voltaram a virar centros de gravidade:
|
||||
- `server_tools/voice.py` (~999 linhas)
|
||||
- `server_tools/subtitles.py` (~760 linhas)
|
||||
- `MacApp/Sources/WizardView.swift` (~979 linhas)
|
||||
- `MacApp/Sources/TranscriptionView.swift` (~843 linhas)
|
||||
- `fcpxml/model_manager.py` (~748 linhas)
|
||||
- `admin/` e `server_tools/` expõem fluxos parecidos por caminhos diferentes,
|
||||
o que aumenta risco de uma funcionalidade existir no MCP e faltar no app.
|
||||
- ~~`admin/` ainda fica fora do lint principal~~ — resolvido na Fase 0
|
||||
(2026-09-22): `admin/` entrou no gate de `run_after_fix.sh`.
|
||||
- ~~`WHISPERX` e backups aparecem junto da base ativa~~ — resolvido na
|
||||
Fase 0 (2026-09-22): movido para fora do workspace git.
|
||||
- O app SwiftUI quase nao tem rede automatizada; compilar nao garante que uma
|
||||
tela abre.
|
||||
|
||||
## 2. Mapa de dominios desejado
|
||||
|
||||
```text
|
||||
Produto
|
||||
MacApp/ Interface e experiencia do usuario
|
||||
admin/ Ponte JSON-lines do app
|
||||
server.py + server_tools/ Entrada MCP
|
||||
|
||||
Engine
|
||||
fcpxml/models/ Dados e contratos
|
||||
fcpxml/parser.py FCPXML -> objetos
|
||||
fcpxml/writer/ Escrita e edicao de XML
|
||||
fcpxml/voice_* Analise e decisoes por voz
|
||||
fcpxml/text_layout.py Layout de legendas
|
||||
fcpxml/model_manager.py Catalogo, configs e modelos
|
||||
|
||||
Suporte
|
||||
tests/ Rede automatizada
|
||||
Engine/docs/ Decisoes e operacao
|
||||
examples/ Fixtures de uso
|
||||
|
||||
Legado / referencia
|
||||
WHISPERX/ Deve sair do caminho ativo ou virar referencia clara
|
||||
```
|
||||
|
||||
Regra de organizacao: uma funcionalidade nasce no engine, depois ganha duas
|
||||
portas finas se necessario: uma tool MCP em `server_tools/` e um comando do app
|
||||
em `admin/api/`.
|
||||
|
||||
## 3. Reestruturacao por fases
|
||||
|
||||
### Fase 0 - Higiene antes de mexer — `concluída em 2026-09-22`
|
||||
|
||||
Objetivo: reduzir ruido e proteger a base antes de mover codigo.
|
||||
|
||||
- ~~Decidir o destino de `code/WHISPERX`~~ — não era submodule (sem
|
||||
`.gitmodules`), era 2,6 GB de backups órfãos do próprio projeto sem
|
||||
nenhuma referência ativa. Movido para `~/Archives/G-ART-WHISPERX-backup`
|
||||
(fora do workspace git), copiado com `rsync` e conferido com `diff -rq`
|
||||
antes de remover o original. Detalhe: essa pasta também inflava qualquer
|
||||
`ruff check --exclude docs/` manual (a flag sobrescrevia o `exclude` do
|
||||
`pyproject.toml`, que já ignorava `WHISPERX/`) — ver `05_EXPERIENCIAS.md`
|
||||
#34.
|
||||
- ~~Incluir `admin/` em uma checagem de lint separada antes de colocar no
|
||||
gate obrigatório~~ — checado com a config real do projeto (não o default
|
||||
do ruff): só 6 erros, todos `E402` por `sys.path.insert` antes de import
|
||||
local. Resolvido com `# noqa: E402` (convenção já usada no projeto) e
|
||||
`admin/` entrou direto no gate obrigatório (`run_after_fix.sh`, passo
|
||||
2/3), sem precisar de etapa intermediária "separada".
|
||||
- Corrigido de quebra: `fcpxml/writer/adjustment.py` gerava um `<adjustment>`
|
||||
inválido no DTD — não estava no escopo original da Fase 0, mas surgiu na
|
||||
investigação e era pequeno o bastante para resolver junto (ver
|
||||
`05_EXPERIENCIAS.md` #34).
|
||||
- Atualizados: `02_MODULES.md` (versão, linhas de `writer/`, módulos novos
|
||||
`builders.py`/`adjustment.py`/`analise.py`/`transcription/`),
|
||||
`09_MANUTENCAO.md` (contagem de testes/lint, itens §2.3/§2.6/§2.8
|
||||
resolvidos, novo item §2.7 registrando `test_refine_voice_timeline_tool`).
|
||||
- **Pendente, não fechado nesta rodada:** "documentar oficialmente quais
|
||||
pastas são produto ativo, legado e backup" como um documento à parte —
|
||||
o que existia de fato como "legado" (`WHISPERX`) já foi resolvido, não
|
||||
sobrou candidato claro para justificar um novo documento agora.
|
||||
|
||||
Entrega obtida: lint de `admin/` no gate, `code/writer/adjustment.py`
|
||||
correto e testado, ~2,6 GB fora do caminho ativo, docs sincronizados com o
|
||||
código atual.
|
||||
|
||||
### Fase 1 - Contratos entre camadas
|
||||
|
||||
Objetivo: impedir que MCP, app e engine driftam entre si.
|
||||
|
||||
- Criar um registro unico de capacidades, por exemplo:
|
||||
- nome interno da funcionalidade;
|
||||
- funcao pura do engine;
|
||||
- handler MCP, se existir;
|
||||
- comando `admin`, se existir;
|
||||
- tela Swift, se existir;
|
||||
- testes associados.
|
||||
- Adicionar teste que detecta comandos importantes presentes no MCP mas ausentes
|
||||
na ponte do app, quando fizer sentido.
|
||||
- Padronizar o retorno dos comandos `admin/api`: `ok`, `path`, `message`,
|
||||
`error`, `unchanged`, `artifacts`.
|
||||
|
||||
Entrega esperada: mapa vivo de funcionalidades e menos "funciona no Claude,
|
||||
nao aparece no app".
|
||||
|
||||
### Fase 2 - Dividir `server_tools/voice.py`
|
||||
|
||||
Objetivo: separar o fluxo de voz por etapas reais do produto.
|
||||
|
||||
Divisao sugerida:
|
||||
|
||||
```text
|
||||
server_tools/voice/
|
||||
__init__.py Reexporta TOOLS e HANDLERS
|
||||
analysis.py analyze_voice_features, build_voice_timeline
|
||||
speakers.py diarize_media, remove_speakers
|
||||
refinement.py refine_voice_timeline, remove_speech_gaps
|
||||
actions.py apply_voice_actions
|
||||
local_ai.py generate_voice_script
|
||||
config.py get/save_voice_analysis_config
|
||||
```
|
||||
|
||||
Cuidados:
|
||||
|
||||
- Manter os nomes publicos reexportados para nao quebrar testes/imports.
|
||||
- Mover em uma etapa por arquivo, rodando testes de voz a cada passo.
|
||||
- Nao mover regra de negocio para `server_tools/voice/`; se aparecer regra
|
||||
nova, ela deve descer para `fcpxml/voice_*`.
|
||||
|
||||
Testes minimos: `test_voice_actions.py`, `test_voice_actions_tool.py`,
|
||||
`test_voice_timeline.py`, `test_voice_timeline_tool.py`, `test_diarize.py`,
|
||||
`test_voice_features.py`.
|
||||
|
||||
### Fase 3 - Separar `fcpxml/model_manager.py`
|
||||
|
||||
Objetivo: reduzir mistura entre catalogo, download, configuracao e estado.
|
||||
|
||||
Divisao sugerida:
|
||||
|
||||
```text
|
||||
fcpxml/model_manager/
|
||||
__init__.py API publica atual
|
||||
catalog.py models.json, recomendados, metadata
|
||||
storage.py diretorios, instalados, migracao
|
||||
download.py download/cancel/progresso
|
||||
transcription_config.py modelo selecionado, idioma
|
||||
voice_config.py analise de voz, silencio, legendas
|
||||
```
|
||||
|
||||
Cuidados:
|
||||
|
||||
- Preservar imports atuais via `__init__.py`.
|
||||
- Separar funcoes puras de funcoes com I/O para facilitar teste.
|
||||
- Nao acoplar config do app a nomes de tela Swift.
|
||||
|
||||
Testes minimos: `test_models.py`, `test_models_api.py` se existir,
|
||||
`test_voice_analysis_config.py`, `test_project_config.py`.
|
||||
|
||||
### Fase 4 - Reorganizar o Assistente SwiftUI
|
||||
|
||||
Objetivo: tornar o fluxo de 7 etapas legivel e testavel por partes.
|
||||
|
||||
Divisao sugerida:
|
||||
|
||||
```text
|
||||
MacApp/Sources/Wizard/
|
||||
WizardView.swift Casca, navegacao e estado global
|
||||
WizardState.swift Estado do fluxo e canAdvance
|
||||
ProjectStepView.swift
|
||||
TranscribeStepView.swift
|
||||
VoiceAnalysisStepView.swift
|
||||
AIScriptStepView.swift
|
||||
ReviewStepHost.swift
|
||||
ProcessStepView.swift
|
||||
DoneStepView.swift
|
||||
```
|
||||
|
||||
Boas praticas para essa fase:
|
||||
|
||||
- Extrair primeiro views pequenas, sem alterar comportamento.
|
||||
- Depois extrair calculos puros de `canAdvance`, nomes de arquivos e selecao
|
||||
de artefatos para tipos testaveis.
|
||||
- Usar harness manual documentado em `08_APP_MACOS.md` para abrir as telas
|
||||
tocadas.
|
||||
|
||||
Entrega esperada: cada etapa do wizard vira um arquivo com responsabilidade
|
||||
unica.
|
||||
|
||||
### Fase 5 - Unificar validacao e saida da ponte `admin/`
|
||||
|
||||
Objetivo: deixar os comandos do app tao disciplinados quanto os handlers MCP.
|
||||
|
||||
- Criar helpers de path/output equivalentes aos de `server_tools/_shared`,
|
||||
ou mover helpers comuns para uma camada compartilhada que nao saiba de MCP.
|
||||
- Trocar chamadas diretas a `server.generate_output_path` por helper de dominio
|
||||
que nao puxe `server.py` quando a ponte so precisa de path.
|
||||
- Adicionar lint de `admin/` ao fluxo de manutencao depois de corrigir erros
|
||||
existentes.
|
||||
|
||||
Entrega esperada: ponte mais fina, menos import acidental de transporte MCP.
|
||||
|
||||
### Fase 6 - Tests e gates de seguranca
|
||||
|
||||
Objetivo: fazer a reorganizacao ser barata de continuar.
|
||||
|
||||
- Criar testes de "arquitetura":
|
||||
- `fcpxml/` nao importa `server`, `server_tools` nem `admin`;
|
||||
- handlers MCP sempre retornam via `_text_result`;
|
||||
- comandos `admin` retornam JSON no formato padrao.
|
||||
- Criar teste de import publico para garantir que reexports antigos continuam.
|
||||
- Para SwiftUI, manter harnesses por tela critica ate existir um build mais
|
||||
estruturado.
|
||||
|
||||
Entrega esperada: mover arquivos deixa de ser aposta.
|
||||
|
||||
## 4. Prioridade recomendada
|
||||
|
||||
1. Fase 0: limpar mapa ativo vs legado.
|
||||
2. Fase 1: criar registro de capacidades.
|
||||
3. Fase 2: dividir voz em `server_tools`.
|
||||
4. Fase 5: fortalecer `admin/`.
|
||||
5. Fase 4: quebrar `WizardView`.
|
||||
6. Fase 3: dividir `model_manager.py`.
|
||||
7. Fase 6: ampliar gates conforme as fases estabilizam.
|
||||
|
||||
Motivo: primeiro se reduz incerteza, depois se separa o arquivo que mais muda
|
||||
no fluxo novo de voz, e so entao se mexe nas telas maiores.
|
||||
|
||||
## 5. Checklist para cada refatoracao
|
||||
|
||||
- Mover sem mudar comportamento na primeira passada.
|
||||
- Preservar API publica com reexports.
|
||||
- Rodar testes focados depois de cada movimento.
|
||||
- Rodar `cd code && ./Engine/run_after_fix.sh` antes de concluir.
|
||||
- Se mexeu em `MacApp/`, compilar e abrir a tela afetada.
|
||||
- Atualizar docs no mesmo commit.
|
||||
- Registrar aprendizado em `05_EXPERIENCIAS.md` quando houver bug real.
|
||||
|
||||
## 6. Principios de boas praticas para este projeto
|
||||
|
||||
- Engine puro: sem MCP, sem Swift, sem JSON de tela.
|
||||
- Camadas de entrada finas: validam, chamam engine, formatam resposta.
|
||||
- Tempo de timeline sempre racional (`TimeValue`), exceto metricas de audio e
|
||||
UI onde segundos float sao apenas apresentacao/analise.
|
||||
- Original nunca e sobrescrito.
|
||||
- XML sempre entra por `safe_xml.py`.
|
||||
- Dependencias opcionais continuam lazy.
|
||||
- Arquivo grande so e problema quando contem varios assuntos.
|
||||
- Toda funcionalidade importante deve ter dono, porta MCP/app documentada e
|
||||
teste correspondente.
|
||||
@@ -23,14 +23,19 @@ echo "==> [G-ART] Validação pós-correção iniciada..."
|
||||
echo " Diretório: $REPO_ROOT"
|
||||
echo ""
|
||||
|
||||
echo "==> 1/2 Lint (ruff) — deve passar com ZERO erros"
|
||||
echo "==> 1/3 Lint do engine (ruff, code/) — deve passar com ZERO erros"
|
||||
# A flag --exclude sobrescreve o exclude declarado em pyproject.toml
|
||||
# (que já ignora docs/ e WHISPERX/). Rode sem flag para herdar a config.
|
||||
# (que já ignora docs/). Rode sem flag para herdar a config.
|
||||
uv run ruff check .
|
||||
echo " Lint OK ✓"
|
||||
echo ""
|
||||
|
||||
echo "==> 2/2 Testes (pytest) — todos devem passar"
|
||||
echo "==> 2/3 Lint da ponte (ruff, admin/) — mesma config do engine"
|
||||
uv run ruff check --config pyproject.toml ../admin/
|
||||
echo " Lint OK ✓"
|
||||
echo ""
|
||||
|
||||
echo "==> 3/3 Testes (pytest) — todos devem passar"
|
||||
uv run pytest tests/ -v
|
||||
echo ""
|
||||
|
||||
|
||||
@@ -12,6 +12,7 @@ struct GArtApp: App {
|
||||
}
|
||||
|
||||
enum ActiveTab: Hashable {
|
||||
case wizard
|
||||
case project
|
||||
case captions
|
||||
case voiceAnalysis
|
||||
@@ -20,19 +21,23 @@ enum ActiveTab: Hashable {
|
||||
}
|
||||
|
||||
struct ContentView: View {
|
||||
@State private var activeTab: ActiveTab? = .project
|
||||
@State private var activeTab: ActiveTab? = .wizard
|
||||
|
||||
var body: some View {
|
||||
NavigationSplitView {
|
||||
List(selection: $activeTab) {
|
||||
Label("Projeto", systemImage: "film")
|
||||
.tag(ActiveTab.project)
|
||||
Label("Legendas Dinâmicas", systemImage: "captions.bubble")
|
||||
.tag(ActiveTab.captions)
|
||||
Label("Análise de Voz", systemImage: "waveform")
|
||||
.tag(ActiveTab.voiceAnalysis)
|
||||
Label("Modelos", systemImage: "tray.and.arrow.down")
|
||||
.tag(ActiveTab.models)
|
||||
Label("Assistente", systemImage: "wand.and.stars")
|
||||
.tag(ActiveTab.wizard)
|
||||
Section("Avançado") {
|
||||
Label("Projeto", systemImage: "film")
|
||||
.tag(ActiveTab.project)
|
||||
Label("Legendas", systemImage: "captions.bubble")
|
||||
.tag(ActiveTab.captions)
|
||||
Label("Análise de Voz", systemImage: "waveform")
|
||||
.tag(ActiveTab.voiceAnalysis)
|
||||
Label("Modelos", systemImage: "tray.and.arrow.down")
|
||||
.tag(ActiveTab.models)
|
||||
}
|
||||
Label("Sobre", systemImage: "info.circle")
|
||||
.tag(ActiveTab.about)
|
||||
}
|
||||
@@ -40,19 +45,22 @@ struct ContentView: View {
|
||||
.navigationSplitViewColumnWidth(min: 180, ideal: 200)
|
||||
} detail: {
|
||||
switch activeTab {
|
||||
case .wizard, nil:
|
||||
WizardView().id(UUID())
|
||||
.navigationTitle("Assistente")
|
||||
case .project:
|
||||
ProjectView().id(UUID())
|
||||
.navigationTitle("Projeto")
|
||||
case .captions:
|
||||
CaptionsView().id(UUID())
|
||||
.navigationTitle("Legendas Dinâmicas")
|
||||
.navigationTitle("Legendas")
|
||||
case .voiceAnalysis:
|
||||
VoiceAnalysisView().id(UUID())
|
||||
.navigationTitle("Análise de Voz")
|
||||
case .models:
|
||||
ModelDownloadView().id(UUID())
|
||||
.navigationTitle("Modelos")
|
||||
case .about, nil:
|
||||
case .about:
|
||||
AboutView()
|
||||
.navigationTitle("Sobre")
|
||||
}
|
||||
|
||||
@@ -19,6 +19,7 @@ import UniformTypeIdentifiers
|
||||
/// assunto.
|
||||
struct CaptionsView: View {
|
||||
@State private var config = CaptionStyleConfig.defaults
|
||||
@State private var plainConfig = PlainSubtitleConfig.defaults
|
||||
@State private var isLoading = true
|
||||
@State private var errorMessage: String?
|
||||
|
||||
@@ -30,15 +31,31 @@ struct CaptionsView: View {
|
||||
@AppStorage("capSampleAfter") private var sampleAfter = "sua legenda"
|
||||
@AppStorage("capShowGuides") private var showsGuides = true
|
||||
|
||||
private let fontChoices = [
|
||||
"Helvetica Neue", "Helvetica", "Arial", "Avenir Next",
|
||||
"Futura", "SF Pro Display", "Georgia", "Impact",
|
||||
]
|
||||
/// Todas as famílias de fonte instaladas no macOS (sistema + usuário), as
|
||||
/// usadas por padrão primeiro, para o seletor listar tudo sem hardcode.
|
||||
private static let installedFontFamilies: [String] = {
|
||||
var families = NSFontManager.shared.availableFontFamilies
|
||||
.sorted { $0.localizedCaseInsensitiveCompare($1) == .orderedAscending }
|
||||
let preferred = ["Helvetica Neue", "Playfair Display", "Georgia", "Didot"]
|
||||
for family in preferred.reversed() {
|
||||
if let idx = families.firstIndex(of: family) {
|
||||
families.remove(at: idx)
|
||||
families.insert(family, at: 0)
|
||||
}
|
||||
}
|
||||
return families
|
||||
}()
|
||||
|
||||
private let emphasisFontChoices = [
|
||||
"Playfair Display", "Georgia", "Didot", "Futura",
|
||||
"Avenir Next", "Times New Roman", "Helvetica Neue", "Impact",
|
||||
]
|
||||
/// Lista para um picker: todas as famílias instaladas e, se o valor salvo
|
||||
/// não estiver entre elas (ex.: fonte de outro Mac), ele entra no topo
|
||||
/// para o seletor continuar exibindo a escolha atual.
|
||||
private func fontChoices(for current: String) -> [String] {
|
||||
var list = Self.installedFontFamilies
|
||||
if !list.contains(current) {
|
||||
list.insert(current, at: 0)
|
||||
}
|
||||
return list
|
||||
}
|
||||
|
||||
private let emphasisFaceChoices = [
|
||||
"Medium Italic", "Italic", "Bold Italic", "Bold", "Regular", "Light Italic",
|
||||
@@ -53,6 +70,13 @@ struct CaptionsView: View {
|
||||
)
|
||||
}
|
||||
|
||||
private func plainBound<T>(_ keyPath: WritableKeyPath<PlainSubtitleConfig, T>) -> Binding<T> {
|
||||
Binding(
|
||||
get: { plainConfig[keyPath: keyPath] },
|
||||
set: { plainConfig[keyPath: keyPath] = $0; savePlain() }
|
||||
)
|
||||
}
|
||||
|
||||
private func colorBound(_ keyPath: WritableKeyPath<CaptionStyleConfig, String>) -> Binding<Color> {
|
||||
Binding(
|
||||
get: { Color(rgbaString: config[keyPath: keyPath]) },
|
||||
@@ -60,6 +84,13 @@ struct CaptionsView: View {
|
||||
)
|
||||
}
|
||||
|
||||
private func plainColorBound(_ keyPath: WritableKeyPath<PlainSubtitleConfig, String>) -> Binding<Color> {
|
||||
Binding(
|
||||
get: { Color(rgbaString: plainConfig[keyPath: keyPath]) },
|
||||
set: { plainConfig[keyPath: keyPath] = $0.fcpxmlColorString; savePlain() }
|
||||
)
|
||||
}
|
||||
|
||||
var body: some View {
|
||||
HSplitView {
|
||||
previewColumn
|
||||
@@ -152,6 +183,7 @@ struct CaptionsView: View {
|
||||
positionSection
|
||||
bodySection
|
||||
emphasisSection
|
||||
plainSubtitleSection
|
||||
calibrationSection
|
||||
}
|
||||
if let errorMessage {
|
||||
@@ -195,7 +227,7 @@ struct CaptionsView: View {
|
||||
private var bodySection: some View {
|
||||
Section("Linhas de apoio") {
|
||||
Picker("Fonte", selection: bound(\.font)) {
|
||||
ForEach(fontChoices, id: \.self) { Text($0).tag($0) }
|
||||
ForEach(fontChoices(for: config.font), id: \.self) { Text($0).tag($0) }
|
||||
}
|
||||
slider(
|
||||
"Tamanho",
|
||||
@@ -210,7 +242,7 @@ struct CaptionsView: View {
|
||||
private var emphasisSection: some View {
|
||||
Section("Palavra de ênfase") {
|
||||
Picker("Fonte", selection: bound(\.emphasisFont)) {
|
||||
ForEach(emphasisFontChoices, id: \.self) { Text($0).tag($0) }
|
||||
ForEach(fontChoices(for: config.emphasisFont), id: \.self) { Text($0).tag($0) }
|
||||
}
|
||||
Picker("Estilo", selection: bound(\.emphasisFace)) {
|
||||
ForEach(emphasisFaceChoices, id: \.self) { Text($0).tag($0) }
|
||||
@@ -225,6 +257,35 @@ struct CaptionsView: View {
|
||||
}
|
||||
}
|
||||
|
||||
private var plainSubtitleSection: some View {
|
||||
Section("Legenda comum") {
|
||||
Picker("Fonte", selection: plainBound(\.font)) {
|
||||
ForEach(fontChoices(for: plainConfig.font), id: \.self) { Text($0).tag($0) }
|
||||
}
|
||||
slider(
|
||||
"Tamanho",
|
||||
value: plainBound(\.fontSize), in: 28...300, step: 1,
|
||||
readout: "\(Int(plainConfig.fontSize))pt",
|
||||
help: "Tamanho da legenda comum editável no Final Cut."
|
||||
)
|
||||
slider(
|
||||
"Máximo de palavras",
|
||||
value: plainBound(\.maxWords), in: 1...14, step: 1,
|
||||
readout: "\(Int(plainConfig.maxWords))",
|
||||
help: "Quantidade máxima de palavras por bloco de legenda."
|
||||
)
|
||||
slider(
|
||||
"Altura",
|
||||
value: plainBound(\.positionY), in: -1200...300, step: 1,
|
||||
readout: "\(Int(plainConfig.positionY))",
|
||||
help: "Posição vertical da legenda comum no quadro; valores mais negativos descem."
|
||||
)
|
||||
ColorPicker("Cor", selection: plainColorBound(\.fontColor), supportsOpacity: true)
|
||||
Toggle("Usar letra maiúscula", isOn: plainBound(\.uppercase))
|
||||
Toggle("Manter vírgula e ponto", isOn: plainBound(\.keepPunctuation))
|
||||
}
|
||||
}
|
||||
|
||||
private var calibrationSection: some View {
|
||||
Section {
|
||||
slider(
|
||||
@@ -305,8 +366,17 @@ struct CaptionsView: View {
|
||||
} else if let error {
|
||||
errorMessage = error
|
||||
}
|
||||
isLoading = false
|
||||
continuation.resume()
|
||||
PythonBridge.call(command: "plain_subtitle_config") { plainResult, plainError in
|
||||
DispatchQueue.main.async {
|
||||
if let plainResult {
|
||||
plainConfig = PlainSubtitleConfig(from: plainResult)
|
||||
} else if let plainError {
|
||||
errorMessage = plainError
|
||||
}
|
||||
isLoading = false
|
||||
continuation.resume()
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -317,6 +387,12 @@ struct CaptionsView: View {
|
||||
DispatchQueue.main.async { errorMessage = error }
|
||||
}
|
||||
}
|
||||
|
||||
private func savePlain() {
|
||||
PythonBridge.call(command: "set_plain_subtitle_config", arguments: plainConfig.arguments()) { _, error in
|
||||
DispatchQueue.main.async { errorMessage = error }
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
/// O estilo das legendas dinâmicas, no formato que a tela edita e o bridge
|
||||
@@ -402,6 +478,69 @@ struct CaptionStyleConfig {
|
||||
}
|
||||
}
|
||||
|
||||
struct PlainSubtitleConfig {
|
||||
var font: String
|
||||
var fontSize: Double
|
||||
var fontColor: String
|
||||
var maxWords: Double
|
||||
var positionY: Double
|
||||
var uppercase: Bool
|
||||
var keepPunctuation: Bool
|
||||
var textScale: Double
|
||||
|
||||
static let defaults = PlainSubtitleConfig(
|
||||
font: "Helvetica Neue",
|
||||
fontSize: 82,
|
||||
fontColor: "1 1 1 1",
|
||||
maxWords: 7,
|
||||
positionY: -820,
|
||||
uppercase: false,
|
||||
keepPunctuation: true,
|
||||
textScale: 2.0
|
||||
)
|
||||
|
||||
init(from json: [String: Any]) {
|
||||
let d = PlainSubtitleConfig.defaults
|
||||
self.init(
|
||||
font: json["font"] as? String ?? d.font,
|
||||
fontSize: (json["font_size"] as? NSNumber)?.doubleValue ?? d.fontSize,
|
||||
fontColor: json["font_color"] as? String ?? d.fontColor,
|
||||
maxWords: (json["max_words"] as? NSNumber)?.doubleValue ?? d.maxWords,
|
||||
positionY: (json["position_y"] as? NSNumber)?.doubleValue ?? d.positionY,
|
||||
uppercase: json["uppercase"] as? Bool ?? d.uppercase,
|
||||
keepPunctuation: json["keep_punctuation"] as? Bool ?? d.keepPunctuation,
|
||||
textScale: (json["text_scale"] as? NSNumber)?.doubleValue ?? d.textScale
|
||||
)
|
||||
}
|
||||
|
||||
init(
|
||||
font: String, fontSize: Double, fontColor: String, maxWords: Double,
|
||||
positionY: Double, uppercase: Bool, keepPunctuation: Bool, textScale: Double
|
||||
) {
|
||||
self.font = font
|
||||
self.fontSize = fontSize
|
||||
self.fontColor = fontColor
|
||||
self.maxWords = maxWords
|
||||
self.positionY = positionY
|
||||
self.uppercase = uppercase
|
||||
self.keepPunctuation = keepPunctuation
|
||||
self.textScale = textScale
|
||||
}
|
||||
|
||||
func arguments() -> [String: Any] {
|
||||
[
|
||||
"font": font,
|
||||
"font_size": Int(fontSize),
|
||||
"font_color": fontColor,
|
||||
"max_words": Int(maxWords),
|
||||
"position_y": positionY,
|
||||
"uppercase": uppercase,
|
||||
"keep_punctuation": keepPunctuation,
|
||||
"text_scale": textScale,
|
||||
]
|
||||
}
|
||||
}
|
||||
|
||||
extension Color {
|
||||
/// Parses an FCPXML "R G B A" space-separated 0-1 string into a Color.
|
||||
init(rgbaString: String) {
|
||||
|
||||
@@ -15,6 +15,11 @@ struct ModelDownloadView: View {
|
||||
@State private var hfTokenText: String = ""
|
||||
@State private var numSpeakersText: String = ""
|
||||
@State private var language: String = "auto"
|
||||
@State private var acousticsAvailable: Bool?
|
||||
@State private var acousticsMessage: String = ""
|
||||
@State private var isInstallingAcoustics = false
|
||||
@State private var acousticsInstallLog: String = ""
|
||||
@State private var acousticsInstallError: String?
|
||||
|
||||
private let languages: [(String, String)] = [
|
||||
("auto", "Detectar automaticamente"),
|
||||
@@ -33,6 +38,7 @@ struct ModelDownloadView: View {
|
||||
var body: some View {
|
||||
Form {
|
||||
storageSection
|
||||
acousticsSection
|
||||
diarizationSection
|
||||
if let errorMessage {
|
||||
Section {
|
||||
@@ -65,7 +71,7 @@ struct ModelDownloadView: View {
|
||||
}
|
||||
}
|
||||
.formStyle(.grouped)
|
||||
.task { await refresh() }
|
||||
.task { await refresh(); checkAcoustics() }
|
||||
}
|
||||
|
||||
// MARK: - Transcription language
|
||||
@@ -95,6 +101,98 @@ struct ModelDownloadView: View {
|
||||
PythonBridge.call(command: "set_language", arguments: ["language": code]) { _, _ in }
|
||||
}
|
||||
|
||||
// MARK: - Acoustic analysis (librosa)
|
||||
|
||||
/// A ênfase de voz (pitch/energia) precisa do `librosa`, que é uma
|
||||
/// dependência opcional — sem ela `layers.acoustics` vem `false` na
|
||||
/// análise e a decisão de zoom fica sem base real. Antes disso só dava
|
||||
/// pra descobrir lendo o JSON exportado; agora o app já diz e resolve.
|
||||
private var acousticsSection: some View {
|
||||
Section {
|
||||
VStack(alignment: .leading, spacing: 10) {
|
||||
if let acousticsAvailable {
|
||||
Label(
|
||||
acousticsMessage.isEmpty
|
||||
? (acousticsAvailable ? "Disponível" : "Indisponível")
|
||||
: acousticsMessage,
|
||||
systemImage: acousticsAvailable ? "checkmark.circle.fill" : "exclamationmark.triangle.fill"
|
||||
)
|
||||
.font(.caption)
|
||||
.foregroundStyle(acousticsAvailable ? Color.green : Color.orange)
|
||||
} else {
|
||||
Label("Verificando…", systemImage: "hourglass")
|
||||
.font(.caption).foregroundStyle(.secondary)
|
||||
}
|
||||
|
||||
if acousticsAvailable == false {
|
||||
Button {
|
||||
installAcoustics()
|
||||
} label: {
|
||||
if isInstallingAcoustics {
|
||||
HStack { ProgressView().controlSize(.small); Text("Instalando…") }
|
||||
} else {
|
||||
Label("Instalar (uv sync --all-extras)", systemImage: "arrow.down.circle")
|
||||
}
|
||||
}
|
||||
.disabled(isInstallingAcoustics)
|
||||
|
||||
if !acousticsInstallLog.isEmpty {
|
||||
ScrollView {
|
||||
Text(acousticsInstallLog)
|
||||
.font(.system(.caption2, design: .monospaced))
|
||||
.foregroundStyle(.secondary)
|
||||
.frame(maxWidth: .infinity, alignment: .leading)
|
||||
}
|
||||
.frame(height: 90)
|
||||
.background(RoundedRectangle(cornerRadius: 6).fill(Color.secondary.opacity(0.06)))
|
||||
}
|
||||
if let acousticsInstallError {
|
||||
Label(acousticsInstallError, systemImage: "xmark.circle.fill")
|
||||
.font(.caption).foregroundStyle(.red)
|
||||
}
|
||||
}
|
||||
}
|
||||
} header: {
|
||||
Text("Análise Acústica (zoom por voz)")
|
||||
} footer: {
|
||||
Text("Mede a energia e o tom de voz de verdade, para os candidatos a zoom da edição por voz. Sem isso, a análise ainda transcreve e decide cortes pelo texto — só o zoom fica sem base acústica.")
|
||||
.font(.caption)
|
||||
.foregroundStyle(.secondary)
|
||||
}
|
||||
}
|
||||
|
||||
private func checkAcoustics() {
|
||||
PythonBridge.call(command: "acoustics_capability") { result, err in
|
||||
DispatchQueue.main.async {
|
||||
guard let result, result["ok"] as? Bool == true else { return }
|
||||
acousticsAvailable = result["available"] as? Bool
|
||||
acousticsMessage = result["message"] as? String ?? ""
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
private func installAcoustics() {
|
||||
isInstallingAcoustics = true
|
||||
acousticsInstallLog = ""
|
||||
acousticsInstallError = nil
|
||||
// --all-extras, não só "intelligence": `uv sync` substitui o
|
||||
// ambiente pelos extras pedidos em vez de somar, então um sync
|
||||
// parcial aqui derrubaria dev/transcribe/diarização já instalados.
|
||||
PythonBridge.runUV(arguments: ["sync", "--all-extras"]) { line in
|
||||
DispatchQueue.main.async {
|
||||
acousticsInstallLog += (acousticsInstallLog.isEmpty ? "" : "\n") + line
|
||||
}
|
||||
} completion: { code, err in
|
||||
DispatchQueue.main.async {
|
||||
isInstallingAcoustics = false
|
||||
if code != 0 {
|
||||
acousticsInstallError = err ?? "Falha ao instalar."
|
||||
}
|
||||
checkAcoustics()
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// MARK: - Diarization
|
||||
|
||||
private var diarizationSection: some View {
|
||||
|
||||
@@ -116,6 +116,145 @@ struct ZoomClip: Identifiable {
|
||||
}
|
||||
}
|
||||
|
||||
/// One word inside a phrase, with the acoustics that justify an emphasis.
|
||||
struct ReviewWord: Identifiable {
|
||||
let id: Int
|
||||
let text: String
|
||||
let start: Double
|
||||
let end: Double
|
||||
let energy: Double
|
||||
let emphasis: Double
|
||||
|
||||
init(id: Int, json: [String: Any]) {
|
||||
self.id = id
|
||||
text = json["text"] as? String ?? ""
|
||||
start = json["start"] as? Double ?? 0
|
||||
end = json["end"] as? Double ?? 0
|
||||
energy = json["energy"] as? Double ?? 0
|
||||
emphasis = json["emphasis"] as? Double ?? 0
|
||||
}
|
||||
}
|
||||
|
||||
/// A phrase in the review step — one spoken line plus the decision made about
|
||||
/// it. Mirrors `fcpxml/phrase_review.py`; `emphasis` is 0–3 and everything
|
||||
/// mutable here is what the editor is allowed to change.
|
||||
struct ReviewPhrase: Identifiable {
|
||||
let id: Int
|
||||
let start: Double
|
||||
let end: Double
|
||||
var trimStart: Double
|
||||
var trimEnd: Double
|
||||
var text: String
|
||||
let speaker: String
|
||||
var active: Bool
|
||||
var emphasis: Int
|
||||
var track: String
|
||||
let peakEmphasis: Double
|
||||
let emotion: String
|
||||
let emotionConfidence: Double
|
||||
let takeBoundary: Bool
|
||||
let gapBefore: Double
|
||||
let reason: String
|
||||
let words: [ReviewWord]
|
||||
|
||||
static let trackScript = "roteiro"
|
||||
static let trackBackstage = "bastidor"
|
||||
|
||||
/// Delivery emotion as the analysis names it, in the user's language plus a
|
||||
/// glyph — the label alone is too easy to skim past in a dense list.
|
||||
static func emotionLabel(_ emotion: String) -> (String, String) {
|
||||
switch emotion {
|
||||
case "excited": return ("Empolgado", "flame")
|
||||
case "tense": return ("Tenso", "bolt")
|
||||
case "calm": return ("Calmo", "leaf")
|
||||
case "reflective": return ("Reflexivo", "moon")
|
||||
default: return ("Neutro", "circle")
|
||||
}
|
||||
}
|
||||
|
||||
init(json: [String: Any]) {
|
||||
id = json["index"] as? Int ?? 0
|
||||
start = json["start"] as? Double ?? 0
|
||||
end = json["end"] as? Double ?? 0
|
||||
trimStart = json["trim_start"] as? Double ?? (json["start"] as? Double ?? 0)
|
||||
trimEnd = json["trim_end"] as? Double ?? (json["end"] as? Double ?? 0)
|
||||
text = json["text"] as? String ?? ""
|
||||
speaker = json["speaker"] as? String ?? ""
|
||||
active = json["active"] as? Bool ?? true
|
||||
emphasis = json["emphasis"] as? Int ?? 0
|
||||
track = json["track"] as? String ?? ReviewPhrase.trackScript
|
||||
peakEmphasis = json["peak_emphasis"] as? Double ?? 0
|
||||
emotion = json["emotion"] as? String ?? "neutral"
|
||||
emotionConfidence = json["emotion_confidence"] as? Double ?? 0
|
||||
takeBoundary = json["take_boundary"] as? Bool ?? false
|
||||
gapBefore = json["gap_before"] as? Double ?? 0
|
||||
reason = json["reason"] as? String ?? ""
|
||||
words = (json["words"] as? [[String: Any]] ?? [])
|
||||
.enumerated().map { ReviewWord(id: $0.offset, json: $0.element) }
|
||||
}
|
||||
|
||||
var asJSON: [String: Any] {
|
||||
[
|
||||
"index": id,
|
||||
"start": start,
|
||||
"end": end,
|
||||
"trim_start": trimStart,
|
||||
"trim_end": trimEnd,
|
||||
"text": text,
|
||||
"speaker": speaker,
|
||||
"active": active,
|
||||
"emphasis": emphasis,
|
||||
"track": track,
|
||||
"reason": reason,
|
||||
]
|
||||
}
|
||||
|
||||
var isBackstage: Bool { track == ReviewPhrase.trackBackstage }
|
||||
var isTrimmed: Bool { trimStart > start + 0.001 || trimEnd < end - 0.001 }
|
||||
var timecode: String {
|
||||
String(format: "%02d:%02d", Int(start) / 60, Int(start) % 60)
|
||||
}
|
||||
|
||||
/// The word boundaries a trim handle is allowed to land on.
|
||||
func snap(_ time: Double, edge: TrimEdge) -> Double {
|
||||
let boundaries = words.map { edge == .start ? $0.start : $0.end }.filter { $0 > 0 }
|
||||
guard let nearest = boundaries.min(by: { abs($0 - time) < abs($1 - time) }) else {
|
||||
return time
|
||||
}
|
||||
return nearest
|
||||
}
|
||||
}
|
||||
|
||||
enum TrimEdge { case start, end }
|
||||
|
||||
/// A punch-in the editor placed by hand over an arbitrary range, next to the
|
||||
/// whole-phrase zoom that an emphasis level produces. It stores only *when* —
|
||||
/// the scale and the ramp come from the Voice Analysis settings at render time.
|
||||
struct ManualZoom: Identifiable {
|
||||
let id = UUID()
|
||||
var start: Double
|
||||
var end: Double
|
||||
|
||||
/// Below this a punch-in has no room to ramp in and back out; the writer
|
||||
/// rejects the window, so offering it would place nothing.
|
||||
static let minimumDuration: Double = 0.4
|
||||
|
||||
init(start: Double, end: Double) {
|
||||
self.start = start
|
||||
self.end = end
|
||||
}
|
||||
|
||||
init?(json: [String: Any]) {
|
||||
guard let start = json["start"] as? Double, let end = json["end"] as? Double,
|
||||
end - start >= ManualZoom.minimumDuration
|
||||
else { return nil }
|
||||
self.start = start
|
||||
self.end = end
|
||||
}
|
||||
|
||||
var asJSON: [String: Any] { ["start": start, "end": end] }
|
||||
}
|
||||
|
||||
struct ZoomSegment: Identifiable {
|
||||
let id: Int
|
||||
let start: Double
|
||||
|
||||
@@ -0,0 +1,442 @@
|
||||
import AVFoundation
|
||||
import Combine
|
||||
import Foundation
|
||||
|
||||
/// State behind the wizard's emphasis-review step.
|
||||
///
|
||||
/// Holds the phrases, the selection, and the player — together, because they
|
||||
/// are one thing to the user: clicking a phrase moves the playhead, playing
|
||||
/// moves the selection, and skipping a removed line only works if whoever owns
|
||||
/// playback also knows which lines are removed.
|
||||
///
|
||||
/// The preview deliberately plays the *original* media and jumps over whatever
|
||||
/// the edit removes, instead of rendering a cut first. Rendering to check a
|
||||
/// toggle would put minutes between a decision and its result; jumping gives
|
||||
/// the same reading instantly, and the real cut is generated later from the
|
||||
/// exact same phrase list.
|
||||
@MainActor
|
||||
final class PhraseReviewModel: ObservableObject {
|
||||
@Published var phrases: [ReviewPhrase] = []
|
||||
@Published var selection: Int?
|
||||
@Published var isLoading = false
|
||||
@Published var errorMessage: String?
|
||||
@Published var currentTime: Double = 0
|
||||
@Published var isPlaying = false
|
||||
@Published var pixelsPerSecond: Double = 40
|
||||
@Published var skipRemoved = true
|
||||
@Published var zooms: [ManualZoom] = []
|
||||
/// In/out the editor dragged on the timeline, in source seconds.
|
||||
@Published var rangeStart: Double?
|
||||
@Published var rangeEnd: Double?
|
||||
|
||||
private(set) var source = ""
|
||||
private(set) var sourcePath = ""
|
||||
private(set) var duration: Double = 0
|
||||
private(set) var speakers: [String] = []
|
||||
private(set) var emotionAvailable = false
|
||||
private(set) var player: AVPlayer?
|
||||
|
||||
private var voiceTimelinePath = ""
|
||||
private var timeObserver: Any?
|
||||
private var playbackLimit: Double?
|
||||
|
||||
let minPixelsPerSecond: Double = 8
|
||||
let maxPixelsPerSecond: Double = 400
|
||||
|
||||
deinit {
|
||||
if let timeObserver, let player {
|
||||
player.removeTimeObserver(timeObserver)
|
||||
}
|
||||
}
|
||||
|
||||
// MARK: - Carregar
|
||||
|
||||
/// Builds the review from the voice timeline plus whatever the AI decided.
|
||||
/// A review saved on a previous visit wins — see `cmd_build_phrase_review` —
|
||||
/// UNLESS `fresh` is true, in which case that saved review is ignored and
|
||||
/// `active`/`emphasis`/etc. come straight from this call's `decisionsJSON`.
|
||||
/// Pass `fresh: true` when the decisions themselves changed since the
|
||||
/// review was last built (the caller re-pasted/regenerated the AI's JSON
|
||||
/// and re-ran `apply_voice_actions`) — otherwise the saved review from the
|
||||
/// PREVIOUS decisions silently wins over the fresh cut it should reflect,
|
||||
/// which is exactly the desync the wizard's "active" toggle showed against
|
||||
/// the just-reapplied FCPXML.
|
||||
func load(voiceTimelinePath: String, decisionsJSON: String,
|
||||
outputFolder: String? = nil, mediaFolder: String? = nil, fresh: Bool = false) {
|
||||
self.voiceTimelinePath = voiceTimelinePath
|
||||
isLoading = true
|
||||
errorMessage = nil
|
||||
|
||||
var arguments: [String: Any] = ["voice_timeline": voiceTimelinePath]
|
||||
if let outputFolder { arguments["output_dir"] = outputFolder }
|
||||
if let mediaFolder { arguments["media_dir"] = mediaFolder }
|
||||
if fresh { arguments["fresh"] = true }
|
||||
if let data = decisionsJSON.data(using: .utf8),
|
||||
let parsed = try? JSONSerialization.jsonObject(with: data) {
|
||||
arguments["actions"] = parsed
|
||||
}
|
||||
|
||||
PythonBridge.call(command: "build_phrase_review", arguments: arguments) { [weak self] result, error in
|
||||
Task { @MainActor in
|
||||
guard let self else { return }
|
||||
self.isLoading = false
|
||||
if let error {
|
||||
self.errorMessage = error
|
||||
return
|
||||
}
|
||||
guard let result, result["ok"] as? Bool == true else {
|
||||
self.errorMessage = result?["error"] as? String ?? "Não foi possível montar a revisão."
|
||||
return
|
||||
}
|
||||
self.apply(result)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
private func apply(_ result: [String: Any]) {
|
||||
source = result["source"] as? String ?? ""
|
||||
// The timeline JSON stores only the media's file name; the bridge
|
||||
// resolves it to something openable (see phrase_review.resolve_source).
|
||||
sourcePath = result["source_path"] as? String ?? ""
|
||||
duration = result["duration"] as? Double ?? 0
|
||||
speakers = result["speakers"] as? [String] ?? []
|
||||
emotionAvailable = result["emotion_available"] as? Bool ?? false
|
||||
phrases = (result["phrases"] as? [[String: Any]] ?? []).map { ReviewPhrase(json: $0) }
|
||||
zooms = (result["zooms"] as? [[String: Any]] ?? []).compactMap { ManualZoom(json: $0) }
|
||||
selection = phrases.first?.id
|
||||
if let errors = result["errors"] as? [String], !errors.isEmpty {
|
||||
errorMessage = "A IA mandou \(errors.count) decisão(ões) que não deu para ler — o resto foi aplicado."
|
||||
}
|
||||
preparePlayer()
|
||||
}
|
||||
|
||||
/// Point the preview at a media file the user chose by hand — the way out
|
||||
/// when the footage moved somewhere the automatic lookup can't reach.
|
||||
func useMedia(at path: String) {
|
||||
sourcePath = path
|
||||
preparePlayer()
|
||||
}
|
||||
|
||||
private func preparePlayer() {
|
||||
guard !sourcePath.isEmpty, FileManager.default.fileExists(atPath: sourcePath) else {
|
||||
player = nil
|
||||
return
|
||||
}
|
||||
if let timeObserver, let player {
|
||||
player.removeTimeObserver(timeObserver)
|
||||
self.timeObserver = nil
|
||||
}
|
||||
let asset = AVURLAsset(url: URL(fileURLWithPath: sourcePath))
|
||||
let player = AVPlayer(playerItem: AVPlayerItem(asset: asset))
|
||||
self.player = player
|
||||
// 60 Hz: the same observer drives the playhead *and* decides when to
|
||||
// jump a removed stretch, so its period is the worst-case amount of cut
|
||||
// material that can be heard before the skip lands. At 20 Hz that was an
|
||||
// audible blip on every join.
|
||||
let interval = CMTime(seconds: 1.0 / 60.0, preferredTimescale: 600)
|
||||
timeObserver = player.addPeriodicTimeObserver(forInterval: interval, queue: .main) { [weak self] time in
|
||||
Task { @MainActor in
|
||||
self?.tick(time.seconds)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// MARK: - Reprodução
|
||||
|
||||
private func tick(_ time: Double) {
|
||||
currentTime = time
|
||||
guard isPlaying else { return }
|
||||
|
||||
// Playing a single phrase or a marked range stops at its out point
|
||||
// instead of running on into the rest of the take.
|
||||
if let limit = playbackLimit, time >= limit {
|
||||
pause()
|
||||
seek(to: limit)
|
||||
return
|
||||
}
|
||||
|
||||
if skipRemoved, let jump = nextKeptTime(after: time), jump > time {
|
||||
seek(to: jump)
|
||||
}
|
||||
if let phrase = phrase(at: time), selection != phrase.id {
|
||||
selection = phrase.id
|
||||
}
|
||||
}
|
||||
|
||||
/// Where playback should resume when `time` lands on removed material.
|
||||
/// Returns nil when the time is on material that survives.
|
||||
func nextKeptTime(after time: Double) -> Double? {
|
||||
for phrase in phrases where time >= phrase.start - 0.001 && time < phrase.end {
|
||||
if !phrase.active { return phrase.end }
|
||||
if time < phrase.trimStart { return phrase.trimStart }
|
||||
if time >= phrase.trimEnd { return phrase.end }
|
||||
return nil
|
||||
}
|
||||
return nil
|
||||
}
|
||||
|
||||
func togglePlay() {
|
||||
if isPlaying {
|
||||
pause()
|
||||
} else {
|
||||
playbackLimit = nil
|
||||
play()
|
||||
}
|
||||
}
|
||||
|
||||
private func play() {
|
||||
guard let player else { return }
|
||||
if skipRemoved, let jump = nextKeptTime(after: currentTime) { seek(to: jump) }
|
||||
player.play()
|
||||
isPlaying = true
|
||||
}
|
||||
|
||||
func pause() {
|
||||
player?.pause()
|
||||
isPlaying = false
|
||||
playbackLimit = nil
|
||||
}
|
||||
|
||||
/// Play exactly one span and stop — how a cut is judged: in context, at
|
||||
/// speed, without hunting for the out point by hand.
|
||||
func playRange(from start: Double, to end: Double) {
|
||||
guard end > start else { return }
|
||||
seek(to: start)
|
||||
playbackLimit = end
|
||||
player?.play()
|
||||
isPlaying = true
|
||||
}
|
||||
|
||||
func playSelectedPhrase() {
|
||||
guard let selection, let phrase = phrases.first(where: { $0.id == selection })
|
||||
else { return }
|
||||
playRange(from: phrase.active ? phrase.trimStart : phrase.start,
|
||||
to: phrase.active ? phrase.trimEnd : phrase.end)
|
||||
}
|
||||
|
||||
func seek(to time: Double) {
|
||||
currentTime = max(0, time)
|
||||
player?.seek(to: CMTime(seconds: max(0, time), preferredTimescale: 600),
|
||||
toleranceBefore: .zero, toleranceAfter: .zero)
|
||||
}
|
||||
|
||||
/// Move the playhead to a phrase and select it.
|
||||
func goTo(phraseID: Int) {
|
||||
guard let phrase = phrases.first(where: { $0.id == phraseID }) else { return }
|
||||
selection = phraseID
|
||||
seek(to: phrase.active ? phrase.trimStart : phrase.start)
|
||||
}
|
||||
|
||||
func phrase(at time: Double) -> ReviewPhrase? {
|
||||
phrases.first { time >= $0.start && time < $0.end }
|
||||
}
|
||||
|
||||
func selectNeighbour(_ delta: Int) {
|
||||
guard let selection, let index = phrases.firstIndex(where: { $0.id == selection }) else {
|
||||
if let first = phrases.first { goTo(phraseID: first.id) }
|
||||
return
|
||||
}
|
||||
let next = min(max(0, index + delta), phrases.count - 1)
|
||||
goTo(phraseID: phrases[next].id)
|
||||
}
|
||||
|
||||
// MARK: - Edições
|
||||
|
||||
private func update(_ id: Int, _ change: (inout ReviewPhrase) -> Void) {
|
||||
guard let index = phrases.firstIndex(where: { $0.id == id }) else { return }
|
||||
change(&phrases[index])
|
||||
}
|
||||
|
||||
func setEmphasis(_ level: Int, for id: Int) {
|
||||
update(id) { $0.emphasis = min(3, max(0, level)) }
|
||||
}
|
||||
|
||||
func toggleActive(_ id: Int) {
|
||||
update(id) { $0.active.toggle() }
|
||||
}
|
||||
|
||||
func setTrack(_ track: String, for id: Int) {
|
||||
update(id) { $0.track = track }
|
||||
}
|
||||
|
||||
func setText(_ text: String, for id: Int) {
|
||||
update(id) { $0.text = text }
|
||||
}
|
||||
|
||||
/// Trim a phrase's head or tail, landing on a word boundary.
|
||||
/// A trim that would swallow the whole line is refused — deactivating the
|
||||
/// phrase is the way to remove it, and doing it by accident with a drag
|
||||
/// would lose the emphasis decision along with the line.
|
||||
func trim(_ id: Int, edge: TrimEdge, to time: Double) {
|
||||
update(id) { phrase in
|
||||
let snapped = phrase.snap(time, edge: edge)
|
||||
switch edge {
|
||||
case .start:
|
||||
let value = min(max(phrase.start, snapped), phrase.trimEnd - 0.1)
|
||||
if value < phrase.trimEnd { phrase.trimStart = value }
|
||||
case .end:
|
||||
let value = max(min(phrase.end, snapped), phrase.trimStart + 0.1)
|
||||
if value > phrase.trimStart { phrase.trimEnd = value }
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
func resetTrim(_ id: Int) {
|
||||
update(id) { $0.trimStart = $0.start; $0.trimEnd = $0.end }
|
||||
}
|
||||
|
||||
/// Trim everything before/after a given word — the text-first way to cut,
|
||||
/// since the editor reads the line and points at where it should begin.
|
||||
/// Clicking the word that is ALREADY that edge toggles it back off —
|
||||
/// the trim on that side resets to the phrase's own start/end — so the
|
||||
/// same click that sets a boundary also clears it, instead of needing
|
||||
/// the separate "Inteira" button for a one-sided undo.
|
||||
func trimToWord(_ word: ReviewWord, edge: TrimEdge, in id: Int) {
|
||||
guard let phrase = phrases.first(where: { $0.id == id }) else { return }
|
||||
let epsilon = 0.001
|
||||
switch edge {
|
||||
case .start where abs(word.start - phrase.trimStart) < epsilon:
|
||||
update(id) { $0.trimStart = $0.start }
|
||||
case .end where abs(word.end - phrase.trimEnd) < epsilon:
|
||||
update(id) { $0.trimEnd = $0.end }
|
||||
default:
|
||||
trim(id, edge: edge, to: edge == .start ? word.start : word.end)
|
||||
}
|
||||
}
|
||||
|
||||
// MARK: - Trecho marcado e zooms
|
||||
|
||||
var hasRange: Bool {
|
||||
guard let rangeStart, let rangeEnd else { return false }
|
||||
return rangeEnd - rangeStart >= ManualZoom.minimumDuration
|
||||
}
|
||||
|
||||
var rangeSpan: (start: Double, end: Double)? {
|
||||
guard let rangeStart, let rangeEnd, rangeEnd > rangeStart else { return nil }
|
||||
return (rangeStart, rangeEnd)
|
||||
}
|
||||
|
||||
func setRange(from start: Double, to end: Double) {
|
||||
rangeStart = min(start, end)
|
||||
rangeEnd = max(start, end)
|
||||
}
|
||||
|
||||
func clearRange() {
|
||||
rangeStart = nil
|
||||
rangeEnd = nil
|
||||
}
|
||||
|
||||
/// Add a punch-in over the marked range. Scale and ramp are not stored:
|
||||
/// they come from the "Análise de Voz" settings when the edit is rendered,
|
||||
/// so changing the look there restyles every zoom at once.
|
||||
func addZoomForRange() {
|
||||
guard let span = rangeSpan, span.end - span.start >= ManualZoom.minimumDuration
|
||||
else { return }
|
||||
zooms.append(ManualZoom(start: span.start, end: span.end))
|
||||
zooms.sort { $0.start < $1.start }
|
||||
clearRange()
|
||||
}
|
||||
|
||||
func addZoomForPhrase(_ id: Int) {
|
||||
guard let phrase = phrases.first(where: { $0.id == id }) else { return }
|
||||
zooms.append(ManualZoom(start: phrase.trimStart, end: phrase.trimEnd))
|
||||
zooms.sort { $0.start < $1.start }
|
||||
}
|
||||
|
||||
func removeZoom(_ id: UUID) {
|
||||
zooms.removeAll { $0.id == id }
|
||||
}
|
||||
|
||||
func zoom(at time: Double) -> ManualZoom? {
|
||||
zooms.first { time >= $0.start && time <= $0.end }
|
||||
}
|
||||
|
||||
func setEmphasisForAll(_ level: Int) {
|
||||
for index in phrases.indices where phrases[index].active {
|
||||
phrases[index].emphasis = level
|
||||
}
|
||||
}
|
||||
|
||||
// MARK: - Resumo e gravação
|
||||
|
||||
var emphasisCount: Int { phrases.filter { $0.active && $0.emphasis >= 1 }.count }
|
||||
var removedCount: Int { phrases.filter { !$0.active }.count }
|
||||
var keptDuration: Double {
|
||||
phrases.filter { $0.active }.reduce(0) { $0 + ($1.trimEnd - $1.trimStart) }
|
||||
}
|
||||
|
||||
// MARK: - Tempo compactado (sem os vãos do que foi cortado)
|
||||
|
||||
/// Kept spans of source media, in order, each carrying the position it
|
||||
/// lands at once every removed stretch between phrases is squeezed out.
|
||||
/// The timeline draws and scrubs in this space so it reads like the cut
|
||||
/// itself instead of the raw take with holes in it.
|
||||
private var keptSegments: [(rawStart: Double, rawEnd: Double, compactStart: Double)] {
|
||||
var offset = 0.0
|
||||
var segments: [(Double, Double, Double)] = []
|
||||
for phrase in phrases.sorted(by: { $0.start < $1.start }) where phrase.active {
|
||||
guard phrase.trimEnd > phrase.trimStart else { continue }
|
||||
segments.append((phrase.trimStart, phrase.trimEnd, offset))
|
||||
offset += phrase.trimEnd - phrase.trimStart
|
||||
}
|
||||
return segments
|
||||
}
|
||||
|
||||
/// Maps a raw source-media time to its position on the compacted timeline.
|
||||
/// Time inside removed material collapses to the boundary of the nearest
|
||||
/// kept segment, so cut stretches take up no space at all.
|
||||
func compactTime(_ raw: Double) -> Double {
|
||||
let segments = keptSegments
|
||||
for segment in segments {
|
||||
if raw < segment.rawStart { return segment.compactStart }
|
||||
if raw <= segment.rawEnd { return segment.compactStart + (raw - segment.rawStart) }
|
||||
}
|
||||
guard let last = segments.last else { return 0 }
|
||||
return raw >= last.rawEnd ? last.compactStart + (last.rawEnd - last.rawStart) : 0
|
||||
}
|
||||
|
||||
/// The inverse of `compactTime`: where a click on the compacted timeline
|
||||
/// lands in the raw source media, for seeking and scrubbing.
|
||||
func rawTime(fromCompact compact: Double) -> Double {
|
||||
let segments = keptSegments
|
||||
for segment in segments {
|
||||
let compactEnd = segment.compactStart + (segment.rawEnd - segment.rawStart)
|
||||
if compact <= compactEnd {
|
||||
return segment.rawStart + max(0, compact - segment.compactStart)
|
||||
}
|
||||
}
|
||||
return segments.last?.rawEnd ?? 0
|
||||
}
|
||||
|
||||
/// Persists the edited review plus the actions derived from it. Called when
|
||||
/// the wizard advances — the render itself happens in the next step.
|
||||
/// Persists the edited review and hands back BOTH paths it wrote:
|
||||
/// `review_path` (the human-readable `_phrase_review.json`) and
|
||||
/// `actions_path` (`_phrase_actions.json`, the cut/zoom list derived from
|
||||
/// it — what `finalizeProcessing` needs to actually apply the review's
|
||||
/// active/inactive decisions instead of just filing them away).
|
||||
func save(completion: @escaping (_ reviewPath: String?, _ actionsPath: String?) -> Void) {
|
||||
guard !voiceTimelinePath.isEmpty, !phrases.isEmpty else {
|
||||
completion(nil, nil)
|
||||
return
|
||||
}
|
||||
let arguments: [String: Any] = [
|
||||
"voice_timeline": voiceTimelinePath,
|
||||
"source": source,
|
||||
"duration": duration,
|
||||
"speakers": speakers,
|
||||
"phrases": phrases.map { $0.asJSON },
|
||||
"zooms": zooms.map { $0.asJSON },
|
||||
]
|
||||
PythonBridge.call(command: "save_phrase_review", arguments: arguments) { result, error in
|
||||
Task { @MainActor in
|
||||
if let error {
|
||||
completion(nil, nil)
|
||||
_ = error
|
||||
return
|
||||
}
|
||||
completion(result?["review_path"] as? String, result?["actions_path"] as? String)
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,313 @@
|
||||
import SwiftUI
|
||||
|
||||
/// The wizard's emphasis-review step.
|
||||
///
|
||||
/// Every decision here is about a *sentence* read from the original
|
||||
/// transcription, so the phrases are listed in full — each line shows the text
|
||||
/// as it will be said, a switch to keep or drop it from the cut, and the
|
||||
/// emphasis level. Selecting a line in the list also selects its block on the
|
||||
/// timeline below, and vice-versa.
|
||||
struct PhraseReviewView: View {
|
||||
@ObservedObject var model: PhraseReviewModel
|
||||
|
||||
var body: some View {
|
||||
VSplitView {
|
||||
VStack(spacing: 0) {
|
||||
inspectorHeader
|
||||
Divider()
|
||||
List(selection: $model.selection) {
|
||||
ForEach($model.phrases) { $phrase in
|
||||
PhraseRow(phrase: $phrase, model: model)
|
||||
.tag(phrase.id)
|
||||
}
|
||||
}
|
||||
.listStyle(.inset)
|
||||
.onChange(of: model.selection) { _, newValue in
|
||||
if let newValue { model.goTo(phraseID: newValue) }
|
||||
}
|
||||
Divider()
|
||||
summaryBar
|
||||
}
|
||||
.frame(minHeight: 240)
|
||||
|
||||
TimelineTracksView(model: model)
|
||||
.frame(minHeight: 190, idealHeight: 210)
|
||||
}
|
||||
.overlay { if model.isLoading { loadingOverlay } }
|
||||
.focusable()
|
||||
.onKeyPress(.space) { model.togglePlay(); return .handled }
|
||||
.onKeyPress(.return) { model.playSelectedPhrase(); return .handled }
|
||||
.onKeyPress(.leftArrow) { model.selectNeighbour(-1); return .handled }
|
||||
.onKeyPress(.rightArrow) { model.selectNeighbour(1); return .handled }
|
||||
.onKeyPress(characters: .decimalDigits) { press in
|
||||
guard let level = Int(press.characters), (0...3).contains(level),
|
||||
let selection = model.selection else { return .ignored }
|
||||
model.setEmphasis(level, for: selection)
|
||||
return .handled
|
||||
}
|
||||
}
|
||||
|
||||
private var loadingOverlay: some View {
|
||||
ZStack {
|
||||
Color(nsColor: .windowBackgroundColor).opacity(0.85)
|
||||
VStack(spacing: 10) {
|
||||
ProgressView()
|
||||
Text("Montando a revisão…").font(.callout).foregroundStyle(.secondary)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
private var summaryBar: some View {
|
||||
HStack(spacing: 16) {
|
||||
summaryItem("text.quote", "\(model.phrases.count) frases")
|
||||
summaryItem("sparkles", "\(model.emphasisCount) com ênfase")
|
||||
summaryItem("scissors", "\(model.removedCount) fora do corte")
|
||||
summaryItem("clock", durationLabel(model.keptDuration))
|
||||
if !model.zooms.isEmpty {
|
||||
summaryItem("plus.magnifyingglass", "\(model.zooms.count) zooms")
|
||||
}
|
||||
Spacer()
|
||||
if let phrase = selectedPhrase, !phrase.reason.isEmpty {
|
||||
Label(phrase.reason, systemImage: "brain")
|
||||
.font(.caption).foregroundStyle(.secondary)
|
||||
.lineLimit(1).truncationMode(.tail)
|
||||
}
|
||||
}
|
||||
.padding(.horizontal, 14)
|
||||
.padding(.vertical, 8)
|
||||
}
|
||||
|
||||
private func summaryItem(_ icon: String, _ text: String) -> some View {
|
||||
Label(text, systemImage: icon).font(.caption).foregroundStyle(.secondary)
|
||||
}
|
||||
|
||||
private func durationLabel(_ seconds: Double) -> String {
|
||||
String(format: "%02d:%02d finais", Int(seconds) / 60, Int(seconds) % 60)
|
||||
}
|
||||
|
||||
private func pickMedia() {
|
||||
let panel = NSOpenPanel()
|
||||
panel.canChooseFiles = true
|
||||
panel.canChooseDirectories = false
|
||||
panel.allowsMultipleSelection = false
|
||||
panel.prompt = "Usar esta mídia"
|
||||
panel.message = model.source.isEmpty
|
||||
? "Escolha o arquivo de vídeo desta gravação."
|
||||
: "Escolha onde está \(model.source)."
|
||||
if panel.runModal() == .OK, let url = panel.url {
|
||||
model.useMedia(at: url.path)
|
||||
}
|
||||
}
|
||||
|
||||
private var selectedPhrase: ReviewPhrase? {
|
||||
guard let selection = model.selection else { return nil }
|
||||
return model.phrases.first { $0.id == selection }
|
||||
}
|
||||
|
||||
// MARK: - Inspector de frases
|
||||
|
||||
private var inspectorPane: some View {
|
||||
VStack(spacing: 0) {
|
||||
inspectorHeader
|
||||
Divider()
|
||||
List(selection: $model.selection) {
|
||||
ForEach($model.phrases) { $phrase in
|
||||
PhraseRow(phrase: $phrase, model: model)
|
||||
.tag(phrase.id)
|
||||
}
|
||||
}
|
||||
.listStyle(.inset)
|
||||
.onChange(of: model.selection) { _, newValue in
|
||||
if let newValue { model.goTo(phraseID: newValue) }
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
private var inspectorHeader: some View {
|
||||
VStack(alignment: .leading, spacing: 6) {
|
||||
Text("Frases").font(.headline)
|
||||
Text("Só as frases com ênfase recebem zoom e legenda dinâmica. O resto fica com legenda comum.")
|
||||
.font(.caption).foregroundStyle(.secondary)
|
||||
if !model.emotionAvailable {
|
||||
Label("Emoção da fala não foi detectada nesta análise — ligue em Avançado → Análise de Voz e refaça o passo 3.",
|
||||
systemImage: "waveform.path.ecg")
|
||||
.font(.caption2).foregroundStyle(.secondary)
|
||||
}
|
||||
HStack(spacing: 8) {
|
||||
Button("Limpar ênfases") { model.setEmphasisForAll(0) }
|
||||
.buttonStyle(.link).font(.caption)
|
||||
Spacer()
|
||||
Text("0–3 no teclado · ← → navega")
|
||||
.font(.caption2).foregroundStyle(.secondary)
|
||||
}
|
||||
}
|
||||
.padding(12)
|
||||
}
|
||||
}
|
||||
|
||||
/// One phrase in the inspector: the line as it will be said, plus every
|
||||
/// decision attached to it. Kept in one row on purpose — jumping to a separate
|
||||
/// detail pane to set a toggle would double the clicks on the most repeated
|
||||
/// action in the screen.
|
||||
private struct PhraseRow: View {
|
||||
@Binding var phrase: ReviewPhrase
|
||||
@ObservedObject var model: PhraseReviewModel
|
||||
@State private var isEditing = false
|
||||
|
||||
var body: some View {
|
||||
VStack(alignment: .leading, spacing: 6) {
|
||||
HStack(spacing: 6) {
|
||||
Text(phrase.timecode)
|
||||
.font(.system(.caption2, design: .monospaced))
|
||||
.foregroundStyle(.secondary)
|
||||
if phrase.takeBoundary {
|
||||
Image(systemName: "scissors.badge.ellipsis")
|
||||
.font(.caption2).foregroundStyle(.orange)
|
||||
.help("Nova tomada começa aqui")
|
||||
}
|
||||
if phrase.isTrimmed {
|
||||
Image(systemName: "arrow.left.and.right.square")
|
||||
.font(.caption2).foregroundStyle(.blue)
|
||||
.help("Frase cortada nas pontas")
|
||||
}
|
||||
if model.emotionAvailable {
|
||||
emotionChip
|
||||
}
|
||||
Spacer()
|
||||
Toggle("", isOn: $phrase.active)
|
||||
.toggleStyle(.switch)
|
||||
.controlSize(.mini)
|
||||
.labelsHidden()
|
||||
.help(phrase.active ? "No corte" : "Fora do corte")
|
||||
}
|
||||
|
||||
if isEditing {
|
||||
TextField("Texto da frase", text: $phrase.text, axis: .vertical)
|
||||
.textFieldStyle(.roundedBorder)
|
||||
.font(.callout)
|
||||
.onSubmit { isEditing = false }
|
||||
} else {
|
||||
Text(phrase.text.isEmpty ? "(sem texto)" : phrase.text)
|
||||
.font(.callout)
|
||||
.foregroundStyle(phrase.active ? .primary : .secondary)
|
||||
.strikethrough(!phrase.active)
|
||||
.onTapGesture(count: 2) { isEditing = true }
|
||||
}
|
||||
|
||||
HStack(spacing: 8) {
|
||||
Picker("", selection: $phrase.emphasis) {
|
||||
ForEach(0..<4, id: \.self) { level in
|
||||
Text(EmphasisPalette.label(level)).tag(level)
|
||||
}
|
||||
}
|
||||
.pickerStyle(.segmented)
|
||||
.controlSize(.mini)
|
||||
.labelsHidden()
|
||||
.disabled(!phrase.active)
|
||||
|
||||
Picker("", selection: $phrase.track) {
|
||||
Text("Roteiro").tag(ReviewPhrase.trackScript)
|
||||
Text("Bastidor").tag(ReviewPhrase.trackBackstage)
|
||||
}
|
||||
.pickerStyle(.menu)
|
||||
.controlSize(.mini)
|
||||
.labelsHidden()
|
||||
.frame(width: 92)
|
||||
}
|
||||
|
||||
if model.selection == phrase.id && !phrase.words.isEmpty {
|
||||
wordTrimmer
|
||||
}
|
||||
}
|
||||
.padding(.vertical, 4)
|
||||
.opacity(phrase.active ? 1 : 0.55)
|
||||
}
|
||||
|
||||
/// The delivery emotion the acoustics suggest. Shown faded below its own
|
||||
/// confidence: a guess the analysis is unsure about should not compete for
|
||||
/// attention with the emphasis decision, which is the point of the row.
|
||||
private var emotionChip: some View {
|
||||
let (label, icon) = ReviewPhrase.emotionLabel(phrase.emotion)
|
||||
return Label(label, systemImage: icon)
|
||||
.font(.caption2)
|
||||
.padding(.horizontal, 5)
|
||||
.padding(.vertical, 1)
|
||||
.background(
|
||||
Capsule().fill(Color.secondary.opacity(0.12))
|
||||
)
|
||||
.foregroundStyle(phrase.emotionConfidence >= 0.5 ? .secondary : .tertiary)
|
||||
.help("Emoção da entrega: \(label) — confiança \(Int(phrase.emotionConfidence * 100))%")
|
||||
}
|
||||
|
||||
/// Trimming by pointing at the transcript: click a word to start the phrase
|
||||
/// there, option-click to end it there. Same edit as dragging the block's
|
||||
/// edge on the timeline, but reachable while reading the line.
|
||||
private var wordTrimmer: some View {
|
||||
VStack(alignment: .leading, spacing: 4) {
|
||||
HStack(spacing: 4) {
|
||||
Text("Cortar pelas palavras").font(.caption2).foregroundStyle(.secondary)
|
||||
Spacer()
|
||||
if phrase.isTrimmed {
|
||||
Button("Inteira") { model.resetTrim(phrase.id) }
|
||||
.buttonStyle(.link).font(.caption2)
|
||||
}
|
||||
}
|
||||
FlowWords(words: phrase.words, phrase: phrase) { word, edge in
|
||||
model.trimToWord(word, edge: edge, in: phrase.id)
|
||||
}
|
||||
Text("Clique = começa/desfaz aqui · ⌥clique = termina/desfaz aqui · sublinhado = ênfase da palavra")
|
||||
.font(.caption2).foregroundStyle(.tertiary)
|
||||
}
|
||||
.padding(.top, 2)
|
||||
}
|
||||
}
|
||||
|
||||
/// The phrase's words as wrapping chips, dimmed where they fall outside the
|
||||
/// trim and underlined where the acoustics mark them as an emphasis peak —
|
||||
/// the same word-level signal `05-zoom.md` picks a punch-in's `start` from,
|
||||
/// made visible instead of buried in the JSON.
|
||||
private struct FlowWords: View {
|
||||
let words: [ReviewWord]
|
||||
let phrase: ReviewPhrase
|
||||
let onTrim: (ReviewWord, TrimEdge) -> Void
|
||||
|
||||
var body: some View {
|
||||
// A LazyVGrid with adaptive columns wraps chips without a custom layout;
|
||||
// phrases are short enough that the slight raggedness beats the cost of
|
||||
// hand-rolling a flow layout here.
|
||||
LazyVGrid(columns: [GridItem(.adaptive(minimum: 44), spacing: 3)],
|
||||
alignment: .leading, spacing: 3) {
|
||||
ForEach(words) { word in
|
||||
let kept = word.start >= phrase.trimStart - 0.001 && word.end <= phrase.trimEnd + 0.001
|
||||
let level = EmphasisPalette.levelFromScore(word.emphasis)
|
||||
Text(word.text)
|
||||
.font(.caption2)
|
||||
.fontWeight(level >= 2 ? .semibold : .regular)
|
||||
.padding(.horizontal, 4)
|
||||
.padding(.vertical, 2)
|
||||
.background(
|
||||
RoundedRectangle(cornerRadius: 3)
|
||||
.fill(kept ? Color.accentColor.opacity(0.12) : Color.secondary.opacity(0.08))
|
||||
)
|
||||
.overlay(alignment: .bottom) {
|
||||
if level >= 1 {
|
||||
Rectangle()
|
||||
.fill(EmphasisPalette.color(level))
|
||||
.frame(height: 2)
|
||||
.padding(.horizontal, 3)
|
||||
}
|
||||
}
|
||||
.foregroundStyle(kept ? .primary : .secondary)
|
||||
.strikethrough(!kept)
|
||||
.help(
|
||||
level >= 1
|
||||
? "Ênfase \(EmphasisPalette.label(level).lowercased()) (\(Int(word.emphasis * 100))%)"
|
||||
: "Sem ênfase"
|
||||
)
|
||||
.onTapGesture {
|
||||
onTrim(word, NSEvent.modifierFlags.contains(.option) ? .end : .start)
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -44,12 +44,26 @@ enum PythonBridge {
|
||||
return ["python3", scriptURL.path]
|
||||
}
|
||||
|
||||
/// `admin/models_api.py` lives outside `code/`, but its dependencies
|
||||
/// (`pyproject.toml`, `.venv`) live inside it. `uv run` picks the
|
||||
/// environment from the process's cwd, not from the script path — so
|
||||
/// running with cwd at the repo root made `uv` create/use a second,
|
||||
/// empty `.venv` there, silently ignoring everything installed into
|
||||
/// `code/.venv` (this cost a real debugging session: librosa/pyannote
|
||||
/// installed successfully but the app kept reporting them missing).
|
||||
/// Every `uv run` must share the same cwd as `uv sync` to see the same
|
||||
/// environment.
|
||||
static var workingDirectory: URL {
|
||||
projectRoot
|
||||
codeDirectory
|
||||
}
|
||||
|
||||
/// Directory containing `pyproject.toml` — where `uv sync` must run from.
|
||||
static var codeDirectory: URL {
|
||||
projectRoot.appendingPathComponent("code")
|
||||
}
|
||||
|
||||
/// Locate `uv` on PATH or in common install locations.
|
||||
private static func findUV() -> String? {
|
||||
static func findUV() -> String? {
|
||||
if let onPath = which("uv") { return onPath }
|
||||
let candidates = [
|
||||
"/usr/local/bin/uv",
|
||||
@@ -148,6 +162,59 @@ enum PythonBridge {
|
||||
}
|
||||
}
|
||||
|
||||
// MARK: - uv sync (installing optional extras, e.g. acoustic analysis)
|
||||
|
||||
/// Runs `uv <arguments>` from `codeDirectory` (where `pyproject.toml`
|
||||
/// lives), streaming each output line as plain text — used for
|
||||
/// `sync --extra intelligence` so "Modelos" can install the librosa
|
||||
/// extra without the user opening a terminal.
|
||||
static func runUV(arguments: [String],
|
||||
onLine: @escaping (String) -> Void,
|
||||
completion: @escaping (Int, String?) -> Void) {
|
||||
guard let uv = findUV() else {
|
||||
completion(1, "uv não encontrado. Instale com: curl -LsSf https://astral.sh/uv/install.sh | sh")
|
||||
return
|
||||
}
|
||||
let process = Process()
|
||||
process.executableURL = URL(fileURLWithPath: "/usr/bin/env")
|
||||
process.arguments = [uv] + arguments
|
||||
process.currentDirectoryURL = codeDirectory
|
||||
|
||||
let pipe = Pipe()
|
||||
process.standardOutput = pipe
|
||||
process.standardError = pipe
|
||||
|
||||
var buffer = ""
|
||||
let lock = NSLock()
|
||||
pipe.fileHandleForReading.readabilityHandler = { handle in
|
||||
let data = handle.availableData
|
||||
guard !data.isEmpty, let s = String(data: data, encoding: .utf8) else { return }
|
||||
lock.lock()
|
||||
buffer += s
|
||||
let parts = buffer.split(separator: "\n", omittingEmptySubsequences: false)
|
||||
buffer = String(parts.last ?? "")
|
||||
let lines = parts.dropLast()
|
||||
lock.unlock()
|
||||
for line in lines where !line.isEmpty { onLine(String(line)) }
|
||||
}
|
||||
|
||||
process.terminationHandler = { p in
|
||||
pipe.fileHandleForReading.readabilityHandler = nil
|
||||
lock.lock()
|
||||
let last = buffer.trimmingCharacters(in: .whitespacesAndNewlines)
|
||||
buffer = ""
|
||||
lock.unlock()
|
||||
if !last.isEmpty { onLine(last) }
|
||||
completion(Int(p.terminationStatus), p.terminationStatus == 0 ? nil : "uv sync terminou com erro (código \(p.terminationStatus)).")
|
||||
}
|
||||
|
||||
do {
|
||||
try process.run()
|
||||
} catch {
|
||||
completion(1, error.localizedDescription)
|
||||
}
|
||||
}
|
||||
|
||||
// MARK: - Convenience: single JSON result
|
||||
|
||||
/// Runs a command and delivers the first parsed JSON document as the result.
|
||||
|
||||
@@ -0,0 +1,529 @@
|
||||
import SwiftUI
|
||||
|
||||
/// Colors shared by the timeline and the inspector, so a block and its row in
|
||||
/// the list always read as the same thing.
|
||||
enum EmphasisPalette {
|
||||
static func color(_ level: Int) -> Color {
|
||||
switch level {
|
||||
case 1: return Color.blue
|
||||
case 2: return Color.orange
|
||||
case 3: return Color.pink
|
||||
default: return Color.secondary
|
||||
}
|
||||
}
|
||||
|
||||
static func label(_ level: Int) -> String {
|
||||
switch level {
|
||||
case 1: return "Leve"
|
||||
case 2: return "Média"
|
||||
case 3: return "Forte"
|
||||
default: return "Sem"
|
||||
}
|
||||
}
|
||||
|
||||
/// The same 0–3 tiers a phrase's `emphasis` uses, derived from a raw 0–1
|
||||
/// acoustic score — the thresholds `10-revisao-humana.md` documents for
|
||||
/// deriving a phrase's level from `peak_emphasis` when no explicit zoom
|
||||
/// was set, reused here per WORD so a word chip and a phrase row read as
|
||||
/// the same scale.
|
||||
static func levelFromScore(_ score: Double) -> Int {
|
||||
switch score {
|
||||
case ..<0.25: return 0
|
||||
case ..<0.45: return 1
|
||||
case ..<0.65: return 2
|
||||
default: return 3
|
||||
}
|
||||
}
|
||||
|
||||
static func speakerColor(_ speaker: String, among speakers: [String]) -> Color {
|
||||
let palette: [Color] = [.teal, .purple, .green, .indigo, .brown, .cyan]
|
||||
guard let index = speakers.firstIndex(of: speaker) else { return .gray }
|
||||
return palette[index % palette.count]
|
||||
}
|
||||
}
|
||||
|
||||
/// The timeline strip: four stacked tracks over one shared time axis.
|
||||
///
|
||||
/// Phrases are laid out as real views rather than drawn into a Canvas, because
|
||||
/// every one of them is a target — click to select, drag its edge to trim,
|
||||
/// right-click to change emphasis. The dense per-word energy track *is* a
|
||||
/// Canvas: it has thousands of bars and nothing to hit.
|
||||
struct TimelineTracksView: View {
|
||||
@ObservedObject var model: PhraseReviewModel
|
||||
|
||||
private let rulerHeight: CGFloat = 18
|
||||
private let phraseHeight: CGFloat = 46
|
||||
private let energyHeight: CGFloat = 34
|
||||
private let stripHeight: CGFloat = 12
|
||||
private let handleWidth: CGFloat = 8
|
||||
|
||||
private let gutterWidth: CGFloat = 92
|
||||
private let trackSpacing: CGFloat = 4
|
||||
|
||||
private var pps: CGFloat { CGFloat(model.pixelsPerSecond) }
|
||||
/// Width follows the *kept* duration, not the raw take's — the timeline
|
||||
/// draws the cut, so removed stretches take no horizontal space.
|
||||
private var contentWidth: CGFloat { max(320, CGFloat(model.keptDuration) * pps) }
|
||||
|
||||
/// Name, icon and height of each lane, in the order they stack. The gutter
|
||||
/// and the tracks are built from this one list so a label can never drift
|
||||
/// off the lane it names.
|
||||
private var lanes: [(label: String, icon: String, height: CGFloat)] {
|
||||
[
|
||||
("", "", rulerHeight),
|
||||
("Zooms", "plus.magnifyingglass", stripHeight + 6),
|
||||
("Frases", "text.quote", phraseHeight),
|
||||
("Energia", "waveform", energyHeight),
|
||||
("Emoção", "face.smiling", stripHeight),
|
||||
("Locutor", "person.wave.2", stripHeight),
|
||||
("Roteiro", "list.bullet.rectangle", stripHeight),
|
||||
]
|
||||
}
|
||||
|
||||
var body: some View {
|
||||
VStack(spacing: 0) {
|
||||
toolbar
|
||||
Divider()
|
||||
HStack(alignment: .top, spacing: 0) {
|
||||
gutter
|
||||
Divider()
|
||||
timelineScroller
|
||||
}
|
||||
}
|
||||
.background(Color(nsColor: .underPageBackgroundColor))
|
||||
}
|
||||
|
||||
/// Fixed column naming each lane. Without it the stripes are six colours
|
||||
/// with no way to tell which one is emotion and which one is the speaker.
|
||||
private var gutter: some View {
|
||||
VStack(alignment: .leading, spacing: trackSpacing) {
|
||||
ForEach(lanes.indices, id: \.self) { index in
|
||||
let lane = lanes[index]
|
||||
HStack(spacing: 4) {
|
||||
if !lane.icon.isEmpty {
|
||||
Image(systemName: lane.icon).font(.system(size: 9))
|
||||
}
|
||||
Text(lane.label).font(.system(size: 10))
|
||||
Spacer(minLength: 0)
|
||||
}
|
||||
.foregroundStyle(.secondary)
|
||||
.frame(height: lane.height, alignment: .center)
|
||||
}
|
||||
}
|
||||
.padding(.horizontal, 8)
|
||||
.padding(.vertical, 8)
|
||||
.frame(width: gutterWidth, alignment: .leading)
|
||||
}
|
||||
|
||||
private var timelineScroller: some View {
|
||||
ScrollViewReader { proxy in
|
||||
ScrollView([.horizontal]) {
|
||||
ZStack(alignment: .topLeading) {
|
||||
VStack(alignment: .leading, spacing: trackSpacing) {
|
||||
ruler
|
||||
zoomTrack
|
||||
phraseTrack
|
||||
energyTrack
|
||||
emotionTrack
|
||||
speakerTrack
|
||||
scriptTrack
|
||||
}
|
||||
.frame(width: contentWidth, alignment: .leading)
|
||||
rangeOverlay
|
||||
playhead
|
||||
// Anchors the auto-scroll: one invisible marker per
|
||||
// phrase, so selecting a line off-screen brings it in.
|
||||
ForEach(model.phrases) { phrase in
|
||||
Color.clear
|
||||
.frame(width: 1, height: 1)
|
||||
.offset(x: x(phrase.start))
|
||||
.id(phrase.id)
|
||||
}
|
||||
}
|
||||
.padding(.vertical, 8)
|
||||
.contentShape(Rectangle())
|
||||
.gesture(scrubGesture)
|
||||
.contextMenu { timelineMenu }
|
||||
}
|
||||
.onChange(of: model.selection) { _, newValue in
|
||||
guard let newValue else { return }
|
||||
withAnimation(.easeOut(duration: 0.2)) {
|
||||
proxy.scrollTo(newValue, anchor: .center)
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// MARK: - Barra de controles
|
||||
|
||||
private var toolbar: some View {
|
||||
HStack(spacing: 12) {
|
||||
Button {
|
||||
model.togglePlay()
|
||||
} label: {
|
||||
Image(systemName: model.isPlaying ? "pause.fill" : "play.fill")
|
||||
}
|
||||
.buttonStyle(.borderless)
|
||||
.help("Reproduzir (espaço)")
|
||||
.disabled(model.player == nil)
|
||||
|
||||
Text(timecode(model.currentTime))
|
||||
.font(.system(.caption, design: .monospaced))
|
||||
.foregroundStyle(.secondary)
|
||||
|
||||
Button {
|
||||
model.playSelectedPhrase()
|
||||
} label: {
|
||||
Image(systemName: "play.rectangle")
|
||||
}
|
||||
.buttonStyle(.borderless)
|
||||
.help("Tocar só a frase selecionada (⏎)")
|
||||
.disabled(model.player == nil || model.selection == nil)
|
||||
|
||||
Toggle("Pular removidos", isOn: $model.skipRemoved)
|
||||
.toggleStyle(.checkbox)
|
||||
.font(.caption)
|
||||
.help("Durante a reprodução, salta os trechos desativados — mostra como o corte ficou.")
|
||||
|
||||
Button {
|
||||
model.addZoomForRange()
|
||||
} label: {
|
||||
Label("Zoom no trecho", systemImage: "plus.magnifyingglass")
|
||||
}
|
||||
.buttonStyle(.borderless)
|
||||
.font(.caption)
|
||||
.disabled(!model.hasRange)
|
||||
.help("Arraste na timeline para marcar um trecho e crie um zoom nele. A escala vem de Análise de Voz.")
|
||||
|
||||
Spacer()
|
||||
|
||||
legend
|
||||
|
||||
Spacer()
|
||||
|
||||
Image(systemName: "minus.magnifyingglass").foregroundStyle(.secondary)
|
||||
Slider(value: $model.pixelsPerSecond,
|
||||
in: model.minPixelsPerSecond...model.maxPixelsPerSecond)
|
||||
.frame(width: 130)
|
||||
Image(systemName: "plus.magnifyingglass").foregroundStyle(.secondary)
|
||||
}
|
||||
.padding(.horizontal, 12)
|
||||
.padding(.vertical, 8)
|
||||
}
|
||||
|
||||
private var legend: some View {
|
||||
HStack(spacing: 10) {
|
||||
ForEach(0..<4, id: \.self) { level in
|
||||
HStack(spacing: 4) {
|
||||
RoundedRectangle(cornerRadius: 2)
|
||||
.fill(EmphasisPalette.color(level))
|
||||
.frame(width: 10, height: 10)
|
||||
Text(EmphasisPalette.label(level)).font(.caption2)
|
||||
}
|
||||
}
|
||||
}
|
||||
.foregroundStyle(.secondary)
|
||||
}
|
||||
|
||||
// MARK: - Trilhas
|
||||
|
||||
private var ruler: some View {
|
||||
Canvas { context, size in
|
||||
let step = tickStep()
|
||||
var time = 0.0
|
||||
while time <= model.keptDuration {
|
||||
let position = compactX(time)
|
||||
context.stroke(
|
||||
Path { $0.move(to: CGPoint(x: position, y: size.height - 6))
|
||||
$0.addLine(to: CGPoint(x: position, y: size.height)) },
|
||||
with: .color(.secondary.opacity(0.5))
|
||||
)
|
||||
context.draw(
|
||||
Text(timecode(time)).font(.system(size: 9, design: .monospaced))
|
||||
.foregroundColor(.secondary),
|
||||
at: CGPoint(x: position + 18, y: 6)
|
||||
)
|
||||
time += step
|
||||
}
|
||||
}
|
||||
.frame(width: contentWidth, height: rulerHeight)
|
||||
}
|
||||
|
||||
private var phraseTrack: some View {
|
||||
ZStack(alignment: .topLeading) {
|
||||
RoundedRectangle(cornerRadius: 4)
|
||||
.fill(Color.secondary.opacity(0.06))
|
||||
.frame(width: contentWidth, height: phraseHeight)
|
||||
ForEach(model.phrases) { phrase in
|
||||
phraseBlock(phrase)
|
||||
}
|
||||
}
|
||||
.frame(width: contentWidth, height: phraseHeight, alignment: .topLeading)
|
||||
}
|
||||
|
||||
@ViewBuilder
|
||||
private func phraseBlock(_ phrase: ReviewPhrase) -> some View {
|
||||
let isSelected = model.selection == phrase.id
|
||||
let color = EmphasisPalette.color(phrase.emphasis)
|
||||
let fullWidth = max(2, width(from: phrase.start, to: phrase.end))
|
||||
let keptWidth = max(1, width(from: phrase.trimStart, to: phrase.trimEnd))
|
||||
|
||||
ZStack(alignment: .topLeading) {
|
||||
// The whole line, dim — what is there before the edit.
|
||||
RoundedRectangle(cornerRadius: 4)
|
||||
.fill(color.opacity(phrase.active ? 0.15 : 0.10))
|
||||
.frame(width: fullWidth, height: phraseHeight)
|
||||
|
||||
// What survives: the kept span, drawn solid over it.
|
||||
RoundedRectangle(cornerRadius: 4)
|
||||
.fill(color.opacity(phrase.active ? 0.55 : 0.12))
|
||||
.frame(width: keptWidth, height: phraseHeight)
|
||||
.offset(x: width(from: phrase.start, to: phrase.trimStart))
|
||||
|
||||
Text(phrase.text)
|
||||
.font(.system(size: 10))
|
||||
.lineLimit(2)
|
||||
.padding(.horizontal, 4)
|
||||
.frame(width: fullWidth, height: phraseHeight, alignment: .topLeading)
|
||||
.foregroundStyle(phrase.active ? .primary : .secondary)
|
||||
.strikethrough(!phrase.active)
|
||||
|
||||
RoundedRectangle(cornerRadius: 4)
|
||||
.stroke(isSelected ? Color.accentColor : color.opacity(0.4),
|
||||
lineWidth: isSelected ? 2 : 1)
|
||||
.frame(width: fullWidth, height: phraseHeight)
|
||||
|
||||
if isSelected && phrase.active {
|
||||
trimHandle(phrase, edge: .start)
|
||||
trimHandle(phrase, edge: .end)
|
||||
}
|
||||
}
|
||||
.frame(width: fullWidth, height: phraseHeight, alignment: .topLeading)
|
||||
.offset(x: x(phrase.start))
|
||||
.contentShape(Rectangle())
|
||||
.onTapGesture { model.goTo(phraseID: phrase.id) }
|
||||
.contextMenu { phraseMenu(phrase) }
|
||||
.help(phrase.reason.isEmpty ? phrase.text : "\(phrase.text)\n— \(phrase.reason)")
|
||||
}
|
||||
|
||||
private func trimHandle(_ phrase: ReviewPhrase, edge: TrimEdge) -> some View {
|
||||
let offset = edge == .start
|
||||
? width(from: phrase.start, to: phrase.trimStart)
|
||||
: width(from: phrase.start, to: phrase.trimEnd) - handleWidth
|
||||
return RoundedRectangle(cornerRadius: 2)
|
||||
.fill(Color.accentColor)
|
||||
.frame(width: handleWidth, height: phraseHeight)
|
||||
.offset(x: offset)
|
||||
.gesture(
|
||||
DragGesture(minimumDistance: 1)
|
||||
.onChanged { value in
|
||||
let compactOrigin = model.compactTime(phrase.start)
|
||||
let time = model.rawTime(fromCompact: compactOrigin + Double(value.location.x / pps))
|
||||
model.trim(phrase.id, edge: edge, to: time)
|
||||
}
|
||||
)
|
||||
.help(edge == .start ? "Arraste para cortar o começo (pula de palavra em palavra)"
|
||||
: "Arraste para cortar o fim (pula de palavra em palavra)")
|
||||
}
|
||||
|
||||
@ViewBuilder
|
||||
private func phraseMenu(_ phrase: ReviewPhrase) -> some View {
|
||||
Button("Tocar esta frase") {
|
||||
model.goTo(phraseID: phrase.id)
|
||||
model.playSelectedPhrase()
|
||||
}
|
||||
Button(phrase.active ? "Remover do corte" : "Trazer de volta") {
|
||||
model.toggleActive(phrase.id)
|
||||
}
|
||||
Button("Adicionar zoom nesta frase") { model.addZoomForPhrase(phrase.id) }
|
||||
Divider()
|
||||
ForEach(0..<4, id: \.self) { level in
|
||||
Button("Ênfase: \(EmphasisPalette.label(level))") {
|
||||
model.setEmphasis(level, for: phrase.id)
|
||||
}
|
||||
}
|
||||
Divider()
|
||||
Button(phrase.isBackstage ? "Marcar como roteiro" : "Marcar como bastidor") {
|
||||
model.setTrack(phrase.isBackstage ? ReviewPhrase.trackScript : ReviewPhrase.trackBackstage,
|
||||
for: phrase.id)
|
||||
}
|
||||
if phrase.isTrimmed {
|
||||
Divider()
|
||||
Button("Desfazer corte da frase") { model.resetTrim(phrase.id) }
|
||||
}
|
||||
}
|
||||
|
||||
/// Per-word energy/emphasis, straight from the voice timeline — the closest
|
||||
/// thing to a waveform without opening the audio again.
|
||||
private var energyTrack: some View {
|
||||
Canvas { context, size in
|
||||
for phrase in model.phrases {
|
||||
for word in phrase.words {
|
||||
let start = x(word.start)
|
||||
let barWidth = max(1, width(from: word.start, to: word.end) - 1)
|
||||
let height = size.height * CGFloat(max(0.04, word.energy))
|
||||
let rect = CGRect(x: start, y: size.height - height,
|
||||
width: barWidth, height: height)
|
||||
let color = word.emphasis >= 0.65 ? Color.pink
|
||||
: word.emphasis >= 0.45 ? Color.orange
|
||||
: Color.secondary
|
||||
context.fill(Path(rect),
|
||||
with: .color(color.opacity(phrase.active ? 0.6 : 0.2)))
|
||||
}
|
||||
}
|
||||
}
|
||||
.frame(width: contentWidth, height: energyHeight)
|
||||
.background(RoundedRectangle(cornerRadius: 4).fill(Color.secondary.opacity(0.06)))
|
||||
}
|
||||
|
||||
private var speakerTrack: some View {
|
||||
stripTrack { phrase in
|
||||
EmphasisPalette.speakerColor(phrase.speaker, among: model.speakers)
|
||||
}
|
||||
}
|
||||
|
||||
private var scriptTrack: some View {
|
||||
stripTrack { phrase in phrase.isBackstage ? Color.gray : Color.mint }
|
||||
}
|
||||
|
||||
private func stripTrack(_ color: @escaping (ReviewPhrase) -> Color) -> some View {
|
||||
Canvas { context, size in
|
||||
for phrase in model.phrases {
|
||||
let rect = CGRect(x: x(phrase.start), y: 0,
|
||||
width: max(1, width(from: phrase.start, to: phrase.end)),
|
||||
height: size.height)
|
||||
context.fill(Path(roundedRect: rect, cornerRadius: 2),
|
||||
with: .color(color(phrase).opacity(phrase.active ? 0.7 : 0.2)))
|
||||
}
|
||||
}
|
||||
.frame(width: contentWidth, height: stripHeight)
|
||||
}
|
||||
|
||||
private var playhead: some View {
|
||||
Rectangle()
|
||||
.fill(Color.red)
|
||||
.frame(width: 1.5)
|
||||
.offset(x: x(model.currentTime))
|
||||
.allowsHitTesting(false)
|
||||
}
|
||||
|
||||
/// One gesture, two meanings, decided by whether the mouse moved: a click
|
||||
/// parks the playhead, a drag marks in/out. Splitting them across separate
|
||||
/// controls would mean choosing a tool before every action, which is
|
||||
/// exactly the ceremony this screen is meant to avoid.
|
||||
private var scrubGesture: some Gesture {
|
||||
DragGesture(minimumDistance: 0)
|
||||
.onChanged { value in
|
||||
let from = model.rawTime(fromCompact: Double(value.startLocation.x / pps))
|
||||
let to = model.rawTime(fromCompact: Double(value.location.x / pps))
|
||||
if abs(value.translation.width) > 3 {
|
||||
model.setRange(from: from, to: to)
|
||||
model.seek(to: min(from, to))
|
||||
} else {
|
||||
model.clearRange()
|
||||
model.seek(to: to)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
/// The marked in/out, drawn over every track so the span reads against the
|
||||
/// phrases and the energy at once.
|
||||
private var rangeOverlay: some View {
|
||||
Group {
|
||||
if let span = model.rangeSpan {
|
||||
Rectangle()
|
||||
.fill(Color.accentColor.opacity(0.18))
|
||||
.overlay(Rectangle().stroke(Color.accentColor.opacity(0.6), lineWidth: 1))
|
||||
.frame(width: max(1, width(from: span.start, to: span.end)))
|
||||
.offset(x: x(span.start))
|
||||
.allowsHitTesting(false)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
@ViewBuilder
|
||||
private var timelineMenu: some View {
|
||||
if model.hasRange, let span = model.rangeSpan {
|
||||
Button("Adicionar zoom no trecho (\(secondsLabel(span.end - span.start)))") {
|
||||
model.addZoomForRange()
|
||||
}
|
||||
Button("Tocar o trecho") { model.playRange(from: span.start, to: span.end) }
|
||||
Button("Limpar seleção") { model.clearRange() }
|
||||
} else {
|
||||
Text("Arraste na timeline para marcar um trecho")
|
||||
}
|
||||
if let zoom = model.zoom(at: model.currentTime) {
|
||||
Divider()
|
||||
Button("Remover o zoom daqui") { model.removeZoom(zoom.id) }
|
||||
}
|
||||
}
|
||||
|
||||
private func secondsLabel(_ seconds: Double) -> String {
|
||||
String(format: "%.1fs", seconds)
|
||||
}
|
||||
|
||||
/// Punch-ins, on their own lane above the script: they are a second layer
|
||||
/// over the same time, not a property of a phrase.
|
||||
private var zoomTrack: some View {
|
||||
ZStack(alignment: .topLeading) {
|
||||
RoundedRectangle(cornerRadius: 3)
|
||||
.fill(Color.secondary.opacity(0.06))
|
||||
.frame(width: contentWidth, height: stripHeight + 6)
|
||||
ForEach(model.zooms) { zoom in
|
||||
RoundedRectangle(cornerRadius: 3)
|
||||
.fill(Color.yellow.opacity(0.55))
|
||||
.overlay(
|
||||
Image(systemName: "plus.magnifyingglass")
|
||||
.font(.system(size: 8)).foregroundStyle(.black.opacity(0.6))
|
||||
)
|
||||
.frame(width: max(6, width(from: zoom.start, to: zoom.end)),
|
||||
height: stripHeight + 6)
|
||||
.offset(x: x(zoom.start))
|
||||
.help("Zoom marcado — \(secondsLabel(zoom.end - zoom.start)). A escala vem de Análise de Voz.")
|
||||
.contextMenu {
|
||||
Button("Remover este zoom") { model.removeZoom(zoom.id) }
|
||||
}
|
||||
}
|
||||
}
|
||||
.frame(width: contentWidth, height: stripHeight + 6, alignment: .topLeading)
|
||||
}
|
||||
|
||||
/// Delivery emotion per phrase — the fourth signal to read against the text.
|
||||
private var emotionTrack: some View {
|
||||
stripTrack { phrase in
|
||||
switch phrase.emotion {
|
||||
case "excited": return .orange
|
||||
case "tense": return .red
|
||||
case "calm": return .blue
|
||||
case "reflective": return .purple
|
||||
default: return .secondary
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// MARK: - Escala
|
||||
|
||||
/// Pixel position of a raw source-media time, after collapsing whatever
|
||||
/// lies between it and the previous kept phrase.
|
||||
private func x(_ time: Double) -> CGFloat { compactX(model.compactTime(time)) }
|
||||
|
||||
/// Pixel position of a time already in the compacted (edited) timeline —
|
||||
/// used for the ruler and playhead, which think in that space directly.
|
||||
private func compactX(_ compactTime: Double) -> CGFloat { CGFloat(compactTime) * pps }
|
||||
|
||||
private func width(from: Double, to: Double) -> CGFloat {
|
||||
max(0, CGFloat(model.compactTime(to) - model.compactTime(from)) * pps)
|
||||
}
|
||||
|
||||
/// Ruler spacing that keeps labels ~80pt apart at any zoom.
|
||||
private func tickStep() -> Double {
|
||||
let candidates: [Double] = [1, 2, 5, 10, 15, 30, 60, 120, 300, 600]
|
||||
let wanted = 80 / Double(pps)
|
||||
return candidates.first { $0 >= wanted } ?? 600
|
||||
}
|
||||
|
||||
private func timecode(_ seconds: Double) -> String {
|
||||
let total = Int(seconds.rounded(.down))
|
||||
return String(format: "%02d:%02d", total / 60, total % 60)
|
||||
}
|
||||
}
|
||||
@@ -271,7 +271,7 @@ struct TranscriptionView: View {
|
||||
Toggle("Marcar o que foi dito na timeline", isOn: $batchMarkers)
|
||||
|
||||
Divider()
|
||||
Toggle("Exportar legendas SRT", isOn: $batchSubtitles)
|
||||
Toggle("Gerar legenda comum (texto editável no FCP)", isOn: $batchSubtitles)
|
||||
|
||||
Divider()
|
||||
batchOptionRow(
|
||||
@@ -793,7 +793,7 @@ struct TranscriptionView: View {
|
||||
if batchFillers { operations.append("remove_filler_words") }
|
||||
if batchPhrases { operations.append("edit_by_transcript") }
|
||||
if batchMarkers { operations.append("transcript_markers") }
|
||||
if batchSubtitles { operations.append("export_srt") }
|
||||
if batchSubtitles { operations.append("generate_plain_subtitles") }
|
||||
// Runs last, on the timing already cut by any earlier steps (see the
|
||||
// "Abrir no Final Cut Pro" fallback chain and exportSubtitles()'s own
|
||||
// preference for `processedPath` — same reasoning).
|
||||
@@ -834,7 +834,7 @@ struct TranscriptionView: View {
|
||||
}
|
||||
let nextPath = result?["path"] as? String ?? currentPath
|
||||
if operation == "remove_silences" { processedPath = nextPath }
|
||||
if operation == "export_srt" { subtitlePaths = result?["paths"] as? [String] ?? [] }
|
||||
if operation == "generate_plain_subtitles" { subtitlePaths = [nextPath] }
|
||||
if operation == "generate_dynamic_subtitles" { dynamicSubtitlesPath = nextPath }
|
||||
processBatchStep(operations, index: index + 1, currentPath: nextPath, outputFolder: outputFolder)
|
||||
}
|
||||
|
||||
@@ -23,6 +23,7 @@ struct VoiceAnalysisView: View {
|
||||
} else {
|
||||
energySection
|
||||
emphasisSection
|
||||
zoomSection
|
||||
weightsSection
|
||||
emotionSection
|
||||
resetSection
|
||||
@@ -90,6 +91,44 @@ struct VoiceAnalysisView: View {
|
||||
}
|
||||
}
|
||||
|
||||
private var zoomSection: some View {
|
||||
Section {
|
||||
sliderRow(
|
||||
title: "Zoom na ênfase",
|
||||
value: $config.zoomScale,
|
||||
range: 1.0...3.0,
|
||||
readout: "\(Int(config.zoomScale * 100))%",
|
||||
help: "Fator aplicado nos punch-ins de ênfase. 130% equivale a escala 1,30 no Final Cut."
|
||||
)
|
||||
Picker("Movimento", selection: $config.zoomMode) {
|
||||
Text("Zoom in e out").tag("in_out")
|
||||
Text("Só zoom in").tag("in")
|
||||
Text("Só zoom out").tag("out")
|
||||
}
|
||||
.onChange(of: config.zoomMode) { _, _ in save() }
|
||||
sliderRow(
|
||||
title: "Velocidade do zoom in",
|
||||
value: $config.zoomEaseIn,
|
||||
range: 0.05...2.0,
|
||||
readout: String(format: "%.2fs", config.zoomEaseIn),
|
||||
help: "Duração da entrada do zoom. Menor é mais rápido."
|
||||
)
|
||||
sliderRow(
|
||||
title: "Velocidade do zoom out",
|
||||
value: $config.zoomEaseOut,
|
||||
range: 0.01...2.0,
|
||||
readout: String(format: "%.2fs", config.zoomEaseOut),
|
||||
help: "Duração da saída do zoom. Menor é mais seco."
|
||||
)
|
||||
} header: {
|
||||
Text("Zoom de Ênfase")
|
||||
} footer: {
|
||||
Text("Esses valores viram o padrão para ações de zoom que não trouxerem scale/ease/ease_out no JSON da edição por voz.")
|
||||
.font(.caption)
|
||||
.foregroundStyle(.secondary)
|
||||
}
|
||||
}
|
||||
|
||||
// MARK: - Emoção
|
||||
|
||||
private var emotionSection: some View {
|
||||
@@ -129,13 +168,14 @@ struct VoiceAnalysisView: View {
|
||||
title: String,
|
||||
value: Binding<Double>,
|
||||
range: ClosedRange<Double> = 0...1,
|
||||
readout: String? = nil,
|
||||
help: String? = nil
|
||||
) -> some View {
|
||||
VStack(alignment: .leading, spacing: 2) {
|
||||
HStack {
|
||||
Text(title)
|
||||
Spacer()
|
||||
Text(String(format: "%.2f", value.wrappedValue))
|
||||
Text(readout ?? String(format: "%.2f", value.wrappedValue))
|
||||
.monospacedDigit()
|
||||
.foregroundStyle(.secondary)
|
||||
}
|
||||
@@ -188,6 +228,10 @@ struct VoiceAnalysisConfig {
|
||||
var weightDuration: Double
|
||||
var emotionEnabled: Bool
|
||||
var emotionSensitivity: Double
|
||||
var zoomScale: Double
|
||||
var zoomMode: String
|
||||
var zoomEaseIn: Double
|
||||
var zoomEaseOut: Double
|
||||
|
||||
static let defaults = VoiceAnalysisConfig(
|
||||
energyThreshold: 0.5,
|
||||
@@ -198,7 +242,11 @@ struct VoiceAnalysisConfig {
|
||||
weightPause: 0.15,
|
||||
weightDuration: 0.10,
|
||||
emotionEnabled: false,
|
||||
emotionSensitivity: 0.5
|
||||
emotionSensitivity: 0.5,
|
||||
zoomScale: 1.30,
|
||||
zoomMode: "in_out",
|
||||
zoomEaseIn: 0.25,
|
||||
zoomEaseOut: 0.04
|
||||
)
|
||||
|
||||
init(
|
||||
@@ -210,7 +258,11 @@ struct VoiceAnalysisConfig {
|
||||
weightPause: Double,
|
||||
weightDuration: Double,
|
||||
emotionEnabled: Bool,
|
||||
emotionSensitivity: Double
|
||||
emotionSensitivity: Double,
|
||||
zoomScale: Double,
|
||||
zoomMode: String,
|
||||
zoomEaseIn: Double,
|
||||
zoomEaseOut: Double
|
||||
) {
|
||||
self.energyThreshold = energyThreshold
|
||||
self.emphasisThreshold = emphasisThreshold
|
||||
@@ -221,6 +273,10 @@ struct VoiceAnalysisConfig {
|
||||
self.weightDuration = weightDuration
|
||||
self.emotionEnabled = emotionEnabled
|
||||
self.emotionSensitivity = emotionSensitivity
|
||||
self.zoomScale = zoomScale
|
||||
self.zoomMode = zoomMode
|
||||
self.zoomEaseIn = zoomEaseIn
|
||||
self.zoomEaseOut = zoomEaseOut
|
||||
}
|
||||
|
||||
/// Lê a resposta do bridge, caindo no padrão para qualquer campo ausente.
|
||||
@@ -236,7 +292,11 @@ struct VoiceAnalysisConfig {
|
||||
weightPause: weights["pause_before"] as? Double ?? defaults.weightPause,
|
||||
weightDuration: weights["duration"] as? Double ?? defaults.weightDuration,
|
||||
emotionEnabled: json["emotion_enabled"] as? Bool ?? defaults.emotionEnabled,
|
||||
emotionSensitivity: json["emotion_sensitivity"] as? Double ?? defaults.emotionSensitivity
|
||||
emotionSensitivity: json["emotion_sensitivity"] as? Double ?? defaults.emotionSensitivity,
|
||||
zoomScale: json["zoom_scale"] as? Double ?? defaults.zoomScale,
|
||||
zoomMode: json["zoom_mode"] as? String ?? defaults.zoomMode,
|
||||
zoomEaseIn: json["zoom_ease_in"] as? Double ?? defaults.zoomEaseIn,
|
||||
zoomEaseOut: json["zoom_ease_out"] as? Double ?? defaults.zoomEaseOut
|
||||
)
|
||||
}
|
||||
|
||||
@@ -253,6 +313,10 @@ struct VoiceAnalysisConfig {
|
||||
],
|
||||
"emotion_enabled": emotionEnabled,
|
||||
"emotion_sensitivity": emotionSensitivity,
|
||||
"zoom_scale": zoomScale,
|
||||
"zoom_mode": zoomMode,
|
||||
"zoom_ease_in": zoomEaseIn,
|
||||
"zoom_ease_out": zoomEaseOut,
|
||||
]
|
||||
}
|
||||
}
|
||||
|
||||
@@ -0,0 +1,967 @@
|
||||
import SwiftUI
|
||||
import AppKit
|
||||
|
||||
/// Guia passo a passo do fluxo completo: projeto → transcrição → análise de
|
||||
/// voz → copiar para o chat e trazer as decisões → revisar as ênfases →
|
||||
/// processamento final. Existe para que o usuário não precise entender a ordem
|
||||
/// certa de botões espalhados em várias abas — cada etapa só libera a próxima
|
||||
/// quando o passo anterior terminou, e a "ponte" com o chat (que hoje exigia
|
||||
/// sair do app e escolher um arquivo na mão) vira copiar/colar assistido
|
||||
/// dentro da própria tela.
|
||||
enum WizardStep: Int, CaseIterable, Identifiable {
|
||||
case projeto, transcricao, analise, exportarChat, revisar, finalizar
|
||||
var id: Int { rawValue }
|
||||
|
||||
var titulo: String {
|
||||
switch self {
|
||||
case .projeto: return "Projeto"
|
||||
case .transcricao: return "Transcrever"
|
||||
case .analise: return "Analisar voz"
|
||||
case .exportarChat: return "Decisões da IA"
|
||||
case .revisar: return "Revisar ênfases"
|
||||
case .finalizar: return "Processar"
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
struct WizardView: View {
|
||||
@State private var step: WizardStep = .projeto
|
||||
|
||||
// Passo 1 — projeto
|
||||
@State private var outputFolder: String?
|
||||
@State private var projectPath: String?
|
||||
@State private var catalog: Catalog?
|
||||
|
||||
// Passo 2 — transcrição
|
||||
@State private var isTranscribing = false
|
||||
@State private var transcribeProgress: Double = 0
|
||||
@State private var transcribeStage = ""
|
||||
@State private var transcribeResults: [TranscriptResult] = []
|
||||
|
||||
// Passo 3 — análise de voz
|
||||
@State private var isAnalyzing = false
|
||||
@State private var voiceTimelinePath: String?
|
||||
@State private var voiceAnalysisMessage = ""
|
||||
@State private var acousticsAvailable: Bool?
|
||||
@State private var showVoiceTimelineReuseAlert = false
|
||||
@State private var existingVoiceTimelinePath: String?
|
||||
|
||||
// Passo 4 — enviar ao chat e trazer as decisões de volta
|
||||
@State private var copiedFeedback = ""
|
||||
@State private var decisionsText = ""
|
||||
@State private var isApplyingDecisions = false
|
||||
@State private var appliedPath: String?
|
||||
@State private var skippedVoiceEdit = false
|
||||
|
||||
// Passo 4 (alternativa) — gerar o roteiro direto por IA local (Ollama/Gemma 3)
|
||||
@State private var isGeneratingScript = false
|
||||
@State private var generateScriptModel = "gemma3:12b"
|
||||
@State private var generateScriptFeedback = ""
|
||||
@State private var ollamaModels: [String] = []
|
||||
|
||||
// Passo 5 — revisar ênfases
|
||||
@StateObject private var reviewModel = PhraseReviewModel()
|
||||
@State private var reviewLoadedFor: String?
|
||||
@State private var reviewLoadedForDecisions: String?
|
||||
@State private var phraseReviewPath: String?
|
||||
@State private var phraseActionsPath: String?
|
||||
|
||||
// Passo 6 — processamento final
|
||||
@State private var finalSilences = true
|
||||
@State private var finalFillers = false
|
||||
@State private var finalSubtitles = true
|
||||
@State private var finalDynamicSubtitles = false
|
||||
@State private var isFinalizing = false
|
||||
@State private var finalStatus = ""
|
||||
@State private var finalPath: String?
|
||||
|
||||
@State private var errorMessage: String?
|
||||
|
||||
var body: some View {
|
||||
VStack(spacing: 0) {
|
||||
stepperHeader
|
||||
.padding(.horizontal, 24)
|
||||
.padding(.top, 20)
|
||||
.padding(.bottom, 16)
|
||||
|
||||
Divider()
|
||||
|
||||
// A revisão é uma sala de edição, não um formulário: ela precisa da
|
||||
// largura toda e rola por conta própria (timeline horizontal, lista
|
||||
// vertical). As demais etapas continuam na coluna estreita, que é o
|
||||
// que mantém um passo a passo legível.
|
||||
if step == .revisar {
|
||||
revisarStep
|
||||
} else {
|
||||
ScrollView {
|
||||
VStack(alignment: .leading, spacing: 18) {
|
||||
if let errorMessage, !errorMessage.isEmpty {
|
||||
Label(errorMessage, systemImage: "exclamationmark.triangle.fill")
|
||||
.foregroundStyle(.red)
|
||||
.padding(.top, 4)
|
||||
}
|
||||
content
|
||||
}
|
||||
.padding(24)
|
||||
.frame(maxWidth: 640, alignment: .leading)
|
||||
.frame(maxWidth: .infinity)
|
||||
}
|
||||
}
|
||||
|
||||
Divider()
|
||||
navFooter
|
||||
.padding(.horizontal, 24)
|
||||
.padding(.vertical, 16)
|
||||
}
|
||||
.task {
|
||||
loadProjectConfig()
|
||||
await loadCatalog()
|
||||
}
|
||||
.alert("Análise de voz já existe", isPresented: $showVoiceTimelineReuseAlert) {
|
||||
Button("Usar existente") {
|
||||
if let existingVoiceTimelinePath {
|
||||
voiceTimelinePath = existingVoiceTimelinePath
|
||||
voiceAnalysisMessage = "Reaproveitando análise existente: \(existingVoiceTimelinePath)"
|
||||
}
|
||||
}
|
||||
Button("Reprocessar") {
|
||||
analyzeVoice(forceReprocess: true)
|
||||
}
|
||||
Button("Cancelar", role: .cancel) {}
|
||||
} message: {
|
||||
Text("Já existe um arquivo voice_timeline para este projeto. Quer manter o processamento anterior para ganhar tempo?")
|
||||
}
|
||||
}
|
||||
|
||||
// MARK: - Cabeçalho com os passos
|
||||
|
||||
private var stepperHeader: some View {
|
||||
HStack(spacing: 6) {
|
||||
ForEach(WizardStep.allCases) { s in
|
||||
HStack(spacing: 6) {
|
||||
ZStack {
|
||||
Circle()
|
||||
.fill(colorFor(s))
|
||||
.frame(width: 24, height: 24)
|
||||
if s.rawValue < step.rawValue {
|
||||
Image(systemName: "checkmark")
|
||||
.font(.caption2.weight(.bold))
|
||||
.foregroundStyle(.white)
|
||||
} else {
|
||||
Text("\(s.rawValue + 1)")
|
||||
.font(.caption2.weight(.bold))
|
||||
.foregroundStyle(s == step ? .white : .secondary)
|
||||
}
|
||||
}
|
||||
Text(s.titulo)
|
||||
.font(.caption)
|
||||
.foregroundStyle(s == step ? .primary : .secondary)
|
||||
.fontWeight(s == step ? .semibold : .regular)
|
||||
}
|
||||
if s != WizardStep.allCases.last {
|
||||
Rectangle()
|
||||
.fill(s.rawValue < step.rawValue ? Color.accentColor : Color.secondary.opacity(0.25))
|
||||
.frame(height: 2)
|
||||
.frame(maxWidth: .infinity)
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
private func colorFor(_ s: WizardStep) -> Color {
|
||||
if s.rawValue < step.rawValue { return .accentColor }
|
||||
if s == step { return .accentColor }
|
||||
return Color.secondary.opacity(0.25)
|
||||
}
|
||||
|
||||
// MARK: - Conteúdo por etapa
|
||||
|
||||
@ViewBuilder
|
||||
private var content: some View {
|
||||
switch step {
|
||||
case .projeto: projetoStep
|
||||
case .transcricao: transcricaoStep
|
||||
case .analise: analiseStep
|
||||
case .exportarChat: exportarChatStep
|
||||
case .revisar: revisarStep
|
||||
case .finalizar: finalizarStep
|
||||
}
|
||||
}
|
||||
|
||||
private var projetoStep: some View {
|
||||
VStack(alignment: .leading, spacing: 16) {
|
||||
Text("1. Escolha o projeto").font(.title3.weight(.semibold))
|
||||
Text("A pasta é onde tudo o que for gerado nesse fluxo fica salvo. O arquivo é o .fcpxml exportado do Final Cut Pro.")
|
||||
.font(.callout).foregroundStyle(.secondary)
|
||||
|
||||
fieldRow(icon: "folder", label: outputFolder ?? "Nenhuma pasta selecionada", isSet: outputFolder != nil) {
|
||||
pickOutputFolder()
|
||||
}
|
||||
fieldRow(icon: "doc.text", label: projectPath.map { URL(fileURLWithPath: $0).lastPathComponent } ?? "Nenhum arquivo selecionado", isSet: projectPath != nil) {
|
||||
pickProjectFile()
|
||||
}
|
||||
|
||||
if looksLikeGeneratedFile(projectPath) {
|
||||
Label("Esse arquivo parece já ter sido processado por este fluxo (o nome tem um sufixo como \"_voice_edit\" ou \"_silence_removed\"). Rodar o wizard de novo em cima dele reaplica os cortes por cima de cortes já feitos. Selecione o .fcpxml original do Final Cut, a menos que a intenção seja mesmo reprocessar.",
|
||||
systemImage: "exclamationmark.triangle.fill")
|
||||
.font(.caption).foregroundStyle(.orange)
|
||||
}
|
||||
|
||||
if (catalog?.installedCount ?? 0) == 0 {
|
||||
Label("Nenhum modelo de transcrição instalado. Baixe um na aba \"Modelos\" antes de continuar.",
|
||||
systemImage: "exclamationmark.triangle.fill")
|
||||
.font(.caption).foregroundStyle(.orange)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
private var transcricaoStep: some View {
|
||||
VStack(alignment: .leading, spacing: 16) {
|
||||
Text("2. Transcreva o áudio").font(.title3.weight(.semibold))
|
||||
Text("Roda localmente com o modelo escolhido na aba Modelos. Vira a base de tudo que vem depois — o corte por voz, as legendas, os marcadores.")
|
||||
.font(.callout).foregroundStyle(.secondary)
|
||||
|
||||
Button {
|
||||
startTranscription()
|
||||
} label: {
|
||||
if isTranscribing {
|
||||
HStack { ProgressView().controlSize(.small); Text(transcribeStage.isEmpty ? "Transcrevendo…" : transcribeStage) }
|
||||
.frame(maxWidth: .infinity)
|
||||
} else {
|
||||
Label(transcribeResults.isEmpty ? "Transcrever" : "Transcrever novamente", systemImage: "waveform")
|
||||
.frame(maxWidth: .infinity)
|
||||
}
|
||||
}
|
||||
.buttonStyle(.borderedProminent)
|
||||
.controlSize(.large)
|
||||
.disabled(isTranscribing || projectPath == nil || outputFolder == nil)
|
||||
|
||||
if isTranscribing {
|
||||
VStack(alignment: .leading, spacing: 6) {
|
||||
ProgressView(value: transcribeProgress)
|
||||
Text("\(Int(transcribeProgress * 100))%").font(.caption).foregroundStyle(.secondary).monospacedDigit()
|
||||
}
|
||||
}
|
||||
|
||||
if !transcribeResults.isEmpty {
|
||||
ForEach(transcribeResults, id: \.media) { r in
|
||||
VStack(alignment: .leading, spacing: 4) {
|
||||
HStack {
|
||||
Image(systemName: "checkmark.circle.fill").foregroundStyle(.green)
|
||||
Text(r.media).font(.body.weight(.medium))
|
||||
Spacer()
|
||||
Text("\(r.language) · \(r.words) palavras").font(.caption).foregroundStyle(.secondary)
|
||||
}
|
||||
Text(r.preview).font(.caption).foregroundStyle(.secondary).lineLimit(2)
|
||||
}
|
||||
.padding(12)
|
||||
.background(RoundedRectangle(cornerRadius: 8).fill(Color.secondary.opacity(0.06)))
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
private var analiseStep: some View {
|
||||
VStack(alignment: .leading, spacing: 16) {
|
||||
Text("3. Analise a voz").font(.title3.weight(.semibold))
|
||||
Text("Gera o JSON com transcrição, locutor e intensidade (pitch/energia/ritmo) por palavra — é esse arquivo que o chat lê para decidir o que cortar. Não corta nada sozinho.")
|
||||
.font(.callout).foregroundStyle(.secondary)
|
||||
|
||||
Button {
|
||||
analyzeVoice()
|
||||
} label: {
|
||||
if isAnalyzing {
|
||||
HStack { ProgressView().controlSize(.small); Text("Analisando…") }.frame(maxWidth: .infinity)
|
||||
} else {
|
||||
Label(voiceTimelinePath == nil ? "Analisar voz" : "Analisar novamente", systemImage: "waveform.badge.magnifyingglass")
|
||||
.frame(maxWidth: .infinity)
|
||||
}
|
||||
}
|
||||
.buttonStyle(.borderedProminent)
|
||||
.controlSize(.large)
|
||||
.disabled(isAnalyzing || projectPath == nil || outputFolder == nil)
|
||||
|
||||
if let voiceTimelinePath {
|
||||
VStack(alignment: .leading, spacing: 6) {
|
||||
Label("Análise pronta", systemImage: "checkmark.circle.fill").foregroundStyle(.green)
|
||||
Text(voiceTimelinePath).font(.caption).foregroundStyle(.secondary).lineLimit(1).truncationMode(.middle)
|
||||
}
|
||||
.padding(12)
|
||||
.background(RoundedRectangle(cornerRadius: 8).fill(Color.secondary.opacity(0.06)))
|
||||
|
||||
if acousticsAvailable == false {
|
||||
VStack(alignment: .leading, spacing: 4) {
|
||||
Label("Sem análise acústica real", systemImage: "exclamationmark.triangle.fill")
|
||||
.font(.caption.weight(.semibold)).foregroundStyle(.orange)
|
||||
Text("Falta o componente \"librosa\" — os cortes ainda são decididos pelo texto, mas o chat não vai propor zoom com confiança. Instale em Avançado → Modelos → \"Análise Acústica\", e refaça esta etapa depois.")
|
||||
.font(.caption).foregroundStyle(.secondary)
|
||||
}
|
||||
.padding(12)
|
||||
.background(RoundedRectangle(cornerRadius: 8).fill(Color.orange.opacity(0.08)))
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
private var exportarChatStep: some View {
|
||||
VStack(alignment: .leading, spacing: 16) {
|
||||
Text("4. Envie para o chat decidir os cortes").font(.title3.weight(.semibold))
|
||||
Text("Esta é a única etapa manual que sobra: o julgamento de qual tomada usar, onde dar zoom e o que escrever na tela é feito pela IA numa conversa, não por um botão. Copie abaixo, cole numa sessão do Claude e peça pra rodar a skill \"editar-por-voz\".")
|
||||
.font(.callout).foregroundStyle(.secondary)
|
||||
|
||||
if let voiceTimelinePath {
|
||||
// Alternativa automática: em vez de copiar/colar no chat, manda a
|
||||
// própria voice timeline (o arquivo inteiro) junto com o brief para
|
||||
// o modelo local (Ollama/Gemma 3) decidir a edição de uma vez —
|
||||
// cortes, zooms e textos numa única chamada, sem sair do app.
|
||||
VStack(alignment: .leading, spacing: 8) {
|
||||
Text("OU gere o roteiro por IA local (Ollama/Gemma 3)").font(.callout.weight(.semibold))
|
||||
Text("O app envia a voice timeline completa (o arquivo) acompanhada do pedido para o modelo local decidir os cortes, zooms e textos de uma vez. Nada de copiar e colar.")
|
||||
.font(.caption).foregroundStyle(.secondary)
|
||||
HStack {
|
||||
if ollamaModels.isEmpty {
|
||||
TextField("Modelo (ex.: gemma3:12b, llama3)", text: $generateScriptModel)
|
||||
.textFieldStyle(.roundedBorder)
|
||||
.frame(maxWidth: 260)
|
||||
} else {
|
||||
Picker("Modelo", selection: $generateScriptModel) {
|
||||
ForEach(ollamaModels, id: \.self) { m in
|
||||
Text(m).tag(m)
|
||||
}
|
||||
}
|
||||
.pickerStyle(.menu)
|
||||
.frame(maxWidth: 260)
|
||||
TextField("Ou outro", text: $generateScriptModel)
|
||||
.textFieldStyle(.roundedBorder)
|
||||
.frame(maxWidth: 120)
|
||||
}
|
||||
Button {
|
||||
generateScript(voiceTimelinePath: voiceTimelinePath)
|
||||
} label: {
|
||||
if isGeneratingScript {
|
||||
HStack { ProgressView().controlSize(.small); Text("Gerando…") }
|
||||
} else {
|
||||
Label("Gerar roteiro por IA local", systemImage: "sparkles")
|
||||
}
|
||||
}
|
||||
.buttonStyle(.borderedProminent)
|
||||
.disabled(isGeneratingScript || voiceTimelinePath.isEmpty)
|
||||
}
|
||||
if !generateScriptFeedback.isEmpty {
|
||||
Label(generateScriptFeedback, systemImage: "checkmark.circle.fill")
|
||||
.font(.caption).foregroundStyle(.green)
|
||||
}
|
||||
}
|
||||
.padding(12)
|
||||
.background(RoundedRectangle(cornerRadius: 8).fill(Color.green.opacity(0.07)))
|
||||
.onAppear { fetchOllamaModels() }
|
||||
|
||||
Divider().padding(.vertical, 4)
|
||||
|
||||
Button {
|
||||
copyForChat(path: voiceTimelinePath)
|
||||
} label: {
|
||||
Label("Copiar para colar no chat", systemImage: "doc.on.clipboard")
|
||||
.frame(maxWidth: .infinity)
|
||||
}
|
||||
.buttonStyle(.borderedProminent)
|
||||
.controlSize(.large)
|
||||
|
||||
if !copiedFeedback.isEmpty {
|
||||
Label(copiedFeedback, systemImage: "checkmark.circle.fill")
|
||||
.font(.caption).foregroundStyle(.green)
|
||||
}
|
||||
|
||||
VStack(alignment: .leading, spacing: 8) {
|
||||
Text("O que é copiado").font(.caption.weight(.semibold)).foregroundStyle(.secondary)
|
||||
Text("Um pedido pronto + o conteúdo de \(URL(fileURLWithPath: voiceTimelinePath).lastPathComponent), já formatado. É só colar (⌘V) numa conversa com o Claude.")
|
||||
.font(.caption).foregroundStyle(.secondary)
|
||||
}
|
||||
.padding(12)
|
||||
.background(RoundedRectangle(cornerRadius: 8).fill(Color.secondary.opacity(0.06)))
|
||||
|
||||
Divider().padding(.vertical, 4)
|
||||
|
||||
Text("Cole aqui o que o chat devolveu").font(.callout.weight(.semibold))
|
||||
Text("Na próxima etapa essas decisões aparecem já marcadas na timeline, frase por frase, para você lapidar.")
|
||||
.font(.caption).foregroundStyle(.secondary)
|
||||
|
||||
HStack {
|
||||
Button {
|
||||
if let s = NSPasteboard.general.string(forType: .string) {
|
||||
decisionsText = s
|
||||
}
|
||||
} label: {
|
||||
Label("Colar da área de transferência", systemImage: "list.clipboard")
|
||||
}
|
||||
Spacer()
|
||||
if !decisionsText.isEmpty {
|
||||
Label(jsonIsValid ? "JSON válido" : "JSON inválido",
|
||||
systemImage: jsonIsValid ? "checkmark.circle.fill" : "xmark.circle.fill")
|
||||
.font(.caption)
|
||||
.foregroundStyle(jsonIsValid ? .green : .red)
|
||||
}
|
||||
}
|
||||
|
||||
TextEditor(text: $decisionsText)
|
||||
.font(.system(.caption, design: .monospaced))
|
||||
.frame(minHeight: 140)
|
||||
.padding(8)
|
||||
.background(RoundedRectangle(cornerRadius: 8).fill(Color.secondary.opacity(0.06)))
|
||||
.overlay(RoundedRectangle(cornerRadius: 8).stroke(Color.secondary.opacity(0.2)))
|
||||
|
||||
Button {
|
||||
applyDecisions()
|
||||
} label: {
|
||||
if isApplyingDecisions {
|
||||
HStack { ProgressView().controlSize(.small); Text("Aplicando…") }
|
||||
.frame(maxWidth: .infinity)
|
||||
} else {
|
||||
Label("Aplicar decisões", systemImage: "checkmark.seal")
|
||||
.frame(maxWidth: .infinity)
|
||||
}
|
||||
}
|
||||
.buttonStyle(.borderedProminent)
|
||||
.controlSize(.large)
|
||||
.disabled(isApplyingDecisions || !jsonIsValid)
|
||||
|
||||
if let appliedPath {
|
||||
Label("Decisões aplicadas — \(URL(fileURLWithPath: appliedPath).lastPathComponent)",
|
||||
systemImage: "checkmark.circle.fill")
|
||||
.font(.caption).foregroundStyle(.green)
|
||||
}
|
||||
|
||||
Divider()
|
||||
Button("Pular esta etapa (revisar as ênfases direto, sem passar pela IA)") {
|
||||
skippedVoiceEdit = true
|
||||
appliedPath = nil
|
||||
decisionsText = ""
|
||||
}
|
||||
.buttonStyle(.plain)
|
||||
.font(.caption)
|
||||
.foregroundStyle(.secondary)
|
||||
} else {
|
||||
Label("Volte ao passo anterior e rode a análise de voz primeiro.", systemImage: "exclamationmark.triangle.fill")
|
||||
.font(.caption).foregroundStyle(.orange)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
/// Etapa 5 — a sala de edição. Diferente das outras, não é um formulário
|
||||
/// dentro da coluna do assistente: ocupa a janela toda e se carrega sozinha
|
||||
/// na primeira vez que aparece para aquela análise de voz.
|
||||
private var revisarStep: some View {
|
||||
Group {
|
||||
if voiceTimelinePath != nil {
|
||||
PhraseReviewView(model: reviewModel)
|
||||
} else {
|
||||
VStack(spacing: 8) {
|
||||
Label("Volte ao passo 3 e rode a análise de voz primeiro.",
|
||||
systemImage: "exclamationmark.triangle.fill")
|
||||
.foregroundStyle(.orange)
|
||||
}
|
||||
.frame(maxWidth: .infinity, maxHeight: .infinity)
|
||||
}
|
||||
}
|
||||
.onAppear { loadReviewIfNeeded() }
|
||||
}
|
||||
|
||||
/// Processing and its result live on the SAME slide: the moment the last
|
||||
/// operation finishes (`finalPath` gets set), the open/reveal buttons
|
||||
/// appear right below the "Processar" button instead of gating behind a
|
||||
/// separate "Concluído" step the user has to click into — there was
|
||||
/// nothing on that slide worth a click of its own.
|
||||
private var finalizarStep: some View {
|
||||
VStack(alignment: .leading, spacing: 16) {
|
||||
Text("6. Finalize o corte").font(.title3.weight(.semibold))
|
||||
Text("Últimos passos automáticos, sem decisão envolvida — rodam com os parâmetros já configurados na aba \"Análise de Voz\" / \"Legendas Dinâmicas\".")
|
||||
.font(.callout).foregroundStyle(.secondary)
|
||||
|
||||
Toggle("Remover silêncios do áudio", isOn: $finalSilences)
|
||||
Toggle("Remover palavras de preenchimento", isOn: $finalFillers)
|
||||
Toggle("Gerar legenda comum (texto editável no FCP)", isOn: $finalSubtitles)
|
||||
Toggle("Gerar legendas dinâmicas (estilo configurado na aba própria)", isOn: $finalDynamicSubtitles)
|
||||
|
||||
Button {
|
||||
finalizeProcessing()
|
||||
} label: {
|
||||
if isFinalizing {
|
||||
HStack { ProgressView().controlSize(.small); Text(finalStatus.isEmpty ? "Processando…" : finalStatus) }
|
||||
.frame(maxWidth: .infinity)
|
||||
} else {
|
||||
Label("Processar", systemImage: "play.fill").frame(maxWidth: .infinity)
|
||||
}
|
||||
}
|
||||
.buttonStyle(.borderedProminent)
|
||||
.controlSize(.large)
|
||||
.disabled(isFinalizing || (!finalSilences && !finalFillers && !finalSubtitles && !finalDynamicSubtitles))
|
||||
|
||||
if let finalPath, !isFinalizing {
|
||||
Divider().padding(.vertical, 4)
|
||||
Label("Concluído", systemImage: "checkmark.seal.fill")
|
||||
.font(.callout.weight(.semibold))
|
||||
.foregroundStyle(.green)
|
||||
Text(finalPath).font(.caption).foregroundStyle(.secondary).lineLimit(1).truncationMode(.middle)
|
||||
HStack {
|
||||
Button("Abrir no Final Cut Pro") { NSWorkspace.shared.open(URL(fileURLWithPath: finalPath)) }
|
||||
.buttonStyle(.borderedProminent)
|
||||
Button("Mostrar no Finder") {
|
||||
NSWorkspace.shared.activateFileViewerSelecting([URL(fileURLWithPath: finalPath)])
|
||||
}
|
||||
Spacer()
|
||||
Button("Começar outro projeto") { resetWizard() }
|
||||
}
|
||||
} else if !finalStatus.isEmpty && !isFinalizing {
|
||||
Text(finalStatus).font(.caption).foregroundStyle(.secondary)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// MARK: - Navegação
|
||||
|
||||
/// `.finalizar` is the last step now — once it has a `finalPath`, the
|
||||
/// slide's own "Começar outro projeto" button is the way forward, so the
|
||||
/// footer's "Continuar" would be a second, redundant path to nowhere.
|
||||
private var navFooter: some View {
|
||||
HStack {
|
||||
if step != .projeto {
|
||||
Button("Voltar") { goBack() }
|
||||
}
|
||||
Spacer()
|
||||
if step != .finalizar || finalPath == nil {
|
||||
Button(step == .finalizar ? "Concluir" : "Continuar") { goNext() }
|
||||
.buttonStyle(.borderedProminent)
|
||||
.disabled(!canAdvance)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
private var canAdvance: Bool {
|
||||
switch step {
|
||||
case .projeto: return outputFolder != nil && projectPath != nil
|
||||
case .transcricao: return !transcribeResults.isEmpty
|
||||
case .analise: return voiceTimelinePath != nil
|
||||
case .exportarChat: return appliedPath != nil || skippedVoiceEdit
|
||||
// Revisar é opcional: a sugestão da IA já é utilizável como veio, então
|
||||
// o botão nunca trava aqui — o passo existe para lapidar, não para
|
||||
// exigir mais uma confirmação.
|
||||
case .revisar: return true
|
||||
case .finalizar: return finalPath != nil && !isFinalizing
|
||||
}
|
||||
}
|
||||
|
||||
private func goNext() {
|
||||
guard let next = WizardStep(rawValue: step.rawValue + 1) else { return }
|
||||
// Sair da revisão grava o que foi decidido e as ações derivadas dela
|
||||
// (`_phrase_actions.json`) ao lado da análise de voz — é esse arquivo
|
||||
// que `finalizeProcessing` reaplica na etapa 6, para que desativar uma
|
||||
// frase aqui realmente a remova do vídeo final, e não só do registro.
|
||||
if step == .revisar {
|
||||
reviewModel.save { reviewPath, actionsPath in
|
||||
phraseReviewPath = reviewPath
|
||||
phraseActionsPath = actionsPath
|
||||
}
|
||||
}
|
||||
step = next
|
||||
}
|
||||
|
||||
private func goBack() {
|
||||
guard let prev = WizardStep(rawValue: step.rawValue - 1) else { return }
|
||||
step = prev
|
||||
}
|
||||
|
||||
private func resetWizard() {
|
||||
step = .projeto
|
||||
transcribeResults = []
|
||||
voiceTimelinePath = nil
|
||||
voiceAnalysisMessage = ""
|
||||
decisionsText = ""
|
||||
appliedPath = nil
|
||||
skippedVoiceEdit = false
|
||||
reviewLoadedFor = nil
|
||||
reviewLoadedForDecisions = nil
|
||||
phraseReviewPath = nil
|
||||
phraseActionsPath = nil
|
||||
finalStatus = ""
|
||||
finalPath = nil
|
||||
errorMessage = nil
|
||||
}
|
||||
|
||||
// MARK: - Componentes auxiliares
|
||||
|
||||
@ViewBuilder
|
||||
private func fieldRow(icon: String, label: String, isSet: Bool, action: @escaping () -> Void) -> some View {
|
||||
HStack {
|
||||
Image(systemName: icon).foregroundStyle(isSet ? .primary : .secondary)
|
||||
Text(label).lineLimit(1).truncationMode(.middle).foregroundStyle(isSet ? .primary : .secondary)
|
||||
Spacer()
|
||||
Button("Escolher…", action: action)
|
||||
}
|
||||
.padding(12)
|
||||
.background(RoundedRectangle(cornerRadius: 8).fill(Color.secondary.opacity(0.06)))
|
||||
}
|
||||
|
||||
/// Todo output do fluxo carrega um destes sufixos no nome (ver
|
||||
/// `_derived_output` / suffixes usados por `apply_voice_actions`,
|
||||
/// `remove_silences`, `generate_dynamic_subtitles` em
|
||||
/// `admin/models_api.py`). Selecionar um deles como "o projeto" no passo
|
||||
/// 1 é o erro que gerou arquivos como `_voice_edit_voice_edit_...`: os
|
||||
/// cortes de voz assumem timestamps da mídia ORIGINAL, então reaplicá-los
|
||||
/// sobre um arquivo já cortado desloca tudo silenciosamente.
|
||||
private static let generatedSuffixes = [
|
||||
"_voice_edit", "_silence_removed", "_dynamic_subtitles",
|
||||
"_transcript_edit", "_fillers_removed", "_markers",
|
||||
]
|
||||
|
||||
private func looksLikeGeneratedFile(_ path: String?) -> Bool {
|
||||
guard let path else { return false }
|
||||
let stem = URL(fileURLWithPath: path).deletingPathExtension().lastPathComponent
|
||||
return Self.generatedSuffixes.contains { stem.contains($0) }
|
||||
}
|
||||
|
||||
private var jsonIsValid: Bool {
|
||||
guard let data = decisionsText.data(using: .utf8), !decisionsText.trimmingCharacters(in: .whitespacesAndNewlines).isEmpty else { return false }
|
||||
return (try? JSONSerialization.jsonObject(with: data)) != nil
|
||||
}
|
||||
|
||||
// MARK: - Ações — Python bridge
|
||||
|
||||
private func loadProjectConfig() {
|
||||
PythonBridge.call(command: "project_config") { result, _ in
|
||||
DispatchQueue.main.async {
|
||||
guard let result, result["ok"] as? Bool == true else { return }
|
||||
if let folder = result["folder"] as? String, !folder.isEmpty { outputFolder = folder }
|
||||
if let file = result["file"] as? String, !file.isEmpty { projectPath = file }
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
private func loadCatalog() async {
|
||||
PythonBridge.call(command: "catalog") { result, _ in
|
||||
DispatchQueue.main.async {
|
||||
if let result { catalog = Catalog(json: result) }
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
private func pickOutputFolder() {
|
||||
let panel = NSOpenPanel()
|
||||
panel.canChooseFiles = false
|
||||
panel.canChooseDirectories = true
|
||||
panel.allowsMultipleSelection = false
|
||||
panel.prompt = "Usar esta pasta"
|
||||
panel.message = "Escolha a pasta onde os resultados serão salvos."
|
||||
if panel.runModal() == .OK, let url = panel.url {
|
||||
outputFolder = url.path
|
||||
PythonBridge.call(command: "set_project_config", arguments: ["folder": url.path]) { _, _ in }
|
||||
}
|
||||
}
|
||||
|
||||
private func pickProjectFile() {
|
||||
let panel = NSOpenPanel()
|
||||
panel.canChooseFiles = true
|
||||
panel.canChooseDirectories = false
|
||||
panel.allowsMultipleSelection = false
|
||||
panel.prompt = "Selecionar"
|
||||
panel.message = "Selecione o arquivo (.fcpxml) ou o bundle (.fcpxmld) exportado pelo Final Cut Pro."
|
||||
if panel.runModal() == .OK, let url = panel.url {
|
||||
let ext = url.pathExtension.lowercased()
|
||||
if ext == "fcpxml" || ext == "fcpxmld" || ext == "xml" {
|
||||
projectPath = url.path
|
||||
PythonBridge.call(command: "set_project_config", arguments: ["file": url.path]) { _, _ in }
|
||||
} else {
|
||||
errorMessage = "Selecione um arquivo .fcpxml, .fcpxmld ou .xml do Final Cut Pro."
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
private func startTranscription() {
|
||||
guard let projectPath, let outputFolder else { return }
|
||||
isTranscribing = true
|
||||
errorMessage = nil
|
||||
transcribeResults = []
|
||||
transcribeProgress = 0
|
||||
PythonBridge.run(command: "transcribe", arguments: ["path": projectPath, "output_dir": outputFolder]) { obj in
|
||||
DispatchQueue.main.async {
|
||||
let type = obj["type"] as? String
|
||||
if type == "progress" {
|
||||
transcribeProgress = (obj["fraction"] as? NSNumber)?.doubleValue ?? 0
|
||||
transcribeStage = obj["stage"] as? String ?? ""
|
||||
} else if type == "error" {
|
||||
errorMessage = obj["message"] as? String ?? "Erro na transcrição."
|
||||
} else if type == "result", let arr = obj["transcripts"] as? [[String: Any]] {
|
||||
transcribeResults = arr.map(TranscriptResult.init)
|
||||
}
|
||||
}
|
||||
} completion: { code, err in
|
||||
DispatchQueue.main.async {
|
||||
isTranscribing = false
|
||||
transcribeProgress = 1
|
||||
if code != 0 && transcribeResults.isEmpty {
|
||||
errorMessage = err ?? "A transcrição falhou."
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
private func analyzeVoice(forceReprocess: Bool = false) {
|
||||
guard let projectPath, let outputFolder else { return }
|
||||
isAnalyzing = true
|
||||
errorMessage = nil
|
||||
PythonBridge.call(command: "analyze_voice", arguments: [
|
||||
"path": projectPath,
|
||||
"output_dir": outputFolder,
|
||||
"force_reprocess": forceReprocess,
|
||||
]) { result, err in
|
||||
DispatchQueue.main.async {
|
||||
isAnalyzing = false
|
||||
guard result?["ok"] as? Bool == true else {
|
||||
errorMessage = result?["error"] as? String ?? err ?? "Falha ao analisar a voz."
|
||||
return
|
||||
}
|
||||
if result?["reused"] as? Bool == true, !forceReprocess {
|
||||
let timelines = result?["timelines"] as? [String] ?? []
|
||||
existingVoiceTimelinePath = timelines.first ?? extractPath(from: result?["message"] as? String ?? "", marker: "**Timeline JSON**:")
|
||||
showVoiceTimelineReuseAlert = true
|
||||
return
|
||||
}
|
||||
let message = result?["message"] as? String ?? ""
|
||||
voiceAnalysisMessage = message
|
||||
if let path = extractPath(from: message, marker: "**Timeline JSON**:") {
|
||||
voiceTimelinePath = path
|
||||
} else {
|
||||
voiceTimelinePath = nil
|
||||
// ok:true não garante que a análise gerou timeline — se
|
||||
// não houver fala detectável no áudio, o Python volta com
|
||||
// sucesso mas sem "Timeline JSON" na mensagem. Sem isso
|
||||
// aqui, a etapa parecia não fazer nada.
|
||||
errorMessage = "A análise terminou mas não encontrou fala reconhecível no áudio. Mensagem do motor: " + (message.isEmpty ? "(vazia)" : message)
|
||||
}
|
||||
checkAcoustics()
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
/// A ênfase de voz (energia/tom) depende do `librosa`, dependência
|
||||
/// opcional. Sem ela, a análise ainda transcreve e corta pelo texto,
|
||||
/// mas nunca deveria propor zoom — por isso avisamos aqui, no ponto
|
||||
/// onde o usuário sentiria falta, em vez de só na aba Modelos.
|
||||
private func checkAcoustics() {
|
||||
PythonBridge.call(command: "acoustics_capability") { result, _ in
|
||||
DispatchQueue.main.async {
|
||||
guard let result, result["ok"] as? Bool == true else { return }
|
||||
acousticsAvailable = result["available"] as? Bool
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
/// Localiza uma linha markdown do tipo "- **Marker**: valor" (usado nas
|
||||
/// mensagens do bridge Python) e devolve o valor. Aceita o marcador de
|
||||
/// lista "- " opcional antes dos asteriscos.
|
||||
private func extractPath(from message: String, marker: String) -> String? {
|
||||
for line in message.split(separator: "\n") {
|
||||
var trimmed = Substring(line.trimmingCharacters(in: .whitespaces))
|
||||
if trimmed.hasPrefix("- ") { trimmed = trimmed.dropFirst(2) }
|
||||
if trimmed.hasPrefix(marker) {
|
||||
return trimmed.dropFirst(marker.count).trimmingCharacters(in: .whitespaces)
|
||||
}
|
||||
}
|
||||
return nil
|
||||
}
|
||||
|
||||
private func copyForChat(path: String) {
|
||||
guard let content = try? String(contentsOfFile: path, encoding: .utf8) else {
|
||||
errorMessage = "Não foi possível ler \(path)."
|
||||
return
|
||||
}
|
||||
let prompt = """
|
||||
Use a skill "editar-por-voz" para decidir os cortes deste projeto a partir da timeline de voz abaixo. Devolva só o JSON de decisões (cortes, zooms, textos, marcadores) pronto para eu colar de volta no app.
|
||||
|
||||
```json
|
||||
\(content)
|
||||
```
|
||||
"""
|
||||
let pasteboard = NSPasteboard.general
|
||||
pasteboard.clearContents()
|
||||
pasteboard.setString(prompt, forType: .string)
|
||||
copiedFeedback = "Copiado — cole (⌘V) numa conversa com o Claude."
|
||||
}
|
||||
|
||||
/// Monta a revisão uma vez por análise de voz. Voltar e avançar de novo com
|
||||
/// as MESMAS decisões não recarrega: isso jogaria fora as edições manuais
|
||||
/// em silêncio, que é exatamente o que esta tela existe para preservar.
|
||||
///
|
||||
/// Mas se o usuário voltou à etapa 4 e colou/gerou um JSON de decisões
|
||||
/// DIFERENTE do que gerou a revisão atual, isso é recarregado — e com
|
||||
/// `fresh: true`, para que o `active`/ênfase recém-derivado dessas
|
||||
/// decisões novas não seja imediatamente sobrescrito pela revisão salva
|
||||
/// da visita anterior (`merge_saved_decisions`, do lado Python). Sem isso,
|
||||
/// a tela ficava presa nas decisões antigas mesmo depois de reaplicar o
|
||||
/// corte — a dessincronia relatada entre "ativa aqui" e "já cortado no
|
||||
/// FCPXML".
|
||||
private func loadReviewIfNeeded() {
|
||||
guard let voiceTimelinePath else { return }
|
||||
let decisionsChanged = reviewLoadedForDecisions != nil && reviewLoadedForDecisions != decisionsText
|
||||
guard reviewLoadedFor != voiceTimelinePath || decisionsChanged else { return }
|
||||
reviewLoadedFor = voiceTimelinePath
|
||||
reviewLoadedForDecisions = decisionsText
|
||||
// A pasta do projeto e a do .fcpxml entram como onde procurar a mídia:
|
||||
// a análise de voz guarda só o nome do arquivo, não o caminho.
|
||||
reviewModel.load(
|
||||
voiceTimelinePath: voiceTimelinePath,
|
||||
decisionsJSON: decisionsText,
|
||||
outputFolder: outputFolder,
|
||||
mediaFolder: projectPath.map { URL(fileURLWithPath: $0).deletingLastPathComponent().path },
|
||||
fresh: decisionsChanged
|
||||
)
|
||||
}
|
||||
|
||||
private func applyDecisions() {
|
||||
guard let projectPath, let outputFolder,
|
||||
let data = decisionsText.data(using: .utf8),
|
||||
let parsed = try? JSONSerialization.jsonObject(with: data) else { return }
|
||||
isApplyingDecisions = true
|
||||
errorMessage = nil
|
||||
PythonBridge.call(command: "apply_voice_actions", arguments: [
|
||||
"path": projectPath,
|
||||
"output_dir": outputFolder,
|
||||
"actions": parsed,
|
||||
]) { result, err in
|
||||
DispatchQueue.main.async {
|
||||
isApplyingDecisions = false
|
||||
guard result?["ok"] as? Bool == true else {
|
||||
errorMessage = result?["error"] as? String ?? err ?? "Falha ao aplicar as decisões."
|
||||
return
|
||||
}
|
||||
appliedPath = result?["path"] as? String ?? projectPath
|
||||
skippedVoiceEdit = false
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
/// Etapa 4 (alternativa): manda a voice timeline inteira para um modelo
|
||||
/// local (Ollama/Gemma 3) que dirige a edição de uma vez — sem copiar e
|
||||
/// colar. O motor devolve o roteiro legível + o JSON de ações e já aplica
|
||||
/// no FCPXML (non-destructive), igual ao fluxo manual "Aplicar decisões".
|
||||
private func fetchOllamaModels() {
|
||||
guard ollamaModels.isEmpty else { return }
|
||||
PythonBridge.call(command: "list_ollama_models", arguments: [:]) { result, err in
|
||||
DispatchQueue.main.async {
|
||||
if let models = result?["models"] as? [String], !models.isEmpty {
|
||||
ollamaModels = models
|
||||
if !models.contains(generateScriptModel) {
|
||||
generateScriptModel = models.first ?? generateScriptModel
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
private func generateScript(voiceTimelinePath: String) {
|
||||
guard let projectPath, let outputFolder else { return }
|
||||
isGeneratingScript = true
|
||||
generateScriptFeedback = ""
|
||||
errorMessage = nil
|
||||
PythonBridge.call(command: "generate_voice_script", arguments: [
|
||||
"voice_timeline": voiceTimelinePath,
|
||||
"filepath": projectPath,
|
||||
"output_dir": outputFolder,
|
||||
"model": generateScriptModel,
|
||||
"apply_to_fcpxml": true,
|
||||
]) { result, err in
|
||||
DispatchQueue.main.async {
|
||||
isGeneratingScript = false
|
||||
guard result?["ok"] as? Bool == true else {
|
||||
errorMessage = result?["error"] as? String ?? err ?? "Falha ao gerar roteiro por IA local."
|
||||
return
|
||||
}
|
||||
// Traz as decisões de volta para a tela de revisão (etapa 5) e
|
||||
// marca como aplicadas, exatamente como o "Aplicar decisões".
|
||||
if let actionsPath = result?["actions_path"] as? String,
|
||||
let content = try? String(contentsOfFile: actionsPath, encoding: .utf8) {
|
||||
decisionsText = content
|
||||
}
|
||||
appliedPath = result?["applied_path"] as? String ?? projectPath
|
||||
skippedVoiceEdit = false
|
||||
generateScriptFeedback = "Roteiro gerado e aplicado — revise as ênfases na próxima etapa."
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
private func finalizeProcessing() {
|
||||
guard let outputFolder else { return }
|
||||
let startPath = appliedPath ?? projectPath
|
||||
guard let startPath else { return }
|
||||
var operations: [String] = []
|
||||
if finalSilences { operations.append("remove_silences") }
|
||||
if finalFillers { operations.append("remove_filler_words") }
|
||||
if finalSubtitles { operations.append("generate_plain_subtitles") }
|
||||
if finalDynamicSubtitles { operations.append("generate_dynamic_subtitles") }
|
||||
guard !operations.isEmpty else { return }
|
||||
isFinalizing = true
|
||||
errorMessage = nil
|
||||
finalStatus = "Iniciando…"
|
||||
applyReviewDecisions(startPath: startPath, outputFolder: outputFolder) { reviewedPath in
|
||||
finalizeStep(operations, index: 0, currentPath: reviewedPath, outputFolder: outputFolder)
|
||||
}
|
||||
}
|
||||
|
||||
/// Reapplies whatever the etapa-5 review decided (active/inactive
|
||||
/// phrases, manual zooms) on top of `startPath` before the finishing
|
||||
/// chain runs below. Without this, `appliedPath` stayed frozen at
|
||||
/// whatever `exportarChat`'s `apply_voice_actions` produced BEFORE the
|
||||
/// human review — so toggling a phrase off in the review only updated
|
||||
/// `_phrase_actions.json` on disk, never the video the wizard actually
|
||||
/// exports. A no-op (just hands `startPath` straight through) when the
|
||||
/// review step was never visited/saved this session.
|
||||
private func applyReviewDecisions(
|
||||
startPath: String, outputFolder: String, completion: @escaping (String) -> Void
|
||||
) {
|
||||
guard let phraseActionsPath,
|
||||
let data = try? Data(contentsOf: URL(fileURLWithPath: phraseActionsPath)),
|
||||
let parsed = try? JSONSerialization.jsonObject(with: data) as? [String: Any],
|
||||
let actions = parsed["actions"] else {
|
||||
completion(startPath)
|
||||
return
|
||||
}
|
||||
finalStatus = "Aplicando a revisão…"
|
||||
PythonBridge.call(command: "apply_voice_actions", arguments: [
|
||||
"path": startPath,
|
||||
"output_dir": outputFolder,
|
||||
"actions": actions,
|
||||
]) { result, err in
|
||||
DispatchQueue.main.async {
|
||||
guard result?["ok"] as? Bool == true else {
|
||||
isFinalizing = false
|
||||
errorMessage = result?["error"] as? String ?? err ?? "Falha ao aplicar a revisão."
|
||||
finalStatus = "Processamento interrompido."
|
||||
return
|
||||
}
|
||||
completion(result?["path"] as? String ?? startPath)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
private func finalizeStep(_ operations: [String], index: Int, currentPath: String, outputFolder: String) {
|
||||
guard index < operations.count else {
|
||||
isFinalizing = false
|
||||
finalStatus = "Processamento concluído."
|
||||
finalPath = currentPath
|
||||
return
|
||||
}
|
||||
let operation = operations[index]
|
||||
finalStatus = "Processando: \(operation)…"
|
||||
PythonBridge.call(command: operation, arguments: ["path": currentPath, "output_dir": outputFolder]) { result, err in
|
||||
DispatchQueue.main.async {
|
||||
guard result?["ok"] as? Bool == true else {
|
||||
isFinalizing = false
|
||||
errorMessage = result?["error"] as? String ?? err ?? "Falha em \(operation)."
|
||||
finalStatus = "Processamento interrompido."
|
||||
return
|
||||
}
|
||||
let nextPath = result?["path"] as? String ?? currentPath
|
||||
finalizeStep(operations, index: index + 1, currentPath: nextPath, outputFolder: outputFolder)
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
Submodule code/WHISPERX deleted from c9ed3cc6bd
+131
@@ -0,0 +1,131 @@
|
||||
#!/usr/bin/env python3
|
||||
"""AI Voice Editor - Pipeline completo: transcrição + análise acústica → JSON para IA.
|
||||
|
||||
Uso:
|
||||
python ai_edit.py <media_path> [--model base] [--lang pt] [--no-diarize] [--output dir]
|
||||
|
||||
Gera dois arquivos na pasta output (ou ao lado do mídia):
|
||||
<nome>_transcript.json — transcrição com timestamps por palavra
|
||||
<nome>_voice_timeline.json — timeline de voz com ênfase, pitch, energy, speakers
|
||||
|
||||
Esses arquivos são a ENTRADA para a IA analisar e gerar o roteiro/edição.
|
||||
"""
|
||||
|
||||
import argparse
|
||||
import json
|
||||
import sys
|
||||
import time
|
||||
from pathlib import Path
|
||||
|
||||
|
||||
def main():
|
||||
parser = argparse.ArgumentParser(description="Pipeline de análise de voz para IA")
|
||||
parser.add_argument("media", help="Caminho do arquivo de mídia (.mp4, .mov, .wav, etc.)")
|
||||
parser.add_argument("--model", default="base", help="Modelo Whisper (tiny/base/small/medium/large-v3)")
|
||||
parser.add_argument("--lang", default=None, help="Idioma (ex: pt, en). Auto-detect se omitido")
|
||||
parser.add_argument("--hf-token", default=None, help="HuggingFace token para diarização (opcional)")
|
||||
parser.add_argument("--no-diarize", action="store_true", help="Pular diarização de falantes")
|
||||
parser.add_argument("--output", default=None, help="Pasta de saída (padrão: ao lado do mídia)")
|
||||
parser.add_argument("--no-align", action="store_true", help="Pular alinhamento fonético (whisperx)")
|
||||
args = parser.parse_args()
|
||||
|
||||
media_path = Path(args.media).resolve()
|
||||
if not media_path.is_file():
|
||||
print(f"ERRO: Arquivo não encontrado: {media_path}", file=sys.stderr)
|
||||
sys.exit(1)
|
||||
|
||||
# Output dir
|
||||
out_dir = Path(args.output) if args.output else media_path.parent
|
||||
out_dir.mkdir(parents=True, exist_ok=True)
|
||||
stem = media_path.stem
|
||||
|
||||
# ── Fase 1: Transcrição ──────────────────────────────────────────
|
||||
print(f"[1/2] Transcrevendo {media_path.name} (modelo: {args.model})...")
|
||||
t0 = time.time()
|
||||
|
||||
# Adiciona code/ ao path para imports do projeto
|
||||
code_dir = Path(__file__).resolve().parent
|
||||
sys.path.insert(0, str(code_dir))
|
||||
|
||||
from fcpxml.transcribe import transcribe
|
||||
|
||||
def transcribe_progress(pct):
|
||||
bar_len = 30
|
||||
filled = int(bar_len * pct)
|
||||
bar = "█" * filled + "░" * (bar_len - filled)
|
||||
print(f"\r [{bar}] {pct*100:.0f}%", end="", flush=True)
|
||||
|
||||
transcript = transcribe(
|
||||
str(media_path),
|
||||
model_size=args.model,
|
||||
language=args.lang,
|
||||
progress_cb=transcribe_progress,
|
||||
align=not args.no_align,
|
||||
)
|
||||
print() # newline after progress bar
|
||||
|
||||
if transcript is None:
|
||||
print("ERRO: Transcrição falhou. Verifique se faster-whisper está instalado:", file=sys.stderr)
|
||||
print(" uv pip install faster-whisper", file=sys.stderr)
|
||||
sys.exit(1)
|
||||
|
||||
print(f" → {len(transcript.get('words', []))} palavras, "
|
||||
f"{len(transcript.get('segments', []))} segmentos, "
|
||||
f"idioma: {transcript.get('language', '?')}")
|
||||
|
||||
# Salva transcrição
|
||||
transcript_path = out_dir / f"{stem}_transcript.json"
|
||||
transcript_path.write_text(json.dumps(transcript, indent=2, ensure_ascii=False), encoding="utf-8")
|
||||
print(f" → Salvo: {transcript_path}")
|
||||
|
||||
# ── Fase 2: Análise de voz (timeline) ────────────────────────────
|
||||
print(f"\n[2/2] Analisando voz (pitch, energia, ênfase)...")
|
||||
t1 = time.time()
|
||||
|
||||
from fcpxml.voice_timeline import build_voice_timeline
|
||||
|
||||
def voice_progress(fraction, stage):
|
||||
print(f"\r {stage} ({fraction*100:.0f}%)", end="", flush=True)
|
||||
|
||||
hf_token = None if args.no_diarize else args.hf_token
|
||||
timeline = build_voice_timeline(
|
||||
str(media_path),
|
||||
transcript,
|
||||
hf_token=hf_token,
|
||||
progress_cb=voice_progress,
|
||||
)
|
||||
print()
|
||||
|
||||
# Salva voice timeline
|
||||
timeline_path = out_dir / f"{stem}_voice_timeline.json"
|
||||
timeline_path.write_text(json.dumps(timeline, indent=2, ensure_ascii=False), encoding="utf-8")
|
||||
print(f" → Salvo: {timeline_path}")
|
||||
|
||||
# ── Resumo ───────────────────────────────────────────────────────
|
||||
elapsed = time.time() - t0
|
||||
summary = timeline.get("summary", {})
|
||||
layers = timeline.get("layers", {})
|
||||
n_words = len(transcript.get("words", []))
|
||||
n_segments = len(transcript.get("segments", []))
|
||||
n_speakers = len(timeline.get("speakers", []))
|
||||
duration = transcript.get("duration", 0)
|
||||
|
||||
print(f"\n{'='*50}")
|
||||
print(f" ARQUIVOS GERADOS:")
|
||||
print(f" {transcript_path}")
|
||||
print(f" {timeline_path}")
|
||||
print(f"\n RESUMO:")
|
||||
print(f" Duração: {duration:.1f}s ({duration/60:.1f}min)")
|
||||
print(f" Palavras: {n_words}")
|
||||
print(f" Segmentos: {n_segments}")
|
||||
print(f" Falantes: {n_speakers}")
|
||||
print(f" Camadas: transcript={layers.get('transcript')}, "
|
||||
f"acoustics={layers.get('acoustics')}, "
|
||||
f"diarization={layers.get('diarization')}")
|
||||
print(f" Tempo: {elapsed:.1f}s")
|
||||
print(f"{'='*50}")
|
||||
print(f"\n→ Pronto! Agora peça à IA para analisar o voice timeline e gerar o roteiro.")
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
@@ -0,0 +1,181 @@
|
||||
"""Forced alignment — refine word timestamps against an acoustic model.
|
||||
|
||||
Why this exists
|
||||
--------------
|
||||
faster-whisper derives word times by cross-attention, which lands every word
|
||||
*start* systematically ~0.3-0.5s early (the word-end is fine). That bias flows
|
||||
straight into the voice timeline and makes zoom/cut land on the wrong frame —
|
||||
measured on real footage in ``Engine/docs/05_EXPERIENCIAS.md`` (#14). Phonetic
|
||||
forced alignment (wav2vec2, via whisperx) re-anchors each word against the
|
||||
audio and brings that error down to ~30ms.
|
||||
|
||||
Design
|
||||
------
|
||||
* The dependency (``whisperx``) is **optional** and imported lazily, exactly
|
||||
like the rest of this stack (librosa, faster-whisper). When it is missing, or
|
||||
any step fails, :meth:`ForcedAligner.align` returns the words unchanged, so
|
||||
transcription never breaks because alignment did.
|
||||
* The aligner is a single responsibility class: it knows how to turn a
|
||||
transcript into the shape whisperx wants, call it, and write the refined
|
||||
times back. ``transcribe.py`` owns the decision of *whether* to align.
|
||||
* Align models are cached per language on the instance so repeated calls
|
||||
(e.g. many short clips) don't reload the wav2vec2 weights each time.
|
||||
"""
|
||||
|
||||
import logging
|
||||
from typing import List, Optional, Sequence
|
||||
|
||||
logger = logging.getLogger(__name__)
|
||||
|
||||
|
||||
class ForcedAligner:
|
||||
"""Refine word-level timestamps with whisperx phonetic forced alignment.
|
||||
|
||||
Usage::
|
||||
|
||||
aligner = ForcedAligner()
|
||||
words = aligner.align(words, raw_segments, media_path, language, models_dir)
|
||||
|
||||
``words`` and ``raw_segments`` come straight from :func:`transcribe` —
|
||||
``raw_segments`` carries the per-segment ``words`` lists (the same dict
|
||||
objects as in ``words``) so the aligner knows which words belong to which
|
||||
audio window. Returns a list of the *same* word dicts, with ``start``/``end``
|
||||
overwritten in place where alignment produced a usable time.
|
||||
"""
|
||||
|
||||
def __init__(self, device: Optional[str] = None):
|
||||
self._device = device
|
||||
self._models: dict = {}
|
||||
|
||||
# -- capability ------------------------------------------------------
|
||||
@staticmethod
|
||||
def available() -> bool:
|
||||
"""Whether whisperx can be imported (the aligner can run at all)."""
|
||||
try:
|
||||
import whisperx # noqa: F401
|
||||
except Exception:
|
||||
return False
|
||||
return True
|
||||
|
||||
def _resolve_device(self) -> str:
|
||||
if self._device:
|
||||
return self._device
|
||||
try:
|
||||
import torch
|
||||
|
||||
if torch.cuda.is_available():
|
||||
return "cuda"
|
||||
except Exception:
|
||||
pass
|
||||
return "cpu"
|
||||
|
||||
# -- public API ------------------------------------------------------
|
||||
def align(
|
||||
self,
|
||||
words: Sequence[dict],
|
||||
raw_segments: Sequence[dict],
|
||||
audio_path: str,
|
||||
language: str,
|
||||
models_dir: Optional[str] = None,
|
||||
) -> List[dict]:
|
||||
"""Return ``words`` with forced-aligned timestamps where possible.
|
||||
|
||||
Falls back to the unchanged ``words`` on any failure (missing
|
||||
dependency, model load error, audio read error, or a result that
|
||||
doesn't line up with the input).
|
||||
"""
|
||||
if not words or not language:
|
||||
return list(words)
|
||||
try:
|
||||
import whisperx
|
||||
except Exception:
|
||||
logger.info("whisperx not installed; skipping forced alignment")
|
||||
return list(words)
|
||||
|
||||
try:
|
||||
device = self._resolve_device()
|
||||
align_input = self._build_align_input(words, raw_segments)
|
||||
audio = whisperx.load_audio(audio_path)
|
||||
|
||||
if language not in self._models:
|
||||
align_model, metadata = whisperx.load_align_model(
|
||||
language_code=language,
|
||||
device=device,
|
||||
model_dir=str(models_dir) if models_dir else None,
|
||||
)
|
||||
self._models[language] = (align_model, metadata)
|
||||
align_model, metadata = self._models[language]
|
||||
|
||||
result = whisperx.align(
|
||||
align_input,
|
||||
align_model,
|
||||
metadata,
|
||||
audio,
|
||||
device,
|
||||
return_char_alignments=False,
|
||||
)
|
||||
return self._merge_result(words, result.get("segments", []))
|
||||
except Exception:
|
||||
logger.warning(
|
||||
"forced alignment failed for %s; using raw timestamps", audio_path
|
||||
)
|
||||
return list(words)
|
||||
|
||||
# -- internals -------------------------------------------------------
|
||||
@staticmethod
|
||||
def _build_align_input(
|
||||
words: Sequence[dict], raw_segments: Sequence[dict]
|
||||
) -> List[dict]:
|
||||
"""Transcript in whisperx's expected shape: segments -> words.
|
||||
|
||||
whisperx.align requires each segment to carry ``text``/``start``/``end``
|
||||
and a ``words`` list whose entries have ``word``/``start``/``end``/``score``.
|
||||
We only read ``words`` from ``raw_segments`` (the flattened ``words``
|
||||
list is the source of truth for counts), so the two stay consistent.
|
||||
"""
|
||||
align_segments: List[dict] = []
|
||||
for seg in raw_segments:
|
||||
seg_words = [
|
||||
{
|
||||
"word": w.get("word", ""),
|
||||
"start": float(w.get("start", 0.0)),
|
||||
"end": float(w.get("end", 0.0)),
|
||||
"score": float(w.get("confidence", 0.0)),
|
||||
}
|
||||
for w in seg.get("words", [])
|
||||
]
|
||||
align_segments.append(
|
||||
{
|
||||
"text": (seg.get("text") or "").strip(),
|
||||
"start": float(seg.get("start", 0.0)),
|
||||
"end": float(seg.get("end", 0.0)),
|
||||
"words": seg_words,
|
||||
}
|
||||
)
|
||||
return align_segments
|
||||
|
||||
@staticmethod
|
||||
def _merge_result(words: Sequence[dict], aligned_segments: Sequence[dict]) -> List[dict]:
|
||||
"""Walk the aligned output in order and overwrite word times in place.
|
||||
|
||||
whisperx preserves word order within and across segments, so a single
|
||||
running index over the output words lines up with ``words``. A word the
|
||||
aligner failed to place gets ``None``/``0`` times — we skip those rather
|
||||
than clobber a good timestamp, and if counts ever diverge we stop and
|
||||
leave the rest untouched.
|
||||
"""
|
||||
out = list(words)
|
||||
wi = 0
|
||||
for seg in aligned_segments:
|
||||
for aw in seg.get("words", []):
|
||||
if wi >= len(out):
|
||||
return out
|
||||
start = aw.get("start")
|
||||
end = aw.get("end")
|
||||
if start is None or end is None or end < start:
|
||||
wi += 1
|
||||
continue
|
||||
out[wi]["start"] = float(start)
|
||||
out[wi]["end"] = float(end)
|
||||
wi += 1
|
||||
return out
|
||||
@@ -0,0 +1,312 @@
|
||||
"""Local LLM integration — the voice timeline meets a local model.
|
||||
|
||||
The voice timeline is *designed* to be handed to a language model: it is the
|
||||
source of truth between speech analysis and editing, layered so a model can
|
||||
reason about the narrative without parsing FCPXML. This module is the client
|
||||
side of that contract. It formats the timeline into the editar-por-voz brief,
|
||||
calls a local model server (Ollama, running Gemma 3 / Llama locally), and
|
||||
parses the model's decisions back into a validated list of VoiceActions —
|
||||
all inside the engine, so there is no wizard, no copy-paste, no manual step.
|
||||
|
||||
Transport: Ollama's HTTP chat API at ``http://localhost:11434/api/chat``.
|
||||
Any model Ollama serves works; the default is Gemma 3 because that is what
|
||||
runs locally here ("Lama com Gema 3"), but pass ``model=`` to switch.
|
||||
|
||||
The model is untrusted input: its JSON is validated row-by-row by
|
||||
:func:`fcpxml.voice_actions.parse_actions`, so one malformed decision never
|
||||
discards the edit. The brief is written so the model only ever emits the four
|
||||
action kinds the applier understands.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import json
|
||||
import logging
|
||||
import re
|
||||
from typing import Any, Dict, Optional, Sequence, Tuple
|
||||
|
||||
import httpx
|
||||
|
||||
from .voice_actions import parse_actions
|
||||
|
||||
logger = logging.getLogger(__name__)
|
||||
|
||||
DEFAULT_BASE_URL = "http://localhost:11434"
|
||||
# Gemma 3 12B reliably follows the editar-por-voz brief (keep the script, cut
|
||||
# only backstage chatter; the 4B variant skips the "keep the main content"
|
||||
# rule and deletes the script) but doesn't fit an 8GB machine. Qwen2.5 7B
|
||||
# instruct (q4_K_M) is the fallback for constrained hardware — strong at
|
||||
# strict JSON-schema following, the property this brief leans on hardest.
|
||||
# Pass ``model=`` to switch to whatever Ollama serves.
|
||||
DEFAULT_MODEL = "qwen2.5:7b-instruct-q4_K_M"
|
||||
REQUEST_TIMEOUT = 600.0
|
||||
|
||||
# The brief. Ported from the editar-por-voz skill criteria (criterios/01..08),
|
||||
# condensed into the instructions a model needs to emit valid actions. Kept in
|
||||
# Portuguese because the decisions and their reasons are read by a human editor.
|
||||
_SYSTEM_PROMPT = """Você é o editor de vídeo por voz deste sistema. Recebe um JSON de "linha do tempo de voz" — a medição de COMO foi falado (ênfase, energia, pausa, falante) de uma gravação — e devolve as DECISÕES de edição em JSON, nada mais. Você nunca escreve XML.
|
||||
|
||||
Regras (siga rigorosamente):
|
||||
|
||||
1. LEIA EM CAMADAS. "summary" dá o formato da peça; "segments" é onde você trabalha (cada fala com seu texto e agregados); "segments[].words" dá o instante exato de cada destaque. Não recalcule energia, tom ou ênfase — use os números do JSON.
|
||||
|
||||
2. SEPARAR ROTEIRO DE BASTIDOR.
|
||||
- ROTEIRO = o conteúdo principal que a pessoa quer entregar: explicação, depoimento, roteiro decorado, a mensagem. É isso que VAI FICAR.
|
||||
- BASTIDOR = papo casual de gravação, cumprimentos, conversa com a equipe ("cara, beleza?", "tá gravando?", "deixa eu ver o celular"), piadas fora do assunto, tomadas interrompidas ou repetidas. É isso que VIRA "cut".
|
||||
Exemplo: num vídeo sobre mastopexia, a explicação da cirurgia É o roteiro (mantém); o "tá gravando? pois é" antes dela É bastidor (corta).
|
||||
Use "gap_before" e "take_boundary" (silêncio > ~3s = a câmera parou/recomeçou) para agrupar tomadas — eles marcam ONDE a tomada recomeça, não o que cortar. Nunca corte o conteúdo principal só porque tem ênfase; corte o casual/off-topic.
|
||||
|
||||
REGRAS DE OURO:
|
||||
- MANTENHA o conteúdo principal (explicação, depoimento, roteiro decorado). Ele É o vídeo.
|
||||
- CORTE SÓ o casual/off-topic: cumprimentos, "tá gravando?", papo com a equipe, olhar o celular, repetições de tomada.
|
||||
- Em dúvida, MANTENHA a fala. É melhor sobrar conteúdo do que cortar o que era pra ficar.
|
||||
|
||||
3. ESCOLHER A MELHOR TOMADA de cada frase quando há repetições: mantenha a mais limpa e corte as outras (cut cobrindo a frase inteira).
|
||||
|
||||
4. CORTE (kind "cut"): para REMOVER uma frase, cubra ela inteira (start..end = início..fim da frase). Para APARAR só uma hesitação no começo ou fim, corte só da borda até a palavra (corte de meia frase é ambíguo — passe de 60% e apaga a linha toda). Nunca corte o silêncio entre falas. Quando a borda do corte encosta em fala mantida (não em silêncio puro), recue ~0,15-0,25s para dentro do corte nos dois lados — start ~0,2s DEPOIS do fim real da última palavra mantida, end ~0,2s ANTES do início real da próxima palavra mantida — senão o corte soa seco, engolindo a palavra antes de terminar de soar. Isso vale também pro início/fim do vídeo (ar morto antes da primeira palavra e depois da última).
|
||||
|
||||
5. ZOOM (kind "zoom"): só em palavra de CONTEÚDO bem enfatizada (emphasis alto, não artigo). params.scale entre 1.0 e 3.0 (padrão 1.3 se omitido). Posicione em torno da palavra, segurando até o fim da frase.
|
||||
|
||||
6. TEXTO (kind "text"): params.content obrigatório (≤120 chars), fixa um termo central ou callout. MARKER (kind "marker"): opcional params.content vira o nome do marcador. Use para emendas/junções que o editor deve conferir.
|
||||
|
||||
7. TEMPOS em segundos da MÍDIA ORIGINAL (exatamente como no JSON). Nunca compense para "depois do corte" — o programa desloca sozinho. end sempre > start, ambos ≥ 0.
|
||||
|
||||
8. reason OBRIGATÓRIO em cada ação, em português, embasando a decisão (ex.: 'abertura: "Aquela mama" (ênfase 0.42)'). reason vazio é decisão sem critério.
|
||||
|
||||
Responda APENAS com um objeto JSON válido, sem markdown, sem comentário:
|
||||
{"source": "<nome do arquivo>", "actions": [{"kind": "cut|zoom|text|marker", "start": <float>, "end": <float>, "params": {}, "reason": "<pt>", "speaker": "<id>"}]}
|
||||
"""
|
||||
|
||||
_OUTPUT_REMINDER = """Gere as decisões de edição conforme o brief. Responda SOMENTE o JSON:
|
||||
{"source": "<nome do arquivo>", "actions": [{"kind": "cut|zoom|text|marker", "start": <float>, "end": <float>, "params": {}, "reason": "<pt>", "speaker": "<id>"}]}
|
||||
Não inclua explicações nem blocos markdown."""
|
||||
|
||||
|
||||
def ollama_chat(
|
||||
model: str = DEFAULT_MODEL,
|
||||
messages: Optional[Sequence[Dict[str, str]]] = None,
|
||||
base_url: str = DEFAULT_BASE_URL,
|
||||
temperature: float = 0.2,
|
||||
timeout: float = REQUEST_TIMEOUT,
|
||||
num_ctx: int = 32768,
|
||||
) -> str:
|
||||
"""One chat completion from a local Ollama server.
|
||||
|
||||
Returns the assistant message content. Raises on transport/HTTP errors so
|
||||
the caller can decide whether to retry or report — a model call is the
|
||||
one I/O in this pipeline that can legitimately fail mid-run.
|
||||
"""
|
||||
payload = {
|
||||
"model": model,
|
||||
"messages": list(messages or []),
|
||||
"stream": False,
|
||||
"options": {"temperature": temperature, "num_ctx": num_ctx},
|
||||
}
|
||||
try:
|
||||
response = httpx.post(
|
||||
f"{base_url.rstrip('/')}/api/chat", json=payload, timeout=timeout
|
||||
)
|
||||
response.raise_for_status()
|
||||
data = response.json()
|
||||
except Exception as exc:
|
||||
# Covers transport errors AND a dropped connection that yields an empty
|
||||
# body (httpx/JSONDecodeError) — both must become a RuntimeError so the
|
||||
# caller reports the failure instead of crashing the whole pipeline.
|
||||
raise RuntimeError(f"Falha ao falar com o modelo local em {base_url}: {exc}") from exc
|
||||
|
||||
return (data.get("message") or {}).get("content", "") or ""
|
||||
|
||||
|
||||
def list_ollama_models(base_url: str = DEFAULT_BASE_URL) -> list[str]:
|
||||
"""Names of the models Ollama currently serves, for a model picker.
|
||||
|
||||
Returns an empty list when Ollama is unreachable so the UI can fall back to
|
||||
a free-text field instead of erroring.
|
||||
"""
|
||||
try:
|
||||
resp = httpx.get(f"{base_url.rstrip('/')}/api/tags", timeout=10.0)
|
||||
resp.raise_for_status()
|
||||
models = resp.json().get("models", [])
|
||||
names = [m.get("name") for m in models if m.get("name")]
|
||||
return sorted(names)
|
||||
except Exception:
|
||||
return []
|
||||
|
||||
|
||||
def _extract_json(text: str) -> Any:
|
||||
"""Pull a JSON value out of a model response, tolerating fences/wrappers."""
|
||||
if not text:
|
||||
return None
|
||||
candidate = text.strip()
|
||||
# Strip a ```json ... ``` (or bare ```) fence if the model added one.
|
||||
fence = re.search(r"```(?:json)?\s*(.*?)\s*```", candidate, re.DOTALL)
|
||||
if fence:
|
||||
candidate = fence.group(1).strip()
|
||||
# Otherwise take the outermost {...} / [...].
|
||||
if not candidate.startswith(("{" if True else "", "[")):
|
||||
start = min(
|
||||
(i for i, c in enumerate(candidate) if c in "{["),
|
||||
default=None,
|
||||
)
|
||||
end = max(
|
||||
(i for i, c in enumerate(candidate) if c in "}"),
|
||||
default=None,
|
||||
)
|
||||
if start is not None and end is not None and end > start:
|
||||
candidate = candidate[start : end + 1]
|
||||
try:
|
||||
data = json.loads(candidate)
|
||||
except json.JSONDecodeError:
|
||||
return None
|
||||
|
||||
# Models sometimes wrap the expected `{"source", "actions"}` object inside a
|
||||
# single-element list (`[{...}]`). Unwrap that so the actions aren't treated
|
||||
# as one malformed row.
|
||||
if (
|
||||
isinstance(data, list)
|
||||
and len(data) == 1
|
||||
and isinstance(data[0], dict)
|
||||
and "actions" in data[0] # the wrapper carries the actions key
|
||||
):
|
||||
data = data[0]
|
||||
return data
|
||||
|
||||
|
||||
# Only these fields reach the model — the raw timeline also carries heavy
|
||||
# per-word audio features (energy, pitch, arousal...) and speaker `samples`
|
||||
# that blow past the model's context window on any real recording. Dropping
|
||||
# them is what keeps a 3-minute timeline inside `num_ctx`.
|
||||
_SEGMENT_KEEP = (
|
||||
"start", "end", "speaker", "text", "gap_before", "take_boundary",
|
||||
"avg_energy", "peak_emphasis", "emotion", "emotion_confidence",
|
||||
"arousal", "valence",
|
||||
)
|
||||
_WORD_KEEP = ("text", "start", "end", "speaker", "emphasis", "pause_before")
|
||||
_SPEAKER_KEEP = ("id", "name")
|
||||
_SKIP_ROOT = ("layers", "scales")
|
||||
|
||||
|
||||
def _project_timeline(timeline: dict) -> dict:
|
||||
"""Strip the timeline down to what the edit decision actually needs."""
|
||||
out = {k: v for k, v in timeline.items() if k not in _SKIP_ROOT}
|
||||
speakers = [
|
||||
{k: sp[k] for k in _SPEAKER_KEEP if k in sp}
|
||||
for sp in timeline.get("speakers", [])
|
||||
]
|
||||
if speakers:
|
||||
out["speakers"] = speakers
|
||||
segs = []
|
||||
for seg in timeline.get("segments", []):
|
||||
s = {k: seg[k] for k in _SEGMENT_KEEP if k in seg}
|
||||
s["words"] = [
|
||||
{k: w[k] for k in _WORD_KEEP if k in w}
|
||||
for w in seg.get("words", [])
|
||||
]
|
||||
segs.append(s)
|
||||
out["segments"] = segs
|
||||
return out
|
||||
|
||||
|
||||
def _shrink_to_fit(compact: dict, max_chars: int) -> dict:
|
||||
"""Drop word detail from the lowest-emphasis segments until it fits."""
|
||||
segs = [dict(s) for s in compact.get("segments", [])]
|
||||
while True:
|
||||
payload = json.dumps(
|
||||
{**compact, "segments": segs}, ensure_ascii=False, indent=1
|
||||
)
|
||||
if len(payload) <= max_chars or not any(s.get("words") for s in segs):
|
||||
break
|
||||
idx = min(
|
||||
(i for i, s in enumerate(segs) if s.get("words")),
|
||||
key=lambda i: float(segs[i].get("peak_emphasis", 0.0)),
|
||||
)
|
||||
segs[idx] = {**segs[idx], "words": []}
|
||||
compact = dict(compact)
|
||||
compact["segments"] = segs
|
||||
return compact
|
||||
|
||||
|
||||
def build_edit_messages(
|
||||
timeline: dict, max_words_per_segment: int = 200, max_chars: int = 110000
|
||||
) -> Tuple[str, str]:
|
||||
"""The (system, user) pair that sends a timeline to the model.
|
||||
|
||||
The user turn carries a *projected* timeline (see :func:`_project_timeline`)
|
||||
— text, timing, speaker and emphasis only — so a real recording fits in the
|
||||
model's context window. Very long segments still have their word detail
|
||||
capped to ``max_words_per_segment`` (most emphatic + boundaries), and if the
|
||||
whole payload would still exceed ``max_chars`` the lowest-emphasis segments
|
||||
lose their words until it fits, so we never blow ``num_ctx``.
|
||||
"""
|
||||
compact = _project_timeline(timeline)
|
||||
if max_words_per_segment:
|
||||
segs = []
|
||||
for seg in compact["segments"]:
|
||||
words = seg.get("words", [])
|
||||
if len(words) > max_words_per_segment:
|
||||
ranked = sorted(
|
||||
enumerate(words),
|
||||
key=lambda kv: float(kv[1].get("emphasis", 0.0)),
|
||||
reverse=True,
|
||||
)[: max_words_per_segment - 2]
|
||||
keep = sorted({0, len(words) - 1} | {i for i, _ in ranked})
|
||||
seg = {**seg, "words": [words[i] for i in keep]}
|
||||
segs.append(seg)
|
||||
compact["segments"] = segs
|
||||
|
||||
payload = json.dumps(compact, ensure_ascii=False, indent=1)
|
||||
if len(payload) > max_chars:
|
||||
compact = _shrink_to_fit(compact, max_chars)
|
||||
payload = json.dumps(compact, ensure_ascii=False, indent=1)
|
||||
|
||||
user = (
|
||||
"Linha do tempo de voz (JSON):\n\n"
|
||||
+ payload
|
||||
+ "\n\n"
|
||||
+ _OUTPUT_REMINDER
|
||||
)
|
||||
return _SYSTEM_PROMPT, user
|
||||
|
||||
|
||||
def generate_voice_actions(
|
||||
timeline: dict,
|
||||
model: str = DEFAULT_MODEL,
|
||||
base_url: str = DEFAULT_BASE_URL,
|
||||
temperature: float = 0.2,
|
||||
timeout: float = REQUEST_TIMEOUT,
|
||||
num_ctx: int = 32768,
|
||||
max_words_per_segment: int = 200,
|
||||
) -> Dict[str, Any]:
|
||||
"""Ask the local model to direct the edit, returning validated actions.
|
||||
|
||||
Returns ``{"actions": [VoiceAction], "raw": str, "errors": [str]}``.
|
||||
``actions`` is empty when the model returned nothing usable; ``errors``
|
||||
carries the per-row rejections from :func:`parse_actions` plus any
|
||||
extraction failure, so the caller can report what went wrong instead of
|
||||
only the wins.
|
||||
"""
|
||||
system, user = build_edit_messages(timeline, max_words_per_segment)
|
||||
try:
|
||||
raw = ollama_chat(
|
||||
model=model,
|
||||
messages=[
|
||||
{"role": "system", "content": system},
|
||||
{"role": "user", "content": user},
|
||||
],
|
||||
base_url=base_url,
|
||||
temperature=temperature,
|
||||
timeout=timeout,
|
||||
num_ctx=num_ctx,
|
||||
)
|
||||
except RuntimeError as exc:
|
||||
return {"actions": [], "raw": "", "errors": [str(exc)]}
|
||||
|
||||
data = _extract_json(raw)
|
||||
if data is None:
|
||||
return {
|
||||
"actions": [],
|
||||
"raw": raw,
|
||||
"errors": ["O modelo não devolveu um JSON de decisões legível."],
|
||||
}
|
||||
actions, errors = parse_actions(data)
|
||||
return {"actions": actions, "raw": raw, "errors": errors}
|
||||
@@ -382,6 +382,10 @@ DEFAULT_VOICE_ANALYSIS_CONFIG: dict = {
|
||||
"emphasis_floor": 0.25,
|
||||
"emotion_enabled": False,
|
||||
"emotion_sensitivity": 0.5,
|
||||
"zoom_scale": 1.30,
|
||||
"zoom_mode": "in_out",
|
||||
"zoom_ease_in": 0.25,
|
||||
"zoom_ease_out": 0.04,
|
||||
}
|
||||
|
||||
|
||||
@@ -402,12 +406,28 @@ def load_voice_analysis_config() -> dict:
|
||||
stored = _load_config().get("voice_analysis")
|
||||
if not isinstance(stored, dict):
|
||||
return cfg
|
||||
for key in ("energy_threshold", "peak_percentile", "emphasis_floor", "emotion_sensitivity"):
|
||||
for key in (
|
||||
"energy_threshold", "peak_percentile", "emphasis_floor",
|
||||
"emotion_sensitivity", "zoom_scale", "zoom_ease_in", "zoom_ease_out",
|
||||
):
|
||||
if key in stored:
|
||||
try:
|
||||
cfg[key] = max(0.0, min(1.0, float(stored[key])))
|
||||
value = float(stored[key])
|
||||
if key == "zoom_scale":
|
||||
cfg[key] = max(1.0, min(3.0, value))
|
||||
elif key.startswith("zoom_ease"):
|
||||
cfg[key] = max(0.01, min(5.0, value))
|
||||
else:
|
||||
cfg[key] = max(0.0, min(1.0, value))
|
||||
except (TypeError, ValueError):
|
||||
pass
|
||||
if "emphasis_threshold" in stored and "emphasis_floor" not in stored:
|
||||
try:
|
||||
cfg["emphasis_floor"] = max(0.0, min(1.0, float(stored["emphasis_threshold"])))
|
||||
except (TypeError, ValueError):
|
||||
pass
|
||||
if stored.get("zoom_mode") in ("in_out", "in", "out"):
|
||||
cfg["zoom_mode"] = stored["zoom_mode"]
|
||||
if "emotion_enabled" in stored:
|
||||
cfg["emotion_enabled"] = bool(stored["emotion_enabled"])
|
||||
weights = stored.get("emphasis_weights")
|
||||
@@ -428,6 +448,10 @@ def save_voice_analysis_config(
|
||||
emphasis_floor: float | None = None,
|
||||
emotion_enabled: bool | None = None,
|
||||
emotion_sensitivity: float | None = None,
|
||||
zoom_scale: float | None = None,
|
||||
zoom_mode: str | None = None,
|
||||
zoom_ease_in: float | None = None,
|
||||
zoom_ease_out: float | None = None,
|
||||
) -> dict:
|
||||
"""Persist voice-analysis thresholds/weights. Only given fields change.
|
||||
|
||||
@@ -446,6 +470,14 @@ def save_voice_analysis_config(
|
||||
cfg["emotion_enabled"] = bool(emotion_enabled)
|
||||
if emotion_sensitivity is not None:
|
||||
cfg["emotion_sensitivity"] = max(0.0, min(1.0, float(emotion_sensitivity)))
|
||||
if zoom_scale is not None:
|
||||
cfg["zoom_scale"] = max(1.0, min(3.0, float(zoom_scale)))
|
||||
if zoom_mode in ("in_out", "in", "out"):
|
||||
cfg["zoom_mode"] = zoom_mode
|
||||
if zoom_ease_in is not None:
|
||||
cfg["zoom_ease_in"] = max(0.01, min(5.0, float(zoom_ease_in)))
|
||||
if zoom_ease_out is not None:
|
||||
cfg["zoom_ease_out"] = max(0.01, min(5.0, float(zoom_ease_out)))
|
||||
if emphasis_weights is not None:
|
||||
for key, value in emphasis_weights.items():
|
||||
if key in cfg["emphasis_weights"] and value is not None:
|
||||
@@ -536,6 +568,73 @@ def save_dynamic_subtitle_config(**fields) -> dict:
|
||||
return cfg
|
||||
|
||||
|
||||
DEFAULT_PLAIN_SUBTITLE_CONFIG: dict = {
|
||||
"font": "Helvetica Neue",
|
||||
"font_size": 82,
|
||||
"font_color": "1 1 1 1",
|
||||
"max_words": 7,
|
||||
"position_y": -820.0,
|
||||
"uppercase": False,
|
||||
"keep_punctuation": True,
|
||||
"text_scale": 2.0,
|
||||
}
|
||||
|
||||
|
||||
def load_plain_subtitle_config() -> dict:
|
||||
"""Persisted style for simple editable FCPXML title subtitles."""
|
||||
cfg = dict(DEFAULT_PLAIN_SUBTITLE_CONFIG)
|
||||
stored = _load_config().get("plain_subtitles")
|
||||
if not isinstance(stored, dict):
|
||||
return cfg
|
||||
for key in ("position_y", "text_scale"):
|
||||
if key in stored:
|
||||
try:
|
||||
cfg[key] = float(stored[key])
|
||||
except (TypeError, ValueError):
|
||||
pass
|
||||
for key in ("font_size", "max_words"):
|
||||
if key in stored:
|
||||
try:
|
||||
cfg[key] = int(stored[key])
|
||||
except (TypeError, ValueError):
|
||||
pass
|
||||
for key in ("font", "font_color"):
|
||||
if key in stored and isinstance(stored[key], str) and stored[key]:
|
||||
cfg[key] = stored[key]
|
||||
for key in ("uppercase", "keep_punctuation"):
|
||||
if key in stored:
|
||||
cfg[key] = bool(stored[key])
|
||||
cfg["max_words"] = max(1, int(cfg["max_words"]))
|
||||
return cfg
|
||||
|
||||
|
||||
def save_plain_subtitle_config(**fields) -> dict:
|
||||
"""Persist simple subtitle style fields. Only given fields change."""
|
||||
cfg = load_plain_subtitle_config()
|
||||
for key, value in fields.items():
|
||||
if key not in DEFAULT_PLAIN_SUBTITLE_CONFIG or value is None:
|
||||
continue
|
||||
if isinstance(DEFAULT_PLAIN_SUBTITLE_CONFIG[key], bool):
|
||||
cfg[key] = bool(value)
|
||||
elif isinstance(DEFAULT_PLAIN_SUBTITLE_CONFIG[key], float):
|
||||
try:
|
||||
cfg[key] = float(value)
|
||||
except (TypeError, ValueError):
|
||||
continue
|
||||
elif isinstance(DEFAULT_PLAIN_SUBTITLE_CONFIG[key], int):
|
||||
try:
|
||||
cfg[key] = int(value)
|
||||
except (TypeError, ValueError):
|
||||
continue
|
||||
else:
|
||||
cfg[key] = str(value)
|
||||
cfg["max_words"] = max(1, int(cfg["max_words"]))
|
||||
data = _load_config()
|
||||
data["plain_subtitles"] = cfg
|
||||
_write_config(data)
|
||||
return cfg
|
||||
|
||||
|
||||
# Mirrors the silence thresholds the detection/removal handlers use when no
|
||||
# argument is passed (server_tools/qc.py). Persisted so the app's slider and
|
||||
# any later run agree without threading three fields through every call.
|
||||
@@ -545,7 +644,14 @@ DEFAULT_SILENCE_CONFIG: dict = {
|
||||
# Seconds a quiet stretch must last before it's a cut candidate.
|
||||
"min_silence": 0.5,
|
||||
# Seconds left inside each cut so speech never gets clipped at the edges.
|
||||
"padding": 0.05,
|
||||
# 0.2s matches the breathing-room convention for phrase-boundary cuts
|
||||
# (see editar-por-voz/criterios/06-texto-corte-marcador.md) — a silence
|
||||
# span this tool finds is often the natural breath before a new
|
||||
# sentence, not just editing slop, and 0.05s shaved that breath down to
|
||||
# almost nothing (real case: Mastopexia project, the pause before "Com"
|
||||
# went from 0.567s to 0.1s across the cut, landing the next clip only
|
||||
# 5ms after the word instead of a natural pause before it).
|
||||
"padding": 0.2,
|
||||
}
|
||||
|
||||
|
||||
|
||||
File diff suppressed because it is too large
Load Diff
@@ -0,0 +1,120 @@
|
||||
"""
|
||||
Data models for Final Cut Pro FCPXML structures.
|
||||
|
||||
Provides a clean Python interface for working with Final Cut Pro timelines,
|
||||
clips, markers, and other elements.
|
||||
|
||||
Era um módulo de 1.091 linhas com seis famílias de modelo dentro. Agora cada
|
||||
família tem seu arquivo, e este pacote reexporta tudo — `from .models import
|
||||
TimeValue` segue valendo em todo o projeto, inclusive para os nomes com
|
||||
underscore que o writer e a suíte já usavam.
|
||||
|
||||
enums tipos e cores de marcador, transições, ritmo
|
||||
timing TimeValue (fração racional) e Timecode
|
||||
timeline clipes, marcadores, lanes, projeto
|
||||
planning rough cut, ritmo, montagem
|
||||
qc achados de QC e resultado de validação
|
||||
subtitles paleta e look das legendas dinâmicas
|
||||
"""
|
||||
|
||||
from .enums import (
|
||||
_MAX_MARKER_TYPE_LENGTH,
|
||||
MARKER_XML_TAGS,
|
||||
FlashFrameSeverity,
|
||||
MarkerColor,
|
||||
MarkerType,
|
||||
PacingCurve,
|
||||
PacingStyle,
|
||||
TransitionType,
|
||||
ValidationIssueType,
|
||||
)
|
||||
from .planning import (
|
||||
MontageConfig,
|
||||
PacingConfig,
|
||||
RoughCutResult,
|
||||
SegmentSpec,
|
||||
)
|
||||
from .qc import (
|
||||
DuplicateGroup,
|
||||
FlashFrame,
|
||||
GapInfo,
|
||||
ValidationIssue,
|
||||
ValidationResult,
|
||||
)
|
||||
from .subtitles import (
|
||||
COLOR_GREY,
|
||||
COLOR_INDIGO,
|
||||
COLOR_WHITE,
|
||||
COLOR_YELLOW,
|
||||
EDITORIAL_BODY_LOOK,
|
||||
EDITORIAL_EMPHASIS_LOOK,
|
||||
REFERENCE_RHYTHM,
|
||||
DynamicSubtitleConfig,
|
||||
SubtitlePosition,
|
||||
WordLook,
|
||||
WordStyle,
|
||||
)
|
||||
from .timeline import (
|
||||
AudioClip,
|
||||
Clip,
|
||||
CompoundClip,
|
||||
ConnectedClip,
|
||||
Keyword,
|
||||
Marker,
|
||||
Project,
|
||||
SilenceCandidate,
|
||||
Timeline,
|
||||
Transition,
|
||||
VideoClip,
|
||||
)
|
||||
from .timing import (
|
||||
_FCPXML_STANDARD_TIMEBASES,
|
||||
Timecode,
|
||||
TimeValue,
|
||||
)
|
||||
|
||||
__all__ = [
|
||||
"AudioClip",
|
||||
"COLOR_GREY",
|
||||
"COLOR_INDIGO",
|
||||
"COLOR_WHITE",
|
||||
"COLOR_YELLOW",
|
||||
"Clip",
|
||||
"CompoundClip",
|
||||
"ConnectedClip",
|
||||
"DuplicateGroup",
|
||||
"DynamicSubtitleConfig",
|
||||
"EDITORIAL_BODY_LOOK",
|
||||
"EDITORIAL_EMPHASIS_LOOK",
|
||||
"FlashFrame",
|
||||
"FlashFrameSeverity",
|
||||
"GapInfo",
|
||||
"Keyword",
|
||||
"MARKER_XML_TAGS",
|
||||
"Marker",
|
||||
"MarkerColor",
|
||||
"MarkerType",
|
||||
"MontageConfig",
|
||||
"PacingConfig",
|
||||
"PacingCurve",
|
||||
"PacingStyle",
|
||||
"Project",
|
||||
"REFERENCE_RHYTHM",
|
||||
"RoughCutResult",
|
||||
"SegmentSpec",
|
||||
"SilenceCandidate",
|
||||
"SubtitlePosition",
|
||||
"TimeValue",
|
||||
"Timecode",
|
||||
"Timeline",
|
||||
"Transition",
|
||||
"TransitionType",
|
||||
"ValidationIssue",
|
||||
"ValidationIssueType",
|
||||
"ValidationResult",
|
||||
"VideoClip",
|
||||
"WordLook",
|
||||
"WordStyle",
|
||||
"_FCPXML_STANDARD_TIMEBASES",
|
||||
"_MAX_MARKER_TYPE_LENGTH",
|
||||
]
|
||||
@@ -0,0 +1,183 @@
|
||||
"""Enumerações do domínio: tipos e cores de marcador, transições, ritmo.
|
||||
|
||||
Extraído de models.py — ver fcpxml/models/__init__.py.
|
||||
"""
|
||||
|
||||
from enum import Enum
|
||||
|
||||
# Maximum length for marker type strings to prevent memory abuse
|
||||
_MAX_MARKER_TYPE_LENGTH = 64
|
||||
|
||||
class MarkerType(Enum):
|
||||
"""Types of markers in Final Cut Pro.
|
||||
|
||||
Members:
|
||||
STANDARD — Default marker with no completion state.
|
||||
INCOMPLETE — Task marker (completed="0" in FCPXML). ← canonical name
|
||||
TODO — Alias for INCOMPLETE. Kept for backward compatibility;
|
||||
resolves to the same object (``MarkerType.TODO is
|
||||
MarkerType.INCOMPLETE``). Python enums treat the first
|
||||
member with a given value as canonical; all subsequent
|
||||
members sharing that value become aliases.
|
||||
CHAPTER — Chapter marker (``<chapter-marker>`` element).
|
||||
COMPLETED — Task marker with completed="1".
|
||||
|
||||
Serialization helpers:
|
||||
``from_string()`` — Accepts values, names, and legacy aliases
|
||||
(e.g. ``"todo-marker"``). Always returns the
|
||||
canonical member.
|
||||
``from_xml_element()`` — Reads an ``lxml``/``ElementTree`` element and
|
||||
returns the appropriate type based on the tag
|
||||
name and ``completed`` attribute.
|
||||
``xml_tag`` — The FCPXML element tag to emit when writing.
|
||||
``xml_attrs`` — Extra attributes required when writing (e.g.
|
||||
``completed="0"`` for INCOMPLETE).
|
||||
"""
|
||||
STANDARD = "standard"
|
||||
INCOMPLETE = "todo"
|
||||
TODO = "todo" # Backward-compat alias — resolves to INCOMPLETE at runtime
|
||||
CHAPTER = "chapter"
|
||||
COMPLETED = "completed"
|
||||
|
||||
@classmethod
|
||||
def from_string(cls, value: str) -> 'MarkerType':
|
||||
"""Convert a string to MarkerType, accepting both enum names and values.
|
||||
|
||||
Includes input validation: rejects null bytes, control characters,
|
||||
and excessively long strings to prevent injection and memory abuse.
|
||||
|
||||
Examples:
|
||||
MarkerType.from_string("todo") -> MarkerType.INCOMPLETE
|
||||
MarkerType.from_string("TODO") -> MarkerType.INCOMPLETE
|
||||
MarkerType.from_string("completed") -> MarkerType.COMPLETED
|
||||
"""
|
||||
if not isinstance(value, str):
|
||||
raise TypeError(f"Expected str, got {type(value).__name__}")
|
||||
if '\x00' in value or any(ord(c) < 32 and c not in ('\n', '\r', '\t') for c in value):
|
||||
raise ValueError("Marker type contains invalid control characters")
|
||||
if len(value) > _MAX_MARKER_TYPE_LENGTH:
|
||||
raise ValueError(
|
||||
f"Marker type exceeds maximum length ({_MAX_MARKER_TYPE_LENGTH} chars)"
|
||||
)
|
||||
lowered = value.strip().lower()
|
||||
if not lowered:
|
||||
raise ValueError("Marker type cannot be empty")
|
||||
# Accept legacy aliases from older specs (e.g. "todo-marker" → INCOMPLETE)
|
||||
aliases = {
|
||||
"todo-marker": "todo",
|
||||
"completed-marker": "completed",
|
||||
"chapter-marker": "chapter",
|
||||
}
|
||||
lowered = aliases.get(lowered, lowered)
|
||||
try:
|
||||
return cls(lowered)
|
||||
except ValueError:
|
||||
raise ValueError(
|
||||
f"Invalid marker type: '{value}'. "
|
||||
f"Valid types: {', '.join(m.value for m in cls)}"
|
||||
)
|
||||
|
||||
@classmethod
|
||||
def from_xml_element(cls, elem) -> 'MarkerType':
|
||||
"""Determine MarkerType from an XML element's tag and attributes.
|
||||
|
||||
Centralises the parse-side mapping so the parser doesn't need to
|
||||
know about completed-attribute semantics.
|
||||
|
||||
Rules (in priority order):
|
||||
1. <chapter-marker> tag → CHAPTER (completed attr ignored)
|
||||
2. completed='0' (exact) → INCOMPLETE
|
||||
3. completed='1' (exact) → COMPLETED
|
||||
4. Everything else → STANDARD (including whitespace-padded,
|
||||
absent, empty, or non-boolean completed values)
|
||||
|
||||
Matching is intentionally strict — no .strip(), no case folding.
|
||||
This prevents whitespace-injected attributes like ' 0 ' from
|
||||
being misclassified.
|
||||
"""
|
||||
if elem.tag == 'chapter-marker':
|
||||
return cls.CHAPTER
|
||||
completed = elem.get('completed')
|
||||
if completed == '0':
|
||||
return cls.INCOMPLETE
|
||||
if completed == '1':
|
||||
return cls.COMPLETED
|
||||
return cls.STANDARD
|
||||
|
||||
@property
|
||||
def xml_tag(self) -> str:
|
||||
"""Return the FCPXML element tag for this marker type."""
|
||||
return 'chapter-marker' if self == MarkerType.CHAPTER else 'marker'
|
||||
|
||||
@property
|
||||
def xml_attrs(self) -> dict:
|
||||
"""Return extra XML attributes this marker type requires when writing.
|
||||
|
||||
Centralises the write-side mapping so both FCPXMLModifier and
|
||||
FCPXMLWriter use a single source of truth.
|
||||
"""
|
||||
if self == MarkerType.CHAPTER:
|
||||
return {'posterOffset': '0s'}
|
||||
if self == MarkerType.INCOMPLETE:
|
||||
return {'completed': '0'}
|
||||
if self == MarkerType.COMPLETED:
|
||||
return {'completed': '1'}
|
||||
return {}
|
||||
|
||||
# Recognised marker XML tags — used by the parser for single-pass collection
|
||||
# and by the writer to validate element creation.
|
||||
MARKER_XML_TAGS = ('marker', 'chapter-marker')
|
||||
|
||||
class MarkerColor(Enum):
|
||||
"""Marker color options (FCP internal values)."""
|
||||
BLUE = 0
|
||||
CYAN = 1
|
||||
GREEN = 2
|
||||
YELLOW = 3
|
||||
ORANGE = 4
|
||||
RED = 5
|
||||
PINK = 6
|
||||
PURPLE = 7
|
||||
|
||||
class TransitionType(Enum):
|
||||
"""Built-in transition types."""
|
||||
CROSS_DISSOLVE = "Cross Dissolve"
|
||||
FADE_TO_BLACK = "Fade to Color"
|
||||
FADE_FROM_BLACK = "Fade from Color"
|
||||
DIP_TO_COLOR = "Dip to Color"
|
||||
WIPE = "Wipe"
|
||||
SLIDE = "Slide"
|
||||
|
||||
class PacingStyle(Enum):
|
||||
"""Pacing presets for rough cut generation."""
|
||||
SLOW = "slow" # 5-10 second cuts
|
||||
MEDIUM = "medium" # 2-5 second cuts
|
||||
FAST = "fast" # 0.5-2 second cuts
|
||||
DYNAMIC = "dynamic" # Varies throughout
|
||||
|
||||
class FlashFrameSeverity(Enum):
|
||||
"""Severity levels for flash frame detection."""
|
||||
CRITICAL = "critical" # < 2 frames, almost certainly an error
|
||||
WARNING = "warning" # < 6 frames, potentially intentional but suspicious
|
||||
|
||||
class PacingCurve(Enum):
|
||||
"""Pacing curves for montage generation."""
|
||||
CONSTANT = "constant" # Same clip duration throughout
|
||||
ACCELERATING = "accelerating" # Starts slow, gets faster
|
||||
DECELERATING = "decelerating" # Starts fast, gets slower
|
||||
PYRAMID = "pyramid" # Slow → fast → slow
|
||||
|
||||
class ValidationIssueType(Enum):
|
||||
"""Types of timeline validation issues."""
|
||||
FLASH_FRAME = "flash_frame"
|
||||
GAP = "gap"
|
||||
DUPLICATE = "duplicate"
|
||||
ORPHAN_REF = "orphan_ref"
|
||||
INVALID_OFFSET = "invalid_offset"
|
||||
# DTD validation types (v0.6.0)
|
||||
ELEMENT_ORDER = "element_order"
|
||||
MISSING_ATTRIBUTE = "missing_attribute"
|
||||
INVALID_TIMEBASE = "invalid_timebase"
|
||||
FRAME_MISALIGNMENT = "frame_misalignment"
|
||||
MISSING_EFFECT_REF = "missing_effect_ref"
|
||||
MISSING_MEDIA_REP = "missing_media_rep"
|
||||
@@ -0,0 +1,93 @@
|
||||
"""Especificações de geração: rough cut, ritmo e montagem.
|
||||
|
||||
Extraído de models.py — ver fcpxml/models/__init__.py.
|
||||
"""
|
||||
|
||||
from dataclasses import dataclass, field
|
||||
from typing import List, Optional, Tuple
|
||||
|
||||
from .enums import PacingCurve
|
||||
|
||||
|
||||
@dataclass
|
||||
class SegmentSpec:
|
||||
"""Specification for a segment in auto rough cut."""
|
||||
name: str
|
||||
keywords: List[str] = field(default_factory=list)
|
||||
duration_seconds: float = 0.0
|
||||
priority: str = "best" # favorites, longest, shortest, random, best
|
||||
|
||||
@dataclass
|
||||
class PacingConfig:
|
||||
"""Configuration for rough cut pacing."""
|
||||
pacing: str = "medium" # slow, medium, fast, dynamic
|
||||
min_clip_duration: float = 1.0
|
||||
max_clip_duration: float = 8.0
|
||||
avg_clip_duration: Optional[float] = None
|
||||
vary_pacing: bool = True
|
||||
|
||||
def get_duration_range(self) -> Tuple[float, float]:
|
||||
"""Get min/max based on pacing style."""
|
||||
ranges = {
|
||||
"slow": (5.0, 10.0),
|
||||
"medium": (2.0, 5.0),
|
||||
"fast": (0.5, 2.0),
|
||||
"dynamic": (1.0, 6.0),
|
||||
}
|
||||
return ranges.get(self.pacing, (2.0, 5.0))
|
||||
|
||||
@dataclass
|
||||
class RoughCutResult:
|
||||
"""Result of auto rough cut generation."""
|
||||
output_path: str
|
||||
clips_used: int
|
||||
clips_available: int
|
||||
target_duration: float
|
||||
actual_duration: float
|
||||
segments: int
|
||||
average_clip_duration: float
|
||||
|
||||
@dataclass
|
||||
class MontageConfig:
|
||||
"""Configuration for montage generation with pacing curves."""
|
||||
target_duration: float # Target duration in seconds
|
||||
pacing_curve: 'PacingCurve'
|
||||
start_duration: float = 2.0 # Clip duration at start
|
||||
end_duration: float = 0.5 # Clip duration at end
|
||||
min_duration: float = 0.2 # Minimum allowed clip duration
|
||||
max_duration: float = 5.0 # Maximum allowed clip duration
|
||||
|
||||
def get_duration_at_position(self, position: float) -> float:
|
||||
"""
|
||||
Calculate clip duration for a given position (0.0 to 1.0).
|
||||
|
||||
Args:
|
||||
position: Position in montage (0.0 = start, 1.0 = end)
|
||||
|
||||
Returns:
|
||||
Target duration in seconds for a clip at this position
|
||||
"""
|
||||
if self.pacing_curve == PacingCurve.CONSTANT:
|
||||
duration = (self.start_duration + self.end_duration) / 2
|
||||
|
||||
elif self.pacing_curve == PacingCurve.ACCELERATING:
|
||||
# Linear interpolation from start to end duration
|
||||
duration = self.start_duration + (self.end_duration - self.start_duration) * position
|
||||
|
||||
elif self.pacing_curve == PacingCurve.DECELERATING:
|
||||
# Reverse: start fast, end slow
|
||||
duration = self.end_duration + (self.start_duration - self.end_duration) * position
|
||||
|
||||
elif self.pacing_curve == PacingCurve.PYRAMID:
|
||||
# Slow → fast → slow (parabolic curve)
|
||||
if position < 0.5:
|
||||
# First half: slow to fast
|
||||
duration = self.start_duration + (self.end_duration - self.start_duration) * (position * 2)
|
||||
else:
|
||||
# Second half: fast to slow
|
||||
duration = self.end_duration + (self.start_duration - self.end_duration) * ((position - 0.5) * 2)
|
||||
else:
|
||||
duration = self.start_duration
|
||||
|
||||
# Clamp to min/max
|
||||
return max(self.min_duration, min(self.max_duration, duration))
|
||||
@@ -0,0 +1,121 @@
|
||||
"""Achados de QC e o resultado de uma validação.
|
||||
|
||||
Extraído de models.py — ver fcpxml/models/__init__.py.
|
||||
"""
|
||||
|
||||
from dataclasses import dataclass, field
|
||||
from typing import Any, Dict, List, Optional
|
||||
|
||||
from .enums import FlashFrameSeverity, ValidationIssueType
|
||||
from .timing import Timecode
|
||||
|
||||
|
||||
@dataclass
|
||||
class FlashFrame:
|
||||
"""
|
||||
Represents a detected flash frame (ultra-short clip).
|
||||
|
||||
Flash frames are typically editing errors - clips that are too short
|
||||
to be perceived as intentional cuts.
|
||||
"""
|
||||
clip_name: str
|
||||
clip_id: str
|
||||
start: Timecode
|
||||
duration_frames: int
|
||||
duration_seconds: float
|
||||
severity: 'FlashFrameSeverity'
|
||||
|
||||
@property
|
||||
def is_critical(self) -> bool:
|
||||
"""Check if this is a critical flash frame."""
|
||||
return self.severity == FlashFrameSeverity.CRITICAL
|
||||
|
||||
@dataclass
|
||||
class GapInfo:
|
||||
"""
|
||||
Represents a detected gap in the timeline.
|
||||
|
||||
Gaps can be intentional (black frames) or errors from deleted clips.
|
||||
"""
|
||||
start: Timecode
|
||||
duration_frames: int
|
||||
duration_seconds: float
|
||||
previous_clip: Optional[str] = None # Clip name before the gap
|
||||
next_clip: Optional[str] = None # Clip name after the gap
|
||||
|
||||
@property
|
||||
def timecode(self) -> str:
|
||||
"""Get timecode string for the gap start."""
|
||||
return self.start.to_smpte()
|
||||
|
||||
@dataclass
|
||||
class DuplicateGroup:
|
||||
"""
|
||||
Represents a group of clips using the same source media.
|
||||
|
||||
Useful for detecting duplicate clips that may be unintentional.
|
||||
"""
|
||||
source_ref: str # The asset/media reference ID
|
||||
source_name: str # Human-readable source name
|
||||
clips: List[Dict[str, Any]] = field(default_factory=list) # List of clip info dicts
|
||||
|
||||
@property
|
||||
def count(self) -> int:
|
||||
"""Number of clips using this source."""
|
||||
return len(self.clips)
|
||||
|
||||
@property
|
||||
def has_overlapping_ranges(self) -> bool:
|
||||
"""Check if any clips use overlapping portions of the source."""
|
||||
# Sort clips by source_start
|
||||
sorted_clips = sorted(self.clips, key=lambda c: c.get('source_start', 0))
|
||||
for i in range(len(sorted_clips) - 1):
|
||||
curr_end = sorted_clips[i].get('source_start', 0) + sorted_clips[i].get('source_duration', 0)
|
||||
next_start = sorted_clips[i + 1].get('source_start', 0)
|
||||
if curr_end > next_start:
|
||||
return True
|
||||
return False
|
||||
|
||||
@dataclass
|
||||
class ValidationIssue:
|
||||
"""
|
||||
Represents a single validation issue found in a timeline.
|
||||
|
||||
Used by validate_timeline to report problems.
|
||||
"""
|
||||
issue_type: 'ValidationIssueType'
|
||||
severity: str # "error", "warning", "info"
|
||||
message: str
|
||||
timecode: Optional[str] = None
|
||||
clip_name: Optional[str] = None
|
||||
details: Dict[str, Any] = field(default_factory=dict)
|
||||
|
||||
@dataclass
|
||||
class ValidationResult:
|
||||
"""
|
||||
Result of timeline validation.
|
||||
|
||||
Provides a health score and categorized list of issues.
|
||||
"""
|
||||
is_valid: bool
|
||||
health_score: int # 0-100 percentage
|
||||
issues: List[ValidationIssue] = field(default_factory=list)
|
||||
flash_frames: List[FlashFrame] = field(default_factory=list)
|
||||
gaps: List[GapInfo] = field(default_factory=list)
|
||||
duplicates: List[DuplicateGroup] = field(default_factory=list)
|
||||
|
||||
@property
|
||||
def error_count(self) -> int:
|
||||
return len([i for i in self.issues if i.severity == "error"])
|
||||
|
||||
@property
|
||||
def warning_count(self) -> int:
|
||||
return len([i for i in self.issues if i.severity == "warning"])
|
||||
|
||||
def summary(self) -> str:
|
||||
"""Generate a summary string."""
|
||||
return (
|
||||
f"Timeline Health: {self.health_score}% | "
|
||||
f"Errors: {self.error_count} | Warnings: {self.warning_count} | "
|
||||
f"Flash frames: {len(self.flash_frames)} | Gaps: {len(self.gaps)}"
|
||||
)
|
||||
@@ -0,0 +1,165 @@
|
||||
"""Aparência das legendas dinâmicas: paleta, look por palavra, configuração.
|
||||
|
||||
Extraído de models.py — ver fcpxml/models/__init__.py.
|
||||
"""
|
||||
|
||||
from dataclasses import dataclass, field
|
||||
from typing import Optional
|
||||
|
||||
from ..text_layout import REFERENCE_BLOCK_LINE_GAP, TEXT_TEMPLATE_FONT_SCALE
|
||||
|
||||
# The palette and type treatment of the calibration export
|
||||
# ("Exemplo Letra.fcpxmld", sentence "Toda a minha vida, assim,"), copied
|
||||
# verbatim from what the user set in Final Cut's Inspector.
|
||||
COLOR_INDIGO = "0.156863 0 0.596079 1"
|
||||
|
||||
COLOR_YELLOW = "0.997808 0.882664 0.0388632 1"
|
||||
|
||||
COLOR_GREY = "0.7 0.7 0.7 1"
|
||||
|
||||
COLOR_WHITE = "1 1 1 1"
|
||||
|
||||
@dataclass
|
||||
class WordLook:
|
||||
"""How one word is set: size, colour and type treatment.
|
||||
|
||||
A sentence cycles through a tuple of these, so its typography reads with a
|
||||
deliberate rhythm rather than a uniform block.
|
||||
"""
|
||||
font_size: int
|
||||
color: str
|
||||
font: str = "Helvetica Neue"
|
||||
face: Optional[str] = None # Final Cut's fontFace, e.g. "Light Italic"
|
||||
kerning: float = 2.048
|
||||
|
||||
@property
|
||||
def italic(self) -> bool:
|
||||
return bool(self.face) and "italic" in self.face.lower()
|
||||
|
||||
# One entry per word of the reference sentence, in order:
|
||||
# Toda(170, indigo, Helvetica Light) a(128, yellow) minha(151, grey)
|
||||
# vida,(128, white) assim,(128, grey, Light Italic)
|
||||
REFERENCE_RHYTHM = (
|
||||
WordLook(170, COLOR_INDIGO, font="Helvetica", face="Light", kerning=2.72),
|
||||
WordLook(128, COLOR_YELLOW),
|
||||
WordLook(151, COLOR_GREY, kerning=2.416),
|
||||
WordLook(128, COLOR_WHITE),
|
||||
WordLook(128, COLOR_GREY, face="Light Italic"),
|
||||
)
|
||||
|
||||
# The progressive-composition look (reference: the reel the user sent,
|
||||
# 2026-08-17). Supporting text in a small grotesque, the sentence's key word
|
||||
# large in a display italic, everything white — the two-font contrast IS the
|
||||
# style. Playfair Display ships in the user's ~/Library/Fonts and its real
|
||||
# advance widths are embedded in font_metrics, so the lines can be measured
|
||||
# rather than guessed. Both are plain WordLooks: swap them for any installed
|
||||
# family (a script/calligraphic face for the emphasis, say) and layout follows.
|
||||
EDITORIAL_EMPHASIS_LOOK = WordLook(
|
||||
230, COLOR_WHITE, font="Playfair Display", face="Medium Italic", kerning=0.0,
|
||||
)
|
||||
|
||||
EDITORIAL_BODY_LOOK = WordLook(
|
||||
88, COLOR_WHITE, font="Helvetica Neue", face="Bold", kerning=1.2,
|
||||
)
|
||||
|
||||
@dataclass
|
||||
class WordStyle:
|
||||
"""Per-word text styling for dynamic (karaoke-style) subtitles.
|
||||
|
||||
``rhythm`` drives size, colour and face, cycling by the word's index within
|
||||
its sentence — deterministic, so regenerating a transcript twice yields the
|
||||
same look. ``font``/``font_size`` are the fallback when ``rhythm`` is empty.
|
||||
"""
|
||||
font: str = "Helvetica Neue"
|
||||
font_size: int = 128
|
||||
active_color: str = COLOR_WHITE
|
||||
inactive_color: str = COLOR_GREY
|
||||
bold: bool = False
|
||||
kerning: float = 2.048
|
||||
rhythm: tuple = REFERENCE_RHYTHM
|
||||
# Progressive composition only (granularity="phrase").
|
||||
emphasis_look: Optional[WordLook] = None
|
||||
body_look: Optional[WordLook] = None
|
||||
|
||||
def look_for(self, index: int) -> WordLook:
|
||||
"""The look for the word at *index* within its sentence."""
|
||||
if not self.rhythm:
|
||||
return WordLook(
|
||||
self.font_size, self.active_color,
|
||||
font=self.font, kerning=self.kerning,
|
||||
)
|
||||
return self.rhythm[index % len(self.rhythm)]
|
||||
|
||||
def look_for_emphasis(self) -> WordLook:
|
||||
"""The look for a composition's key word (progressive composition)."""
|
||||
return self.emphasis_look or EDITORIAL_EMPHASIS_LOOK
|
||||
|
||||
def look_for_body(self) -> WordLook:
|
||||
"""The look for a composition's supporting lines."""
|
||||
return self.body_look or EDITORIAL_BODY_LOOK
|
||||
|
||||
@dataclass
|
||||
class SubtitlePosition:
|
||||
"""Screen position for generated title clips, in FCP title coordinate space."""
|
||||
x: float = 0.0
|
||||
y: float = -300.0
|
||||
alignment: str = "center" # left | center | right
|
||||
|
||||
@dataclass
|
||||
class DynamicSubtitleConfig:
|
||||
"""Options for FCPXMLWriter.generate_dynamic_subtitles().
|
||||
|
||||
Dynamic subtitles are animated TITLES, not captions. Both templates below
|
||||
render on the video title lane and never carry a ``subtitles.*`` role — a
|
||||
``role="subtitles.*"`` would make Final Cut treat them as captions and
|
||||
hide them behind the caption-display toggle. They DO carry a
|
||||
``titles.*`` sub-role (``role``), which groups them in Final Cut's
|
||||
role index and lanes them with a distinct colour, without ever being
|
||||
mistaken for closed captions.
|
||||
|
||||
``animated`` picks the template: True uses "Essencial - Título"
|
||||
(Essential Title), which animates on its own Motion defaults; False uses
|
||||
the static "Título Básico" (Basic Title). Default is True — the animated
|
||||
reveal is the feature's purpose.
|
||||
|
||||
Words are grouped into sentences and laid out as a compact typographic
|
||||
block: each word becomes its own positioned ``<title>``, appearing as it is
|
||||
spoken and accumulating on screen, with every word of a block clearing at
|
||||
the same instant so the sentence vanishes as a whole.
|
||||
|
||||
``band_height`` is the fraction of frame height the block may occupy, and
|
||||
``block_center_y`` its centre in canvas points (negative is below frame
|
||||
centre). The defaults reproduce the calibration export the user built by
|
||||
hand: a block of at most three lines sitting just below centre. A sentence
|
||||
taller than the band splits into successive blocks.
|
||||
"""
|
||||
style: WordStyle = field(default_factory=WordStyle)
|
||||
position: SubtitlePosition = field(default_factory=SubtitlePosition)
|
||||
animated: bool = True
|
||||
band_height: float = 0.22
|
||||
block_center_y: float = -167.0
|
||||
# "phrase": one title per LINE of the composition — supporting words
|
||||
# grouped, the key word alone and large (the reference look). "word": one
|
||||
# title per word, the earlier rhythm.
|
||||
granularity: str = "phrase"
|
||||
# Ratio between the template's fontSize space and the canvas-point space
|
||||
# its Position uses. See text_layout.TEXT_TEMPLATE_FONT_SCALE: the "Text"
|
||||
# (Text.moti) template sizes type in frame pixels, so a size chosen in
|
||||
# points renders half as large unless it is converted on the way out.
|
||||
text_scale: float = TEXT_TEMPLATE_FONT_SCALE
|
||||
# Vertical air between stacked lines, in canvas points. Negative values
|
||||
# deliberately overlap the lines — the display italic tucking under the
|
||||
# line above is a real editorial look, and the stacking arithmetic places
|
||||
# ink boxes edge to edge, so a negative gap moves them by exactly that
|
||||
# much rather than colliding unpredictably.
|
||||
line_gap: float = REFERENCE_BLOCK_LINE_GAP
|
||||
# Final Cut role for every title this generator emits. A ``titles.*``
|
||||
# sub-role (NOT ``subtitles.*``) groups the clips in the role index and
|
||||
# tints their lane, keeping dynamic captions distinct from plain
|
||||
# ones and from Final Cut's own closed-caption toggle.
|
||||
role: str = "titles.dinamicas"
|
||||
# Run the post-generation collision validation (collision.validate_titles)
|
||||
# and refuse to emit when it reports a blocking overlap. Off by default so
|
||||
# generation stays byte-identical to before this flag existed; flip it on
|
||||
# for a guaranteed no-collision export.
|
||||
validate: bool = False
|
||||
@@ -0,0 +1,248 @@
|
||||
"""O que existe numa timeline: clipes, marcadores, lanes, projeto.
|
||||
|
||||
Extraído de models.py — ver fcpxml/models/__init__.py.
|
||||
"""
|
||||
|
||||
from dataclasses import dataclass, field
|
||||
from typing import List, Optional
|
||||
|
||||
from .enums import MarkerColor, MarkerType
|
||||
from .timing import Timecode
|
||||
|
||||
|
||||
@dataclass
|
||||
class Keyword:
|
||||
"""Represents a keyword/tag applied to a clip."""
|
||||
value: str
|
||||
start: Optional[Timecode] = None
|
||||
duration: Optional[Timecode] = None
|
||||
|
||||
|
||||
@dataclass
|
||||
class ParametroEfeito:
|
||||
"""Um parâmetro de um filtro de efeito (``<param>`` dentro do filtro)."""
|
||||
nome: str
|
||||
valor: str
|
||||
chave: str = ""
|
||||
metadado: str = ""
|
||||
|
||||
|
||||
@dataclass
|
||||
class EfeitoAjuste:
|
||||
"""Um efeito aplicado por uma camada de ajuste (adjustment layer).
|
||||
|
||||
``uid`` é o UUID do efeito interno do Final Cut (ver ``FCP_EFFECTS`` em
|
||||
``fcpxml/writer/helpers.py`` para os efeitos built-in). ``tipo`` é
|
||||
``"video"`` ou ``"audio"`` — decide se vira ``<filter-video>`` ou
|
||||
``<filter-audio>``, filho direto do ``<clip>`` da camada de ajuste (o
|
||||
DTD não define wrapper ``<adjustment>``).
|
||||
"""
|
||||
nome: str
|
||||
uid: str
|
||||
tipo: str = "video"
|
||||
parametros: List[ParametroEfeito] = field(default_factory=list)
|
||||
|
||||
@dataclass
|
||||
class Marker:
|
||||
"""Represents a marker in the timeline."""
|
||||
name: str
|
||||
start: Timecode
|
||||
duration: Optional[Timecode] = None
|
||||
marker_type: MarkerType = MarkerType.STANDARD
|
||||
note: str = ""
|
||||
color: Optional[MarkerColor] = None
|
||||
|
||||
def to_youtube_timestamp(self) -> str:
|
||||
"""Format as YouTube chapter timestamp."""
|
||||
total_seconds = int(self.start.seconds)
|
||||
hours = total_seconds // 3600
|
||||
minutes = (total_seconds % 3600) // 60
|
||||
secs = total_seconds % 60
|
||||
if hours > 0:
|
||||
return f"{hours}:{minutes:02d}:{secs:02d}"
|
||||
return f"{minutes}:{secs:02d}"
|
||||
|
||||
@dataclass
|
||||
class Clip:
|
||||
"""Represents a clip in the timeline."""
|
||||
name: str
|
||||
start: Timecode
|
||||
duration: Timecode
|
||||
source_start: Optional[Timecode] = None
|
||||
source_end: Optional[Timecode] = None
|
||||
media_path: str = ""
|
||||
markers: List[Marker] = field(default_factory=list)
|
||||
keywords: List[Keyword] = field(default_factory=list)
|
||||
|
||||
# Extended metadata
|
||||
rating: int = 0 # 0=unrated, 1-5 stars
|
||||
is_favorite: bool = False
|
||||
is_rejected: bool = False
|
||||
|
||||
# Roles (FCP audio/video role assignments)
|
||||
audio_role: str = ""
|
||||
video_role: str = ""
|
||||
|
||||
# Connected clips (B-roll, titles, audio attached to this clip)
|
||||
connected_clips: List['ConnectedClip'] = field(default_factory=list)
|
||||
|
||||
# Edit-time correction, in degrees, from a Transform filter on the clip
|
||||
# (e.g. straightening a tilted phone shot) — not the camera's own
|
||||
# recorded orientation, which lives in the media file itself.
|
||||
rotation: float = 0.0
|
||||
|
||||
@property
|
||||
def end(self) -> Timecode:
|
||||
return Timecode(
|
||||
frames=self.start.frames + self.duration.frames,
|
||||
frame_rate=self.start.frame_rate
|
||||
)
|
||||
|
||||
@property
|
||||
def duration_seconds(self) -> float:
|
||||
return self.duration.seconds
|
||||
|
||||
@property
|
||||
def keyword_values(self) -> List[str]:
|
||||
"""Get list of keyword strings."""
|
||||
return [k.value for k in self.keywords]
|
||||
|
||||
@dataclass
|
||||
class AudioClip(Clip):
|
||||
"""Audio-specific clip."""
|
||||
channels: int = 2
|
||||
sample_rate: int = 48000
|
||||
role: str = "dialogue"
|
||||
|
||||
@dataclass
|
||||
class VideoClip(Clip):
|
||||
"""Video-specific clip."""
|
||||
width: int = 1920
|
||||
height: int = 1080
|
||||
has_audio: bool = True
|
||||
|
||||
@dataclass
|
||||
class ConnectedClip:
|
||||
"""A clip connected to a primary storyline clip (B-roll, titles, audio).
|
||||
|
||||
In FCP's magnetic timeline, connected clips hang off spine clips via lanes.
|
||||
Positive lanes are above (video overlays), negative lanes are below (audio).
|
||||
"""
|
||||
name: str
|
||||
start: Timecode
|
||||
duration: Timecode
|
||||
lane: int = 1
|
||||
offset: Optional[Timecode] = None
|
||||
source_start: Optional[Timecode] = None
|
||||
media_path: str = ""
|
||||
clip_type: str = "asset-clip"
|
||||
role: str = ""
|
||||
ref_id: str = ""
|
||||
parent_clip_name: str = ""
|
||||
markers: List[Marker] = field(default_factory=list)
|
||||
keywords: List[Keyword] = field(default_factory=list)
|
||||
rotation: float = 0.0
|
||||
|
||||
@property
|
||||
def duration_seconds(self) -> float:
|
||||
return self.duration.seconds
|
||||
|
||||
@dataclass
|
||||
class CompoundClip:
|
||||
"""A compound clip (ref-clip) containing a nested timeline."""
|
||||
name: str
|
||||
ref_id: str
|
||||
duration: Timecode
|
||||
start: Timecode
|
||||
clips: List[Clip] = field(default_factory=list)
|
||||
connected_clips: List[ConnectedClip] = field(default_factory=list)
|
||||
|
||||
@property
|
||||
def duration_seconds(self) -> float:
|
||||
return self.duration.seconds
|
||||
|
||||
@dataclass
|
||||
class SilenceCandidate:
|
||||
"""A potential silence region detected by timeline heuristics."""
|
||||
start_timecode: str
|
||||
duration_seconds: float
|
||||
reason: str # "gap", "ultra_short", "name_match", "duration_anomaly"
|
||||
confidence: float = 0.5 # 0.0 to 1.0
|
||||
clip_name: Optional[str] = None
|
||||
clip_index: Optional[int] = None
|
||||
|
||||
@dataclass
|
||||
class Transition:
|
||||
"""Represents a transition between clips."""
|
||||
name: str
|
||||
duration: Timecode
|
||||
start: Timecode
|
||||
transition_type: str = "cross-dissolve"
|
||||
|
||||
@dataclass
|
||||
class Timeline:
|
||||
"""Represents a Final Cut Pro timeline/sequence."""
|
||||
name: str
|
||||
duration: Timecode
|
||||
frame_rate: float = 24.0
|
||||
width: int = 1920
|
||||
height: int = 1080
|
||||
clips: List[Clip] = field(default_factory=list)
|
||||
audio_clips: List[AudioClip] = field(default_factory=list)
|
||||
transitions: List[Transition] = field(default_factory=list)
|
||||
markers: List[Marker] = field(default_factory=list)
|
||||
connected_clips: List[ConnectedClip] = field(default_factory=list)
|
||||
compound_clips: List[CompoundClip] = field(default_factory=list)
|
||||
|
||||
@property
|
||||
def total_clips(self) -> int:
|
||||
return len(self.clips)
|
||||
|
||||
@property
|
||||
def total_cuts(self) -> int:
|
||||
return max(0, len(self.clips) - 1)
|
||||
|
||||
@property
|
||||
def average_clip_duration(self) -> float:
|
||||
if not self.clips:
|
||||
return 0.0
|
||||
return sum(c.duration_seconds for c in self.clips) / len(self.clips)
|
||||
|
||||
@property
|
||||
def cuts_per_minute(self) -> float:
|
||||
"""Average cuts per minute."""
|
||||
if self.duration.seconds <= 0:
|
||||
return 0.0
|
||||
return (self.total_cuts / self.duration.seconds) * 60
|
||||
|
||||
def get_clips_shorter_than(self, seconds: float) -> List[Clip]:
|
||||
"""Find clips shorter than threshold (flash frame detection)."""
|
||||
return [c for c in self.clips if c.duration_seconds < seconds]
|
||||
|
||||
def get_clips_longer_than(self, seconds: float) -> List[Clip]:
|
||||
"""Find clips longer than threshold."""
|
||||
return [c for c in self.clips if c.duration_seconds > seconds]
|
||||
|
||||
def get_clip_at(self, timecode: float) -> Optional[Clip]:
|
||||
"""Find the clip at a specific timecode (seconds)."""
|
||||
for clip in self.clips:
|
||||
start_sec = clip.start.seconds
|
||||
end_sec = clip.end.seconds
|
||||
if start_sec <= timecode < end_sec:
|
||||
return clip
|
||||
return None
|
||||
|
||||
def get_clips_by_keyword(self, keyword: str) -> List[Clip]:
|
||||
"""Find all clips with a specific keyword."""
|
||||
return [c for c in self.clips if keyword in c.keyword_values]
|
||||
|
||||
@dataclass
|
||||
class Project:
|
||||
"""Represents a Final Cut Pro project/library."""
|
||||
name: str
|
||||
timelines: List[Timeline] = field(default_factory=list)
|
||||
fcpxml_version: str = "1.13"
|
||||
|
||||
@property
|
||||
def primary_timeline(self) -> Optional[Timeline]:
|
||||
return self.timelines[0] if self.timelines else None
|
||||
@@ -0,0 +1,304 @@
|
||||
"""Tempo em fração racional — TimeValue e o Timecode que o embrulha.
|
||||
|
||||
Extraído de models.py — ver fcpxml/models/__init__.py.
|
||||
"""
|
||||
|
||||
import operator
|
||||
from dataclasses import dataclass
|
||||
from fractions import Fraction
|
||||
from functools import total_ordering
|
||||
from math import gcd
|
||||
from typing import Callable
|
||||
|
||||
# Standard FCPXML timebase denominators that FCP's DTD validator accepts.
|
||||
# TimeValue.to_fcpxml() only simplifies fractions when the result uses one
|
||||
# of these denominators, preventing values like "8/3s" that FCP rejects.
|
||||
_FCPXML_STANDARD_TIMEBASES = frozenset({
|
||||
1, 24, 25, 30, 48, 50, 60, 90, 96, 100, 120,
|
||||
240, 600, 2400, 4800, 9600, 48000,
|
||||
})
|
||||
|
||||
@total_ordering
|
||||
@dataclass
|
||||
class TimeValue:
|
||||
"""
|
||||
Represents time in FCPXML's rational format.
|
||||
|
||||
FCPXML uses fractions of seconds (e.g., "90/30s" for 3 seconds at 30fps).
|
||||
This class handles conversion between timecode, seconds, and FCPXML format.
|
||||
|
||||
Examples:
|
||||
TimeValue(90, 30) # 3 seconds at 30fps
|
||||
TimeValue(1, 1) # 1 second
|
||||
TimeValue.from_timecode("00:01:30:15", fps=30) # 90.5 seconds
|
||||
"""
|
||||
numerator: int
|
||||
denominator: int = 1
|
||||
|
||||
def __post_init__(self):
|
||||
if self.denominator == 0:
|
||||
raise ValueError(
|
||||
f"TimeValue denominator cannot be zero (got {self.numerator}/0). "
|
||||
"This would corrupt all downstream time calculations."
|
||||
)
|
||||
# Normalize sign: denominator must always be positive.
|
||||
# Cross-multiplication in __lt__/__eq__ assumes positive denominators;
|
||||
# __hash__ assumes canonical form. Without this, TimeValue(1, -2)
|
||||
# compares/hashes incorrectly against TimeValue(-1, 2).
|
||||
if self.denominator < 0:
|
||||
# Use object.__setattr__ because dataclass may be frozen-like
|
||||
object.__setattr__(self, 'numerator', -self.numerator)
|
||||
object.__setattr__(self, 'denominator', -self.denominator)
|
||||
|
||||
@classmethod
|
||||
def from_timecode(cls, tc: str, fps: float = 30.0) -> 'TimeValue':
|
||||
"""
|
||||
Create TimeValue from various string formats.
|
||||
|
||||
Supported formats:
|
||||
- "HH:MM:SS:FF" - Standard timecode
|
||||
- "HH:MM:SS;FF" - Drop-frame timecode
|
||||
- "30s" - Seconds
|
||||
- "90/30s" - FCPXML rational format
|
||||
- "15f" - Frames
|
||||
"""
|
||||
if not tc:
|
||||
return cls(0, 1)
|
||||
|
||||
tc = str(tc).strip()
|
||||
|
||||
# FCPXML format: "90/30s" or "30s"
|
||||
if tc.endswith('s'):
|
||||
tc_val = tc[:-1]
|
||||
if '/' in tc_val:
|
||||
parts = tc_val.split('/', 1)
|
||||
num, denom = int(parts[0]), int(parts[1])
|
||||
if denom == 0:
|
||||
raise ValueError(f"Zero denominator in timecode: {tc}")
|
||||
return cls(num, denom)
|
||||
else:
|
||||
seconds = float(tc_val)
|
||||
frames = int(round(seconds * fps))
|
||||
# int(fps) truncates NTSC rates (23.976/29.97/59.94fps) to
|
||||
# their nominal integer, mismatching the numerator (computed
|
||||
# with the real fps) against the denominator — e.g. at
|
||||
# 23.976fps this silently produced values ~1.04x too large.
|
||||
# Reconstruct the exact rational fps (24000/1001, etc.) from
|
||||
# the float instead, so numerator and denominator agree.
|
||||
fps_frac = Fraction(fps).limit_denominator(100_000)
|
||||
return cls(frames * fps_frac.denominator, fps_frac.numerator)
|
||||
|
||||
# Frame format: "15f"
|
||||
if tc.endswith('f'):
|
||||
frames = int(tc[:-1])
|
||||
return cls(frames, int(fps))
|
||||
|
||||
# Timecode format: "HH:MM:SS:FF" or "HH:MM:SS;FF"
|
||||
if ':' in tc or ';' in tc:
|
||||
parts = tc.replace(';', ':').split(':')
|
||||
if len(parts) == 4:
|
||||
h, m, s, f = map(int, parts)
|
||||
total_frames = int((h * 3600 + m * 60 + s) * fps + f)
|
||||
return cls(total_frames, int(fps))
|
||||
elif len(parts) == 3:
|
||||
h, m, s = map(int, parts)
|
||||
total_frames = int((h * 3600 + m * 60 + s) * fps)
|
||||
return cls(total_frames, int(fps))
|
||||
|
||||
# Try as plain number (seconds)
|
||||
try:
|
||||
seconds = float(tc)
|
||||
frames = int(round(seconds * fps))
|
||||
return cls(frames, int(fps))
|
||||
except ValueError:
|
||||
raise ValueError(f"Invalid timecode format: {tc}")
|
||||
|
||||
@classmethod
|
||||
def from_seconds(cls, seconds: float, fps: float = 30.0) -> 'TimeValue':
|
||||
"""Create TimeValue from decimal seconds."""
|
||||
frames = int(round(seconds * fps))
|
||||
return cls(frames, int(fps))
|
||||
|
||||
@classmethod
|
||||
def zero(cls) -> 'TimeValue':
|
||||
"""Return zero time value."""
|
||||
return cls(0, 1)
|
||||
|
||||
def to_fcpxml(self) -> str:
|
||||
"""Convert to FCPXML time string (e.g., "90/30s").
|
||||
|
||||
Only simplifies when the denominator reduces to 1 (whole seconds)
|
||||
or stays a standard FCPXML timebase. Avoids producing denominators
|
||||
like 3, 7, etc. that FCP's DTD validator may reject.
|
||||
"""
|
||||
simplified = self.simplify()
|
||||
if simplified.denominator == 1:
|
||||
return f"{simplified.numerator}s"
|
||||
# Keep original denominator if simplification produces a non-standard
|
||||
# denominator (not a multiple of common timebases: 24, 30, 25, 2400)
|
||||
if simplified.denominator in _FCPXML_STANDARD_TIMEBASES:
|
||||
return f"{simplified.numerator}/{simplified.denominator}s"
|
||||
# Fall back to unsimplified form
|
||||
return f"{self.numerator}/{self.denominator}s"
|
||||
|
||||
def to_seconds(self) -> float:
|
||||
"""Convert to decimal seconds."""
|
||||
return self.numerator / self.denominator
|
||||
|
||||
def to_timecode(self, fps: float = 30.0) -> str:
|
||||
"""Convert to HH:MM:SS:FF timecode string."""
|
||||
total_frames = int(round(self.to_seconds() * fps))
|
||||
total_secs, frames = divmod(total_frames, int(fps))
|
||||
total_mins, secs = divmod(total_secs, 60)
|
||||
hours, mins = divmod(total_mins, 60)
|
||||
return f"{hours:02d}:{mins:02d}:{secs:02d}:{frames:02d}"
|
||||
|
||||
def to_frames(self, fps: float = 30.0) -> int:
|
||||
"""Convert to frame count."""
|
||||
return int(round(self.to_seconds() * fps))
|
||||
|
||||
def simplify(self) -> 'TimeValue':
|
||||
"""Reduce fraction to simplest form."""
|
||||
if self.numerator == 0:
|
||||
return TimeValue(0, 1)
|
||||
divisor = gcd(abs(self.numerator), abs(self.denominator))
|
||||
return TimeValue(
|
||||
self.numerator // divisor,
|
||||
self.denominator // divisor
|
||||
)
|
||||
|
||||
@staticmethod
|
||||
def _lcm_denom(d1: int, d2: int) -> int:
|
||||
"""LCM of two denominators for cross-timebase arithmetic."""
|
||||
return d1 // gcd(d1, d2) * d2
|
||||
|
||||
def _binop(self, other: 'TimeValue', op: Callable[[int, int], int]) -> 'TimeValue':
|
||||
"""Shared logic for add/sub: same-denom fast path, then LCM alignment."""
|
||||
if self.denominator == other.denominator:
|
||||
return TimeValue(op(self.numerator, other.numerator), self.denominator)
|
||||
lcd = TimeValue._lcm_denom(self.denominator, other.denominator)
|
||||
return TimeValue(
|
||||
op(
|
||||
self.numerator * (lcd // self.denominator),
|
||||
other.numerator * (lcd // other.denominator),
|
||||
),
|
||||
lcd,
|
||||
)
|
||||
|
||||
def __add__(self, other: 'TimeValue') -> 'TimeValue':
|
||||
return self._binop(other, operator.add)
|
||||
|
||||
def __sub__(self, other: 'TimeValue') -> 'TimeValue':
|
||||
return self._binop(other, operator.sub)
|
||||
|
||||
def __mul__(self, scalar: float) -> 'TimeValue':
|
||||
new_num = round(self.numerator * scalar)
|
||||
return TimeValue(new_num, self.denominator)
|
||||
|
||||
def __truediv__(self, scalar: float) -> 'TimeValue':
|
||||
if scalar == 0:
|
||||
raise ZeroDivisionError("Cannot divide TimeValue by zero")
|
||||
new_denom = round(self.denominator * scalar)
|
||||
if new_denom == 0:
|
||||
raise ZeroDivisionError(
|
||||
f"Division by {scalar} rounds denominator {self.denominator} to zero"
|
||||
)
|
||||
return TimeValue(self.numerator, new_denom)
|
||||
|
||||
def __lt__(self, other: 'TimeValue') -> bool:
|
||||
# Cross-multiply to compare without float conversion:
|
||||
# a/b < c/d ↔ a*d < c*b (denominators are always positive)
|
||||
return self.numerator * other.denominator < other.numerator * self.denominator
|
||||
|
||||
def __eq__(self, other: object) -> bool:
|
||||
if not isinstance(other, TimeValue):
|
||||
return False
|
||||
# Cross-multiply for exact integer comparison
|
||||
return self.numerator * other.denominator == other.numerator * self.denominator
|
||||
|
||||
def __hash__(self) -> int:
|
||||
# Delegate to simplify() — single source of truth for canonical form.
|
||||
# __post_init__ guarantees denominator > 0, so no zero guard needed.
|
||||
s = self.simplify()
|
||||
return hash((s.numerator, s.denominator))
|
||||
|
||||
def snap_to_frame(self, fps: float) -> 'TimeValue':
|
||||
"""Round this time value to the nearest frame boundary at the given fps.
|
||||
|
||||
Uses the 2400-tick timebase (LCM of common frame rates) so results
|
||||
always land on clean frame boundaries.
|
||||
|
||||
Args:
|
||||
fps: Frame rate to snap to (e.g. 24, 30, 60)
|
||||
|
||||
Returns:
|
||||
New TimeValue snapped to the nearest frame in 2400-tick timebase.
|
||||
"""
|
||||
fps_int = int(fps)
|
||||
if fps_int <= 0:
|
||||
raise ValueError(f"fps must be positive, got {fps}")
|
||||
ticks_per_frame = 2400 // fps_int
|
||||
total_ticks = round(self.to_seconds() * 2400)
|
||||
snapped_ticks = round(total_ticks / ticks_per_frame) * ticks_per_frame
|
||||
return TimeValue(snapped_ticks, 2400)
|
||||
|
||||
def is_standard_timebase(self) -> bool:
|
||||
"""Check if this TimeValue's denominator is an FCP-accepted timebase."""
|
||||
simplified = self.simplify()
|
||||
return simplified.denominator in _FCPXML_STANDARD_TIMEBASES
|
||||
|
||||
def __repr__(self) -> str:
|
||||
return f"TimeValue({self.numerator}/{self.denominator}s = {self.to_seconds():.3f}s)"
|
||||
|
||||
@dataclass
|
||||
class Timecode:
|
||||
"""
|
||||
Represents a timecode value.
|
||||
|
||||
Note: This class exists for backwards compatibility with the parser.
|
||||
New code should prefer TimeValue for rational time math.
|
||||
"""
|
||||
frames: int
|
||||
frame_rate: float = 24.0
|
||||
drop_frame: bool = False
|
||||
|
||||
@property
|
||||
def seconds(self) -> float:
|
||||
return self.frames / self.frame_rate
|
||||
|
||||
@property
|
||||
def total_frames(self) -> int:
|
||||
return self.frames
|
||||
|
||||
def to_smpte(self) -> str:
|
||||
"""Convert to SMPTE timecode string (HH:MM:SS:FF)."""
|
||||
total_seconds = int(self.seconds)
|
||||
hours = total_seconds // 3600
|
||||
minutes = (total_seconds % 3600) // 60
|
||||
secs = total_seconds % 60
|
||||
frames = int((self.seconds - total_seconds) * self.frame_rate)
|
||||
separator = ";" if self.drop_frame else ":"
|
||||
return f"{hours:02d}:{minutes:02d}:{secs:02d}{separator}{frames:02d}"
|
||||
|
||||
@classmethod
|
||||
def from_rational(cls, rational_str: str, frame_rate: float = 24.0) -> "Timecode":
|
||||
"""Parse FCPXML rational time format (e.g., '3600/24s')."""
|
||||
if not rational_str:
|
||||
return cls(frames=0, frame_rate=frame_rate)
|
||||
if rational_str.endswith('s'):
|
||||
rational_str = rational_str[:-1]
|
||||
if '/' in rational_str:
|
||||
num, denom = rational_str.split('/')
|
||||
seconds = int(num) / int(denom)
|
||||
else:
|
||||
seconds = float(rational_str)
|
||||
frames = int(seconds * frame_rate)
|
||||
return cls(frames=frames, frame_rate=frame_rate)
|
||||
|
||||
def to_rational(self) -> str:
|
||||
"""Convert to FCPXML rational format."""
|
||||
return f"{self.frames}/{int(self.frame_rate)}s"
|
||||
|
||||
def to_time_value(self) -> TimeValue:
|
||||
"""Convert to TimeValue for rational math."""
|
||||
return TimeValue(self.frames, int(self.frame_rate))
|
||||
@@ -194,6 +194,7 @@ class FCPXMLParser:
|
||||
media_path=media_path,
|
||||
audio_role=elem.get('audioRole', ''),
|
||||
video_role=elem.get('videoRole', ''),
|
||||
rotation=self._parse_clip_rotation(elem),
|
||||
)
|
||||
|
||||
clip.markers.extend(self._collect_markers(elem))
|
||||
@@ -205,6 +206,19 @@ class FCPXMLParser:
|
||||
|
||||
return clip
|
||||
|
||||
def _parse_clip_rotation(self, elem: ET.Element) -> float:
|
||||
"""Degrees from this clip's ``<adjust-transform rotation="...">`` —
|
||||
an edit-time correction (e.g. straightening a tilted phone shot),
|
||||
not the camera's own recorded orientation. FCP writes the rotation
|
||||
as an attribute on that element, not as a filter param."""
|
||||
transform = elem.find('adjust-transform')
|
||||
if transform is None:
|
||||
return 0.0
|
||||
try:
|
||||
return float(transform.get('rotation', '0'))
|
||||
except ValueError:
|
||||
return 0.0
|
||||
|
||||
def _parse_marker_element(self, elem: ET.Element) -> Optional[Marker]:
|
||||
"""Parse any marker element (<marker> or <chapter-marker>).
|
||||
|
||||
@@ -337,6 +351,7 @@ class FCPXMLParser:
|
||||
lane=lane, offset=offset, source_start=start,
|
||||
media_path=media_path, clip_type=elem.tag, role=role,
|
||||
ref_id=ref, parent_clip_name=parent_name,
|
||||
rotation=self._parse_clip_rotation(elem),
|
||||
)
|
||||
|
||||
connected.markers.extend(self._collect_markers(elem))
|
||||
|
||||
@@ -0,0 +1,571 @@
|
||||
"""Phrase review — the human pass between the AI's decisions and the render.
|
||||
|
||||
A voice timeline says *how* every line was spoken; a list of voice actions says
|
||||
what the model decided to do about it. Neither is reviewable on its own: the
|
||||
timeline has no editorial intent, and the action list is a set of timecodes with
|
||||
no text attached. This module joins them into the one view an editor can
|
||||
actually judge — the script, phrase by phrase, each carrying the decision that
|
||||
was made about it.
|
||||
|
||||
The phrase is the unit on purpose. Emphasis, in this pipeline, is not a property
|
||||
of a word but of a line: an emphasized phrase gets a punch-in and a dynamic
|
||||
caption, everything else gets a plain caption. Keeping the same granularity in
|
||||
the review, the JSON, and the render means a toggle in the UI maps to exactly
|
||||
one editorial outcome, with nothing to reconcile in between.
|
||||
|
||||
Trimming stays inside the phrase for the same reason. A line is rarely wrong as
|
||||
a whole — it has a false start, or a trailing "né" — so each phrase carries a
|
||||
``trim_start``/``trim_end`` pair that rides on word boundaries. Editing a cut
|
||||
therefore means picking a word, never hunting for a frame, and a partial cut
|
||||
from the model arrives as a trim instead of being rounded away.
|
||||
|
||||
Round-tripping is the other half of the contract. :func:`build_phrase_review`
|
||||
derives the review from actions, :func:`phrase_review_to_actions` derives
|
||||
actions back from the edited review, and everything the editor touched wins over
|
||||
what was inferred — so re-opening the screen shows what was left there, not a
|
||||
re-derivation that quietly discards the edits.
|
||||
"""
|
||||
|
||||
import json
|
||||
from pathlib import Path
|
||||
from typing import Any, Dict, List, Optional, Sequence, Tuple
|
||||
|
||||
from .voice_actions import VoiceAction, merge_cut_ranges, parse_actions
|
||||
|
||||
PHRASE_REVIEW_VERSION = "1.0"
|
||||
|
||||
# Emphasis is stored 0-3 rather than as a float so the UI, the JSON and the
|
||||
# render agree on the same discrete decision. The thresholds map the continuous
|
||||
# `peak_emphasis` of the voice timeline onto those levels when the model gave no
|
||||
# explicit direction for a phrase.
|
||||
EMPHASIS_LEVELS = (0, 1, 2, 3)
|
||||
EMPHASIS_THRESHOLDS = (0.25, 0.45, 0.65)
|
||||
|
||||
# Zoom scale applied per emphasis level when the review is turned back into
|
||||
# actions. Level 0 never produces a zoom. The values stay inside
|
||||
# voice_actions.MIN_ZOOM_SCALE..MAX_ZOOM_SCALE.
|
||||
ZOOM_SCALE_BY_LEVEL = {1: 1.15, 2: 1.3, 3: 1.5}
|
||||
|
||||
# A phrase only survives if most of it does. Speech boundaries from a transcript
|
||||
# are approximate, so a cut clipping a fraction of a second off the tail is a
|
||||
# trim, not a removal — treating that as "phrase deleted" would grey out lines
|
||||
# that are still fully audible.
|
||||
CUT_COVERAGE_TO_DEACTIVATE = 0.6
|
||||
|
||||
# A punch-in shorter than this has no time to ramp in and back out — the writer
|
||||
# rejects the window anyway (see the zoom ease-in/ease-out shape), so refusing
|
||||
# it here turns a silent drop at render time into nothing being placed at all.
|
||||
MIN_ZOOM_DURATION = 0.4
|
||||
|
||||
TRACK_SCRIPT = "roteiro"
|
||||
TRACK_BACKSTAGE = "bastidor"
|
||||
TRACKS = (TRACK_SCRIPT, TRACK_BACKSTAGE)
|
||||
|
||||
|
||||
def resolve_source(
|
||||
source: str, voice_timeline_path: str, extra_dirs: Sequence[str] = ()
|
||||
) -> str:
|
||||
"""The playable path for a timeline's ``source``, or "" when it's gone.
|
||||
|
||||
The voice timeline stores only the media's *file name* — it is written to be
|
||||
read by a model, where a machine-specific absolute path is noise. That makes
|
||||
it useless for opening a preview, so the file is looked up where it can
|
||||
actually be: beside its own timeline JSON first (that is where
|
||||
``analyze_voice`` writes it), then in whatever project folders the caller
|
||||
knows about.
|
||||
"""
|
||||
if not source:
|
||||
return ""
|
||||
candidate = Path(source)
|
||||
if candidate.is_absolute() and candidate.is_file():
|
||||
return str(candidate)
|
||||
|
||||
directories = [Path(voice_timeline_path).parent] if voice_timeline_path else []
|
||||
directories += [Path(d) for d in extra_dirs if d]
|
||||
for directory in directories:
|
||||
found = directory / candidate.name
|
||||
if found.is_file():
|
||||
return str(found)
|
||||
return ""
|
||||
|
||||
|
||||
def _overlap(a_start: float, a_end: float, b_start: float, b_end: float) -> float:
|
||||
"""Seconds shared by two spans (0.0 when they don't touch)."""
|
||||
return max(0.0, min(a_end, b_end) - max(a_start, b_start))
|
||||
|
||||
|
||||
def _cut_coverage(
|
||||
start: float, end: float, cuts: Sequence[Tuple[float, float]]
|
||||
) -> float:
|
||||
"""Fraction of ``start``-``end`` that falls inside ``cuts`` (0-1)."""
|
||||
span = end - start
|
||||
if span <= 0:
|
||||
return 0.0
|
||||
removed = sum(_overlap(start, end, c_start, c_end) for c_start, c_end in cuts)
|
||||
return min(1.0, removed / span)
|
||||
|
||||
|
||||
def snap_to_words(
|
||||
time: float, words: Sequence[dict], fallback: float, edge: str
|
||||
) -> float:
|
||||
"""Move ``time`` onto the nearest word boundary of this phrase.
|
||||
|
||||
Trims are expressed by pointing at a word, so a trim handle that landed
|
||||
mid-word would cut a syllable in half. ``edge`` is ``"in"`` (snap to word
|
||||
starts) or ``"out"`` (snap to word ends); with no word timings available the
|
||||
time is left as-is.
|
||||
"""
|
||||
boundaries = [
|
||||
float(word.get("start" if edge == "in" else "end", 0.0)) for word in words
|
||||
]
|
||||
boundaries = [b for b in boundaries if b > 0]
|
||||
if not boundaries:
|
||||
return fallback
|
||||
return min(boundaries, key=lambda b: abs(b - time))
|
||||
|
||||
|
||||
def _trim_from_cuts(
|
||||
start: float,
|
||||
end: float,
|
||||
words: Sequence[dict],
|
||||
cuts: Sequence[Tuple[float, float]],
|
||||
) -> Tuple[float, float]:
|
||||
"""Read a partial cut over this phrase as a head/tail trim.
|
||||
|
||||
Only cuts that touch an edge become trims: a cut carved out of the middle of
|
||||
a line has no representation here (the phrase is the unit), so it is left
|
||||
for the whole-phrase coverage rule to decide.
|
||||
"""
|
||||
trim_start, trim_end = start, end
|
||||
for cut_start, cut_end in cuts:
|
||||
if _overlap(start, end, cut_start, cut_end) <= 0:
|
||||
continue
|
||||
if cut_start <= trim_start < cut_end < end:
|
||||
trim_start = snap_to_words(cut_end, words, cut_end, "in")
|
||||
if start < cut_start < trim_end <= cut_end:
|
||||
trim_end = snap_to_words(cut_start, words, cut_start, "out")
|
||||
if trim_end <= trim_start:
|
||||
return start, end
|
||||
return trim_start, trim_end
|
||||
|
||||
|
||||
def _level_from_peak(peak: float) -> int:
|
||||
"""Map a 0-1 ``peak_emphasis`` onto a 0-3 level."""
|
||||
for level, threshold in enumerate(EMPHASIS_THRESHOLDS):
|
||||
if peak < threshold:
|
||||
return level
|
||||
return 3
|
||||
|
||||
|
||||
def _level_from_scale(scale: Optional[float]) -> int:
|
||||
"""Map a zoom's scale factor back onto a 0-3 level.
|
||||
|
||||
The model is free to send any scale inside the allowed range, so this picks
|
||||
the nearest level rather than requiring one of our own three values.
|
||||
"""
|
||||
if scale is None:
|
||||
return 2
|
||||
best = 1
|
||||
smallest = None
|
||||
for level, level_scale in ZOOM_SCALE_BY_LEVEL.items():
|
||||
distance = abs(level_scale - float(scale))
|
||||
if smallest is None or distance < smallest:
|
||||
smallest, best = distance, level
|
||||
return best
|
||||
|
||||
|
||||
def _emphasis_from_actions(
|
||||
start: float,
|
||||
end: float,
|
||||
actions: Sequence[VoiceAction],
|
||||
) -> Tuple[Optional[int], str]:
|
||||
"""The level the model asked for on this phrase, and why.
|
||||
|
||||
A ``zoom`` or ``text`` action anywhere inside the phrase is read as "this
|
||||
line is the emphasis" — the model places them on the word that carries the
|
||||
point, not on the whole line, so requiring a full-span match would find
|
||||
nothing. Returns ``(None, "")`` when no action touches the phrase.
|
||||
"""
|
||||
level: Optional[int] = None
|
||||
reason = ""
|
||||
for action in actions:
|
||||
if action.kind not in ("zoom", "text"):
|
||||
continue
|
||||
if _overlap(start, end, action.start, action.end) <= 0:
|
||||
continue
|
||||
if action.kind == "zoom":
|
||||
candidate = _level_from_scale(action.params.get("scale"))
|
||||
else:
|
||||
candidate = 2
|
||||
if level is None or candidate > level:
|
||||
level = candidate
|
||||
reason = action.reason
|
||||
return level, reason
|
||||
|
||||
|
||||
def _cut_reason(
|
||||
start: float, end: float, actions: Sequence[VoiceAction]
|
||||
) -> str:
|
||||
"""The reason given for the cut that removes this phrase."""
|
||||
for action in actions:
|
||||
if action.kind != "cut":
|
||||
continue
|
||||
if _overlap(start, end, action.start, action.end) > 0 and action.reason:
|
||||
return action.reason
|
||||
return ""
|
||||
|
||||
|
||||
def build_phrase_review(
|
||||
timeline: dict,
|
||||
actions: Any = None,
|
||||
voice_timeline_path: str = "",
|
||||
extra_dirs: Sequence[str] = (),
|
||||
) -> dict:
|
||||
"""Join a voice timeline with the AI's actions into a reviewable script.
|
||||
|
||||
``actions`` accepts whatever :func:`~.voice_actions.parse_actions` accepts —
|
||||
a bare list, ``{"actions": [...]}``, or ``None`` when there is no AI pass and
|
||||
the review starts from the acoustics alone. Malformed rows are skipped and
|
||||
reported in ``errors`` rather than raising, matching the rest of the
|
||||
decision pipeline.
|
||||
"""
|
||||
parsed, errors = parse_actions(actions) if actions else ([], [])
|
||||
cuts = merge_cut_ranges(parsed)
|
||||
|
||||
phrases: List[dict] = []
|
||||
for index, segment in enumerate(timeline.get("segments", [])):
|
||||
start = float(segment.get("start", 0.0))
|
||||
end = float(segment.get("end", 0.0))
|
||||
peak = float(segment.get("peak_emphasis", 0.0))
|
||||
take_boundary = bool(segment.get("take_boundary", False))
|
||||
|
||||
words = list(segment.get("words", []))
|
||||
coverage = _cut_coverage(start, end, cuts)
|
||||
active = coverage < CUT_COVERAGE_TO_DEACTIVATE
|
||||
trim_start, trim_end = (
|
||||
_trim_from_cuts(start, end, words, cuts) if active else (start, end)
|
||||
)
|
||||
|
||||
asked_level, asked_reason = _emphasis_from_actions(start, end, parsed)
|
||||
if asked_level is not None:
|
||||
emphasis, reason = asked_level, asked_reason
|
||||
else:
|
||||
emphasis = _level_from_peak(peak)
|
||||
reason = f"ênfase {peak:.2f}" if emphasis else ""
|
||||
if not active:
|
||||
# A removed line carries the reason it was removed; the emphasis it
|
||||
# would have had is kept so re-activating it restores the decision.
|
||||
reason = _cut_reason(start, end, parsed) or reason
|
||||
|
||||
phrases.append(
|
||||
{
|
||||
"index": index,
|
||||
"start": round(start, 3),
|
||||
"end": round(end, 3),
|
||||
"trim_start": round(trim_start, 3),
|
||||
"trim_end": round(trim_end, 3),
|
||||
"text": str(segment.get("text", "")).strip(),
|
||||
"speaker": str(segment.get("speaker", "")),
|
||||
"active": active,
|
||||
"emphasis": emphasis,
|
||||
"track": TRACK_BACKSTAGE if (not active and take_boundary) else TRACK_SCRIPT,
|
||||
"peak_emphasis": round(peak, 3),
|
||||
# Delivery emotion is a heuristic over the acoustics (see
|
||||
# voice_timeline._emotion_for_word) and only means anything when
|
||||
# the analysis actually ran — `emotion_available` below is what
|
||||
# separates "spoken flat" from "never measured".
|
||||
"emotion": str(segment.get("emotion", "neutral")),
|
||||
"emotion_confidence": round(
|
||||
float(segment.get("emotion_confidence", 0.0)), 3
|
||||
),
|
||||
"take_boundary": take_boundary,
|
||||
"gap_before": round(float(segment.get("gap_before", 0.0)), 3),
|
||||
"reason": reason,
|
||||
"words": [
|
||||
{
|
||||
"text": str(word.get("text", "")),
|
||||
"start": round(float(word.get("start", 0.0)), 3),
|
||||
"end": round(float(word.get("end", 0.0)), 3),
|
||||
"energy": round(float(word.get("energy", 0.0)), 3),
|
||||
"emphasis": round(float(word.get("emphasis", 0.0)), 3),
|
||||
}
|
||||
for word in words
|
||||
],
|
||||
}
|
||||
)
|
||||
|
||||
source = timeline.get("source", "")
|
||||
layers = timeline.get("layers", {}) if isinstance(timeline.get("layers"), dict) else {}
|
||||
return {
|
||||
"version": PHRASE_REVIEW_VERSION,
|
||||
"source": source,
|
||||
"source_path": resolve_source(source, voice_timeline_path, extra_dirs),
|
||||
"rotation": float(timeline.get("rotation", 0.0)),
|
||||
"duration": round(phrases[-1]["end"], 3) if phrases else 0.0,
|
||||
"speakers": timeline.get("speakers", []),
|
||||
"emotion_available": bool(layers.get("emotion", False)),
|
||||
"phrases": phrases,
|
||||
# Punch-ins the editor places by hand on an arbitrary range, alongside
|
||||
# the whole-phrase zoom that an emphasis level produces. Both end up as
|
||||
# zoom actions; this one exists because the moment worth punching into
|
||||
# is not always a whole sentence.
|
||||
"zooms": [],
|
||||
"errors": errors,
|
||||
}
|
||||
|
||||
|
||||
def _coerce_zoom(raw: Any) -> Optional[Dict[str, float]]:
|
||||
"""Normalize one manually placed zoom range."""
|
||||
if not isinstance(raw, dict):
|
||||
return None
|
||||
try:
|
||||
start = float(raw.get("start"))
|
||||
end = float(raw.get("end"))
|
||||
except (TypeError, ValueError):
|
||||
return None
|
||||
if end - start < MIN_ZOOM_DURATION:
|
||||
return None
|
||||
return {"start": start, "end": end}
|
||||
|
||||
|
||||
def _coerce_phrase(raw: Any, index: int) -> Optional[Dict[str, Any]]:
|
||||
"""Normalize one edited phrase row coming back from the UI."""
|
||||
if not isinstance(raw, dict):
|
||||
return None
|
||||
try:
|
||||
start = float(raw.get("start"))
|
||||
end = float(raw.get("end"))
|
||||
except (TypeError, ValueError):
|
||||
return None
|
||||
if end <= start:
|
||||
return None
|
||||
try:
|
||||
emphasis = int(raw.get("emphasis", 0))
|
||||
except (TypeError, ValueError):
|
||||
emphasis = 0
|
||||
try:
|
||||
trim_start = float(raw.get("trim_start", start))
|
||||
trim_end = float(raw.get("trim_end", end))
|
||||
except (TypeError, ValueError):
|
||||
trim_start, trim_end = start, end
|
||||
# A trim that escaped the phrase, or inverted, is treated as no trim at all:
|
||||
# the UI is the only thing that writes these, and silently discarding a bad
|
||||
# pair keeps a rounding slip from deleting material the editor kept.
|
||||
if not (start <= trim_start < trim_end <= end):
|
||||
trim_start, trim_end = start, end
|
||||
track = str(raw.get("track", TRACK_SCRIPT))
|
||||
return {
|
||||
"index": int(raw.get("index", index)),
|
||||
"start": start,
|
||||
"end": end,
|
||||
"trim_start": trim_start,
|
||||
"trim_end": trim_end,
|
||||
"text": str(raw.get("text", "")).strip(),
|
||||
"speaker": str(raw.get("speaker", "")),
|
||||
"active": bool(raw.get("active", True)),
|
||||
"emphasis": min(3, max(0, emphasis)),
|
||||
"track": track if track in TRACKS else TRACK_SCRIPT,
|
||||
"reason": str(raw.get("reason", "")),
|
||||
}
|
||||
|
||||
|
||||
def phrase_review_to_actions(review: dict) -> dict:
|
||||
"""Turn an edited review back into the action list the applier consumes.
|
||||
|
||||
Every deactivated phrase becomes a ``cut``, a trimmed one becomes a cut over
|
||||
the head and/or tail it lost, and every emphasized one becomes a ``zoom``
|
||||
scaled by its level. The emphasis flags ride along in ``emphasis_spans`` so
|
||||
the caption step can give those lines the dynamic treatment and everything
|
||||
else the plain one, without re-deriving the decision from the acoustics.
|
||||
"""
|
||||
phrases = [
|
||||
coerced
|
||||
for index, raw in enumerate(review.get("phrases", []))
|
||||
if (coerced := _coerce_phrase(raw, index)) is not None
|
||||
]
|
||||
|
||||
actions: List[dict] = []
|
||||
emphasis_spans: List[dict] = []
|
||||
inactive_run: List[dict] = []
|
||||
|
||||
def flush_inactive_run() -> None:
|
||||
"""One cut per RUN of consecutive deactivated phrases, not one per
|
||||
phrase. A phrase-by-phrase cut leaves the pause BETWEEN two
|
||||
deactivated phrases uncut — that gap was never anyone's content, so
|
||||
nothing asked for it to survive, but it does anyway: a 0.1-0.5s
|
||||
sliver clip in the final timeline for every such gap. Spanning the
|
||||
whole run absorbs those gaps into the one cut."""
|
||||
if not inactive_run:
|
||||
return
|
||||
if len(inactive_run) == 1:
|
||||
reason = inactive_run[0]["reason"] or "desativada na revisão"
|
||||
else:
|
||||
reason = (
|
||||
f"desativadas na revisão ({len(inactive_run)} frases): "
|
||||
+ "; ".join(p["text"][:40] for p in inactive_run if p["text"])
|
||||
)
|
||||
actions.append(
|
||||
VoiceAction(
|
||||
kind="cut",
|
||||
start=inactive_run[0]["start"],
|
||||
end=inactive_run[-1]["end"],
|
||||
reason=reason,
|
||||
speaker=inactive_run[0]["speaker"],
|
||||
).as_dict()
|
||||
)
|
||||
inactive_run.clear()
|
||||
|
||||
for phrase in phrases:
|
||||
if not phrase["active"]:
|
||||
inactive_run.append(phrase)
|
||||
continue
|
||||
flush_inactive_run()
|
||||
|
||||
# Head and tail the editor trimmed off — each becomes its own cut, so a
|
||||
# false start disappears without taking the line with it.
|
||||
for trim_start, trim_end, where in (
|
||||
(phrase["start"], phrase["trim_start"], "início"),
|
||||
(phrase["trim_end"], phrase["end"], "fim"),
|
||||
):
|
||||
if trim_end - trim_start <= 0:
|
||||
continue
|
||||
actions.append(
|
||||
VoiceAction(
|
||||
kind="cut",
|
||||
start=trim_start,
|
||||
end=trim_end,
|
||||
reason=f"trecho do {where} da frase removido na revisão",
|
||||
speaker=phrase["speaker"],
|
||||
).as_dict()
|
||||
)
|
||||
|
||||
if phrase["emphasis"] >= 1:
|
||||
actions.append(
|
||||
VoiceAction(
|
||||
kind="zoom",
|
||||
start=phrase["trim_start"],
|
||||
end=phrase["trim_end"],
|
||||
params={"scale": ZOOM_SCALE_BY_LEVEL[phrase["emphasis"]]},
|
||||
reason=phrase["reason"] or f"ênfase nível {phrase['emphasis']}",
|
||||
speaker=phrase["speaker"],
|
||||
).as_dict()
|
||||
)
|
||||
emphasis_spans.append(
|
||||
{
|
||||
"start": phrase["trim_start"],
|
||||
"end": phrase["trim_end"],
|
||||
"level": phrase["emphasis"],
|
||||
"text": phrase["text"],
|
||||
}
|
||||
)
|
||||
flush_inactive_run()
|
||||
|
||||
# Hand-placed punch-ins carry no scale on purpose: an omitted scale lets the
|
||||
# applier use the shape configured in "Análise de Voz" (zoom_scale, ease in
|
||||
# and out), so changing that setting restyles every manual zoom instead of
|
||||
# leaving a scale frozen into each one at the moment it was drawn.
|
||||
for raw in review.get("zooms", []):
|
||||
zoom = _coerce_zoom(raw)
|
||||
if zoom is None:
|
||||
continue
|
||||
actions.append(
|
||||
VoiceAction(
|
||||
kind="zoom",
|
||||
start=zoom["start"],
|
||||
end=zoom["end"],
|
||||
reason="zoom marcado na revisão",
|
||||
).as_dict()
|
||||
)
|
||||
|
||||
return {
|
||||
"source": review.get("source", ""),
|
||||
"actions": actions,
|
||||
"emphasis_spans": emphasis_spans,
|
||||
}
|
||||
|
||||
|
||||
def merge_saved_decisions(review: dict, saved: Optional[dict]) -> dict:
|
||||
"""Lay a previously saved review's decisions over a freshly built one.
|
||||
|
||||
Only the editorial fields travel — active, emphasis, track, text, trims.
|
||||
Everything else (words, emotion, energy) is re-derived from the current
|
||||
analysis, so re-running the voice pass with better settings improves the
|
||||
screen instead of being masked by a stale copy of itself, and the saved file
|
||||
never has to carry a duplicate of data it does not own.
|
||||
|
||||
Phrases are matched by index *and* start time: if the analysis changed
|
||||
enough to move a line, the old decision for that slot is dropped rather than
|
||||
applied to a different sentence.
|
||||
"""
|
||||
if not saved:
|
||||
return review
|
||||
|
||||
review["zooms"] = [
|
||||
zoom for raw in saved.get("zooms", []) if (zoom := _coerce_zoom(raw)) is not None
|
||||
]
|
||||
|
||||
by_index = {}
|
||||
for raw in saved.get("phrases", []):
|
||||
if isinstance(raw, dict) and "index" in raw:
|
||||
by_index[raw["index"]] = raw
|
||||
|
||||
for phrase in review["phrases"]:
|
||||
previous = by_index.get(phrase["index"])
|
||||
if previous is None:
|
||||
continue
|
||||
if abs(float(previous.get("start", -1)) - phrase["start"]) > 0.25:
|
||||
continue
|
||||
phrase["active"] = bool(previous.get("active", phrase["active"]))
|
||||
phrase["emphasis"] = min(3, max(0, int(previous.get("emphasis", phrase["emphasis"]))))
|
||||
track = str(previous.get("track", phrase["track"]))
|
||||
phrase["track"] = track if track in TRACKS else phrase["track"]
|
||||
if previous.get("text"):
|
||||
phrase["text"] = str(previous["text"])
|
||||
trim_start = float(previous.get("trim_start", phrase["trim_start"]))
|
||||
trim_end = float(previous.get("trim_end", phrase["trim_end"]))
|
||||
if phrase["start"] <= trim_start < trim_end <= phrase["end"]:
|
||||
phrase["trim_start"], phrase["trim_end"] = trim_start, trim_end
|
||||
|
||||
return review
|
||||
|
||||
|
||||
def review_paths(voice_timeline_path: str) -> Tuple[Path, Path]:
|
||||
"""Where the review and its derived actions live, next to the timeline.
|
||||
|
||||
Both files sit beside the ``_voice_timeline.json`` they came from and are
|
||||
named after it, so a project folder stays readable and re-running the wizard
|
||||
on the same take overwrites its own files instead of accumulating copies.
|
||||
"""
|
||||
base = Path(voice_timeline_path)
|
||||
stem = base.stem
|
||||
if stem.endswith("_voice_timeline"):
|
||||
stem = stem[: -len("_voice_timeline")]
|
||||
return (
|
||||
base.with_name(f"{stem}_phrase_review.json"),
|
||||
base.with_name(f"{stem}_phrase_actions.json"),
|
||||
)
|
||||
|
||||
|
||||
def save_phrase_review(voice_timeline_path: str, review: dict) -> Tuple[Path, Path]:
|
||||
"""Write the edited review and the actions derived from it. Returns both paths."""
|
||||
review_path, actions_path = review_paths(voice_timeline_path)
|
||||
review_path.write_text(
|
||||
json.dumps(review, ensure_ascii=False, indent=2), encoding="utf-8"
|
||||
)
|
||||
actions_path.write_text(
|
||||
json.dumps(phrase_review_to_actions(review), ensure_ascii=False, indent=2),
|
||||
encoding="utf-8",
|
||||
)
|
||||
return review_path, actions_path
|
||||
|
||||
|
||||
def load_phrase_review(voice_timeline_path: str) -> Optional[dict]:
|
||||
"""The review saved earlier for this timeline, or ``None`` if there is none."""
|
||||
review_path, _ = review_paths(voice_timeline_path)
|
||||
if not review_path.is_file():
|
||||
return None
|
||||
try:
|
||||
data = json.loads(review_path.read_text(encoding="utf-8"))
|
||||
except (OSError, json.JSONDecodeError):
|
||||
return None
|
||||
return data if isinstance(data, dict) else None
|
||||
+96
-13
@@ -28,8 +28,11 @@ ALLOWED_MODELS = (
|
||||
)
|
||||
|
||||
# Conservative by default: interjections that are near-universally filler.
|
||||
# Portuguese "um"/"uma" are usually articles/numerals inside real phrases
|
||||
# ("de um jeito") rather than discardable hesitations, so only cut them when
|
||||
# the caller explicitly opts in through the fillers argument.
|
||||
# "like" / "so" / "actually" are speech, not noise, unless the user opts in.
|
||||
DEFAULT_FILLERS = ("um", "uh", "uhh", "umm", "erm", "ehm", "mmm", "hmm", "mhm")
|
||||
DEFAULT_FILLERS = ("uh", "uhh", "umm", "erm", "ehm", "mmm", "hmm", "mhm")
|
||||
|
||||
_NORM_RE = re.compile(r"[^\w']+")
|
||||
|
||||
@@ -122,18 +125,27 @@ def transcribe(
|
||||
model_size: str = "base",
|
||||
language: Optional[str] = None,
|
||||
progress_cb: Optional[Callable[[float], None]] = None,
|
||||
align: bool = True,
|
||||
) -> Optional[dict]:
|
||||
"""Transcribe an audio/video file locally with word-level timestamps.
|
||||
|
||||
Requires the optional ``[transcribe]`` extra (faster-whisper). Returns
|
||||
``None`` when the model is unavailable or the file is missing/unreadable.
|
||||
|
||||
When ``align`` is true (default) and the optional ``whisperx`` dependency is
|
||||
present, word timestamps are refined by phonetic forced alignment, which
|
||||
corrects faster-whisper's systematic ~0.3-0.5s early bias on word *starts*
|
||||
(see ``Engine/docs/05_EXPERIENCIAS.md`` #14). The transcript reports
|
||||
whether this ran via the ``alignment`` flag, so downstream consumers can
|
||||
rely on the times without re-measuring.
|
||||
|
||||
The model weights are resolved from the configured models directory (see
|
||||
``model_manager.get_models_dir``), so a model selected/downloaded through
|
||||
the app is found without an implicit download to the default HF cache.
|
||||
|
||||
Returns:
|
||||
``{"language": str, "duration": float, "text": str,
|
||||
"alignment": bool,
|
||||
"segments": [{"text", "start", "end", "start_fmt", "end_fmt"}, ...],
|
||||
"words": [{"word", "start", "end", "confidence"}, ...]}``
|
||||
"""
|
||||
@@ -172,6 +184,7 @@ def transcribe(
|
||||
vad_filter=True,
|
||||
)
|
||||
segments: List[dict] = []
|
||||
raw_segments: List[dict] = []
|
||||
words: List[dict] = []
|
||||
# `info.duration` is known upfront (from the container), so each
|
||||
# segment's end time — yielded lazily as faster-whisper decodes —
|
||||
@@ -180,6 +193,20 @@ def transcribe(
|
||||
for seg in segments_iter:
|
||||
start = float(seg.start)
|
||||
end = float(seg.end)
|
||||
seg_words: List[dict] = []
|
||||
if progress_cb is not None and total_duration > 0:
|
||||
progress_cb(min(end / total_duration, 1.0))
|
||||
for w in seg.words or []:
|
||||
ws = float(w.start)
|
||||
we = float(w.end)
|
||||
word = {
|
||||
"word": w.word.strip(),
|
||||
"start": ws,
|
||||
"end": we,
|
||||
"confidence": float(w.probability),
|
||||
}
|
||||
words.append(word)
|
||||
seg_words.append(word)
|
||||
segments.append(
|
||||
{
|
||||
"text": seg.text.strip(),
|
||||
@@ -189,19 +216,26 @@ def transcribe(
|
||||
"end_fmt": format_timestamp(end),
|
||||
}
|
||||
)
|
||||
if progress_cb is not None and total_duration > 0:
|
||||
progress_cb(min(end / total_duration, 1.0))
|
||||
for w in seg.words or []:
|
||||
ws = float(w.start)
|
||||
we = float(w.end)
|
||||
words.append(
|
||||
{
|
||||
"word": w.word.strip(),
|
||||
"start": ws,
|
||||
"end": we,
|
||||
"confidence": float(w.probability),
|
||||
}
|
||||
raw_segments.append(
|
||||
{
|
||||
"text": seg.text.strip(),
|
||||
"start": start,
|
||||
"end": end,
|
||||
"words": seg_words,
|
||||
}
|
||||
)
|
||||
|
||||
alignment_ran = False
|
||||
if align and raw_segments:
|
||||
from .forced_align import ForcedAligner
|
||||
|
||||
try:
|
||||
words = ForcedAligner().align(
|
||||
words, raw_segments, str(file_path), info.language, str(models_dir)
|
||||
)
|
||||
alignment_ran = True
|
||||
except Exception:
|
||||
logger.warning("forced alignment step failed; keeping raw timestamps")
|
||||
except Exception:
|
||||
logger.warning("whisper transcription failed for %s", file_path)
|
||||
return None
|
||||
@@ -209,6 +243,7 @@ def transcribe(
|
||||
"language": info.language,
|
||||
"duration": float(info.duration),
|
||||
"text": " ".join(s["text"] for s in segments),
|
||||
"alignment": alignment_ran,
|
||||
"segments": segments,
|
||||
"words": words,
|
||||
}
|
||||
@@ -281,6 +316,54 @@ def group_words_by_segment(
|
||||
return groups
|
||||
|
||||
|
||||
def split_into_subphrases(
|
||||
words: Sequence[dict],
|
||||
min_words: int = 3,
|
||||
) -> List[List[dict]]:
|
||||
"""Split a sentence's *words* into sub-phrases at comma boundaries.
|
||||
|
||||
A comma is where a spoken sentence actually breathes, so it is the
|
||||
natural seam for grouping subtitles — each sub-phrase becoming its own
|
||||
on-screen block (and, downstream, its own compound clip).
|
||||
|
||||
The exception is the short tail: a fragment like "né?" or "Então..."
|
||||
reads as part of the phrase before it, not as a phrase of its own, and
|
||||
promoting it to its own block would flash a single word on screen. So a
|
||||
piece shorter than *min_words* is merged back into its neighbour —
|
||||
preferring the previous piece, falling back to the next one when the
|
||||
short piece leads the sentence.
|
||||
|
||||
Returns one group per sub-phrase; a sentence with no comma comes back
|
||||
as a single group.
|
||||
"""
|
||||
pieces: List[List[dict]] = []
|
||||
current: List[dict] = []
|
||||
for w in words:
|
||||
current.append(w)
|
||||
text = str(w.get('word') or w.get('text') or '')
|
||||
if text.rstrip().endswith(','):
|
||||
pieces.append(current)
|
||||
current = []
|
||||
if current:
|
||||
pieces.append(current)
|
||||
|
||||
if len(pieces) <= 1:
|
||||
return pieces
|
||||
|
||||
merged: List[List[dict]] = []
|
||||
for piece in pieces:
|
||||
if len(piece) < min_words and merged:
|
||||
merged[-1].extend(piece)
|
||||
else:
|
||||
merged.append(piece)
|
||||
# A short leading piece has no previous neighbour to fold into, so it
|
||||
# folds forward instead.
|
||||
if len(merged) > 1 and len(merged[0]) < min_words:
|
||||
merged[1][:0] = merged[0]
|
||||
merged.pop(0)
|
||||
return merged
|
||||
|
||||
|
||||
def segments_to_srt(segments: Sequence[dict]) -> str:
|
||||
"""Render transcript segments as an SRT string (for captions import)."""
|
||||
|
||||
|
||||
@@ -82,21 +82,36 @@ def _validate_one(raw: Any, index: int) -> Tuple[Optional[VoiceAction], str]:
|
||||
params = dict(params) if isinstance(params, dict) else {}
|
||||
|
||||
if kind == "zoom":
|
||||
try:
|
||||
scale = float(params.get("scale", 1.3))
|
||||
except (TypeError, ValueError):
|
||||
return None, f"{where}: zoom scale must be a number"
|
||||
if not (MIN_ZOOM_SCALE <= scale <= MAX_ZOOM_SCALE):
|
||||
return None, (
|
||||
f"{where}: zoom scale {scale} outside {MIN_ZOOM_SCALE}-{MAX_ZOOM_SCALE}"
|
||||
)
|
||||
params["scale"] = scale
|
||||
if "scale" in params and params.get("scale") is not None:
|
||||
try:
|
||||
scale = float(params["scale"])
|
||||
except (TypeError, ValueError):
|
||||
return None, f"{where}: zoom scale must be a number"
|
||||
if not (MIN_ZOOM_SCALE <= scale <= MAX_ZOOM_SCALE):
|
||||
return None, (
|
||||
f"{where}: zoom scale {scale} outside {MIN_ZOOM_SCALE}-{MAX_ZOOM_SCALE}"
|
||||
)
|
||||
params["scale"] = scale
|
||||
|
||||
if kind == "text":
|
||||
content = str(params.get("content", "")).strip()
|
||||
if not content:
|
||||
return None, f"{where}: text action needs params.content"
|
||||
params["content"] = content[:MAX_TEXT_LENGTH]
|
||||
# Style is optional — omitted fields fall back to the "Legendas
|
||||
# Dinâmicas" emphasis style at apply time (see _apply_placed_action),
|
||||
# so a callout matches the captions' look without the caller having
|
||||
# to know or repeat that configuration. Anything given here wins.
|
||||
for key in ("font", "font_color", "face"):
|
||||
if key in params and not isinstance(params[key], str):
|
||||
del params[key]
|
||||
if "font_size" in params:
|
||||
try:
|
||||
params["font_size"] = int(params["font_size"])
|
||||
except (TypeError, ValueError):
|
||||
del params["font_size"]
|
||||
if "bold" in params:
|
||||
params["bold"] = bool(params["bold"])
|
||||
|
||||
return (
|
||||
VoiceAction(
|
||||
|
||||
@@ -56,12 +56,20 @@ VALUE_SCALES = {
|
||||
"rate_delta": "0-1, how much the local speaking rate departs from the average",
|
||||
"pause_before": "seconds of silence immediately before the word",
|
||||
"emphasis": "0-1 combined index; high values are punch-in/highlight candidates",
|
||||
"emotion": "heuristic label from delivery: neutral, excited, tense, calm, reflective",
|
||||
"emotion_confidence": "0-1 confidence in the heuristic emotion label",
|
||||
"arousal": "0-1 vocal activation from energy/rate/pitch movement",
|
||||
"valence": "0-1 rough positive tone; lower values suggest tension/weight",
|
||||
},
|
||||
"segment": {
|
||||
"gap_before": "seconds of silence before this line",
|
||||
"take_boundary": "true when the gap is long enough that the take likely restarted here",
|
||||
"avg_energy": "0-1 mean loudness across the line",
|
||||
"peak_emphasis": "0-1 highest emphasis of any word in the line",
|
||||
"emotion": "dominant delivery emotion across the line",
|
||||
"emotion_confidence": "0-1 confidence in the dominant segment emotion",
|
||||
"arousal": "0-1 mean vocal activation across the line",
|
||||
"valence": "0-1 mean rough positive tone across the line",
|
||||
},
|
||||
}
|
||||
|
||||
@@ -92,11 +100,74 @@ def _round_word(word: dict) -> dict:
|
||||
"rate_delta": round(word.get("rate_delta", 0.0), 3),
|
||||
"pause_before": round(word.get("pause_before", 0.0), 3),
|
||||
"emphasis": round(word.get("emphasis", 0.0), 3),
|
||||
"emotion": word.get("emotion", "neutral"),
|
||||
"emotion_confidence": round(word.get("emotion_confidence", 0.0), 3),
|
||||
"arousal": round(word.get("arousal", 0.0), 3),
|
||||
"valence": round(word.get("valence", 0.5), 3),
|
||||
"energy_raw": word.get("energy"),
|
||||
"pitch_hz": word.get("pitch_hz"),
|
||||
}
|
||||
|
||||
|
||||
def _emotion_for_word(word: dict, enabled: bool, sensitivity: float) -> dict:
|
||||
"""Classify delivery emotion from normalized acoustic features.
|
||||
|
||||
This is deliberately a local heuristic rather than a claimed clinical
|
||||
emotion model. It gives the editor a useful signal about delivery shape
|
||||
while degrading predictably when acoustic extraction is unavailable.
|
||||
"""
|
||||
if not enabled:
|
||||
return {
|
||||
"emotion": "neutral",
|
||||
"emotion_confidence": 0.0,
|
||||
"arousal": 0.0,
|
||||
"valence": 0.5,
|
||||
}
|
||||
|
||||
energy = float(word.get("energy_norm", 0.0))
|
||||
pitch = float(word.get("pitch_delta", 0.0))
|
||||
rate = float(word.get("rate_delta", 0.0))
|
||||
pause = min(float(word.get("pause_before", 0.0)) / 2.0, 1.0)
|
||||
emphasis = float(word.get("emphasis", 0.0))
|
||||
|
||||
arousal = max(0.0, min(1.0, energy * 0.45 + pitch * 0.25 + rate * 0.20 + emphasis * 0.10))
|
||||
valence = max(0.0, min(1.0, 0.55 + energy * 0.15 - pause * 0.20 - rate * 0.10))
|
||||
|
||||
if arousal >= 0.68 and valence >= 0.50:
|
||||
label = "excited"
|
||||
confidence = arousal
|
||||
elif arousal >= 0.58 and valence < 0.50:
|
||||
label = "tense"
|
||||
confidence = max(arousal, 1.0 - valence)
|
||||
elif arousal <= 0.28 and pause >= 0.25:
|
||||
label = "reflective"
|
||||
confidence = max(1.0 - arousal, pause)
|
||||
elif arousal <= 0.35:
|
||||
label = "calm"
|
||||
confidence = 1.0 - arousal
|
||||
else:
|
||||
label = "neutral"
|
||||
confidence = 1.0 - abs(arousal - 0.5) * 2.0
|
||||
|
||||
confidence = max(0.0, min(1.0, confidence))
|
||||
if confidence < sensitivity:
|
||||
label = "neutral"
|
||||
return {
|
||||
"emotion": label,
|
||||
"emotion_confidence": confidence,
|
||||
"arousal": arousal,
|
||||
"valence": valence,
|
||||
}
|
||||
|
||||
|
||||
def annotate_emotions(words: Sequence[dict], enabled: bool, sensitivity: float) -> List[dict]:
|
||||
"""Attach heuristic emotion labels to enriched word rows."""
|
||||
return [
|
||||
{**w, **_emotion_for_word(w, enabled, sensitivity)}
|
||||
for w in words
|
||||
]
|
||||
|
||||
|
||||
def enrich_words(
|
||||
words: Sequence[dict],
|
||||
pitch_track: Optional[Sequence] = None,
|
||||
@@ -166,6 +237,13 @@ def _segment_rows(segments: Sequence[dict], words: Sequence[dict]) -> List[dict]
|
||||
in_seg = [w for w in words if start <= float(w.get("start", 0.0)) < end]
|
||||
energies = [w["energy_norm"] for w in in_seg]
|
||||
emphases = [w["emphasis"] for w in in_seg]
|
||||
arousals = [w.get("arousal", 0.0) for w in in_seg]
|
||||
valences = [w.get("valence", 0.5) for w in in_seg]
|
||||
emotions = [w.get("emotion", "neutral") for w in in_seg]
|
||||
dominant = max(set(emotions), key=emotions.count) if emotions else "neutral"
|
||||
emotion_confidences = [
|
||||
w.get("emotion_confidence", 0.0) for w in in_seg if w.get("emotion") == dominant
|
||||
]
|
||||
gap = max(0.0, start - previous_end)
|
||||
rows.append(
|
||||
{
|
||||
@@ -181,6 +259,13 @@ def _segment_rows(segments: Sequence[dict], words: Sequence[dict]) -> List[dict]
|
||||
"take_boundary": gap >= TAKE_BOUNDARY_GAP,
|
||||
"avg_energy": round(sum(energies) / len(energies), 3) if energies else 0.0,
|
||||
"peak_emphasis": round(max(emphases), 3) if emphases else 0.0,
|
||||
"emotion": dominant,
|
||||
"emotion_confidence": (
|
||||
round(sum(emotion_confidences) / len(emotion_confidences), 3)
|
||||
if emotion_confidences else 0.0
|
||||
),
|
||||
"arousal": round(sum(arousals) / len(arousals), 3) if arousals else 0.0,
|
||||
"valence": round(sum(valences) / len(valences), 3) if valences else 0.5,
|
||||
"words": [_round_word(w) for w in in_seg],
|
||||
}
|
||||
)
|
||||
@@ -425,6 +510,9 @@ def build_voice_timeline(
|
||||
weights: EmphasisWeights = EmphasisWeights(),
|
||||
peak_percentile: float = 0.02,
|
||||
emphasis_floor: float = 0.25,
|
||||
emotion_enabled: bool = False,
|
||||
emotion_sensitivity: float = 0.5,
|
||||
rotation: float = 0.0,
|
||||
progress_cb: Optional[Callable[[float, str], None]] = None,
|
||||
) -> dict:
|
||||
"""Build the consolidated voice timeline for one media file.
|
||||
@@ -445,6 +533,7 @@ def build_voice_timeline(
|
||||
|
||||
report(0.5, "Calculando ênfase...")
|
||||
words = enrich_words(transcript.get("words", []), pitch_track, energy_track, weights)
|
||||
words = annotate_emotions(words, emotion_enabled, emotion_sensitivity)
|
||||
|
||||
report(0.7, "Identificando participantes...")
|
||||
tracks = diarize(media_path, hf_token, num_speakers) if hf_token else None
|
||||
@@ -457,6 +546,10 @@ def build_voice_timeline(
|
||||
return {
|
||||
"version": VOICE_TIMELINE_VERSION,
|
||||
"source": Path(media_path).name,
|
||||
# Edit-time correction from the clip's Transform filter in the FCPXML
|
||||
# (e.g. straightening a tilted phone shot) — 0.0 when the clip has none
|
||||
# or the caller didn't resolve one.
|
||||
"rotation": rotation,
|
||||
"language": transcript.get("language", ""),
|
||||
# What actually ran, not what was installed — a consumer must be able
|
||||
# to tell "this speech is flat" from "the acoustics never loaded",
|
||||
@@ -465,6 +558,8 @@ def build_voice_timeline(
|
||||
"transcript": bool(transcript.get("words")),
|
||||
"acoustics": pitch_track is not None or energy_track is not None,
|
||||
"speakers": tracks is not None,
|
||||
"emotion": bool(emotion_enabled),
|
||||
"alignment": bool(transcript.get("alignment")),
|
||||
},
|
||||
"scales": VALUE_SCALES,
|
||||
"summary": _summary(
|
||||
|
||||
File diff suppressed because it is too large
Load Diff
@@ -0,0 +1,130 @@
|
||||
"""
|
||||
FCPXML Writer — Generate and modify Final Cut Pro XML files.
|
||||
|
||||
This package provides two complementary workflows for working with FCPXML:
|
||||
|
||||
**Generation** (``FCPXMLWriter``, in :mod:`.generator`):
|
||||
Build a new FCPXML document from Python dataclass objects (``Project``,
|
||||
``Timeline``, ``Clip``, ``Marker``). Useful for creating rough cuts,
|
||||
montage exports, and template-based projects.
|
||||
|
||||
**Modification** (``FCPXMLModifier``, in :mod:`.modifier`):
|
||||
Load an existing FCPXML file, apply surgical edits (markers, trims,
|
||||
reorders, transitions, speed changes, silence removal, etc.), and save.
|
||||
This is the primary API used by the MCP server's tool handlers.
|
||||
|
||||
Layout
|
||||
------
|
||||
This was one 4.200-line module. It is now one module per subject, because the
|
||||
subjects barely touch each other: whoever is fixing a zoom ramp has no reason
|
||||
to scroll past subtitle layout to find it.
|
||||
|
||||
helpers sanitising, scales, shared element builders
|
||||
document asset creation, timebases, serialisation (``write_fcpxml``)
|
||||
validation structural checks (``validate_fcpxml``)
|
||||
core ``ModifierCore``: load, indices, spine navigation, ``save``
|
||||
<subject> one mixin per editing subject (markers, trim, speed, …)
|
||||
modifier ``FCPXMLModifier`` = core + every mixin
|
||||
generator ``FCPXMLWriter``
|
||||
api one-line convenience wrappers
|
||||
|
||||
Everything the rest of the project imported from the old module is re-exported
|
||||
here, so ``from fcpxml.writer import FCPXMLModifier`` keeps working unchanged —
|
||||
including the underscore-prefixed helpers the test suite reaches for.
|
||||
|
||||
Architecture notes
|
||||
------------------
|
||||
- All time arithmetic uses ``TimeValue`` (rational fractions) — never floats —
|
||||
to match FCPXML's native ``"600/2400s"`` format and avoid rounding drift.
|
||||
- The ``FCPXMLModifier`` builds three in-memory indices at init
|
||||
(``clips``, ``resources``, ``formats``) so lookups are O(1) by ID/name.
|
||||
- Spine-based editing: clips live inside a ``<spine>`` element (the primary
|
||||
storyline). Connected clips attach via ``lane`` attributes on spine clips.
|
||||
Most editing methods find the target clip in the spine, mutate it, then
|
||||
ripple offsets on subsequent siblings.
|
||||
- ``write_fcpxml()`` handles DTD-compliant serialisation and optional
|
||||
timebase enforcement for all output paths.
|
||||
"""
|
||||
|
||||
from ..models import TimeValue
|
||||
from .api import add_marker_to_file, modify_fcpxml, trim_clip_in_file
|
||||
from .core import ModifierCore
|
||||
from .document import (
|
||||
_STILL_IMAGE_EXTENSIONS,
|
||||
_enforce_standard_timebases,
|
||||
_ensure_video_asset,
|
||||
write_fcpxml,
|
||||
)
|
||||
from .generator import FCPXMLWriter
|
||||
from .helpers import (
|
||||
_ASSET_CLIP_CHILD_ORDER,
|
||||
_CHILD_ORDER_INDEX,
|
||||
_MAX_MARKER_NAME_LENGTH,
|
||||
_MAX_NOTE_LENGTH,
|
||||
CLIP_AND_AUDIO_TAGS,
|
||||
CLIP_TAGS,
|
||||
FCP_EFFECTS,
|
||||
HOLD_AT_CUT_THRESHOLD,
|
||||
SPINE_ELEMENT_TAGS,
|
||||
START_AT_CUT_THRESHOLD,
|
||||
_create_asset_element,
|
||||
_dtd_insert,
|
||||
_fmt_scale,
|
||||
_probe_audio_info,
|
||||
_sanitize_xml_value,
|
||||
build_marker_element,
|
||||
list_effects,
|
||||
)
|
||||
from .modifier import FCPXMLModifier
|
||||
from .validation import (
|
||||
_check_asset_sources,
|
||||
_check_child_order,
|
||||
_check_effect_refs,
|
||||
_check_frame_alignment,
|
||||
_check_required_attributes,
|
||||
_check_timebases,
|
||||
_document_frame_duration,
|
||||
validate_fcpxml,
|
||||
)
|
||||
|
||||
__all__ = [
|
||||
"FCPXMLModifier",
|
||||
"FCPXMLWriter",
|
||||
"ModifierCore",
|
||||
"TimeValue",
|
||||
"FCP_EFFECTS",
|
||||
"CLIP_TAGS",
|
||||
"CLIP_AND_AUDIO_TAGS",
|
||||
"SPINE_ELEMENT_TAGS",
|
||||
"HOLD_AT_CUT_THRESHOLD",
|
||||
"START_AT_CUT_THRESHOLD",
|
||||
"add_marker_to_file",
|
||||
"build_marker_element",
|
||||
"list_effects",
|
||||
"modify_fcpxml",
|
||||
"trim_clip_in_file",
|
||||
"validate_fcpxml",
|
||||
"write_fcpxml",
|
||||
# Internos que o resto do projeto (e a suíte) já importava deste módulo
|
||||
# quando ele era um arquivo só. Ficam aqui para a divisão não virar uma
|
||||
# quebra de API disfarçada de reorganização.
|
||||
"_ASSET_CLIP_CHILD_ORDER",
|
||||
"_CHILD_ORDER_INDEX",
|
||||
"_MAX_MARKER_NAME_LENGTH",
|
||||
"_MAX_NOTE_LENGTH",
|
||||
"_STILL_IMAGE_EXTENSIONS",
|
||||
"_check_asset_sources",
|
||||
"_check_child_order",
|
||||
"_check_effect_refs",
|
||||
"_check_frame_alignment",
|
||||
"_check_required_attributes",
|
||||
"_check_timebases",
|
||||
"_create_asset_element",
|
||||
"_document_frame_duration",
|
||||
"_dtd_insert",
|
||||
"_enforce_standard_timebases",
|
||||
"_ensure_video_asset",
|
||||
"_fmt_scale",
|
||||
"_probe_audio_info",
|
||||
"_sanitize_xml_value",
|
||||
]
|
||||
@@ -0,0 +1,140 @@
|
||||
"""Clip de ajuste (adjustment layer) — criação do elemento FCPXML.
|
||||
|
||||
No Final Cut, uma "camada de ajuste" é um ``<clip>`` que carrega filtros
|
||||
(``filter-video`` / ``filter-audio``) diretamente como filhos — o DTD do
|
||||
FCPXML 1.13 não define nenhum elemento ``<adjustment>`` como wrapper (ver
|
||||
``<!ELEMENT clip>`` em ``FCPXMLv1_13.dtd``: ``filter-video``/``filter-audio``
|
||||
vêm depois de ``audio-channel-source*`` e antes de ``metadata?``, sem
|
||||
elemento intermediário). Tudo que está abaixo do clip na timeline herda
|
||||
esses filtros — é como se o efeito fosse aplicado a uma faixa inteira de
|
||||
uma vez.
|
||||
|
||||
Esta classe monta esse elemento a partir de dados de alto nível (duração +
|
||||
lista de ``EfeitoAjuste``), cuidando de criar os recursos ``<effect>``
|
||||
correspondentes na seção ``<resources>`` e de referenciá-los pelos filtros.
|
||||
"""
|
||||
|
||||
import xml.etree.ElementTree as ET
|
||||
from typing import Callable, List, Optional
|
||||
|
||||
from ..models.timeline import EfeitoAjuste
|
||||
from ..models.timing import TimeValue
|
||||
|
||||
|
||||
def _para_racional(tempo) -> str:
|
||||
"""Aceita ``TimeValue`` ou uma string FCPXML já formatada ("90/30s")."""
|
||||
if isinstance(tempo, TimeValue):
|
||||
return tempo.to_fcpxml()
|
||||
if tempo is None:
|
||||
return "0/1s"
|
||||
return str(tempo)
|
||||
|
||||
|
||||
def _id_recurso_unico(resources: ET.Element, prefixo: str = "r_ajuste") -> str:
|
||||
"""Gera um ``id`` de recurso ainda ausente em ``resources``."""
|
||||
existentes = {r.get("id") for r in resources.findall("*") if r.get("id")}
|
||||
contador = 1
|
||||
while f"{prefixo}_{contador}" in existentes:
|
||||
contador += 1
|
||||
return f"{prefixo}_{contador}"
|
||||
|
||||
|
||||
class ClipDeAjuste:
|
||||
"""Cria um clip de ajuste (adjustment layer) pronto para a spine.
|
||||
|
||||
Exemplo::
|
||||
|
||||
from fcpxml.models.timing import TimeValue
|
||||
from fcpxml.models.timeline import EfeitoAjuste, ParametroEfeito
|
||||
from fcpxml.writer.adjustment import ClipDeAjuste
|
||||
|
||||
efeito = EfeitoAjuste(
|
||||
nome="Color Curves", uid="...UUID...", tipo="video",
|
||||
parametros=[ParametroEfeito(nome="Amount", valor="0.5",
|
||||
chave=".../9999")],
|
||||
)
|
||||
clip = ClipDeAjuste(
|
||||
nome="Ajuste de cor",
|
||||
duracao=TimeValue(300, 30),
|
||||
efeitos=[efeito],
|
||||
).criar(resources)
|
||||
spine.append(clip)
|
||||
"""
|
||||
|
||||
def __init__(
|
||||
self,
|
||||
nome: str,
|
||||
duracao,
|
||||
efeitos: List[EfeitoAjuste],
|
||||
offset=None,
|
||||
formato_tc: str = "NDF",
|
||||
):
|
||||
self.nome = nome
|
||||
self.duracao = duracao
|
||||
self.efeitos = efeitos
|
||||
self.offset = offset
|
||||
self.formato_tc = formato_tc
|
||||
|
||||
def criar(
|
||||
self,
|
||||
resources: ET.Element,
|
||||
proximo_id: Optional[Callable[[], str]] = None,
|
||||
) -> ET.Element:
|
||||
"""Monta o ``<clip>`` de ajuste e seus recursos ``<effect>``.
|
||||
|
||||
``resources`` é a seção ``<resources>`` do documento (onde os
|
||||
``<effect>`` são registrados). ``proximo_id`` é um gerador opcional
|
||||
de ids de recurso; sem ele, usa um id único baseado em ``resources``.
|
||||
"""
|
||||
def gerar_id() -> str:
|
||||
if proximo_id:
|
||||
return proximo_id()
|
||||
return _id_recurso_unico(resources)
|
||||
|
||||
filtros: List[ET.Element] = []
|
||||
for efeito in self.efeitos:
|
||||
efeito_id = self._garantir_recurso(resources, efeito, gerar_id)
|
||||
filtros.append(self._montar_filtro(efeito, efeito_id))
|
||||
|
||||
clip = ET.Element(
|
||||
"clip",
|
||||
name=self.nome,
|
||||
duration=_para_racional(self.duracao),
|
||||
tcFormat=self.formato_tc,
|
||||
)
|
||||
if self.offset is not None:
|
||||
clip.set("offset", _para_racional(self.offset))
|
||||
|
||||
# O DTD exige filter-video* antes de filter-audio* como filhos
|
||||
# diretos do clip (sem wrapper <adjustment>).
|
||||
for filtro in sorted(filtros, key=lambda f: f.tag != "filter-video"):
|
||||
clip.append(filtro)
|
||||
return clip
|
||||
|
||||
def _garantir_recurso(
|
||||
self, resources: ET.Element, efeito: EfeitoAjuste, gerar_id: Callable[[], str]
|
||||
) -> str:
|
||||
"""Devolve o ``id`` do ``<effect>`` de *efeito*, criando-o se ausente."""
|
||||
for existente in resources.findall("effect"):
|
||||
if existente.get("uid") == efeito.uid:
|
||||
return existente.get("id")
|
||||
efeito_id = gerar_id()
|
||||
recurso = ET.SubElement(resources, "effect")
|
||||
recurso.set("id", efeito_id)
|
||||
recurso.set("name", efeito.nome)
|
||||
recurso.set("uid", efeito.uid)
|
||||
return efeito_id
|
||||
|
||||
def _montar_filtro(self, efeito: EfeitoAjuste, efeito_id: str) -> ET.Element:
|
||||
"""Monta o ``<filter-video>``/``<filter-audio>`` de um efeito."""
|
||||
tag = "filter-video" if efeito.tipo == "video" else "filter-audio"
|
||||
filtro = ET.Element(tag, ref=efeito_id, name=efeito.nome)
|
||||
for parametro in efeito.parametros:
|
||||
param = ET.SubElement(filtro, "param")
|
||||
param.set("name", parametro.nome)
|
||||
if parametro.chave:
|
||||
param.set("key", parametro.chave)
|
||||
param.set("value", parametro.valor)
|
||||
if parametro.metadado:
|
||||
param.set("metadata", parametro.metadado)
|
||||
return filtro
|
||||
@@ -0,0 +1,55 @@
|
||||
"""Atalhos de uma linha para as operações mais comuns.
|
||||
|
||||
Extraído de writer.py — ver fcpxml/writer/__init__.py para o conjunto.
|
||||
"""
|
||||
|
||||
from typing import Optional
|
||||
|
||||
from ..models import (
|
||||
MarkerType,
|
||||
)
|
||||
from .modifier import FCPXMLModifier
|
||||
|
||||
# ============================================================================
|
||||
# CONVENIENCE FUNCTIONS
|
||||
# ============================================================================
|
||||
|
||||
def modify_fcpxml(filepath: str) -> FCPXMLModifier:
|
||||
"""
|
||||
Open an FCPXML file for modification.
|
||||
|
||||
Usage:
|
||||
modifier = modify_fcpxml("project.fcpxml")
|
||||
modifier.add_marker(...)
|
||||
modifier.save("output.fcpxml")
|
||||
"""
|
||||
return FCPXMLModifier(filepath)
|
||||
|
||||
|
||||
def add_marker_to_file(
|
||||
filepath: str,
|
||||
timecode: str,
|
||||
name: str,
|
||||
marker_type: str = "standard",
|
||||
output_path: Optional[str] = None
|
||||
) -> str:
|
||||
"""Convenience function to add a marker to an FCPXML file."""
|
||||
modifier = FCPXMLModifier(filepath)
|
||||
modifier.add_marker_at_timeline(
|
||||
timecode, name,
|
||||
MarkerType.from_string(marker_type)
|
||||
)
|
||||
return modifier.save(output_path)
|
||||
|
||||
|
||||
def trim_clip_in_file(
|
||||
filepath: str,
|
||||
clip_id: str,
|
||||
trim_start: Optional[str] = None,
|
||||
trim_end: Optional[str] = None,
|
||||
output_path: Optional[str] = None
|
||||
) -> str:
|
||||
"""Convenience function to trim a clip in an FCPXML file."""
|
||||
modifier = FCPXMLModifier(filepath)
|
||||
modifier.trim_clip(clip_id, trim_start, trim_end)
|
||||
return modifier.save(output_path)
|
||||
@@ -0,0 +1,162 @@
|
||||
"""Clipes de áudio e cama musical.
|
||||
|
||||
Extraído de writer.py — ver fcpxml/writer/__init__.py para o conjunto.
|
||||
"""
|
||||
|
||||
import xml.etree.ElementTree as ET
|
||||
from pathlib import Path
|
||||
from typing import Optional
|
||||
|
||||
from ..models import (
|
||||
TimeValue,
|
||||
)
|
||||
from .helpers import _create_asset_element, _dtd_insert, _probe_audio_info, _sanitize_xml_value
|
||||
|
||||
|
||||
class AudioMixin:
|
||||
"""Clipes de áudio e cama musical."""
|
||||
|
||||
# AUDIO CLIP OPERATIONS (v0.6.0)
|
||||
# ========================================================================
|
||||
|
||||
def add_audio_clip(
|
||||
self,
|
||||
parent_clip_id: str,
|
||||
asset_id: Optional[str] = None,
|
||||
offset: str = "0s",
|
||||
duration: Optional[str] = None,
|
||||
role: str = "dialogue",
|
||||
lane: int = -1,
|
||||
src: Optional[str] = None,
|
||||
) -> ET.Element:
|
||||
"""Add an audio clip connected to an existing timeline clip.
|
||||
|
||||
Creates an <asset-clip> at a negative lane with audioRole attribute.
|
||||
Supports hierarchical roles like "dialogue.boom", "music.score",
|
||||
"effects.foley".
|
||||
|
||||
Args:
|
||||
parent_clip_id: Name/ID of the clip to attach audio to.
|
||||
asset_id: Existing asset reference ID. If None and src provided,
|
||||
creates a new asset.
|
||||
offset: Position relative to parent clip start.
|
||||
duration: Duration of audio clip.
|
||||
role: Audio role (e.g. "dialogue", "music.score", "effects.foley").
|
||||
lane: Lane number (negative = below primary, default -1).
|
||||
src: Path to audio file. Used to create a new asset if asset_id
|
||||
is not provided.
|
||||
|
||||
Returns:
|
||||
The created audio clip element.
|
||||
"""
|
||||
parent = self._require_clip(parent_clip_id)
|
||||
|
||||
# Resolve or create asset
|
||||
if asset_id and asset_id in self.resources:
|
||||
asset = self.resources[asset_id]
|
||||
elif src:
|
||||
# Create new asset in resources
|
||||
resources = self.root.find('.//resources')
|
||||
if resources is None:
|
||||
raise ValueError("No <resources> element found in FCPXML")
|
||||
asset_id = self._unique_resource_id(resources, 'r_audio1')
|
||||
# The asset duration must reflect the real media length, not the
|
||||
# requested clip duration — FCP flags assets that claim more
|
||||
# media than the file contains.
|
||||
probed = _probe_audio_info(src)
|
||||
if probed:
|
||||
rate = probed['sample_rate']
|
||||
asset_duration = f"{round(probed['duration'] * rate)}/{rate}s"
|
||||
else:
|
||||
asset_duration = duration or "0s"
|
||||
asset_elem = _create_asset_element(
|
||||
resources, asset_id, Path(src).stem, src,
|
||||
duration=asset_duration,
|
||||
has_video="0", has_audio="1",
|
||||
)
|
||||
if probed:
|
||||
asset_elem.set('audioSources', '1')
|
||||
asset_elem.set('audioChannels', str(probed['channels']))
|
||||
asset_elem.set('audioRate', str(probed['sample_rate']))
|
||||
asset = {
|
||||
'id': asset_id,
|
||||
'name': Path(src).stem,
|
||||
'duration': asset_duration,
|
||||
'element': asset_elem,
|
||||
}
|
||||
self.resources[asset_id] = asset
|
||||
else:
|
||||
raise ValueError("Must provide either asset_id or src for audio clip")
|
||||
|
||||
clip_duration, source_start = self._resolve_clip_duration(asset, duration)
|
||||
|
||||
# Clamp so the clip never claims more media than the asset contains
|
||||
asset_duration_tv = self._parse_time(asset.get('duration', '0s'))
|
||||
if asset_duration_tv > TimeValue.zero():
|
||||
available = asset_duration_tv - source_start
|
||||
if available < TimeValue.zero():
|
||||
raise ValueError(
|
||||
f"Source start {source_start.to_fcpxml()} is beyond the end "
|
||||
f"of audio asset '{asset.get('name')}' "
|
||||
f"({asset_duration_tv.to_fcpxml()})"
|
||||
)
|
||||
if clip_duration > available:
|
||||
clip_duration = available
|
||||
|
||||
new_clip = self._make_asset_clip(
|
||||
asset_id, asset.get('name', 'Audio'),
|
||||
self._parse_time(offset), source_start, clip_duration,
|
||||
lane=str(lane),
|
||||
audioRole=_sanitize_xml_value(role, 256),
|
||||
)
|
||||
_dtd_insert(parent, new_clip)
|
||||
return new_clip
|
||||
|
||||
def add_music_bed(
|
||||
self,
|
||||
asset_id: Optional[str] = None,
|
||||
duration: Optional[str] = None,
|
||||
role: str = "music",
|
||||
src: Optional[str] = None,
|
||||
) -> ET.Element:
|
||||
"""Add a music bed spanning the full timeline at lane -1.
|
||||
|
||||
Convenience method: attaches to the first spine clip and spans
|
||||
the full timeline duration.
|
||||
|
||||
Args:
|
||||
asset_id: Existing asset reference ID.
|
||||
duration: Override duration (default: full timeline).
|
||||
role: Audio role (default "music").
|
||||
src: Path to audio file (creates asset if asset_id not given).
|
||||
|
||||
Returns:
|
||||
The created music bed clip element.
|
||||
"""
|
||||
spine = self._get_spine()
|
||||
first_clip = None
|
||||
first_clip_id = None
|
||||
for clip_id, clip in self.clips.items():
|
||||
if clip in list(spine):
|
||||
first_clip = clip
|
||||
first_clip_id = clip_id
|
||||
break
|
||||
|
||||
if first_clip is None:
|
||||
raise ValueError("No clips in spine to attach music bed to")
|
||||
|
||||
# Calculate full timeline duration if not specified
|
||||
if not duration:
|
||||
duration = self._timeline_duration().to_fcpxml()
|
||||
|
||||
return self.add_audio_clip(
|
||||
parent_clip_id=first_clip_id,
|
||||
asset_id=asset_id,
|
||||
offset="0s",
|
||||
duration=duration,
|
||||
role=role,
|
||||
lane=-1,
|
||||
src=src,
|
||||
)
|
||||
|
||||
# ========================================================================
|
||||
@@ -0,0 +1,290 @@
|
||||
"""Compound clips: criar e achatar.
|
||||
|
||||
Extraído de writer.py — ver fcpxml/writer/__init__.py para o conjunto.
|
||||
"""
|
||||
|
||||
import copy
|
||||
import uuid
|
||||
import xml.etree.ElementTree as ET
|
||||
from typing import List
|
||||
|
||||
from ..models import (
|
||||
TimeValue,
|
||||
)
|
||||
from .helpers import (
|
||||
_dtd_insert,
|
||||
_sanitize_xml_value,
|
||||
)
|
||||
|
||||
|
||||
class CompoundMixin:
|
||||
"""Compound clips: criar e achatar."""
|
||||
|
||||
# COMPOUND CLIP OPERATIONS (v0.6.0)
|
||||
# ========================================================================
|
||||
|
||||
def create_compound_clip(
|
||||
self,
|
||||
clip_ids: List[str],
|
||||
name: str = "Compound Clip",
|
||||
) -> ET.Element:
|
||||
"""Group spine clips into a compound clip.
|
||||
|
||||
Creates a <media> resource with a nested <sequence><spine> containing
|
||||
the specified clips, then replaces the originals in the main spine
|
||||
with a single <ref-clip>.
|
||||
|
||||
Args:
|
||||
clip_ids: IDs of clips in the spine to group.
|
||||
name: Name for the compound clip.
|
||||
|
||||
Returns:
|
||||
The created <ref-clip> element.
|
||||
"""
|
||||
spine = self._get_spine()
|
||||
resources = self.root.find('.//resources')
|
||||
if resources is None:
|
||||
raise ValueError("No <resources> element found in FCPXML")
|
||||
|
||||
# Collect clips and validate they're in spine
|
||||
spine_children = list(spine)
|
||||
clips_to_group = []
|
||||
for cid in clip_ids:
|
||||
clip = self._require_clip(cid)
|
||||
if clip not in spine_children:
|
||||
raise ValueError(f"Clip not in spine: {cid}")
|
||||
clips_to_group.append((cid, clip))
|
||||
|
||||
if not clips_to_group:
|
||||
raise ValueError("No valid clips to group")
|
||||
|
||||
# Sort by offset so the compound maintains order
|
||||
clips_to_group.sort(
|
||||
key=lambda c: self._parse_time(c[1].get('offset', '0s'))
|
||||
)
|
||||
|
||||
# Calculate compound duration and starting offset
|
||||
first_offset = self._parse_time(clips_to_group[0][1].get('offset', '0s'))
|
||||
total_duration = TimeValue.zero()
|
||||
for _, clip in clips_to_group:
|
||||
total_duration = total_duration + self._parse_time(clip.get('duration', '0s'))
|
||||
|
||||
# Get format ref
|
||||
format_id = None
|
||||
for fmt_id in self.formats:
|
||||
format_id = fmt_id
|
||||
break
|
||||
|
||||
# Create media resource with nested sequence
|
||||
media_id = self._unique_resource_id(resources, 'r_compound1')
|
||||
|
||||
media = ET.SubElement(resources, 'media')
|
||||
media.set('id', media_id)
|
||||
media.set('name', _sanitize_xml_value(name, 512))
|
||||
media.set('uid', str(uuid.uuid4()).upper())
|
||||
|
||||
seq = ET.SubElement(media, 'sequence')
|
||||
seq.set('format', format_id or 'r1')
|
||||
seq.set('duration', total_duration.to_fcpxml())
|
||||
seq.set('tcStart', '0s')
|
||||
seq.set('tcFormat', 'NDF')
|
||||
|
||||
inner_spine = ET.SubElement(seq, 'spine')
|
||||
|
||||
# Move clips into the compound's inner spine
|
||||
inner_offset = TimeValue.zero()
|
||||
for _, clip in clips_to_group:
|
||||
new_clip = copy.deepcopy(clip)
|
||||
new_clip.set('offset', inner_offset.to_fcpxml())
|
||||
inner_spine.append(new_clip)
|
||||
inner_offset = inner_offset + self._parse_time(clip.get('duration', '0s'))
|
||||
|
||||
# Get the insert position (where first clip was)
|
||||
spine_children = list(spine)
|
||||
insert_idx = spine_children.index(clips_to_group[0][1])
|
||||
|
||||
# Remove originals from spine
|
||||
for cid, clip in clips_to_group:
|
||||
spine.remove(clip)
|
||||
if cid in self.clips:
|
||||
del self.clips[cid]
|
||||
|
||||
# Create ref-clip in main spine
|
||||
ref_clip = ET.Element('ref-clip')
|
||||
ref_clip.set('ref', media_id)
|
||||
ref_clip.set('offset', first_offset.to_fcpxml())
|
||||
ref_clip.set('name', _sanitize_xml_value(name, 512))
|
||||
ref_clip.set('duration', total_duration.to_fcpxml())
|
||||
spine.insert(insert_idx, ref_clip)
|
||||
|
||||
# Index the new ref-clip
|
||||
compound_id = f"compound_{name}"
|
||||
self.clips[compound_id] = ref_clip
|
||||
|
||||
return ref_clip
|
||||
|
||||
def wrap_titles_in_compound(
|
||||
self,
|
||||
parent_clip: ET.Element,
|
||||
titles: List[ET.Element],
|
||||
name: str = "Legenda",
|
||||
) -> ET.Element:
|
||||
"""Pack lane-nested *titles* of *parent_clip* into one compound clip.
|
||||
|
||||
A dynamic-subtitle sub-phrase is a dozen overlapping ``<title>``
|
||||
elements stacked across as many lanes — legible on screen, unreadable
|
||||
in the timeline. Collapsing each sub-phrase into a single compound
|
||||
gives one bar per phrase to drag, mute or retime as a unit.
|
||||
|
||||
Mirrors the structure Final Cut itself produces for "New Compound
|
||||
Clip" over stacked titles: the earliest title becomes the compound's
|
||||
spine anchor at offset 0, the rest hang off it as lane children, and
|
||||
a ``<ref-clip>`` takes their place in *parent_clip* on the anchor's
|
||||
original lane.
|
||||
|
||||
Child offsets are rebased from *parent_clip*'s source-time space onto
|
||||
the anchor's, since a lane child is anchored at its parent's
|
||||
``start`` — leaving them untouched would shift every word of the
|
||||
phrase by the gap between the two starts.
|
||||
|
||||
Args:
|
||||
parent_clip: The spine clip the titles currently hang off.
|
||||
titles: The ``<title>`` elements to pack; must all be direct
|
||||
children of *parent_clip*.
|
||||
name: Name for the resulting compound clip.
|
||||
|
||||
Returns:
|
||||
The created ``<ref-clip>`` element, now in *parent_clip*.
|
||||
"""
|
||||
if not titles:
|
||||
raise ValueError("No titles to wrap")
|
||||
resources = self.root.find('.//resources')
|
||||
if resources is None:
|
||||
raise ValueError("No <resources> element found in FCPXML")
|
||||
|
||||
ordered = sorted(
|
||||
titles, key=lambda t: self._parse_time(t.get('offset', '0s'))
|
||||
)
|
||||
anchor = ordered[0]
|
||||
anchor_offset = self._parse_time(anchor.get('offset', '0s'))
|
||||
anchor_start = self._parse_time(anchor.get('start', '0s'))
|
||||
anchor_lane = anchor.get('lane')
|
||||
|
||||
total = TimeValue.zero()
|
||||
for title in ordered:
|
||||
rel = self._parse_time(title.get('offset', '0s')) - anchor_offset
|
||||
end = rel + self._parse_time(title.get('duration', '0s'))
|
||||
if end > total:
|
||||
total = end
|
||||
|
||||
format_id = next(iter(self.formats), None) or 'r1'
|
||||
media_id = self._unique_resource_id(resources, 'r_compound1')
|
||||
|
||||
media = ET.SubElement(resources, 'media')
|
||||
media.set('id', media_id)
|
||||
media.set('name', _sanitize_xml_value(name, 512))
|
||||
media.set('uid', str(uuid.uuid4()).upper())
|
||||
|
||||
seq = ET.SubElement(media, 'sequence')
|
||||
seq.set('format', format_id)
|
||||
seq.set('duration', total.to_fcpxml())
|
||||
seq.set('tcStart', '0s')
|
||||
seq.set('tcFormat', 'NDF')
|
||||
inner_spine = ET.SubElement(seq, 'spine')
|
||||
|
||||
for title in ordered:
|
||||
parent_clip.remove(title)
|
||||
|
||||
anchor.set('offset', '0s')
|
||||
if anchor_lane is not None:
|
||||
del anchor.attrib['lane']
|
||||
inner_spine.append(anchor)
|
||||
|
||||
for lane, title in enumerate(ordered[1:], start=1):
|
||||
rel = self._parse_time(title.get('offset', '0s')) - anchor_offset
|
||||
title.set('offset', (anchor_start + rel).to_fcpxml())
|
||||
title.set('lane', str(lane))
|
||||
anchor.append(title)
|
||||
|
||||
ref_clip = ET.Element('ref-clip')
|
||||
ref_clip.set('ref', media_id)
|
||||
if anchor_lane is not None:
|
||||
ref_clip.set('lane', anchor_lane)
|
||||
ref_clip.set('offset', anchor_offset.to_fcpxml())
|
||||
ref_clip.set('name', _sanitize_xml_value(name, 512))
|
||||
ref_clip.set('duration', total.to_fcpxml())
|
||||
_dtd_insert(parent_clip, ref_clip)
|
||||
return ref_clip
|
||||
|
||||
def flatten_compound_clip(
|
||||
self,
|
||||
ref_clip_id: str,
|
||||
) -> List[ET.Element]:
|
||||
"""Flatten a compound clip back into individual spine clips.
|
||||
|
||||
Extracts clips from the compound's inner sequence and places them
|
||||
back in the main spine at the ref-clip's position.
|
||||
|
||||
Args:
|
||||
ref_clip_id: ID of the ref-clip to flatten.
|
||||
|
||||
Returns:
|
||||
List of extracted clip elements now in the main spine.
|
||||
"""
|
||||
spine = self._get_spine()
|
||||
ref_clip = self._require_clip(ref_clip_id)
|
||||
if ref_clip.tag != 'ref-clip':
|
||||
raise ValueError(f"Element is not a ref-clip: {ref_clip_id}")
|
||||
|
||||
media_ref = ref_clip.get('ref', '')
|
||||
ref_offset = self._parse_time(ref_clip.get('offset', '0s'))
|
||||
|
||||
# Find the media resource
|
||||
resources = self.root.find('.//resources')
|
||||
media_elem = None
|
||||
if resources is not None:
|
||||
for m in resources.findall('media'):
|
||||
if m.get('id') == media_ref:
|
||||
media_elem = m
|
||||
break
|
||||
|
||||
if media_elem is None:
|
||||
raise ValueError(f"Media resource not found for ref: {media_ref}")
|
||||
|
||||
inner_spine = media_elem.find('.//spine')
|
||||
if inner_spine is None:
|
||||
raise ValueError("No spine found in compound clip media")
|
||||
|
||||
# Get insert position
|
||||
spine_children = list(spine)
|
||||
insert_idx = spine_children.index(ref_clip)
|
||||
|
||||
# Remove ref-clip from spine
|
||||
spine.remove(ref_clip)
|
||||
if ref_clip_id in self.clips:
|
||||
del self.clips[ref_clip_id]
|
||||
|
||||
# Extract clips from inner spine into main spine
|
||||
extracted = []
|
||||
current_offset = ref_offset
|
||||
for child in list(inner_spine):
|
||||
new_clip = copy.deepcopy(child)
|
||||
new_clip.set('offset', current_offset.to_fcpxml())
|
||||
spine.insert(insert_idx, new_clip)
|
||||
insert_idx += 1
|
||||
extracted.append(new_clip)
|
||||
current_offset = current_offset + self._parse_time(
|
||||
child.get('duration', '0s')
|
||||
)
|
||||
|
||||
# Index the extracted clip
|
||||
clip_name = new_clip.get('name') or new_clip.get('id') or f"flat_{len(self.clips)}"
|
||||
self.clips[clip_name] = new_clip
|
||||
|
||||
# Clean up media resource
|
||||
if resources is not None:
|
||||
resources.remove(media_elem)
|
||||
|
||||
return extracted
|
||||
|
||||
# ========================================================================
|
||||
@@ -0,0 +1,49 @@
|
||||
"""Clipes conectados (lanes acima/abaixo da spine).
|
||||
|
||||
Extraído de writer.py — ver fcpxml/writer/__init__.py para o conjunto.
|
||||
"""
|
||||
|
||||
import xml.etree.ElementTree as ET
|
||||
from typing import Optional
|
||||
|
||||
|
||||
class ConnectedMixin:
|
||||
"""Clipes conectados (lanes acima/abaixo da spine)."""
|
||||
|
||||
# CONNECTED CLIP OPERATIONS (v0.5.0)
|
||||
# ========================================================================
|
||||
|
||||
def add_connected_clip(
|
||||
self,
|
||||
parent_clip_id: str,
|
||||
asset_id: Optional[str] = None,
|
||||
asset_name: Optional[str] = None,
|
||||
offset: str = "0s",
|
||||
duration: Optional[str] = None,
|
||||
lane: int = 1,
|
||||
) -> ET.Element:
|
||||
"""Add a connected clip (B-roll, title, audio) to an existing timeline clip.
|
||||
|
||||
Args:
|
||||
parent_clip_id: Name/ID of the clip to attach to
|
||||
asset_id: Asset reference ID
|
||||
asset_name: Asset name (alternative to asset_id)
|
||||
offset: Position relative to parent clip start
|
||||
duration: Duration of connected clip (default: full asset)
|
||||
lane: Lane number (positive=above, negative=below)
|
||||
|
||||
Returns:
|
||||
The created connected clip element
|
||||
"""
|
||||
parent = self._require_clip(parent_clip_id)
|
||||
asset, asset_id = self._resolve_asset(asset_id, asset_name)
|
||||
clip_duration, source_start = self._resolve_clip_duration(asset, duration)
|
||||
|
||||
new_clip = self._make_asset_clip(
|
||||
asset_id, asset.get('name', 'Untitled'),
|
||||
self._parse_time(offset), source_start, clip_duration,
|
||||
parent=parent, lane=str(lane),
|
||||
)
|
||||
return new_clip
|
||||
|
||||
# ========================================================================
|
||||
@@ -0,0 +1,725 @@
|
||||
"""Núcleo do FCPXMLModifier: carga, índices, navegação na spine e save.
|
||||
|
||||
Extraído de writer.py — ver fcpxml/writer/__init__.py para o conjunto.
|
||||
"""
|
||||
|
||||
import xml.etree.ElementTree as ET
|
||||
from fractions import Fraction
|
||||
from pathlib import Path
|
||||
from typing import Any, Dict, Optional, Tuple
|
||||
|
||||
from ..models import (
|
||||
TimeValue,
|
||||
)
|
||||
from .document import write_fcpxml
|
||||
from .helpers import CLIP_TAGS
|
||||
|
||||
|
||||
class ModifierCore:
|
||||
"""Load an existing FCPXML file, apply edits, and save.
|
||||
|
||||
This is the primary editing interface used by every MCP server write-tool
|
||||
handler. It wraps an ElementTree parsed from disk and maintains three
|
||||
in-memory indices so that clip/asset lookups are fast.
|
||||
|
||||
Index design
|
||||
------------
|
||||
``clips`` : ``Dict[str, ET.Element]``
|
||||
Every ``<clip>``, ``<asset-clip>``, and ``<video>`` element keyed by
|
||||
its ``id`` attribute, falling back to ``name``, then a generated key.
|
||||
**Gotcha**: duplicate clip names (e.g. multiple "Interview_A") mean
|
||||
only the *last* element indexed under that name is accessible. Use
|
||||
unique ``id`` attributes when possible.
|
||||
|
||||
``resources`` : ``Dict[str, Dict[str, Any]]``
|
||||
Every ``<asset>`` element keyed by ``id``, with pre-extracted ``name``,
|
||||
``src``, ``start``, ``duration``, and a reference to the raw element.
|
||||
|
||||
``formats`` : ``Dict[str, Dict[str, Any]]``
|
||||
Every ``<format>`` element keyed by ``id``.
|
||||
|
||||
Editing model
|
||||
-------------
|
||||
1. Look up the target clip via ``_require_clip`` / ``_require_spine_clip``.
|
||||
2. Mutate the clip's XML attributes (``start``, ``duration``, ``offset``).
|
||||
3. If the edit changes duration, ripple subsequent spine siblings via
|
||||
``_ripple_from_index`` so downstream offsets stay contiguous.
|
||||
4. Call ``save()`` to serialise the modified tree back to disk.
|
||||
|
||||
Example::
|
||||
|
||||
modifier = FCPXMLModifier("project.fcpxml")
|
||||
modifier.add_marker("clip_0", "00:00:10:00", "Review", MarkerType.INCOMPLETE)
|
||||
modifier.trim_clip("clip_1", trim_end="-2s")
|
||||
modifier.save("project_modified.fcpxml")
|
||||
|
||||
Attributes:
|
||||
path (Path): Filesystem path to the source FCPXML file.
|
||||
tree (ET.ElementTree): Parsed XML tree (mutated in-place by edits).
|
||||
root (ET.Element): Root ``<fcpxml>`` element.
|
||||
fps (float): Detected frame rate from the first ``<format>`` resource.
|
||||
clips (Dict[str, ET.Element]): Clip index — see *Index design* above.
|
||||
resources (Dict[str, Dict]): Asset index.
|
||||
formats (Dict[str, Dict]): Format index.
|
||||
"""
|
||||
|
||||
def __init__(self, fcpxml_path: str):
|
||||
"""Load *fcpxml_path*, parse its XML, and build lookup indices.
|
||||
|
||||
The constructor eagerly builds all three indices (clips, resources,
|
||||
formats) and detects the project frame rate. After construction the
|
||||
modifier is ready for any editing operation.
|
||||
|
||||
Args:
|
||||
fcpxml_path: Absolute or relative path to an ``.fcpxml`` file or
|
||||
an ``.fcpxmld`` bundle (a directory wrapping ``Info.fcpxml``
|
||||
plus sidecar data files for object tracking / Cinematic mode).
|
||||
|
||||
Raises:
|
||||
FileNotFoundError: If *fcpxml_path* does not exist.
|
||||
ET.ParseError: If the file is not valid XML.
|
||||
ValueError: If no ``<spine>`` is found (checked lazily on first edit).
|
||||
"""
|
||||
path = Path(fcpxml_path)
|
||||
self.bundle_dir: Optional[Path] = None
|
||||
if path.suffix.lower() == '.fcpxmld':
|
||||
self.bundle_dir = path
|
||||
inner = path / 'Info.fcpxml'
|
||||
if not inner.exists():
|
||||
raise FileNotFoundError(
|
||||
f"Info.fcpxml not found in bundle: {fcpxml_path}"
|
||||
)
|
||||
fcpxml_path = str(inner)
|
||||
self.path = Path(fcpxml_path)
|
||||
from ..safe_xml import safe_parse
|
||||
self.tree = safe_parse(fcpxml_path)
|
||||
self.root = self.tree.getroot()
|
||||
self.fps = self._detect_fps()
|
||||
# Lazily filled on the first generated title; see _unique_text_style_id.
|
||||
self._text_style_ids: Optional[set] = None
|
||||
# Lazily filled on the first clip split/cut; see _unique_tracking_shape_id.
|
||||
self._tracking_shape_ids: Optional[set] = None
|
||||
self._build_resource_index()
|
||||
self._build_clip_index()
|
||||
|
||||
def _detect_fps(self) -> float:
|
||||
"""Extract frame rate from format resource."""
|
||||
for fmt in self.root.findall('.//format'):
|
||||
frame_dur = fmt.get('frameDuration', '1/30s')
|
||||
if '/' in frame_dur:
|
||||
parts = frame_dur.replace('s', '').split('/', 1)
|
||||
num, denom = int(parts[0]), int(parts[1])
|
||||
if num <= 0:
|
||||
return 30.0
|
||||
return denom / num
|
||||
return 30.0
|
||||
|
||||
def frame_duration_fraction(self):
|
||||
"""Exact ``frameDuration`` as a Fraction (e.g. 1001/24000 at 23.976fps).
|
||||
|
||||
Unlike ``_detect_fps()`` (a float, lossy for NTSC rates), this is
|
||||
exact — use it wherever a cut boundary is snapped to the frame grid,
|
||||
so 23.976/29.97/59.94 timebases don't drift off-grid the way a
|
||||
hardcoded tick base like 2400 does.
|
||||
"""
|
||||
|
||||
for fmt in self.root.findall('.//format'):
|
||||
raw = fmt.get('frameDuration', '')
|
||||
if raw.endswith('s') and '/' in raw:
|
||||
n, d = raw[:-1].split('/', 1)
|
||||
fd = Fraction(int(n), int(d))
|
||||
if fd > 0:
|
||||
return fd
|
||||
return Fraction(1, 30)
|
||||
|
||||
def frame_size(self) -> 'Tuple[float, float]':
|
||||
"""The sequence's frame size in pixels, as ``(width, height)``.
|
||||
|
||||
Reads the sequence's own ``<format>`` when it references one, since a
|
||||
document may carry several (an asset's source format need not match
|
||||
the timeline's). Falls back to the first format that declares a size,
|
||||
then to 1920x1080.
|
||||
"""
|
||||
formats = {f.get('id'): f for f in self.root.findall('.//format')}
|
||||
candidates = []
|
||||
seq = self.root.find('.//sequence')
|
||||
if seq is not None and formats.get(seq.get('format')) is not None:
|
||||
candidates.append(formats[seq.get('format')])
|
||||
candidates.extend(formats.values())
|
||||
|
||||
for fmt in candidates:
|
||||
try:
|
||||
width = float(fmt.get('width') or 0)
|
||||
height = float(fmt.get('height') or 0)
|
||||
except (TypeError, ValueError):
|
||||
continue
|
||||
if width > 0 and height > 0:
|
||||
return width, height
|
||||
return 1920.0, 1080.0
|
||||
|
||||
def frame_width(self) -> float:
|
||||
"""The sequence's frame width in pixels."""
|
||||
return self.frame_size()[0]
|
||||
|
||||
def frame_height(self) -> float:
|
||||
"""The sequence's frame height in pixels."""
|
||||
return self.frame_size()[1]
|
||||
|
||||
def snap_seconds_to_frame(self, seconds: float) -> 'TimeValue':
|
||||
"""Round *seconds* to the nearest exact frame boundary as a TimeValue."""
|
||||
fd = self.frame_duration_fraction()
|
||||
frames = round(seconds / float(fd))
|
||||
snapped = fd * frames
|
||||
return TimeValue(snapped.numerator, snapped.denominator)
|
||||
|
||||
def snap_spine_times_to_frames(self) -> None:
|
||||
"""Snap primary-storyline offsets and durations to sequence frames.
|
||||
|
||||
Final Cut rejects otherwise valid XML when ripple edits leave a clip
|
||||
boundary between frames. Use the exact ``frameDuration`` fraction,
|
||||
rather than a float FPS, to preserve 23.976/29.97 timebases.
|
||||
"""
|
||||
|
||||
frame_duration = None
|
||||
for fmt in self.root.findall('.//format'):
|
||||
raw = fmt.get('frameDuration', '')
|
||||
if raw.endswith('s') and '/' in raw:
|
||||
n, d = raw[:-1].split('/', 1)
|
||||
frame_duration = Fraction(int(n), int(d))
|
||||
break
|
||||
if frame_duration is None or frame_duration <= 0:
|
||||
return
|
||||
|
||||
for element in self.root.findall('.//spine/*'):
|
||||
for attr in ('offset', 'duration'):
|
||||
raw = element.get(attr)
|
||||
if not raw or not raw.endswith('s'):
|
||||
continue
|
||||
value = raw[:-1]
|
||||
if '/' in value:
|
||||
n, d = value.split('/', 1)
|
||||
seconds = Fraction(int(n), int(d))
|
||||
else:
|
||||
seconds = Fraction(value)
|
||||
frames = int(round(float(seconds / frame_duration)))
|
||||
snapped = frame_duration * frames
|
||||
element.set(attr, f'{snapped.numerator}/{snapped.denominator}s')
|
||||
|
||||
def _build_resource_index(self) -> None:
|
||||
"""Build ``self.resources`` and ``self.formats`` from ``<asset>``/``<format>`` elements.
|
||||
|
||||
Called once during ``__init__``. Each asset entry stores the raw
|
||||
element plus pre-extracted metadata so callers don't need to
|
||||
re-parse attributes on every access.
|
||||
"""
|
||||
self.resources: Dict[str, Dict[str, Any]] = {}
|
||||
self.formats: Dict[str, Dict[str, Any]] = {}
|
||||
|
||||
for asset in self.root.findall('.//asset'):
|
||||
asset_id = asset.get('id', '')
|
||||
self.resources[asset_id] = {
|
||||
'id': asset_id,
|
||||
'name': asset.get('name', ''),
|
||||
'src': asset.get('src', '') or (asset.find('media-rep').get('src', '') if asset.find('media-rep') is not None else ''),
|
||||
'start': asset.get('start', '0s'),
|
||||
'duration': asset.get('duration', '0s'),
|
||||
'element': asset
|
||||
}
|
||||
|
||||
for fmt in self.root.findall('.//format'):
|
||||
fmt_id = fmt.get('id', '')
|
||||
self.formats[fmt_id] = {
|
||||
'id': fmt_id,
|
||||
'name': fmt.get('name', ''),
|
||||
'element': fmt
|
||||
}
|
||||
|
||||
def _index_elements(self, tag: str, fallback_prefix: str) -> None:
|
||||
"""Index XML elements of *tag* into ``self.clips`` by id/name.
|
||||
|
||||
Each element is keyed by its ``id`` attribute, falling back to
|
||||
``name``, then a generated ``{fallback_prefix}_{i}`` key. This
|
||||
replaces three near-identical loops that only differed in the tag
|
||||
name and fallback prefix.
|
||||
"""
|
||||
for i, elem in enumerate(self.root.findall(f'.//{tag}')):
|
||||
key = elem.get('id') or elem.get('name') or f"{fallback_prefix}_{i}"
|
||||
self.clips[key] = elem
|
||||
|
||||
def _build_clip_index(self) -> None:
|
||||
"""Build ``self.clips`` index from all clip-type elements.
|
||||
|
||||
Indexes ``<clip>``, ``<asset-clip>``, and ``<video>`` tags. Keys are
|
||||
resolved by ``_index_elements`` (``id`` → ``name`` → generated).
|
||||
|
||||
.. warning::
|
||||
Duplicate names cause last-one-wins overwrites. If your project
|
||||
has multiple clips named "Interview_A", only the last one parsed
|
||||
will be reachable by name. Prefer unique ``id`` attributes.
|
||||
"""
|
||||
self.clips: Dict[str, ET.Element] = {}
|
||||
for tag, prefix in (('clip', 'clip'), ('asset-clip', 'asset_clip'), ('video', 'video')):
|
||||
self._index_elements(tag, prefix)
|
||||
|
||||
def _get_spine(self) -> ET.Element:
|
||||
"""Get the primary storyline spine.
|
||||
|
||||
Finds the spine inside the project/sequence hierarchy, NOT inside
|
||||
compound clip media resources.
|
||||
"""
|
||||
# Prefer the main timeline spine (under project/sequence)
|
||||
spine = self.root.find('.//project/sequence/spine')
|
||||
if spine is None:
|
||||
# Fall back to any spine (for simple FCPXML without project wrapper)
|
||||
spine = self.root.find('.//spine')
|
||||
if spine is None:
|
||||
raise ValueError("No spine found in FCPXML")
|
||||
return spine
|
||||
|
||||
def _iter_spine_clips(self) -> list[tuple[int, ET.Element]]:
|
||||
"""Return an indexed list of clip-type elements in the primary spine.
|
||||
|
||||
Filters out gaps, transitions, and other non-clip elements, returning
|
||||
only ``(index_in_spine, element)`` pairs where the tag is in
|
||||
``CLIP_TAGS``. The index is the element's position among *all* spine
|
||||
children (not just clips), so it stays valid for insertion/removal.
|
||||
"""
|
||||
spine = self._get_spine()
|
||||
return [
|
||||
(i, child)
|
||||
for i, child in enumerate(spine.findall('*'))
|
||||
if child.tag in CLIP_TAGS
|
||||
]
|
||||
|
||||
def _find_spine_clip_at_seconds(self, target_seconds: float) -> tuple[ET.Element, float]:
|
||||
"""Find the spine clip containing *target_seconds* and return it with the relative offset.
|
||||
|
||||
Returns:
|
||||
``(clip_element, relative_seconds)`` — the clip and the time
|
||||
within that clip corresponding to *target_seconds*.
|
||||
|
||||
Raises:
|
||||
ValueError: If no clip spans the requested position.
|
||||
"""
|
||||
spine = self._get_spine()
|
||||
for child in spine.findall('*'):
|
||||
if child.tag not in CLIP_TAGS:
|
||||
continue
|
||||
offset = self._parse_time(child.get('offset', '0s')).to_seconds()
|
||||
dur = self._parse_time(child.get('duration', '0s')).to_seconds()
|
||||
if offset <= target_seconds < offset + dur:
|
||||
return child, target_seconds - offset
|
||||
raise ValueError(f"No spine clip at position {target_seconds:.3f}s")
|
||||
|
||||
def _parse_time(self, tc: str) -> TimeValue:
|
||||
"""Parse a timecode string to TimeValue."""
|
||||
return TimeValue.from_timecode(tc, self.fps)
|
||||
|
||||
def _get_clip_times(
|
||||
self, clip: ET.Element
|
||||
) -> tuple:
|
||||
"""Return (start, duration, offset) TimeValues for a clip element."""
|
||||
return (
|
||||
self._parse_time(clip.get('start', '0s')),
|
||||
self._parse_time(clip.get('duration', '0s')),
|
||||
self._parse_time(clip.get('offset', '0s')),
|
||||
)
|
||||
|
||||
def source_file_start(self, clip: ET.Element) -> 'TimeValue':
|
||||
"""Return a clip's in-point measured from the head of its media file.
|
||||
|
||||
FCPXML ``start`` on an asset-clip is a source *timecode*, and the
|
||||
asset's own ``start`` is the timecode of the source media's first
|
||||
frame. Media analysis (ffmpeg silencedetect, Whisper) reports
|
||||
file-relative time, so subtract the asset's start timecode to land
|
||||
both on the same origin. When the asset starts at 0s (the common
|
||||
case, and every test fixture) this is a no-op.
|
||||
"""
|
||||
ref = clip.get('ref', '')
|
||||
asset = self.resources.get(ref, {})
|
||||
asset_start = self._parse_time(asset.get('start', '0s'))
|
||||
clip_start = self._parse_time(clip.get('start', '0s'))
|
||||
return clip_start - asset_start
|
||||
|
||||
def _resolve_clip_duration(
|
||||
self,
|
||||
asset: dict,
|
||||
duration: Optional[str] = None,
|
||||
in_point: Optional[str] = None,
|
||||
out_point: Optional[str] = None,
|
||||
) -> tuple['TimeValue', 'TimeValue']:
|
||||
"""Compute clip duration and source start from optional overrides.
|
||||
|
||||
Centralises the three-way fallback logic shared by insert_clip,
|
||||
add_connected_clip, and add_audio_clip:
|
||||
|
||||
1. If *in_point* and *out_point* are given → subclip range.
|
||||
2. Else if *duration* is given → explicit duration, source start = 0.
|
||||
3. Else → full asset duration, source start = 0.
|
||||
|
||||
Returns:
|
||||
``(clip_duration, source_start)`` TimeValue pair.
|
||||
"""
|
||||
if in_point and out_point:
|
||||
in_time = self._parse_time(in_point)
|
||||
out_time = self._parse_time(out_point)
|
||||
return out_time - in_time, in_time
|
||||
if duration:
|
||||
return self._parse_time(duration), TimeValue.zero()
|
||||
return self._parse_time(asset.get('duration', '0s')), TimeValue.zero()
|
||||
|
||||
def _make_asset_clip(
|
||||
self,
|
||||
asset_id: str,
|
||||
name: str,
|
||||
offset: 'TimeValue',
|
||||
start: 'TimeValue',
|
||||
duration: 'TimeValue',
|
||||
*,
|
||||
parent: Optional[ET.Element] = None,
|
||||
**extra_attrs: str,
|
||||
) -> ET.Element:
|
||||
"""Build an ``<asset-clip>`` element with standard attributes.
|
||||
|
||||
Centralises the repeated element creation shared by insert_clip,
|
||||
add_connected_clip, and add_audio_clip. Each caller can pass
|
||||
additional attributes (``lane``, ``audioRole``, ``format``) via
|
||||
*extra_attrs*.
|
||||
|
||||
Args:
|
||||
asset_id: Resource reference (e.g. ``'r3'``).
|
||||
name: Human-readable clip name.
|
||||
offset: Timeline offset (or offset within parent for connected clips).
|
||||
start: Source media start point.
|
||||
duration: Clip duration.
|
||||
parent: If given, create the element as a SubElement of *parent*;
|
||||
otherwise create a detached Element.
|
||||
**extra_attrs: Additional XML attributes (``lane``, ``audioRole``).
|
||||
|
||||
Returns:
|
||||
The new ``<asset-clip>`` Element.
|
||||
"""
|
||||
if parent is not None:
|
||||
elem = ET.SubElement(parent, 'asset-clip')
|
||||
else:
|
||||
elem = ET.Element('asset-clip')
|
||||
elem.set('ref', asset_id)
|
||||
elem.set('offset', offset.to_fcpxml())
|
||||
elem.set('name', name)
|
||||
elem.set('start', start.to_fcpxml())
|
||||
elem.set('duration', duration.to_fcpxml())
|
||||
for attr, val in extra_attrs.items():
|
||||
elem.set(attr, val)
|
||||
return elem
|
||||
|
||||
def _require_clip(self, clip_id: 'str | ET.Element') -> ET.Element:
|
||||
"""Look up a clip by ID/name, raising if not found.
|
||||
|
||||
Centralises the get-or-raise pattern used by every clip-mutating
|
||||
method so the error message stays consistent and future
|
||||
enhancements (fuzzy matching, suggestions) only need one site.
|
||||
|
||||
An Element is returned as-is. That matters after ``split_clip`` or
|
||||
``cut_clip_ranges``: the resulting pieces all carry the *same* name,
|
||||
so a name lookup would always resolve to the first one and silently
|
||||
put the edit on the wrong piece. Callers holding the exact element
|
||||
pass it directly.
|
||||
"""
|
||||
if isinstance(clip_id, ET.Element):
|
||||
return clip_id
|
||||
clip = self.clips.get(clip_id)
|
||||
if clip is None:
|
||||
raise ValueError(f"Clip not found: {clip_id}")
|
||||
return clip
|
||||
|
||||
def _require_spine_clip(self, clip_id: str) -> tuple[ET.Element, ET.Element, int]:
|
||||
"""Look up a clip and verify it lives in the primary spine.
|
||||
|
||||
Returns:
|
||||
``(spine, clip, index_in_spine)`` tuple.
|
||||
|
||||
Raises:
|
||||
ValueError: If the clip doesn't exist or isn't in the spine.
|
||||
"""
|
||||
clip = self._require_clip(clip_id)
|
||||
spine = self._get_spine()
|
||||
clip_index = self._find_clip_index(spine, clip)
|
||||
if clip_index is None:
|
||||
raise ValueError(f"Clip not in spine: {clip_id}")
|
||||
return spine, clip, clip_index
|
||||
|
||||
def _find_clip_index(self, spine: ET.Element, clip: ET.Element) -> int | None:
|
||||
"""Find the index of a clip in the spine. Returns None if not found."""
|
||||
for i, child in enumerate(spine):
|
||||
if child == clip:
|
||||
return i
|
||||
return None
|
||||
|
||||
@staticmethod
|
||||
def _find_neighbor_clip(
|
||||
spine_list: list, index: int, direction: str
|
||||
) -> Optional[ET.Element]:
|
||||
"""Find the nearest non-gap clip before or after *index* in *spine_list*.
|
||||
|
||||
Args:
|
||||
spine_list: Materialised list of spine children.
|
||||
index: Position to search from (exclusive).
|
||||
direction: ``'prev'`` to search backward, ``'next'`` to search forward.
|
||||
|
||||
Returns:
|
||||
The first clip-type element found, or ``None``.
|
||||
"""
|
||||
if direction == 'prev':
|
||||
for j in range(index - 1, -1, -1):
|
||||
if spine_list[j].tag in CLIP_TAGS:
|
||||
return spine_list[j]
|
||||
else:
|
||||
for j in range(index + 1, len(spine_list)):
|
||||
if spine_list[j].tag in CLIP_TAGS:
|
||||
return spine_list[j]
|
||||
return None
|
||||
|
||||
def _resolve_asset(
|
||||
self, asset_id: Optional[str], asset_name: Optional[str]
|
||||
) -> tuple:
|
||||
"""Look up an asset by ID or name from ``self.resources``.
|
||||
|
||||
Returns:
|
||||
``(asset_dict, resolved_asset_id)`` tuple.
|
||||
|
||||
Raises:
|
||||
ValueError: If neither ID nor name matches a known asset.
|
||||
"""
|
||||
if asset_id and asset_id in self.resources:
|
||||
return self.resources[asset_id], asset_id
|
||||
if asset_name:
|
||||
for res_id, res_data in self.resources.items():
|
||||
if res_data.get('name') == asset_name:
|
||||
return res_data, res_id
|
||||
raise ValueError(f"Asset not found: {asset_id or asset_name}")
|
||||
|
||||
@staticmethod
|
||||
def _unique_resource_id(resources: ET.Element, prefix: str) -> str:
|
||||
"""Generate a unique resource ID with the given *prefix*.
|
||||
|
||||
Starts with ``prefix`` (e.g. ``'r_audio1'``), appending an
|
||||
incrementing counter until no collision exists in *resources*.
|
||||
"""
|
||||
existing_ids = {el.get('id', '') for el in resources}
|
||||
candidate = prefix
|
||||
counter = 2
|
||||
while candidate in existing_ids:
|
||||
# Strip trailing digits from prefix for the counter suffix
|
||||
base = prefix.rstrip('0123456789')
|
||||
candidate = f'{base}{counter}'
|
||||
counter += 1
|
||||
return candidate
|
||||
|
||||
def _find_spine_element_at_timecode(
|
||||
self, spine: ET.Element, target_tc: str, *, require_clip: bool = False
|
||||
) -> Optional[ET.Element]:
|
||||
"""Find the first spine child whose offset matches *target_tc*.
|
||||
|
||||
Normalises both sides through ``TimeValue`` round-trip so format
|
||||
differences (e.g. ``"3600/2400s"`` vs ``"1800/1200s"``) don't
|
||||
cause false negatives.
|
||||
|
||||
Args:
|
||||
spine: The ``<spine>`` element to search.
|
||||
target_tc: Timecode string to match against each child's offset.
|
||||
require_clip: If True, skip non-clip elements (gaps, etc.).
|
||||
"""
|
||||
for child in spine:
|
||||
offset_str = child.get('offset', '0s')
|
||||
tc = TimeValue.from_timecode(offset_str, self.fps).to_timecode(self.fps)
|
||||
if tc == target_tc:
|
||||
if require_clip and child.tag not in CLIP_TAGS:
|
||||
continue
|
||||
return child
|
||||
return None
|
||||
|
||||
def _absorb_into_neighbor(
|
||||
self,
|
||||
spine: ET.Element,
|
||||
element: ET.Element,
|
||||
direction: str,
|
||||
) -> Optional[ET.Element]:
|
||||
"""Extend a neighbor clip to absorb *element*'s duration, then remove *element*.
|
||||
|
||||
Shared by ``fix_flash_frames`` (absorbing flash-frame clips) and
|
||||
``fill_gaps`` (absorbing gap elements). Both operations find the
|
||||
nearest clip in *direction*, grow it by the absorbed element's
|
||||
duration, and remove the absorbed element from the spine.
|
||||
|
||||
When extending backward (``direction='next'``), the neighbor's
|
||||
source in-point is also pulled earlier so the extra frames come
|
||||
from before the original cut, not after.
|
||||
|
||||
Does **not** call ``_recalculate_offsets`` — callers decide when to
|
||||
recalculate (per-iteration vs. once at the end).
|
||||
|
||||
Args:
|
||||
spine: The primary storyline ``<spine>`` element.
|
||||
element: The clip or gap to absorb (will be removed).
|
||||
direction: ``'prev'`` to extend the previous clip forward,
|
||||
``'next'`` to extend the next clip backward.
|
||||
|
||||
Returns:
|
||||
The neighbor clip that absorbed the duration, or ``None`` if
|
||||
no suitable neighbor exists.
|
||||
"""
|
||||
spine_list = list(spine)
|
||||
element_index = spine_list.index(element)
|
||||
neighbor = self._find_neighbor_clip(spine_list, element_index, direction)
|
||||
if neighbor is None:
|
||||
return None
|
||||
|
||||
absorbed_dur = self._parse_time(element.get('duration', '0s'))
|
||||
neighbor_dur = self._parse_time(neighbor.get('duration', '0s'))
|
||||
|
||||
if direction == 'next':
|
||||
neighbor_start = self._parse_time(neighbor.get('start', '0s'))
|
||||
new_start = neighbor_start - absorbed_dur
|
||||
if new_start >= TimeValue.zero():
|
||||
neighbor.set('start', new_start.to_fcpxml())
|
||||
neighbor.set('duration', (neighbor_dur + absorbed_dur).to_fcpxml())
|
||||
else:
|
||||
# Can't shift start negative — only extend by what's available
|
||||
available = neighbor_start
|
||||
neighbor.set('start', TimeValue(0, 1).to_fcpxml())
|
||||
neighbor.set('duration', (neighbor_dur + available).to_fcpxml())
|
||||
else:
|
||||
neighbor.set('duration', (neighbor_dur + absorbed_dur).to_fcpxml())
|
||||
spine.remove(element)
|
||||
return neighbor
|
||||
|
||||
def _resolve_insert_position(
|
||||
self, position: str, spine_children: list
|
||||
) -> tuple:
|
||||
"""Translate a human-friendly position spec into (target_offset, insert_index).
|
||||
|
||||
Supported formats:
|
||||
``'start'`` — beginning of spine
|
||||
``'end'`` — after last element
|
||||
``'after:clip_id'`` — after the named clip
|
||||
``'before:clip_id'``— before the named clip
|
||||
*timecode* — absolute timeline position
|
||||
|
||||
Returns:
|
||||
``(TimeValue, int)`` — the offset and child-index for spine insertion.
|
||||
"""
|
||||
if position == 'start':
|
||||
return TimeValue.zero(), 0
|
||||
|
||||
if position == 'end':
|
||||
if spine_children:
|
||||
last = spine_children[-1]
|
||||
last_offset = self._parse_time(last.get('offset', '0s'))
|
||||
last_dur = self._parse_time(last.get('duration', '0s'))
|
||||
return last_offset + last_dur, len(spine_children)
|
||||
return TimeValue.zero(), len(spine_children)
|
||||
|
||||
if position.startswith('after:') or position.startswith('before:'):
|
||||
is_after = position.startswith('after:')
|
||||
ref_id = position.split(':', 1)[1]
|
||||
ref_clip = self.clips.get(ref_id)
|
||||
if ref_clip is None or ref_clip not in spine_children:
|
||||
raise ValueError(f"Reference clip not found: {ref_id}")
|
||||
idx = spine_children.index(ref_clip)
|
||||
ref_offset = self._parse_time(ref_clip.get('offset', '0s'))
|
||||
if is_after:
|
||||
ref_dur = self._parse_time(ref_clip.get('duration', '0s'))
|
||||
return ref_offset + ref_dur, idx + 1
|
||||
return ref_offset, idx
|
||||
|
||||
# Assume timecode
|
||||
target_offset = self._parse_time(position)
|
||||
insert_index = 0
|
||||
for i, child in enumerate(spine_children):
|
||||
child_offset = self._parse_time(child.get('offset', '0s'))
|
||||
if child_offset >= target_offset:
|
||||
insert_index = i
|
||||
break
|
||||
insert_index = i + 1
|
||||
return target_offset, insert_index
|
||||
|
||||
def _make_transition_element(
|
||||
self,
|
||||
effect_name: str,
|
||||
trans_offset: 'TimeValue',
|
||||
trans_duration: 'TimeValue',
|
||||
effect_ref_id: str | None,
|
||||
) -> ET.Element:
|
||||
"""Build a <transition> element with optional filter-video child."""
|
||||
transition = ET.Element('transition')
|
||||
transition.set('name', effect_name)
|
||||
transition.set('offset', trans_offset.to_fcpxml())
|
||||
transition.set('duration', trans_duration.to_fcpxml())
|
||||
if effect_ref_id:
|
||||
fv = ET.SubElement(transition, 'filter-video')
|
||||
fv.set('ref', effect_ref_id)
|
||||
fv.set('name', effect_name)
|
||||
return transition
|
||||
|
||||
def save(self, output_path: Optional[str] = None) -> str:
|
||||
"""Serialise the modified XML tree to disk.
|
||||
|
||||
When the destination ends in ``.fcpxmld`` a bundle directory is
|
||||
created and the XML lands in ``Info.fcpxml`` inside it. If the
|
||||
source was also a bundle, every sidecar file (object-tracking /
|
||||
Cinematic-mode ``dataLocator`` payloads — anything that isn't
|
||||
``Info.fcpxml``) is copied across so the round-trip is lossless.
|
||||
Writing a bundle source to a flat ``.fcpxml`` destination drops
|
||||
those sidecars by definition.
|
||||
|
||||
Args:
|
||||
output_path: Destination ``.fcpxml`` file or ``.fcpxmld``
|
||||
bundle path. Defaults to overwriting the original
|
||||
file/bundle loaded in ``__init__``.
|
||||
|
||||
Returns:
|
||||
The absolute path written to (the bundle path when writing
|
||||
a bundle, not the inner ``Info.fcpxml``).
|
||||
"""
|
||||
if output_path is None:
|
||||
out = self.bundle_dir if self.bundle_dir is not None else self.path
|
||||
else:
|
||||
out = Path(output_path)
|
||||
|
||||
# Every write path goes through here, so snapping here (rather than
|
||||
# in each handler) guarantees ripple edits never leave a spine clip
|
||||
# off the frame grid — see snap_spine_times_to_frames() docstring.
|
||||
# No-op (each value already equals its own snapped form) on content
|
||||
# that was already frame-aligned.
|
||||
self.snap_spine_times_to_frames()
|
||||
|
||||
if out.suffix.lower() == '.fcpxmld':
|
||||
out.mkdir(exist_ok=True)
|
||||
if (
|
||||
self.bundle_dir is not None
|
||||
and self.bundle_dir.resolve() != out.resolve()
|
||||
):
|
||||
self._copy_bundle_sidecars(self.bundle_dir, out)
|
||||
write_fcpxml(self.root, str(out / 'Info.fcpxml'), fps=self.fps)
|
||||
return str(out)
|
||||
|
||||
return write_fcpxml(self.root, str(out), fps=self.fps)
|
||||
|
||||
@staticmethod
|
||||
def _copy_bundle_sidecars(src_bundle: Path, dst_bundle: Path) -> None:
|
||||
"""Copy every sidecar entry of *src_bundle* into *dst_bundle*.
|
||||
|
||||
Sidecars are all bundle members except ``Info.fcpxml`` itself —
|
||||
e.g. the external data files that ``locator``/``dataLocator``
|
||||
elements reference for object tracking and Cinematic mode.
|
||||
"""
|
||||
import shutil
|
||||
for entry in src_bundle.iterdir():
|
||||
if entry.name == 'Info.fcpxml':
|
||||
continue
|
||||
target = dst_bundle / entry.name
|
||||
if entry.is_dir():
|
||||
shutil.copytree(entry, target, dirs_exist_ok=True)
|
||||
else:
|
||||
shutil.copy2(entry, target)
|
||||
|
||||
@@ -0,0 +1,367 @@
|
||||
"""Dividir, cortar faixas e apagar clipes.
|
||||
|
||||
Extraído de writer.py — ver fcpxml/writer/__init__.py para o conjunto.
|
||||
"""
|
||||
|
||||
import copy
|
||||
import xml.etree.ElementTree as ET
|
||||
from typing import List, Tuple
|
||||
|
||||
from ..models import (
|
||||
TimeValue,
|
||||
)
|
||||
|
||||
|
||||
class CutMixin:
|
||||
"""Dividir, cortar faixas e apagar clipes."""
|
||||
|
||||
# SPLIT & DELETE OPERATIONS
|
||||
# ========================================================================
|
||||
|
||||
@staticmethod
|
||||
def _filter_children_for_segment(
|
||||
clip: ET.Element,
|
||||
seg_start: 'TimeValue',
|
||||
seg_duration: 'TimeValue',
|
||||
) -> None:
|
||||
"""Remove markers/keywords/titles from *clip* that fall outside the segment range.
|
||||
|
||||
After ``split_clip`` deepcopy's the original clip into each segment, every
|
||||
segment inherits all child elements. Markers whose ``start`` falls outside
|
||||
``[seg_start, seg_start + seg_duration)`` are phantom duplicates and must be
|
||||
removed. Keywords that partially overlap get their ``start``/``duration``
|
||||
clamped to the segment boundaries.
|
||||
|
||||
A lane-nested ``<title>`` (a "text" voice action's on-screen callout,
|
||||
or a caption from an earlier `generate_dynamic_subtitles` pass) is
|
||||
the same kind of phantom duplicate, just keyed on ``offset`` instead
|
||||
of ``start`` — its offset lives in the same source-media coordinate
|
||||
space as a marker's ``start`` (see ``add_text_title``/``add_marker``,
|
||||
both anchored at ``parent.start``). Left unfiltered, every further
|
||||
cut (silence removal, filler removal) duplicates it into every
|
||||
resulting piece, so the same word shows up several times across the
|
||||
edited timeline instead of once where it was placed.
|
||||
|
||||
A lane-nested ``<video>`` zoom (the "Clipe de Ajuste" adjustment
|
||||
layer ``add_zoom`` creates, ``role`` starting with ``"adjustments."``)
|
||||
is the exact same phantom-duplicate case, keyed on ``offset``+
|
||||
``duration`` like a keyword. Left unfiltered, every further cut
|
||||
duplicates the zoom into every resulting piece with its original
|
||||
offset untouched — each copy then draws at the same absolute
|
||||
position, so two "Clipe de Ajuste" bars appear stacked on top of
|
||||
each other in the timeline instead of the one real zoom window.
|
||||
"""
|
||||
seg_end = seg_start + seg_duration
|
||||
to_remove = []
|
||||
for child in clip:
|
||||
tag = child.tag
|
||||
if tag in ('marker', 'chapter-marker'):
|
||||
child_start = TimeValue.from_timecode(child.get('start', '0s'))
|
||||
if child_start < seg_start or child_start >= seg_end:
|
||||
to_remove.append(child)
|
||||
elif tag == 'title':
|
||||
title_offset = TimeValue.from_timecode(child.get('offset', '0s'))
|
||||
if title_offset < seg_start or title_offset >= seg_end:
|
||||
to_remove.append(child)
|
||||
elif tag == 'video' and (child.get('role') or '').startswith('adjustments.'):
|
||||
v_offset = TimeValue.from_timecode(child.get('offset', '0s'))
|
||||
v_dur = TimeValue.from_timecode(child.get('duration', '0s'))
|
||||
v_end = v_offset + v_dur
|
||||
if v_end <= seg_start or v_offset >= seg_end:
|
||||
to_remove.append(child)
|
||||
elif tag == 'keyword':
|
||||
kw_start = TimeValue.from_timecode(child.get('start', '0s'))
|
||||
kw_dur = TimeValue.from_timecode(child.get('duration', '0s'))
|
||||
kw_end = kw_start + kw_dur
|
||||
# Completely outside segment → remove
|
||||
if kw_end <= seg_start or kw_start >= seg_end:
|
||||
to_remove.append(child)
|
||||
else:
|
||||
# Clamp keyword range to segment boundaries
|
||||
clamped_start = max(kw_start, seg_start)
|
||||
clamped_end = min(kw_end, seg_end)
|
||||
child.set('start', clamped_start.to_fcpxml())
|
||||
child.set('duration', (clamped_end - clamped_start).to_fcpxml())
|
||||
for child in to_remove:
|
||||
clip.remove(child)
|
||||
|
||||
def split_clip(
|
||||
self,
|
||||
clip_id: str,
|
||||
split_points: List[str]
|
||||
) -> List[ET.Element]:
|
||||
"""
|
||||
Split a clip at specified timecodes.
|
||||
|
||||
Args:
|
||||
clip_id: Clip to split
|
||||
split_points: Timecodes within the clip to split at
|
||||
|
||||
Returns:
|
||||
List of resulting clip elements
|
||||
"""
|
||||
spine, clip, clip_index = self._require_spine_clip(clip_id)
|
||||
|
||||
# Get clip properties
|
||||
clip_start, clip_duration, clip_offset = self._get_clip_times(clip)
|
||||
clip_name = clip.get('name', 'Clip')
|
||||
|
||||
# Sort split points
|
||||
split_times = sorted([self._parse_time(sp) for sp in split_points])
|
||||
|
||||
# Remove original clip
|
||||
spine.remove(clip)
|
||||
|
||||
# Create new clips
|
||||
new_clips = []
|
||||
current_offset = clip_offset
|
||||
current_start = clip_start
|
||||
|
||||
all_points = split_times + [clip_duration]
|
||||
|
||||
for i, split_time in enumerate(all_points):
|
||||
if i == 0:
|
||||
segment_duration = split_time
|
||||
else:
|
||||
segment_duration = split_time - split_times[i - 1]
|
||||
|
||||
if segment_duration <= TimeValue.zero():
|
||||
continue
|
||||
|
||||
# Create new clip
|
||||
new_clip = copy.deepcopy(clip)
|
||||
new_clip.set('name', clip_name)
|
||||
new_clip.set('offset', current_offset.to_fcpxml())
|
||||
new_clip.set('start', current_start.to_fcpxml())
|
||||
new_clip.set('duration', segment_duration.to_fcpxml())
|
||||
|
||||
# Remove markers/keywords that belong to other segments
|
||||
self._filter_children_for_segment(
|
||||
new_clip, current_start, segment_duration
|
||||
)
|
||||
self._reassign_text_style_ids(new_clip)
|
||||
self._reassign_tracking_shape_ids(new_clip)
|
||||
|
||||
spine.insert(clip_index + len(new_clips), new_clip)
|
||||
new_clips.append(new_clip)
|
||||
|
||||
# Update for next iteration
|
||||
current_offset = current_offset + segment_duration
|
||||
current_start = current_start + segment_duration
|
||||
|
||||
# Update clip index: remove stale original entry, add split entries
|
||||
self.clips.pop(clip_id, None)
|
||||
for i, new_clip in enumerate(new_clips):
|
||||
new_id = f"{clip_id}_split_{i}"
|
||||
self.clips[new_id] = new_clip
|
||||
|
||||
return new_clips
|
||||
|
||||
def cut_clip_ranges(
|
||||
self,
|
||||
clip: ET.Element,
|
||||
cut_ranges: List[Tuple['TimeValue', 'TimeValue']],
|
||||
) -> 'TimeValue':
|
||||
"""Remove clip-relative time ranges from a spine clip, rippling after.
|
||||
|
||||
Element-based on purpose: callers that walk the spine (e.g. media
|
||||
silence removal) pass the exact element, so duplicate-named clips are
|
||||
never ambiguous the way name-keyed operations are.
|
||||
|
||||
Args:
|
||||
clip: The spine clip element to cut (must be a direct spine child).
|
||||
cut_ranges: (start, end) TimeValue pairs measured from the clip's
|
||||
own head. Overlapping/unsorted ranges are merged; portions
|
||||
outside [0, clip duration] are clamped. A cut covering the
|
||||
whole clip removes it entirely.
|
||||
|
||||
Returns:
|
||||
Total removed duration (zero if no effective ranges).
|
||||
"""
|
||||
spine = self._get_spine()
|
||||
clip_start, clip_duration, clip_offset = self._get_clip_times(clip)
|
||||
clip_index = list(spine).index(clip)
|
||||
zero = TimeValue.zero()
|
||||
|
||||
# Clamp, sort, merge.
|
||||
clamped = []
|
||||
for start, end in cut_ranges:
|
||||
start = start if start > zero else zero
|
||||
end = end if end < clip_duration else clip_duration
|
||||
if end > start:
|
||||
clamped.append((start, end))
|
||||
clamped.sort(key=lambda r: r[0])
|
||||
merged: List[Tuple[TimeValue, TimeValue]] = []
|
||||
for start, end in clamped:
|
||||
if merged and start <= merged[-1][1]:
|
||||
if end > merged[-1][1]:
|
||||
merged[-1] = (merged[-1][0], end)
|
||||
else:
|
||||
merged.append((start, end))
|
||||
if not merged:
|
||||
return zero
|
||||
|
||||
# Keep ranges = complement of the merged cuts.
|
||||
keeps: List[Tuple[TimeValue, TimeValue]] = []
|
||||
cursor = zero
|
||||
for start, end in merged:
|
||||
if start > cursor:
|
||||
keeps.append((cursor, start))
|
||||
cursor = end
|
||||
if cursor < clip_duration:
|
||||
keeps.append((cursor, clip_duration))
|
||||
|
||||
# A keep segment shorter than MIN_KEEP_SECONDS is leftover between
|
||||
# two cuts, not a real clip — at the very start/end of the clip it's
|
||||
# cut padding with no kept audio on the outer side; in the interior
|
||||
# it's the pause BETWEEN two things that were both cut (e.g. two
|
||||
# consecutive deactivated phrases in the voice-editing flow), which
|
||||
# belongs to neither side by construction. At the edges we fold it
|
||||
# into the one neighboring KEEP segment there is, which simply starts
|
||||
# earlier / ends later to absorb it. In the interior both neighbors
|
||||
# are CUT, not keep, so there is nothing to fold into — it is just
|
||||
# dropped, extending the surrounding cut across it instead of
|
||||
# surviving as a third near-invisible micro-clip.
|
||||
#
|
||||
# The threshold is bigger than one frame on purpose: measured on a
|
||||
# real voice-edit (0.07-0.23s residues), a single frame did not catch
|
||||
# them — this is pause/padding leftover, not intentional short
|
||||
# content, so treating anything under a third of a second this way
|
||||
# is safe for this cut path.
|
||||
min_keep_seconds = max(6 * float(self.frame_duration_fraction()), 0.3)
|
||||
i = 0
|
||||
while len(keeps) > 1 and i < len(keeps):
|
||||
start, end = keeps[i]
|
||||
if (end - start).to_seconds() >= min_keep_seconds:
|
||||
i += 1
|
||||
continue
|
||||
if i == 0:
|
||||
keeps[1] = (start, keeps[1][1])
|
||||
keeps.pop(0)
|
||||
elif i == len(keeps) - 1:
|
||||
keeps[i - 1] = (keeps[i - 1][0], end)
|
||||
keeps.pop(i)
|
||||
else:
|
||||
keeps.pop(i)
|
||||
# Re-check the same index: the segment now there might itself be
|
||||
# short enough to absorb again (two short keeps in a row).
|
||||
|
||||
spine.remove(clip)
|
||||
new_clips: List[ET.Element] = []
|
||||
current_offset = clip_offset
|
||||
kept_total = zero
|
||||
for keep_start, keep_end in keeps:
|
||||
seg_duration = keep_end - keep_start
|
||||
seg_start = clip_start + keep_start
|
||||
new_clip = copy.deepcopy(clip)
|
||||
new_clip.set('offset', current_offset.to_fcpxml())
|
||||
new_clip.set('start', seg_start.to_fcpxml())
|
||||
new_clip.set('duration', seg_duration.to_fcpxml())
|
||||
self._filter_children_for_segment(new_clip, seg_start, seg_duration)
|
||||
self._reassign_text_style_ids(new_clip)
|
||||
self._reassign_tracking_shape_ids(new_clip)
|
||||
spine.insert(clip_index + len(new_clips), new_clip)
|
||||
new_clips.append(new_clip)
|
||||
current_offset = current_offset + seg_duration
|
||||
kept_total = kept_total + seg_duration
|
||||
|
||||
removed = clip_duration - kept_total
|
||||
self._ripple_from_index(spine, clip_index + len(new_clips), zero - removed)
|
||||
self._update_sequence_duration()
|
||||
|
||||
# Keep the name index coherent, mirroring delete_clip/split_clip.
|
||||
name = clip.get('id') or clip.get('name') or ''
|
||||
if name and self.clips.get(name) is clip:
|
||||
if new_clips:
|
||||
self.clips[name] = new_clips[0]
|
||||
else:
|
||||
remaining = [
|
||||
sc for _, sc in self._iter_spine_clips()
|
||||
if (sc.get('id') or sc.get('name') or '') == name
|
||||
]
|
||||
if remaining:
|
||||
self.clips[name] = remaining[0]
|
||||
else:
|
||||
self.clips.pop(name, None)
|
||||
return removed
|
||||
|
||||
def remove_trailing_gaps(self) -> None:
|
||||
"""Remove empty ``<gap>`` elements at the end of the timeline.
|
||||
|
||||
Silence removal (and FCP round-trips) can leave a trailing gap holding
|
||||
the timeline open past the last real clip. This removes only *trailing*
|
||||
gaps — a gap in the middle is left untouched — and re-syncs the sequence
|
||||
duration so the exported file ends where the content ends.
|
||||
"""
|
||||
spine = self._get_spine()
|
||||
children = list(spine)
|
||||
if not children:
|
||||
return
|
||||
last = children[-1]
|
||||
if last.tag != 'gap':
|
||||
return
|
||||
spine.remove(last)
|
||||
self._update_sequence_duration()
|
||||
|
||||
def delete_clip(
|
||||
self,
|
||||
clip_ids: List[str],
|
||||
ripple: bool = True
|
||||
) -> None:
|
||||
"""
|
||||
Delete clips from timeline.
|
||||
|
||||
Uses spine iteration instead of the name-indexed dict so that
|
||||
duplicate-named clips (e.g. four ``Interview_A``) are resolved
|
||||
correctly — always targeting the *first* spine match rather than
|
||||
the last-indexed entry.
|
||||
|
||||
Args:
|
||||
clip_ids: Clips to delete
|
||||
ripple: If True, shift subsequent clips. If False, leave gaps.
|
||||
"""
|
||||
spine = self._get_spine()
|
||||
|
||||
for clip_id in clip_ids:
|
||||
# Walk spine directly to find the first clip matching this name,
|
||||
# avoiding the last-one-wins problem in self.clips.
|
||||
target = None
|
||||
for _spine_idx, spine_clip in self._iter_spine_clips():
|
||||
name = spine_clip.get('id') or spine_clip.get('name') or ''
|
||||
if name == clip_id:
|
||||
target = spine_clip
|
||||
break
|
||||
|
||||
if target is None:
|
||||
continue
|
||||
|
||||
_, clip_duration, clip_offset = self._get_clip_times(target)
|
||||
clip_index = list(spine).index(target)
|
||||
|
||||
if ripple:
|
||||
spine.remove(target)
|
||||
self._ripple_from_index(
|
||||
spine, clip_index, TimeValue.zero() - clip_duration
|
||||
)
|
||||
else:
|
||||
# Replace with gap
|
||||
gap = ET.Element('gap')
|
||||
gap.set('name', 'Gap')
|
||||
gap.set('offset', clip_offset.to_fcpxml())
|
||||
gap.set('duration', clip_duration.to_fcpxml())
|
||||
|
||||
spine.remove(target)
|
||||
spine.insert(clip_index, gap)
|
||||
|
||||
# Re-index: if other spine clips share this name, point the
|
||||
# dict entry at the next one; otherwise remove entirely.
|
||||
remaining = [
|
||||
sc for _, sc in self._iter_spine_clips()
|
||||
if (sc.get('id') or sc.get('name') or '') == clip_id
|
||||
]
|
||||
if remaining:
|
||||
self.clips[clip_id] = remaining[0]
|
||||
else:
|
||||
self.clips.pop(clip_id, None)
|
||||
|
||||
# ========================================================================
|
||||
@@ -0,0 +1,170 @@
|
||||
"""Escrita do documento FCPXML: assets de vídeo, timebases e serialização.
|
||||
|
||||
Extraído de writer.py — ver fcpxml/writer/__init__.py para o conjunto.
|
||||
"""
|
||||
|
||||
import logging
|
||||
import subprocess
|
||||
import xml.etree.ElementTree as ET
|
||||
from pathlib import Path
|
||||
from typing import Optional
|
||||
|
||||
from ..models import (
|
||||
TimeValue,
|
||||
)
|
||||
from .validation import validate_fcpxml
|
||||
|
||||
_log = logging.getLogger(__name__)
|
||||
|
||||
# ============================================================================
|
||||
# STILL IMAGE AUTO-CONVERSION (v0.6.0)
|
||||
# ============================================================================
|
||||
|
||||
_STILL_IMAGE_EXTENSIONS = {'.png', '.jpg', '.jpeg', '.tiff', '.tif', '.bmp'}
|
||||
|
||||
|
||||
def _ensure_video_asset(
|
||||
src_path: str,
|
||||
duration: float = 10.0,
|
||||
fps: int = 24,
|
||||
width: int = 1920,
|
||||
height: int = 1080,
|
||||
) -> str:
|
||||
"""Convert a still image to a video file if needed.
|
||||
|
||||
Detects still images by extension and converts them to MOV using ffmpeg.
|
||||
Video files are returned as-is.
|
||||
|
||||
Args:
|
||||
src_path: Path to the source media file.
|
||||
duration: Duration in seconds for the still-to-video conversion.
|
||||
fps: Frame rate for the output video.
|
||||
width: Output width (even number).
|
||||
height: Output height (even number).
|
||||
|
||||
Returns:
|
||||
Path to the video file (original path if already video, new .mov path
|
||||
if converted from still).
|
||||
|
||||
Raises:
|
||||
FileNotFoundError: If ffmpeg is not installed.
|
||||
"""
|
||||
# Validate numeric parameters to prevent ffmpeg abuse / resource exhaustion.
|
||||
if not isinstance(duration, (int, float)) or duration <= 0 or duration > 3600:
|
||||
raise ValueError(f"duration must be 0 < d <= 3600, got {duration!r}")
|
||||
if not isinstance(fps, int) or fps < 1 or fps > 240:
|
||||
raise ValueError(f"fps must be 1–240, got {fps!r}")
|
||||
if not isinstance(width, int) or width < 2 or width > 7680 or width % 2:
|
||||
raise ValueError(f"width must be even, 2–7680, got {width!r}")
|
||||
if not isinstance(height, int) or height < 2 or height > 4320 or height % 2:
|
||||
raise ValueError(f"height must be even, 2–4320, got {height!r}")
|
||||
|
||||
path = Path(src_path)
|
||||
if path.suffix.lower() not in _STILL_IMAGE_EXTENSIONS:
|
||||
return src_path
|
||||
|
||||
output_path = path.with_suffix('.mov')
|
||||
if output_path.exists():
|
||||
return str(output_path)
|
||||
|
||||
# Build ffmpeg command: still image → video with specified duration
|
||||
cmd = [
|
||||
'ffmpeg', '-y',
|
||||
'-loop', '1',
|
||||
'-i', str(path),
|
||||
'-c:v', 'prores_ks',
|
||||
'-profile:v', '0',
|
||||
'-t', str(duration),
|
||||
'-r', str(fps),
|
||||
'-vf', f'scale={width}:{height}:force_original_aspect_ratio=decrease,'
|
||||
f'pad={width}:{height}:(ow-iw)/2:(oh-ih)/2',
|
||||
'-pix_fmt', 'yuva444p10le',
|
||||
str(output_path),
|
||||
]
|
||||
try:
|
||||
subprocess.run(cmd, check=True, capture_output=True, timeout=120)
|
||||
except FileNotFoundError:
|
||||
raise FileNotFoundError(
|
||||
"ffmpeg not found. Install ffmpeg to use still image auto-conversion: "
|
||||
"brew install ffmpeg"
|
||||
)
|
||||
except subprocess.TimeoutExpired:
|
||||
raise RuntimeError(
|
||||
f"Image conversion timed out after 120s: {path}"
|
||||
)
|
||||
except subprocess.CalledProcessError as e:
|
||||
stderr_msg = e.stderr.decode(errors='replace') if e.stderr else str(e)
|
||||
raise RuntimeError(f"ffmpeg conversion failed: {stderr_msg}")
|
||||
return str(output_path)
|
||||
|
||||
|
||||
def _enforce_standard_timebases(root: ET.Element) -> None:
|
||||
"""Walk all elements and snap time attributes to standard FCPXML timebases.
|
||||
|
||||
Targets offset, start, duration, and tcStart attributes. Values that
|
||||
already use a standard denominator are left untouched.
|
||||
"""
|
||||
time_attrs = ('offset', 'start', 'duration', 'tcStart')
|
||||
for elem in root.iter():
|
||||
for attr in time_attrs:
|
||||
val = elem.get(attr)
|
||||
if val and val.endswith('s') and '/' in val:
|
||||
try:
|
||||
tv = TimeValue.from_timecode(val)
|
||||
if not tv.is_standard_timebase():
|
||||
# Snap to nearest frame at 2400 ticks/sec
|
||||
snapped = tv.snap_to_frame(24)
|
||||
elem.set(attr, snapped.to_fcpxml())
|
||||
except (ValueError, ZeroDivisionError):
|
||||
pass # Skip unparseable values
|
||||
|
||||
|
||||
def write_fcpxml(
|
||||
root: ET.Element,
|
||||
filepath: str,
|
||||
enforce_timebases: bool = False,
|
||||
strict: bool = False,
|
||||
fps: Optional[float] = None,
|
||||
) -> str:
|
||||
"""Format an ElementTree root as pretty-printed FCPXML and write to disk.
|
||||
|
||||
Handles XML declaration, DOCTYPE insertion, and blank-line cleanup
|
||||
consistently across all FCPXML output paths (modifier, writer, rough cut).
|
||||
|
||||
Args:
|
||||
root: The <fcpxml> root Element to serialize.
|
||||
filepath: Destination file path.
|
||||
enforce_timebases: If True, snap all time values to standard FCPXML
|
||||
timebases before writing. Default False for backward compat.
|
||||
strict: If True, raise ValueError on validation errors.
|
||||
If False (default), log warnings.
|
||||
fps: Frame rate for the frame-alignment validation check. Defaults
|
||||
to 24 when omitted — pass the sequence's real (float) rate so
|
||||
NTSC projects (23.976/29.97/59.94fps) don't get spurious
|
||||
"not frame-aligned at 24fps" warnings for values that are
|
||||
exactly aligned at their own true rate.
|
||||
|
||||
Returns:
|
||||
The filepath written to.
|
||||
"""
|
||||
if enforce_timebases:
|
||||
_enforce_standard_timebases(root)
|
||||
|
||||
# Auto-validate before writing
|
||||
issues = validate_fcpxml(root, fps=fps if fps is not None else 24.0)
|
||||
if issues:
|
||||
errors = [i for i in issues if i.severity == "error"]
|
||||
warnings = [i for i in issues if i.severity == "warning"]
|
||||
for w in warnings:
|
||||
_log.warning("FCPXML validation: %s", w.message)
|
||||
if errors and strict:
|
||||
msg = "; ".join(e.message for e in errors)
|
||||
raise ValueError(f"FCPXML validation failed: {msg}")
|
||||
for e in errors:
|
||||
_log.error("FCPXML validation: %s", e.message)
|
||||
|
||||
from ..safe_xml import serialize_xml
|
||||
|
||||
return serialize_xml(root, filepath, doctype='<!DOCTYPE fcpxml>')
|
||||
|
||||
|
||||
@@ -0,0 +1,147 @@
|
||||
"""FCPXMLWriter: gera um documento novo a partir de objetos Python.
|
||||
|
||||
Extraído de writer.py — ver fcpxml/writer/__init__.py para o conjunto.
|
||||
"""
|
||||
|
||||
import uuid
|
||||
import xml.etree.ElementTree as ET
|
||||
from datetime import datetime
|
||||
|
||||
from ..models import (
|
||||
Marker,
|
||||
Project,
|
||||
Timecode,
|
||||
)
|
||||
from .document import write_fcpxml
|
||||
from .helpers import build_marker_element
|
||||
|
||||
# ============================================================================
|
||||
# FCPXML GENERATOR - Create from Python objects
|
||||
# ============================================================================
|
||||
|
||||
class FCPXMLWriter:
|
||||
"""Generate a new FCPXML document from Python dataclass objects.
|
||||
|
||||
Converts a ``Project`` (containing ``Timeline`` → ``Clip`` → ``Marker``
|
||||
hierarchies) into a spec-compliant FCPXML v1.11 element tree and writes
|
||||
it to disk. Used by ``RoughCutGenerator`` and the ``generate_*`` MCP
|
||||
tools to create fresh timelines from scratch.
|
||||
|
||||
Unlike ``FCPXMLModifier`` (which mutates existing XML), this class
|
||||
*creates* XML from structured Python objects.
|
||||
|
||||
Example::
|
||||
|
||||
from fcpxml.models import Project, Timeline, Clip, Timecode
|
||||
project = Project(name="My Edit", timelines=[...])
|
||||
writer = FCPXMLWriter()
|
||||
writer.write_project(project, "output.fcpxml")
|
||||
"""
|
||||
|
||||
def __init__(self, version: str = "1.13"):
|
||||
"""Initialize writer targeting the given FCPXML version."""
|
||||
self.version = version
|
||||
self.resource_counter = 1
|
||||
|
||||
def _next_resource_id(self) -> str:
|
||||
"""Return an auto-incrementing resource ID (r1, r2, ...)."""
|
||||
rid = f"r{self.resource_counter}"
|
||||
self.resource_counter += 1
|
||||
return rid
|
||||
|
||||
def _generate_uid(self) -> str:
|
||||
"""Generate a unique identifier for FCPXML elements."""
|
||||
return str(uuid.uuid4()).upper()
|
||||
|
||||
def _tc_to_rational(self, tc: Timecode) -> str:
|
||||
"""Convert a Timecode to FCPXML rational time string (e.g. '48/24s')."""
|
||||
return f"{tc.frames}/{int(tc.frame_rate)}s"
|
||||
|
||||
def write_project(self, project: Project, filepath: str):
|
||||
"""Write a project to an FCPXML file."""
|
||||
root = self._build_fcpxml(project)
|
||||
write_fcpxml(root, filepath)
|
||||
|
||||
def _build_fcpxml(self, project: Project) -> ET.Element:
|
||||
"""Build the full FCPXML element tree: fcpxml > resources + library > event > project."""
|
||||
root = ET.Element('fcpxml', version=self.version)
|
||||
resources = ET.SubElement(root, 'resources')
|
||||
resource_map = {}
|
||||
|
||||
if project.timelines:
|
||||
timeline = project.timelines[0]
|
||||
format_id = self._next_resource_id()
|
||||
ET.SubElement(resources, 'format',
|
||||
id=format_id,
|
||||
name=f"FFVideoFormat{timeline.height}p{int(timeline.frame_rate)}",
|
||||
frameDuration=f"1/{int(timeline.frame_rate)}s",
|
||||
width=str(timeline.width), height=str(timeline.height))
|
||||
resource_map['_format'] = format_id
|
||||
|
||||
library = ET.SubElement(root, 'library',
|
||||
location=f"file:///Users/editor/Movies/{project.name}.fcpbundle/")
|
||||
event = ET.SubElement(library, 'event', name=project.name, uid=self._generate_uid())
|
||||
|
||||
for timeline in project.timelines:
|
||||
self._add_timeline(event, timeline, resources, resource_map)
|
||||
return root
|
||||
|
||||
def _add_timeline(self, event, timeline, resources, resource_map):
|
||||
"""Add a timeline as a project > sequence > spine structure under the event."""
|
||||
project_elem = ET.SubElement(event, 'project',
|
||||
name=timeline.name, uid=self._generate_uid(),
|
||||
modDate=datetime.now().strftime("%Y-%m-%d %H:%M:%S -0500"))
|
||||
|
||||
format_id = resource_map.get('_format', 'r1')
|
||||
sequence = ET.SubElement(project_elem, 'sequence',
|
||||
format=format_id, duration=self._tc_to_rational(timeline.duration),
|
||||
tcStart="0s", tcFormat="NDF", audioLayout="stereo", audioRate="48k")
|
||||
|
||||
spine = ET.SubElement(sequence, 'spine')
|
||||
for clip in timeline.clips:
|
||||
self._add_clip(spine, clip, resources, resource_map)
|
||||
for marker in timeline.markers:
|
||||
self._add_marker(sequence, marker)
|
||||
|
||||
def _add_clip(self, spine, clip, resources, resource_map):
|
||||
"""Add a clip as an asset-clip element, creating its asset resource if needed."""
|
||||
if clip.media_path and clip.media_path not in resource_map:
|
||||
asset_id = self._next_resource_id()
|
||||
ET.SubElement(resources, 'asset', id=asset_id, name=clip.name,
|
||||
uid=self._generate_uid(), src=clip.media_path, start="0s",
|
||||
duration=self._tc_to_rational(clip.duration), hasVideo="1", hasAudio="1")
|
||||
resource_map[clip.media_path] = asset_id
|
||||
|
||||
asset_id = resource_map.get(clip.media_path, 'r1')
|
||||
format_id = resource_map.get('_format', 'r1')
|
||||
clip_elem = ET.SubElement(spine, 'asset-clip',
|
||||
ref=asset_id, offset=self._tc_to_rational(clip.start), name=clip.name,
|
||||
start=self._tc_to_rational(clip.source_start) if clip.source_start else "0s",
|
||||
duration=self._tc_to_rational(clip.duration), format=format_id, tcFormat="NDF")
|
||||
|
||||
for marker in clip.markers:
|
||||
self._add_marker(clip_elem, marker)
|
||||
for keyword in clip.keywords:
|
||||
self._add_keyword(clip_elem, keyword)
|
||||
|
||||
def _add_marker(self, parent: ET.Element, marker: Marker):
|
||||
"""Add a marker or chapter-marker element to a parent clip or sequence."""
|
||||
build_marker_element(
|
||||
parent=parent,
|
||||
marker_type=marker.marker_type,
|
||||
start=self._tc_to_rational(marker.start),
|
||||
duration=self._tc_to_rational(marker.duration) if marker.duration else "1/24s",
|
||||
name=marker.name,
|
||||
note=marker.note or None,
|
||||
)
|
||||
|
||||
def _add_keyword(self, parent, keyword):
|
||||
"""Add a keyword element with optional start/duration range to a parent clip."""
|
||||
attrs = {'value': keyword.value}
|
||||
if keyword.start:
|
||||
attrs['start'] = self._tc_to_rational(keyword.start)
|
||||
if keyword.duration:
|
||||
attrs['duration'] = self._tc_to_rational(keyword.duration)
|
||||
ET.SubElement(parent, 'keyword', **attrs)
|
||||
|
||||
|
||||
@@ -0,0 +1,279 @@
|
||||
"""Ajudantes de nível de módulo do writer: sanitização, escalas, elementos base.
|
||||
|
||||
Extraído de writer.py — ver fcpxml/writer/__init__.py para o conjunto.
|
||||
"""
|
||||
|
||||
import subprocess
|
||||
import xml.etree.ElementTree as ET
|
||||
from pathlib import Path
|
||||
from typing import Any, Dict, List, Optional
|
||||
|
||||
from ..models import (
|
||||
MarkerType,
|
||||
)
|
||||
|
||||
# Maximum lengths for XML attribute values to prevent memory abuse
|
||||
_MAX_MARKER_NAME_LENGTH = 1024
|
||||
_MAX_NOTE_LENGTH = 4096
|
||||
|
||||
# ============================================================================
|
||||
# EFFECT RESOURCE REGISTRY (v0.6.0)
|
||||
# ============================================================================
|
||||
|
||||
# FCP built-in transition/filter effect UUIDs extracted from Filters.bundle.
|
||||
# Maps slug → (display_name, uuid).
|
||||
FCP_EFFECTS: Dict[str, tuple] = {
|
||||
# Dissolves
|
||||
'cross-dissolve': ('Cross Dissolve', '4731E73A-8DAC-4113-9A30-AE85B1761265'),
|
||||
'fade': ('Fade', '8154D0DA-C99B-4EF8-8FF8-006FE5ED57F1'),
|
||||
'dip-to-color': ('Dip to Color', 'F779C565-486D-4633-8035-0374B4DB8F5C'),
|
||||
'noise-dissolve': ('Noise Dissolve', 'ABFED81E-35D9-429C-AB47-438C1FB5D9DE'),
|
||||
# Wipes
|
||||
'edge-wipe': ('Edge Wipe', '857E2FBA-98DB-411B-A88C-CE6ABC1F65D8'),
|
||||
'slide': ('Slide', '6AAB0D54-FCD8-4EBD-A62D-D352A5ED1648'),
|
||||
'band-wipe': ('Band Wipe', 'A4E0B8E4-E916-474B-A14C-E3A9E0B1A3C1'),
|
||||
'center-wipe': ('Center Wipe', 'B3F2D4A1-7C8E-4B9D-A5F6-D1E2C3B4A5D6'),
|
||||
'checker-wipe': ('Checker Wipe', 'C4D3E2F1-8A7B-4C6D-B5E4-F2A1D3C4B5E6'),
|
||||
'clock-wipe': ('Clock Wipe', 'D5E4F3A2-9B8C-4D7E-C6F5-A3B2E4D5C6F7'),
|
||||
'gradient-wipe': ('Gradient Wipe', 'E6F5A4B3-AC9D-4E8F-D7A6-B4C3F5E6D7A8'),
|
||||
'inset-wipe': ('Inset Wipe', 'F7A6B5C4-BD0E-4F9A-E8B7-C5D4A6F7E8B9'),
|
||||
'star-wipe': ('Star Wipe', 'A8B7C6D5-CE1F-4A0B-F9C8-D6E5B7A8F9C0'),
|
||||
# Legacy aliases — map common shorthand to canonical slugs
|
||||
'fade-to-black': ('Fade', '8154D0DA-C99B-4EF8-8FF8-006FE5ED57F1'),
|
||||
'fade-from-black': ('Fade', '8154D0DA-C99B-4EF8-8FF8-006FE5ED57F1'),
|
||||
'wipe': ('Edge Wipe', '857E2FBA-98DB-411B-A88C-CE6ABC1F65D8'),
|
||||
'dissolve': ('Cross Dissolve', '4731E73A-8DAC-4113-9A30-AE85B1761265'),
|
||||
}
|
||||
|
||||
|
||||
def list_effects() -> List[Dict[str, str]]:
|
||||
"""Return a list of all available FCP transition effects.
|
||||
|
||||
Each entry contains slug, display_name, and uuid.
|
||||
Legacy aliases are excluded to avoid duplicates.
|
||||
"""
|
||||
seen_uuids: set = set()
|
||||
effects = []
|
||||
for slug, (name, uid) in FCP_EFFECTS.items():
|
||||
if uid in seen_uuids:
|
||||
continue
|
||||
seen_uuids.add(uid)
|
||||
effects.append({'slug': slug, 'name': name, 'uuid': uid})
|
||||
return effects
|
||||
|
||||
# Named constants for clip-tag sets used across operations.
|
||||
# Using named tuples prevents inconsistent ad-hoc tag lists and ensures
|
||||
# new clip types only need adding in one place.
|
||||
CLIP_TAGS = ('clip', 'asset-clip', 'video', 'ref-clip')
|
||||
CLIP_AND_AUDIO_TAGS = ('clip', 'asset-clip', 'video', 'audio', 'ref-clip')
|
||||
SPINE_ELEMENT_TAGS = ('clip', 'asset-clip', 'video', 'audio', 'gap', 'transition', 'ref-clip')
|
||||
|
||||
|
||||
def _sanitize_xml_value(value: str, max_length: int = _MAX_MARKER_NAME_LENGTH) -> str:
|
||||
"""Sanitize a string value before writing it into an XML attribute.
|
||||
|
||||
Strips null bytes, control characters (except tab/newline/CR), and
|
||||
enforces a length limit to prevent memory abuse or malformed XML.
|
||||
"""
|
||||
if not isinstance(value, str):
|
||||
return str(value)
|
||||
# Remove null bytes and non-printable control characters
|
||||
cleaned = ''.join(
|
||||
c for c in value
|
||||
if c in ('\t', '\n', '\r') or ord(c) >= 32
|
||||
)
|
||||
if len(cleaned) > max_length:
|
||||
cleaned = cleaned[:max_length]
|
||||
return cleaned
|
||||
|
||||
|
||||
# FCPXML DTD child element ordering for asset-clip / clip elements.
|
||||
# Elements MUST appear in this order for DTD validation.
|
||||
# See: https://developer.apple.com/documentation/professional-video-applications/fcpxml-reference
|
||||
_ASSET_CLIP_CHILD_ORDER = [
|
||||
'note',
|
||||
'conform-rate', 'timeMap',
|
||||
'adjust-crop', 'adjust-corners', 'adjust-conform', 'adjust-transform',
|
||||
'adjust-blend', 'adjust-stabilization', 'adjust-rollingShutter',
|
||||
'adjust-360-transform', 'adjust-reorient', 'adjust-orientation',
|
||||
'adjust-volume', 'adjust-panner',
|
||||
# anchor items (connected clips, titles, etc.)
|
||||
'audio', 'video', 'clip', 'title', 'caption',
|
||||
'mc-clip', 'ref-clip', 'sync-clip', 'asset-clip', 'audition', 'spine',
|
||||
# marker items
|
||||
'marker', 'chapter-marker', 'rating', 'keyword', 'analysis-marker',
|
||||
# trailing
|
||||
'audio-channel-source',
|
||||
'filter-video', 'filter-video-mask',
|
||||
'filter-audio',
|
||||
'metadata',
|
||||
]
|
||||
|
||||
# Build a priority lookup: tag → index for fast comparison
|
||||
_CHILD_ORDER_INDEX = {tag: i for i, tag in enumerate(_ASSET_CLIP_CHILD_ORDER)}
|
||||
|
||||
|
||||
# How close to the end of a clip a zoom must finish for the return to be
|
||||
# skipped. Within this margin the cut arrives before the eye registers the
|
||||
# move back, so the return reads as a twitch rather than a resolution.
|
||||
HOLD_AT_CUT_THRESHOLD = 1.0
|
||||
|
||||
# How close to the start of a clip a zoom must begin for the ramp-in to be
|
||||
# skipped and the shot to simply open already zoomed. Tighter than the end
|
||||
# margin on purpose: at the end the cut hides an unfinished return, but at
|
||||
# the start a ramp is visible from frame one and reads as the shot settling.
|
||||
START_AT_CUT_THRESHOLD = 0.5
|
||||
|
||||
|
||||
def _fmt_scale(value: float) -> str:
|
||||
"""Format a scale factor without trailing float noise (1.0 -> "1")."""
|
||||
return f"{value:.6f}".rstrip("0").rstrip(".") or "0"
|
||||
|
||||
|
||||
def _dtd_insert(parent: ET.Element, child: ET.Element) -> ET.Element:
|
||||
"""Insert a child element into parent at the correct DTD-ordered position.
|
||||
|
||||
Instead of blindly appending (which can violate DTD ordering),
|
||||
this finds the right insertion point based on the FCPXML DTD's
|
||||
required element sequence for asset-clip / clip elements.
|
||||
|
||||
Unknown tags are appended at the end.
|
||||
"""
|
||||
child_priority = _CHILD_ORDER_INDEX.get(child.tag, len(_ASSET_CLIP_CHILD_ORDER))
|
||||
|
||||
# Find the first existing child whose priority is greater than ours
|
||||
insert_idx = len(parent)
|
||||
for i, existing in enumerate(parent):
|
||||
existing_priority = _CHILD_ORDER_INDEX.get(existing.tag, len(_ASSET_CLIP_CHILD_ORDER))
|
||||
if existing_priority > child_priority:
|
||||
insert_idx = i
|
||||
break
|
||||
|
||||
parent.insert(insert_idx, child)
|
||||
return child
|
||||
|
||||
|
||||
def build_marker_element(
|
||||
parent: ET.Element,
|
||||
marker_type: MarkerType,
|
||||
start: str,
|
||||
duration: str,
|
||||
name: str,
|
||||
note: Optional[str] = None,
|
||||
) -> ET.Element:
|
||||
"""Create a marker or chapter-marker XML element under *parent*.
|
||||
|
||||
Single source of truth for marker element construction — used by both
|
||||
FCPXMLModifier (edit-existing workflow) and FCPXMLWriter (generate-new
|
||||
workflow). Centralises tag selection, type-specific attributes, note
|
||||
guards, and input sanitization so changes only need to happen once.
|
||||
"""
|
||||
elem = ET.Element(marker_type.xml_tag)
|
||||
elem.set('start', start)
|
||||
elem.set('duration', duration)
|
||||
elem.set('value', _sanitize_xml_value(name, _MAX_MARKER_NAME_LENGTH))
|
||||
for attr, val in marker_type.xml_attrs.items():
|
||||
elem.set(attr, val)
|
||||
if note and marker_type != MarkerType.CHAPTER:
|
||||
elem.set('note', _sanitize_xml_value(note, _MAX_NOTE_LENGTH))
|
||||
_dtd_insert(parent, elem)
|
||||
return elem
|
||||
|
||||
|
||||
def _create_asset_element(
|
||||
resources: ET.Element,
|
||||
asset_id: str,
|
||||
name: str,
|
||||
src: str,
|
||||
duration: str = "0s",
|
||||
start: str = "0s",
|
||||
has_video: str = "1",
|
||||
has_audio: str = "1",
|
||||
uid: Optional[str] = None,
|
||||
) -> ET.Element:
|
||||
"""Create an <asset> element with <media-rep> child instead of src attribute.
|
||||
|
||||
FCP's DTD prefers <media-rep kind="original-media" src="..."/> children
|
||||
over the src attribute on <asset>. This helper produces the preferred form.
|
||||
|
||||
Args:
|
||||
resources: Parent <resources> element to append to.
|
||||
asset_id: Resource ID (e.g. "r3").
|
||||
name: Human-readable asset name.
|
||||
src: File path or URL for the media source.
|
||||
duration: Asset duration in FCPXML rational format.
|
||||
start: Asset start time.
|
||||
has_video: "1" if asset has video track.
|
||||
has_audio: "1" if asset has audio track.
|
||||
uid: Optional UUID; auto-generated if not provided.
|
||||
|
||||
Returns:
|
||||
The created <asset> Element.
|
||||
"""
|
||||
import uuid as _uuid
|
||||
asset = ET.SubElement(resources, 'asset')
|
||||
asset.set('id', asset_id)
|
||||
asset.set('name', _sanitize_xml_value(name, 512))
|
||||
asset.set('uid', uid or str(_uuid.uuid4()).upper())
|
||||
asset.set('start', start)
|
||||
asset.set('duration', duration)
|
||||
asset.set('hasVideo', has_video)
|
||||
asset.set('hasAudio', has_audio)
|
||||
# Use media-rep child instead of src attribute
|
||||
media_rep = ET.SubElement(asset, 'media-rep')
|
||||
media_rep.set('kind', 'original-media')
|
||||
media_rep.set('src', src)
|
||||
return asset
|
||||
|
||||
|
||||
def _probe_audio_info(src: str) -> Optional[Dict[str, Any]]:
|
||||
"""Probe an audio file for its real duration, sample rate, and channels.
|
||||
|
||||
Tries ffprobe first, then falls back to the stdlib ``wave`` module for
|
||||
.wav files. Returns ``None`` when the file can't be probed, so callers
|
||||
can fall back to caller-supplied durations.
|
||||
|
||||
Returns:
|
||||
``{'duration': float, 'sample_rate': int, 'channels': int}`` or None.
|
||||
"""
|
||||
path = Path(src)
|
||||
if not path.is_file():
|
||||
return None
|
||||
try:
|
||||
result = subprocess.run(
|
||||
['ffprobe', '-v', 'error', '-select_streams', 'a:0',
|
||||
'-show_entries', 'stream=sample_rate,channels,duration',
|
||||
'-show_entries', 'format=duration',
|
||||
'-of', 'json', str(path)],
|
||||
capture_output=True, text=True, timeout=15,
|
||||
)
|
||||
if result.returncode == 0:
|
||||
import json
|
||||
data = json.loads(result.stdout)
|
||||
streams = data.get('streams') or [{}]
|
||||
stream = streams[0]
|
||||
duration = stream.get('duration') or data.get('format', {}).get('duration')
|
||||
if duration:
|
||||
return {
|
||||
'duration': float(duration),
|
||||
'sample_rate': int(stream.get('sample_rate') or 48000),
|
||||
'channels': int(stream.get('channels') or 2),
|
||||
}
|
||||
except (OSError, subprocess.TimeoutExpired, ValueError):
|
||||
pass
|
||||
if path.suffix.lower() == '.wav':
|
||||
try:
|
||||
import wave
|
||||
with wave.open(str(path), 'rb') as wf:
|
||||
rate = wf.getframerate()
|
||||
if rate > 0:
|
||||
return {
|
||||
'duration': wf.getnframes() / rate,
|
||||
'sample_rate': rate,
|
||||
'channels': wf.getnchannels(),
|
||||
}
|
||||
except (OSError, wave.Error, EOFError):
|
||||
pass
|
||||
return None
|
||||
|
||||
|
||||
@@ -0,0 +1,78 @@
|
||||
"""Inserir clipes na spine.
|
||||
|
||||
Extraído de writer.py — ver fcpxml/writer/__init__.py para o conjunto.
|
||||
"""
|
||||
|
||||
import xml.etree.ElementTree as ET
|
||||
from typing import Optional
|
||||
|
||||
|
||||
class InsertMixin:
|
||||
"""Inserir clipes na spine."""
|
||||
|
||||
# INSERT CLIP OPERATIONS
|
||||
# ========================================================================
|
||||
|
||||
def insert_clip(
|
||||
self,
|
||||
position: str,
|
||||
asset_id: Optional[str] = None,
|
||||
asset_name: Optional[str] = None,
|
||||
duration: Optional[str] = None,
|
||||
in_point: Optional[str] = None,
|
||||
out_point: Optional[str] = None,
|
||||
ripple: bool = True
|
||||
) -> ET.Element:
|
||||
"""
|
||||
Insert a library clip onto the timeline.
|
||||
|
||||
Args:
|
||||
position: Where to insert - 'start', 'end', timecode, or 'after:clip_id'
|
||||
asset_id: Asset reference ID (e.g., 'r3')
|
||||
asset_name: Asset name (alternative to asset_id)
|
||||
duration: Duration of clip (if not using in/out points)
|
||||
in_point: Source in-point for subclip
|
||||
out_point: Source out-point for subclip
|
||||
ripple: Whether to shift subsequent clips
|
||||
|
||||
Returns:
|
||||
The created clip element
|
||||
"""
|
||||
asset, asset_id = self._resolve_asset(asset_id, asset_name)
|
||||
clip_duration, source_start = self._resolve_clip_duration(
|
||||
asset, duration, in_point, out_point
|
||||
)
|
||||
|
||||
# Get spine and calculate insert position
|
||||
spine = self._get_spine()
|
||||
spine_children = list(spine)
|
||||
target_offset, insert_index = self._resolve_insert_position(
|
||||
position, spine_children
|
||||
)
|
||||
|
||||
# Build extra attrs — include format from first available format
|
||||
extra: dict[str, str] = {}
|
||||
for fmt_id in self.formats:
|
||||
extra['format'] = fmt_id
|
||||
break
|
||||
|
||||
new_clip = self._make_asset_clip(
|
||||
asset_id, asset.get('name', 'Untitled'),
|
||||
target_offset, source_start, clip_duration,
|
||||
**extra,
|
||||
)
|
||||
|
||||
# Insert into spine
|
||||
spine.insert(insert_index, new_clip)
|
||||
|
||||
# Ripple subsequent clips if needed
|
||||
if ripple and insert_index < len(spine_children):
|
||||
self._ripple_from_index(spine, insert_index + 1, clip_duration)
|
||||
|
||||
# Add to clip index
|
||||
clip_id = f"inserted_{len(self.clips)}"
|
||||
self.clips[clip_id] = new_clip
|
||||
|
||||
return new_clip
|
||||
|
||||
# ========================================================================
|
||||
@@ -0,0 +1,165 @@
|
||||
"""Marcadores: um, por timecode, e em lote.
|
||||
|
||||
Extraído de writer.py — ver fcpxml/writer/__init__.py para o conjunto.
|
||||
"""
|
||||
|
||||
import xml.etree.ElementTree as ET
|
||||
from typing import Any, Dict, List, Optional
|
||||
|
||||
from ..models import (
|
||||
MarkerColor,
|
||||
MarkerType,
|
||||
TimeValue,
|
||||
)
|
||||
from .helpers import build_marker_element
|
||||
|
||||
|
||||
class MarkersMixin:
|
||||
"""Marcadores: um, por timecode, e em lote."""
|
||||
|
||||
# ========================================================================
|
||||
# MARKER OPERATIONS
|
||||
# ========================================================================
|
||||
|
||||
def add_marker(
|
||||
self,
|
||||
clip_id: 'str | ET.Element',
|
||||
timecode: str,
|
||||
name: str,
|
||||
marker_type: "MarkerType | str" = MarkerType.STANDARD,
|
||||
color: Optional[MarkerColor] = None,
|
||||
note: Optional[str] = None
|
||||
) -> ET.Element:
|
||||
"""
|
||||
Add a marker to a clip.
|
||||
|
||||
Args:
|
||||
clip_id: Target clip identifier (name or ID)
|
||||
timecode: Position within clip (relative to clip start)
|
||||
name: Marker label
|
||||
marker_type: STANDARD, TODO, COMPLETED, or CHAPTER (enum or string)
|
||||
color: Optional marker color
|
||||
note: Optional marker note
|
||||
|
||||
Returns:
|
||||
The created marker element
|
||||
"""
|
||||
clip = self._require_clip(clip_id)
|
||||
|
||||
if isinstance(marker_type, str):
|
||||
marker_type = MarkerType.from_string(marker_type)
|
||||
|
||||
time_value = self._parse_time(timecode)
|
||||
|
||||
return build_marker_element(
|
||||
parent=clip,
|
||||
marker_type=marker_type,
|
||||
start=time_value.to_fcpxml(),
|
||||
duration=f"1/{int(self.fps)}s",
|
||||
name=name,
|
||||
note=note,
|
||||
)
|
||||
|
||||
def add_marker_at_timeline(
|
||||
self,
|
||||
timecode: str,
|
||||
name: str,
|
||||
marker_type: "MarkerType | str" = MarkerType.STANDARD,
|
||||
color: Optional[MarkerColor] = None,
|
||||
note: Optional[str] = None
|
||||
) -> ET.Element:
|
||||
"""Add a marker at a timeline position (finds the containing clip).
|
||||
|
||||
Uses ``_find_spine_clip_at_seconds`` to walk the spine directly,
|
||||
avoiding the name-indexed ``self.clips`` dict which silently drops
|
||||
duplicate-named clips.
|
||||
"""
|
||||
if isinstance(marker_type, str):
|
||||
marker_type = MarkerType.from_string(marker_type)
|
||||
time_value = self._parse_time(timecode)
|
||||
target_seconds = time_value.to_seconds()
|
||||
|
||||
clip, relative_seconds = self._find_spine_clip_at_seconds(target_seconds)
|
||||
relative_tc = TimeValue.from_seconds(relative_seconds, self.fps)
|
||||
|
||||
return build_marker_element(
|
||||
parent=clip,
|
||||
marker_type=marker_type,
|
||||
start=relative_tc.to_fcpxml(),
|
||||
duration=f"1/{int(self.fps)}s",
|
||||
name=name,
|
||||
note=note,
|
||||
)
|
||||
|
||||
def batch_add_markers(
|
||||
self,
|
||||
markers: List[Dict[str, Any]],
|
||||
auto_at_cuts: bool = False,
|
||||
auto_at_intervals: Optional[str] = None
|
||||
) -> List[ET.Element]:
|
||||
"""
|
||||
Add multiple markers at once.
|
||||
|
||||
Args:
|
||||
markers: List of marker specs [{timecode, name, marker_type, color}]
|
||||
auto_at_cuts: Add marker at every cut point
|
||||
auto_at_intervals: Add markers at regular intervals (e.g., "00:00:30:00")
|
||||
|
||||
Returns:
|
||||
List of created marker elements
|
||||
"""
|
||||
created = []
|
||||
|
||||
# Handle explicit markers
|
||||
for m in markers:
|
||||
marker = self.add_marker_at_timeline(
|
||||
timecode=m['timecode'],
|
||||
name=m['name'],
|
||||
marker_type=MarkerType.from_string(m.get('marker_type', 'standard')),
|
||||
color=MarkerColor[m['color'].upper()] if m.get('color') else None,
|
||||
note=m.get('note')
|
||||
)
|
||||
created.append(marker)
|
||||
|
||||
# Auto-detect at cuts — add a marker at the start of every spine clip.
|
||||
if auto_at_cuts:
|
||||
for i, clip in self._iter_spine_clips():
|
||||
clip_start = clip.get('start', '0s')
|
||||
marker = build_marker_element(
|
||||
parent=clip,
|
||||
marker_type=MarkerType.STANDARD,
|
||||
start=clip_start,
|
||||
duration=f"1/{int(self.fps)}s",
|
||||
name=f"Cut {i+1}",
|
||||
)
|
||||
created.append(marker)
|
||||
|
||||
# Auto-detect at intervals — place markers at regular time steps.
|
||||
if auto_at_intervals:
|
||||
interval = self._parse_time(auto_at_intervals).to_seconds()
|
||||
total_duration = self._timeline_duration().to_seconds()
|
||||
if total_duration > 0:
|
||||
|
||||
current = interval
|
||||
count = 1
|
||||
while current < total_duration:
|
||||
try:
|
||||
clip, relative = self._find_spine_clip_at_seconds(current)
|
||||
except ValueError:
|
||||
current += interval
|
||||
count += 1
|
||||
continue
|
||||
rel_tv = TimeValue.from_seconds(relative, self.fps)
|
||||
marker = build_marker_element(
|
||||
parent=clip,
|
||||
marker_type=MarkerType.STANDARD,
|
||||
start=rel_tv.to_fcpxml(),
|
||||
duration=f"1/{int(self.fps)}s",
|
||||
name=f"Marker {count}",
|
||||
)
|
||||
created.append(marker)
|
||||
current += interval
|
||||
count += 1
|
||||
|
||||
return created
|
||||
|
||||
@@ -0,0 +1,64 @@
|
||||
"""FCPXMLModifier — a edição de FCPXML montada a partir de um mixin por assunto.
|
||||
|
||||
A classe era um bloco de 3.300 linhas com dezoito assuntos dentro. Ela continua
|
||||
sendo uma classe só para quem chama — `modifier.add_marker(...)` não mudou — mas
|
||||
cada assunto agora mora no seu próprio arquivo e pode ser lido inteiro sem rolar
|
||||
por marcadores, velocidade e legendas até achar o trecho procurado.
|
||||
|
||||
Mixins em vez de objetos separados por uma razão concreta: todas essas operações
|
||||
mexem no *mesmo* documento e dependem dos mesmos índices e da mesma navegação na
|
||||
spine (`_require_clip`, `_iter_spine_clips`, `_ripple_after_clip`). Separá-las em
|
||||
objetos independentes obrigaria cada um a carregar uma referência de volta ao
|
||||
documento e transformaria toda chamada interna em travessia de fronteira, sem
|
||||
nada em troca — a divisão que importa aqui é de *leitura*, não de estado.
|
||||
|
||||
A ordem abaixo é irrelevante para o comportamento: nenhum mixin sobrescreve
|
||||
método de outro; cada um contribui com um conjunto disjunto de operações.
|
||||
"""
|
||||
|
||||
from .audio import AudioMixin
|
||||
from .compound import CompoundMixin
|
||||
from .connected import ConnectedMixin
|
||||
from .core import ModifierCore
|
||||
from .cut import CutMixin
|
||||
from .insert import InsertMixin
|
||||
from .markers import MarkersMixin
|
||||
from .rapid import RapidMixin
|
||||
from .reformat import ReformatMixin
|
||||
from .relink import RelinkMixin
|
||||
from .reorder import ReorderMixin
|
||||
from .roles import RolesMixin
|
||||
from .selection import SelectionMixin
|
||||
from .silence import SilenceMixin
|
||||
from .speed import SpeedMixin
|
||||
from .titles import TitlesMixin
|
||||
from .transitions import TransitionsMixin
|
||||
from .trim import TrimMixin
|
||||
|
||||
|
||||
class FCPXMLModifier(
|
||||
RelinkMixin,
|
||||
MarkersMixin,
|
||||
TrimMixin,
|
||||
ReorderMixin,
|
||||
TransitionsMixin,
|
||||
SpeedMixin,
|
||||
CutMixin,
|
||||
RapidMixin,
|
||||
SelectionMixin,
|
||||
InsertMixin,
|
||||
ConnectedMixin,
|
||||
TitlesMixin,
|
||||
AudioMixin,
|
||||
CompoundMixin,
|
||||
RolesMixin,
|
||||
ReformatMixin,
|
||||
SilenceMixin,
|
||||
ModifierCore,
|
||||
):
|
||||
"""Carrega um FCPXML, aplica edições cirúrgicas e salva.
|
||||
|
||||
Interface de escrita usada por todos os handlers do servidor MCP. A
|
||||
documentação de cada operação está no mixin correspondente; o
|
||||
carregamento, os índices e o `save` estão em `core.ModifierCore`.
|
||||
"""
|
||||
@@ -0,0 +1,240 @@
|
||||
"""Corte rápido: flash frames, rapid trim, preencher buracos.
|
||||
|
||||
Extraído de writer.py — ver fcpxml/writer/__init__.py para o conjunto.
|
||||
"""
|
||||
|
||||
from typing import Any, Dict, List, Optional
|
||||
|
||||
|
||||
class RapidMixin:
|
||||
"""Corte rápido: flash frames, rapid trim, preencher buracos."""
|
||||
|
||||
# SPEED CUTTING OPERATIONS (v0.3.0)
|
||||
# ========================================================================
|
||||
|
||||
def fix_flash_frames(
|
||||
self,
|
||||
mode: str = 'auto',
|
||||
threshold_frames: int = 6,
|
||||
critical_threshold_frames: int = 2
|
||||
) -> List[Dict[str, Any]]:
|
||||
"""
|
||||
Automatically fix flash frames (ultra-short clips).
|
||||
|
||||
Args:
|
||||
mode: How to fix flash frames:
|
||||
- 'extend_previous': Extend the previous clip to cover the flash frame
|
||||
- 'extend_next': Extend the next clip backward to cover the flash frame
|
||||
- 'delete': Remove the flash frame entirely (ripple)
|
||||
- 'auto': Use smart logic (extend prev for critical, delete for warning)
|
||||
threshold_frames: Frames below this are considered flash frames
|
||||
critical_threshold_frames: Frames below this are critical (default: 2)
|
||||
|
||||
Returns:
|
||||
List of fixed flash frames with details
|
||||
"""
|
||||
spine = self._get_spine()
|
||||
fixed = []
|
||||
|
||||
# Collect flash frames first (can't modify while iterating)
|
||||
flash_frames = []
|
||||
for i, clip in self._iter_spine_clips():
|
||||
duration = self._parse_time(clip.get('duration', '0s'))
|
||||
duration_frames = duration.to_frames(self.fps)
|
||||
|
||||
if duration_frames < threshold_frames:
|
||||
is_critical = duration_frames < critical_threshold_frames
|
||||
flash_frames.append({
|
||||
'index': i,
|
||||
'clip': clip,
|
||||
'clip_id': clip.get('name') or clip.get('id') or f"clip_{i}",
|
||||
'duration_frames': duration_frames,
|
||||
'is_critical': is_critical
|
||||
})
|
||||
|
||||
# Process in reverse order to maintain indices
|
||||
for ff in reversed(flash_frames):
|
||||
clip = ff['clip']
|
||||
_, _, clip_offset = self._get_clip_times(clip)
|
||||
|
||||
# Determine actual mode
|
||||
actual_mode = mode
|
||||
if mode == 'auto':
|
||||
# Critical: try to extend previous, otherwise delete
|
||||
# Warning: delete
|
||||
actual_mode = 'extend_previous' if ff['is_critical'] else 'delete'
|
||||
|
||||
result = {
|
||||
'clip_name': ff['clip_id'],
|
||||
'duration_frames': ff['duration_frames'],
|
||||
'was_critical': ff['is_critical'],
|
||||
'action': actual_mode,
|
||||
'timecode': clip_offset.to_timecode(self.fps)
|
||||
}
|
||||
|
||||
direction = {'extend_previous': 'prev', 'extend_next': 'next'}.get(actual_mode)
|
||||
if direction:
|
||||
neighbor = self._absorb_into_neighbor(spine, clip, direction)
|
||||
if neighbor is not None:
|
||||
self._recalculate_offsets(spine)
|
||||
result['extended_clip'] = neighbor.get('name', direction.title())
|
||||
else:
|
||||
spine.remove(clip)
|
||||
self._recalculate_offsets(spine)
|
||||
else: # delete
|
||||
spine.remove(clip)
|
||||
self._recalculate_offsets(spine)
|
||||
|
||||
fixed.append(result)
|
||||
|
||||
# Rebuild clip index
|
||||
self._build_clip_index()
|
||||
|
||||
return fixed
|
||||
|
||||
def rapid_trim(
|
||||
self,
|
||||
max_duration: Optional[str] = None,
|
||||
min_duration: Optional[str] = None,
|
||||
keywords: Optional[List[str]] = None,
|
||||
trim_from: str = 'end'
|
||||
) -> List[Dict[str, Any]]:
|
||||
"""
|
||||
Batch trim clips to enforce duration limits.
|
||||
|
||||
Args:
|
||||
max_duration: Maximum clip duration (e.g., '2s', '00:00:02:00')
|
||||
min_duration: Minimum clip duration (clips shorter are extended/left alone)
|
||||
keywords: Only trim clips with these keywords (None = all clips)
|
||||
trim_from: Where to trim - 'start', 'end', or 'center'
|
||||
|
||||
Returns:
|
||||
List of trimmed clips with before/after durations
|
||||
"""
|
||||
trimmed = []
|
||||
|
||||
max_dur = self._parse_time(max_duration) if max_duration else None
|
||||
min_dur = self._parse_time(min_duration) if min_duration else None
|
||||
|
||||
for _i, clip in self._iter_spine_clips():
|
||||
|
||||
clip_name = clip.get('name') or clip.get('id') or 'Unknown'
|
||||
|
||||
# Check keyword filter
|
||||
if keywords:
|
||||
clip_keywords = set()
|
||||
for kw_elem in clip.findall('keyword'):
|
||||
clip_keywords.add(kw_elem.get('value', ''))
|
||||
if not clip_keywords.intersection(set(keywords)):
|
||||
continue
|
||||
|
||||
current_start, current_duration, _ = self._get_clip_times(clip)
|
||||
original_duration = current_duration.to_seconds()
|
||||
|
||||
# Skip clips shorter than min_duration (leave them alone)
|
||||
if min_dur and current_duration < min_dur:
|
||||
continue
|
||||
|
||||
# Check max duration
|
||||
if max_dur and current_duration > max_dur:
|
||||
excess = current_duration - max_dur
|
||||
|
||||
if trim_from == 'end':
|
||||
# Keep start, reduce duration
|
||||
clip.set('duration', max_dur.to_fcpxml())
|
||||
|
||||
elif trim_from == 'start':
|
||||
# Increase start, reduce duration
|
||||
new_start = current_start + excess
|
||||
clip.set('start', new_start.to_fcpxml())
|
||||
clip.set('duration', max_dur.to_fcpxml())
|
||||
|
||||
elif trim_from == 'center':
|
||||
# Trim equal amounts from both ends
|
||||
half_excess = excess * 0.5
|
||||
new_start = current_start + half_excess
|
||||
clip.set('start', new_start.to_fcpxml())
|
||||
clip.set('duration', max_dur.to_fcpxml())
|
||||
|
||||
trimmed.append({
|
||||
'clip_name': clip_name,
|
||||
'original_duration': original_duration,
|
||||
'new_duration': max_dur.to_seconds(),
|
||||
'trim_from': trim_from,
|
||||
'action': 'trimmed'
|
||||
})
|
||||
|
||||
# Recalculate offsets
|
||||
self._recalculate_offsets(self._get_spine())
|
||||
|
||||
return trimmed
|
||||
|
||||
def fill_gaps(
|
||||
self,
|
||||
mode: str = 'extend_previous',
|
||||
max_gap: Optional[str] = None
|
||||
) -> List[Dict[str, Any]]:
|
||||
"""
|
||||
Fill gaps in the timeline.
|
||||
|
||||
Args:
|
||||
mode: How to fill gaps:
|
||||
- 'extend_previous': Extend previous clip to fill gap
|
||||
- 'extend_next': Extend next clip backward to fill gap
|
||||
- 'delete': Remove gap elements and ripple
|
||||
max_gap: Only fill gaps smaller than this (None = all gaps)
|
||||
|
||||
Returns:
|
||||
List of filled gaps with details
|
||||
"""
|
||||
spine = self._get_spine()
|
||||
filled = []
|
||||
max_gap_time = self._parse_time(max_gap) if max_gap else None
|
||||
|
||||
# Find all gaps
|
||||
gaps_to_process = []
|
||||
for i, child in enumerate(list(spine)):
|
||||
if child.tag == 'gap':
|
||||
gap_duration = self._parse_time(child.get('duration', '0s'))
|
||||
gap_offset = self._parse_time(child.get('offset', '0s'))
|
||||
|
||||
# Check max_gap filter
|
||||
if max_gap_time and gap_duration > max_gap_time:
|
||||
continue
|
||||
|
||||
gaps_to_process.append({
|
||||
'element': child,
|
||||
'index': i,
|
||||
'duration': gap_duration,
|
||||
'offset': gap_offset
|
||||
})
|
||||
|
||||
# Process in reverse to maintain indices
|
||||
for gap_info in reversed(gaps_to_process):
|
||||
gap = gap_info['element']
|
||||
gap_duration = gap_info['duration']
|
||||
gap_offset = gap_info['offset']
|
||||
|
||||
result = {
|
||||
'timecode': gap_offset.to_timecode(self.fps),
|
||||
'duration_frames': gap_duration.to_frames(self.fps),
|
||||
'duration_seconds': gap_duration.to_seconds(),
|
||||
'action': mode
|
||||
}
|
||||
|
||||
direction = {'extend_previous': 'prev', 'extend_next': 'next'}.get(mode)
|
||||
if direction:
|
||||
neighbor = self._absorb_into_neighbor(spine, gap, direction)
|
||||
if neighbor is not None:
|
||||
result['extended_clip'] = neighbor.get('name', direction.title())
|
||||
filled.append(result)
|
||||
else: # delete
|
||||
spine.remove(gap)
|
||||
filled.append(result)
|
||||
|
||||
# Recalculate offsets
|
||||
self._recalculate_offsets(spine)
|
||||
|
||||
return filled
|
||||
|
||||
# ========================================================================
|
||||
@@ -0,0 +1,43 @@
|
||||
"""Reenquadrar a resolução do projeto.
|
||||
|
||||
Extraído de writer.py — ver fcpxml/writer/__init__.py para o conjunto.
|
||||
"""
|
||||
|
||||
|
||||
|
||||
class ReformatMixin:
|
||||
"""Reenquadrar a resolução do projeto."""
|
||||
|
||||
# REFORMAT OPERATIONS (v0.5.0)
|
||||
# ========================================================================
|
||||
|
||||
SOCIAL_FORMATS = {
|
||||
"9:16": (1080, 1920),
|
||||
"1:1": (1080, 1080),
|
||||
"4:5": (1080, 1350),
|
||||
"16:9": (1920, 1080),
|
||||
"4:3": (1440, 1080),
|
||||
}
|
||||
|
||||
def reformat_resolution(self, width: int, height: int) -> None:
|
||||
"""Change the timeline format to a new resolution.
|
||||
|
||||
Updates the format resource dimensions. FCP handles spatial
|
||||
conforming (letterbox/pillarbox) on import.
|
||||
|
||||
Args:
|
||||
width: Target width in pixels
|
||||
height: Target height in pixels
|
||||
"""
|
||||
for fmt in self.root.findall('.//format'):
|
||||
fmt.set('width', str(width))
|
||||
fmt.set('height', str(height))
|
||||
old_name = fmt.get('name', '')
|
||||
if old_name:
|
||||
fmt.set('name', f"FFVideoFormat{width}x{height}")
|
||||
|
||||
sequence = self.root.find('.//sequence')
|
||||
if sequence is not None and sequence.get('format'):
|
||||
pass # format ref stays the same, dimensions updated in-place
|
||||
|
||||
# ========================================================================
|
||||
@@ -0,0 +1,94 @@
|
||||
"""Repontar a mídia de um projeto para novos arquivos.
|
||||
|
||||
Extraído de writer.py — ver fcpxml/writer/__init__.py para o conjunto.
|
||||
"""
|
||||
|
||||
from pathlib import Path
|
||||
from typing import Any, Dict
|
||||
|
||||
|
||||
class RelinkMixin:
|
||||
"""Repontar a mídia de um projeto para novos arquivos."""
|
||||
|
||||
# ========================================================================
|
||||
# MEDIA RELINK
|
||||
# ========================================================================
|
||||
|
||||
def relink_media(
|
||||
self,
|
||||
find: str,
|
||||
replace: str,
|
||||
dry_run: bool = False,
|
||||
) -> Dict[str, Any]:
|
||||
"""Bulk-rewrite media source paths (programmatic relink).
|
||||
|
||||
Rewrites the ``src`` of every ``<asset>`` / ``<media-rep>`` whose
|
||||
path starts with *find*, substituting *replace* — the standard
|
||||
technique for relinking a moved or renamed media folder without
|
||||
opening Final Cut Pro. FCP relinks via the ``media-rep`` file URL
|
||||
on import; the device-specific bookmark blob is left untouched
|
||||
(FCP regenerates it).
|
||||
|
||||
*find* / *replace* accept plain paths (``/Volumes/OldDrive``) or
|
||||
``file://`` URLs; percent-encoding in existing URLs is handled.
|
||||
Matching is prefix-based on whole path segments, so ``/Media/A``
|
||||
matches ``/Media/A/clip.mov`` but not ``/Media/AB/clip.mov``.
|
||||
|
||||
Args:
|
||||
find: Old path prefix to match.
|
||||
replace: New path prefix to substitute.
|
||||
dry_run: When True, report what would change without
|
||||
mutating the tree.
|
||||
|
||||
Returns:
|
||||
Summary dict: ``total_assets``, ``relinked`` (reference
|
||||
count), ``dry_run``, and ``changes`` — a list of
|
||||
``{asset, old, new, target_exists}`` entries
|
||||
(``target_exists`` checks the new path on this machine).
|
||||
"""
|
||||
from urllib.parse import quote, unquote, urlparse
|
||||
|
||||
def _to_path(value: str) -> str:
|
||||
if value.startswith('file://'):
|
||||
return unquote(urlparse(value).path)
|
||||
return value
|
||||
|
||||
find_path = _to_path(find).rstrip('/')
|
||||
replace_path = _to_path(replace).rstrip('/')
|
||||
if not find_path:
|
||||
raise ValueError("relink_media: 'find' must be a non-empty path prefix")
|
||||
|
||||
changes = []
|
||||
for asset_id, info in self.resources.items():
|
||||
elem = info['element']
|
||||
targets = [(elem, elem.get('src'))]
|
||||
media_rep = elem.find('media-rep')
|
||||
if media_rep is not None:
|
||||
targets.append((media_rep, media_rep.get('src')))
|
||||
|
||||
for node, old_src in targets:
|
||||
if not old_src:
|
||||
continue
|
||||
was_url = old_src.startswith('file://')
|
||||
old_path = _to_path(old_src)
|
||||
if old_path != find_path and not old_path.startswith(find_path + '/'):
|
||||
continue
|
||||
new_path = replace_path + old_path[len(find_path):]
|
||||
new_src = 'file://' + quote(new_path) if was_url else new_path
|
||||
if not dry_run:
|
||||
node.set('src', new_src)
|
||||
info['src'] = new_src
|
||||
changes.append({
|
||||
'asset': info.get('name') or asset_id,
|
||||
'old': old_src,
|
||||
'new': new_src,
|
||||
'target_exists': Path(new_path).exists(),
|
||||
})
|
||||
|
||||
return {
|
||||
'total_assets': len(self.resources),
|
||||
'relinked': len(changes),
|
||||
'dry_run': dry_run,
|
||||
'changes': changes,
|
||||
}
|
||||
|
||||
@@ -0,0 +1,126 @@
|
||||
"""Reordenar clipes e recalcular offsets/duração.
|
||||
|
||||
Extraído de writer.py — ver fcpxml/writer/__init__.py para o conjunto.
|
||||
"""
|
||||
|
||||
import xml.etree.ElementTree as ET
|
||||
from typing import List
|
||||
|
||||
from ..models import (
|
||||
TimeValue,
|
||||
)
|
||||
from .helpers import SPINE_ELEMENT_TAGS
|
||||
|
||||
|
||||
class ReorderMixin:
|
||||
"""Reordenar clipes e recalcular offsets/duração."""
|
||||
|
||||
# REORDER OPERATIONS
|
||||
# ========================================================================
|
||||
|
||||
def reorder_clips(
|
||||
self,
|
||||
clip_ids: List[str],
|
||||
target_position: str,
|
||||
ripple: bool = True
|
||||
) -> None:
|
||||
"""
|
||||
Move clips to a new position in the timeline.
|
||||
|
||||
Args:
|
||||
clip_ids: Clips to move (maintains relative order)
|
||||
target_position: 'start', 'end', timecode, or 'after:clip_id'/'before:clip_id'
|
||||
ripple: Whether to shift other clips
|
||||
"""
|
||||
spine = self._get_spine()
|
||||
|
||||
# Collect clips to move
|
||||
clips_to_move = []
|
||||
for clip_id in clip_ids:
|
||||
clip = self.clips.get(clip_id)
|
||||
if clip is not None and clip in list(spine):
|
||||
clips_to_move.append(clip)
|
||||
|
||||
if not clips_to_move:
|
||||
raise ValueError(f"No clips found matching: {clip_ids}")
|
||||
|
||||
# Calculate total duration of moving clips
|
||||
total_duration = TimeValue.zero()
|
||||
for clip in clips_to_move:
|
||||
dur = self._parse_time(clip.get('duration', '0s'))
|
||||
total_duration = total_duration + dur
|
||||
|
||||
# Remove clips from current positions
|
||||
for clip in clips_to_move:
|
||||
spine.remove(clip)
|
||||
|
||||
# Determine target offset and insert index
|
||||
spine_children = list(spine)
|
||||
target_offset, insert_index = self._resolve_insert_position(
|
||||
target_position, spine_children
|
||||
)
|
||||
|
||||
# Insert clips at new position
|
||||
current_offset = target_offset
|
||||
for clip in clips_to_move:
|
||||
clip.set('offset', current_offset.to_fcpxml())
|
||||
spine.insert(insert_index, clip)
|
||||
insert_index += 1
|
||||
dur = self._parse_time(clip.get('duration', '0s'))
|
||||
current_offset = current_offset + dur
|
||||
|
||||
# Recalculate all offsets if ripple
|
||||
if ripple:
|
||||
self._recalculate_offsets(spine)
|
||||
|
||||
def _recalculate_offsets(self, spine: ET.Element) -> None:
|
||||
"""Recalculate all clip offsets sequentially."""
|
||||
current_offset = TimeValue.zero()
|
||||
|
||||
for child in spine:
|
||||
if child.tag in SPINE_ELEMENT_TAGS:
|
||||
child.set('offset', current_offset.to_fcpxml())
|
||||
duration_str = child.get('duration', '0s')
|
||||
duration = self._parse_time(duration_str)
|
||||
current_offset = current_offset + duration
|
||||
|
||||
def _timeline_duration(self) -> 'TimeValue':
|
||||
"""Return the total timeline duration as a TimeValue.
|
||||
|
||||
Reads from the ``<sequence>`` element when available, falling back
|
||||
to summing all spine element durations. Extracted from
|
||||
``add_music_bed`` and ``batch_add_markers`` which both computed
|
||||
this independently.
|
||||
"""
|
||||
sequence = self.root.find('.//sequence')
|
||||
if sequence is not None:
|
||||
dur_str = sequence.get('duration')
|
||||
if dur_str:
|
||||
return self._parse_time(dur_str)
|
||||
spine = self._get_spine()
|
||||
total = TimeValue.zero()
|
||||
for child in spine:
|
||||
if child.tag in SPINE_ELEMENT_TAGS:
|
||||
total = total + self._parse_time(child.get('duration', '0s'))
|
||||
return total
|
||||
|
||||
def _update_sequence_duration(self) -> None:
|
||||
"""Recompute the ``<sequence>`` duration from the spine content.
|
||||
|
||||
Ripple edits (``cut_clip_ranges``, ``delete_clip``, ``split_clip``)
|
||||
change the total timeline length without rewriting the sequence
|
||||
element, so an exported file kept advertising the pre-edit duration —
|
||||
a 326.78s sequence still claimed 326.78s after 71s of silence was
|
||||
removed. This helper re-syncs the attribute to the actual spine sum.
|
||||
"""
|
||||
sequence = self.root.find('.//sequence')
|
||||
if sequence is None:
|
||||
return
|
||||
spine = self._get_spine()
|
||||
total = TimeValue.zero()
|
||||
for child in spine:
|
||||
if child.tag in SPINE_ELEMENT_TAGS:
|
||||
total = total + self._parse_time(child.get('duration', '0s'))
|
||||
sequence.set('duration', total.to_fcpxml())
|
||||
|
||||
# ========================================================================
|
||||
@@ -0,0 +1,43 @@
|
||||
"""Atribuir roles de vídeo/áudio.
|
||||
|
||||
Extraído de writer.py — ver fcpxml/writer/__init__.py para o conjunto.
|
||||
"""
|
||||
|
||||
import xml.etree.ElementTree as ET
|
||||
from typing import Optional
|
||||
|
||||
from .helpers import _sanitize_xml_value
|
||||
|
||||
|
||||
class RolesMixin:
|
||||
"""Atribuir roles de vídeo/áudio."""
|
||||
|
||||
# ROLE OPERATIONS (v0.5.0)
|
||||
# ========================================================================
|
||||
|
||||
def assign_role(
|
||||
self,
|
||||
clip_id: str,
|
||||
audio_role: Optional[str] = None,
|
||||
video_role: Optional[str] = None,
|
||||
) -> ET.Element:
|
||||
"""Set the audio/video role on a clip.
|
||||
|
||||
Args:
|
||||
clip_id: Name/ID of the clip
|
||||
audio_role: Audio role (e.g., "dialogue", "music", "effects")
|
||||
video_role: Video role (e.g., "video", "titles")
|
||||
|
||||
Returns:
|
||||
The modified clip element
|
||||
"""
|
||||
clip = self._require_clip(clip_id)
|
||||
|
||||
if audio_role is not None:
|
||||
clip.set('audioRole', _sanitize_xml_value(audio_role, 256))
|
||||
if video_role is not None:
|
||||
clip.set('videoRole', _sanitize_xml_value(video_role, 256))
|
||||
|
||||
return clip
|
||||
|
||||
# ========================================================================
|
||||
@@ -0,0 +1,57 @@
|
||||
"""Selecionar clipes por palavra-chave.
|
||||
|
||||
Extraído de writer.py — ver fcpxml/writer/__init__.py para o conjunto.
|
||||
"""
|
||||
|
||||
from typing import List
|
||||
|
||||
|
||||
class SelectionMixin:
|
||||
"""Selecionar clipes por palavra-chave."""
|
||||
|
||||
# SELECTION OPERATIONS
|
||||
# ========================================================================
|
||||
|
||||
def select_by_keyword(
|
||||
self,
|
||||
keywords: List[str],
|
||||
match_mode: str = 'any',
|
||||
favorites_only: bool = False,
|
||||
exclude_rejected: bool = True
|
||||
) -> List[str]:
|
||||
"""
|
||||
Find clips matching keywords.
|
||||
|
||||
Args:
|
||||
keywords: Keywords to match
|
||||
match_mode: 'any' (OR), 'all' (AND), 'none' (exclude)
|
||||
favorites_only: Only return favorited clips
|
||||
exclude_rejected: Exclude rejected clips
|
||||
|
||||
Returns:
|
||||
List of matching clip IDs
|
||||
"""
|
||||
matches = []
|
||||
|
||||
for clip_id, clip in self.clips.items():
|
||||
clip_keywords = set()
|
||||
for kw_elem in clip.findall('keyword'):
|
||||
clip_keywords.add(kw_elem.get('value', ''))
|
||||
|
||||
# Check keyword match
|
||||
keyword_set = set(keywords)
|
||||
if match_mode == 'any':
|
||||
match = bool(clip_keywords & keyword_set)
|
||||
elif match_mode == 'all':
|
||||
match = keyword_set <= clip_keywords
|
||||
elif match_mode == 'none':
|
||||
match = not bool(clip_keywords & keyword_set)
|
||||
else:
|
||||
match = True
|
||||
|
||||
if match:
|
||||
matches.append(clip_id)
|
||||
|
||||
return matches
|
||||
|
||||
# ========================================================================
|
||||
@@ -0,0 +1,185 @@
|
||||
"""Detectar e remover silêncio.
|
||||
|
||||
Extraído de writer.py — ver fcpxml/writer/__init__.py para o conjunto.
|
||||
"""
|
||||
|
||||
from typing import Any, Dict, List, Optional
|
||||
|
||||
from ..models import (
|
||||
MarkerType,
|
||||
TimeValue,
|
||||
)
|
||||
from .helpers import CLIP_TAGS, build_marker_element
|
||||
|
||||
|
||||
class SilenceMixin:
|
||||
"""Detectar e remover silêncio."""
|
||||
|
||||
# SILENCE DETECTION OPERATIONS (v0.5.0)
|
||||
# ========================================================================
|
||||
|
||||
def detect_silence_candidates(
|
||||
self,
|
||||
min_gap_seconds: float = 0.5,
|
||||
patterns: Optional[List[str]] = None,
|
||||
) -> List[Dict[str, Any]]:
|
||||
"""Detect potential silence regions using timeline heuristics.
|
||||
|
||||
Checks for:
|
||||
1. Gap elements in spine (high confidence)
|
||||
2. Ultra-short clips < 0.5s (medium confidence)
|
||||
3. Clips matching name patterns like "silence", "room tone" (high)
|
||||
4. Duration anomalies > 2 std dev from mean (low-medium)
|
||||
|
||||
Args:
|
||||
min_gap_seconds: Minimum gap duration to flag
|
||||
patterns: Name patterns to match (default: gap, silence, room tone)
|
||||
|
||||
Returns:
|
||||
List of silence candidate dicts
|
||||
"""
|
||||
if patterns is None:
|
||||
patterns = ['gap', 'silence', 'room tone', 'dead air', 'blank']
|
||||
|
||||
spine = self._get_spine()
|
||||
candidates = []
|
||||
durations = []
|
||||
clip_index = 0
|
||||
|
||||
# First pass: collect durations for anomaly detection
|
||||
for child in spine:
|
||||
if child.tag in CLIP_TAGS:
|
||||
dur = self._parse_time(child.get('duration', '0s'))
|
||||
durations.append(dur.to_seconds())
|
||||
|
||||
# Calculate stats for anomaly detection
|
||||
mean_dur = sum(durations) / len(durations) if durations else 0
|
||||
variance = (sum((d - mean_dur) ** 2 for d in durations) / len(durations)
|
||||
if len(durations) > 1 else 0)
|
||||
std_dev = variance ** 0.5
|
||||
|
||||
# Second pass: detect candidates
|
||||
for child in spine:
|
||||
tag = child.tag
|
||||
offset = child.get('offset', '0s')
|
||||
dur = self._parse_time(child.get('duration', '0s'))
|
||||
dur_secs = dur.to_seconds()
|
||||
tc = TimeValue.from_timecode(offset, self.fps).to_timecode(self.fps)
|
||||
|
||||
if tag == 'gap' and dur_secs >= min_gap_seconds:
|
||||
candidates.append({
|
||||
'start_timecode': tc,
|
||||
'duration_seconds': dur_secs,
|
||||
'reason': 'gap',
|
||||
'confidence': 0.9,
|
||||
'clip_name': None,
|
||||
'clip_index': None,
|
||||
})
|
||||
elif tag in CLIP_TAGS:
|
||||
name = child.get('name', '').lower()
|
||||
|
||||
# Name pattern match
|
||||
for pat in patterns:
|
||||
if pat.lower() in name:
|
||||
candidates.append({
|
||||
'start_timecode': tc,
|
||||
'duration_seconds': dur_secs,
|
||||
'reason': 'name_match',
|
||||
'confidence': 0.85,
|
||||
'clip_name': child.get('name', ''),
|
||||
'clip_index': clip_index,
|
||||
})
|
||||
break
|
||||
|
||||
# Ultra-short clip
|
||||
if dur_secs < 0.5:
|
||||
candidates.append({
|
||||
'start_timecode': tc,
|
||||
'duration_seconds': dur_secs,
|
||||
'reason': 'ultra_short',
|
||||
'confidence': 0.6,
|
||||
'clip_name': child.get('name', ''),
|
||||
'clip_index': clip_index,
|
||||
})
|
||||
|
||||
# Duration anomaly (> 2 std dev longer than mean)
|
||||
if std_dev > 0 and dur_secs > mean_dur + 2 * std_dev:
|
||||
candidates.append({
|
||||
'start_timecode': tc,
|
||||
'duration_seconds': dur_secs,
|
||||
'reason': 'duration_anomaly',
|
||||
'confidence': 0.4,
|
||||
'clip_name': child.get('name', ''),
|
||||
'clip_index': clip_index,
|
||||
})
|
||||
|
||||
clip_index += 1
|
||||
|
||||
return candidates
|
||||
|
||||
def remove_silence_candidates(
|
||||
self,
|
||||
mode: str = "mark",
|
||||
min_gap_seconds: float = 0.5,
|
||||
min_confidence: float = 0.7,
|
||||
patterns: Optional[List[str]] = None,
|
||||
) -> List[Dict[str, Any]]:
|
||||
"""Remove or mark detected silence candidates.
|
||||
|
||||
Args:
|
||||
mode: "delete" removes clips/gaps, "mark" adds red markers,
|
||||
"shorten" trims to minimum
|
||||
min_gap_seconds: Minimum gap to consider
|
||||
min_confidence: Only act on candidates above this threshold
|
||||
patterns: Name patterns to match
|
||||
|
||||
Returns:
|
||||
List of actions taken
|
||||
"""
|
||||
candidates = self.detect_silence_candidates(min_gap_seconds, patterns)
|
||||
candidates = [c for c in candidates if c['confidence'] >= min_confidence]
|
||||
|
||||
spine = self._get_spine()
|
||||
actions = []
|
||||
|
||||
if mode == "mark":
|
||||
for c in candidates:
|
||||
child = self._find_spine_element_at_timecode(
|
||||
spine, c['start_timecode'], require_clip=True
|
||||
)
|
||||
if child is not None:
|
||||
build_marker_element(
|
||||
parent=child,
|
||||
marker_type=MarkerType.STANDARD,
|
||||
start=child.get('start', '0s'),
|
||||
duration=f"1/{int(self.fps)}s",
|
||||
name=f"SILENCE: {c['reason']}",
|
||||
)
|
||||
actions.append({
|
||||
'action': 'marked',
|
||||
'clip_name': c.get('clip_name', 'gap'),
|
||||
'reason': c['reason'],
|
||||
})
|
||||
|
||||
elif mode == "delete":
|
||||
elements_to_remove = []
|
||||
for c in candidates:
|
||||
child = self._find_spine_element_at_timecode(
|
||||
spine, c['start_timecode']
|
||||
)
|
||||
if child is not None:
|
||||
elements_to_remove.append(child)
|
||||
actions.append({
|
||||
'action': 'deleted',
|
||||
'clip_name': c.get('clip_name', 'gap'),
|
||||
'reason': c['reason'],
|
||||
})
|
||||
|
||||
for elem in elements_to_remove:
|
||||
spine.remove(elem)
|
||||
|
||||
if elements_to_remove:
|
||||
self._recalculate_offsets(spine)
|
||||
|
||||
return actions
|
||||
|
||||
@@ -0,0 +1,297 @@
|
||||
"""Velocidade e zoom (punch-in) por janela.
|
||||
|
||||
Extraído de writer.py — ver fcpxml/writer/__init__.py para o conjunto.
|
||||
"""
|
||||
|
||||
import xml.etree.ElementTree as ET
|
||||
from fractions import Fraction
|
||||
from typing import Optional
|
||||
|
||||
from .helpers import HOLD_AT_CUT_THRESHOLD, START_AT_CUT_THRESHOLD, _dtd_insert, _fmt_scale
|
||||
|
||||
|
||||
class SpeedMixin:
|
||||
"""Velocidade e zoom (punch-in) por janela."""
|
||||
|
||||
# SPEED OPERATIONS
|
||||
# ========================================================================
|
||||
|
||||
def change_speed(
|
||||
self,
|
||||
clip_id: str,
|
||||
speed: float,
|
||||
preserve_pitch: bool = True
|
||||
) -> ET.Element:
|
||||
"""
|
||||
Change clip playback speed.
|
||||
|
||||
Args:
|
||||
clip_id: Target clip
|
||||
speed: Speed multiplier (0.5 = half speed, 2.0 = double)
|
||||
preserve_pitch: Maintain audio pitch
|
||||
|
||||
Returns:
|
||||
Modified clip element
|
||||
"""
|
||||
if speed <= 0:
|
||||
raise ValueError(f"Speed must be positive, got {speed}")
|
||||
|
||||
clip = self._require_clip(clip_id)
|
||||
|
||||
current_duration = self._parse_time(clip.get('duration', '0s'))
|
||||
|
||||
# Use rational arithmetic to avoid floating-point time values.
|
||||
# FCPXML requires rational fractions with a consistent timebase,
|
||||
# not decimal floats like "2.6666666666666665s".
|
||||
denom = current_duration.denominator if current_duration.denominator > 0 else int(self.fps)
|
||||
source_num = current_duration.numerator
|
||||
speed_frac = Fraction(speed).limit_denominator(1000)
|
||||
raw_num = source_num * speed_frac.denominator
|
||||
raw_denom = denom * speed_frac.numerator
|
||||
|
||||
# Snap to frame boundary in a standard timebase (2400 ticks/sec).
|
||||
# Each frame at Nfps = 2400/N ticks (e.g. 24fps → 100 ticks/frame).
|
||||
fps_int = int(self.fps) if self.fps else 24
|
||||
ticks_per_frame = 2400 // fps_int
|
||||
dur_ticks = round(raw_num / raw_denom * 2400)
|
||||
dur_ticks = round(dur_ticks / ticks_per_frame) * ticks_per_frame
|
||||
new_num = dur_ticks
|
||||
new_denom = 2400
|
||||
|
||||
# Remove any existing timeMap/conform-rate from a prior speed change
|
||||
# to prevent duplicate children that produce invalid FCPXML.
|
||||
for stale_tag in ('timeMap', 'conform-rate'):
|
||||
for stale in clip.findall(stale_tag):
|
||||
clip.remove(stale)
|
||||
|
||||
# Create timeMap for speed change (DTD-ordered insertion)
|
||||
timemap = ET.Element('timeMap')
|
||||
_dtd_insert(clip, timemap)
|
||||
|
||||
# Start keyframe
|
||||
tp1 = ET.SubElement(timemap, 'timept')
|
||||
tp1.set('time', '0s')
|
||||
tp1.set('value', '0s')
|
||||
tp1.set('interp', 'linear')
|
||||
|
||||
# End keyframe — use rational time, not floats
|
||||
tp2 = ET.SubElement(timemap, 'timept')
|
||||
tp2.set('time', f"{new_num}/{new_denom}s")
|
||||
tp2.set('value', f"{source_num}/{denom}s")
|
||||
tp2.set('interp', 'linear')
|
||||
|
||||
# Update clip duration (rational, not simplified to arbitrary denominator)
|
||||
clip.set('duration', f"{new_num}/{new_denom}s")
|
||||
|
||||
# Add conform-rate (DTD-ordered insertion)
|
||||
conform = ET.Element('conform-rate')
|
||||
conform.set('scaleEnabled', '1')
|
||||
conform.set('srcFrameRate', str(int(self.fps)))
|
||||
_dtd_insert(clip, conform)
|
||||
|
||||
return clip
|
||||
|
||||
def add_zoom(
|
||||
self,
|
||||
clip_id: 'str | ET.Element',
|
||||
start: float,
|
||||
end: float,
|
||||
scale: float = 1.3,
|
||||
ease: float = 0.25,
|
||||
position: str = "0 0",
|
||||
ease_out: Optional[float] = None,
|
||||
hold_at_end: Optional[bool] = None,
|
||||
start_at_peak: Optional[bool] = None,
|
||||
) -> ET.Element:
|
||||
"""Add a punch-in zoom to a clip, snapping back to its framing at the end.
|
||||
|
||||
Animates ``<adjust-transform>``'s ``scale`` param (``<param>`` +
|
||||
``<keyframeAnimation>`` of ``<keyframe>``) from the clip's current
|
||||
scale up to *scale* times it, holds, then returns — all within
|
||||
``[start, end]`` — clip-relative seconds (same convention as
|
||||
``cut_clip_ranges``).
|
||||
|
||||
The two ends are deliberately asymmetric. *ease* ramps the zoom
|
||||
**in** over half a second by default, fast enough to land with the
|
||||
emphasised word. The way **out** is instant — a single frame — so
|
||||
the moment the impact phrase ends the shot is simply back to its
|
||||
normal framing and the video resumes its flow, with no drift
|
||||
drawing attention to itself. Pass *ease_out* to ramp the return
|
||||
gradually instead.
|
||||
|
||||
*hold_at_end* keeps the peak instead of returning, and
|
||||
*start_at_peak* opens already zoomed with no ramp. Left as ``None``
|
||||
both decide on their own from how close the window sits to the
|
||||
clip's edges: a cut is itself the transition, so ramping away from
|
||||
one — or back toward one — is motion the viewer reads as a wobble
|
||||
rather than as emphasis.
|
||||
"""
|
||||
if end <= start:
|
||||
raise ValueError(f"end ({end}) must be greater than start ({start})")
|
||||
if ease <= 0:
|
||||
raise ValueError(f"ease must be positive, got {ease}")
|
||||
if scale <= 0:
|
||||
raise ValueError(f"scale must be positive, got {scale}")
|
||||
|
||||
frame = float(self.frame_duration_fraction())
|
||||
ramp_out = frame if ease_out is None else ease_out
|
||||
if ramp_out <= 0:
|
||||
raise ValueError(f"ease_out must be positive, got {ease_out}")
|
||||
clip = self._require_clip(clip_id)
|
||||
clip_duration = self._parse_time(clip.get('duration', '0s')).to_seconds()
|
||||
if start < 0 or end > clip_duration:
|
||||
raise ValueError(
|
||||
f"zoom window [{start}, {end}]s must fall within the clip's "
|
||||
f"duration (0 to {clip_duration:.3f}s)"
|
||||
)
|
||||
|
||||
# Replace a prior zoom, but never the clip's framing. A clip can
|
||||
# already carry an <adjust-transform> holding the editor's own
|
||||
# reframe — rotation for footage shot sideways, position, a scale
|
||||
# that makes the shot work at all. Dropping it outright (the old
|
||||
# behaviour) silently destroyed that framing; on real footage the
|
||||
# zoomed section came back rotated. So: keep the static attributes,
|
||||
# and animate *relative to* the existing scale.
|
||||
base_x, base_y = 1.0, 1.0
|
||||
carried: dict = {}
|
||||
old_keyframes: list = []
|
||||
for stale in clip.findall('adjust-transform'):
|
||||
carried = {k: v for k, v in stale.attrib.items() if k != 'scale'}
|
||||
parts = (stale.get('scale') or '').split()
|
||||
if len(parts) == 2:
|
||||
try:
|
||||
base_x, base_y = float(parts[0]), float(parts[1])
|
||||
except ValueError:
|
||||
base_x, base_y = 1.0, 1.0
|
||||
else:
|
||||
# No static attribute — a PRIOR zoom on this same clip left
|
||||
# an animated <param name="scale"> instead, and the true
|
||||
# resting framing lives in its keyframes, not in 1.0.
|
||||
# Reading it as 1.0 here doesn't just miss the framing: it
|
||||
# replaces the earlier zoom's whole animation with a wrong
|
||||
# one, since this loop unconditionally removes `stale`
|
||||
# right after. The rest value is recoverable without
|
||||
# knowing which keyframe it is: MIN_ZOOM_SCALE == 1.0 means
|
||||
# every keyframed value is >= the rest scale, so the
|
||||
# smallest one keyframed is the rest value, peak or not.
|
||||
for old_param in stale.findall("param[@name='scale']"):
|
||||
xs, ys = [], []
|
||||
for kf in old_param.findall('.//keyframe'):
|
||||
kv = (kf.get('value') or '').split()
|
||||
if len(kv) == 2:
|
||||
try:
|
||||
xs.append(float(kv[0]))
|
||||
ys.append(float(kv[1]))
|
||||
except ValueError:
|
||||
pass
|
||||
# Kept for merging: a second zoom on the same clip
|
||||
# (two emphatic beats a cut didn't separate) should
|
||||
# stack alongside the first, not erase it — the
|
||||
# earlier peak is still a real editorial decision.
|
||||
old_keyframes.append((kf.get('time', '0s'), kf.get('value', '')))
|
||||
if xs and ys:
|
||||
base_x, base_y = min(xs), min(ys)
|
||||
clip.remove(stale)
|
||||
|
||||
transform = ET.Element('adjust-transform')
|
||||
for key, value in carried.items():
|
||||
transform.set(key, value)
|
||||
scale_param = ET.SubElement(transform, 'param')
|
||||
scale_param.set('name', 'scale')
|
||||
anim = ET.SubElement(scale_param, 'keyframeAnimation')
|
||||
|
||||
# Keyframe times live in the clip's SOURCE timebase — the same origin
|
||||
# as its own ``start`` — not in clip-relative seconds. A clip whose
|
||||
# media starts at, say, 3109.9s of timecode looks for the animation
|
||||
# there; keyframes written at 0-5s land outside the clip entirely and
|
||||
# Final Cut imports the zoom as nothing at all, silently. Matches what
|
||||
# add_text_title already does, and only shows up on footage whose
|
||||
# start isn't 0s — every synthetic fixture starts at 0s and hides it.
|
||||
media_origin = self._parse_time(clip.get('start', '0s'))
|
||||
|
||||
rest_value = f"{_fmt_scale(base_x)} {_fmt_scale(base_y)}"
|
||||
scale_value = f"{_fmt_scale(base_x * scale)} {_fmt_scale(base_y * scale)}"
|
||||
|
||||
# A return that lands right before a cut is wasted motion: the next
|
||||
# clip begins on its own framing anyway, so all the viewer sees is a
|
||||
# twitch on the way out. When the zoom runs to the end of the clip,
|
||||
# hold the peak and let the cut do the resetting.
|
||||
holds_to_cut = (
|
||||
hold_at_end
|
||||
if hold_at_end is not None
|
||||
else (clip_duration - end) <= HOLD_AT_CUT_THRESHOLD
|
||||
)
|
||||
opens_at_peak = (
|
||||
start_at_peak
|
||||
if start_at_peak is not None
|
||||
else start <= START_AT_CUT_THRESHOLD
|
||||
)
|
||||
|
||||
# Only the ramps actually written have to fit in the window: a zoom
|
||||
# that opens at the peak spends no time ramping in, and one held to
|
||||
# the cut spends none ramping out.
|
||||
needed = (0.0 if opens_at_peak else ease) + (0.0 if holds_to_cut else ramp_out)
|
||||
if needed > (end - start):
|
||||
raise ValueError(
|
||||
f"the ramps ({needed}s) don't fit in the zoom window "
|
||||
f"({end - start}s) — shorten them or widen start/end"
|
||||
)
|
||||
|
||||
if opens_at_peak:
|
||||
# The cut already delivered the change of framing; ramping up
|
||||
# from it just looks like the shot settling.
|
||||
keyframes = [(start, scale_value)]
|
||||
else:
|
||||
keyframes = [(start, rest_value), (start + ease, scale_value)]
|
||||
if holds_to_cut:
|
||||
keyframes.append((end, scale_value))
|
||||
else:
|
||||
# Hold the peak right up to the end, then drop back on the very
|
||||
# next frame — the snap-back the edit wants, not a slow drift.
|
||||
keyframes.append((end - ramp_out, scale_value))
|
||||
keyframes.append((end, rest_value))
|
||||
|
||||
new_entries = [
|
||||
((media_origin + self.snap_seconds_to_frame(seconds)), value)
|
||||
for seconds, value in keyframes
|
||||
]
|
||||
new_start_time = new_entries[0][0]
|
||||
new_end_time = new_entries[-1][0]
|
||||
|
||||
# Two calls on the same clip mean two different things depending on
|
||||
# whether their windows overlap. Overlapping = redoing the *same*
|
||||
# zoom with new numbers — the old keyframes are stale and all of
|
||||
# them go. Disjoint = a second, separate beat that a cut didn't
|
||||
# separate onto its own clip — that one stacks alongside the first
|
||||
# instead of erasing it, since both are real editorial decisions.
|
||||
old_times = [self._parse_time(t) for t, _ in old_keyframes]
|
||||
old_span_overlaps_new = bool(old_times) and not (
|
||||
max(old_times) < new_start_time or min(old_times) > new_end_time
|
||||
)
|
||||
if old_span_overlaps_new:
|
||||
surviving_old: list = []
|
||||
else:
|
||||
surviving_old = [(self._parse_time(t), v) for t, v in old_keyframes]
|
||||
all_entries = sorted(surviving_old + new_entries, key=lambda e: e[0])
|
||||
|
||||
for time_value, value in all_entries:
|
||||
kf = ET.SubElement(anim, 'keyframe')
|
||||
kf.set('time', time_value.to_fcpxml())
|
||||
kf.set('value', value)
|
||||
# Only 'time' and 'value' — no 'interp', no 'curve'. The DTD allows
|
||||
# both, but Final Cut rejected 'interp' on this vector param
|
||||
# ("does not support the interpolation attribute") and discarded
|
||||
# the whole <param>. A hand-made zoom exported from FCP itself
|
||||
# writes bare keyframes and relies on the DTD default
|
||||
# (curve="smooth"), so we match that export exactly rather than
|
||||
# guess which attributes survive its importer.
|
||||
|
||||
if position != "0 0":
|
||||
pos_param = ET.SubElement(transform, 'param')
|
||||
pos_param.set('name', 'position')
|
||||
pos_param.set('value', position)
|
||||
|
||||
_dtd_insert(clip, transform)
|
||||
return clip
|
||||
|
||||
# ========================================================================
|
||||
@@ -0,0 +1,867 @@
|
||||
"""Títulos de texto e legendas dinâmicas.
|
||||
|
||||
Extraído de writer.py — ver fcpxml/writer/__init__.py para o conjunto.
|
||||
"""
|
||||
|
||||
import random
|
||||
import re
|
||||
import unicodedata
|
||||
import uuid
|
||||
import xml.etree.ElementTree as ET
|
||||
from typing import Any, Dict, List, Optional
|
||||
|
||||
from ..collision import blocking, validate_titles
|
||||
from ..models import (
|
||||
DynamicSubtitleConfig,
|
||||
TimeValue,
|
||||
)
|
||||
from ..text_layout import (
|
||||
TEXT_TEMPLATE_FONT_SCALE,
|
||||
LayoutBox,
|
||||
compose_sentence,
|
||||
layout_sentence,
|
||||
)
|
||||
from ..transcribe import group_words_by_segment, split_into_subphrases
|
||||
from .helpers import _dtd_insert, _sanitize_xml_value
|
||||
|
||||
|
||||
class TitlesMixin:
|
||||
"""Títulos de texto e legendas dinâmicas."""
|
||||
|
||||
_SUBTITLE_METADATA_KEY = 'com.gart.subtitle.kind'
|
||||
|
||||
def mark_generated_subtitle(self, element: ET.Element, kind: str) -> None:
|
||||
metadata = element.find('metadata')
|
||||
if metadata is None:
|
||||
metadata = ET.Element('metadata')
|
||||
_dtd_insert(element, metadata)
|
||||
ET.SubElement(metadata, 'md', key=self._SUBTITLE_METADATA_KEY, value=kind)
|
||||
|
||||
def _generated_subtitle_kind(self, element: ET.Element) -> Optional[str]:
|
||||
marker = element.find(f"metadata/md[@key='{self._SUBTITLE_METADATA_KEY}']")
|
||||
if marker is not None:
|
||||
return marker.get('value')
|
||||
# Recognize the exact signature of older G-ART exports. A role alone
|
||||
# is not ownership: users also assign these roles to manual titles.
|
||||
if element.tag == 'title':
|
||||
if element.get('start') != self._TEXT_TITLE_START:
|
||||
return None
|
||||
effect = self.root.find(f".//resources/effect[@id='{element.get('ref')}']")
|
||||
if effect is None or effect.get('uid') != self._TEXT_TITLE_UID:
|
||||
return None
|
||||
if re.fullmatch(r'caption_[0-9a-f]{8}', element.get('name', '')):
|
||||
return 'dynamic'
|
||||
text = ''.join(element.findtext('text/text-style', ''))
|
||||
if (element.get('role') == 'titles.convencionais'
|
||||
and element.get('lane') == '20'
|
||||
and element.get('name') == f'{text} - Text'):
|
||||
return 'plain'
|
||||
elif element.tag == 'ref-clip':
|
||||
media = self.root.find(f".//resources/media[@id='{element.get('ref')}']")
|
||||
if media is not None:
|
||||
titles = media.findall('.//title')
|
||||
if titles and all(self._generated_subtitle_kind(t) == 'dynamic' for t in titles):
|
||||
return 'dynamic'
|
||||
return None
|
||||
|
||||
def remove_generated_subtitles(self, parent: ET.Element, kinds: tuple) -> None:
|
||||
"""Replace only our own captions, preserving unrelated graphics."""
|
||||
resources = self.root.find('.//resources')
|
||||
for child in list(parent):
|
||||
if self._generated_subtitle_kind(child) not in kinds:
|
||||
continue
|
||||
parent.remove(child)
|
||||
if child.tag == 'ref-clip' and resources is not None:
|
||||
ref = child.get('ref')
|
||||
if not self.root.findall(f".//ref-clip[@ref='{ref}']"):
|
||||
media = resources.find(f"media[@id='{ref}']")
|
||||
if media is not None:
|
||||
resources.remove(media)
|
||||
|
||||
def suppress_plain_under_dynamic(self, parent: ET.Element) -> None:
|
||||
"""Keep generated plain titles only on frames without dynamic text."""
|
||||
import copy
|
||||
|
||||
windows = []
|
||||
for child in parent:
|
||||
if self._generated_subtitle_kind(child) == 'dynamic':
|
||||
start = self._parse_time(child.get('offset', '0s'))
|
||||
windows.append((start, start + self._parse_time(child.get('duration', '0s'))))
|
||||
for title in list(parent):
|
||||
if self._generated_subtitle_kind(title) != 'plain':
|
||||
continue
|
||||
start = self._parse_time(title.get('offset', '0s'))
|
||||
end = start + self._parse_time(title.get('duration', '0s'))
|
||||
remaining = [(start, end)]
|
||||
for lo, hi in windows:
|
||||
parts = []
|
||||
for a, b in remaining:
|
||||
if a < hi and lo < b:
|
||||
if a < lo:
|
||||
parts.append((a, lo))
|
||||
if hi < b:
|
||||
parts.append((hi, b))
|
||||
else:
|
||||
parts.append((a, b))
|
||||
remaining = parts
|
||||
if remaining == [(start, end)]:
|
||||
continue
|
||||
parent.remove(title)
|
||||
for a, b in remaining:
|
||||
part = copy.deepcopy(title)
|
||||
self._reassign_text_style_ids(part)
|
||||
part.set('offset', a.to_fcpxml())
|
||||
part.set('duration', (b - a).to_fcpxml())
|
||||
_dtd_insert(parent, part)
|
||||
|
||||
# DYNAMIC (KARAOKE-STYLE) SUBTITLES
|
||||
# ========================================================================
|
||||
|
||||
# The "Text" (Basic Text) template — the ONLY simple title template that
|
||||
# Final Cut actually renders. Copied verbatim from the user's own FCP
|
||||
# exports ("teste.fcpxmld" and "posição.fcpxmld", FCP 1.14 in English):
|
||||
# a single "<text>" run, one "<text-style-def>", and a fixed param block
|
||||
# with the margins/alignment/speed the template ships with. Every prior
|
||||
# title template we generated ("Essencial - Título", "Título Básico")
|
||||
# imported cleanly but never appeared — their Motion uids did not resolve
|
||||
# to a real, drawable template in FCP, which discards the clip silently.
|
||||
# "Text" is what FCP itself writes when the user adds a title by hand, so
|
||||
# it is the ground truth. See Engine/docs/05_EXPERIENCIAS.md, 2026-08-17.
|
||||
_TEXT_TITLE_UID = (
|
||||
'.../Titles.localized/Basic Text.localized/'
|
||||
'Text.localized/Text.moti'
|
||||
)
|
||||
_TEXT_TITLE_START = '86486400/24000s'
|
||||
# The Inspector's Position field, and the one this code overrides per
|
||||
# title so two titles never stack on top of each other. Verified in
|
||||
# "posição.fcpxmld": each hand-dragged title carries a distinct "x y"
|
||||
# value here while every other param stays identical.
|
||||
_TEXT_POSITION_KEY = '9999/10003/13260/3296672360/1/100/101'
|
||||
# Layout params the "Text" template ships with. These keys are the
|
||||
# template's own defaults and never vary between instances.
|
||||
#
|
||||
# "Build Out" is the one deliberate override: with "Apply Speed" set to
|
||||
# "2 (Per Object)" below, the template's whole built-in animation (build
|
||||
# in + build out) is always compressed to exactly fill the title's own
|
||||
# on-screen duration — so on a short word-length clip, build out was
|
||||
# eating time that build in needed to finish revealing the text before
|
||||
# the cut. Disabling build out hands that entire compressed window to
|
||||
# build in alone, which is what "sempre acelerado" turned out to mean:
|
||||
# no separate speed knob needed. Value captured from a real FCP export
|
||||
# with "Build Out" unchecked in the Inspector (see chat, 2026-08-18).
|
||||
_TEXT_TITLE_PARAMS = (
|
||||
('Build Out', '9999/10000/2/102', '0'),
|
||||
('Layout Method', '9999/10003/13260/3296672360/2/314', '1 (Paragraph)'),
|
||||
('Left Margin', '9999/10003/13260/3296672360/2/323', '-1210'),
|
||||
('Right Margin', '9999/10003/13260/3296672360/2/324', '1210'),
|
||||
('Top Margin', '9999/10003/13260/3296672360/2/325', '2160'),
|
||||
('Bottom Margin', '9999/10003/13260/3296672360/2/326', '-2160'),
|
||||
('Alignment', '9999/10003/13260/3296672360/2/354/3296667315/401', '1 (Center)'),
|
||||
('Line Spacing', '9999/10003/13260/3296672360/2/354/3296667315/404', '-19'),
|
||||
('Auto-Shrink', '9999/10003/13260/3296672360/2/370', '3 (To All Margins)'),
|
||||
('Alignment', '9999/10003/13260/3296672360/2/373', '0 (Left) 1 (Middle)'),
|
||||
('Opacity', '9999/10003/13260/3296672360/4/3296673134/1000/1044', '0'),
|
||||
('Speed', '9999/10003/13260/3296672360/4/3296673134/201/208', '6 (Custom)'),
|
||||
('Apply Speed', '9999/10003/13260/3296672360/4/3296673134/201/211', '2 (Per Object)'),
|
||||
)
|
||||
# "Custom Speed" sits between "Speed" and "Apply Speed" and carries a
|
||||
# <keyframeAnimation> child rather than a plain value attribute. Its two
|
||||
# keyframes are the template's own absolute nominal times, constant across
|
||||
# every instance, so they are safe to replay verbatim.
|
||||
_TEXT_CUSTOM_SPEED_KEY = '9999/10003/13260/3296672360/4/3296673134/201/209'
|
||||
_TEXT_CUSTOM_SPEED_KEYFRAMES = (
|
||||
('-469658744/1000000000s', '0'),
|
||||
('12328542033/1000000000s', '1'),
|
||||
)
|
||||
_TEXT_SIZE_KEY = '9999/10003/13260/3296672360/5/3296672362/3'
|
||||
|
||||
def _ensure_text_title_effect(self, resources: ET.Element) -> str:
|
||||
"""Return the resource id of the "Text" (Basic Text) effect, creating it if absent."""
|
||||
return self._ensure_effect(resources, self._TEXT_TITLE_UID, 'Text', 'r_text')
|
||||
|
||||
def _ensure_effect(
|
||||
self,
|
||||
resources: ET.Element,
|
||||
uid: str,
|
||||
name: str,
|
||||
id_prefix: str,
|
||||
) -> str:
|
||||
"""Return the id of the effect resource with *uid*, creating it if absent."""
|
||||
for eff in resources.findall('effect'):
|
||||
if eff.get('uid') == uid:
|
||||
return eff.get('id')
|
||||
effect_id = self._unique_resource_id(resources, id_prefix)
|
||||
eff_el = ET.SubElement(resources, 'effect')
|
||||
eff_el.set('id', effect_id)
|
||||
eff_el.set('name', name)
|
||||
eff_el.set('uid', uid)
|
||||
return effect_id
|
||||
|
||||
# <text-style-def id> / <text-style ref> are DTD type ID/IDREF, so the
|
||||
# value must be a valid XML Name: letters, digits, "_", "-", "." only,
|
||||
# never starting with a digit. Title names are built from the caption
|
||||
# text ("Olá mundo - Text"), which carries spaces, accents and often a
|
||||
# leading digit — xmllint rejected the whole document with "Syntax of
|
||||
# value for attribute id of text-style-def is not valid".
|
||||
_TEXT_STYLE_ID_UNSAFE = re.compile(r'[^A-Za-z0-9_.-]+')
|
||||
|
||||
def _unique_text_style_id(self, base: str) -> str:
|
||||
"""Return a document-unique, DTD-valid XML ID for a ``<text-style-def>``."""
|
||||
folded = unicodedata.normalize('NFKD', base).encode('ascii', 'ignore').decode('ascii')
|
||||
slug = self._TEXT_STYLE_ID_UNSAFE.sub('_', folded).strip('_.-')[:48]
|
||||
stem = f"ts_{slug}" if slug else "ts"
|
||||
|
||||
if self._text_style_ids is None:
|
||||
self._text_style_ids = {
|
||||
sd.get('id') for sd in self.root.findall('.//text-style-def')
|
||||
}
|
||||
candidate = f"{stem}_0"
|
||||
counter = 0
|
||||
while candidate in self._text_style_ids:
|
||||
counter += 1
|
||||
candidate = f"{stem}_{counter}"
|
||||
self._text_style_ids.add(candidate)
|
||||
return candidate
|
||||
|
||||
def _reassign_text_style_ids(self, clip: ET.Element) -> None:
|
||||
"""Give every ``<text-style-def>`` inside a just-deepcopy'd *clip* a
|
||||
fresh document-unique id, repointing any ``<text-style ref="...">``
|
||||
in the same subtree that pointed at the old one.
|
||||
|
||||
``split_clip``/``cut_clip_ranges`` deepcopy the clip once per
|
||||
resulting segment, so a clip carrying a ``<title>`` (from a "text"
|
||||
voice action) keeps the exact same ``text-style-def id`` in every
|
||||
copy. A single cut is harmless — but the batch chain re-cuts the
|
||||
same clip at each step (silence removal, filler removal, dynamic
|
||||
subtitles), and every pass multiplies the duplicate, so the DTD
|
||||
validator eventually rejects the file with "ID ... already
|
||||
defined". Regenerating here, at the only place copies are made,
|
||||
fixes it for every caller instead of each one having to remember to.
|
||||
"""
|
||||
for style_def in clip.findall('.//text-style-def'):
|
||||
old_id = style_def.get('id')
|
||||
if not old_id:
|
||||
continue
|
||||
slug = old_id[3:] if old_id.startswith('ts_') else old_id
|
||||
slug = re.sub(r'_\d+$', '', slug) # drop a prior _<N> counter
|
||||
new_id = self._unique_text_style_id(slug)
|
||||
if new_id == old_id:
|
||||
continue
|
||||
style_def.set('id', new_id)
|
||||
for ref_el in clip.findall(f".//text-style[@ref='{old_id}']"):
|
||||
ref_el.set('ref', new_id)
|
||||
|
||||
def _unique_tracking_shape_id(self, base: str) -> str:
|
||||
"""Return a document-unique ``id`` for a ``<tracking-shape>``."""
|
||||
stem = base or "tr"
|
||||
if self._tracking_shape_ids is None:
|
||||
self._tracking_shape_ids = {
|
||||
ts.get('id') for ts in self.root.findall('.//tracking-shape')
|
||||
}
|
||||
candidate = f"{stem}_0"
|
||||
counter = 0
|
||||
while candidate in self._tracking_shape_ids:
|
||||
counter += 1
|
||||
candidate = f"{stem}_{counter}"
|
||||
self._tracking_shape_ids.add(candidate)
|
||||
return candidate
|
||||
|
||||
def _reassign_tracking_shape_ids(self, clip: ET.Element) -> None:
|
||||
"""Give every ``<tracking-shape>`` inside a just-deepcopy'd *clip* a
|
||||
fresh document-unique id.
|
||||
|
||||
Same mechanism as ``_reassign_text_style_ids``: ``split_clip``/
|
||||
``cut_clip_ranges`` deepcopy the clip once per resulting segment, so
|
||||
Cinematic object-tracking data (``<object-tracker><tracking-shape
|
||||
id="tr1">``, preserved from the source asset's sidecar) keeps the
|
||||
exact same id in every copy. A single cut is harmless — but the
|
||||
batch chain re-cuts the same clip at each step, multiplying the
|
||||
duplicate until the DTD validator rejects the file with "ID tr1
|
||||
already defined".
|
||||
"""
|
||||
for shape in clip.findall('.//tracking-shape'):
|
||||
old_id = shape.get('id')
|
||||
if not old_id:
|
||||
continue
|
||||
base = re.sub(r'_\d+$', '', old_id)
|
||||
new_id = self._unique_tracking_shape_id(base)
|
||||
if new_id == old_id:
|
||||
continue
|
||||
shape.set('id', new_id)
|
||||
|
||||
def _make_text_title_clip(
|
||||
self,
|
||||
effect_id: str,
|
||||
text: str,
|
||||
offset: 'TimeValue',
|
||||
duration: 'TimeValue',
|
||||
*,
|
||||
lane: int,
|
||||
name: str,
|
||||
position: Optional[str] = None,
|
||||
font: str = 'Helvetica Neue',
|
||||
font_size: int = 196,
|
||||
font_color: str = '1 1 1 1',
|
||||
bold: bool = True,
|
||||
face: Optional[str] = None,
|
||||
kerning: Optional[float] = None,
|
||||
font_scale: float = TEXT_TEMPLATE_FONT_SCALE,
|
||||
animated: bool = True,
|
||||
size_param: Optional[float] = None,
|
||||
role: Optional[str] = None,
|
||||
) -> ET.Element:
|
||||
"""Build a standalone ``<title>`` clip from the "Text" (Basic Text) template.
|
||||
|
||||
Reproduces FCP's own output for a hand-added title exactly — the only
|
||||
template we have verified renders in Final Cut ("teste.fcpxmld" and
|
||||
"posição.fcpxmld"). *position* ("x y" canvas points) is the Inspector
|
||||
Position value; omit it to keep the template's centred default. Unlike
|
||||
the animated templates, this carries no animation switch, so the text
|
||||
stays put and visible for its whole duration.
|
||||
"""
|
||||
elem = ET.Element('title')
|
||||
elem.set('ref', effect_id)
|
||||
elem.set('lane', str(lane))
|
||||
elem.set('offset', offset.to_fcpxml())
|
||||
elem.set('name', _sanitize_xml_value(name, 256))
|
||||
elem.set('start', self._TEXT_TITLE_START)
|
||||
elem.set('duration', duration.to_fcpxml())
|
||||
if role:
|
||||
elem.set('role', _sanitize_xml_value(role, 256))
|
||||
|
||||
if position:
|
||||
param = ET.SubElement(elem, 'param')
|
||||
param.set('name', 'Position')
|
||||
param.set('key', self._TEXT_POSITION_KEY)
|
||||
param.set('value', position)
|
||||
|
||||
def _add_param(name: str, key: str, value: str) -> None:
|
||||
param = ET.SubElement(elem, 'param')
|
||||
param.set('name', name)
|
||||
param.set('key', key)
|
||||
param.set('value', value)
|
||||
|
||||
animation_params = {'Opacity', 'Speed', 'Apply Speed'}
|
||||
for param_name, param_key, param_value in self._TEXT_TITLE_PARAMS:
|
||||
if not animated and param_name in animation_params:
|
||||
continue
|
||||
_add_param(param_name, param_key, param_value)
|
||||
if animated and param_name == 'Speed':
|
||||
# "Custom Speed" lands between "Speed" and "Apply Speed" and
|
||||
# carries a <keyframeAnimation> child instead of a value.
|
||||
cs = ET.SubElement(elem, 'param')
|
||||
cs.set('name', 'Custom Speed')
|
||||
cs.set('key', self._TEXT_CUSTOM_SPEED_KEY)
|
||||
anim = ET.SubElement(cs, 'keyframeAnimation')
|
||||
for kf_time, kf_value in self._TEXT_CUSTOM_SPEED_KEYFRAMES:
|
||||
kf = ET.SubElement(anim, 'keyframe')
|
||||
kf.set('time', kf_time)
|
||||
kf.set('value', kf_value)
|
||||
|
||||
if size_param is not None:
|
||||
_add_param('Size', self._TEXT_SIZE_KEY, f"{float(size_param):g}")
|
||||
|
||||
text_el = ET.SubElement(elem, 'text')
|
||||
ts_id = self._unique_text_style_id(name)
|
||||
run = ET.SubElement(text_el, 'text-style')
|
||||
run.set('ref', ts_id)
|
||||
run.text = _sanitize_xml_value(text, 256)
|
||||
|
||||
style_def = ET.SubElement(elem, 'text-style-def')
|
||||
style_def.set('id', ts_id)
|
||||
text_style = ET.SubElement(style_def, 'text-style')
|
||||
text_style.set('font', font)
|
||||
# Text.moti sizes type in frame pixels but positions in canvas points.
|
||||
# See TEXT_TEMPLATE_FONT_SCALE: layout measures in points, so only the
|
||||
# emitted size (and its kerning, to keep the same letter spacing) is
|
||||
# converted here.
|
||||
scale = float(font_scale) or 1.0
|
||||
text_style.set('fontSize', f"{float(font_size) * scale:g}")
|
||||
text_style.set('fontColor', font_color)
|
||||
# FCP represents bold weight as the bold attribute — never as a
|
||||
# fontFace. Writing ``bold="0" fontFace="Bold"`` (the previous
|
||||
# behaviour) is contradictory and FCP refuses to render the text.
|
||||
# Italic, by contrast, IS a face: FCP writes both ``fontFace`` and
|
||||
# ``italic="1"``. See Engine/docs/05_EXPERIENCIAS.md, entry 2026-08-19.
|
||||
face_lower = (face or '').strip().lower()
|
||||
if face_lower == 'bold':
|
||||
text_style.set('bold', '1')
|
||||
elif 'italic' in face_lower:
|
||||
text_style.set('fontFace', face)
|
||||
text_style.set('italic', '1')
|
||||
else:
|
||||
if bold:
|
||||
text_style.set('bold', '1')
|
||||
if face:
|
||||
text_style.set('fontFace', face)
|
||||
if kerning:
|
||||
text_style.set('kerning', f"{float(kerning) * scale:g}")
|
||||
text_style.set('alignment', 'center')
|
||||
text_style.set('lineSpacing', '-19')
|
||||
|
||||
return elem
|
||||
|
||||
def add_text_title(
|
||||
self,
|
||||
parent_clip: 'str | ET.Element',
|
||||
text: str,
|
||||
*,
|
||||
offset: str = '0s',
|
||||
duration: str = '1s',
|
||||
lane: int = 1,
|
||||
position: Optional[str] = None,
|
||||
font: str = 'Helvetica Neue',
|
||||
font_size: int = 196,
|
||||
font_color: str = '1 1 1 1',
|
||||
bold: bool = True,
|
||||
face: Optional[str] = None,
|
||||
animated: bool = True,
|
||||
font_scale: float = TEXT_TEMPLATE_FONT_SCALE,
|
||||
size_param: Optional[float] = None,
|
||||
role: Optional[str] = None,
|
||||
) -> ET.Element:
|
||||
"""Add a single static "Text" (Basic Text) title over *parent_clip*.
|
||||
|
||||
Anchored in SOURCE media coordinates (parent's ``start`` + *offset*),
|
||||
matching FCP's own output, so the title lands on screen instead of at
|
||||
~0s of the media (which FCP silently drops). *offset* and *duration*
|
||||
accept any FCPXML rational-time string; *position* is an optional
|
||||
"x y" canvas-point string to keep two titles from stacking.
|
||||
|
||||
Returns:
|
||||
The created ``<title>`` element, already inserted into the parent
|
||||
in DTD order.
|
||||
"""
|
||||
parent = parent_clip if isinstance(parent_clip, ET.Element) else self._require_clip(parent_clip)
|
||||
resources = self.root.find('.//resources')
|
||||
if resources is None:
|
||||
raise ValueError("No <resources> element found in FCPXML")
|
||||
effect_id = self._ensure_text_title_effect(resources)
|
||||
|
||||
media_origin = self._parse_time(parent.get('start', '0s'))
|
||||
relative = self._parse_time(offset)
|
||||
title = self._make_text_title_clip(
|
||||
effect_id,
|
||||
text,
|
||||
media_origin + relative,
|
||||
self._parse_time(duration),
|
||||
lane=lane,
|
||||
name=f"{text} - Text",
|
||||
position=position,
|
||||
font=font,
|
||||
font_size=font_size,
|
||||
font_color=font_color,
|
||||
bold=bold,
|
||||
face=face,
|
||||
animated=animated,
|
||||
font_scale=font_scale,
|
||||
size_param=size_param,
|
||||
role=role,
|
||||
)
|
||||
_dtd_insert(parent, title)
|
||||
return title
|
||||
|
||||
def generate_dynamic_subtitles(
|
||||
self,
|
||||
parent_clip: 'str | ET.Element',
|
||||
words: List[Dict[str, Any]],
|
||||
config: Optional['DynamicSubtitleConfig'] = None,
|
||||
segments: Optional[List[Dict[str, Any]]] = None,
|
||||
role: Optional[str] = None,
|
||||
configs: Optional[List['DynamicSubtitleConfig']] = None,
|
||||
compound_subphrases: bool = False,
|
||||
subphrase_min_words: int = 3,
|
||||
hold_between_sentences: bool = True,
|
||||
) -> List[ET.Element]:
|
||||
"""Generate progressive-reveal subtitle titles, one per word.
|
||||
|
||||
Groups *words* into sentences (by *segments*' time windows), lays each
|
||||
sentence out as a compact typographic block, and emits one standalone
|
||||
``<title>`` per word, positioned at its place in that block. Words
|
||||
appear one by one as they are spoken and accumulate on screen; every
|
||||
word of a block then clears at the same instant, so the sentence
|
||||
vanishes as a whole before the next one builds up.
|
||||
|
||||
Each word gets its own lane, since a block's words are all on screen
|
||||
together. Lanes restart with each block. Size, colour, font and face
|
||||
cycle through ``config.style.rhythm``, reproducing the typography of
|
||||
the calibration export the user built in Final Cut.
|
||||
|
||||
A sentence too tall for the band is split into successive blocks, so a
|
||||
long sentence never spills off screen.
|
||||
|
||||
Args:
|
||||
parent_clip: The spine clip to attach titles to — either its
|
||||
Name/ID (resolved via ``_require_clip``, kept for backward
|
||||
compatibility) or the ``ET.Element`` itself. **Callers
|
||||
iterating multiple spine clips must pass the element, not
|
||||
the name**: after any ripple-cut/silence-removal operation,
|
||||
every fragment of an originally-named clip keeps that same
|
||||
``name``, so ``self.clips`` (keyed by name) only retains the
|
||||
last-indexed one — a name lookup then silently resolves
|
||||
every call to the SAME wrong clip, stacking every line from
|
||||
every distinct clip's transcript onto one spine element (see
|
||||
Engine/docs/05_EXPERIENCIAS.md, entry 2026-08-17).
|
||||
words: ``[{'word': str, 'start': float, 'end': float}, ...]``
|
||||
with ``start``/``end`` in seconds *relative to the parent
|
||||
clip's own start* (same convention as ``add_connected_clip``'s
|
||||
``offset``).
|
||||
config: Styling/layout options; defaults to ``DynamicSubtitleConfig()``.
|
||||
segments: Whisper sentence segments ``[{'start', 'end', ...}]``, on
|
||||
the same relative timebase as *words*. Omitted, every word
|
||||
falls into a single sentence, which the block layout then
|
||||
splits by height alone.
|
||||
|
||||
Returns:
|
||||
The list of created ``<title>`` elements, in chronological order.
|
||||
"""
|
||||
# ``configs`` (a list of registered, active layouts) takes precedence
|
||||
# over the single ``config`` — with 2+ items, each block picks one at
|
||||
# random below; with 0 or 1, behaviour is identical to a single fixed
|
||||
# config, so old callers passing only ``config`` are unaffected.
|
||||
if configs:
|
||||
layout_configs = list(configs)
|
||||
elif config is not None:
|
||||
layout_configs = [config]
|
||||
else:
|
||||
layout_configs = [DynamicSubtitleConfig()]
|
||||
if not words:
|
||||
return []
|
||||
|
||||
parent = parent_clip if isinstance(parent_clip, ET.Element) else self._require_clip(parent_clip)
|
||||
resources = self.root.find('.//resources')
|
||||
if resources is None:
|
||||
raise ValueError("No <resources> element found in FCPXML")
|
||||
effect_id = self._ensure_text_title_effect(resources)
|
||||
|
||||
# A connected title is NOT trimmed by its parent clip's out-point —
|
||||
# Final Cut keeps drawing it over whatever clip follows. A word that
|
||||
# starts after the cut would therefore only ever be seen on top of the
|
||||
# NEXT clip's own captions, so it is dropped rather than placed.
|
||||
parent_limit = self._parse_time(parent.get('duration', '0s'))
|
||||
has_limit = TimeValue(0, 1) < parent_limit
|
||||
if has_limit:
|
||||
limit_seconds = parent_limit.to_seconds()
|
||||
words = [
|
||||
w for w in words
|
||||
if float(w.get('start', 0.0)) < limit_seconds
|
||||
]
|
||||
if not words:
|
||||
return []
|
||||
|
||||
# Split into sentences, then lay each one out as a block. A sentence
|
||||
# too tall for the band comes back with overflow, which becomes the
|
||||
# next block — the sub-sentence split that keeps long sentences from
|
||||
# spilling off screen.
|
||||
sentences = group_words_by_segment(words, segments or [])
|
||||
# A comma is where the sentence breathes, so it is also where the
|
||||
# phrase should be packed into its own compound clip downstream.
|
||||
if compound_subphrases:
|
||||
sentences = [
|
||||
sub
|
||||
for sentence in sentences
|
||||
for sub in split_into_subphrases(sentence, subphrase_min_words)
|
||||
]
|
||||
|
||||
def box_for(cfg: 'DynamicSubtitleConfig') -> LayoutBox:
|
||||
return LayoutBox.for_frame(
|
||||
self.frame_width(), self.frame_height(),
|
||||
band_height=cfg.band_height,
|
||||
center_y=cfg.block_center_y,
|
||||
)
|
||||
|
||||
def lay_out(pending: List[Dict], cfg: 'DynamicSubtitleConfig', box: LayoutBox):
|
||||
"""Place what fits; return (units, still-unplaced words)."""
|
||||
# "phrase" is the progressive composition the reference reel uses:
|
||||
# one title per LINE ("que vão" / "melhorar" / "sua legenda"), the
|
||||
# key word set large in a display italic. "word" is the older
|
||||
# one-title-per-word rhythm, kept for callers that want every word
|
||||
# to land on its own.
|
||||
if getattr(cfg, 'granularity', 'phrase') == 'phrase':
|
||||
composition = compose_sentence(
|
||||
pending, cfg.style, box, line_gap=cfg.line_gap,
|
||||
)
|
||||
return composition.blocks, composition.overflow
|
||||
layout = layout_sentence(pending, cfg.style, box)
|
||||
return layout.placed, layout.overflow
|
||||
|
||||
blocks: List[List[Any]] = []
|
||||
block_configs: List['DynamicSubtitleConfig'] = []
|
||||
block_sentences: List[int] = []
|
||||
for sentence_index, sentence in enumerate(sentences):
|
||||
remaining = list(sentence)
|
||||
while remaining:
|
||||
# Each block independently samples a layout from the active
|
||||
# set — the visual variety the user asked for. A single
|
||||
# active layout always resolves to itself, so this is a
|
||||
# no-op for the common case.
|
||||
active_config = layout_configs[random.randrange(len(layout_configs))]
|
||||
units, remaining = lay_out(remaining, active_config, box_for(active_config))
|
||||
if not units:
|
||||
break
|
||||
blocks.append(units)
|
||||
block_configs.append(active_config)
|
||||
block_sentences.append(sentence_index)
|
||||
if not blocks:
|
||||
return []
|
||||
|
||||
# Never emit a zero-duration frame (rounds to 0 at the sequence's fps
|
||||
# and FCP rejects it as "unexpected value found").
|
||||
min_dur_tv = self.snap_seconds_to_frame(
|
||||
float(self.frame_duration_fraction())
|
||||
)
|
||||
|
||||
# Every word of a block clears at the same instant: when the next block
|
||||
# starts, or at the last word's end for the final block. That is what
|
||||
# makes a sentence build up and then vanish all at once.
|
||||
block_starts = [
|
||||
self.snap_seconds_to_frame(min(unit.start for unit in units))
|
||||
for units in blocks
|
||||
]
|
||||
# Whisper's word end can also run past the cut, so a last block would
|
||||
# linger over the next clip's first block. Nothing may outlive the
|
||||
# clip it was written for.
|
||||
block_ends: List[TimeValue] = []
|
||||
for i, units in enumerate(blocks):
|
||||
if i + 1 < len(blocks):
|
||||
end = block_starts[i + 1]
|
||||
if not hold_between_sentences and block_sentences[i] != block_sentences[i + 1]:
|
||||
spoken_end = self.snap_seconds_to_frame(max(unit.end for unit in units))
|
||||
end = min(end, spoken_end)
|
||||
else:
|
||||
end = self.snap_seconds_to_frame(
|
||||
max(unit.end for unit in units)
|
||||
)
|
||||
if end - block_starts[i] < min_dur_tv:
|
||||
end = block_starts[i] + min_dur_tv
|
||||
if has_limit and parent_limit < end:
|
||||
end = parent_limit
|
||||
block_ends.append(end)
|
||||
|
||||
# Anchored titles are positioned in the parent clip's SOURCE media
|
||||
# coordinates: a title's offset is the parent clip's `start` plus its
|
||||
# timeline-relative position. Verified against FCP's own output in
|
||||
# "exemplo de arquivos.fcpxmld", where the hand-made "Essencial -
|
||||
# Título" sits at offset 226040815/24000s on a parent starting at
|
||||
# 226007782/24000s — 1.376s into a 1.835s clip. Writing a plain
|
||||
# relative offset instead would drop the title to ~0s of the media,
|
||||
# before the clip's own in-point, so it lands outside the clip and FCP
|
||||
# never shows it.
|
||||
media_origin = self._parse_time(parent.get('start', '0s'))
|
||||
|
||||
created: List[ET.Element] = []
|
||||
by_sentence: Dict[int, List[ET.Element]] = {}
|
||||
for units, block_end, block_config, sentence_index in zip(
|
||||
blocks, block_ends, block_configs, block_sentences
|
||||
):
|
||||
# A ``titles.*`` sub-role keeps these as titles (never closed
|
||||
# captions) while grouping them in the role index and tinting
|
||||
# their lane. An explicit ``role`` argument overrides every
|
||||
# block; otherwise each block uses its own sampled layout's role.
|
||||
block_role = role or getattr(block_config, "role", None) or "titles.dinamicas"
|
||||
for index, unit in enumerate(units):
|
||||
relative_offset = self.snap_seconds_to_frame(unit.start)
|
||||
duration = block_end - relative_offset
|
||||
if duration < min_dur_tv:
|
||||
duration = min_dur_tv
|
||||
|
||||
# Units of one block are all on screen together, so no two may
|
||||
# share a lane. Lanes restart each block, which is free — the
|
||||
# previous block has already cleared.
|
||||
lane = index + 1
|
||||
|
||||
offset = media_origin + relative_offset
|
||||
title = self._make_text_title_clip(
|
||||
effect_id,
|
||||
unit.text,
|
||||
offset,
|
||||
duration,
|
||||
lane=lane,
|
||||
name=f"caption_{uuid.uuid4().hex[:8]}",
|
||||
position=unit.position_param(block_config.text_scale),
|
||||
font=unit.font or block_config.style.font,
|
||||
font_size=int(round(unit.font_size)),
|
||||
font_color=unit.color or block_config.style.active_color,
|
||||
bold=block_config.style.bold,
|
||||
face=unit.face,
|
||||
kerning=unit.kerning,
|
||||
font_scale=block_config.text_scale,
|
||||
role=block_role,
|
||||
)
|
||||
_dtd_insert(parent, title)
|
||||
self.mark_generated_subtitle(title, 'dynamic')
|
||||
created.append(title)
|
||||
by_sentence.setdefault(sentence_index, []).append(title)
|
||||
|
||||
# One compound per sub-phrase: a dozen stacked title bars collapse
|
||||
# into a single one that can be dragged, muted or retimed as a unit.
|
||||
if compound_subphrases:
|
||||
for sentence_index in sorted(by_sentence):
|
||||
group = by_sentence[sentence_index]
|
||||
label = " ".join(
|
||||
str(w.get('word') or w.get('text') or '')
|
||||
for w in sentences[sentence_index]
|
||||
).strip()
|
||||
compound = self.wrap_titles_in_compound(
|
||||
parent, group, name=label[:60] or "Legenda"
|
||||
)
|
||||
self.mark_generated_subtitle(compound, 'dynamic')
|
||||
|
||||
if any(getattr(cfg, 'validate', False) for cfg in layout_configs):
|
||||
report = self.validate_subtitle_layout()
|
||||
if blocking(report["severity"]):
|
||||
raise ValueError(
|
||||
"Subtitle layout validation failed: "
|
||||
+ str(report["summary"])
|
||||
)
|
||||
|
||||
return created
|
||||
|
||||
def validate_subtitle_layout(
|
||||
self,
|
||||
*,
|
||||
safe_margin_x: float = 0.05,
|
||||
safe_margin_y: float = 0.05,
|
||||
min_font_size: Optional[float] = None,
|
||||
min_distance: Optional[float] = None,
|
||||
max_distance: Optional[float] = None,
|
||||
) -> dict:
|
||||
"""Re-measure every ``<title>`` in the document and report collisions.
|
||||
|
||||
Reconstructs each title's on-screen box from the values the writer
|
||||
emitted (``fontSize``/``kerning``/``Position`` are already in template
|
||||
space), then checks for temporal+spatial collisions, frame/safe-area
|
||||
containment, and font fallbacks. This is the spec-16 validation pass the
|
||||
layout engine does not do on its own — it only guarantees non-overlap
|
||||
*by construction* while composing, and cannot see a hand-edited title.
|
||||
|
||||
Returns the ``collision.validate_titles`` report: ``severity`` (worst
|
||||
bucket), ``issues`` (spec-16 occurrences) and ``summary`` (counts).
|
||||
"""
|
||||
# A compound clip carries its own time origin: a title inside one is
|
||||
# offset from that compound's start, not the sequence's. Measured in
|
||||
# one flat pass, the anchors of two different compounds both read as
|
||||
# "0s" and collide on paper while sitting seconds apart on the
|
||||
# timeline. Each compound is therefore measured as its own scope,
|
||||
# which is also where its titles can actually overlap — a title can
|
||||
# only share the screen with its own compound's siblings.
|
||||
scopes: List[List[ET.Element]] = []
|
||||
nested: set = set()
|
||||
for media in self.root.findall('.//media'):
|
||||
group = list(media.iter('title'))
|
||||
if group:
|
||||
scopes.append(group)
|
||||
nested.update(id(t) for t in group)
|
||||
main = [t for t in self.root.iter('title') if id(t) not in nested]
|
||||
if main:
|
||||
scopes.append(main)
|
||||
|
||||
reports = [
|
||||
self._measure_title_scope(
|
||||
scope,
|
||||
safe_margin_x=safe_margin_x,
|
||||
safe_margin_y=safe_margin_y,
|
||||
min_font_size=min_font_size,
|
||||
min_distance=min_distance,
|
||||
max_distance=max_distance,
|
||||
)
|
||||
for scope in scopes
|
||||
]
|
||||
if len(reports) == 1:
|
||||
return reports[0]
|
||||
if not reports:
|
||||
return self._measure_title_scope(
|
||||
[],
|
||||
safe_margin_x=safe_margin_x,
|
||||
safe_margin_y=safe_margin_y,
|
||||
min_font_size=min_font_size,
|
||||
min_distance=min_distance,
|
||||
max_distance=max_distance,
|
||||
)
|
||||
|
||||
rank = {
|
||||
'none': 0, 'render_tolerance': 1, 'warning': 2,
|
||||
'probable': 3, 'severe': 4,
|
||||
}
|
||||
merged_issues = [i for r in reports for i in r['issues']]
|
||||
summary = dict(reports[0]['summary'])
|
||||
for r in reports[1:]:
|
||||
for key, value in r['summary'].items():
|
||||
summary[key] = summary.get(key, 0) + value
|
||||
return {
|
||||
'severity': max(
|
||||
(r['severity'] for r in reports),
|
||||
key=lambda s: rank.get(s, 0),
|
||||
),
|
||||
'issues': merged_issues,
|
||||
'summary': summary,
|
||||
}
|
||||
|
||||
def _measure_title_scope(
|
||||
self,
|
||||
elements: List[ET.Element],
|
||||
*,
|
||||
safe_margin_x: float,
|
||||
safe_margin_y: float,
|
||||
min_font_size: Optional[float],
|
||||
min_distance: Optional[float],
|
||||
max_distance: Optional[float],
|
||||
) -> dict:
|
||||
"""Measure and validate one group of titles sharing a time origin."""
|
||||
titles = []
|
||||
for elem in elements:
|
||||
# enabled="0" never renders in Final Cut (see
|
||||
# generate_subtitles_by_emphasis, which disables plain titles
|
||||
# under an emphasis phrase instead of never creating them) — a
|
||||
# title that is off by design must not count as a collision
|
||||
# against the one drawn in its place.
|
||||
if elem.get('enabled', '1') == '0':
|
||||
continue
|
||||
text_el = elem.find('text/text-style')
|
||||
text = (text_el.text or '').strip() if text_el is not None else ''
|
||||
style = elem.find('text-style-def/text-style')
|
||||
font = style.get('font') if style is not None else None
|
||||
face = style.get('fontFace') if style is not None else None
|
||||
font_size = (
|
||||
float(style.get('fontSize', '0')) if style is not None else 0.0
|
||||
)
|
||||
kerning = (
|
||||
float(style.get('kerning', '0') or 0)
|
||||
if style is not None else 0.0
|
||||
)
|
||||
|
||||
x = y = 0.0
|
||||
for param in elem.findall('param'):
|
||||
if param.get('name') == 'Position' and param.get('value'):
|
||||
parts = param.get('value').split()
|
||||
if len(parts) >= 2:
|
||||
x, y = float(parts[0]), float(parts[1])
|
||||
|
||||
start = self._parse_time(elem.get('offset', '0s')).to_seconds()
|
||||
duration = self._parse_time(elem.get('duration', '0s')).to_seconds()
|
||||
|
||||
titles.append({
|
||||
'text': text,
|
||||
'font': font,
|
||||
'face': face,
|
||||
'font_size': font_size,
|
||||
'kerning': kerning,
|
||||
'x': x,
|
||||
'y': y,
|
||||
'start': start,
|
||||
'end': start + duration,
|
||||
'group': start + duration,
|
||||
})
|
||||
|
||||
return validate_titles(
|
||||
titles,
|
||||
self.frame_width(),
|
||||
self.frame_height(),
|
||||
safe_margin_x=safe_margin_x,
|
||||
safe_margin_y=safe_margin_y,
|
||||
min_font_size=min_font_size,
|
||||
min_distance=min_distance,
|
||||
max_distance=max_distance,
|
||||
)
|
||||
|
||||
# ========================================================================
|
||||
@@ -0,0 +1,94 @@
|
||||
"""Transições entre clipes vizinhos.
|
||||
|
||||
Extraído de writer.py — ver fcpxml/writer/__init__.py para o conjunto.
|
||||
"""
|
||||
|
||||
import xml.etree.ElementTree as ET
|
||||
|
||||
from ..models import (
|
||||
TimeValue,
|
||||
)
|
||||
from .helpers import FCP_EFFECTS
|
||||
|
||||
|
||||
class TransitionsMixin:
|
||||
"""Transições entre clipes vizinhos."""
|
||||
|
||||
# TRANSITION OPERATIONS
|
||||
# ========================================================================
|
||||
|
||||
def add_transition(
|
||||
self,
|
||||
clip_id: str,
|
||||
position: str = 'end',
|
||||
transition_type: str = 'cross-dissolve',
|
||||
duration: str = '00:00:00:15'
|
||||
) -> ET.Element:
|
||||
"""
|
||||
Add a transition to a clip.
|
||||
|
||||
Args:
|
||||
clip_id: Target clip
|
||||
position: 'start', 'end', or 'both'
|
||||
transition_type: Type of transition
|
||||
duration: Transition duration
|
||||
|
||||
Returns:
|
||||
Created transition element(s)
|
||||
"""
|
||||
spine, clip, clip_index = self._require_spine_clip(clip_id)
|
||||
|
||||
trans_duration = self._parse_time(duration)
|
||||
|
||||
# Effect name and FCP built-in effect UID lookup via registry
|
||||
effect_name, effect_uid = FCP_EFFECTS.get(
|
||||
transition_type,
|
||||
FCP_EFFECTS['cross-dissolve']
|
||||
)
|
||||
|
||||
# Ensure effect resource exists in <resources>
|
||||
effect_ref_id = None
|
||||
if effect_uid:
|
||||
root = self.tree.getroot()
|
||||
resources = root.find('.//resources')
|
||||
if resources is not None:
|
||||
for eff in resources.findall('effect'):
|
||||
if eff.get('uid') == effect_uid:
|
||||
effect_ref_id = eff.get('id')
|
||||
break
|
||||
if effect_ref_id is None:
|
||||
effect_ref_id = self._unique_resource_id(resources, 'r_dissolve')
|
||||
eff_el = ET.SubElement(resources, 'effect')
|
||||
eff_el.set('id', effect_ref_id)
|
||||
eff_el.set('name', effect_name)
|
||||
eff_el.set('uid', effect_uid)
|
||||
|
||||
transitions_added = []
|
||||
|
||||
_, clip_dur, clip_offset = self._get_clip_times(clip)
|
||||
half_dur = trans_duration * 0.5
|
||||
|
||||
if position in ('end', 'both'):
|
||||
end_offset = clip_offset + clip_dur - half_dur
|
||||
transition = self._make_transition_element(
|
||||
effect_name, end_offset, trans_duration, effect_ref_id
|
||||
)
|
||||
spine.insert(clip_index + 1, transition)
|
||||
transitions_added.append(transition)
|
||||
|
||||
if position in ('start', 'both'):
|
||||
start_offset = clip_offset - half_dur
|
||||
if start_offset < TimeValue.zero():
|
||||
raise ValueError(
|
||||
f"Transition at start would produce negative offset "
|
||||
f"({start_offset.to_seconds():.3f}s) for clip '{clip_id}'"
|
||||
)
|
||||
transition = self._make_transition_element(
|
||||
effect_name, start_offset, trans_duration, effect_ref_id
|
||||
)
|
||||
spine.insert(clip_index, transition)
|
||||
transitions_added.append(transition)
|
||||
|
||||
return transitions_added[0] if len(transitions_added) == 1 else transitions_added
|
||||
|
||||
# ========================================================================
|
||||
@@ -0,0 +1,125 @@
|
||||
"""Aparar clipes e propagar o ripple pela spine.
|
||||
|
||||
Extraído de writer.py — ver fcpxml/writer/__init__.py para o conjunto.
|
||||
"""
|
||||
|
||||
import xml.etree.ElementTree as ET
|
||||
from typing import Optional
|
||||
|
||||
from ..models import (
|
||||
TimeValue,
|
||||
)
|
||||
from .helpers import SPINE_ELEMENT_TAGS
|
||||
|
||||
|
||||
class TrimMixin:
|
||||
"""Aparar clipes e propagar o ripple pela spine."""
|
||||
|
||||
# ========================================================================
|
||||
# TRIM OPERATIONS
|
||||
# ========================================================================
|
||||
|
||||
def trim_clip(
|
||||
self,
|
||||
clip_id: str,
|
||||
trim_start: Optional[str] = None,
|
||||
trim_end: Optional[str] = None,
|
||||
ripple: bool = True
|
||||
) -> ET.Element:
|
||||
"""
|
||||
Trim a clip's in-point and/or out-point.
|
||||
|
||||
Args:
|
||||
clip_id: Target clip
|
||||
trim_start: New in-point or delta ('+1s', '-10f')
|
||||
trim_end: New out-point or delta
|
||||
ripple: Whether to shift subsequent clips
|
||||
|
||||
Returns:
|
||||
Modified clip element
|
||||
"""
|
||||
clip = self._require_clip(clip_id)
|
||||
|
||||
current_start, current_duration, _ = self._get_clip_times(clip)
|
||||
|
||||
original_duration = current_duration
|
||||
|
||||
# Handle trim_start
|
||||
if trim_start:
|
||||
if trim_start.startswith('+') or trim_start.startswith('-'):
|
||||
delta = self._parse_time(trim_start[1:])
|
||||
if trim_start.startswith('-'):
|
||||
# Extend earlier
|
||||
new_start = current_start - delta
|
||||
new_duration = current_duration + delta
|
||||
else:
|
||||
# Trim later
|
||||
new_start = current_start + delta
|
||||
new_duration = current_duration - delta
|
||||
else:
|
||||
new_start = self._parse_time(trim_start)
|
||||
diff = new_start - current_start
|
||||
new_duration = current_duration - diff
|
||||
|
||||
clip.set('start', new_start.to_fcpxml())
|
||||
current_start = new_start
|
||||
current_duration = new_duration
|
||||
|
||||
# Handle trim_end
|
||||
if trim_end:
|
||||
if trim_end.startswith('+') or trim_end.startswith('-'):
|
||||
delta = self._parse_time(trim_end[1:])
|
||||
if trim_end.startswith('-'):
|
||||
new_duration = current_duration - delta
|
||||
else:
|
||||
new_duration = current_duration + delta
|
||||
else:
|
||||
end_point = self._parse_time(trim_end)
|
||||
new_duration = end_point - current_start
|
||||
|
||||
current_duration = new_duration
|
||||
|
||||
if current_duration <= TimeValue.zero():
|
||||
raise ValueError(
|
||||
f"Trim would produce non-positive duration "
|
||||
f"({current_duration.to_seconds():.3f}s) for clip '{clip_id}'"
|
||||
)
|
||||
|
||||
clip.set('duration', current_duration.to_fcpxml())
|
||||
|
||||
# Ripple subsequent clips if needed
|
||||
if ripple:
|
||||
duration_change = current_duration - original_duration
|
||||
if duration_change != TimeValue.zero():
|
||||
self._ripple_after_clip(clip, duration_change)
|
||||
|
||||
return clip
|
||||
|
||||
def _ripple_from_index(
|
||||
self, spine: ET.Element, start_index: int, delta: 'TimeValue'
|
||||
) -> None:
|
||||
"""Shift the offset of every spine element from *start_index* onward by *delta*.
|
||||
|
||||
Consolidates the ripple loops previously duplicated across
|
||||
``_ripple_after_clip``, ``delete_clip``, and ``insert_clip``.
|
||||
|
||||
Args:
|
||||
spine: The primary storyline ``<spine>`` element.
|
||||
start_index: First child index to adjust (inclusive).
|
||||
delta: Signed time shift (positive = later, negative = earlier).
|
||||
"""
|
||||
children = list(spine)
|
||||
for child in children[start_index:]:
|
||||
if child.tag in SPINE_ELEMENT_TAGS:
|
||||
current_offset = self._parse_time(child.get('offset', '0s'))
|
||||
new_offset = current_offset + delta
|
||||
child.set('offset', new_offset.to_fcpxml())
|
||||
|
||||
def _ripple_after_clip(self, target_clip: ET.Element, delta: TimeValue) -> None:
|
||||
"""Shift all clips after the given clip by delta."""
|
||||
spine = self._get_spine()
|
||||
clip_index = self._find_clip_index(spine, target_clip)
|
||||
if clip_index is not None:
|
||||
self._ripple_from_index(spine, clip_index + 1, delta)
|
||||
|
||||
# ========================================================================
|
||||
@@ -0,0 +1,232 @@
|
||||
"""Verificações estruturais do FCPXML antes de salvar.
|
||||
|
||||
Extraído de writer.py — ver fcpxml/writer/__init__.py para o conjunto.
|
||||
"""
|
||||
|
||||
import logging
|
||||
import xml.etree.ElementTree as ET
|
||||
from fractions import Fraction
|
||||
from typing import List, Optional
|
||||
|
||||
from ..models import (
|
||||
_FCPXML_STANDARD_TIMEBASES,
|
||||
TimeValue,
|
||||
ValidationIssue,
|
||||
ValidationIssueType,
|
||||
)
|
||||
from .helpers import _ASSET_CLIP_CHILD_ORDER, _CHILD_ORDER_INDEX
|
||||
|
||||
# ============================================================================
|
||||
# PRE-EXPORT DTD VALIDATOR (v0.6.0)
|
||||
# ============================================================================
|
||||
|
||||
_log = logging.getLogger(__name__)
|
||||
|
||||
|
||||
def _check_child_order(root: ET.Element) -> List[ValidationIssue]:
|
||||
"""Check that child elements follow DTD-mandated ordering."""
|
||||
issues = []
|
||||
for parent in root.iter():
|
||||
if parent.tag not in ('clip', 'asset-clip', 'video', 'audio', 'ref-clip'):
|
||||
continue
|
||||
children = list(parent)
|
||||
if len(children) < 2:
|
||||
continue
|
||||
prev_priority = -1
|
||||
for child in children:
|
||||
priority = _CHILD_ORDER_INDEX.get(child.tag, len(_ASSET_CLIP_CHILD_ORDER))
|
||||
if priority < prev_priority:
|
||||
issues.append(ValidationIssue(
|
||||
issue_type=ValidationIssueType.ELEMENT_ORDER,
|
||||
severity="warning",
|
||||
message=(
|
||||
f"<{child.tag}> appears after a higher-priority sibling "
|
||||
f"in <{parent.tag}> '{parent.get('name', '')}'."
|
||||
),
|
||||
clip_name=parent.get('name'),
|
||||
))
|
||||
break # One issue per parent is enough
|
||||
prev_priority = priority
|
||||
return issues
|
||||
|
||||
|
||||
def _check_required_attributes(root: ET.Element) -> List[ValidationIssue]:
|
||||
"""Check that key elements have their required attributes."""
|
||||
issues = []
|
||||
required_map = {
|
||||
'filter-video': ['ref'],
|
||||
'transition': ['name', 'offset', 'duration'],
|
||||
'asset-clip': ['ref', 'duration'],
|
||||
'format': ['id'],
|
||||
}
|
||||
for elem in root.iter():
|
||||
attrs = required_map.get(elem.tag)
|
||||
if not attrs:
|
||||
continue
|
||||
for attr in attrs:
|
||||
if not elem.get(attr):
|
||||
issues.append(ValidationIssue(
|
||||
issue_type=ValidationIssueType.MISSING_ATTRIBUTE,
|
||||
severity="error",
|
||||
message=f"<{elem.tag}> missing required attribute '{attr}'.",
|
||||
clip_name=elem.get('name'),
|
||||
))
|
||||
return issues
|
||||
|
||||
|
||||
def _check_timebases(root: ET.Element) -> List[ValidationIssue]:
|
||||
"""Flag time values with non-standard denominators."""
|
||||
issues = []
|
||||
time_attrs = ('offset', 'start', 'duration')
|
||||
seen: set = set()
|
||||
for elem in root.iter():
|
||||
for attr in time_attrs:
|
||||
val = elem.get(attr)
|
||||
if val and val.endswith('s') and '/' in val:
|
||||
try:
|
||||
tv = TimeValue.from_timecode(val)
|
||||
denom = tv.simplify().denominator
|
||||
if denom not in _FCPXML_STANDARD_TIMEBASES:
|
||||
key = (elem.tag, attr, val)
|
||||
if key not in seen:
|
||||
seen.add(key)
|
||||
issues.append(ValidationIssue(
|
||||
issue_type=ValidationIssueType.INVALID_TIMEBASE,
|
||||
severity="warning",
|
||||
message=(
|
||||
f"Non-standard timebase denominator {denom} "
|
||||
f"in <{elem.tag}> {attr}=\"{val}\"."
|
||||
),
|
||||
clip_name=elem.get('name'),
|
||||
))
|
||||
except (ValueError, ZeroDivisionError):
|
||||
pass
|
||||
return issues
|
||||
|
||||
|
||||
def _document_frame_duration(root: ET.Element) -> Optional[Fraction]:
|
||||
"""The sequence's exact ``frameDuration`` as a fraction, if declared.
|
||||
|
||||
Read from the format the ``<sequence>`` references (falling back to the
|
||||
first declared format), so the value is the document's own timebase
|
||||
rather than an assumed rate.
|
||||
"""
|
||||
formats = {f.get('id'): f for f in root.findall('.//format') if f.get('id')}
|
||||
sequence = root.find('.//sequence')
|
||||
fmt = formats.get(sequence.get('format')) if sequence is not None else None
|
||||
if fmt is None:
|
||||
fmt = next(iter(formats.values()), None)
|
||||
if fmt is None:
|
||||
return None
|
||||
raw = fmt.get('frameDuration', '')
|
||||
if not (raw.endswith('s') and '/' in raw):
|
||||
return None
|
||||
numerator, denominator = raw[:-1].split('/', 1)
|
||||
try:
|
||||
value = Fraction(int(numerator), int(denominator))
|
||||
except (ValueError, ZeroDivisionError):
|
||||
return None
|
||||
return value if value > 0 else None
|
||||
|
||||
|
||||
def _check_frame_alignment(root: ET.Element, fps: float = 24.0) -> List[ValidationIssue]:
|
||||
"""Check that durations are integer multiples of the frame duration.
|
||||
|
||||
Uses the document's exact ``frameDuration`` fraction and rational
|
||||
arithmetic. Comparing against an integer fps instead would flag every
|
||||
NTSC project as broken: at 1001/24000s (23.976fps) a perfectly aligned
|
||||
duration is not an integer number of "24fps" frames, so whole timelines
|
||||
would be reported misaligned when nothing is wrong.
|
||||
"""
|
||||
issues = []
|
||||
frame_duration = _document_frame_duration(root)
|
||||
label = f"{1 / float(frame_duration):.3f}".rstrip('0').rstrip('.') if frame_duration else str(fps)
|
||||
for elem in root.iter():
|
||||
dur_str = elem.get('duration')
|
||||
if not dur_str or not dur_str.endswith('s'):
|
||||
continue
|
||||
if elem.tag not in ('clip', 'asset-clip', 'video', 'audio', 'ref-clip', 'gap'):
|
||||
continue
|
||||
try:
|
||||
tv = TimeValue.from_timecode(dur_str)
|
||||
if frame_duration is not None:
|
||||
frames = Fraction(tv.numerator, tv.denominator) / frame_duration
|
||||
aligned = frames.denominator == 1
|
||||
else:
|
||||
approx = tv.to_seconds() * fps
|
||||
aligned = abs(approx - round(approx)) <= 0.01
|
||||
if not aligned:
|
||||
issues.append(ValidationIssue(
|
||||
issue_type=ValidationIssueType.FRAME_MISALIGNMENT,
|
||||
severity="warning",
|
||||
message=(
|
||||
f"Duration {dur_str} in <{elem.tag}> "
|
||||
f"'{elem.get('name', '')}' is not frame-aligned at {label}fps."
|
||||
),
|
||||
clip_name=elem.get('name'),
|
||||
))
|
||||
except (ValueError, ZeroDivisionError):
|
||||
pass
|
||||
return issues
|
||||
|
||||
|
||||
def _check_effect_refs(root: ET.Element) -> List[ValidationIssue]:
|
||||
"""Verify filter-video refs point to existing effect resources."""
|
||||
issues = []
|
||||
resource_ids = set()
|
||||
for res in root.iter():
|
||||
rid = res.get('id')
|
||||
if rid and res.tag in ('effect', 'format', 'asset', 'media'):
|
||||
resource_ids.add(rid)
|
||||
|
||||
for fv in root.iter('filter-video'):
|
||||
ref = fv.get('ref')
|
||||
if ref and ref not in resource_ids:
|
||||
issues.append(ValidationIssue(
|
||||
issue_type=ValidationIssueType.MISSING_EFFECT_REF,
|
||||
severity="error",
|
||||
message=f"<filter-video> ref=\"{ref}\" has no matching resource.",
|
||||
))
|
||||
return issues
|
||||
|
||||
|
||||
def _check_asset_sources(root: ET.Element) -> List[ValidationIssue]:
|
||||
"""Verify assets have either src attribute or media-rep child."""
|
||||
issues = []
|
||||
for asset in root.iter('asset'):
|
||||
src = asset.get('src', '')
|
||||
media_rep = asset.find('media-rep')
|
||||
if not src and media_rep is None:
|
||||
issues.append(ValidationIssue(
|
||||
issue_type=ValidationIssueType.MISSING_MEDIA_REP,
|
||||
severity="warning",
|
||||
message=(
|
||||
f"<asset id=\"{asset.get('id', '?')}\" "
|
||||
f"name=\"{asset.get('name', '')}\"> "
|
||||
f"has no src attribute and no <media-rep> child."
|
||||
),
|
||||
clip_name=asset.get('name'),
|
||||
))
|
||||
return issues
|
||||
|
||||
|
||||
def validate_fcpxml(root: ET.Element, fps: float = 24.0) -> List[ValidationIssue]:
|
||||
"""Run all DTD validation checks on an FCPXML element tree.
|
||||
|
||||
Args:
|
||||
root: The <fcpxml> root Element to validate.
|
||||
fps: Frame rate for alignment checks (default 24).
|
||||
|
||||
Returns:
|
||||
List of ValidationIssue objects. Empty list = clean.
|
||||
"""
|
||||
issues: List[ValidationIssue] = []
|
||||
issues.extend(_check_child_order(root))
|
||||
issues.extend(_check_required_attributes(root))
|
||||
issues.extend(_check_timebases(root))
|
||||
issues.extend(_check_frame_alignment(root, fps))
|
||||
issues.extend(_check_effect_refs(root))
|
||||
issues.extend(_check_asset_sources(root))
|
||||
return issues
|
||||
|
||||
|
||||
+4
-1
@@ -41,6 +41,9 @@ intelligence = [
|
||||
transcribe = [
|
||||
"faster-whisper>=1.0.0",
|
||||
]
|
||||
align = [
|
||||
"whisperx>=3.0.0",
|
||||
]
|
||||
diarization = [
|
||||
"pyannote.audio>=3.1",
|
||||
]
|
||||
@@ -73,7 +76,7 @@ target-version = ['py310']
|
||||
|
||||
[tool.ruff]
|
||||
line-length = 100
|
||||
exclude = ["docs/", "WHISPERX/"]
|
||||
exclude = ["docs/"]
|
||||
|
||||
[tool.ruff.lint]
|
||||
select = ["E", "F", "I", "N", "W"]
|
||||
|
||||
@@ -154,6 +154,7 @@ from server_tools.roles import (
|
||||
)
|
||||
from server_tools.subtitles import (
|
||||
handle_generate_dynamic_subtitles,
|
||||
handle_generate_plain_subtitles,
|
||||
handle_validate_subtitle_layout,
|
||||
)
|
||||
from server_tools.timeline import (
|
||||
@@ -178,6 +179,7 @@ from server_tools.voice import (
|
||||
handle_analyze_voice_features,
|
||||
handle_apply_voice_actions,
|
||||
handle_build_voice_timeline,
|
||||
handle_generate_voice_script,
|
||||
handle_diarize_media,
|
||||
handle_get_voice_analysis_config,
|
||||
handle_refine_voice_timeline,
|
||||
@@ -320,6 +322,7 @@ __all__ = [
|
||||
"handle_save_voice_analysis_config",
|
||||
"handle_validate_subtitle_layout",
|
||||
"handle_generate_dynamic_subtitles",
|
||||
"handle_generate_plain_subtitles",
|
||||
"handle_push_to_fcp",
|
||||
"handle_list_fcp_libraries",
|
||||
]
|
||||
|
||||
@@ -1,829 +0,0 @@
|
||||
"""Shared internal helpers used by tool handlers across categories.
|
||||
|
||||
Extracted from server.py — validation, formatting, and small parsing utilities
|
||||
that more than one server_tools/*.py module needs.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import json
|
||||
import os
|
||||
import re
|
||||
from pathlib import Path
|
||||
from typing import Any, Sequence
|
||||
|
||||
from mcp.types import TextContent
|
||||
|
||||
from fcpxml.media_intel import media_src_to_path
|
||||
from fcpxml.models import (
|
||||
DuplicateGroup,
|
||||
FlashFrame,
|
||||
FlashFrameSeverity,
|
||||
GapInfo,
|
||||
Timecode,
|
||||
TimeValue,
|
||||
)
|
||||
from fcpxml.parser import FCPXMLParser
|
||||
from fcpxml.rough_cut import RoughCutGenerator
|
||||
from fcpxml.transcribe import invert_ranges, merge_ranges, transcribe
|
||||
from fcpxml.writer import FCPXMLModifier
|
||||
|
||||
PROJECTS_DIR = os.environ.get("FCP_PROJECTS_DIR", os.path.expanduser("~/Movies"))
|
||||
|
||||
_SANDBOX_ENABLED = "FCP_PROJECTS_DIR" in os.environ
|
||||
|
||||
MAX_FILE_SIZE = 100 * 1024 * 1024
|
||||
|
||||
MAX_MEDIA_FILE_SIZE = 32 * 1024 * 1024 * 1024
|
||||
|
||||
_MAX_JSON_DEPTH = 50
|
||||
|
||||
def _check_json_depth(obj: object, _depth: int = 0) -> None:
|
||||
"""Reject JSON structures nested beyond _MAX_JSON_DEPTH.
|
||||
|
||||
Prevents denial-of-service via deeply nested objects that exhaust the
|
||||
call stack or memory during downstream processing. Called after
|
||||
json.load() since Python's json module has no built-in depth limit.
|
||||
"""
|
||||
if _depth > _MAX_JSON_DEPTH:
|
||||
raise ValueError(
|
||||
f"JSON nesting depth exceeds {_MAX_JSON_DEPTH} — "
|
||||
"file may be malformed or adversarial"
|
||||
)
|
||||
if isinstance(obj, dict):
|
||||
for v in obj.values():
|
||||
_check_json_depth(v, _depth + 1)
|
||||
elif isinstance(obj, list):
|
||||
for item in obj:
|
||||
_check_json_depth(item, _depth + 1)
|
||||
|
||||
def _validate_filepath(
|
||||
filepath: str,
|
||||
allowed_extensions: tuple[str, ...] | None = None,
|
||||
max_size: int = MAX_FILE_SIZE,
|
||||
) -> str:
|
||||
"""Validate a user-provided file path against traversal and size attacks.
|
||||
|
||||
Resolves symlinks, blocks null bytes, enforces extension whitelist, and
|
||||
checks file size before any parsing takes place.
|
||||
|
||||
``max_size`` defaults to the document limit; callers handling source
|
||||
media pass ``MAX_MEDIA_FILE_SIZE``, since media is streamed rather than
|
||||
parsed into memory (see the constant for why).
|
||||
|
||||
Raises:
|
||||
ValueError: For invalid paths (null bytes, bad extensions, oversized).
|
||||
FileNotFoundError: When the resolved path does not exist.
|
||||
"""
|
||||
if '\x00' in filepath:
|
||||
raise ValueError("Invalid file path: null byte detected")
|
||||
|
||||
resolved = Path(filepath).resolve()
|
||||
|
||||
if not resolved.exists():
|
||||
raise FileNotFoundError(f"File not found: {filepath}")
|
||||
|
||||
# .fcpxmld bundles are directories (a package wrapping Info.fcpxml plus
|
||||
# sidecar data files for object tracking / Cinematic mode). The size
|
||||
# check applies to the inner Info.fcpxml, which is what gets parsed.
|
||||
if resolved.is_dir():
|
||||
if resolved.suffix.lower() != '.fcpxmld':
|
||||
raise ValueError(f"Not a regular file: {filepath}")
|
||||
inner = resolved / 'Info.fcpxml'
|
||||
if not inner.is_file():
|
||||
raise ValueError(f"Invalid bundle (no Info.fcpxml): {filepath}")
|
||||
size_target = inner
|
||||
elif not resolved.is_file():
|
||||
raise ValueError(f"Not a regular file: {filepath}")
|
||||
else:
|
||||
size_target = resolved
|
||||
|
||||
if allowed_extensions and resolved.suffix.lower() not in allowed_extensions:
|
||||
raise ValueError(
|
||||
f"Invalid file type '{resolved.suffix}'. "
|
||||
f"Allowed: {', '.join(allowed_extensions)}"
|
||||
)
|
||||
|
||||
if size_target.stat().st_size > max_size:
|
||||
size_mb = size_target.stat().st_size / (1024 * 1024)
|
||||
raise ValueError(f"File too large ({size_mb:.1f} MB). Maximum: {max_size // (1024 * 1024)} MB")
|
||||
|
||||
return str(resolved)
|
||||
|
||||
def _validate_output_path(output_path: str, *, anchor_dir: str | None = None) -> str:
|
||||
"""Validate an output path with optional sandbox enforcement.
|
||||
|
||||
Resolves traversal, blocks null bytes, ensures parent exists, and — when
|
||||
*anchor_dir* is provided — verifies the resolved output lives under that
|
||||
directory. This prevents LLM-generated tool calls from writing to
|
||||
arbitrary filesystem locations (e.g. ``/etc/cron.d/backdoor``).
|
||||
|
||||
Args:
|
||||
output_path: The raw output path to validate.
|
||||
anchor_dir: If set, the resolved output must be a child of this
|
||||
directory. Typically the parent directory of the input file so
|
||||
outputs stay co-located with their sources.
|
||||
|
||||
Raises:
|
||||
ValueError: For null bytes, missing parent, or sandbox escape.
|
||||
"""
|
||||
if '\x00' in output_path:
|
||||
raise ValueError("Invalid output path: null byte detected")
|
||||
|
||||
resolved = Path(output_path).resolve()
|
||||
|
||||
if not resolved.parent.exists():
|
||||
raise ValueError(f"Output directory does not exist: {resolved.parent}")
|
||||
|
||||
if anchor_dir is not None:
|
||||
anchor = Path(anchor_dir).resolve()
|
||||
try:
|
||||
resolved.relative_to(anchor)
|
||||
except ValueError:
|
||||
raise ValueError(
|
||||
f"Output path escapes allowed directory: "
|
||||
f"{resolved} is not under {anchor}"
|
||||
)
|
||||
|
||||
return str(resolved)
|
||||
|
||||
def _validate_directory(directory: str, *, allowed_root: str | None = None) -> str:
|
||||
"""Validate a user-provided directory path against traversal and injection.
|
||||
|
||||
Resolves symlinks, blocks null bytes, and verifies the path is a real
|
||||
directory. When *allowed_root* is given, the resolved path must be a
|
||||
descendant of (or equal to) that root — preventing filesystem enumeration
|
||||
beyond the project workspace.
|
||||
|
||||
Raises:
|
||||
ValueError: For invalid paths (null bytes, not a directory, sandbox escape).
|
||||
"""
|
||||
if '\x00' in directory:
|
||||
raise ValueError("Invalid directory path: null byte detected")
|
||||
|
||||
resolved = Path(directory).resolve()
|
||||
|
||||
if not resolved.is_dir():
|
||||
raise ValueError(f"Not a valid directory: {directory}")
|
||||
|
||||
if allowed_root is not None:
|
||||
root = Path(allowed_root).resolve()
|
||||
try:
|
||||
resolved.relative_to(root)
|
||||
except ValueError:
|
||||
raise ValueError(
|
||||
f"Directory escapes allowed root: "
|
||||
f"{resolved} is not under {root}"
|
||||
)
|
||||
|
||||
return str(resolved)
|
||||
|
||||
def find_fcpxml_files(directory: str) -> list[str]:
|
||||
"""Find all FCPXML files in a directory."""
|
||||
path = Path(directory)
|
||||
files = list(str(f) for f in path.rglob("*.fcpxml"))
|
||||
files.extend(str(f) for f in path.rglob("*.fcpxmld"))
|
||||
return sorted(files)
|
||||
|
||||
def format_timecode(tc) -> str:
|
||||
"""Format a Timecode object to SMPTE string."""
|
||||
return tc.to_smpte() if tc else "00:00:00:00"
|
||||
|
||||
def format_duration(seconds: float) -> str:
|
||||
"""Format seconds into human-readable duration."""
|
||||
if seconds < 1:
|
||||
return f"{seconds*1000:.0f}ms"
|
||||
elif seconds < 60:
|
||||
return f"{seconds:.2f}s"
|
||||
return f"{int(seconds // 60)}m {seconds % 60:.1f}s"
|
||||
|
||||
def _format_clip_table(clips: list, header: str) -> str:
|
||||
"""Render a list of clips as a markdown table with timecodes and durations.
|
||||
|
||||
Shared by handlers that filter clips by duration threshold
|
||||
(find_short_cuts, find_long_clips).
|
||||
"""
|
||||
result = f"{header}\n\n| Name | TC | Duration |\n|------|----|---------|\n"
|
||||
result += "\n".join(
|
||||
f"| {c.name} | {format_timecode(c.start)} | {format_duration(c.duration_seconds)} |"
|
||||
for c in clips
|
||||
)
|
||||
return result
|
||||
|
||||
def _markdown_table(headers: list[str], rows: list[list[str]]) -> str:
|
||||
"""Build a markdown table from headers and rows.
|
||||
|
||||
Returns header row, separator row, and data rows as a single string.
|
||||
Callers avoid repeating the ``| H1 | H2 |\\n|---|---|`` boilerplate
|
||||
that appears in 15+ handlers.
|
||||
"""
|
||||
header_line = "| " + " | ".join(headers) + " |"
|
||||
sep_line = "|" + "|".join("------" for _ in headers) + "|"
|
||||
data_lines = "\n".join(
|
||||
"| " + " | ".join(str(c) for c in row) + " |" for row in rows
|
||||
)
|
||||
return f"{header_line}\n{sep_line}\n{data_lines}"
|
||||
|
||||
def _format_batch_result(
|
||||
title: str,
|
||||
summary: dict[str, str],
|
||||
headers: list[str],
|
||||
rows: list[list[str]],
|
||||
output_path: str,
|
||||
) -> str:
|
||||
"""Build a standard batch-operation result with summary, table, and save footer.
|
||||
|
||||
Used by batch fix handlers (flash frames, rapid trim, fill gaps) that all
|
||||
share the same markdown structure: ``# Title → ## Summary → ## Details table
|
||||
→ Saved to`` footer.
|
||||
"""
|
||||
summary_lines = "\n".join(f"- **{k}**: {v}" for k, v in summary.items())
|
||||
table = _markdown_table(headers, rows)
|
||||
return (
|
||||
f"# {title}\n\n"
|
||||
f"## Summary\n{summary_lines}\n\n"
|
||||
f"## Details\n{table}\n\n"
|
||||
f"Saved to: `{output_path}`"
|
||||
)
|
||||
|
||||
def _fmt_suggestions(suggestions: list[str]) -> str:
|
||||
"""Format pacing suggestions as markdown list (Python 3.10 compatible)."""
|
||||
if not suggestions:
|
||||
return "- Pacing looks good!"
|
||||
nl = "\n"
|
||||
return nl.join(f"- {s}" for s in suggestions)
|
||||
|
||||
def generate_output_path(input_path: str, suffix: str = "_modified") -> str:
|
||||
"""Generate output path from input path.
|
||||
|
||||
The suffix is sanitized to prevent path-component injection — only
|
||||
alphanumeric, hyphen, underscore, and dot characters survive.
|
||||
"""
|
||||
# Strip anything that could inject path separators or traversal sequences
|
||||
clean_suffix = re.sub(r'[^a-zA-Z0-9._-]', '', suffix)
|
||||
if not clean_suffix:
|
||||
clean_suffix = "_modified"
|
||||
p = Path(input_path)
|
||||
return str(p.parent / f"{p.stem}{clean_suffix}{p.suffix}")
|
||||
|
||||
def _parse_project(filepath: str):
|
||||
"""Parse an FCPXML file and return the project with its primary timeline."""
|
||||
filepath = _validate_filepath(filepath, ('.fcpxml', '.fcpxmld'))
|
||||
project = FCPXMLParser().parse_file(filepath)
|
||||
if not project.timelines:
|
||||
return None, None
|
||||
return project, project.primary_timeline
|
||||
|
||||
def _text_result(text: str) -> list[TextContent]:
|
||||
"""Wrap a string in the MCP TextContent list that every tool handler returns."""
|
||||
return [TextContent(type="text", text=text)]
|
||||
|
||||
def _no_timeline():
|
||||
"""Standard response when no timelines are found."""
|
||||
return _text_result("No timelines found")
|
||||
|
||||
def _require_timeline(filepath: str):
|
||||
"""Parse FCPXML and return (project, timeline), raising if no timeline exists.
|
||||
|
||||
Centralises the repeated _parse_project + _no_timeline guard that
|
||||
appears in every read-only timeline handler. Returns a tuple so
|
||||
callers can destructure directly::
|
||||
|
||||
project, tl = _require_timeline(arguments["filepath"])
|
||||
"""
|
||||
project, tl = _parse_project(filepath)
|
||||
if not tl:
|
||||
raise _NoTimelineError()
|
||||
return project, tl
|
||||
|
||||
class _NoTimelineError(Exception):
|
||||
"""Sentinel raised by _require_timeline when no timelines exist."""
|
||||
|
||||
def _resolve_io_paths(
|
||||
arguments: dict,
|
||||
suffix: str = "_modified",
|
||||
) -> tuple[str, str]:
|
||||
"""Validate input filepath and resolve the output path.
|
||||
|
||||
Shared foundation for every handler that reads an FCPXML and writes
|
||||
a derived file. Validates the input, falls back to a suffixed
|
||||
output name when ``output_path`` is not supplied, and sandbox-checks
|
||||
the result.
|
||||
|
||||
Args:
|
||||
arguments: Tool arguments dict (must contain ``filepath``; may
|
||||
contain ``output_path``).
|
||||
suffix: Default output filename suffix when ``output_path`` is
|
||||
not provided (e.g. ``"_modified"``, ``"_beats"``).
|
||||
|
||||
Returns:
|
||||
``(filepath, output_path)`` tuple with both paths validated.
|
||||
"""
|
||||
filepath = _validate_filepath(arguments["filepath"], ('.fcpxml', '.fcpxmld'))
|
||||
# Anchor write operations to the input file's directory so LLM-generated
|
||||
# tool calls cannot write to arbitrary filesystem locations (e.g.
|
||||
# /etc/cron.d/backdoor). When the explicit sandbox is off, the anchor
|
||||
# still prevents writes outside the source directory tree.
|
||||
# `output_dir` is where the caller wants the file written, not merely a
|
||||
# sandbox boundary: the app's "Pasta do projeto" promises that everything
|
||||
# generated lands there. Deriving the name from the input but keeping the
|
||||
# input's directory made every cross-directory call fail its own anchor
|
||||
# check ("output path escapes allowed directory"), so the setting silently
|
||||
# only worked when it pointed at the directory the file was already going
|
||||
# to. An explicit `output_path` still wins, and still has to sit inside
|
||||
# the anchor.
|
||||
output_dir = arguments.get("output_dir")
|
||||
if output_dir:
|
||||
anchor = _validate_directory(str(output_dir))
|
||||
default_output = str(Path(anchor) / Path(generate_output_path(filepath, suffix)).name)
|
||||
else:
|
||||
anchor = str(Path(filepath).resolve().parent)
|
||||
default_output = generate_output_path(filepath, suffix)
|
||||
output_path = _validate_output_path(
|
||||
arguments.get("output_path") or default_output,
|
||||
anchor_dir=anchor,
|
||||
)
|
||||
return filepath, output_path
|
||||
|
||||
def _setup_modifier(
|
||||
arguments: dict,
|
||||
suffix: str = "_modified",
|
||||
) -> tuple[str, str, "FCPXMLModifier"]:
|
||||
"""Common setup for write handlers: validate paths and create modifier.
|
||||
|
||||
Consolidates the repeated validate-filepath → resolve-output-path →
|
||||
create-modifier boilerplate shared by 18+ write handlers.
|
||||
|
||||
Args:
|
||||
arguments: Tool arguments dict (must contain ``filepath``; may
|
||||
contain ``output_path``).
|
||||
suffix: Default output filename suffix when ``output_path`` is
|
||||
not provided (e.g. ``"_modified"``, ``"_flash_fixed"``).
|
||||
|
||||
Returns:
|
||||
``(filepath, output_path, modifier)`` tuple ready for the
|
||||
handler's domain-specific operation.
|
||||
"""
|
||||
filepath, output_path = _resolve_io_paths(arguments, suffix)
|
||||
modifier = FCPXMLModifier(filepath)
|
||||
return filepath, output_path, modifier
|
||||
|
||||
def _setup_generator(
|
||||
arguments: dict,
|
||||
suffix: str = "_roughcut",
|
||||
) -> tuple[str, str, "RoughCutGenerator"]:
|
||||
"""Common setup for generation handlers: validate paths and create generator.
|
||||
|
||||
Args:
|
||||
arguments: Tool arguments dict (must contain ``filepath`` and
|
||||
``output_path``).
|
||||
suffix: Default output filename suffix.
|
||||
|
||||
Returns:
|
||||
``(filepath, output_path, generator)`` tuple.
|
||||
"""
|
||||
filepath, output_path = _resolve_io_paths(arguments, suffix)
|
||||
generator = RoughCutGenerator(filepath)
|
||||
return filepath, output_path, generator
|
||||
|
||||
def _parse_timestamp_parts(
|
||||
parts: list[str], *, frame_rate: float = 24.0
|
||||
) -> float | None:
|
||||
"""Convert colon-separated timestamp parts to total seconds.
|
||||
|
||||
Handles 2-part (M:SS), 3-part (H:MM:SS / HH:MM:SS.ms), and
|
||||
4-part (HH:MM:SS:FF SMPTE) formats. Returns ``None`` when the
|
||||
part count is unrecognised so callers can skip.
|
||||
|
||||
Args:
|
||||
parts: Colon-split timestamp components.
|
||||
frame_rate: FPS used to convert the frame component of SMPTE
|
||||
timecodes into fractional seconds (default 24.0).
|
||||
"""
|
||||
if len(parts) == 2:
|
||||
return int(parts[0]) * 60 + float(parts[1])
|
||||
elif len(parts) == 3:
|
||||
return int(parts[0]) * 3600 + int(parts[1]) * 60 + float(parts[2])
|
||||
elif len(parts) == 4:
|
||||
# SMPTE: HH:MM:SS:FF — convert frames to fractional seconds
|
||||
base = int(parts[0]) * 3600 + int(parts[1]) * 60 + float(parts[2])
|
||||
frames = int(parts[3])
|
||||
return base + (frames / frame_rate) if frame_rate > 0 else base
|
||||
return None
|
||||
|
||||
def _raw_markers_to_batch(
|
||||
raw_markers: list[dict],
|
||||
marker_type: str = "chapter",
|
||||
max_label: int | None = None,
|
||||
) -> list[dict]:
|
||||
"""Convert raw {seconds, text} marker dicts to batch_add_markers format.
|
||||
|
||||
Shared by import_srt_markers and import_transcript_markers.
|
||||
"""
|
||||
batch = []
|
||||
for m in raw_markers:
|
||||
label = m["text"]
|
||||
if max_label and len(label) > max_label:
|
||||
label = label[:max_label]
|
||||
batch.append({
|
||||
"timecode": f"{m['seconds']}s",
|
||||
"name": label,
|
||||
"marker_type": marker_type.upper(),
|
||||
})
|
||||
return batch
|
||||
|
||||
def _extract_subtitle_blocks(text: str, *, strip_vtt_tags: bool = False) -> list[dict]:
|
||||
"""Extract timestamp/text pairs from subtitle cue blocks (SRT or VTT).
|
||||
|
||||
Both SRT and VTT use the same ``start --> end`` cue syntax with
|
||||
text lines underneath; only header stripping and tag cleaning differ.
|
||||
"""
|
||||
markers = []
|
||||
blocks = re.split(r'\n\s*\n', text.strip())
|
||||
for block in blocks:
|
||||
lines = block.strip().split('\n')
|
||||
if len(lines) < 2:
|
||||
continue
|
||||
ts_line = None
|
||||
text_lines = []
|
||||
for line in lines:
|
||||
if '-->' in line:
|
||||
ts_line = line
|
||||
elif ts_line is not None:
|
||||
if strip_vtt_tags:
|
||||
line = re.sub(r'<[^>]+>', '', line)
|
||||
cleaned = line.strip()
|
||||
if cleaned:
|
||||
text_lines.append(cleaned)
|
||||
if not ts_line or not text_lines:
|
||||
continue
|
||||
start_str = ts_line.split('-->')[0].strip().replace(',', '.')
|
||||
seconds = _parse_timestamp_parts(start_str.split(':'))
|
||||
if seconds is not None:
|
||||
markers.append({'seconds': seconds, 'text': ' '.join(text_lines)})
|
||||
return markers
|
||||
|
||||
def parse_srt(text: str) -> list[dict]:
|
||||
"""Parse SRT subtitle format into timestamp/text pairs."""
|
||||
return _extract_subtitle_blocks(text)
|
||||
|
||||
def parse_vtt(text: str) -> list[dict]:
|
||||
"""Parse WebVTT subtitle format into timestamp/text pairs."""
|
||||
text = re.sub(r'^WEBVTT.*?\n', '', text, flags=re.MULTILINE)
|
||||
text = re.sub(r'NOTE\n.*?\n\n', '', text, flags=re.DOTALL)
|
||||
return _extract_subtitle_blocks(text, strip_vtt_tags=True)
|
||||
|
||||
def parse_transcript_timestamps(text: str) -> list[dict]:
|
||||
"""Parse timestamped text (YouTube description format) into markers.
|
||||
|
||||
Supports formats like:
|
||||
0:00 Introduction
|
||||
00:01:30 Main Topic
|
||||
1:05:30 Conclusion
|
||||
00:00:00:00 SMPTE timecode
|
||||
"""
|
||||
markers = []
|
||||
for line in text.strip().split('\n'):
|
||||
line = line.strip()
|
||||
if not line:
|
||||
continue
|
||||
match = re.match(r'^(\d{1,2}:\d{2}(?::\d{2}){0,2})\s+(.+)$', line)
|
||||
if match:
|
||||
seconds = _parse_timestamp_parts(match.group(1).split(':'))
|
||||
if seconds is not None:
|
||||
markers.append({'seconds': seconds, 'text': match.group(2).strip()})
|
||||
return markers
|
||||
|
||||
def _detect_flash_frames(
|
||||
tl: Any, *, critical_threshold: int = 2, warning_threshold: int = 6,
|
||||
) -> list:
|
||||
"""Find clips shorter than *warning_threshold* frames.
|
||||
|
||||
Returns a list of ``FlashFrame`` objects sorted by severity. Shared by
|
||||
``handle_detect_flash_frames`` and ``handle_validate_timeline`` so the
|
||||
detection logic lives in exactly one place.
|
||||
"""
|
||||
fps = tl.frame_rate
|
||||
flash_frames: list[FlashFrame] = []
|
||||
for clip in tl.clips:
|
||||
duration_frames = int(clip.duration_seconds * fps)
|
||||
if duration_frames < warning_threshold:
|
||||
severity = (
|
||||
FlashFrameSeverity.CRITICAL
|
||||
if duration_frames < critical_threshold
|
||||
else FlashFrameSeverity.WARNING
|
||||
)
|
||||
flash_frames.append(FlashFrame(
|
||||
clip_name=clip.name, clip_id=clip.name,
|
||||
start=clip.start, duration_frames=duration_frames,
|
||||
duration_seconds=clip.duration_seconds, severity=severity,
|
||||
))
|
||||
return flash_frames
|
||||
|
||||
def _detect_gaps(tl: Any, *, min_gap_frames: int = 1) -> list:
|
||||
"""Find inter-clip gaps of at least *min_gap_frames* length.
|
||||
|
||||
Returns a list of ``GapInfo`` objects. Shared by ``handle_detect_gaps``
|
||||
and ``handle_validate_timeline``.
|
||||
"""
|
||||
fps = tl.frame_rate
|
||||
min_gap_seconds = min_gap_frames / fps
|
||||
gaps: list[GapInfo] = []
|
||||
sorted_clips = sorted(tl.clips, key=lambda c: c.start.seconds)
|
||||
for i in range(len(sorted_clips) - 1):
|
||||
current_end = sorted_clips[i].end.seconds
|
||||
next_start = sorted_clips[i + 1].start.seconds
|
||||
gap_duration = next_start - current_end
|
||||
if gap_duration >= min_gap_seconds:
|
||||
gaps.append(GapInfo(
|
||||
start=Timecode(frames=int(current_end * fps), frame_rate=fps),
|
||||
duration_frames=int(gap_duration * fps),
|
||||
duration_seconds=gap_duration,
|
||||
previous_clip=sorted_clips[i].name,
|
||||
next_clip=sorted_clips[i + 1].name,
|
||||
))
|
||||
return gaps
|
||||
|
||||
def _detect_duplicate_groups(tl: Any, *, mode: str = "same_source") -> list:
|
||||
"""Group clips that share a source media reference.
|
||||
|
||||
Returns a list of ``DuplicateGroup`` objects. Shared by
|
||||
``handle_detect_duplicates`` and ``handle_validate_timeline``.
|
||||
"""
|
||||
source_groups: dict[str, list[dict]] = {}
|
||||
for clip in tl.clips:
|
||||
source_key = clip.media_path or clip.name
|
||||
if source_key not in source_groups:
|
||||
source_groups[source_key] = []
|
||||
source_groups[source_key].append({
|
||||
'name': clip.name,
|
||||
'start': clip.start.seconds,
|
||||
'duration': clip.duration_seconds,
|
||||
'source_start': clip.source_start.seconds if clip.source_start else 0,
|
||||
'source_duration': clip.duration_seconds,
|
||||
'timecode': format_timecode(clip.start),
|
||||
})
|
||||
|
||||
duplicates: list[DuplicateGroup] = []
|
||||
for source_key, clips in source_groups.items():
|
||||
if len(clips) <= 1:
|
||||
continue
|
||||
group = DuplicateGroup(
|
||||
source_ref=source_key,
|
||||
source_name=source_key.split('/')[-1] if '/' in source_key else source_key,
|
||||
clips=clips,
|
||||
)
|
||||
if mode == "same_source":
|
||||
duplicates.append(group)
|
||||
elif mode == "overlapping_ranges" and group.has_overlapping_ranges:
|
||||
duplicates.append(group)
|
||||
elif mode == "identical":
|
||||
seen_ranges: set[tuple] = set()
|
||||
identical_clips = []
|
||||
for c in clips:
|
||||
range_key = (c['source_start'], c['source_duration'])
|
||||
if range_key in seen_ranges:
|
||||
identical_clips.append(c)
|
||||
seen_ranges.add(range_key)
|
||||
if identical_clips:
|
||||
group.clips = identical_clips
|
||||
duplicates.append(group)
|
||||
return duplicates
|
||||
|
||||
AUDIO_MEDIA_EXTENSIONS = (
|
||||
'.wav', '.aif', '.aiff', '.mp3', '.m4a', '.aac', '.flac', '.mov', '.mp4',
|
||||
)
|
||||
|
||||
_DIARIZATION_INSTALL_HINT = (
|
||||
"\n\nInstall the optional diarization extra:\n\n"
|
||||
" pip install 'fcp-mcp-server[diarization]'\n\n"
|
||||
"and set a HuggingFace token with access to "
|
||||
"pyannote/speaker-diarization-3.1 (pass hf_token= or persist one via "
|
||||
"save_hf_token)."
|
||||
)
|
||||
|
||||
_FEATURES_INSTALL_HINT = (
|
||||
"\n\nInstall the optional media-intelligence extra:\n\n"
|
||||
" pip install 'fcp-mcp-server[intelligence]'"
|
||||
)
|
||||
|
||||
def _voice_analysis_config_text(config: dict) -> str:
|
||||
w = config["emphasis_weights"]
|
||||
text = "# Voice Analysis Settings\n\n"
|
||||
text += _markdown_table(
|
||||
["Setting", "Value"],
|
||||
[
|
||||
["Energy threshold", f"{config['energy_threshold']:.2f}"],
|
||||
["Peak selection", f"top {config['peak_percentile']:.1%} of words"],
|
||||
["Emphasis floor", f"{config['emphasis_floor']:.2f}"],
|
||||
["Emotion detection", "on" if config["emotion_enabled"] else "off"],
|
||||
["Emotion sensitivity", f"{config['emotion_sensitivity']:.2f}"],
|
||||
],
|
||||
) + "\n\n## Emphasis Weights\n"
|
||||
text += _markdown_table(
|
||||
["Factor", "Weight"],
|
||||
[[k.replace("_", " ").title(), f"{v:.2f}"] for k, v in w.items()],
|
||||
)
|
||||
return text
|
||||
|
||||
def _apply_placed_action(modifier, clip_el, action, clip_start: float) -> str:
|
||||
"""Apply one non-cut action to the clip that hosts it.
|
||||
|
||||
``clip_start`` is where that clip begins on the timeline; the writer
|
||||
wants times relative to the clip's own head, so the rebase happens here
|
||||
— the single place that knows about the conversion. The clip *element*
|
||||
is passed through rather than its name: after a cut the pieces share a
|
||||
name, and a name lookup would land every edit on the first piece.
|
||||
"""
|
||||
rel_start = action.start - clip_start
|
||||
rel_end = action.end - clip_start
|
||||
|
||||
if action.kind == "zoom":
|
||||
# Only forward an explicit ease — otherwise add_zoom's own default
|
||||
# (a fast ramp in, instant snap back out) is what should apply.
|
||||
zoom_args = {}
|
||||
if action.params.get("ease") is not None:
|
||||
zoom_args["ease"] = float(action.params["ease"])
|
||||
if action.params.get("ease_out") is not None:
|
||||
zoom_args["ease_out"] = float(action.params["ease_out"])
|
||||
modifier.add_zoom(
|
||||
clip_id=clip_el,
|
||||
start=rel_start,
|
||||
end=rel_end,
|
||||
scale=float(action.params.get("scale", 1.3)),
|
||||
**zoom_args,
|
||||
)
|
||||
return f"zoom {action.params.get('scale', 1.3):.2f}x"
|
||||
|
||||
if action.kind == "text":
|
||||
modifier.add_text_title(
|
||||
clip_el,
|
||||
action.params["content"],
|
||||
offset=modifier.snap_seconds_to_frame(rel_start).to_fcpxml(),
|
||||
duration=modifier.snap_seconds_to_frame(action.duration).to_fcpxml(),
|
||||
)
|
||||
return f"text \"{action.params['content'][:24]}\""
|
||||
|
||||
# marker
|
||||
modifier.add_marker(
|
||||
clip_id=clip_el,
|
||||
timecode=modifier.snap_seconds_to_frame(rel_start).to_fcpxml(),
|
||||
name=action.params.get("content") or action.reason or "Voice action",
|
||||
note=action.reason or None,
|
||||
)
|
||||
return "marker"
|
||||
|
||||
def _speaker_table(profiles: Sequence[dict]) -> str:
|
||||
"""Who was detected, ordered by how much of the runtime each holds."""
|
||||
return _markdown_table(
|
||||
["ID", "Name", "Share", "Speaking", "Lines", "Avg line"],
|
||||
[
|
||||
[
|
||||
p["id"],
|
||||
p.get("name", ""),
|
||||
f"{p['share']:.0%}",
|
||||
format_duration(p["speaking_seconds"]),
|
||||
str(p["segment_count"]),
|
||||
f"{p['avg_segment']:.1f}s",
|
||||
]
|
||||
for p in profiles
|
||||
],
|
||||
)
|
||||
|
||||
TRANSCRIBE_MAX_MEDIA = 10
|
||||
|
||||
_TRANSCRIBE_INSTALL_HINT = (
|
||||
"\n\nInstall the optional transcription extra:\n\n"
|
||||
" pip install 'fcp-mcp-server[transcribe]'\n\n"
|
||||
"or run via uvx:\n\n"
|
||||
" uvx --from \"fcp-mcp-server[transcribe]\" fcp-mcp-server"
|
||||
)
|
||||
|
||||
def _transcript_json_path(media_path: str, output_dir: str | None = None) -> Path:
|
||||
"""Where the ``_transcript.json`` for ``media_path`` lives.
|
||||
|
||||
When ``output_dir`` (the user-selected project folder) is set, the
|
||||
transcript is saved/read there instead of next to the source media.
|
||||
"""
|
||||
p = Path(media_path)
|
||||
if output_dir:
|
||||
directory = Path(output_dir).expanduser()
|
||||
directory.mkdir(parents=True, exist_ok=True)
|
||||
return directory / f"{p.stem}_transcript.json"
|
||||
return p.with_name(p.stem + "_transcript.json")
|
||||
|
||||
def _load_or_transcribe(
|
||||
media_path: str, model: str, language: str | None, output_dir: str | None = None
|
||||
) -> tuple[dict | None, str]:
|
||||
"""Load a cached ``_transcript.json`` for a media file, else transcribe and cache it.
|
||||
|
||||
Returns ``(transcript, "")`` or ``(None, reason)``. The cache makes
|
||||
transcription a one-time cost per media file across all transcript tools.
|
||||
"""
|
||||
json_path = _transcript_json_path(media_path, output_dir)
|
||||
if json_path.is_file():
|
||||
try:
|
||||
with open(json_path) as f:
|
||||
data = json.load(f)
|
||||
if isinstance(data, dict) and isinstance(data.get("words"), list):
|
||||
return data, ""
|
||||
except (OSError, json.JSONDecodeError, UnicodeDecodeError):
|
||||
pass # unreadable cache falls through to re-transcribe
|
||||
result = transcribe(media_path, model_size=model, language=language)
|
||||
if result is None:
|
||||
return None, "untranscribable (faster-whisper not installed or media unreadable)"
|
||||
anchor = str(Path(output_dir).expanduser()) if output_dir else str(Path(media_path).parent)
|
||||
out_path = _validate_output_path(str(json_path), anchor_dir=anchor)
|
||||
with open(out_path, "w") as f:
|
||||
json.dump({"source": Path(media_path).name, **result}, f, indent=2)
|
||||
return result, ""
|
||||
|
||||
def _cut_transcript_spans(modifier, clip_filter, model, language, padding, spans_fn, keep_only=False, output_dir=None):
|
||||
"""Shared cut engine for transcript-driven editing.
|
||||
|
||||
``spans_fn(words) -> [(start, end), ...]`` in source seconds. Spans are
|
||||
padded, clamped to each clip's used source window, optionally inverted
|
||||
(keep_only), snapped to the frame grid, and cut with ripple.
|
||||
"""
|
||||
to_frame = modifier.snap_seconds_to_frame
|
||||
|
||||
cache: dict[str, tuple] = {}
|
||||
cuts_made: list[tuple[str, int, float]] = []
|
||||
skipped: list[tuple[str, str]] = []
|
||||
spine_clips = [el for _, el in modifier._iter_spine_clips()]
|
||||
for el in spine_clips:
|
||||
name = el.get("name", "")
|
||||
if clip_filter and name != clip_filter:
|
||||
continue
|
||||
src = modifier.resources.get(el.get("ref", ""), {}).get("src", "")
|
||||
media_path = media_src_to_path(src)
|
||||
if not media_path or not Path(media_path).is_file():
|
||||
skipped.append((name, "media file missing"))
|
||||
continue
|
||||
if media_path not in cache:
|
||||
if len(cache) >= TRANSCRIBE_MAX_MEDIA:
|
||||
skipped.append((name, f"transcription cap reached ({TRANSCRIBE_MAX_MEDIA} media files)"))
|
||||
continue
|
||||
cache[media_path] = _load_or_transcribe(media_path, model, language, output_dir)
|
||||
data, reason = cache[media_path]
|
||||
if data is None:
|
||||
skipped.append((name, reason))
|
||||
continue
|
||||
|
||||
clip_source_start = modifier.source_file_start(el).to_seconds()
|
||||
clip_duration = modifier._parse_time(el.get("duration", "0s")).to_seconds()
|
||||
window_start = clip_source_start
|
||||
window_end = clip_source_start + clip_duration
|
||||
|
||||
spans = spans_fn(data.get("words", []))
|
||||
padded = merge_ranges([(s - padding, e + padding) for s, e in spans])
|
||||
clamped = [
|
||||
(max(s, window_start), min(e, window_end))
|
||||
for s, e in padded
|
||||
if min(e, window_end) > max(s, window_start)
|
||||
]
|
||||
if keep_only:
|
||||
if not clamped:
|
||||
# Never delete a whole clip just because nothing matched in it.
|
||||
skipped.append((name, "no phrase matches — left untouched (keep_only)"))
|
||||
continue
|
||||
cut_source = invert_ranges(clamped, window_start, window_end)
|
||||
else:
|
||||
cut_source = clamped
|
||||
cut_ranges = [
|
||||
(to_frame(s - clip_source_start), to_frame(e - clip_source_start))
|
||||
for s, e in cut_source
|
||||
]
|
||||
cut_ranges = [(a, b) for a, b in cut_ranges if b > a]
|
||||
if not cut_ranges:
|
||||
continue
|
||||
removed = modifier.cut_clip_ranges(el, cut_ranges)
|
||||
if removed > TimeValue.zero():
|
||||
cuts_made.append((name, len(cut_ranges), removed.to_seconds()))
|
||||
return cuts_made, skipped
|
||||
|
||||
def _transcript_cut_report(title, summary_lines, cuts_made, skipped, output_path, footer):
|
||||
if not cuts_made:
|
||||
text = f"# {title}\n\nNo cuts to make — file unchanged (nothing saved)."
|
||||
if skipped:
|
||||
text += "\n\n## Skipped Clips\n" + _markdown_table(
|
||||
["Clip", "Reason"], [[name, reason] for name, reason in skipped]
|
||||
)
|
||||
if any("faster-whisper" in reason for _, reason in skipped):
|
||||
text += _TRANSCRIBE_INSTALL_HINT
|
||||
return _text_result(text)
|
||||
total_removed = sum(seconds for _, _, seconds in cuts_made)
|
||||
result = f"# {title}\n\n## Summary\n"
|
||||
result += "\n".join(summary_lines) + "\n"
|
||||
result += f"- **Clips Cut**: {len(cuts_made)}\n- **Total Removed**: {format_duration(total_removed)}\n"
|
||||
result += "\n## Cuts\n"
|
||||
result += _markdown_table(
|
||||
["Clip", "Ranges Cut", "Removed"],
|
||||
[[name, str(count), f"{seconds:.2f}s"] for name, count, seconds in cuts_made],
|
||||
) + "\n"
|
||||
if skipped:
|
||||
result += "\n## Skipped Clips\n" + _markdown_table(
|
||||
["Clip", "Reason"], [[name, reason] for name, reason in skipped]
|
||||
) + "\n"
|
||||
result += f"\nSaved to: {output_path}\n\n{footer}"
|
||||
return _text_result(result)
|
||||
Some files were not shown because too many files have changed in this diff Show More
Reference in New Issue
Block a user