Compare commits

..
Author SHA1 Message Date
João HenriqueandClaude Sonnet 5 e9a17c1b62 feat(rag): provisiona busca RAG do G-ART e corrige indexação que abortava em chunk grande
Cria rag/ (schema, busca híbrida densa+lexical com RRF em search.py/
search_gart.sh, SETUP.md) — o projeto já tinha admin/update_rag.py para
indexar, mas nenhuma forma de consultar o índice. Corrige admin/update_rag.py:
um chunk denso em tokens (code/fcpxml/font_metrics.py) estourava o contexto
do modelo de embedding e derrubava a transação inteira; agora só aquele
chunk é pulado. Banco rag_gart provisionado no rag-hub-db compartilhado e
primeira indexação completa rodada (304 arquivos, 1702 chunks).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-23 09:38:46 -04:00
João HenriqueandClaude Sonnet 5 d13f643ebc chore(fase0): higiene do repositório + corrige gitignore que escondia fcpxml/models/
Fase 0 do roteiro de reestruturação (Engine/docs/10_MAPA_REESTRUTURACAO.md):
move code/WHISPERX (2,6 GB de backups órfãos, sem uso ativo, sem
.gitmodules) para ~/Archives/G-ART-WHISPERX-backup fora do workspace git;
traz admin/ para o gate de lint de run_after_fix.sh; corrige
fcpxml/writer/adjustment.py, que gerava um wrapper <adjustment> inexistente
no DTD 1.13 (filtros agora vão direto no <clip>, na ordem exigida), com
teste de regressão novo.

Achado à parte: .gitignore tinha uma regra solta "models/" (pensada só
para o cache do Whisper em code/models/) que também escondia do git todo o
pacote fcpxml/models/ — nunca commitado, sem proteção nenhuma. Corrigida
para /code/models/, ancorada na raiz.

Docs atualizados no mesmo commit (02_MODULES, 09_MANUTENCAO,
10_MAPA_REESTRUTURACAO, 05_EXPERIENCIAS #34 e #36), conforme a regra do
CLAUDE.md.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-23 08:28:44 -04:00
João HenriqueandClaude Sonnet 5 0fdfe33613 fix(legendas): impede legenda comum sob composição dinâmica e duplicação ao regerar
Marca cada título gerado (dynamic/plain) em metadata para que regenerar
substitua a saída anterior em vez de empilhar, e usa os spans de ênfase
revisados (não os segmentos brutos do Whisper) como janela da composição
dinâmica, evitando que ela invada o trecho de legenda comum seguinte.
suppress_plain_under_dynamic corta qualquer sobra visível como rede de
segurança.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-22 21:51:02 -04:00
João HenriqueandClaude Sonnet 5 32d78d0f8d fix(qc): padding padrão do remove_media_silence de 0,05s para 0,2s
0,05s existia como margem de segurança contra cortar a palavra em cima,
mas um silêncio que essa ferramenta encontra costuma ser o respiro
natural antes de uma frase nova, não sujeira de edição — e 0,05s raspava
esse respiro quase todo.

Caso real (projeto Mastopexia): a pausa antes de "Com" tinha 0,567s no
áudio original; com padding 0,05 sobrou só ~0,1s no total (0,05 de cada
lado), colando o clipe seguinte a 5ms da palavra em vez de deixar uma
pausa perceptível. 0,2s alinha com a convenção já documentada para folga
em corte de fronteira de frase (editar-por-voz/06-texto-corte-marcador.md).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-26 17:15:44 -04:00
João HenriqueandClaude Sonnet 5 c99274895c feat(legendas): liga compound_subphrases por padrão no pipeline
generate_dynamic_subtitles e a metade dinâmica de generate_subtitles_by_emphasis
passam a empacotar cada sub-frase da legenda dinâmica num compound clip por
padrão (compound_subphrases=True), completando o wrap_titles_in_compound
e split_into_subphrases do commit anterior — que ainda não tinham chamador
em produção.

Também torna validate_subtitle_layout ciente de compound clips: media cada
grupo (spine principal + cada <media> de compound) no seu próprio espaço de
tempo, em vez de uma varredura .//title global — sem isso, âncoras de
compounds diferentes liam offset "0s" e acusavam colisão espacial entre
frases que nunca dividem a tela, só porque compartilham o mesmo zero de
tempo local.

Testado ponta a ponta na gravação real (Mastopexia): 12 compounds, 41
títulos todos empacotados, zero soltos, zero IDs duplicados, DTD válida.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-26 17:05:20 -04:00
João HenriqueandClaude Opus 5 688bdeddb6 feat(legendas): sub-frases por vírgula e empacotamento em compound clip
split_into_subphrases divide a frase na vírgula — onde a fala respira —
mas funde de volta o pedaço curto ("né?", "Então..."), que lê como parte
da frase anterior e não como bloco próprio.

wrap_titles_in_compound empacota os títulos de uma sub-frase num compound
clip, replicando a estrutura que o próprio Final Cut produz: o primeiro
título vira âncora do spine em offset 0, os demais penduram nele por lane,
e um ref-clip toma o lugar deles na lane original. Os offsets dos filhos
são rebaseados para o espaço de tempo da âncora, senão cada palavra
escorregaria pela diferença entre os dois start.

Junto: _filter_children_for_segment passa a filtrar também o <video> do
Clipe de Ajuste. Sem isso, cada corte subsequente duplicava o zoom em
todos os pedaços resultantes com o offset original intacto, e as cópias
desenhavam empilhadas na mesma posição da timeline.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-26 16:35:21 -04:00
João HenriqueandClaude Sonnet 5 635d1bb553 fix(fcpxml): tracking-shape id duplicado ao cortar clipe com Cinematic
Object-tracker/tracking-shape (dado de rastreamento de objeto preservado
do asset original) mantinha o mesmo id em cada deepcopy feito por
split_clip/cut_clip_ranges, e o FCP acabava rejeitando o arquivo com "ID
tr1 already defined" depois de vários cortes. Mesmo mecanismo do bug já
corrigido para text-style-def, agora coberto também para tracking-shape.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-26 16:03:31 -04:00
João HenriqueandClaude Sonnet 5 2ad5854570 fix(voz): sobra de fatia interior no corte, e legenda comum poluindo com clipe desativado
Dois problemas reais vistos no projeto Mastopexia:

1. cut_clip_ranges só absorvia um keep-segment curto no INÍCIO/FIM do
   clipe (a lógica já existente do #6). Um keep curto no MEIO (entre dois
   cuts, sem nenhum vizinho mantido pra herdar) nunca era absorvido —
   sobrava como clipe de vídeo de 0,07-0,23s na timeline. Generalizado
   pra qualquer posição, com limiar maior (6 frames / 0,3s, medido no
   material real) — no meio, o pedacinho é descartado (vira parte do
   corte ao redor), nas bordas continua sendo herdado pelo vizinho.

2. generate_subtitles_by_emphasis gerava a legenda comum inteira e
   desativava (enabled="0") onde a dinâmica cobre. Título desativado
   continua aparecendo como clipe riscado na timeline do Final Cut mesmo
   sem renderizar — um corte com bastante ênfase virava dezenas de clipes
   mortos poluindo a trilha (visto ao vivo pelo usuário: "ficou uma
   bosta"). Trocado por não gerar o bloco comum ali, em vez de gerar e
   desativar. Custo: reativar ênfase manualmente depois exige regenerar a
   legenda comum daquele trecho, não só reabilitar.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-21 18:59:21 -04:00
João HenriqueandClaude Sonnet 5 8257155fd3 fix(voz): frases desativadas em sequência deixavam fatias sobrando no corte
phrase_review_to_actions() cortava cada frase desativada isoladamente
(start..end da própria frase) — quando várias seguidas estavam desativadas,
a pausa ENTRE elas não pertencia a nenhuma frase e sobrevivia como um
clipe minúsculo (0,1-0,5s) na timeline final. Confirmado no projeto
Mastopexia real: 29 cuts individuais geravam mais de uma dezena de fatias
sub-segundo; agrupar frases desativadas consecutivas num único cut (do
início da primeira ao fim da última) reduziu para 3 cuts e 4 fatias
residuais (menores, provavelmente do padding do remove_media_silence —
registrado como dívida separada em 09_MANUTENCAO.md §2.5).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-21 18:49:13 -04:00
João HenriqueandClaude Sonnet 5 fd791e116a docs(skill): corte deve deixar folga na borda que encosta em fala mantida
Cortes escritos rente ao timestamp da palavra soavam secos (relatado no
projeto Mastopexia) — o critério e o prompt do modelo local mandavam cobrir
a frase inteira sem orientar a borda que toca fala mantida. Adiciona a regra
de recuar ~0,15-0,25s nas duas pontas quando o corte encosta em conteúdo
que fica, tanto no skill (06-texto-corte-marcador.md) quanto no prompt
embutido do Ollama (llm_local.py) — pra não precisar ajustar na mão de novo.

Também registra em 05_EXPERIENCIAS.md/09_MANUTENCAO.md a dívida de
resolve_actions não tolerar margem quando zoom/marker encosta na borda
de um corte (contornado manualmente, não corrigido em código ainda), e
atualiza a lista de dívidas abertas (etapa 6/offset de whisper já resolvidos
nesta sessão).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-21 18:31:27 -04:00
45 changed files with 4409 additions and 216 deletions
@@ -53,6 +53,34 @@ Acima de 3s a pausa deixa de contar como ênfase por construção — medido em
material real, lacunas de 6–9s rankeavam como os momentos mais enfáticos da material real, lacunas de 6–9s rankeavam como os momentos mais enfáticos da
gravação só porque a escala saturava. gravação só porque a escala saturava.
### Nunca corte rente à palavra — deixe uma folga
Um `cut` cujo `start`/`end` cai exatamente no timestamp da palavra (fim da
última palavra mantida = início do corte) produz um corte seco: a palavra é
engolida antes de terminar de soar, e a fala seguinte começa sem nenhum ar.
Isso é diferente de cortar a pausa curta (que seria apagar a própria ênfase,
proibido acima) — aqui a pausa **já existe** entre o fim de um bloco mantido
e o início do próximo, e o corte está comendo justamente essa margem.
Ao escrever a borda de um `cut` que encosta em fala mantida (não em silêncio
puro), recue **~0,15–0,25s** para dentro do próprio corte, nos dois lados:
- o `start` do corte fica ~0,2s **depois** do fim real da última palavra
mantida;
- o `end` do corte fica ~0,2s **antes** do início real da próxima palavra
mantida.
Caso real (projeto Mastopexia): um corte escrito rente (`10.77 → 95.50`,
exatamente nos timestamps de palavra) soava abrupto nas duas emendas.
Recuado para `10.97 → 95.30`, cada lado ganhou ~0,2s de respiro sem alterar
o que é dito — e não empurra o próximo zoom/marcador contra a borda do corte
(ver `05-zoom.md` sobre janelas encostadas em corte).
Isso vale também para o **início e o fim do vídeo**: ar morto antes da
primeira palavra e depois da última também leva `cut`, com a mesma folga —
não é "silêncio dentro da fala" (isso é `remove_media_silence`), é o mesmo
corte de tomada/bastidor que você já está decidindo.
### O que continua NÃO sendo seu trabalho ### O que continua NÃO sendo seu trabalho
| Tarefa | Ferramenta | Por quê | | Tarefa | Ferramenta | Por quê |
+6 -2
View File
@@ -30,6 +30,7 @@ Thumbs.db
# Env files (NUNCA commitar — contêm segredos) # Env files (NUNCA commitar — contêm segredos)
*.env *.env
.env .env
admin/gart-rag.env
# Graphify output (gerado, não rastrear) # Graphify output (gerado, não rastrear)
graphify-out/ graphify-out/
@@ -37,6 +38,9 @@ graphify-out/
# FCPXML bundles de exemplo (podem ser grandes) # FCPXML bundles de exemplo (podem ser grandes)
*.fcpxmld/ *.fcpxmld/
# WhisperX models cache # Cache de modelos Whisper baixados (código/models, ~11 GB, HuggingFace hub
models/ # format). Âncora em /code/models/ — NUNCA "models/" solto: isso também
# ignorava fcpxml/models/, o pacote de dados do engine (ver
# Engine/docs/05_EXPERIENCIAS.md #36).
/code/models/
whisper/ whisper/
+12
View File
@@ -0,0 +1,12 @@
# Copie para admin/gart-rag.env e preencha a senha. Este arquivo é apenas um
# modelo; admin/gart-rag.env é ignorado pelo git.
RAG_DB_HOST=127.0.0.1
RAG_DB_PORT=55435
RAG_DB_NAME=rag_gart
RAG_DB_SCHEMA=gart
RAG_DB_USER=gart_rag_indexer
RAG_DB_PASSWORD=
# Ollama que fornece nomic-embed-text.
OLLAMA_URL=http://127.0.0.1:11434
RAG_EMBED_MODEL=nomic-embed-text
+80 -5
View File
@@ -68,6 +68,39 @@ Commands:
-> {"ok": true, "review_path", "actions_path", "emphasis_count", -> {"ok": true, "review_path", "actions_path", "emphasis_count",
"removed_count"} "removed_count"}
build_speaker_review {"voice_timeline": "..._voice_timeline.json", "fresh": false}
Runs right after `analyze_voice` (wizard step 3): who was detected
(with speaking share and sample lines, default active/kept) plus one
row per transcript segment, for a naming + mute + strike-line screen
before anything reaches the AI. A review saved earlier is merged
back on top unless `fresh` is true.
-> {"ok": true, "reused": bool, "source", "duration", "video_type",
"speakers": [{id, name, display_name, speaking_seconds, share,
segment_count, avg_segment, word_count, samples,
active}],
"segments": [{id, start, end, speaker, text, excluded, words}]}
recalc_speaker_review {"voice_timeline": "...", "speakers": [...],
"segments": [...], "video_type": ""}
Reruns the same filter+recompute `save_speaker_review` persists to
_voice_timeline_clean.json, but writes nothing — a live preview for
the "Recalcular" button so muting a speaker or striking a line
updates the emphasis/peak numbers shown (pure math over the words
that survived; no new audio pass).
-> {"ok": true, "duration", "peak_count",
"segments": [{start, end, speaker, text, words, ...}]}
save_speaker_review {"voice_timeline": "...", "speakers": [...],
"segments": [...], "video_type": "", "source": "...",
"duration": 0.0}
Writes _speaker_review.json (the decisions) and
_voice_timeline_clean.json (inactive speakers + struck lines
removed) — the raw _voice_timeline.json is never touched. From here
on, `copyForChat` and `generate_voice_script` should prefer the
_clean file when it exists.
-> {"ok": true, "review_path", "clean_path", "active_speakers",
"muted_speakers", "excluded_segments"}
generate_voice_script {"media_path": "...", "voice_timeline": "...", "filepath": "...", generate_voice_script {"media_path": "...", "voice_timeline": "...", "filepath": "...",
"model": "gemma3:12b", "base_url": "http://localhost:11434", "model": "gemma3:12b", "base_url": "http://localhost:11434",
"model_size": "base", "language": "pt"|"auto"|null, "model_size": "base", "language": "pt"|"auto"|null,
@@ -90,15 +123,46 @@ Commands:
-> {"ok": true, "models": ["gemma3:12b", ...]} -> {"ok": true, "models": ["gemma3:12b", ...]}
dynamic_subtitle_config {} dynamic_subtitle_config {}
Compat: style of the FIRST ACTIVE registered layout (no id/name/
active). Prefer list_dynamic_subtitle_layouts for the app's UI.
-> {"ok": true, "band_height", "block_center_y", "line_gap", "font", -> {"ok": true, "band_height", "block_center_y", "line_gap", "font",
"font_size", "emphasis_font", "emphasis_face", "emphasis_size", "font_size", "emphasis_font", "emphasis_face", "emphasis_size",
"active_color", "emphasis_color", "text_scale"} "active_color", "emphasis_color", "text_scale"}
set_dynamic_subtitle_config {<any of the fields above>} set_dynamic_subtitle_config {<any of the fields above>}
Persists only the given fields to ~/.fcp-mcp-server/config.json. Compat: persists style fields onto the first active layout. Prefer
generate_dynamic_subtitles reads this as its own fallback default. update_dynamic_subtitle_layout for the app's UI.
-> {"ok": true, <same shape as dynamic_subtitle_config>} -> {"ok": true, <same shape as dynamic_subtitle_config>}
list_dynamic_subtitle_layouts {}
All registered "Legendas Dinâmicas" layouts. A single global config
used to hold ONE style; it's now a list of named, independently
toggleable layouts. With 2+ marked `active`, generate_dynamic_subtitles
randomly samples one per subtitle block, alternating styles through
the video. With 0 active, the first registered layout is used.
-> {"ok": true, "layouts": [{"id", "name", "active", "band_height",
"block_center_y", "line_gap", "font", "font_size",
"emphasis_font", "emphasis_face", "emphasis_size",
"active_color", "emphasis_color", "text_scale", "role"}, ...]}
create_dynamic_subtitle_layout {"name": "...", <any style field above>}
Registers a new layout, active by default. Omitted style fields fall
back to the same defaults as the very first layout.
-> {"ok": true, "layout": {...}}
update_dynamic_subtitle_layout {"id": "...", <name/active/style fields>}
Updates only the given fields of one registered layout.
-> {"ok": true, "layout": {...}} or {"ok": false, "error": "..."}
delete_dynamic_subtitle_layout {"id": "..."}
Removes a layout. If it was the last one, a "Padrão" layout is
recreated automatically so there is always at least one registered.
-> {"ok": true, "layouts": [...]}
set_dynamic_subtitle_layout_active {"id": "...", "active": true}
Toggles one layout's active flag.
-> {"ok": true, "layout": {...}}
silence_config {} silence_config {}
-> {"ok": true, "noise_db": -30.0, "min_silence": 0.5, "padding": 0.05} -> {"ok": true, "noise_db": -30.0, "min_silence": 0.5, "padding": 0.05}
@@ -108,7 +172,10 @@ Commands:
-> {"ok": true, <same shape as silence_config>} -> {"ok": true, <same shape as silence_config>}
transcribe {"path": "...", "model": "small", "language": "pt"|null, transcribe {"path": "...", "model": "small", "language": "pt"|null,
"hf_token": "..."|null, "num_speakers": ""|null} "hf_token": "..."|null, "num_speakers": ""|null,
"force": true|false}
`force: true` ignores the existing transcript cache and overwrites it
with a fresh transcription. The default is false.
-> JSON-lines: -> JSON-lines:
{"type":"progress","fraction":0.5,"stage":"Transcrevendo..."} {"type":"progress","fraction":0.5,"stage":"Transcrevendo..."}
{"type":"result","transcripts":[{"media","language","words", {"type":"result","transcripts":[{"media","language","words",
@@ -184,7 +251,7 @@ _REPO_ROOT = str(Path(__file__).resolve().parent.parent)
if _REPO_ROOT not in sys.path: if _REPO_ROOT not in sys.path:
sys.path.insert(0, _REPO_ROOT) sys.path.insert(0, _REPO_ROOT)
from admin.api import ( from admin.api import ( # noqa: E402
editing, editing,
models, models,
project, project,
@@ -194,7 +261,7 @@ from admin.api import (
voice, voice,
zoom, zoom,
) )
from admin.api.shared import emit # noqa: F401 from admin.api.shared import emit # noqa: E402,F401
def main() -> int: def main() -> int:
@@ -238,6 +305,11 @@ def main() -> int:
"analyze_voice": voice.cmd_analyze_voice, "analyze_voice": voice.cmd_analyze_voice,
"dynamic_subtitle_config": subtitles.cmd_dynamic_subtitle_config, "dynamic_subtitle_config": subtitles.cmd_dynamic_subtitle_config,
"set_dynamic_subtitle_config": subtitles.cmd_set_dynamic_subtitle_config, "set_dynamic_subtitle_config": subtitles.cmd_set_dynamic_subtitle_config,
"list_dynamic_subtitle_layouts": subtitles.cmd_list_dynamic_subtitle_layouts,
"create_dynamic_subtitle_layout": subtitles.cmd_create_dynamic_subtitle_layout,
"update_dynamic_subtitle_layout": subtitles.cmd_update_dynamic_subtitle_layout,
"delete_dynamic_subtitle_layout": subtitles.cmd_delete_dynamic_subtitle_layout,
"set_dynamic_subtitle_layout_active": subtitles.cmd_set_dynamic_subtitle_layout_active,
"plain_subtitle_config": subtitles.cmd_plain_subtitle_config, "plain_subtitle_config": subtitles.cmd_plain_subtitle_config,
"set_plain_subtitle_config": subtitles.cmd_set_plain_subtitle_config, "set_plain_subtitle_config": subtitles.cmd_set_plain_subtitle_config,
"apply_voice_actions": voice.cmd_apply_voice_actions, "apply_voice_actions": voice.cmd_apply_voice_actions,
@@ -245,6 +317,9 @@ def main() -> int:
"list_ollama_models": voice.cmd_list_ollama_models, "list_ollama_models": voice.cmd_list_ollama_models,
"build_phrase_review": review.cmd_build_phrase_review, "build_phrase_review": review.cmd_build_phrase_review,
"save_phrase_review": review.cmd_save_phrase_review, "save_phrase_review": review.cmd_save_phrase_review,
"build_speaker_review": review.cmd_build_speaker_review,
"recalc_speaker_review": review.cmd_recalc_speaker_review,
"save_speaker_review": review.cmd_save_speaker_review,
"project_config": project.cmd_project_config, "project_config": project.cmd_project_config,
"set_project_config": project.cmd_set_project_config, "set_project_config": project.cmd_set_project_config,
"silence_config": editing.cmd_silence_config, "silence_config": editing.cmd_silence_config,
+4 -4
View File
@@ -28,8 +28,8 @@ _CODE_DIR = str(Path(__file__).resolve().parent.parent / "code")
if _CODE_DIR not in sys.path: if _CODE_DIR not in sys.path:
sys.path.insert(0, _CODE_DIR) sys.path.insert(0, _CODE_DIR)
from fcpxml.media_intel import media_src_to_path from fcpxml.media_intel import media_src_to_path # noqa: E402
from fcpxml.model_manager import ( from fcpxml.model_manager import ( # noqa: E402
download_model, download_model,
get_models_dir, get_models_dir,
is_model_downloaded, is_model_downloaded,
@@ -40,8 +40,8 @@ from fcpxml.model_manager import (
save_models_dir, save_models_dir,
save_selected_model, save_selected_model,
) )
from fcpxml.parser import parse_fcpxml from fcpxml.parser import parse_fcpxml # noqa: E402
from fcpxml.transcribe import transcribe from fcpxml.transcribe import transcribe # noqa: E402
logger = logging.getLogger(__name__) logger = logging.getLogger(__name__)
+68
View File
@@ -0,0 +1,68 @@
#!/bin/zsh
# Atualiza incrementalmente a RAG do G-ART usando o banco compartilhado.
#
# Credenciais: defina RAG_DB_PASSWORD no ambiente ou crie
# admin/gart-rag.env (ignorado pelo git). O arquivo pode conter também
# RAG_DB_USER, RAG_DB_PORT, OLLAMA_URL e RAG_EMBED_MODEL.
set -euo pipefail
ROOT="$(cd "$(dirname "$0")/.." && pwd -P)"
ENV_FILE="$ROOT/admin/gart-rag.env"
if [[ -f "$ENV_FILE" ]]; then
set -a
source "$ENV_FILE"
set +a
fi
PYTHON="${RAG_PYTHON:-}"
if [[ -z "$PYTHON" ]]; then
for candidate in "$ROOT/admin/.venv/bin/python3" "$ROOT/rag/.venv/bin/python3"; do
if [[ -x "$candidate" ]]; then PYTHON="$candidate"; break; fi
done
fi
PYTHON="${PYTHON:-$(command -v python3)}"
if ! "$PYTHON" -c 'import psycopg2, requests' >/dev/null 2>&1; then
echo "ERRO: o Python da RAG precisa dos pacotes psycopg2 e requests." >&2
echo "Instale-os no ambiente indicado por RAG_PYTHON e tente novamente." >&2
exit 1
fi
if [[ -z "${RAG_DB_PASSWORD:-}" ]]; then
echo "ERRO: defina RAG_DB_PASSWORD ou configure $ENV_FILE" >&2
exit 1
fi
HOST="${RAG_VPS_HOST:-179.197.228.240}"
LOCAL_PORT="${RAG_DB_PORT:-55435}"
REMOTE_PORT="${RAG_REMOTE_PORT:-55435}"
TUNNEL_PID=""
cleanup() {
if [[ -n "$TUNNEL_PID" ]] && kill -0 "$TUNNEL_PID" 2>/dev/null; then
kill "$TUNNEL_PID" 2>/dev/null || true
wait "$TUNNEL_PID" 2>/dev/null || true
fi
}
trap cleanup EXIT
if ! nc -z 127.0.0.1 "$LOCAL_PORT" 2>/dev/null; then
echo "==> Abrindo túnel RAG (127.0.0.1:$LOCAL_PORT)..."
ssh -N -o ExitOnForwardFailure=yes -o ServerAliveInterval=30 \
-o ServerAliveCountMax=3 -L "127.0.0.1:$LOCAL_PORT:127.0.0.1:$REMOTE_PORT" \
"${RAG_VPS_USER:-root}@$HOST" &
TUNNEL_PID=$!
for _ in {1..20}; do
nc -z 127.0.0.1 "$LOCAL_PORT" 2>/dev/null && break
kill -0 "$TUNNEL_PID" 2>/dev/null || break
sleep 0.25
done
fi
if ! nc -z 127.0.0.1 "$LOCAL_PORT" 2>/dev/null; then
echo "ERRO: não foi possível abrir o túnel RAG." >&2
exit 1
fi
echo "==> Atualizando RAG do G-ART (incremental)..."
cd "$ROOT"
exec "$PYTHON" "$ROOT/admin/update_rag.py"
+216
View File
@@ -0,0 +1,216 @@
#!/usr/bin/env python3
"""Atualiza incrementalmente o índice RAG do G-ART.
As credenciais são fornecidas pelo ambiente; este arquivo nunca deve conter
senha. O indexador usa o banco ``rag_gart`` e o schema ``gart`` por padrão.
"""
from __future__ import annotations
import hashlib
import os
import re
import sys
from pathlib import Path
import psycopg2
import requests
ROOT = Path(__file__).resolve().parents[1]
DB_NAME = os.environ.get("RAG_DB_NAME", "rag_gart")
DB_SCHEMA = os.environ.get("RAG_DB_SCHEMA", "gart")
OLLAMA_URL = os.environ.get("OLLAMA_URL", "http://127.0.0.1:11434")
EMBED_MODEL = os.environ.get("RAG_EMBED_MODEL", "nomic-embed-text")
EMBED_DIM = int(os.environ.get("RAG_EMBED_DIM", "768"))
INCLUDE_EXTENSIONS = {
".command", ".md", ".py", ".sh", ".sql", ".swift", ".txt", ".yml", ".yaml",
}
EXCLUDE_DIRS = {
".git", ".venv", ".pytest_cache", ".ruff_cache", "__pycache__", "build",
"dist", "node_modules", "graphify-out", "bm", "models", "whisper",
}
EXCLUDE_FILES = {".env", "admin/genial-crm.env", "admin/genial-crm.local.env"}
CHUNK_LINES = 60
CHUNK_OVERLAP = 10
CHUNK_MAX_CHARS = 5000
def _sql_id(value: str) -> str:
return '"' + value.replace('"', '""') + '"'
def _connect():
password = os.environ.get("RAG_DB_PASSWORD")
if not password:
raise RuntimeError("RAG_DB_PASSWORD não foi definida")
return psycopg2.connect(
host=os.environ.get("RAG_DB_HOST", "127.0.0.1"),
port=os.environ.get("RAG_DB_PORT", "55435"),
dbname=DB_NAME,
user=os.environ.get("RAG_DB_USER", "gart_rag_indexer"),
password=password,
connect_timeout=5,
)
def _iter_files():
for path in ROOT.rglob("*"):
if not path.is_file() or path.suffix.lower() not in INCLUDE_EXTENSIONS:
continue
rel = path.relative_to(ROOT).as_posix()
parts = set(path.relative_to(ROOT).parts)
if parts & EXCLUDE_DIRS or rel in EXCLUDE_FILES or path.name in EXCLUDE_FILES:
continue
if any(part.startswith(".") for part in path.relative_to(ROOT).parts[:-1]):
continue
yield path, rel
def _chunks(text: str):
lines = text.splitlines()
if not lines:
return []
step = max(1, CHUNK_LINES - CHUNK_OVERLAP)
result = []
for start in range(0, len(lines), step):
window_start = start
buffer = []
size = 0
for offset, line in enumerate(lines[start:start + CHUNK_LINES]):
if buffer and size + len(line) + 1 > CHUNK_MAX_CHARS:
result.append((window_start + 1, window_start + len(buffer), "\n".join(buffer).strip()))
buffer = []
window_start = start + offset
size = 0
buffer.append(line)
size += len(line) + 1
if buffer:
result.append((window_start + 1, window_start + len(buffer), "\n".join(buffer).strip()))
if start + CHUNK_LINES >= len(lines):
break
return [(start, end, content) for start, end, content in result if content]
def _facts(text: str, rel_path: str):
lines = text.splitlines()
summary = next(
(line.strip().lstrip("#! ").strip() for line in lines[:30] if line.strip()),
None,
)
symbols = re.findall(
r"^\s*(?:class|def|async\s+def|func|struct|enum|protocol|actor|interface)\s+([A-Za-z_]\w*)",
text,
re.MULTILINE,
)
parts = Path(rel_path).parts
module = parts[0] if len(parts) > 1 else None
return module, Path(rel_path).stem, summary, sorted(set(symbols)), len(lines)
class ChunkTooLargeError(Exception):
"""Chunk excede o contexto do modelo de embedding (ver EXCLUDE_FILES/CHUNK_MAX_CHARS)."""
def _embed(text: str):
response = requests.post(
f"{OLLAMA_URL.rstrip('/')}/api/embeddings",
json={"model": EMBED_MODEL, "prompt": f"search_document: {text}"},
timeout=60,
)
if response.status_code == 500 and "context length" in response.text.lower():
raise ChunkTooLargeError(response.text)
response.raise_for_status()
vector = response.json()["embedding"]
if len(vector) != EMBED_DIM:
raise ValueError(f"embedding com {len(vector)} dimensões; esperado {EMBED_DIM}")
return vector
def _hash(text: str) -> str:
return hashlib.md5(text.encode("utf-8")).hexdigest()
def index():
schema = _sql_id(DB_SCHEMA)
conn = _connect()
conn.autocommit = False
indexed = skipped = deleted = chunks_written = 0
seen = set()
try:
with conn.cursor() as cur:
for path, rel_path in sorted(_iter_files(), key=lambda item: item[1]):
try:
text = path.read_text(encoding="utf-8", errors="ignore")
except OSError as exc:
print(f"[RAG] ignorado {rel_path}: {exc}", file=sys.stderr)
continue
seen.add(rel_path)
digest = _hash(text)
cur.execute(f"SELECT content_hash FROM {schema}.indexed_files WHERE file_path = %s", (rel_path,))
row = cur.fetchone()
if row and row[0] == digest:
skipped += 1
continue
module, main_type, summary, symbols, n_lines = _facts(text, rel_path)
cur.execute(f"DELETE FROM {schema}.code_chunks WHERE file_path = %s", (rel_path,))
for index_number, (start, end, content) in enumerate(_chunks(text)):
try:
vector = _embed(content)
except ChunkTooLargeError:
# Chunks densos em tokens (ex: tabelas de dados numéricas
# como font_metrics.py) podem passar de CHUNK_MAX_CHARS em
# caracteres mas estourar o contexto do modelo em tokens.
# Pular o chunk em vez de abortar a transação inteira.
print(f"[RAG] chunk grande demais, pulado: {rel_path}:{start}-{end}", file=sys.stderr)
continue
cur.execute(
f"""INSERT INTO {schema}.code_chunks
(file_path, content, chunk_index, embedding, content_hash,
file_mtime, start_line, end_line, symbols, module, kind)
VALUES (%s, %s, %s, %s, %s, %s, %s, %s, %s, %s, %s)""",
(rel_path, content, index_number, vector, digest,
path.stat().st_mtime, start, end, ", ".join(symbols), module, "window"),
)
chunks_written += 1
cur.execute(
f"""INSERT INTO {schema}.file_index
(file_path, module, main_type, public_symbols, summary, n_lines, content_hash)
VALUES (%s, %s, %s, %s, %s, %s, %s)
ON CONFLICT (file_path) DO UPDATE SET
module = EXCLUDED.module, main_type = EXCLUDED.main_type,
public_symbols = EXCLUDED.public_symbols, summary = EXCLUDED.summary,
n_lines = EXCLUDED.n_lines, content_hash = EXCLUDED.content_hash,
updated_at = CURRENT_TIMESTAMP""",
(rel_path, module, main_type, symbols, summary, n_lines, digest),
)
cur.execute(
f"""INSERT INTO {schema}.indexed_files (file_path, content_hash)
VALUES (%s, %s)
ON CONFLICT (file_path) DO UPDATE SET
content_hash = EXCLUDED.content_hash, updated_at = CURRENT_TIMESTAMP""",
(rel_path, digest),
)
indexed += 1
print(f"[RAG] {rel_path}")
cur.execute(f"SELECT file_path FROM {schema}.indexed_files")
for (rel_path,) in cur.fetchall():
if rel_path not in seen:
cur.execute(f"DELETE FROM {schema}.code_chunks WHERE file_path = %s", (rel_path,))
cur.execute(f"DELETE FROM {schema}.file_index WHERE file_path = %s", (rel_path,))
cur.execute(f"DELETE FROM {schema}.indexed_files WHERE file_path = %s", (rel_path,))
deleted += 1
print(f"[RAG] removido {rel_path}")
conn.commit()
except Exception:
conn.rollback()
raise
finally:
conn.close()
print(f"[RAG] concluído: {indexed} atualizado(s), {skipped} sem mudança, {deleted} removido(s), {chunks_written} chunk(s).")
if __name__ == "__main__":
index()
+65 -7
View File
@@ -7,7 +7,7 @@ Mapa módulo a módulo do núcleo Python: onde cada coisa mora e o que ela faz.
A API pública é reexportada em `fcpxml/__init__.py` — essa é a fonte da verdade A API pública é reexportada em `fcpxml/__init__.py` — essa é a fonte da verdade
do `__all__`. do `__all__`.
Versão: `0.6.35` · Última varredura: 2026-08-19 Versão: `0.13.1` · Última varredura: 2026-09-22
> **Por que existem pacotes aqui.** `writer.py` tinha 4.199 linhas e `models.py` > **Por que existem pacotes aqui.** `writer.py` tinha 4.199 linhas e `models.py`
> 1.091, cada um com muitos assuntos dentro. Viraram pacotes com um módulo por > 1.091, cada um com muitos assuntos dentro. Viraram pacotes com um módulo por
@@ -21,13 +21,16 @@ Versão: `0.6.35` · Última varredura: 2026-08-19
| Módulo / pacote | Linhas | Papel | | Módulo / pacote | Linhas | Papel |
|-----------------|-------:|-------| |-----------------|-------:|-------|
| `writer/` | 4.687 | **Edição e escrita de FCPXML** — o coração | | `writer/` | 5.377 | **Edição e escrita de FCPXML** — o coração |
| `models/` | 1.195 | Data classes e enums | | `models/` | 1.195 | Data classes e enums |
| `text_layout.py` | 901 | Diagramação das legendas dinâmicas | | `text_layout.py` | 901 | Diagramação das legendas dinâmicas |
| `rough_cut.py` | 798 | Geração de timelines novas | | `rough_cut.py` | 798 | Geração de timelines novas |
| `model_manager.py` | 748 | Modelos Whisper: catálogo, download, config | | `model_manager.py` | 748 | Modelos Whisper: catálogo, download, config |
| `voice_timeline.py` | 600 | O JSON de voz que a IA lê | | `voice_timeline.py` | 600 | O JSON de voz que a IA lê |
| `analise.py` | 218 | `AnalisadorDeArquivo` — orquestra transcrição/diarização/ênfase/emoção e monta o `_voice_timeline.json`; usado por `voice_timeline.py` |
| `phrase_review.py` | 547 | Revisão de frases (etapa 5 do assistente) | | `phrase_review.py` | 547 | Revisão de frases (etapa 5 do assistente) |
| `speaker_review.py` | 208 | Revisão de falantes (etapa 3 do assistente) |
| `transcription/` | 325 | Pacote: `engine.py` (adapter faster-whisper), `segments.py`/`text.py`/`timestamps.py` (operações puras sobre transcript); `transcribe.py` é a fachada de compatibilidade |
| `collision.py` | 472 | Colisão entre títulos na tela | | `collision.py` | 472 | Colisão entre títulos na tela |
| `font_metrics.py` | 445 | Largura real de glifos por fonte | | `font_metrics.py` | 445 | Largura real de glifos por fonte |
| `templates.py` | 387 | Templates de timeline | | `templates.py` | 387 | Templates de timeline |
@@ -36,7 +39,7 @@ Versão: `0.6.35` · Última varredura: 2026-08-19
| `forced_align.py` | 181 | Alinhamento forçado opcional (whisperx/wav2vec2) que corrige o viés de ~0,4s no início das palavras | | `forced_align.py` | 181 | Alinhamento forçado opcional (whisperx/wav2vec2) que corrige o viés de ~0,4s no início das palavras |
| `live.py` | 273 | Modo Live (push_to_fcp) | | `live.py` | 273 | Modo Live (push_to_fcp) |
| `diff.py` | 269 | Comparação de timelines | | `diff.py` | 269 | Comparação de timelines |
| `voice_actions.py` | 263 | Decisões de edição (cut/zoom/text/marker) | | `voice_actions.py` | 319 | Decisões de edição (cut/zoom/text/marker) |
| `export.py` | 226 | Export Resolve v1.9 + FCP7 XMEML v5 | | `export.py` | 226 | Export Resolve v1.9 + FCP7 XMEML v5 |
| `voice_features.py` | 220 | Pitch, energia, ritmo, pausas | | `voice_features.py` | 220 | Pitch, energia, ritmo, pausas |
| `diarize.py` | 180 | Quem falou (pyannote) | | `diarize.py` | 180 | Quem falou (pyannote) |
@@ -55,9 +58,10 @@ editorial, todos operando sobre o mesmo documento e os mesmos índices.
| Módulo | Linhas | Conteúdo | | Módulo | Linhas | Conteúdo |
|--------|-------:|----------| |--------|-------:|----------|
| `core.py` | 723 | `ModifierCore`: carga, índices, navegação na spine, `save` | | `core.py` | 723 | `ModifierCore`: carga, índices, navegação na spine, `save` |
| `titles.py` | 600 | Títulos de texto e legendas dinâmicas | | `titles.py` | 867 | Títulos de texto e legendas dinâmicas |
| `cut.py` | 333 | Dividir, cortar faixas, apagar | | `cut.py` | 333 | Dividir, cortar faixas, apagar |
| `speed.py` | 297 | Velocidade e zoom (punch-in) | | `speed.py` | 94 | Velocidade de reprodução |
| `zoom.py` | 204 | Zoom (punch-in) via clipe de ajuste conectado |
| `helpers.py` | 279 | Sanitização, escalas, construtores de elemento | | `helpers.py` | 279 | Sanitização, escalas, construtores de elemento |
| `rapid.py` | 240 | Flash frames, rapid trim, preencher buracos | | `rapid.py` | 240 | Flash frames, rapid trim, preencher buracos |
| `validation.py` | 232 | Verificações estruturais antes de salvar | | `validation.py` | 232 | Verificações estruturais antes de salvar |
@@ -66,7 +70,9 @@ editorial, todos operando sobre o mesmo documento e os mesmos índices.
| `document.py` | 170 | Assets de vídeo, timebases, `write_fcpxml` | | `document.py` | 170 | Assets de vídeo, timebases, `write_fcpxml` |
| `markers.py` | 165 | Marcadores: um, por timecode, em lote | | `markers.py` | 165 | Marcadores: um, por timecode, em lote |
| `audio.py` | 162 | Clipes de áudio e cama musical | | `audio.py` | 162 | Clipes de áudio e cama musical |
| `generator.py` | 147 | `FCPXMLWriter` — cria documento do zero | | `generator.py` | 94 | `FCPXMLWriter` — orquestra a criação do zero (estado + delegação) |
| `builders.py` | 174 | Um builder por tipo de elemento: `FormatBuilder`, `AssetBuilder`, `MarkerBuilder`, `KeywordBuilder`, `ClipBuilder`, `SequenceBuilder`, `LibraryBuilder` |
| `adjustment.py` | 140 | `ClipDeAjuste` — camada de ajuste (filtros `filter-video`/`filter-audio` direto no `<clip>`, sem uso ainda em `server_tools`/`admin/api`) |
| `reorder.py` | 126 | Reordenar e recalcular offsets | | `reorder.py` | 126 | Reordenar e recalcular offsets |
| `trim.py` | 125 | Aparar e propagar o ripple | | `trim.py` | 125 | Aparar e propagar o ripple |
| `transitions.py` | 94 | Transições entre vizinhos | | `transitions.py` | 94 | Transições entre vizinhos |
@@ -121,7 +127,10 @@ voice_features.py pitch, energia, ritmo, pausas
▼ ▼
emphasis.py combina tudo num índice 0–1 por palavra emphasis.py combina tudo num índice 0–1 por palavra
▼ ▼
voice_timeline.py monta o _voice_timeline.json ◄── é isto que a IA lê voice_timeline.py monta o _voice_timeline.json ◄── a análise crua
▼
speaker_review.py (opcional) filtra falante mutado + linha riscada
→ _voice_timeline_clean.json ◄── é isto que a IA prefere
▼ ▼
[decisão: skill "editar-por-voz", ou a mão do usuário] [decisão: skill "editar-por-voz", ou a mão do usuário]
▼ ▼
@@ -155,6 +164,36 @@ Saída em camadas, para um modelo raciocinar do topo e descer só onde importa:
`layers` existe para separar *"a fala é monótona"* de *"a análise acústica nunca `layers` existe para separar *"a fala é monótona"* de *"a análise acústica nunca
carregou"* — os dois deixam os mesmos zeros nos dados. carregou"* — os dois deixam os mesmos zeros nos dados.
### `speaker_review.py` — a triagem antes da IA
Roda logo após `analyze_voice` (etapa 3 do assistente, tela `SpeakerReviewView`
no app): lista quem foi detectado (`speaker_profiles`, com % de fala e falas
de amostra) e a transcrição segmento a segmento, para o usuário nomear cada
falante, mutar quem não interessa (ex.: o entrevistador) e riscar linhas soltas
antes de qualquer IA ver o arquivo. `build_speaker_review` nunca toca a
timeline crua; `apply_speaker_review`/`write_clean_voice_timeline` produzem
uma cópia separada, `_voice_timeline_clean.json`, reaproveitando
`enrich_words`/`_segment_rows`/`_summary` de `voice_timeline.py` para
recalcular a ênfase só sobre quem sobrou — mesma lógica de `restrict_to_kept`,
por falante/segmento em vez de por intervalo de tempo. O merge de decisões
salvas segue o padrão de `phrase_review.merge_saved_decisions`: sempre
reconstrói da análise atual, só as escolhas humanas persistem.
A skill "editar-por-voz", `generate_voice_script` e `cmd_build_phrase_review`
(etapa 4/5, `admin/api/review.py`) preferem o `_clean` quando ele existe; sem
revisão salva, seguem lendo o `_voice_timeline.json` normal — a etapa 3 é
sempre opcional. Os três pontos de leitura precisam concordar nessa
preferência: se um deles voltar a ler o arquivo cru direto, a revisão de
falantes vira letra morta sem nenhum erro visível (ver `05_EXPERIENCIAS.md`).
Cada linha em `build_speaker_review` carrega suas `words` originais (ênfase
por palavra), para a tela desenhar os mesmos chips da etapa 5 sem esperar um
recálculo. `apply_speaker_review` é reaproveitada por dois caminhos: gravar
(`write_clean_voice_timeline`, via `save_speaker_review`) e só **prever**
(`cmd_recalc_speaker_review`, sem tocar disco) — o botão "Recalcular" da tela
usa o segundo caminho para atualizar a ênfase só sobre quem sobreviveu ao
corte, sem reprocessar áudio.
### `phrase_review.py` — a revisão humana ### `phrase_review.py` — a revisão humana
Junta o timeline de voz com as ações da IA numa lista de frases editáveis, e Junta o timeline de voz com as ações da IA numa lista de frases editáveis, e
@@ -175,6 +214,25 @@ converte de volta. Frase inativa vira `cut`; ênfase ≥ 1 vira `zoom` mais um
Estes três não estão divididos porque **cada um já é um assunto só**. O Estes três não estão divididos porque **cada um já é um assunto só**. O
`text_layout.py` tem 901 linhas de um problema coeso: diagramação. `text_layout.py` tem 901 linhas de um problema coeso: diagramação.
### Separação de role entre legendas dinâmicas e convencionais
As duas categorias são ambas `<title>` conectados, mas recebem **roles
diferentes** para ficarem didáticas na timeline do FCP (cada role ganha cor
própria no índice). O atributo usado em `<title>` é `role` (CDATA) — **nunca**
`videoRole`, que é DTD-inválido para títulos (ver `05_EXPERIENCIAS.md`,
entrada 32).
| Categoria | `role` | De onde vem |
|-----------|--------|-------------|
| Legendas dinâmicas | `titles.dinamicas` | `DynamicSubtitleConfig.role` / `load_dynamic_subtitle_config()["role"]` |
| Legendas convencionais | `titles.convencionais` | `load_plain_subtitle_config()["role"]` |
A cor do texto em si continua nos configs de fonte (abas do app), não no
role. Os geradores `generate_dynamic_subtitles` (writer/titles.py),
`handle_generate_plain_subtitles` e `handle_generate_subtitles_by_emphasis`
(server_tools/subtitles.py) aplicam o role em cada `<title>` criado; o
parâmetro `role` das ferramentas MCP sobrescreve o default.
--- ---
## Armadilhas do FCPXML (custaram sessões de depuração) ## Armadilhas do FCPXML (custaram sessões de depuração)
+53 -14
View File
@@ -1,6 +1,6 @@
# 03 — Camada MCP (`server.py` + `server_tools/`) — 77 ferramentas # 03 — Camada MCP (`server.py` + `server_tools/`) — 78 ferramentas
> **Escopo:** As 77 ferramentas MCP: helpers, categorias e como criar uma nova. > **Escopo:** As 78 ferramentas MCP: helpers, categorias e como criar uma nova.
> **Não cobre:** Lógica de edição, que mora no engine (→ 02) · comandos do app (→ 08) > **Não cobre:** Lógica de edição, que mora no engine (→ 02) · comandos do app (→ 08)
`server.py` (592 linhas) é só o transporte: dispatch por dicionário `server.py` (592 linhas) é só o transporte: dispatch por dicionário
@@ -82,8 +82,21 @@ continua funcionando. A coluna diz o módulo real, para quando você precisar
### Voz (análise → decisão → aplicação) ### Voz (análise → decisão → aplicação)
`analyze_voice_features`, `build_voice_timeline`, `refine_voice_timeline`, `analyze_voice_features`, `build_voice_timeline`, `refine_voice_timeline`,
`remove_speakers`, `apply_voice_actions`, `generate_voice_script`, `remove_speakers`, `remove_speech_gaps`, `apply_voice_actions`,
`get_voice_analysis_config`, `save_voice_analysis_config`. `generate_voice_script`, `get_voice_analysis_config`,
`save_voice_analysis_config`.
`remove_speech_gaps` corta pelo que a **transcrição** já sabe que não tem
fala — lê `words[].start/end` do `_voice_timeline.json` (função
`speech_gap_cut_actions`, em `fcpxml/voice_actions.py`) em vez de medir
volume. É o complemento correto para o caso que `remove_media_silence`
(silêncio por dB, ver seção "Silêncio e beats") não cobre: um trecho sem
fala mas com som real acima do limiar (respiração, ruído de roupa, batida) —
`remove_media_silence` nunca vai cortar isso porque tecnicamente não é
silêncio. Não corta a lacuna antes da primeiríssima palavra (pode ser quase
o arquivo inteiro, antes da tomada realmente começar) — isso continua
decisão manual na Fase 6 do `apply_voice_actions`
(`.claude/skills/editar-por-voz/criterios/06-texto-corte-marcador.md`).
O fluxo manual é: `build_voice_timeline` mede (caro, roda uma vez) → O fluxo manual é: `build_voice_timeline` mede (caro, roda uma vez) →
o modelo decide os cortes → **`refine_voice_timeline` renormaliza sobre o que o modelo decide os cortes → **`refine_voice_timeline` renormaliza sobre o que
@@ -124,16 +137,18 @@ severidade probable/severe → investigar CADA colisão pela fração exata do
XML antes de mudar código (ver checklist abaixo) XML antes de mudar código (ver checklist abaixo)
``` ```
**`generate_subtitles_by_emphasis`** gera as duas legendas numa passada só — **`generate_subtitles_by_emphasis`** gera as duas legendas numa passada só,
mas não divide as palavras entre elas. A comum é gerada **completa, do início dividindo por palavra: a dinâmica cobre as frases marcadas como ênfase na
ao fim do clipe**, sempre; a dinâmica é gerada só sobre as frases marcadas etapa 5 (zoom aplicado, nível ≥ 1); a comum cobre **todo o resto** — um bloco
como ênfase na etapa 5 (zoom aplicado, nível ≥ 1); e onde a dinâmica cobre um comum simplesmente não é criado onde a dinâmica já cobre. A primeira versão
trecho, os títulos comuns daquele trecho recebem `enabled="0"` — continuam no gerava a comum inteira e desativava (`enabled="0"`) o que ficava sob a
XML (editáveis/reativáveis no Final Cut), só não são desenhados. É a tradução dinâmica, mas um título desativado continua aparecendo como clipe riscado na
literal de `10-revisao-humana.md` (skill `editar-por-voz`): "a frase de timeline do Final Cut mesmo sem renderizar — um corte com bastante ênfase
ênfase recebe zoom E legenda dinâmica; as demais recebem legenda comum" — enchia a trilha de clipes mortos. Trocado por não gerar ali: o preço é que,
sem nunca deixar um vão sem legenda nenhuma se a ênfase for desativada depois se a ênfase for desativada à mão depois, a legenda comum daquele trecho
(a comum já estava lá, só desligada). A decisão vem de precisa ser regenerada, não só reativada. É a tradução de `10-revisao-humana.md`
(skill `editar-por-voz`): "a frase de ênfase recebe zoom E legenda dinâmica;
as demais recebem legenda comum". A decisão vem de
`<mídia>_phrase_actions.json["emphasis_spans"]`, escrito por `<mídia>_phrase_actions.json["emphasis_spans"]`, escrito por
`save_phrase_review` quando o editor termina a etapa 5 — sem esse arquivo (ou `save_phrase_review` quando o editor termina a etapa 5 — sem esse arquivo (ou
sem `zoom`/`text` marcados na revisão), a tool gera só a comum, tudo ligado, sem `zoom`/`text` marcados na revisão), a tool gera só a comum, tudo ligado,
@@ -141,6 +156,30 @@ e avisa no relatório ("Sem revisão de ênfase"). Não expõe overrides de esti
por chamada — usa a config salva ("Legendas Dinâmicas"/plain); para estilo por chamada — usa a config salva ("Legendas Dinâmicas"/plain); para estilo
pontual, use `generate_dynamic_subtitles`/`generate_plain_subtitles` direto. pontual, use `generate_dynamic_subtitles`/`generate_plain_subtitles` direto.
**Separação por role (didática na timeline):** `generate_dynamic_subtitles` e
`generate_plain_subtitles` (e a metade dinâmica/comum do `by_emphasis`) aplicam
`role="titles.dinamicas"` e `role="titles.convencionais"` em cada `<title>`
criado — sub-roles de `titles`, **nunca** `subtitles.*` (que esconderia o título
atrás de Code). O parâmetro `role` de cada ferramenta MCP sobrescreve o default
(vindo de `load_dynamic_subtitle_config()["role"]` /
`load_plain_subtitle_config()["role"]`). Ver `02_MODULES.md` (seção "Separação
de role") e `05_EXPERIENCIAS.md` entrada 32 (DTD: `<title>` leva `role`, não
`videoRole`).
**Compound clip por sub-frase (padrão em `generate_dynamic_subtitles` e na
metade dinâmica do `by_emphasis`):** `compound_subphrases=True` divide cada
frase em sub-frases pela vírgula (`transcribe.split_into_subphrases`) e
empacota os `<title>` de cada uma num `<ref-clip>` — a dúzia de títulos
empilhados por lane que uma frase gera vira uma barra só, arrastável/mutável
como unidade. Exceção: um trecho curto depois da vírgula ("né?", "Então...",
< 3 palavras) funde de volta na sub-frase anterior em vez de virar compound
próprio — soa como parte da mesma respiração, não uma frase nova. A estrutura
replica o que o próprio Final Cut gera em "New Compound Clip": o título mais
cedo vira âncora do spine interno em offset 0, os demais penduram nele por
lane. `validate_subtitle_layout` mede cada compound no seu próprio espaço de
tempo — sem isso, âncoras de compounds diferentes leem "0s" e colidem no
papel mesmo estando segundos distantes na timeline real.
**Antes de atribuir uma colisão ao gerador, confirme que é o gerador.** **Antes de atribuir uma colisão ao gerador, confirme que é o gerador.**
Um `<title>` de nome estranho (`ref` diferente, params tipo `Auto-Shrink`/ Um `<title>` de nome estranho (`ref` diferente, params tipo `Auto-Shrink`/
`Left Margin` que `_make_text_title_clip` nunca escreve) é conteúdo humano `Left Margin` que `_make_text_title_clip` nunca escreve) é conteúdo humano
+479
View File
@@ -44,6 +44,16 @@ que merece entrada.
| 24 | 2026-08-19 | `admin/test_models_api.py` existia mas estava fora de `testpaths` — 13 testes que nunca rodaram | `resolvido` | | 24 | 2026-08-19 | `admin/test_models_api.py` existia mas estava fora de `testpaths` — 13 testes que nunca rodaram | `resolvido` |
| 25 | 2026-08-20 | `admin/api/shared.py` apontava para `admin/code` (inexistente) após a divisão — install editável mascarou o bug em toda validação anterior | `resolvido` | | 25 | 2026-08-20 | `admin/api/shared.py` apontava para `admin/code` (inexistente) após a divisão — install editável mascarou o bug em toda validação anterior | `resolvido` |
| 26 | 2026-08-21 | `generate_voice_script` (IA local/Ollama) caía com "Falha ao gerar roteiro por IA local" — prompt embutia a timeline inteira (47k tokens) e estourava `num_ctx`; e `response.json()` de conexão caída escapava como `JSONDecodeError` | `resolvido` | | 26 | 2026-08-21 | `generate_voice_script` (IA local/Ollama) caía com "Falha ao gerar roteiro por IA local" — prompt embutia a timeline inteira (47k tokens) e estourava `num_ctx`; e `response.json()` de conexão caída escapava como `JSONDecodeError` | `resolvido` |
| 27 | 2026-08-21 | Cortes escritos rente ao timestamp da palavra soam secos — critério da skill e prompt do modelo local não instruíam folga na borda | `resolvido` |
| 28 | 2026-08-21 | Frases desativadas em sequência deixavam fatias de 0,1-0,5s sobrando entre clipes | `resolvido` |
| 29 | 2026-08-21 | `remove_media_silence` (dB) não corta lacuna sem fala mas com som real — trecho sobrevivia intacto na timeline final | `resolvido` |
| 30 | 2026-08-24 | `add_zoom` animava `<adjust-transform>` direto no clipe, diferente de como o FCP realmente exporta zoom (clipe de ajuste conectado) | `resolvido` |
| 31 | 2026-08-24 | Revisão de falantes ("Quem fica na edição") salvava certo, mas etapa 4 (Revisão de frases) lia a timeline crua, ignorando falantes mutados/linhas riscadas | `resolvido` |
| 32 | 2026-08-24 | Separar legendas dinâmicas de convencionais por role: `<title>` aceita `role` (CDATA), NÃO `videoRole` — este último é DTD-inválido para títulos e quebra a validação | `resolvido` |
| 33 | 2026-09-22 | Legenda comum sobreposta à composição dinâmica em `generate_subtitles_by_emphasis`; regenerar acumulava títulos em vez de substituir | `resolvido` |
| 34 | 2026-09-22 | `ClipDeAjuste` gerava wrapper `<adjustment>` inválido no DTD; `code/WHISPERX` era 2,6 GB de backup órfão que inflava o lint quando rodado com `--exclude` explícito | `resolvido` |
| 35 | 2026-09-22 | Indexação RAG (`admin/update_rag.py`) abortava a transação inteira ao achar um chunk que estoura o contexto do modelo de embedding | `resolvido` |
| 36 | 2026-09-23 | `.gitignore` com regra `models/` solta escondia do git o pacote inteiro `fcpxml/models/` (dados do engine), não só o cache do Whisper | `resolvido` |
> Mantenha o índice acima sempre sincronizado com as entradas mais recentes. > Mantenha o índice acima sempre sincronizado com as entradas mais recentes.
@@ -1414,3 +1424,472 @@ o outro; percentil entrega um punhado útil nos dois casos.
> de análise cru no prompt; projetar só o que a decisão usa. E qualquer parse > de análise cru no prompt; projetar só o que a decisão usa. E qualquer parse
> de resposta de servidor local deve tratar body vazio/quebrado como erro de > de resposta de servidor local deve tratar body vazio/quebrado como erro de
> transporte, não como sucesso mudo. > transporte, não como sucesso mudo.
---
### 2026-08-21 — Cortes escritos rente ao timestamp da palavra soam secos
- **Sintoma:** usuário revisou o corte final (projeto Mastopexia) e reportou
"os cortes estão muito secos, principalmente no final de frase — falta um
tempinho a mais pra concluir as palavras". Também notou que o ar morto
antes da primeira fala do vídeo não tinha sido cortado.
- **Causa:** o critério `06-texto-corte-marcador.md` (e o prompt embutido do
modelo local em `fcpxml/llm_local.py`) instruíam cobrir a frase inteira
(`start..end = início..fim da frase`) ao escrever um `cut`, sem nenhuma
orientação sobre a borda que encosta em fala **mantida** (não em silêncio
puro). Um `cut` com `start` exatamente no fim da última palavra mantida
engole essa palavra antes dela terminar de soar; um `cut` com `end` no
início exato da próxima engole o ataque da fala seguinte. É um problema
diferente de cortar a pausa curta (proibido, é a própria ênfase) — aqui a
pausa natural entre os blocos já existe, e o corte estava comendo essa
margem sozinho.
- **Correção:**
- `06-texto-corte-marcador.md` ganhou a seção "Nunca corte rente à
palavra — deixe uma folga": recuar `start`/`end` do corte em ~0,15–0,25s
para dentro do próprio corte nas bordas que tocam fala mantida (não em
silêncio puro), incluindo o início/fim do vídeo.
- `fcpxml/llm_local.py::_SYSTEM_PROMPT` (item 4) recebeu a mesma
instrução, para o modelo local gerar decisões já com a folga.
- **Validação manual:** reaplicado no projeto Mastopexia real —
`10.77 → 95.50` (rente) virou `10.97 → 95.30` (folga de ~0,2s nas duas
pontas), e as 4 emendas seguintes receberam o mesmo tratamento; zoom/texto/
marcador continuaram longe o suficiente da nova borda do corte — a folga
também evita o problema relacionado (não corrigido em código, só
contornado manualmente nesta sessão): um `zoom`/`marker` cuja borda cai
exatamente em cima do início/fim de um `cut` é descartado por
`resolve_actions` como "apontando para material cortado", mesmo quando a
intenção era ficar bem ao lado. Vale registrar como dívida: `resolve_actions`
poderia tolerar uma margem de meio-frame antes de considerar a ação "dentro"
do corte.
- **Estado:** `resolvido`
> **Aprendizado:** "cobrir a frase inteira" não é a instrução completa para
> um corte — a frase que **sobra** ao lado do corte também precisa de uma
> borda que respire. Regra prática: só cortar rente ao timestamp quando a
> borda encosta em silêncio real (`gap_before` grande) ou em conteúdo que
> também será descartado; encostando em fala mantida, sempre recuar.
---
### 2026-08-21 — Frases desativadas em sequência deixavam fatias de 0,1-0,5s sobrando
- **Sintoma:** usuário viu, no Final Cut, um clipe minúsculo sobrando entre
dois clipes normais na timeline (projeto Mastopexia, confirmado por
screenshot). Investigação achou 29 `cut`s individuais no
`_phrase_actions.json` gerado pela etapa 5, e a timeline final saiu com
mais de uma dezena de fatias de 0,1-0,5s entre clipes.
- **Causa:** `phrase_review_to_actions()` (`fcpxml/phrase_review.py`) gerava
**um `cut` por frase desativada**, cobrindo só `[phrase.start, phrase.end]`.
Quando duas ou mais frases seguidas estão desativadas, a pausa **entre**
elas nunca pertence a nenhuma frase — não é coberta por nenhum `cut` — e
sobrevive como um clipe próprio, minúsculo, que ninguém pediu para manter.
- **Correção:** `phrase_review_to_actions()` agora agrupa frases desativadas
**consecutivas** (`flush_inactive_run()`) e emite um único `cut` cobrindo do
início da primeira ao fim da última do grupo, absorvendo as pausas entre
elas. Uma frase ativa no meio ainda quebra o grupo — cuts continuam
separados quando há conteúdo mantido entre eles.
- **Validação:** `tests/test_phrase_review.py` ganhou
`test_consecutive_inactive_phrases_merge_into_one_cut`,
`test_inactive_run_at_the_end_still_flushes` e
`test_isolated_inactive_phrases_stay_separate_cuts`. No projeto Mastopexia
real, 29 cuts individuais viraram 3 cuts mescladas; a contagem de fatias
sub-segundo na timeline final caiu de mais de uma dezena para 4 (resíduo
menor, provavelmente do padding do `remove_media_silence` na emenda entre
clipes — não investigado a fundo nesta sessão, ver `09_MANUTENCAO.md`).
- **Estado:** `resolvido` (a causa principal); a sobra residual do
`remove_media_silence` continua como dívida separada.
> **Aprendizado:** "cortar cada frase desativada" não é a mesma coisa que
> "cortar o trecho desativado" quando frases se sucedem sem conteúdo mantido
> entre elas — a pausa entre duas coisas descartadas também precisa ser
> descartada, e ninguém a cobre por definição se o corte for por frase.
---
### 2026-08-21 — `remove_media_silence` (dB) não pega lacuna sem fala com som real
- **Sintoma:** usuário viu, no projeto Mastopexia real, um trecho de ~1,9s
sem fala (imagem parada antes da tomada começar) que sobreviveu intacto
na timeline final — depois de `apply_voice_actions`, `remove_media_silence`
e `generate_dynamic_subtitles` já terem rodado. Achou que era bug de ordem
no encadeamento das etapas ("corta e depois volta").
- **Investigação:** não era ordem. Extraído o áudio real do trecho
(`ffmpeg -af volumedetect`): `mean_volume -21.4dB`, `max_volume 0.0dB` —
longe do limiar padrão de silêncio (-30dB). Rodado `detect_silence` nos
mesmos limiares do sistema (-30/-25/-20/-16dB): nenhum sinaliza o trecho.
O trecho tem som real (roupa, respiração, ambiente) mas nenhuma palavra —
exatamente o caso que `06-texto-corte-marcador.md` já descrevia
("ausência de fala não é ausência de som"), só que sem ferramenta para
agir sobre ele: `remove_media_silence` só enxerga volume, nunca vai
cortar algo que soa alto mas não tem fala.
- **Correção:** nova função pura `speech_gap_cut_actions()` em
`fcpxml/voice_actions.py` — gera `cut`s a partir dos gaps entre
`words[].start/end` do `_voice_timeline.json` (tempo de fonte, como todo
`VoiceAction`), com a mesma folga por dentro (`padding`) que
`speaker_cut_actions()` já usava. Nova tool MCP `remove_speech_gaps`
(`server_tools/voice.py`, mesmo molde de `remove_speakers`): resolve o
`media_path`, lê a timeline, gera as ações e reaplica via
`handle_apply_voice_actions` — não duplica a lógica de corte no FCPXML.
Deliberadamente não corta a lacuna antes da primeiríssima palavra (pode
ser quase o arquivo inteiro, antes da tomada começar de verdade).
- **Ordem revista:** `apply_voice_actions → remove_speech_gaps →
remove_media_silence → generate_dynamic_subtitles` — a lacuna "sem fala"
some primeiro (cobertura ampla, por transcrição), o que sobra de silêncio
técnico *dentro* da fala é apertado depois.
- **Validação:** `tests/test_voice_actions.py::TestSpeechGapCutActions`
(gap acima/abaixo do limiar, lacuna antes da 1ª palavra nunca cortada,
segmentos com palavras sobrepostas não quebram, timeline vazia). Suíte
completa (1508 testes) roda limpa.
- **Estado:** `resolvido`
> **Aprendizado:** um detector de silêncio por dB nunca vai cobrir "sem fala
> com som" — são categorias diferentes, não uma questão de calibrar o
> limiar. Quando já existe transcrição confiável, ela é a fonte melhor para
> "onde não tem fala": não depende de threshold nenhum, só da própria
> palavra existir ou não naquele instante.
---
### 2026-08-24 — Zoom era `<adjust-transform>` no próprio clipe; FCP exporta como clipe de ajuste
- **Sintoma:** usuário pediu para o zoom parar de mexer diretamente no
clipe da timeline e passar a usar um "adjustment clip" com crop
animado — o jeito como ele já fazia zoom manualmente no FCP.
- **Investigação:** não havia amostra real no projeto para confirmar a
forma exata do XML (`adjust-crop`? um `<clip>` com `<adjustment>` como
`fcpxml/writer/adjustment.py` já fazia para filtros?). O usuário enviou
um `.fcpxmld` exportado pelo próprio FCP com um zoom manual
(`exemplo zoom.fcpxmld`), que revelou a forma real: um `<video ref="...">`
referenciando o efeito nativo `FFAdjustmentEffect` ("Clipe de Ajuste"),
anexado numa lane acima do clipe, com seu **próprio** `<adjust-transform>`
animando `scale` de `1 1` até o pico — não `adjust-crop`, e não o wrapper
`<adjustment>` que `adjustment.py` usa (que, conferido contra o DTD real
da Apple, **não existe** — aquele módulo gera XML inválido; ver dívida
em `09_MANUTENCAO.md`). Cruzado com o DTD oficial (`FCPXMLv1_13.dtd`, uma
cópia local encontrada fora do projeto): `<video>` é `%anchor_item;`
válido sem precisar de asset, e `adjust-transform` é filho direto seu.
- **Correção:** `add_zoom` (extraído para `fcpxml/writer/zoom.py`, deixou
de compartilhar módulo com `change_speed`) agora cria um `<video>`
conectado em vez de animar o clipe base. Isso **simplificou** a lógica
antiga: como o clipe de ajuste composita por cima da imagem já
reenquadrada, não precisa mais ler/preservar rotação, posição ou escala
do clipe original (a classe de teste inteira sobre "preservar
enquadramento" — e o bug histórico #15 que ela cobria — deixou de fazer
sentido); e dois zooms disjuntos no mesmo clipe agora são dois `<video>`
irmãos, não um merge de keyframes num `<adjust-transform>` só.
- **Validação:** os 22 testes de zoom em `test_writer.py` reescritos contra
a nova forma (`clip.find('video').find('adjust-transform')...`), mais
`test_voice_actions_tool.py`. Offset/duration da timeline gerada
conferidos byte a byte contra os números reais do `.fcpxmld` de exemplo
(bateram exatamente). Suíte completa roda limpa.
- **Estado:** `resolvido`
> **Aprendizado:** para decisões de forma exata de XML, um exemplo real
> exportado pelo próprio FCP vale mais que qualquer inferência — a diferença
> entre `adjust-crop`, o wrapper inválido de `adjustment.py` e a forma real
> (`<video ref="FFAdjustmentEffect">`) não dava para cravar sem um dos dois
> (amostra real ou o DTD oficial da Apple, que também foi cruzado aqui).
> Peça o exemplo antes de implementar às cegas.
---
### 2026-08-24 — Revisão de falantes salvava certo, mas a etapa 4 nunca lia o resultado
- **Sintoma:** usuário desmarcou falas de bastidor na tela "Quem fica na edição"
(etapa 3, `SpeakerReviewView`) e clicou "Salvar seleção", mas as falas
desmarcadas continuavam voltando na revisão de frases (etapa 4) e no roteiro
final gerado a partir dela.
- **Causa raiz:** `save_speaker_review` (`fcpxml/speaker_review.py`) e o
`_voice_timeline_clean.json` que ela grava estavam **corretos** — conferido
num projeto real: 37 segmentos na timeline crua, 15 marcados `excluded` na
revisão salva, 22 sobrando no `_clean.json` (37-15=22, bate exato). O bug
estava um passo adiante: `cmd_build_phrase_review`
(`admin/api/review.py`), que monta a etapa 4, abria
`args.get("voice_timeline")` — o arquivo **cru** — direto, sem nunca checar
se existia um `_voice_timeline_clean.json` ao lado. `generate_voice_script`
(`server_tools/voice.py`) e `copyForChat` (`WizardView.swift`) já faziam
essa checagem corretamente; só a etapa 4 ficou de fora.
- **Onde:** `admin/api/review.py::cmd_build_phrase_review`.
- **Por que passou despercebido:** a tela de revisão de falantes em si
funcionava e mostrava "Salvo" — o problema só aparecia num passo seguinte
e sem nenhum erro, então parecia que "a seleção não estava sendo salva"
quando na verdade ela salvava certo e era ignorada mais adiante.
- **Solução adotada:** `cmd_build_phrase_review` agora resolve
`speaker_review.clean_voice_timeline_path(timeline_path)` primeiro e lê
esse arquivo quando ele existe, caindo para o cru só na ausência dele —
mesma checagem que os outros dois pontos já faziam.
- **Aprendizado:** quando existem **múltiplos pontos de leitura** de um
mesmo artefato derivado (aqui: três lugares que podem preferir
`_voice_timeline_clean.json` sobre o cru), adicionar a checagem em um novo
ponto de leitura não é opcional — ela precisa ser replicada em todos, ou o
comportamento diverge silenciosamente conforme o caminho que o app tomar.
Vale grepar por todo lugar que abre o arquivo "canônico" sempre que um
arquivo "_clean"/derivado for introduzido.
- **Estado:** `resolvido` — corrigido em `admin/api/review.py`, suíte
completa (1506 de 1508 testes; as 2 falhas restantes são de ambiente —
WhisperX/torchcodec sem libs de sistema, sem relação com a mudança) e
lint do arquivo alterado limpos.
---
### Entrada 32 — 2026-08-24: `<title>` leva `role`, nunca `videoRole`
**Sintoma:** ao atribuir role de vídeo a legendas geradas (para separar
legendas dinâmicas de convencionais na timeline), a validação contra o DTD
FCPXML v1.13 quebrou com `No declaration for attribute videoRole of element
title`.
**Causa:** no DTD da Apple, `<title>` (`<!ATTLIST title %clip_attrs;>` +
`<!ATTLIST title role CDATA #IMPLIED>`) **não** declara `videoRole`. Esse
atributo existe em `<video>`, `<asset-clip>`, `<clip>` etc., mas não em
títulos. `<title>` usa o atributo genérico `role` (CDATA). Confirmado no
`FCPXMLv1_13.dtd` linhas 566–569.
**Decisão:** legendas dinâmicas e convencionais recebem `role="titles.dinamicas"`
e `role="titles.convencionais"` (sub-roles de `titles`, NUNCA `subtitles.*` —
ver entrada sobre roteamento de captions). O campo de config e o parâmetro dos
geradores chama-se `role` (não `video_role`). `assign_role` (mixin `RolesMixin`)
continua correto para clips/vídeos, pois seta `videoRole` neles — não confundir
os dois caminhos.
**Lição:** antes de setar `videoRole` num elemento qualquer, conferir o DTD:
títulos usam `role`. Teste de regressão em `tests/test_dynamic_subtitles.py`
(`test_titles_carry_title_subrole`) garante `titles.*` e bloqueia `subtitles.*`.
> **Nota de reconciliação:** entradas antigas deste arquivo (2026-08-17)
> afirmavam "nenhum título gerado carrega `role`" e tinham o teste
> `test_titles_carry_no_caption_role`. Aquilo referia-se **especificamente**
> a `role="subtitles.*"` (que roteia o título para a pista de captions e o
> esconde). A regra continua válida: proibido `subtitles.*`. O que mudou é que
> agora aplicamos `role="titles.*"` (sub-role de título, válido no DTD e útil
> para separar dinâmicas de convencionais na timeline). O teste foi renomeado
> para `test_titles_carry_title_subrole` e passa a exigir `titles.*` + bloquear
> `subtitles.*`.
---
### Entrada 33 — 2026-09-22: legenda comum sob a composição dinâmica; regenerar acumulava títulos
**Sintoma:** num corte real (Mastopexia), aos 11s a legenda comum "mamas
também mudam. É" aparecia simultaneamente com a composição dinâmica de
ênfase, poluindo o quadro com texto duplicado. Gerar novamente as legendas
(dinâmica ou convencional) sobre um clipe já legendado empilhava um segundo
conjunto de títulos por cima do anterior em vez de substituí-lo.
**Causa raiz — duas falhas distintas:**
1. **Sem marcação de autoria.** Os três handlers de legenda
(`handle_generate_dynamic_subtitles`, `handle_generate_plain_subtitles`,
`handle_generate_subtitles_by_emphasis`) só *adicionavam* títulos —
nenhum removia o que uma chamada anterior tinha gerado. Sem uma forma de
distinguir "título que este programa gerou" de "título que o editor
inseriu manualmente no FCP", uma regeneração não tinha como saber o que é
seguro apagar.
2. **Janela da legenda de ênfase maior que a fala.** Em
`handle_generate_subtitles_by_emphasis`, o cálculo de fim de bloco usava
os segmentos brutos do Whisper (`data["segments"]`) para decidir até onde
a composição dinâmica se estende — não os spans de ênfase revisados
(`spans`). Um segmento do Whisper cobre a frase inteira; a ênfase cobre só
o trecho grifado. A dinâmica então ficava "seguindo" além do próprio
áudio que a originou, invadindo o intervalo onde a legenda comum já
deveria estar sozinha.
- **Onde:** `code/fcpxml/writer/titles.py` (`TitlesMixin`) e
`code/server_tools/subtitles.py` (os três handlers de geração).
- **Solução adotada:**
- Todo título/composição gerado por este programa carrega uma marca em
`<metadata><md key="com.gart.subtitle.kind" value="dynamic|plain">`
(`mark_generated_subtitle`). Um heurístico de compatibilidade
(`_generated_subtitle_kind`) reconhece a assinatura exata de exports
antigos sem a marca (efeito/uid/start de texto do G-ART + padrão de nome),
para não tratar título manual do editor como "nosso" por engano.
- Cada handler chama `remove_generated_subtitles(el, kinds)` no início,
apagando só os títulos com a marca do próprio tipo que está sendo
regerado — títulos manuais e do outro tipo ficam intactos.
- `generate_dynamic_subtitles` ganhou o parâmetro `hold_between_sentences`
(default `True`, preserva o comportamento anterior nas chamadas normais).
`handle_generate_subtitles_by_emphasis` passa `hold_between_sentences=False`
e usa os `spans` de ênfase revisados como `emphasis_segments` (em vez dos
segmentos brutos do Whisper) — a composição dinâmica agora encerra no fim
real da palavra falada quando o próximo bloco pertence a outra frase, e
nunca ultrapassa a janela de ênfase que a gerou.
- `suppress_plain_under_dynamic` recorta (fatiando o clipe do título, sem
duplicar `text-style`) qualquer legenda comum gerada cujo intervalo caia
dentro de uma composição dinâmica ainda ativa — mesmo que o cálculo de
janela de algum outro caminho volte a divergir no futuro, isso funciona
como rede de segurança contra sobreposição visível.
- **Aprendizado:** um gerador que pode ser chamado de novo sobre a mesma
timeline **precisa** de uma forma de reconhecer sua própria saída anterior
antes de decidir "substituir" — sem isso, "regerar" e "empilhar" são
indistinguíveis. E ao derivar o fim de uma janela temporal a partir de uma
fonte (segmentos do Whisper, spans de ênfase, etc.), confirme que a fonte
escolhida tem a granularidade do fenômeno que está sendo delimitado — usar
a fonte "mais larga disponível" por conveniência cria sobra sistemática.
- **Teste de regressão:**
`code/tests/test_subtitle_overlap_regression.py` — roda o handler real
(`handle_generate_subtitles_by_emphasis`) contra um intervalo de ênfase
seguido de uma lacuna de fala comum, e confere que nenhuma composição
dinâmica sobrepõe uma legenda comum; e que chamar o mesmo handler duas
vezes não duplica títulos gerados nem remove um título manual inserido
entre as duas chamadas.
- **Estado:** `resolvido` — 224 testes das suítes de legenda/writer
passando (incl. o novo regressivo); suíte completa 1540 passando, 8
skipped, 1 falha e 1 erro de ambiente sem relação com a mudança (WhisperX/
`extract_pitch` ausente, torchcodec sem libs de sistema); lint dos arquivos
alterados limpo.
---
### Entrada 34 — 2026-09-22: `<adjustment>` inválido no DTD e `WHISPERX` órfão inflando o lint
**Sintoma 1:** `fcpxml/writer/adjustment.py` (`ClipDeAjuste`, código de uma
sessão anterior não commitado) montava
`<clip><adjustment><filter-video .../></adjustment></clip>` para camadas de
ajuste. Nada usava o módulo ainda (sem chamada em `server_tools`/
`admin/api`), mas ficava pronto para alguém reusar do jeito errado.
**Causa 1:** o DTD real da Apple (`FCPXMLv1_13.dtd`) não define nenhum
elemento `<adjustment>`. A produção real de `<clip>` é
`(note?, %timing-params;, %intrinsic-params;, (spine|(%clip_item;)|caption)*,
(%marker_item;)*, audio-channel-source*, (%video_filter_item;)*,
filter-audio*, metadata?)` — ou seja, `filter-video`/`filter-audio` são
filhos diretos do `<clip>`, sem wrapper, e nessa ordem (vídeo antes de
áudio).
**Solução 1:** `ClipDeAjuste.criar()` agora anexa os filtros direto no
`<clip>`, ordenados com vídeo antes de áudio
(`sorted(filtros, key=lambda f: f.tag != "filter-video")`). Teste de
regressão novo: `tests/test_writer_adjustment.py` (sem wrapper, ordem
correta, um `<effect>` por `uid` em `resources`).
**Sintoma 2 (achado ao investigar o mesmo módulo):** um `ruff check .
--exclude docs/` rodado manualmente no início desta sessão acusou **510
erros** — muito acima do que a suíte normalmente reporta.
**Causa 2:** `code/WHISPERX` era uma pasta `.git` solta de **2,6 GB** dentro
de `code/` (não um submodule registrado — sem `.gitmodules`), contendo
cópias/backups congelados do próprio projeto, incluindo uma cópia inteira e
antiga de `fcp-mcp-server-main` dentro de si mesma. O `pyproject.toml` já
excluía `WHISPERX/` do lint por padrão (`[tool.ruff] exclude = ["docs/",
"WHISPERX/"]`), mas passar `--exclude docs/` na linha de comando
**sobrescreve** esse `exclude` em vez de complementá-lo — foi assim que o
lint passou a varrer os 2,6 GB de código velho lá dentro. Confirmado por
grep que só 3 arquivos no código ativo referenciam "WHISPERX", todos em
comentários explicativos (`fcpxml/diarize.py`, `tests/test_diarize.py`,
`admin/api/shared.py`) — nenhum import ou caminho real dependia da pasta.
**Solução 2:** pasta movida para `~/Archives/G-ART-WHISPERX-backup` (fora do
workspace git), copiada com `rsync -a --no-perms` e conferida com
`diff -rq` antes de remover o original. `WHISPERX/` também saiu do
`exclude` do ruff em `code/pyproject.toml` (não faz mais sentido excluir um
caminho que não existe mais em `code/`).
**Aprendizado:** (1) um wrapper de elemento "que faz sentido conceitualmente"
não substitui checar o DTD real antes de escrever o gerador — o padrão do
projeto (`dtd.py`, DTDs em `bm/*/FCPXMLv1_13.dtd`) existe exatamente para
isso. (2) uma flag de linha de comando como `--exclude` em ferramentas de
lint tipicamente **substitui** a config do projeto, não a estende — rodar
`ruff check .` sem flags (herdando `pyproject.toml`) é o comando correto
para refletir o gate real; qualquer variação manual com `--exclude` pode
mentir sobre o estado do lint. (3) uma pasta de backup improvisada dentro do
diretório ativo do projeto (mesmo que "só para não perder nada") é dívida
que cresce sem ninguém perceber — 2,6 GB não apareceram de uma vez.
**Estado:** `resolvido` — `tests/test_writer_adjustment.py` (3 testes)
passando; `admin/` trazido ao lint gate no mesmo commit (ver
`09_MANUTENCAO.md` §2.3); suíte completa 1543 passando, 8 skipped, 1 falha
+ 1 erro pré-existentes de outro trabalho em andamento (sem relação com
esta correção).
### Entrada 35 — 2026-09-22: chunk grande demais derrubava a indexação RAG inteira
**Sintoma:** `admin/update_rag.command` (primeira indexação completa do
G-ART, banco `rag_gart` recém-provisionado) morria sempre no mesmo ponto com
`requests.exceptions.HTTPError: 500 Server Error` na chamada ao Ollama —
sempre logo após imprimir `code/fcpxml/export.py`, ou seja, no arquivo
seguinte na ordem alfabética.
**Causa raiz:** `code/fcpxml/font_metrics.py` é uma tabela de larguras de
glifo (`METRICS = {...}`), texto extremamente denso em tokens (muitos
números/pontuação curtos) — um chunk de ~4900 caracteres (dentro do limite
`CHUNK_MAX_CHARS = 5000`) virou 2653 tokens no tokenizer do
`nomic-embed-text`, estourando o contexto de 2048 tokens do servidor Ollama
local (`llama.cpp`: "input length exceeds the context length"). Reproduzido
isolando o arquivo e chamando `/api/embeddings` chunk a chunk — 6 dos 9
chunks falhavam. `CHUNK_MAX_CHARS` mede caracteres, não tokens; assume
implicitamente ~1 token por poucos caracteres, o que não vale para conteúdo
não-prosa (tabelas numéricas, JSON denso).
Segundo problema, apontado por que a primeira tentativa não recuperou nada:
`admin/update_rag.py::index()` roda a varredura inteira (centenas de
arquivos) em **uma única transação**, com `commit()` só no fim e
`rollback()` em qualquer exceção — um único chunk problemático em um único
arquivo descartava a indexação inteira, mesmo que os outros 300+ arquivos
já tivessem embedado e inserido com sucesso.
**Solução:** `_embed()` agora detecta essa resposta específica do Ollama
(`ChunkTooLarge`, checado por `500` + `"context length"` no corpo) e o loop
principal captura essa exceção por chunk, pula só aquele chunk (aviso em
stderr) e continua o arquivo — sem abortar a transação. Não trunca nem
reduz `CHUNK_MAX_CHARS` globalmente (afetaria todo o corpus por causa de
poucos arquivos atípicos); a lacuna fica só nos poucos chunks realmente
grandes demais, e o resto do arquivo ainda fica pesquisável.
**Aprendizado:** um limite de chunk em caracteres é uma aproximação, não uma
garantia de contexto — arquivos de dados brutos (tabelas, mapeamentos
numéricos, JSON/CSV embutido em `.py`) tokenizam bem mais denso que prosa ou
código comum e podem violar o limite do modelo mesmo dentro do teto de
caracteres. Uma indexação em lote sobre centenas de arquivos não deve ficar
tudo-ou-nada numa única transação: uma falha isolada e recuperável (chunk
específico, arquivo específico) deve ser contida ali, não descartar o
trabalho inteiro já validado.
**Estado:** `resolvido` — indexação completa rodou até o fim: 304 arquivos,
1702 chunks, 0 removidos. Também nesta sessão: criada a pasta `rag/` na raiz
(schema, busca híbrida `search.py`/`search_gart.sh`, `SETUP.md`) — ver
`rag/README.md` para a divisão de responsabilidades com `admin/update_rag.py`.
---
### Entrada 36 — 2026-09-23: `.gitignore` escondia `fcpxml/models/` inteiro do git
**Sintoma:** ao investigar por que um `git diff` de um arquivo recém-editado
(`fcpxml/models/timeline.py`, durante a correção da Entrada 34) não mostrava
nada, `git status` também não listava o arquivo como modificado nem como
untracked — como se ele simplesmente não existisse para o git.
**Causa:** `.gitignore` tinha a regra solta `models/` (comentada como
"WhisperX models cache", pensada para ignorar o cache de ~11 GB de modelos
Whisper baixados em `code/models/`). Uma regra sem `/` inicial no
`.gitignore` casa com **qualquer diretório com esse nome em qualquer
profundidade** — não só `code/models/`, mas também `code/fcpxml/models/`, o
pacote de data classes (`TimeValue`, `Clip`, `Timeline`, `Marker`, etc.) que
sustenta todo o engine. Confirmado: `git ls-tree -r HEAD` não tem nenhum
`fcpxml/models.py` nem `fcpxml/models/` em nenhum commit do histórico — o
pacote inteiro (1.234 linhas, 7 módulos) só existia em disco, sem nenhuma
proteção de versionamento, desde que a divisão de `models.py` em pacote foi
feita (sessão anterior, nunca commitada).
**Risco:** qualquer operação que limpa arquivos não rastreados
(`git clean -fd`, reinstalar do zero, trocar de máquina via `git clone`)
apagaria essa base sem chance de recuperação — nenhum commit para reverter.
**Solução:** regra trocada para `/code/models/` (ancorada na raiz do repo,
só o cache real), preservando `whisper/` (sem uso hoje, mas inofensiva) e
tudo mais. Confirmado com `git check-ignore -v`: `fcpxml/models/timeline.py`
não é mais ignorado; `code/models/models--Systran--faster-whisper-base`
continua ignorado. `fcpxml/models/` passou a aparecer como `??` no
`git status` — visível, pronto para ser commitado quando o dono do trabalho
revisar.
**Aprendizado:** regra de `.gitignore` sem `/` inicial (ex.: `models/`) casa
em qualquer profundidade da árvore — é fácil escrever pensando só no caso
que motivou a regra (um cache na raiz) e esquecer que o mesmo nome de pasta
pode existir, com sentido completamente diferente, dentro do código-fonte.
Regra de bolso: nomes de pasta genéricos (`models/`, `build/`, `cache/`,
`data/`) no `.gitignore` deveriam quase sempre vir ancorados (`/caminho/
exato/`), a menos que a intenção seja mesmo ignorar toda ocorrência do nome
em qualquer lugar da árvore.
**Estado:** `resolvido` — regra corrigida, `fcpxml/models/` confirmado
visível ao git (não commitado ainda; fica para quem já está com esse
trabalho em andamento decidir quando commitar). Nenhum código alterado,
só o `.gitignore`.
+63 -27
View File
@@ -7,7 +7,7 @@ Este é o documento de rota. Os outros descrevem o que **é**; este diz o que
**fazer** e por onde começar quando chega uma implementação, uma melhoria ou **fazer** e por onde começar quando chega uma implementação, uma melhoria ou
uma correção. uma correção.
Última varredura: 2026-08-19 · 1.466 testes · lint zerado Última varredura: 2026-09-22 · 1.543 testes passando (+1 falha pré-existente em `test_forced_align.py` e +1 erro pré-existente em `test_refine_voice_timeline_tool.py`, ver §2.6) · lint zerado em `code/` e em `admin/` (fora de server.py/ai_edit.py/fcpxml/analise.py, pré-existentes — outro trabalho em andamento na branch)
--- ---
@@ -33,42 +33,76 @@ vai para `fcpxml/`.
Ordenado por quanto atrapalha, não por esforço. Ordenado por quanto atrapalha, não por esforço.
### 2.1 A etapa 6 ignora a revisão de ênfases ### 2.1 `resolve_actions` não tolera margem no encosto de zoom/marker contra um corte
O usuário lapida as frases na etapa 5, o `_phrase_review.json` é gravado — e a Um `zoom`/`marker` cuja borda cai exatamente em cima do `start`/`end` de um
etapa 6 ainda processa como antes. Falta ligar: **zoom e legenda dinâmica só `cut` é descartado como "apontando para material cortado" — mesmo quando a
nas frases de ênfase, legenda comum no resto**. É a continuação natural do intenção era ficar bem ao lado. Contornado manualmente no projeto Mastopexia
trabalho da etapa 5 e o item mais valioso da lista. (recuando as bordas na mão); a correção estrutural é dar a `resolve_actions`
→ `MacApp/Sources/WizardView.swift` (`finalizeProcessing`), `admin/api/subtitles.py`, uma margem de tolerância (meio frame) antes de considerar uma ação "dentro"
`fcpxml/phrase_review.py` (`emphasis_spans` já é produzido e ninguém consome). do corte. → `fcpxml/voice_actions.py` (`resolve_actions`/`shift_after_cuts`),
`05_EXPERIENCIAS.md` #27.
### 2.2 Offset de ~400 ms no timing por palavra ### 2.2 `MacApp/` não tem teste automatizado
O faster-whisper sem alinhamento forçado erra o início de cada palavra em
~0,4 s. Isso desloca zoom, corte e `gap_before` de uma vez. Há paliativo
aplicado por projeto; a correção estrutural é ligar o **WhisperX** (ou
alinhamento equivalente) em `transcribe.py`, o que levaria o erro para ~30 ms.
Custo real: regerar todos os `_transcript.json` e `_voice_timeline.json`
existentes. → `05_EXPERIENCIAS.md` #14, estado `parcialmente resolvido`.
### 2.3 `MacApp/` não tem teste automatizado
5.500 linhas de Swift sem uma asserção. A rede hoje é o harness manual (§4) e 5.500 linhas de Swift sem uma asserção. A rede hoje é o harness manual (§4) e
o olho do usuário. Não é para sair criando suíte de UI — mas lógica pura que o olho do usuário. Não é para sair criando suíte de UI — mas lógica pura que
foi parar na camada de tela (cálculo de trim, mapeamento de tempo) deveria foi parar na camada de tela (cálculo de trim, mapeamento de tempo) deveria
descer para o Python, onde já existe rede. descer para o Python, onde já existe rede.
### 2.4 `admin/` fica fora do lint ### 2.3 ~~`admin/` fica fora do lint~~ — resolvido em 2026-09-22
`run_after_fix.sh` roda o ruff de dentro de `code/`, então `admin/` — 1.751 `run_after_fix.sh` agora roda um segundo passo (`ruff check --config
linhas de código que o app depende para funcionar — nunca é verificado. pyproject.toml ../admin/`) com a mesma config do engine. Precisou de
Incluir mexe no gate, então é decisão consciente, não esquecimento. `# noqa: E402` em 6 imports de `admin/models_api.py`/`admin/models_gui.py`
(padrão `sys.path.insert` antes do import local, convenção já usada no
projeto). Lint de `admin/` está zerado.
### 2.5 Confirmações visuais pendentes no FCP ### 2.4 Confirmações visuais pendentes no FCP
Várias entradas do `05_EXPERIENCIAS.md` estão marcadas como resolvidas *no XML* Várias entradas do `05_EXPERIENCIAS.md` estão marcadas como resolvidas *no XML*
— testes verdes, DTD válido — mas **pendentes de importação real no Final Cut**. — testes verdes, DTD válido — mas **pendentes de importação real no Final Cut**.
XML válido não é o mesmo que XML que renderiza como o esperado. Ao mexer em XML válido não é o mesmo que XML que renderiza como o esperado. Ao mexer em
legenda, zoom ou keyframe, a confirmação final é abrir no FCP. legenda, zoom ou keyframe, a confirmação final é abrir no FCP.
### 2.6 Submódulo `WHISPERX` com conteúdo modificado e não commitado ### 2.5 `remove_media_silence` deixa fatias sub-segundo nas emendas entre clipes
Está fora dos commits de propósito, porque ninguém verificou o que mudou lá Mesmo depois de corrigir o merge de cortes consecutivos (`05_EXPERIENCIAS.md`
dentro. Precisa ser olhado e resolvido — ou commitado, ou revertido. #28), sobraram 4 clipes de 0,07-0,23s no projeto Mastopexia real, todos bem
na emenda entre dois clipes vizinhos — mesma família do #6 (clipe-fantasma de
1 frame por padding sem vizinho na borda), mas não confirmado se é a mesma
causa raiz. Não investigado a fundo ainda.
→ `fcpxml/writer/cut.py` (`cut_clip_ranges`, `min_keep_seconds`), padding do
`remove_media_silence`.
### 2.6 ~~`fcpxml/writer/adjustment.py` gerava um wrapper `<adjustment>` inválido~~ — resolvido em 2026-09-22
`ClipDeAjuste` embrulhava filtros num `<clip><adjustment>...</adjustment></clip>`,
que não existe no DTD real da Apple. Corrigido para anexar
`filter-video`/`filter-audio` direto como filhos do `<clip>` (na ordem que o
DTD exige: vídeo antes de áudio). Teste de regressão em
`tests/test_writer_adjustment.py`. Segue sem uso em `server_tools`/`admin/api`
— só deixou de estar pronto pra alguém reusar do jeito errado.
→ `05_EXPERIENCIAS.md` #34.
### 2.7 `test_refine_voice_timeline_tool.py` quebrado: `voice_timeline.extract_pitch` ausente
`TestRefineVoiceTimelineHandler::test_max_zooms_caps_the_list` tenta
`monkeypatch.setattr(vt, "extract_pitch", ...)` mas `fcpxml/voice_timeline.py`
não tem mais (ou nunca teve, nesta branch) essa função. Pertence ao trabalho
de análise de voz já em andamento nesta branch (`voice_timeline.py`
modificado, não commitado) — não investigado a fundo, só registrado aqui
para não se perder.
→ `fcpxml/voice_timeline.py`, `tests/test_refine_voice_timeline_tool.py`.
### 2.8 ~~Submódulo `WHISPERX` com conteúdo modificado e não commitado~~ — resolvido em 2026-09-22
Não era um submódulo git registrado (sem `.gitmodules`) — era uma pasta
`.git` solta de 2,6 GB dentro de `code/`, com cópias/backups congelados do
próprio projeto (`WHISPERX_backup_88476/`, uma cópia inteira e antiga de
`fcp-mcp-server-main`). Só 3 referências no código ativo, todas em
comentários (`fcpxml/diarize.py`, `tests/test_diarize.py`,
`admin/api/shared.py`), nenhum import ou caminho dependia dela. Além do
peso morto, ela também inflava qualquer lint rodado com `--exclude`
explícito (que sobrescreve o `exclude` do `pyproject.toml`) — foi assim que
um `ruff check . --exclude docs/` chegou a acusar 510 erros, quase todos
dentro dela. Movida para `~/Archives/G-ART-WHISPERX-backup` (fora do
workspace git), copiada e verificada (`diff -rq`) antes de remover o
original. `WHISPERX/` também saiu do `exclude` do ruff em
`code/pyproject.toml` — não faz mais sentido excluir um caminho que não
existe mais dentro de `code/`.
--- ---
@@ -98,8 +132,10 @@ quanto arquivo gigante.
## 4. Checklist antes de dar algo por pronto ## 4. Checklist antes de dar algo por pronto
```bash ```bash
cd code && ./Engine/run_after_fix.sh # lint zerado + 1.466 testes cd code && ./Engine/run_after_fix.sh # lint zerado + 1.498 testes
admin/run_app.command # se mexeu no app (padrão de revisão) admin/run_app.command # se mexeu no app (padrão de revisão)
admin/run.command # app + atualização incremental da RAG
rag/search_gart.sh "consulta" # busca híbrida no índice RAG (ver rag/README.md)
``` ```
E, além do script: E, além do script:
@@ -125,7 +161,7 @@ E, além do script:
| FCP recusa o arquivo ao importar | `id` inválido, ordem de filhos, timebase | `writer/validation.py`, `dtd.py` | | FCP recusa o arquivo ao importar | `id` inválido, ordem de filhos, timebase | `writer/validation.py`, `dtd.py` |
| Título importa mas não aparece | Template/uid Motion que não resolve | `writer/titles.py` | | Título importa mas não aparece | Template/uid Motion que não resolve | `writer/titles.py` |
| Corte no lugar errado | Tempo pós-corte usado como se fosse original | `voice_actions.py` (`shift_after_cuts`) | | Corte no lugar errado | Tempo pós-corte usado como se fosse original | `voice_actions.py` (`shift_after_cuts`) |
| Zoom no lugar errado | Idem, ou offset de timing do Whisper | §2.2 | | Zoom/marker sumindo perto de um corte | Borda encostando exatamente no `cut` | §2.1 |
| Legenda sobrepondo | Layout ou conteúdo antigo no arquivo | `collision.py`, `text_layout.py` | | Legenda sobrepondo | Layout ou conteúdo antigo no arquivo | `collision.py`, `text_layout.py` |
| "Ênfase" apontando para palavra à toa | Falta renormalizar após o corte | `refine_voice_timeline` | | "Ênfase" apontando para palavra à toa | Falta renormalizar após o corte | `refine_voice_timeline` |
| App diz que falta librosa/pyannote | `uv run` com cwd errado | `PythonBridge.swift` (§3 do doc 08) | | App diz que falta librosa/pyannote | `uv run` com cwd errado | `PythonBridge.swift` (§3 do doc 08) |
+269
View File
@@ -0,0 +1,269 @@
# 10 - Mapa de Reestruturacao de Funcionalidades
> Escopo: roteiro pratico para reorganizar o codigo sem quebrar o produto.
> Baseado na varredura de 2026-08-24 sobre engine Python, ponte do app,
> ferramentas MCP e app SwiftUI.
## 1. Diagnostico rapido
O projeto ja tem uma arquitetura-alvo correta: `fcpxml/` como engine puro,
`server.py` + `server_tools/` como camada MCP, `admin/` como ponte JSON-lines
do app e `MacApp/` como interface. A melhoria agora nao e "reinventar" a
arquitetura, e reduzir os pontos onde as responsabilidades ainda se misturam.
### Pontos fortes
- Engine Python bem testado e com regra clara: logica de timeline fica em
`fcpxml/`.
- `writer/` ja foi quebrado em mixins por assunto, preservando API publica.
- `server.py` funciona como composition root e usa dispatch por dicionario.
- Documentacao interna registra decisoes, armadilhas e padroes do projeto.
- Fluxos criticos tem testes extensos em `code/tests/`.
### Dores atuais
- Alguns arquivos voltaram a virar centros de gravidade:
- `server_tools/voice.py` (~999 linhas)
- `server_tools/subtitles.py` (~760 linhas)
- `MacApp/Sources/WizardView.swift` (~979 linhas)
- `MacApp/Sources/TranscriptionView.swift` (~843 linhas)
- `fcpxml/model_manager.py` (~748 linhas)
- `admin/` e `server_tools/` expõem fluxos parecidos por caminhos diferentes,
o que aumenta risco de uma funcionalidade existir no MCP e faltar no app.
- ~~`admin/` ainda fica fora do lint principal~~ — resolvido na Fase 0
(2026-09-22): `admin/` entrou no gate de `run_after_fix.sh`.
- ~~`WHISPERX` e backups aparecem junto da base ativa~~ — resolvido na
Fase 0 (2026-09-22): movido para fora do workspace git.
- O app SwiftUI quase nao tem rede automatizada; compilar nao garante que uma
tela abre.
## 2. Mapa de dominios desejado
```text
Produto
MacApp/ Interface e experiencia do usuario
admin/ Ponte JSON-lines do app
server.py + server_tools/ Entrada MCP
Engine
fcpxml/models/ Dados e contratos
fcpxml/parser.py FCPXML -> objetos
fcpxml/writer/ Escrita e edicao de XML
fcpxml/voice_* Analise e decisoes por voz
fcpxml/text_layout.py Layout de legendas
fcpxml/model_manager.py Catalogo, configs e modelos
Suporte
tests/ Rede automatizada
Engine/docs/ Decisoes e operacao
examples/ Fixtures de uso
Legado / referencia
WHISPERX/ Deve sair do caminho ativo ou virar referencia clara
```
Regra de organizacao: uma funcionalidade nasce no engine, depois ganha duas
portas finas se necessario: uma tool MCP em `server_tools/` e um comando do app
em `admin/api/`.
## 3. Reestruturacao por fases
### Fase 0 - Higiene antes de mexer — `concluída em 2026-09-22`
Objetivo: reduzir ruido e proteger a base antes de mover codigo.
- ~~Decidir o destino de `code/WHISPERX`~~ — não era submodule (sem
`.gitmodules`), era 2,6 GB de backups órfãos do próprio projeto sem
nenhuma referência ativa. Movido para `~/Archives/G-ART-WHISPERX-backup`
(fora do workspace git), copiado com `rsync` e conferido com `diff -rq`
antes de remover o original. Detalhe: essa pasta também inflava qualquer
`ruff check --exclude docs/` manual (a flag sobrescrevia o `exclude` do
`pyproject.toml`, que já ignorava `WHISPERX/`) — ver `05_EXPERIENCIAS.md`
#34.
- ~~Incluir `admin/` em uma checagem de lint separada antes de colocar no
gate obrigatório~~ — checado com a config real do projeto (não o default
do ruff): só 6 erros, todos `E402` por `sys.path.insert` antes de import
local. Resolvido com `# noqa: E402` (convenção já usada no projeto) e
`admin/` entrou direto no gate obrigatório (`run_after_fix.sh`, passo
2/3), sem precisar de etapa intermediária "separada".
- Corrigido de quebra: `fcpxml/writer/adjustment.py` gerava um `<adjustment>`
inválido no DTD — não estava no escopo original da Fase 0, mas surgiu na
investigação e era pequeno o bastante para resolver junto (ver
`05_EXPERIENCIAS.md` #34).
- Atualizados: `02_MODULES.md` (versão, linhas de `writer/`, módulos novos
`builders.py`/`adjustment.py`/`analise.py`/`transcription/`),
`09_MANUTENCAO.md` (contagem de testes/lint, itens §2.3/§2.6/§2.8
resolvidos, novo item §2.7 registrando `test_refine_voice_timeline_tool`).
- **Pendente, não fechado nesta rodada:** "documentar oficialmente quais
pastas são produto ativo, legado e backup" como um documento à parte —
o que existia de fato como "legado" (`WHISPERX`) já foi resolvido, não
sobrou candidato claro para justificar um novo documento agora.
Entrega obtida: lint de `admin/` no gate, `code/writer/adjustment.py`
correto e testado, ~2,6 GB fora do caminho ativo, docs sincronizados com o
código atual.
### Fase 1 - Contratos entre camadas
Objetivo: impedir que MCP, app e engine driftam entre si.
- Criar um registro unico de capacidades, por exemplo:
- nome interno da funcionalidade;
- funcao pura do engine;
- handler MCP, se existir;
- comando `admin`, se existir;
- tela Swift, se existir;
- testes associados.
- Adicionar teste que detecta comandos importantes presentes no MCP mas ausentes
na ponte do app, quando fizer sentido.
- Padronizar o retorno dos comandos `admin/api`: `ok`, `path`, `message`,
`error`, `unchanged`, `artifacts`.
Entrega esperada: mapa vivo de funcionalidades e menos "funciona no Claude,
nao aparece no app".
### Fase 2 - Dividir `server_tools/voice.py`
Objetivo: separar o fluxo de voz por etapas reais do produto.
Divisao sugerida:
```text
server_tools/voice/
__init__.py Reexporta TOOLS e HANDLERS
analysis.py analyze_voice_features, build_voice_timeline
speakers.py diarize_media, remove_speakers
refinement.py refine_voice_timeline, remove_speech_gaps
actions.py apply_voice_actions
local_ai.py generate_voice_script
config.py get/save_voice_analysis_config
```
Cuidados:
- Manter os nomes publicos reexportados para nao quebrar testes/imports.
- Mover em uma etapa por arquivo, rodando testes de voz a cada passo.
- Nao mover regra de negocio para `server_tools/voice/`; se aparecer regra
nova, ela deve descer para `fcpxml/voice_*`.
Testes minimos: `test_voice_actions.py`, `test_voice_actions_tool.py`,
`test_voice_timeline.py`, `test_voice_timeline_tool.py`, `test_diarize.py`,
`test_voice_features.py`.
### Fase 3 - Separar `fcpxml/model_manager.py`
Objetivo: reduzir mistura entre catalogo, download, configuracao e estado.
Divisao sugerida:
```text
fcpxml/model_manager/
__init__.py API publica atual
catalog.py models.json, recomendados, metadata
storage.py diretorios, instalados, migracao
download.py download/cancel/progresso
transcription_config.py modelo selecionado, idioma
voice_config.py analise de voz, silencio, legendas
```
Cuidados:
- Preservar imports atuais via `__init__.py`.
- Separar funcoes puras de funcoes com I/O para facilitar teste.
- Nao acoplar config do app a nomes de tela Swift.
Testes minimos: `test_models.py`, `test_models_api.py` se existir,
`test_voice_analysis_config.py`, `test_project_config.py`.
### Fase 4 - Reorganizar o Assistente SwiftUI
Objetivo: tornar o fluxo de 7 etapas legivel e testavel por partes.
Divisao sugerida:
```text
MacApp/Sources/Wizard/
WizardView.swift Casca, navegacao e estado global
WizardState.swift Estado do fluxo e canAdvance
ProjectStepView.swift
TranscribeStepView.swift
VoiceAnalysisStepView.swift
AIScriptStepView.swift
ReviewStepHost.swift
ProcessStepView.swift
DoneStepView.swift
```
Boas praticas para essa fase:
- Extrair primeiro views pequenas, sem alterar comportamento.
- Depois extrair calculos puros de `canAdvance`, nomes de arquivos e selecao
de artefatos para tipos testaveis.
- Usar harness manual documentado em `08_APP_MACOS.md` para abrir as telas
tocadas.
Entrega esperada: cada etapa do wizard vira um arquivo com responsabilidade
unica.
### Fase 5 - Unificar validacao e saida da ponte `admin/`
Objetivo: deixar os comandos do app tao disciplinados quanto os handlers MCP.
- Criar helpers de path/output equivalentes aos de `server_tools/_shared`,
ou mover helpers comuns para uma camada compartilhada que nao saiba de MCP.
- Trocar chamadas diretas a `server.generate_output_path` por helper de dominio
que nao puxe `server.py` quando a ponte so precisa de path.
- Adicionar lint de `admin/` ao fluxo de manutencao depois de corrigir erros
existentes.
Entrega esperada: ponte mais fina, menos import acidental de transporte MCP.
### Fase 6 - Tests e gates de seguranca
Objetivo: fazer a reorganizacao ser barata de continuar.
- Criar testes de "arquitetura":
- `fcpxml/` nao importa `server`, `server_tools` nem `admin`;
- handlers MCP sempre retornam via `_text_result`;
- comandos `admin` retornam JSON no formato padrao.
- Criar teste de import publico para garantir que reexports antigos continuam.
- Para SwiftUI, manter harnesses por tela critica ate existir um build mais
estruturado.
Entrega esperada: mover arquivos deixa de ser aposta.
## 4. Prioridade recomendada
1. Fase 0: limpar mapa ativo vs legado.
2. Fase 1: criar registro de capacidades.
3. Fase 2: dividir voz em `server_tools`.
4. Fase 5: fortalecer `admin/`.
5. Fase 4: quebrar `WizardView`.
6. Fase 3: dividir `model_manager.py`.
7. Fase 6: ampliar gates conforme as fases estabilizam.
Motivo: primeiro se reduz incerteza, depois se separa o arquivo que mais muda
no fluxo novo de voz, e so entao se mexe nas telas maiores.
## 5. Checklist para cada refatoracao
- Mover sem mudar comportamento na primeira passada.
- Preservar API publica com reexports.
- Rodar testes focados depois de cada movimento.
- Rodar `cd code && ./Engine/run_after_fix.sh` antes de concluir.
- Se mexeu em `MacApp/`, compilar e abrir a tela afetada.
- Atualizar docs no mesmo commit.
- Registrar aprendizado em `05_EXPERIENCIAS.md` quando houver bug real.
## 6. Principios de boas praticas para este projeto
- Engine puro: sem MCP, sem Swift, sem JSON de tela.
- Camadas de entrada finas: validam, chamam engine, formatam resposta.
- Tempo de timeline sempre racional (`TimeValue`), exceto metricas de audio e
UI onde segundos float sao apenas apresentacao/analise.
- Original nunca e sobrescrito.
- XML sempre entra por `safe_xml.py`.
- Dependencias opcionais continuam lazy.
- Arquivo grande so e problema quando contem varios assuntos.
- Toda funcionalidade importante deve ter dono, porta MCP/app documentada e
teste correspondente.
+8 -3
View File
@@ -23,14 +23,19 @@ echo "==> [G-ART] Validação pós-correção iniciada..."
echo " Diretório: $REPO_ROOT" echo " Diretório: $REPO_ROOT"
echo "" echo ""
echo "==> 1/2 Lint (ruff) — deve passar com ZERO erros" echo "==> 1/3 Lint do engine (ruff, code/) — deve passar com ZERO erros"
# A flag --exclude sobrescreve o exclude declarado em pyproject.toml # A flag --exclude sobrescreve o exclude declarado em pyproject.toml
# (que já ignora docs/ e WHISPERX/). Rode sem flag para herdar a config. # (que já ignora docs/). Rode sem flag para herdar a config.
uv run ruff check . uv run ruff check .
echo " Lint OK ✓" echo " Lint OK ✓"
echo "" echo ""
echo "==> 2/2 Testes (pytest) — todos devem passar" echo "==> 2/3 Lint da ponte (ruff, admin/) — mesma config do engine"
uv run ruff check --config pyproject.toml ../admin/
echo " Lint OK ✓"
echo ""
echo "==> 3/3 Testes (pytest) — todos devem passar"
uv run pytest tests/ -v uv run pytest tests/ -v
echo "" echo ""
Submodule code/WHISPERX deleted from c9ed3cc6bd
+1 -1
View File
@@ -63,7 +63,7 @@ Regras (siga rigorosamente):
3. ESCOLHER A MELHOR TOMADA de cada frase quando há repetições: mantenha a mais limpa e corte as outras (cut cobrindo a frase inteira). 3. ESCOLHER A MELHOR TOMADA de cada frase quando há repetições: mantenha a mais limpa e corte as outras (cut cobrindo a frase inteira).
4. CORTE (kind "cut"): para REMOVER uma frase, cubra ela inteira (start..end = início..fim da frase). Para APARAR só uma hesitação no começo ou fim, corte só da borda até a palavra (corte de meia frase é ambíguo — passe de 60% e apaga a linha toda). Nunca corte o silêncio entre falas. 4. CORTE (kind "cut"): para REMOVER uma frase, cubra ela inteira (start..end = início..fim da frase). Para APARAR só uma hesitação no começo ou fim, corte só da borda até a palavra (corte de meia frase é ambíguo — passe de 60% e apaga a linha toda). Nunca corte o silêncio entre falas. Quando a borda do corte encosta em fala mantida (não em silêncio puro), recue ~0,15-0,25s para dentro do corte nos dois lados — start ~0,2s DEPOIS do fim real da última palavra mantida, end ~0,2s ANTES do início real da próxima palavra mantida — senão o corte soa seco, engolindo a palavra antes de terminar de soar. Isso vale também pro início/fim do vídeo (ar morto antes da primeira palavra e depois da última).
5. ZOOM (kind "zoom"): só em palavra de CONTEÚDO bem enfatizada (emphasis alto, não artigo). params.scale entre 1.0 e 3.0 (padrão 1.3 se omitido). Posicione em torno da palavra, segurando até o fim da frase. 5. ZOOM (kind "zoom"): só em palavra de CONTEÚDO bem enfatizada (emphasis alto, não artigo). params.scale entre 1.0 e 3.0 (padrão 1.3 se omitido). Posicione em torno da palavra, segurando até o fim da frase.
+8 -1
View File
@@ -644,7 +644,14 @@ DEFAULT_SILENCE_CONFIG: dict = {
# Seconds a quiet stretch must last before it's a cut candidate. # Seconds a quiet stretch must last before it's a cut candidate.
"min_silence": 0.5, "min_silence": 0.5,
# Seconds left inside each cut so speech never gets clipped at the edges. # Seconds left inside each cut so speech never gets clipped at the edges.
"padding": 0.05, # 0.2s matches the breathing-room convention for phrase-boundary cuts
# (see editar-por-voz/criterios/06-texto-corte-marcador.md) — a silence
# span this tool finds is often the natural breath before a new
# sentence, not just editing slop, and 0.05s shaved that breath down to
# almost nothing (real case: Mastopexia project, the pause before "Com"
# went from 0.567s to 0.1s across the cut, landing the next clip only
# 5ms after the word instead of a natural pause before it).
"padding": 0.2,
} }
+120
View File
@@ -0,0 +1,120 @@
"""
Data models for Final Cut Pro FCPXML structures.
Provides a clean Python interface for working with Final Cut Pro timelines,
clips, markers, and other elements.
Era um módulo de 1.091 linhas com seis famílias de modelo dentro. Agora cada
família tem seu arquivo, e este pacote reexporta tudo — `from .models import
TimeValue` segue valendo em todo o projeto, inclusive para os nomes com
underscore que o writer e a suíte já usavam.
enums tipos e cores de marcador, transições, ritmo
timing TimeValue (fração racional) e Timecode
timeline clipes, marcadores, lanes, projeto
planning rough cut, ritmo, montagem
qc achados de QC e resultado de validação
subtitles paleta e look das legendas dinâmicas
"""
from .enums import (
_MAX_MARKER_TYPE_LENGTH,
MARKER_XML_TAGS,
FlashFrameSeverity,
MarkerColor,
MarkerType,
PacingCurve,
PacingStyle,
TransitionType,
ValidationIssueType,
)
from .planning import (
MontageConfig,
PacingConfig,
RoughCutResult,
SegmentSpec,
)
from .qc import (
DuplicateGroup,
FlashFrame,
GapInfo,
ValidationIssue,
ValidationResult,
)
from .subtitles import (
COLOR_GREY,
COLOR_INDIGO,
COLOR_WHITE,
COLOR_YELLOW,
EDITORIAL_BODY_LOOK,
EDITORIAL_EMPHASIS_LOOK,
REFERENCE_RHYTHM,
DynamicSubtitleConfig,
SubtitlePosition,
WordLook,
WordStyle,
)
from .timeline import (
AudioClip,
Clip,
CompoundClip,
ConnectedClip,
Keyword,
Marker,
Project,
SilenceCandidate,
Timeline,
Transition,
VideoClip,
)
from .timing import (
_FCPXML_STANDARD_TIMEBASES,
Timecode,
TimeValue,
)
__all__ = [
"AudioClip",
"COLOR_GREY",
"COLOR_INDIGO",
"COLOR_WHITE",
"COLOR_YELLOW",
"Clip",
"CompoundClip",
"ConnectedClip",
"DuplicateGroup",
"DynamicSubtitleConfig",
"EDITORIAL_BODY_LOOK",
"EDITORIAL_EMPHASIS_LOOK",
"FlashFrame",
"FlashFrameSeverity",
"GapInfo",
"Keyword",
"MARKER_XML_TAGS",
"Marker",
"MarkerColor",
"MarkerType",
"MontageConfig",
"PacingConfig",
"PacingCurve",
"PacingStyle",
"Project",
"REFERENCE_RHYTHM",
"RoughCutResult",
"SegmentSpec",
"SilenceCandidate",
"SubtitlePosition",
"TimeValue",
"Timecode",
"Timeline",
"Transition",
"TransitionType",
"ValidationIssue",
"ValidationIssueType",
"ValidationResult",
"VideoClip",
"WordLook",
"WordStyle",
"_FCPXML_STANDARD_TIMEBASES",
"_MAX_MARKER_TYPE_LENGTH",
]
+183
View File
@@ -0,0 +1,183 @@
"""Enumerações do domínio: tipos e cores de marcador, transições, ritmo.
Extraído de models.py — ver fcpxml/models/__init__.py.
"""
from enum import Enum
# Maximum length for marker type strings to prevent memory abuse
_MAX_MARKER_TYPE_LENGTH = 64
class MarkerType(Enum):
"""Types of markers in Final Cut Pro.
Members:
STANDARD — Default marker with no completion state.
INCOMPLETE — Task marker (completed="0" in FCPXML). ← canonical name
TODO — Alias for INCOMPLETE. Kept for backward compatibility;
resolves to the same object (``MarkerType.TODO is
MarkerType.INCOMPLETE``). Python enums treat the first
member with a given value as canonical; all subsequent
members sharing that value become aliases.
CHAPTER — Chapter marker (``<chapter-marker>`` element).
COMPLETED — Task marker with completed="1".
Serialization helpers:
``from_string()`` — Accepts values, names, and legacy aliases
(e.g. ``"todo-marker"``). Always returns the
canonical member.
``from_xml_element()`` — Reads an ``lxml``/``ElementTree`` element and
returns the appropriate type based on the tag
name and ``completed`` attribute.
``xml_tag`` — The FCPXML element tag to emit when writing.
``xml_attrs`` — Extra attributes required when writing (e.g.
``completed="0"`` for INCOMPLETE).
"""
STANDARD = "standard"
INCOMPLETE = "todo"
TODO = "todo" # Backward-compat alias — resolves to INCOMPLETE at runtime
CHAPTER = "chapter"
COMPLETED = "completed"
@classmethod
def from_string(cls, value: str) -> 'MarkerType':
"""Convert a string to MarkerType, accepting both enum names and values.
Includes input validation: rejects null bytes, control characters,
and excessively long strings to prevent injection and memory abuse.
Examples:
MarkerType.from_string("todo") -> MarkerType.INCOMPLETE
MarkerType.from_string("TODO") -> MarkerType.INCOMPLETE
MarkerType.from_string("completed") -> MarkerType.COMPLETED
"""
if not isinstance(value, str):
raise TypeError(f"Expected str, got {type(value).__name__}")
if '\x00' in value or any(ord(c) < 32 and c not in ('\n', '\r', '\t') for c in value):
raise ValueError("Marker type contains invalid control characters")
if len(value) > _MAX_MARKER_TYPE_LENGTH:
raise ValueError(
f"Marker type exceeds maximum length ({_MAX_MARKER_TYPE_LENGTH} chars)"
)
lowered = value.strip().lower()
if not lowered:
raise ValueError("Marker type cannot be empty")
# Accept legacy aliases from older specs (e.g. "todo-marker" → INCOMPLETE)
aliases = {
"todo-marker": "todo",
"completed-marker": "completed",
"chapter-marker": "chapter",
}
lowered = aliases.get(lowered, lowered)
try:
return cls(lowered)
except ValueError:
raise ValueError(
f"Invalid marker type: '{value}'. "
f"Valid types: {', '.join(m.value for m in cls)}"
)
@classmethod
def from_xml_element(cls, elem) -> 'MarkerType':
"""Determine MarkerType from an XML element's tag and attributes.
Centralises the parse-side mapping so the parser doesn't need to
know about completed-attribute semantics.
Rules (in priority order):
1. <chapter-marker> tag → CHAPTER (completed attr ignored)
2. completed='0' (exact) → INCOMPLETE
3. completed='1' (exact) → COMPLETED
4. Everything else → STANDARD (including whitespace-padded,
absent, empty, or non-boolean completed values)
Matching is intentionally strict — no .strip(), no case folding.
This prevents whitespace-injected attributes like ' 0 ' from
being misclassified.
"""
if elem.tag == 'chapter-marker':
return cls.CHAPTER
completed = elem.get('completed')
if completed == '0':
return cls.INCOMPLETE
if completed == '1':
return cls.COMPLETED
return cls.STANDARD
@property
def xml_tag(self) -> str:
"""Return the FCPXML element tag for this marker type."""
return 'chapter-marker' if self == MarkerType.CHAPTER else 'marker'
@property
def xml_attrs(self) -> dict:
"""Return extra XML attributes this marker type requires when writing.
Centralises the write-side mapping so both FCPXMLModifier and
FCPXMLWriter use a single source of truth.
"""
if self == MarkerType.CHAPTER:
return {'posterOffset': '0s'}
if self == MarkerType.INCOMPLETE:
return {'completed': '0'}
if self == MarkerType.COMPLETED:
return {'completed': '1'}
return {}
# Recognised marker XML tags — used by the parser for single-pass collection
# and by the writer to validate element creation.
MARKER_XML_TAGS = ('marker', 'chapter-marker')
class MarkerColor(Enum):
"""Marker color options (FCP internal values)."""
BLUE = 0
CYAN = 1
GREEN = 2
YELLOW = 3
ORANGE = 4
RED = 5
PINK = 6
PURPLE = 7
class TransitionType(Enum):
"""Built-in transition types."""
CROSS_DISSOLVE = "Cross Dissolve"
FADE_TO_BLACK = "Fade to Color"
FADE_FROM_BLACK = "Fade from Color"
DIP_TO_COLOR = "Dip to Color"
WIPE = "Wipe"
SLIDE = "Slide"
class PacingStyle(Enum):
"""Pacing presets for rough cut generation."""
SLOW = "slow" # 5-10 second cuts
MEDIUM = "medium" # 2-5 second cuts
FAST = "fast" # 0.5-2 second cuts
DYNAMIC = "dynamic" # Varies throughout
class FlashFrameSeverity(Enum):
"""Severity levels for flash frame detection."""
CRITICAL = "critical" # < 2 frames, almost certainly an error
WARNING = "warning" # < 6 frames, potentially intentional but suspicious
class PacingCurve(Enum):
"""Pacing curves for montage generation."""
CONSTANT = "constant" # Same clip duration throughout
ACCELERATING = "accelerating" # Starts slow, gets faster
DECELERATING = "decelerating" # Starts fast, gets slower
PYRAMID = "pyramid" # Slow → fast → slow
class ValidationIssueType(Enum):
"""Types of timeline validation issues."""
FLASH_FRAME = "flash_frame"
GAP = "gap"
DUPLICATE = "duplicate"
ORPHAN_REF = "orphan_ref"
INVALID_OFFSET = "invalid_offset"
# DTD validation types (v0.6.0)
ELEMENT_ORDER = "element_order"
MISSING_ATTRIBUTE = "missing_attribute"
INVALID_TIMEBASE = "invalid_timebase"
FRAME_MISALIGNMENT = "frame_misalignment"
MISSING_EFFECT_REF = "missing_effect_ref"
MISSING_MEDIA_REP = "missing_media_rep"
+93
View File
@@ -0,0 +1,93 @@
"""Especificações de geração: rough cut, ritmo e montagem.
Extraído de models.py — ver fcpxml/models/__init__.py.
"""
from dataclasses import dataclass, field
from typing import List, Optional, Tuple
from .enums import PacingCurve
@dataclass
class SegmentSpec:
"""Specification for a segment in auto rough cut."""
name: str
keywords: List[str] = field(default_factory=list)
duration_seconds: float = 0.0
priority: str = "best" # favorites, longest, shortest, random, best
@dataclass
class PacingConfig:
"""Configuration for rough cut pacing."""
pacing: str = "medium" # slow, medium, fast, dynamic
min_clip_duration: float = 1.0
max_clip_duration: float = 8.0
avg_clip_duration: Optional[float] = None
vary_pacing: bool = True
def get_duration_range(self) -> Tuple[float, float]:
"""Get min/max based on pacing style."""
ranges = {
"slow": (5.0, 10.0),
"medium": (2.0, 5.0),
"fast": (0.5, 2.0),
"dynamic": (1.0, 6.0),
}
return ranges.get(self.pacing, (2.0, 5.0))
@dataclass
class RoughCutResult:
"""Result of auto rough cut generation."""
output_path: str
clips_used: int
clips_available: int
target_duration: float
actual_duration: float
segments: int
average_clip_duration: float
@dataclass
class MontageConfig:
"""Configuration for montage generation with pacing curves."""
target_duration: float # Target duration in seconds
pacing_curve: 'PacingCurve'
start_duration: float = 2.0 # Clip duration at start
end_duration: float = 0.5 # Clip duration at end
min_duration: float = 0.2 # Minimum allowed clip duration
max_duration: float = 5.0 # Maximum allowed clip duration
def get_duration_at_position(self, position: float) -> float:
"""
Calculate clip duration for a given position (0.0 to 1.0).
Args:
position: Position in montage (0.0 = start, 1.0 = end)
Returns:
Target duration in seconds for a clip at this position
"""
if self.pacing_curve == PacingCurve.CONSTANT:
duration = (self.start_duration + self.end_duration) / 2
elif self.pacing_curve == PacingCurve.ACCELERATING:
# Linear interpolation from start to end duration
duration = self.start_duration + (self.end_duration - self.start_duration) * position
elif self.pacing_curve == PacingCurve.DECELERATING:
# Reverse: start fast, end slow
duration = self.end_duration + (self.start_duration - self.end_duration) * position
elif self.pacing_curve == PacingCurve.PYRAMID:
# Slow → fast → slow (parabolic curve)
if position < 0.5:
# First half: slow to fast
duration = self.start_duration + (self.end_duration - self.start_duration) * (position * 2)
else:
# Second half: fast to slow
duration = self.end_duration + (self.start_duration - self.end_duration) * ((position - 0.5) * 2)
else:
duration = self.start_duration
# Clamp to min/max
return max(self.min_duration, min(self.max_duration, duration))
+121
View File
@@ -0,0 +1,121 @@
"""Achados de QC e o resultado de uma validação.
Extraído de models.py — ver fcpxml/models/__init__.py.
"""
from dataclasses import dataclass, field
from typing import Any, Dict, List, Optional
from .enums import FlashFrameSeverity, ValidationIssueType
from .timing import Timecode
@dataclass
class FlashFrame:
"""
Represents a detected flash frame (ultra-short clip).
Flash frames are typically editing errors - clips that are too short
to be perceived as intentional cuts.
"""
clip_name: str
clip_id: str
start: Timecode
duration_frames: int
duration_seconds: float
severity: 'FlashFrameSeverity'
@property
def is_critical(self) -> bool:
"""Check if this is a critical flash frame."""
return self.severity == FlashFrameSeverity.CRITICAL
@dataclass
class GapInfo:
"""
Represents a detected gap in the timeline.
Gaps can be intentional (black frames) or errors from deleted clips.
"""
start: Timecode
duration_frames: int
duration_seconds: float
previous_clip: Optional[str] = None # Clip name before the gap
next_clip: Optional[str] = None # Clip name after the gap
@property
def timecode(self) -> str:
"""Get timecode string for the gap start."""
return self.start.to_smpte()
@dataclass
class DuplicateGroup:
"""
Represents a group of clips using the same source media.
Useful for detecting duplicate clips that may be unintentional.
"""
source_ref: str # The asset/media reference ID
source_name: str # Human-readable source name
clips: List[Dict[str, Any]] = field(default_factory=list) # List of clip info dicts
@property
def count(self) -> int:
"""Number of clips using this source."""
return len(self.clips)
@property
def has_overlapping_ranges(self) -> bool:
"""Check if any clips use overlapping portions of the source."""
# Sort clips by source_start
sorted_clips = sorted(self.clips, key=lambda c: c.get('source_start', 0))
for i in range(len(sorted_clips) - 1):
curr_end = sorted_clips[i].get('source_start', 0) + sorted_clips[i].get('source_duration', 0)
next_start = sorted_clips[i + 1].get('source_start', 0)
if curr_end > next_start:
return True
return False
@dataclass
class ValidationIssue:
"""
Represents a single validation issue found in a timeline.
Used by validate_timeline to report problems.
"""
issue_type: 'ValidationIssueType'
severity: str # "error", "warning", "info"
message: str
timecode: Optional[str] = None
clip_name: Optional[str] = None
details: Dict[str, Any] = field(default_factory=dict)
@dataclass
class ValidationResult:
"""
Result of timeline validation.
Provides a health score and categorized list of issues.
"""
is_valid: bool
health_score: int # 0-100 percentage
issues: List[ValidationIssue] = field(default_factory=list)
flash_frames: List[FlashFrame] = field(default_factory=list)
gaps: List[GapInfo] = field(default_factory=list)
duplicates: List[DuplicateGroup] = field(default_factory=list)
@property
def error_count(self) -> int:
return len([i for i in self.issues if i.severity == "error"])
@property
def warning_count(self) -> int:
return len([i for i in self.issues if i.severity == "warning"])
def summary(self) -> str:
"""Generate a summary string."""
return (
f"Timeline Health: {self.health_score}% | "
f"Errors: {self.error_count} | Warnings: {self.warning_count} | "
f"Flash frames: {len(self.flash_frames)} | Gaps: {len(self.gaps)}"
)
+165
View File
@@ -0,0 +1,165 @@
"""Aparência das legendas dinâmicas: paleta, look por palavra, configuração.
Extraído de models.py — ver fcpxml/models/__init__.py.
"""
from dataclasses import dataclass, field
from typing import Optional
from ..text_layout import REFERENCE_BLOCK_LINE_GAP, TEXT_TEMPLATE_FONT_SCALE
# The palette and type treatment of the calibration export
# ("Exemplo Letra.fcpxmld", sentence "Toda a minha vida, assim,"), copied
# verbatim from what the user set in Final Cut's Inspector.
COLOR_INDIGO = "0.156863 0 0.596079 1"
COLOR_YELLOW = "0.997808 0.882664 0.0388632 1"
COLOR_GREY = "0.7 0.7 0.7 1"
COLOR_WHITE = "1 1 1 1"
@dataclass
class WordLook:
"""How one word is set: size, colour and type treatment.
A sentence cycles through a tuple of these, so its typography reads with a
deliberate rhythm rather than a uniform block.
"""
font_size: int
color: str
font: str = "Helvetica Neue"
face: Optional[str] = None # Final Cut's fontFace, e.g. "Light Italic"
kerning: float = 2.048
@property
def italic(self) -> bool:
return bool(self.face) and "italic" in self.face.lower()
# One entry per word of the reference sentence, in order:
# Toda(170, indigo, Helvetica Light) a(128, yellow) minha(151, grey)
# vida,(128, white) assim,(128, grey, Light Italic)
REFERENCE_RHYTHM = (
WordLook(170, COLOR_INDIGO, font="Helvetica", face="Light", kerning=2.72),
WordLook(128, COLOR_YELLOW),
WordLook(151, COLOR_GREY, kerning=2.416),
WordLook(128, COLOR_WHITE),
WordLook(128, COLOR_GREY, face="Light Italic"),
)
# The progressive-composition look (reference: the reel the user sent,
# 2026-08-17). Supporting text in a small grotesque, the sentence's key word
# large in a display italic, everything white — the two-font contrast IS the
# style. Playfair Display ships in the user's ~/Library/Fonts and its real
# advance widths are embedded in font_metrics, so the lines can be measured
# rather than guessed. Both are plain WordLooks: swap them for any installed
# family (a script/calligraphic face for the emphasis, say) and layout follows.
EDITORIAL_EMPHASIS_LOOK = WordLook(
230, COLOR_WHITE, font="Playfair Display", face="Medium Italic", kerning=0.0,
)
EDITORIAL_BODY_LOOK = WordLook(
88, COLOR_WHITE, font="Helvetica Neue", face="Bold", kerning=1.2,
)
@dataclass
class WordStyle:
"""Per-word text styling for dynamic (karaoke-style) subtitles.
``rhythm`` drives size, colour and face, cycling by the word's index within
its sentence — deterministic, so regenerating a transcript twice yields the
same look. ``font``/``font_size`` are the fallback when ``rhythm`` is empty.
"""
font: str = "Helvetica Neue"
font_size: int = 128
active_color: str = COLOR_WHITE
inactive_color: str = COLOR_GREY
bold: bool = False
kerning: float = 2.048
rhythm: tuple = REFERENCE_RHYTHM
# Progressive composition only (granularity="phrase").
emphasis_look: Optional[WordLook] = None
body_look: Optional[WordLook] = None
def look_for(self, index: int) -> WordLook:
"""The look for the word at *index* within its sentence."""
if not self.rhythm:
return WordLook(
self.font_size, self.active_color,
font=self.font, kerning=self.kerning,
)
return self.rhythm[index % len(self.rhythm)]
def look_for_emphasis(self) -> WordLook:
"""The look for a composition's key word (progressive composition)."""
return self.emphasis_look or EDITORIAL_EMPHASIS_LOOK
def look_for_body(self) -> WordLook:
"""The look for a composition's supporting lines."""
return self.body_look or EDITORIAL_BODY_LOOK
@dataclass
class SubtitlePosition:
"""Screen position for generated title clips, in FCP title coordinate space."""
x: float = 0.0
y: float = -300.0
alignment: str = "center" # left | center | right
@dataclass
class DynamicSubtitleConfig:
"""Options for FCPXMLWriter.generate_dynamic_subtitles().
Dynamic subtitles are animated TITLES, not captions. Both templates below
render on the video title lane and never carry a ``subtitles.*`` role — a
``role="subtitles.*"`` would make Final Cut treat them as captions and
hide them behind the caption-display toggle. They DO carry a
``titles.*`` sub-role (``role``), which groups them in Final Cut's
role index and lanes them with a distinct colour, without ever being
mistaken for closed captions.
``animated`` picks the template: True uses "Essencial - Título"
(Essential Title), which animates on its own Motion defaults; False uses
the static "Título Básico" (Basic Title). Default is True — the animated
reveal is the feature's purpose.
Words are grouped into sentences and laid out as a compact typographic
block: each word becomes its own positioned ``<title>``, appearing as it is
spoken and accumulating on screen, with every word of a block clearing at
the same instant so the sentence vanishes as a whole.
``band_height`` is the fraction of frame height the block may occupy, and
``block_center_y`` its centre in canvas points (negative is below frame
centre). The defaults reproduce the calibration export the user built by
hand: a block of at most three lines sitting just below centre. A sentence
taller than the band splits into successive blocks.
"""
style: WordStyle = field(default_factory=WordStyle)
position: SubtitlePosition = field(default_factory=SubtitlePosition)
animated: bool = True
band_height: float = 0.22
block_center_y: float = -167.0
# "phrase": one title per LINE of the composition — supporting words
# grouped, the key word alone and large (the reference look). "word": one
# title per word, the earlier rhythm.
granularity: str = "phrase"
# Ratio between the template's fontSize space and the canvas-point space
# its Position uses. See text_layout.TEXT_TEMPLATE_FONT_SCALE: the "Text"
# (Text.moti) template sizes type in frame pixels, so a size chosen in
# points renders half as large unless it is converted on the way out.
text_scale: float = TEXT_TEMPLATE_FONT_SCALE
# Vertical air between stacked lines, in canvas points. Negative values
# deliberately overlap the lines — the display italic tucking under the
# line above is a real editorial look, and the stacking arithmetic places
# ink boxes edge to edge, so a negative gap moves them by exactly that
# much rather than colliding unpredictably.
line_gap: float = REFERENCE_BLOCK_LINE_GAP
# Final Cut role for every title this generator emits. A ``titles.*``
# sub-role (NOT ``subtitles.*``) groups the clips in the role index and
# tints their lane, keeping dynamic captions distinct from plain
# ones and from Final Cut's own closed-caption toggle.
role: str = "titles.dinamicas"
# Run the post-generation collision validation (collision.validate_titles)
# and refuse to emit when it reports a blocking overlap. Off by default so
# generation stays byte-identical to before this flag existed; flip it on
# for a guaranteed no-collision export.
validate: bool = False
+248
View File
@@ -0,0 +1,248 @@
"""O que existe numa timeline: clipes, marcadores, lanes, projeto.
Extraído de models.py — ver fcpxml/models/__init__.py.
"""
from dataclasses import dataclass, field
from typing import List, Optional
from .enums import MarkerColor, MarkerType
from .timing import Timecode
@dataclass
class Keyword:
"""Represents a keyword/tag applied to a clip."""
value: str
start: Optional[Timecode] = None
duration: Optional[Timecode] = None
@dataclass
class ParametroEfeito:
"""Um parâmetro de um filtro de efeito (``<param>`` dentro do filtro)."""
nome: str
valor: str
chave: str = ""
metadado: str = ""
@dataclass
class EfeitoAjuste:
"""Um efeito aplicado por uma camada de ajuste (adjustment layer).
``uid`` é o UUID do efeito interno do Final Cut (ver ``FCP_EFFECTS`` em
``fcpxml/writer/helpers.py`` para os efeitos built-in). ``tipo`` é
``"video"`` ou ``"audio"`` — decide se vira ``<filter-video>`` ou
``<filter-audio>``, filho direto do ``<clip>`` da camada de ajuste (o
DTD não define wrapper ``<adjustment>``).
"""
nome: str
uid: str
tipo: str = "video"
parametros: List[ParametroEfeito] = field(default_factory=list)
@dataclass
class Marker:
"""Represents a marker in the timeline."""
name: str
start: Timecode
duration: Optional[Timecode] = None
marker_type: MarkerType = MarkerType.STANDARD
note: str = ""
color: Optional[MarkerColor] = None
def to_youtube_timestamp(self) -> str:
"""Format as YouTube chapter timestamp."""
total_seconds = int(self.start.seconds)
hours = total_seconds // 3600
minutes = (total_seconds % 3600) // 60
secs = total_seconds % 60
if hours > 0:
return f"{hours}:{minutes:02d}:{secs:02d}"
return f"{minutes}:{secs:02d}"
@dataclass
class Clip:
"""Represents a clip in the timeline."""
name: str
start: Timecode
duration: Timecode
source_start: Optional[Timecode] = None
source_end: Optional[Timecode] = None
media_path: str = ""
markers: List[Marker] = field(default_factory=list)
keywords: List[Keyword] = field(default_factory=list)
# Extended metadata
rating: int = 0 # 0=unrated, 1-5 stars
is_favorite: bool = False
is_rejected: bool = False
# Roles (FCP audio/video role assignments)
audio_role: str = ""
video_role: str = ""
# Connected clips (B-roll, titles, audio attached to this clip)
connected_clips: List['ConnectedClip'] = field(default_factory=list)
# Edit-time correction, in degrees, from a Transform filter on the clip
# (e.g. straightening a tilted phone shot) — not the camera's own
# recorded orientation, which lives in the media file itself.
rotation: float = 0.0
@property
def end(self) -> Timecode:
return Timecode(
frames=self.start.frames + self.duration.frames,
frame_rate=self.start.frame_rate
)
@property
def duration_seconds(self) -> float:
return self.duration.seconds
@property
def keyword_values(self) -> List[str]:
"""Get list of keyword strings."""
return [k.value for k in self.keywords]
@dataclass
class AudioClip(Clip):
"""Audio-specific clip."""
channels: int = 2
sample_rate: int = 48000
role: str = "dialogue"
@dataclass
class VideoClip(Clip):
"""Video-specific clip."""
width: int = 1920
height: int = 1080
has_audio: bool = True
@dataclass
class ConnectedClip:
"""A clip connected to a primary storyline clip (B-roll, titles, audio).
In FCP's magnetic timeline, connected clips hang off spine clips via lanes.
Positive lanes are above (video overlays), negative lanes are below (audio).
"""
name: str
start: Timecode
duration: Timecode
lane: int = 1
offset: Optional[Timecode] = None
source_start: Optional[Timecode] = None
media_path: str = ""
clip_type: str = "asset-clip"
role: str = ""
ref_id: str = ""
parent_clip_name: str = ""
markers: List[Marker] = field(default_factory=list)
keywords: List[Keyword] = field(default_factory=list)
rotation: float = 0.0
@property
def duration_seconds(self) -> float:
return self.duration.seconds
@dataclass
class CompoundClip:
"""A compound clip (ref-clip) containing a nested timeline."""
name: str
ref_id: str
duration: Timecode
start: Timecode
clips: List[Clip] = field(default_factory=list)
connected_clips: List[ConnectedClip] = field(default_factory=list)
@property
def duration_seconds(self) -> float:
return self.duration.seconds
@dataclass
class SilenceCandidate:
"""A potential silence region detected by timeline heuristics."""
start_timecode: str
duration_seconds: float
reason: str # "gap", "ultra_short", "name_match", "duration_anomaly"
confidence: float = 0.5 # 0.0 to 1.0
clip_name: Optional[str] = None
clip_index: Optional[int] = None
@dataclass
class Transition:
"""Represents a transition between clips."""
name: str
duration: Timecode
start: Timecode
transition_type: str = "cross-dissolve"
@dataclass
class Timeline:
"""Represents a Final Cut Pro timeline/sequence."""
name: str
duration: Timecode
frame_rate: float = 24.0
width: int = 1920
height: int = 1080
clips: List[Clip] = field(default_factory=list)
audio_clips: List[AudioClip] = field(default_factory=list)
transitions: List[Transition] = field(default_factory=list)
markers: List[Marker] = field(default_factory=list)
connected_clips: List[ConnectedClip] = field(default_factory=list)
compound_clips: List[CompoundClip] = field(default_factory=list)
@property
def total_clips(self) -> int:
return len(self.clips)
@property
def total_cuts(self) -> int:
return max(0, len(self.clips) - 1)
@property
def average_clip_duration(self) -> float:
if not self.clips:
return 0.0
return sum(c.duration_seconds for c in self.clips) / len(self.clips)
@property
def cuts_per_minute(self) -> float:
"""Average cuts per minute."""
if self.duration.seconds <= 0:
return 0.0
return (self.total_cuts / self.duration.seconds) * 60
def get_clips_shorter_than(self, seconds: float) -> List[Clip]:
"""Find clips shorter than threshold (flash frame detection)."""
return [c for c in self.clips if c.duration_seconds < seconds]
def get_clips_longer_than(self, seconds: float) -> List[Clip]:
"""Find clips longer than threshold."""
return [c for c in self.clips if c.duration_seconds > seconds]
def get_clip_at(self, timecode: float) -> Optional[Clip]:
"""Find the clip at a specific timecode (seconds)."""
for clip in self.clips:
start_sec = clip.start.seconds
end_sec = clip.end.seconds
if start_sec <= timecode < end_sec:
return clip
return None
def get_clips_by_keyword(self, keyword: str) -> List[Clip]:
"""Find all clips with a specific keyword."""
return [c for c in self.clips if keyword in c.keyword_values]
@dataclass
class Project:
"""Represents a Final Cut Pro project/library."""
name: str
timelines: List[Timeline] = field(default_factory=list)
fcpxml_version: str = "1.13"
@property
def primary_timeline(self) -> Optional[Timeline]:
return self.timelines[0] if self.timelines else None
+304
View File
@@ -0,0 +1,304 @@
"""Tempo em fração racional — TimeValue e o Timecode que o embrulha.
Extraído de models.py — ver fcpxml/models/__init__.py.
"""
import operator
from dataclasses import dataclass
from fractions import Fraction
from functools import total_ordering
from math import gcd
from typing import Callable
# Standard FCPXML timebase denominators that FCP's DTD validator accepts.
# TimeValue.to_fcpxml() only simplifies fractions when the result uses one
# of these denominators, preventing values like "8/3s" that FCP rejects.
_FCPXML_STANDARD_TIMEBASES = frozenset({
1, 24, 25, 30, 48, 50, 60, 90, 96, 100, 120,
240, 600, 2400, 4800, 9600, 48000,
})
@total_ordering
@dataclass
class TimeValue:
"""
Represents time in FCPXML's rational format.
FCPXML uses fractions of seconds (e.g., "90/30s" for 3 seconds at 30fps).
This class handles conversion between timecode, seconds, and FCPXML format.
Examples:
TimeValue(90, 30) # 3 seconds at 30fps
TimeValue(1, 1) # 1 second
TimeValue.from_timecode("00:01:30:15", fps=30) # 90.5 seconds
"""
numerator: int
denominator: int = 1
def __post_init__(self):
if self.denominator == 0:
raise ValueError(
f"TimeValue denominator cannot be zero (got {self.numerator}/0). "
"This would corrupt all downstream time calculations."
)
# Normalize sign: denominator must always be positive.
# Cross-multiplication in __lt__/__eq__ assumes positive denominators;
# __hash__ assumes canonical form. Without this, TimeValue(1, -2)
# compares/hashes incorrectly against TimeValue(-1, 2).
if self.denominator < 0:
# Use object.__setattr__ because dataclass may be frozen-like
object.__setattr__(self, 'numerator', -self.numerator)
object.__setattr__(self, 'denominator', -self.denominator)
@classmethod
def from_timecode(cls, tc: str, fps: float = 30.0) -> 'TimeValue':
"""
Create TimeValue from various string formats.
Supported formats:
- "HH:MM:SS:FF" - Standard timecode
- "HH:MM:SS;FF" - Drop-frame timecode
- "30s" - Seconds
- "90/30s" - FCPXML rational format
- "15f" - Frames
"""
if not tc:
return cls(0, 1)
tc = str(tc).strip()
# FCPXML format: "90/30s" or "30s"
if tc.endswith('s'):
tc_val = tc[:-1]
if '/' in tc_val:
parts = tc_val.split('/', 1)
num, denom = int(parts[0]), int(parts[1])
if denom == 0:
raise ValueError(f"Zero denominator in timecode: {tc}")
return cls(num, denom)
else:
seconds = float(tc_val)
frames = int(round(seconds * fps))
# int(fps) truncates NTSC rates (23.976/29.97/59.94fps) to
# their nominal integer, mismatching the numerator (computed
# with the real fps) against the denominator — e.g. at
# 23.976fps this silently produced values ~1.04x too large.
# Reconstruct the exact rational fps (24000/1001, etc.) from
# the float instead, so numerator and denominator agree.
fps_frac = Fraction(fps).limit_denominator(100_000)
return cls(frames * fps_frac.denominator, fps_frac.numerator)
# Frame format: "15f"
if tc.endswith('f'):
frames = int(tc[:-1])
return cls(frames, int(fps))
# Timecode format: "HH:MM:SS:FF" or "HH:MM:SS;FF"
if ':' in tc or ';' in tc:
parts = tc.replace(';', ':').split(':')
if len(parts) == 4:
h, m, s, f = map(int, parts)
total_frames = int((h * 3600 + m * 60 + s) * fps + f)
return cls(total_frames, int(fps))
elif len(parts) == 3:
h, m, s = map(int, parts)
total_frames = int((h * 3600 + m * 60 + s) * fps)
return cls(total_frames, int(fps))
# Try as plain number (seconds)
try:
seconds = float(tc)
frames = int(round(seconds * fps))
return cls(frames, int(fps))
except ValueError:
raise ValueError(f"Invalid timecode format: {tc}")
@classmethod
def from_seconds(cls, seconds: float, fps: float = 30.0) -> 'TimeValue':
"""Create TimeValue from decimal seconds."""
frames = int(round(seconds * fps))
return cls(frames, int(fps))
@classmethod
def zero(cls) -> 'TimeValue':
"""Return zero time value."""
return cls(0, 1)
def to_fcpxml(self) -> str:
"""Convert to FCPXML time string (e.g., "90/30s").
Only simplifies when the denominator reduces to 1 (whole seconds)
or stays a standard FCPXML timebase. Avoids producing denominators
like 3, 7, etc. that FCP's DTD validator may reject.
"""
simplified = self.simplify()
if simplified.denominator == 1:
return f"{simplified.numerator}s"
# Keep original denominator if simplification produces a non-standard
# denominator (not a multiple of common timebases: 24, 30, 25, 2400)
if simplified.denominator in _FCPXML_STANDARD_TIMEBASES:
return f"{simplified.numerator}/{simplified.denominator}s"
# Fall back to unsimplified form
return f"{self.numerator}/{self.denominator}s"
def to_seconds(self) -> float:
"""Convert to decimal seconds."""
return self.numerator / self.denominator
def to_timecode(self, fps: float = 30.0) -> str:
"""Convert to HH:MM:SS:FF timecode string."""
total_frames = int(round(self.to_seconds() * fps))
total_secs, frames = divmod(total_frames, int(fps))
total_mins, secs = divmod(total_secs, 60)
hours, mins = divmod(total_mins, 60)
return f"{hours:02d}:{mins:02d}:{secs:02d}:{frames:02d}"
def to_frames(self, fps: float = 30.0) -> int:
"""Convert to frame count."""
return int(round(self.to_seconds() * fps))
def simplify(self) -> 'TimeValue':
"""Reduce fraction to simplest form."""
if self.numerator == 0:
return TimeValue(0, 1)
divisor = gcd(abs(self.numerator), abs(self.denominator))
return TimeValue(
self.numerator // divisor,
self.denominator // divisor
)
@staticmethod
def _lcm_denom(d1: int, d2: int) -> int:
"""LCM of two denominators for cross-timebase arithmetic."""
return d1 // gcd(d1, d2) * d2
def _binop(self, other: 'TimeValue', op: Callable[[int, int], int]) -> 'TimeValue':
"""Shared logic for add/sub: same-denom fast path, then LCM alignment."""
if self.denominator == other.denominator:
return TimeValue(op(self.numerator, other.numerator), self.denominator)
lcd = TimeValue._lcm_denom(self.denominator, other.denominator)
return TimeValue(
op(
self.numerator * (lcd // self.denominator),
other.numerator * (lcd // other.denominator),
),
lcd,
)
def __add__(self, other: 'TimeValue') -> 'TimeValue':
return self._binop(other, operator.add)
def __sub__(self, other: 'TimeValue') -> 'TimeValue':
return self._binop(other, operator.sub)
def __mul__(self, scalar: float) -> 'TimeValue':
new_num = round(self.numerator * scalar)
return TimeValue(new_num, self.denominator)
def __truediv__(self, scalar: float) -> 'TimeValue':
if scalar == 0:
raise ZeroDivisionError("Cannot divide TimeValue by zero")
new_denom = round(self.denominator * scalar)
if new_denom == 0:
raise ZeroDivisionError(
f"Division by {scalar} rounds denominator {self.denominator} to zero"
)
return TimeValue(self.numerator, new_denom)
def __lt__(self, other: 'TimeValue') -> bool:
# Cross-multiply to compare without float conversion:
# a/b < c/d ↔ a*d < c*b (denominators are always positive)
return self.numerator * other.denominator < other.numerator * self.denominator
def __eq__(self, other: object) -> bool:
if not isinstance(other, TimeValue):
return False
# Cross-multiply for exact integer comparison
return self.numerator * other.denominator == other.numerator * self.denominator
def __hash__(self) -> int:
# Delegate to simplify() — single source of truth for canonical form.
# __post_init__ guarantees denominator > 0, so no zero guard needed.
s = self.simplify()
return hash((s.numerator, s.denominator))
def snap_to_frame(self, fps: float) -> 'TimeValue':
"""Round this time value to the nearest frame boundary at the given fps.
Uses the 2400-tick timebase (LCM of common frame rates) so results
always land on clean frame boundaries.
Args:
fps: Frame rate to snap to (e.g. 24, 30, 60)
Returns:
New TimeValue snapped to the nearest frame in 2400-tick timebase.
"""
fps_int = int(fps)
if fps_int <= 0:
raise ValueError(f"fps must be positive, got {fps}")
ticks_per_frame = 2400 // fps_int
total_ticks = round(self.to_seconds() * 2400)
snapped_ticks = round(total_ticks / ticks_per_frame) * ticks_per_frame
return TimeValue(snapped_ticks, 2400)
def is_standard_timebase(self) -> bool:
"""Check if this TimeValue's denominator is an FCP-accepted timebase."""
simplified = self.simplify()
return simplified.denominator in _FCPXML_STANDARD_TIMEBASES
def __repr__(self) -> str:
return f"TimeValue({self.numerator}/{self.denominator}s = {self.to_seconds():.3f}s)"
@dataclass
class Timecode:
"""
Represents a timecode value.
Note: This class exists for backwards compatibility with the parser.
New code should prefer TimeValue for rational time math.
"""
frames: int
frame_rate: float = 24.0
drop_frame: bool = False
@property
def seconds(self) -> float:
return self.frames / self.frame_rate
@property
def total_frames(self) -> int:
return self.frames
def to_smpte(self) -> str:
"""Convert to SMPTE timecode string (HH:MM:SS:FF)."""
total_seconds = int(self.seconds)
hours = total_seconds // 3600
minutes = (total_seconds % 3600) // 60
secs = total_seconds % 60
frames = int((self.seconds - total_seconds) * self.frame_rate)
separator = ";" if self.drop_frame else ":"
return f"{hours:02d}:{minutes:02d}:{secs:02d}{separator}{frames:02d}"
@classmethod
def from_rational(cls, rational_str: str, frame_rate: float = 24.0) -> "Timecode":
"""Parse FCPXML rational time format (e.g., '3600/24s')."""
if not rational_str:
return cls(frames=0, frame_rate=frame_rate)
if rational_str.endswith('s'):
rational_str = rational_str[:-1]
if '/' in rational_str:
num, denom = rational_str.split('/')
seconds = int(num) / int(denom)
else:
seconds = float(rational_str)
frames = int(seconds * frame_rate)
return cls(frames=frames, frame_rate=frame_rate)
def to_rational(self) -> str:
"""Convert to FCPXML rational format."""
return f"{self.frames}/{int(self.frame_rate)}s"
def to_time_value(self) -> TimeValue:
"""Convert to TimeValue for rational math."""
return TimeValue(self.frames, int(self.frame_rate))
+32 -9
View File
@@ -386,18 +386,40 @@ def phrase_review_to_actions(review: dict) -> dict:
actions: List[dict] = [] actions: List[dict] = []
emphasis_spans: List[dict] = [] emphasis_spans: List[dict] = []
inactive_run: List[dict] = []
def flush_inactive_run() -> None:
"""One cut per RUN of consecutive deactivated phrases, not one per
phrase. A phrase-by-phrase cut leaves the pause BETWEEN two
deactivated phrases uncut — that gap was never anyone's content, so
nothing asked for it to survive, but it does anyway: a 0.1-0.5s
sliver clip in the final timeline for every such gap. Spanning the
whole run absorbs those gaps into the one cut."""
if not inactive_run:
return
if len(inactive_run) == 1:
reason = inactive_run[0]["reason"] or "desativada na revisão"
else:
reason = (
f"desativadas na revisão ({len(inactive_run)} frases): "
+ "; ".join(p["text"][:40] for p in inactive_run if p["text"])
)
actions.append(
VoiceAction(
kind="cut",
start=inactive_run[0]["start"],
end=inactive_run[-1]["end"],
reason=reason,
speaker=inactive_run[0]["speaker"],
).as_dict()
)
inactive_run.clear()
for phrase in phrases: for phrase in phrases:
if not phrase["active"]: if not phrase["active"]:
actions.append( inactive_run.append(phrase)
VoiceAction(
kind="cut",
start=phrase["start"],
end=phrase["end"],
reason=phrase["reason"] or "desativada na revisão",
speaker=phrase["speaker"],
).as_dict()
)
continue continue
flush_inactive_run()
# Head and tail the editor trimmed off — each becomes its own cut, so a # Head and tail the editor trimmed off — each becomes its own cut, so a
# false start disappears without taking the line with it. # false start disappears without taking the line with it.
@@ -436,6 +458,7 @@ def phrase_review_to_actions(review: dict) -> dict:
"text": phrase["text"], "text": phrase["text"],
} }
) )
flush_inactive_run()
# Hand-placed punch-ins carry no scale on purpose: an omitted scale lets the # Hand-placed punch-ins carry no scale on purpose: an omitted scale lets the
# applier use the shape configured in "Análise de Voz" (zoom_scale, ease in # applier use the shape configured in "Análise de Voz" (zoom_scale, ease in
+48
View File
@@ -316,6 +316,54 @@ def group_words_by_segment(
return groups return groups
def split_into_subphrases(
words: Sequence[dict],
min_words: int = 3,
) -> List[List[dict]]:
"""Split a sentence's *words* into sub-phrases at comma boundaries.
A comma is where a spoken sentence actually breathes, so it is the
natural seam for grouping subtitles — each sub-phrase becoming its own
on-screen block (and, downstream, its own compound clip).
The exception is the short tail: a fragment like "né?" or "Então..."
reads as part of the phrase before it, not as a phrase of its own, and
promoting it to its own block would flash a single word on screen. So a
piece shorter than *min_words* is merged back into its neighbour —
preferring the previous piece, falling back to the next one when the
short piece leads the sentence.
Returns one group per sub-phrase; a sentence with no comma comes back
as a single group.
"""
pieces: List[List[dict]] = []
current: List[dict] = []
for w in words:
current.append(w)
text = str(w.get('word') or w.get('text') or '')
if text.rstrip().endswith(','):
pieces.append(current)
current = []
if current:
pieces.append(current)
if len(pieces) <= 1:
return pieces
merged: List[List[dict]] = []
for piece in pieces:
if len(piece) < min_words and merged:
merged[-1].extend(piece)
else:
merged.append(piece)
# A short leading piece has no previous neighbour to fold into, so it
# folds forward instead.
if len(merged) > 1 and len(merged[0]) < min_words:
merged[1][:0] = merged[0]
merged.pop(0)
return merged
def segments_to_srt(segments: Sequence[dict]) -> str: def segments_to_srt(segments: Sequence[dict]) -> str:
"""Render transcript segments as an SRT string (for captions import).""" """Render transcript segments as an SRT string (for captions import)."""
+140
View File
@@ -0,0 +1,140 @@
"""Clip de ajuste (adjustment layer) — criação do elemento FCPXML.
No Final Cut, uma "camada de ajuste" é um ``<clip>`` que carrega filtros
(``filter-video`` / ``filter-audio``) diretamente como filhos — o DTD do
FCPXML 1.13 não define nenhum elemento ``<adjustment>`` como wrapper (ver
``<!ELEMENT clip>`` em ``FCPXMLv1_13.dtd``: ``filter-video``/``filter-audio``
vêm depois de ``audio-channel-source*`` e antes de ``metadata?``, sem
elemento intermediário). Tudo que está abaixo do clip na timeline herda
esses filtros — é como se o efeito fosse aplicado a uma faixa inteira de
uma vez.
Esta classe monta esse elemento a partir de dados de alto nível (duração +
lista de ``EfeitoAjuste``), cuidando de criar os recursos ``<effect>``
correspondentes na seção ``<resources>`` e de referenciá-los pelos filtros.
"""
import xml.etree.ElementTree as ET
from typing import Callable, List, Optional
from ..models.timeline import EfeitoAjuste
from ..models.timing import TimeValue
def _para_racional(tempo) -> str:
"""Aceita ``TimeValue`` ou uma string FCPXML já formatada ("90/30s")."""
if isinstance(tempo, TimeValue):
return tempo.to_fcpxml()
if tempo is None:
return "0/1s"
return str(tempo)
def _id_recurso_unico(resources: ET.Element, prefixo: str = "r_ajuste") -> str:
"""Gera um ``id`` de recurso ainda ausente em ``resources``."""
existentes = {r.get("id") for r in resources.findall("*") if r.get("id")}
contador = 1
while f"{prefixo}_{contador}" in existentes:
contador += 1
return f"{prefixo}_{contador}"
class ClipDeAjuste:
"""Cria um clip de ajuste (adjustment layer) pronto para a spine.
Exemplo::
from fcpxml.models.timing import TimeValue
from fcpxml.models.timeline import EfeitoAjuste, ParametroEfeito
from fcpxml.writer.adjustment import ClipDeAjuste
efeito = EfeitoAjuste(
nome="Color Curves", uid="...UUID...", tipo="video",
parametros=[ParametroEfeito(nome="Amount", valor="0.5",
chave=".../9999")],
)
clip = ClipDeAjuste(
nome="Ajuste de cor",
duracao=TimeValue(300, 30),
efeitos=[efeito],
).criar(resources)
spine.append(clip)
"""
def __init__(
self,
nome: str,
duracao,
efeitos: List[EfeitoAjuste],
offset=None,
formato_tc: str = "NDF",
):
self.nome = nome
self.duracao = duracao
self.efeitos = efeitos
self.offset = offset
self.formato_tc = formato_tc
def criar(
self,
resources: ET.Element,
proximo_id: Optional[Callable[[], str]] = None,
) -> ET.Element:
"""Monta o ``<clip>`` de ajuste e seus recursos ``<effect>``.
``resources`` é a seção ``<resources>`` do documento (onde os
``<effect>`` são registrados). ``proximo_id`` é um gerador opcional
de ids de recurso; sem ele, usa um id único baseado em ``resources``.
"""
def gerar_id() -> str:
if proximo_id:
return proximo_id()
return _id_recurso_unico(resources)
filtros: List[ET.Element] = []
for efeito in self.efeitos:
efeito_id = self._garantir_recurso(resources, efeito, gerar_id)
filtros.append(self._montar_filtro(efeito, efeito_id))
clip = ET.Element(
"clip",
name=self.nome,
duration=_para_racional(self.duracao),
tcFormat=self.formato_tc,
)
if self.offset is not None:
clip.set("offset", _para_racional(self.offset))
# O DTD exige filter-video* antes de filter-audio* como filhos
# diretos do clip (sem wrapper <adjustment>).
for filtro in sorted(filtros, key=lambda f: f.tag != "filter-video"):
clip.append(filtro)
return clip
def _garantir_recurso(
self, resources: ET.Element, efeito: EfeitoAjuste, gerar_id: Callable[[], str]
) -> str:
"""Devolve o ``id`` do ``<effect>`` de *efeito*, criando-o se ausente."""
for existente in resources.findall("effect"):
if existente.get("uid") == efeito.uid:
return existente.get("id")
efeito_id = gerar_id()
recurso = ET.SubElement(resources, "effect")
recurso.set("id", efeito_id)
recurso.set("name", efeito.nome)
recurso.set("uid", efeito.uid)
return efeito_id
def _montar_filtro(self, efeito: EfeitoAjuste, efeito_id: str) -> ET.Element:
"""Monta o ``<filter-video>``/``<filter-audio>`` de um efeito."""
tag = "filter-video" if efeito.tipo == "video" else "filter-audio"
filtro = ET.Element(tag, ref=efeito_id, name=efeito.nome)
for parametro in efeito.parametros:
param = ET.SubElement(filtro, "param")
param.set("name", parametro.nome)
if parametro.chave:
param.set("key", parametro.chave)
param.set("value", parametro.valor)
if parametro.metadado:
param.set("metadata", parametro.metadado)
return filtro
+94
View File
@@ -12,6 +12,7 @@ from ..models import (
TimeValue, TimeValue,
) )
from .helpers import ( from .helpers import (
_dtd_insert,
_sanitize_xml_value, _sanitize_xml_value,
) )
@@ -122,6 +123,99 @@ class CompoundMixin:
return ref_clip return ref_clip
def wrap_titles_in_compound(
self,
parent_clip: ET.Element,
titles: List[ET.Element],
name: str = "Legenda",
) -> ET.Element:
"""Pack lane-nested *titles* of *parent_clip* into one compound clip.
A dynamic-subtitle sub-phrase is a dozen overlapping ``<title>``
elements stacked across as many lanes — legible on screen, unreadable
in the timeline. Collapsing each sub-phrase into a single compound
gives one bar per phrase to drag, mute or retime as a unit.
Mirrors the structure Final Cut itself produces for "New Compound
Clip" over stacked titles: the earliest title becomes the compound's
spine anchor at offset 0, the rest hang off it as lane children, and
a ``<ref-clip>`` takes their place in *parent_clip* on the anchor's
original lane.
Child offsets are rebased from *parent_clip*'s source-time space onto
the anchor's, since a lane child is anchored at its parent's
``start`` — leaving them untouched would shift every word of the
phrase by the gap between the two starts.
Args:
parent_clip: The spine clip the titles currently hang off.
titles: The ``<title>`` elements to pack; must all be direct
children of *parent_clip*.
name: Name for the resulting compound clip.
Returns:
The created ``<ref-clip>`` element, now in *parent_clip*.
"""
if not titles:
raise ValueError("No titles to wrap")
resources = self.root.find('.//resources')
if resources is None:
raise ValueError("No <resources> element found in FCPXML")
ordered = sorted(
titles, key=lambda t: self._parse_time(t.get('offset', '0s'))
)
anchor = ordered[0]
anchor_offset = self._parse_time(anchor.get('offset', '0s'))
anchor_start = self._parse_time(anchor.get('start', '0s'))
anchor_lane = anchor.get('lane')
total = TimeValue.zero()
for title in ordered:
rel = self._parse_time(title.get('offset', '0s')) - anchor_offset
end = rel + self._parse_time(title.get('duration', '0s'))
if end > total:
total = end
format_id = next(iter(self.formats), None) or 'r1'
media_id = self._unique_resource_id(resources, 'r_compound1')
media = ET.SubElement(resources, 'media')
media.set('id', media_id)
media.set('name', _sanitize_xml_value(name, 512))
media.set('uid', str(uuid.uuid4()).upper())
seq = ET.SubElement(media, 'sequence')
seq.set('format', format_id)
seq.set('duration', total.to_fcpxml())
seq.set('tcStart', '0s')
seq.set('tcFormat', 'NDF')
inner_spine = ET.SubElement(seq, 'spine')
for title in ordered:
parent_clip.remove(title)
anchor.set('offset', '0s')
if anchor_lane is not None:
del anchor.attrib['lane']
inner_spine.append(anchor)
for lane, title in enumerate(ordered[1:], start=1):
rel = self._parse_time(title.get('offset', '0s')) - anchor_offset
title.set('offset', (anchor_start + rel).to_fcpxml())
title.set('lane', str(lane))
anchor.append(title)
ref_clip = ET.Element('ref-clip')
ref_clip.set('ref', media_id)
if anchor_lane is not None:
ref_clip.set('lane', anchor_lane)
ref_clip.set('offset', anchor_offset.to_fcpxml())
ref_clip.set('name', _sanitize_xml_value(name, 512))
ref_clip.set('duration', total.to_fcpxml())
_dtd_insert(parent_clip, ref_clip)
return ref_clip
def flatten_compound_clip( def flatten_compound_clip(
self, self,
ref_clip_id: str, ref_clip_id: str,
+2
View File
@@ -97,6 +97,8 @@ class ModifierCore:
self.fps = self._detect_fps() self.fps = self._detect_fps()
# Lazily filled on the first generated title; see _unique_text_style_id. # Lazily filled on the first generated title; see _unique_text_style_id.
self._text_style_ids: Optional[set] = None self._text_style_ids: Optional[set] = None
# Lazily filled on the first clip split/cut; see _unique_tracking_shape_id.
self._tracking_shape_ids: Optional[set] = None
self._build_resource_index() self._build_resource_index()
self._build_clip_index() self._build_clip_index()
+50 -16
View File
@@ -41,6 +41,15 @@ class CutMixin:
cut (silence removal, filler removal) duplicates it into every cut (silence removal, filler removal) duplicates it into every
resulting piece, so the same word shows up several times across the resulting piece, so the same word shows up several times across the
edited timeline instead of once where it was placed. edited timeline instead of once where it was placed.
A lane-nested ``<video>`` zoom (the "Clipe de Ajuste" adjustment
layer ``add_zoom`` creates, ``role`` starting with ``"adjustments."``)
is the exact same phantom-duplicate case, keyed on ``offset``+
``duration`` like a keyword. Left unfiltered, every further cut
duplicates the zoom into every resulting piece with its original
offset untouched — each copy then draws at the same absolute
position, so two "Clipe de Ajuste" bars appear stacked on top of
each other in the timeline instead of the one real zoom window.
""" """
seg_end = seg_start + seg_duration seg_end = seg_start + seg_duration
to_remove = [] to_remove = []
@@ -54,6 +63,12 @@ class CutMixin:
title_offset = TimeValue.from_timecode(child.get('offset', '0s')) title_offset = TimeValue.from_timecode(child.get('offset', '0s'))
if title_offset < seg_start or title_offset >= seg_end: if title_offset < seg_start or title_offset >= seg_end:
to_remove.append(child) to_remove.append(child)
elif tag == 'video' and (child.get('role') or '').startswith('adjustments.'):
v_offset = TimeValue.from_timecode(child.get('offset', '0s'))
v_dur = TimeValue.from_timecode(child.get('duration', '0s'))
v_end = v_offset + v_dur
if v_end <= seg_start or v_offset >= seg_end:
to_remove.append(child)
elif tag == 'keyword': elif tag == 'keyword':
kw_start = TimeValue.from_timecode(child.get('start', '0s')) kw_start = TimeValue.from_timecode(child.get('start', '0s'))
kw_dur = TimeValue.from_timecode(child.get('duration', '0s')) kw_dur = TimeValue.from_timecode(child.get('duration', '0s'))
@@ -125,6 +140,7 @@ class CutMixin:
new_clip, current_start, segment_duration new_clip, current_start, segment_duration
) )
self._reassign_text_style_ids(new_clip) self._reassign_text_style_ids(new_clip)
self._reassign_tracking_shape_ids(new_clip)
spine.insert(clip_index + len(new_clips), new_clip) spine.insert(clip_index + len(new_clips), new_clip)
new_clips.append(new_clip) new_clips.append(new_clip)
@@ -195,23 +211,40 @@ class CutMixin:
if cursor < clip_duration: if cursor < clip_duration:
keeps.append((cursor, clip_duration)) keeps.append((cursor, clip_duration))
# A keep segment shorter than a couple frames at the very start or # A keep segment shorter than MIN_KEEP_SECONDS is leftover between
# end of the clip is just leftover cut padding with no neighboring # two cuts, not a real clip — at the very start/end of the clip it's
# kept audio on its outer side (the silence butts against the clip's # cut padding with no kept audio on the outer side; in the interior
# own edge) — not a real clip. Rather than emit it as its own # it's the pause BETWEEN two things that were both cut (e.g. two
# near-invisible micro-clip, fold it into the adjacent real segment, # consecutive deactivated phrases in the voice-editing flow), which
# which simply starts earlier / ends later to absorb it. # belongs to neither side by construction. At the edges we fold it
min_keep_seconds = 2 * float(self.frame_duration_fraction()) # into the one neighboring KEEP segment there is, which simply starts
if len(keeps) > 1: # earlier / ends later to absorb it. In the interior both neighbors
first_start, first_end = keeps[0] # are CUT, not keep, so there is nothing to fold into — it is just
if (first_end - first_start).to_seconds() < min_keep_seconds: # dropped, extending the surrounding cut across it instead of
keeps[1] = (first_start, keeps[1][1]) # surviving as a third near-invisible micro-clip.
#
# The threshold is bigger than one frame on purpose: measured on a
# real voice-edit (0.07-0.23s residues), a single frame did not catch
# them — this is pause/padding leftover, not intentional short
# content, so treating anything under a third of a second this way
# is safe for this cut path.
min_keep_seconds = max(6 * float(self.frame_duration_fraction()), 0.3)
i = 0
while len(keeps) > 1 and i < len(keeps):
start, end = keeps[i]
if (end - start).to_seconds() >= min_keep_seconds:
i += 1
continue
if i == 0:
keeps[1] = (start, keeps[1][1])
keeps.pop(0) keeps.pop(0)
if len(keeps) > 1: elif i == len(keeps) - 1:
last_start, last_end = keeps[-1] keeps[i - 1] = (keeps[i - 1][0], end)
if (last_end - last_start).to_seconds() < min_keep_seconds: keeps.pop(i)
keeps[-2] = (keeps[-2][0], last_end) else:
keeps.pop() keeps.pop(i)
# Re-check the same index: the segment now there might itself be
# short enough to absorb again (two short keeps in a row).
spine.remove(clip) spine.remove(clip)
new_clips: List[ET.Element] = [] new_clips: List[ET.Element] = []
@@ -226,6 +259,7 @@ class CutMixin:
new_clip.set('duration', seg_duration.to_fcpxml()) new_clip.set('duration', seg_duration.to_fcpxml())
self._filter_children_for_segment(new_clip, seg_start, seg_duration) self._filter_children_for_segment(new_clip, seg_start, seg_duration)
self._reassign_text_style_ids(new_clip) self._reassign_text_style_ids(new_clip)
self._reassign_tracking_shape_ids(new_clip)
spine.insert(clip_index + len(new_clips), new_clip) spine.insert(clip_index + len(new_clips), new_clip)
new_clips.append(new_clip) new_clips.append(new_clip)
current_offset = current_offset + seg_duration current_offset = current_offset + seg_duration
+287 -27
View File
@@ -3,6 +3,7 @@
Extraído de writer.py — ver fcpxml/writer/__init__.py para o conjunto. Extraído de writer.py — ver fcpxml/writer/__init__.py para o conjunto.
""" """
import random
import re import re
import unicodedata import unicodedata
import uuid import uuid
@@ -20,13 +21,99 @@ from ..text_layout import (
compose_sentence, compose_sentence,
layout_sentence, layout_sentence,
) )
from ..transcribe import group_words_by_segment from ..transcribe import group_words_by_segment, split_into_subphrases
from .helpers import _dtd_insert, _sanitize_xml_value from .helpers import _dtd_insert, _sanitize_xml_value
class TitlesMixin: class TitlesMixin:
"""Títulos de texto e legendas dinâmicas.""" """Títulos de texto e legendas dinâmicas."""
_SUBTITLE_METADATA_KEY = 'com.gart.subtitle.kind'
def mark_generated_subtitle(self, element: ET.Element, kind: str) -> None:
metadata = element.find('metadata')
if metadata is None:
metadata = ET.Element('metadata')
_dtd_insert(element, metadata)
ET.SubElement(metadata, 'md', key=self._SUBTITLE_METADATA_KEY, value=kind)
def _generated_subtitle_kind(self, element: ET.Element) -> Optional[str]:
marker = element.find(f"metadata/md[@key='{self._SUBTITLE_METADATA_KEY}']")
if marker is not None:
return marker.get('value')
# Recognize the exact signature of older G-ART exports. A role alone
# is not ownership: users also assign these roles to manual titles.
if element.tag == 'title':
if element.get('start') != self._TEXT_TITLE_START:
return None
effect = self.root.find(f".//resources/effect[@id='{element.get('ref')}']")
if effect is None or effect.get('uid') != self._TEXT_TITLE_UID:
return None
if re.fullmatch(r'caption_[0-9a-f]{8}', element.get('name', '')):
return 'dynamic'
text = ''.join(element.findtext('text/text-style', ''))
if (element.get('role') == 'titles.convencionais'
and element.get('lane') == '20'
and element.get('name') == f'{text} - Text'):
return 'plain'
elif element.tag == 'ref-clip':
media = self.root.find(f".//resources/media[@id='{element.get('ref')}']")
if media is not None:
titles = media.findall('.//title')
if titles and all(self._generated_subtitle_kind(t) == 'dynamic' for t in titles):
return 'dynamic'
return None
def remove_generated_subtitles(self, parent: ET.Element, kinds: tuple) -> None:
"""Replace only our own captions, preserving unrelated graphics."""
resources = self.root.find('.//resources')
for child in list(parent):
if self._generated_subtitle_kind(child) not in kinds:
continue
parent.remove(child)
if child.tag == 'ref-clip' and resources is not None:
ref = child.get('ref')
if not self.root.findall(f".//ref-clip[@ref='{ref}']"):
media = resources.find(f"media[@id='{ref}']")
if media is not None:
resources.remove(media)
def suppress_plain_under_dynamic(self, parent: ET.Element) -> None:
"""Keep generated plain titles only on frames without dynamic text."""
import copy
windows = []
for child in parent:
if self._generated_subtitle_kind(child) == 'dynamic':
start = self._parse_time(child.get('offset', '0s'))
windows.append((start, start + self._parse_time(child.get('duration', '0s'))))
for title in list(parent):
if self._generated_subtitle_kind(title) != 'plain':
continue
start = self._parse_time(title.get('offset', '0s'))
end = start + self._parse_time(title.get('duration', '0s'))
remaining = [(start, end)]
for lo, hi in windows:
parts = []
for a, b in remaining:
if a < hi and lo < b:
if a < lo:
parts.append((a, lo))
if hi < b:
parts.append((hi, b))
else:
parts.append((a, b))
remaining = parts
if remaining == [(start, end)]:
continue
parent.remove(title)
for a, b in remaining:
part = copy.deepcopy(title)
self._reassign_text_style_ids(part)
part.set('offset', a.to_fcpxml())
part.set('duration', (b - a).to_fcpxml())
_dtd_insert(parent, part)
# DYNAMIC (KARAOKE-STYLE) SUBTITLES # DYNAMIC (KARAOKE-STYLE) SUBTITLES
# ======================================================================== # ========================================================================
@@ -164,6 +251,44 @@ class TitlesMixin:
for ref_el in clip.findall(f".//text-style[@ref='{old_id}']"): for ref_el in clip.findall(f".//text-style[@ref='{old_id}']"):
ref_el.set('ref', new_id) ref_el.set('ref', new_id)
def _unique_tracking_shape_id(self, base: str) -> str:
"""Return a document-unique ``id`` for a ``<tracking-shape>``."""
stem = base or "tr"
if self._tracking_shape_ids is None:
self._tracking_shape_ids = {
ts.get('id') for ts in self.root.findall('.//tracking-shape')
}
candidate = f"{stem}_0"
counter = 0
while candidate in self._tracking_shape_ids:
counter += 1
candidate = f"{stem}_{counter}"
self._tracking_shape_ids.add(candidate)
return candidate
def _reassign_tracking_shape_ids(self, clip: ET.Element) -> None:
"""Give every ``<tracking-shape>`` inside a just-deepcopy'd *clip* a
fresh document-unique id.
Same mechanism as ``_reassign_text_style_ids``: ``split_clip``/
``cut_clip_ranges`` deepcopy the clip once per resulting segment, so
Cinematic object-tracking data (``<object-tracker><tracking-shape
id="tr1">``, preserved from the source asset's sidecar) keeps the
exact same id in every copy. A single cut is harmless — but the
batch chain re-cuts the same clip at each step, multiplying the
duplicate until the DTD validator rejects the file with "ID tr1
already defined".
"""
for shape in clip.findall('.//tracking-shape'):
old_id = shape.get('id')
if not old_id:
continue
base = re.sub(r'_\d+$', '', old_id)
new_id = self._unique_tracking_shape_id(base)
if new_id == old_id:
continue
shape.set('id', new_id)
def _make_text_title_clip( def _make_text_title_clip(
self, self,
effect_id: str, effect_id: str,
@@ -183,6 +308,7 @@ class TitlesMixin:
font_scale: float = TEXT_TEMPLATE_FONT_SCALE, font_scale: float = TEXT_TEMPLATE_FONT_SCALE,
animated: bool = True, animated: bool = True,
size_param: Optional[float] = None, size_param: Optional[float] = None,
role: Optional[str] = None,
) -> ET.Element: ) -> ET.Element:
"""Build a standalone ``<title>`` clip from the "Text" (Basic Text) template. """Build a standalone ``<title>`` clip from the "Text" (Basic Text) template.
@@ -200,6 +326,8 @@ class TitlesMixin:
elem.set('name', _sanitize_xml_value(name, 256)) elem.set('name', _sanitize_xml_value(name, 256))
elem.set('start', self._TEXT_TITLE_START) elem.set('start', self._TEXT_TITLE_START)
elem.set('duration', duration.to_fcpxml()) elem.set('duration', duration.to_fcpxml())
if role:
elem.set('role', _sanitize_xml_value(role, 256))
if position: if position:
param = ET.SubElement(elem, 'param') param = ET.SubElement(elem, 'param')
@@ -290,6 +418,7 @@ class TitlesMixin:
animated: bool = True, animated: bool = True,
font_scale: float = TEXT_TEMPLATE_FONT_SCALE, font_scale: float = TEXT_TEMPLATE_FONT_SCALE,
size_param: Optional[float] = None, size_param: Optional[float] = None,
role: Optional[str] = None,
) -> ET.Element: ) -> ET.Element:
"""Add a single static "Text" (Basic Text) title over *parent_clip*. """Add a single static "Text" (Basic Text) title over *parent_clip*.
@@ -327,6 +456,7 @@ class TitlesMixin:
animated=animated, animated=animated,
font_scale=font_scale, font_scale=font_scale,
size_param=size_param, size_param=size_param,
role=role,
) )
_dtd_insert(parent, title) _dtd_insert(parent, title)
return title return title
@@ -337,6 +467,11 @@ class TitlesMixin:
words: List[Dict[str, Any]], words: List[Dict[str, Any]],
config: Optional['DynamicSubtitleConfig'] = None, config: Optional['DynamicSubtitleConfig'] = None,
segments: Optional[List[Dict[str, Any]]] = None, segments: Optional[List[Dict[str, Any]]] = None,
role: Optional[str] = None,
configs: Optional[List['DynamicSubtitleConfig']] = None,
compound_subphrases: bool = False,
subphrase_min_words: int = 3,
hold_between_sentences: bool = True,
) -> List[ET.Element]: ) -> List[ET.Element]:
"""Generate progressive-reveal subtitle titles, one per word. """Generate progressive-reveal subtitle titles, one per word.
@@ -380,8 +515,16 @@ class TitlesMixin:
Returns: Returns:
The list of created ``<title>`` elements, in chronological order. The list of created ``<title>`` elements, in chronological order.
""" """
if config is None: # ``configs`` (a list of registered, active layouts) takes precedence
config = DynamicSubtitleConfig() # over the single ``config`` — with 2+ items, each block picks one at
# random below; with 0 or 1, behaviour is identical to a single fixed
# config, so old callers passing only ``config`` are unaffected.
if configs:
layout_configs = list(configs)
elif config is not None:
layout_configs = [config]
else:
layout_configs = [DynamicSubtitleConfig()]
if not words: if not words:
return [] return []
@@ -411,35 +554,54 @@ class TitlesMixin:
# next block — the sub-sentence split that keeps long sentences from # next block — the sub-sentence split that keeps long sentences from
# spilling off screen. # spilling off screen.
sentences = group_words_by_segment(words, segments or []) sentences = group_words_by_segment(words, segments or [])
box = LayoutBox.for_frame( # A comma is where the sentence breathes, so it is also where the
self.frame_width(), self.frame_height(), # phrase should be packed into its own compound clip downstream.
band_height=config.band_height, if compound_subphrases:
center_y=config.block_center_y, sentences = [
) sub
# "phrase" is the progressive composition the reference reel uses: one for sentence in sentences
# title per LINE ("que vão" / "melhorar" / "sua legenda"), the key word for sub in split_into_subphrases(sentence, subphrase_min_words)
# set large in a display italic. "word" is the older one-title-per-word ]
# rhythm, kept for callers that want every word to land on its own.
phrase_mode = getattr(config, 'granularity', 'phrase') == 'phrase'
def lay_out(pending: List[Dict]): def box_for(cfg: 'DynamicSubtitleConfig') -> LayoutBox:
return LayoutBox.for_frame(
self.frame_width(), self.frame_height(),
band_height=cfg.band_height,
center_y=cfg.block_center_y,
)
def lay_out(pending: List[Dict], cfg: 'DynamicSubtitleConfig', box: LayoutBox):
"""Place what fits; return (units, still-unplaced words).""" """Place what fits; return (units, still-unplaced words)."""
if phrase_mode: # "phrase" is the progressive composition the reference reel uses:
# one title per LINE ("que vão" / "melhorar" / "sua legenda"), the
# key word set large in a display italic. "word" is the older
# one-title-per-word rhythm, kept for callers that want every word
# to land on its own.
if getattr(cfg, 'granularity', 'phrase') == 'phrase':
composition = compose_sentence( composition = compose_sentence(
pending, config.style, box, line_gap=config.line_gap, pending, cfg.style, box, line_gap=cfg.line_gap,
) )
return composition.blocks, composition.overflow return composition.blocks, composition.overflow
layout = layout_sentence(pending, config.style, box) layout = layout_sentence(pending, cfg.style, box)
return layout.placed, layout.overflow return layout.placed, layout.overflow
blocks: List[List[Any]] = [] blocks: List[List[Any]] = []
for sentence in sentences: block_configs: List['DynamicSubtitleConfig'] = []
block_sentences: List[int] = []
for sentence_index, sentence in enumerate(sentences):
remaining = list(sentence) remaining = list(sentence)
while remaining: while remaining:
units, remaining = lay_out(remaining) # Each block independently samples a layout from the active
# set — the visual variety the user asked for. A single
# active layout always resolves to itself, so this is a
# no-op for the common case.
active_config = layout_configs[random.randrange(len(layout_configs))]
units, remaining = lay_out(remaining, active_config, box_for(active_config))
if not units: if not units:
break break
blocks.append(units) blocks.append(units)
block_configs.append(active_config)
block_sentences.append(sentence_index)
if not blocks: if not blocks:
return [] return []
@@ -463,6 +625,9 @@ class TitlesMixin:
for i, units in enumerate(blocks): for i, units in enumerate(blocks):
if i + 1 < len(blocks): if i + 1 < len(blocks):
end = block_starts[i + 1] end = block_starts[i + 1]
if not hold_between_sentences and block_sentences[i] != block_sentences[i + 1]:
spoken_end = self.snap_seconds_to_frame(max(unit.end for unit in units))
end = min(end, spoken_end)
else: else:
end = self.snap_seconds_to_frame( end = self.snap_seconds_to_frame(
max(unit.end for unit in units) max(unit.end for unit in units)
@@ -485,7 +650,15 @@ class TitlesMixin:
media_origin = self._parse_time(parent.get('start', '0s')) media_origin = self._parse_time(parent.get('start', '0s'))
created: List[ET.Element] = [] created: List[ET.Element] = []
for units, block_end in zip(blocks, block_ends): by_sentence: Dict[int, List[ET.Element]] = {}
for units, block_end, block_config, sentence_index in zip(
blocks, block_ends, block_configs, block_sentences
):
# A ``titles.*`` sub-role keeps these as titles (never closed
# captions) while grouping them in the role index and tinting
# their lane. An explicit ``role`` argument overrides every
# block; otherwise each block uses its own sampled layout's role.
block_role = role or getattr(block_config, "role", None) or "titles.dinamicas"
for index, unit in enumerate(units): for index, unit in enumerate(units):
relative_offset = self.snap_seconds_to_frame(unit.start) relative_offset = self.snap_seconds_to_frame(unit.start)
duration = block_end - relative_offset duration = block_end - relative_offset
@@ -505,19 +678,36 @@ class TitlesMixin:
duration, duration,
lane=lane, lane=lane,
name=f"caption_{uuid.uuid4().hex[:8]}", name=f"caption_{uuid.uuid4().hex[:8]}",
position=unit.position_param(config.text_scale), position=unit.position_param(block_config.text_scale),
font=unit.font or config.style.font, font=unit.font or block_config.style.font,
font_size=int(round(unit.font_size)), font_size=int(round(unit.font_size)),
font_color=unit.color or config.style.active_color, font_color=unit.color or block_config.style.active_color,
bold=config.style.bold, bold=block_config.style.bold,
face=unit.face, face=unit.face,
kerning=unit.kerning, kerning=unit.kerning,
font_scale=config.text_scale, font_scale=block_config.text_scale,
role=block_role,
) )
_dtd_insert(parent, title) _dtd_insert(parent, title)
self.mark_generated_subtitle(title, 'dynamic')
created.append(title) created.append(title)
by_sentence.setdefault(sentence_index, []).append(title)
if getattr(config, 'validate', False): # One compound per sub-phrase: a dozen stacked title bars collapse
# into a single one that can be dragged, muted or retimed as a unit.
if compound_subphrases:
for sentence_index in sorted(by_sentence):
group = by_sentence[sentence_index]
label = " ".join(
str(w.get('word') or w.get('text') or '')
for w in sentences[sentence_index]
).strip()
compound = self.wrap_titles_in_compound(
parent, group, name=label[:60] or "Legenda"
)
self.mark_generated_subtitle(compound, 'dynamic')
if any(getattr(cfg, 'validate', False) for cfg in layout_configs):
report = self.validate_subtitle_layout() report = self.validate_subtitle_layout()
if blocking(report["severity"]): if blocking(report["severity"]):
raise ValueError( raise ValueError(
@@ -548,8 +738,78 @@ class TitlesMixin:
Returns the ``collision.validate_titles`` report: ``severity`` (worst Returns the ``collision.validate_titles`` report: ``severity`` (worst
bucket), ``issues`` (spec-16 occurrences) and ``summary`` (counts). bucket), ``issues`` (spec-16 occurrences) and ``summary`` (counts).
""" """
# A compound clip carries its own time origin: a title inside one is
# offset from that compound's start, not the sequence's. Measured in
# one flat pass, the anchors of two different compounds both read as
# "0s" and collide on paper while sitting seconds apart on the
# timeline. Each compound is therefore measured as its own scope,
# which is also where its titles can actually overlap — a title can
# only share the screen with its own compound's siblings.
scopes: List[List[ET.Element]] = []
nested: set = set()
for media in self.root.findall('.//media'):
group = list(media.iter('title'))
if group:
scopes.append(group)
nested.update(id(t) for t in group)
main = [t for t in self.root.iter('title') if id(t) not in nested]
if main:
scopes.append(main)
reports = [
self._measure_title_scope(
scope,
safe_margin_x=safe_margin_x,
safe_margin_y=safe_margin_y,
min_font_size=min_font_size,
min_distance=min_distance,
max_distance=max_distance,
)
for scope in scopes
]
if len(reports) == 1:
return reports[0]
if not reports:
return self._measure_title_scope(
[],
safe_margin_x=safe_margin_x,
safe_margin_y=safe_margin_y,
min_font_size=min_font_size,
min_distance=min_distance,
max_distance=max_distance,
)
rank = {
'none': 0, 'render_tolerance': 1, 'warning': 2,
'probable': 3, 'severe': 4,
}
merged_issues = [i for r in reports for i in r['issues']]
summary = dict(reports[0]['summary'])
for r in reports[1:]:
for key, value in r['summary'].items():
summary[key] = summary.get(key, 0) + value
return {
'severity': max(
(r['severity'] for r in reports),
key=lambda s: rank.get(s, 0),
),
'issues': merged_issues,
'summary': summary,
}
def _measure_title_scope(
self,
elements: List[ET.Element],
*,
safe_margin_x: float,
safe_margin_y: float,
min_font_size: Optional[float],
min_distance: Optional[float],
max_distance: Optional[float],
) -> dict:
"""Measure and validate one group of titles sharing a time origin."""
titles = [] titles = []
for elem in self.root.iter('title'): for elem in elements:
# enabled="0" never renders in Final Cut (see # enabled="0" never renders in Final Cut (see
# generate_subtitles_by_emphasis, which disables plain titles # generate_subtitles_by_emphasis, which disables plain titles
# under an emphasis phrase instead of never creating them) — a # under an emphasis phrase instead of never creating them) — a
+1 -1
View File
@@ -76,7 +76,7 @@ target-version = ['py310']
[tool.ruff] [tool.ruff]
line-length = 100 line-length = 100
exclude = ["docs/", "WHISPERX/"] exclude = ["docs/"]
[tool.ruff.lint] [tool.ruff.lint]
select = ["E", "F", "I", "N", "W"] select = ["E", "F", "I", "N", "W"]
+1 -1
View File
@@ -155,7 +155,7 @@ TOOLS = [
"filepath": {"type": "string", "description": "Path to FCPXML file"}, "filepath": {"type": "string", "description": "Path to FCPXML file"},
"noise_db": {"type": "number", "description": "Silence threshold in dBFS, -120 to 0. Falls back to the saved silence settings (default -30)"}, "noise_db": {"type": "number", "description": "Silence threshold in dBFS, -120 to 0. Falls back to the saved silence settings (default -30)"},
"min_silence": {"type": "number", "description": "Minimum silence duration in seconds to cut. Falls back to the saved silence settings (default 0.5)"}, "min_silence": {"type": "number", "description": "Minimum silence duration in seconds to cut. Falls back to the saved silence settings (default 0.5)"},
"padding": {"type": "number", "description": "Seconds of silence to keep on each side of a cut so edits breathe (max 5). Falls back to the saved silence settings (default 0.05)"}, "padding": {"type": "number", "description": "Seconds of silence to keep on each side of a cut so edits breathe (max 5). Falls back to the saved silence settings (default 0.2)"},
"clip_name": {"type": "string", "description": "Only cut silence in the clip with this name"}, "clip_name": {"type": "string", "description": "Only cut silence in the clip with this name"},
"output_path": {"type": "string", "description": "Output path (default: adds _silence_removed suffix)"}, "output_path": {"type": "string", "description": "Output path (default: adds _silence_removed suffix)"},
}, },
+196 -97
View File
@@ -13,7 +13,10 @@ from typing import Sequence
from mcp.types import TextContent, Tool from mcp.types import TextContent, Tool
from fcpxml.media_intel import media_src_to_path from fcpxml.media_intel import media_src_to_path
from fcpxml.model_manager import load_dynamic_subtitle_config, load_plain_subtitle_config from fcpxml.model_manager import (
get_active_dynamic_subtitle_layouts,
load_plain_subtitle_config,
)
from fcpxml.models import DynamicSubtitleConfig, WordLook, WordStyle from fcpxml.models import DynamicSubtitleConfig, WordLook, WordStyle
from fcpxml.writer import FCPXMLModifier from fcpxml.writer import FCPXMLModifier
from server_tools._shared import ( from server_tools._shared import (
@@ -61,6 +64,7 @@ TOOLS = [
"emphasis_size": {"type": "integer", "description": "Key-word size in canvas points, at the 2160x3840 reference frame. Falls back to the saved style (default 265)"}, "emphasis_size": {"type": "integer", "description": "Key-word size in canvas points, at the 2160x3840 reference frame. Falls back to the saved style (default 265)"},
"emphasis_color": {"type": "string", "description": "RGBA (0-1, space-separated) for the key word (phrase mode). Defaults to active_color, so the block reads in a single colour unless the key word is deliberately set apart"}, "emphasis_color": {"type": "string", "description": "RGBA (0-1, space-separated) for the key word (phrase mode). Defaults to active_color, so the block reads in a single colour unless the key word is deliberately set apart"},
"text_scale": {"type": "number", "description": "Ratio between the title template's fontSize space and the canvas-point space it positions in. The \"Text\" template sizes type in frame pixels, so sizes are doubled on the way out. Falls back to the saved style (default 2.0). Lower it only if a template renders type larger than the chosen point size"}, "text_scale": {"type": "number", "description": "Ratio between the title template's fontSize space and the canvas-point space it positions in. The \"Text\" template sizes type in frame pixels, so sizes are doubled on the way out. Falls back to the saved style (default 2.0). Lower it only if a template renders type larger than the chosen point size"},
"role": {"type": "string", "description": "Final Cut role for every generated title (a 'titles.*' sub-role, never 'subtitles.*'). Falls back to the saved style (default 'titles.dinamicas'). Groups the clips in the role index and tints their lane."},
"font": {"type": "string", "description": "Title font family (supporting lines in phrase mode). Falls back to the saved style (default 'Helvetica Neue')"}, "font": {"type": "string", "description": "Title font family (supporting lines in phrase mode). Falls back to the saved style (default 'Helvetica Neue')"},
"font_size": {"type": "integer", "description": "Supporting-line font size in canvas points, at the 2160x3840 reference frame. Falls back to the saved style (default 104)"}, "font_size": {"type": "integer", "description": "Supporting-line font size in canvas points, at the 2160x3840 reference frame. Falls back to the saved style (default 104)"},
"active_color": {"type": "string", "description": "RGBA (0-1, space-separated) for even-indexed lines. Falls back to the saved style (default '1 1 1 1')"}, "active_color": {"type": "string", "description": "RGBA (0-1, space-separated) for even-indexed lines. Falls back to the saved style (default '1 1 1 1')"},
@@ -88,6 +92,7 @@ TOOLS = [
"uppercase": {"type": "boolean", "description": "Render text in uppercase."}, "uppercase": {"type": "boolean", "description": "Render text in uppercase."},
"keep_punctuation": {"type": "boolean", "description": "Keep punctuation such as comma and period."}, "keep_punctuation": {"type": "boolean", "description": "Keep punctuation such as comma and period."},
"text_scale": {"type": "number", "description": "Template font-size scale. Falls back to saved plain-subtitle config."}, "text_scale": {"type": "number", "description": "Template font-size scale. Falls back to saved plain-subtitle config."},
"role": {"type": "string", "description": "Final Cut role for every generated title (a 'titles.*' sub-role, never 'subtitles.*'). Falls back to the saved style (default 'titles.convencionais'). Groups the clips in the role index and tints their lane."},
"output_path": {"type": "string", "description": "Output path (default: adds _plain_subtitles suffix)"}, "output_path": {"type": "string", "description": "Output path (default: adds _plain_subtitles suffix)"},
}, },
"required": ["filepath"] "required": ["filepath"]
@@ -95,7 +100,7 @@ TOOLS = [
), ),
Tool( Tool(
name="generate_subtitles_by_emphasis", name="generate_subtitles_by_emphasis",
description="Generate BOTH subtitle styles over the FULL clip and let them coexist by visibility, not by splitting words: plain static titles (see generate_plain_subtitles) cover every word from start to end; dynamic progressive-composition titles (see generate_dynamic_subtitles) are additionally generated for whichever whole phrases were marked as emphasis in the phrase-review step (etapa 5, zoom applied, level >= 1). Wherever a dynamic phrase is on screen, the plain titles underneath it are set enabled=\"0\" (still present in the FCPXML, editable/re-enable-able in Final Cut, just not rendered) instead of never being generated there — so disabling emphasis later never leaves a silent gap in the plain track. Reads emphasis spans from the media's cached '<media>_phrase_actions.json' (written by save_phrase_review after the app's etapa 5 review) — run the voice-editing wizard through that step first, or nothing is treated as emphasis and every title stays plain and enabled. Style knobs are the saved 'Legendas Dinâmicas'/plain-subtitle configs (~/.fcp-mcp-server/config.json); this tool does not expose per-call style overrides, only the split logic — use generate_dynamic_subtitles/generate_plain_subtitles directly if you need one-off styling.", description="Generate BOTH subtitle styles in one pass, split by word so they never coexist on the same range: dynamic progressive-composition titles (see generate_dynamic_subtitles) cover whichever whole phrases were marked as emphasis in the phrase-review step (etapa 5, zoom applied, level >= 1); plain static titles (see generate_plain_subtitles) cover every OTHER word in the clip. A plain block is simply not created where a dynamic phrase already covers — not created-then-disabled — because a disabled title still shows as its own struck-through clip in Final Cut's timeline even though it never renders, and a heavily emphasized edit ended up with dozens of dead clips cluttering the track. Trade-off: if emphasis is turned off by hand later, the plain line under it has to be regenerated, not just re-enabled. Reads emphasis spans from the media's cached '<media>_phrase_actions.json' (written by save_phrase_review after the app's etapa 5 review) — run the voice-editing wizard through that step first, or nothing is treated as emphasis and every word gets a plain title. Style knobs are the saved 'Legendas Dinâmicas'/plain-subtitle configs (~/.fcp-mcp-server/config.json); this tool does not expose per-call style overrides, only the split logic — use generate_dynamic_subtitles/generate_plain_subtitles directly if you need one-off styling.",
inputSchema={ inputSchema={
"type": "object", "type": "object",
"properties": { "properties": {
@@ -116,6 +121,31 @@ TOOLS = [
_PUNCT_RE = re.compile(r"[^\w\sÀ-ÖØ-öø-ÿ]", re.UNICODE) _PUNCT_RE = re.compile(r"[^\w\sÀ-ÖØ-öø-ÿ]", re.UNICODE)
_DEFAULT_DYNAMIC_ROLE = "titles.dinamicas"
_DEFAULT_PLAIN_ROLE = "titles.convencionais"
def _title_subrole(value: str | None, fallback: str) -> str:
"""Return a Final Cut title sub-role, never a closed-caption role."""
role = str(value or "").strip() or fallback
if role.startswith("subtitles."):
return "titles." + role.removeprefix("subtitles.")
if role == "subtitles":
return fallback
if not role.startswith("titles."):
return fallback
return role
def _separate_subtitle_roles(dynamic_role: str | None, plain_role: str | None) -> tuple[str, str]:
"""Keep normal and dynamic subtitles in distinct Final Cut role lanes."""
dynamic = _title_subrole(dynamic_role, _DEFAULT_DYNAMIC_ROLE)
plain = _title_subrole(plain_role, _DEFAULT_PLAIN_ROLE)
if dynamic == plain:
if dynamic != _DEFAULT_DYNAMIC_ROLE:
return dynamic, _DEFAULT_PLAIN_ROLE
return _DEFAULT_DYNAMIC_ROLE, _DEFAULT_PLAIN_ROLE
return dynamic, plain
def _words_overlapping_clip(words: Sequence[dict], start: float, end: float) -> list[dict]: def _words_overlapping_clip(words: Sequence[dict], start: float, end: float) -> list[dict]:
@@ -171,12 +201,14 @@ def _phrase_actions_path(media_path: str) -> Path:
return Path(media_path).with_name(f"{stem}_phrase_actions.json") return Path(media_path).with_name(f"{stem}_phrase_actions.json")
def _load_emphasis_spans(media_path: str) -> list[dict]: def _load_review_spans(media_path: str, key: str) -> list[dict]:
"""Load emphasis spans (source-media time) saved by the etapa-5 phrase review. """Load one span list (source-media time) saved by the etapa-5 phrase review.
Returns [] if the review was never run for this media — callers should treat ``key`` is ``"emphasis_spans"`` (phrases with `subtitle_dynamic` on) or
that as "nothing is emphasis yet", not as an error, since the wizard's later ``"plain_exclude_spans"`` (phrases with `subtitle_common` off). Returns []
steps are optional. if the review was never run for this media, or saved nothing under that
key — callers should treat that as "nothing marked", not as an error,
since the wizard's later steps are optional.
""" """
path = _phrase_actions_path(media_path) path = _phrase_actions_path(media_path)
if not path.is_file(): if not path.is_file():
@@ -185,10 +217,25 @@ def _load_emphasis_spans(media_path: str) -> list[dict]:
data = json.loads(path.read_text(encoding="utf-8")) data = json.loads(path.read_text(encoding="utf-8"))
except (OSError, json.JSONDecodeError): except (OSError, json.JSONDecodeError):
return [] return []
spans = data.get("emphasis_spans", []) spans = data.get(key, [])
return [s for s in spans if isinstance(s, dict) and "start" in s and "end" in s] return [s for s in spans if isinstance(s, dict) and "start" in s and "end" in s]
def _load_emphasis_spans(media_path: str) -> list[dict]:
"""Spans (source-media time) whose phrase has `subtitle_dynamic` on."""
return _load_review_spans(media_path, "emphasis_spans")
def _load_plain_exclude_spans(media_path: str) -> list[dict]:
"""Spans (source-media time) whose phrase has `subtitle_common` off.
Independent from emphasis spans: a phrase can have `subtitle_common` off
without being emphasized, so plain must be hidden there too even though
no dynamic line is going to cover the gap.
"""
return _load_review_spans(media_path, "plain_exclude_spans")
def _word_in_spans(word_start: float, word_end: float, spans: Sequence[dict]) -> bool: def _word_in_spans(word_start: float, word_end: float, spans: Sequence[dict]) -> bool:
"""A word belongs to an emphasis span if its midpoint falls inside it. """A word belongs to an emphasis span if its midpoint falls inside it.
@@ -234,7 +281,7 @@ def _segments_in_spans(segments: Sequence[dict], spans: Sequence[dict]) -> list[
def _overlaps_any_span(start: float, end: float, spans: Sequence[tuple[float, float]]) -> bool: def _overlaps_any_span(start: float, end: float, spans: Sequence[tuple[float, float]]) -> bool:
"""Half-open interval overlap: a plain title under this window must hide.""" """Half-open interval overlap: a plain title under this window is skipped."""
return any(start < span_end and end > span_start for span_start, span_end in spans) return any(start < span_end and end > span_start for span_start, span_end in spans)
@@ -302,6 +349,46 @@ async def handle_validate_subtitle_layout(arguments: dict) -> Sequence[TextConte
return _text_result("\n".join(lines)) return _text_result("\n".join(lines))
def _build_dynamic_subtitle_config(saved: dict, overrides: dict | None = None) -> DynamicSubtitleConfig:
"""Build a :class:`DynamicSubtitleConfig` from one registered layout dict.
``overrides`` (typically the tool call's own ``arguments``) only makes
sense to apply when there is a single active layout — callers with 2+
active layouts pass ``{}`` so every sampled block uses its layout as
registered, unambiguously.
"""
overrides = overrides or {}
body_color = overrides.get("active_color") or saved["active_color"]
return DynamicSubtitleConfig(
style=WordStyle(
font=overrides.get("font") or saved["font"],
font_size=int(overrides.get("font_size", saved["font_size"])),
active_color=body_color,
inactive_color=overrides.get("inactive_color", "0.7 0.7 0.7 1"),
emphasis_look=WordLook(
int(overrides.get("emphasis_size", saved["emphasis_size"])),
overrides.get("emphasis_color") or saved["emphasis_color"] or body_color,
font=overrides.get("emphasis_font") or saved["emphasis_font"],
face=overrides.get("emphasis_face") or saved["emphasis_face"],
kerning=0.0,
),
body_look=WordLook(
int(overrides.get("font_size", saved["font_size"])),
body_color,
font=overrides.get("font") or saved["font"],
face="Bold",
kerning=1.2,
),
),
band_height=float(overrides.get("band_height", saved["band_height"])),
block_center_y=float(overrides.get("block_center_y", saved["block_center_y"])),
granularity=overrides.get("granularity", "phrase"),
text_scale=float(overrides.get("text_scale", saved["text_scale"])),
line_gap=float(overrides.get("line_gap", saved["line_gap"])),
role=_title_subrole(saved.get("role"), _DEFAULT_DYNAMIC_ROLE),
)
async def handle_generate_dynamic_subtitles(arguments: dict) -> Sequence[TextContent]: async def handle_generate_dynamic_subtitles(arguments: dict) -> Sequence[TextContent]:
"""Generate per-word subtitle titles laid out as a block per sentence. """Generate per-word subtitle titles laid out as a block per sentence.
@@ -318,39 +405,21 @@ async def handle_generate_dynamic_subtitles(arguments: dict) -> Sequence[TextCon
output_dir = arguments.get("output_dir") output_dir = arguments.get("output_dir")
clip_filter = arguments.get("clip_name") clip_filter = arguments.get("clip_name")
# Anything the caller didn't explicitly pass falls back to the style # Anything the caller didn't explicitly pass falls back to the style(s)
# persisted from the "Legendas Dinâmicas" screen (~/.fcp-mcp-server/ # persisted from the "Legendas Dinâmicas" screen (~/.fcp-mcp-server/
# config.json), not a hardcoded default — so the UI is the single place # config.json) — one or more named, active layouts. With exactly one
# that configures the look, and every caller (app, MCP, this session) # active layout, per-call overrides (arguments) still apply, same as
# renders the same thing without threading 11 fields through every call. # before this screen supported multiple layouts. With 2+ active layouts,
saved = load_dynamic_subtitle_config() # each block below randomly samples one of them, so per-call overrides
body_color = arguments.get("active_color") or saved["active_color"] # are ambiguous (which layout would they apply to?) and are ignored —
config = DynamicSubtitleConfig( # register/edit the layouts themselves instead.
style=WordStyle( active_layouts = get_active_dynamic_subtitle_layouts()
font=arguments.get("font") or saved["font"], overrides = arguments if len(active_layouts) == 1 else {}
font_size=int(arguments.get("font_size", saved["font_size"])), configs = [_build_dynamic_subtitle_config(saved, overrides) for saved in active_layouts]
active_color=body_color, single_role_override = (
inactive_color=arguments.get("inactive_color", "0.7 0.7 0.7 1"), _title_subrole(arguments.get("role"), configs[0].role)
emphasis_look=WordLook( if len(active_layouts) == 1 and arguments.get("role")
int(arguments.get("emphasis_size", saved["emphasis_size"])), else None
arguments.get("emphasis_color") or saved["emphasis_color"] or body_color,
font=arguments.get("emphasis_font") or saved["emphasis_font"],
face=arguments.get("emphasis_face") or saved["emphasis_face"],
kerning=0.0,
),
body_look=WordLook(
int(arguments.get("font_size", saved["font_size"])),
body_color,
font=arguments.get("font") or saved["font"],
face="Bold",
kerning=1.2,
),
),
band_height=float(arguments.get("band_height", saved["band_height"])),
block_center_y=float(arguments.get("block_center_y", saved["block_center_y"])),
granularity=arguments.get("granularity", "phrase"),
text_scale=float(arguments.get("text_scale", saved["text_scale"])),
line_gap=float(arguments.get("line_gap", saved["line_gap"])),
) )
filepath, output_path, modifier = _setup_modifier(arguments, "_dynamic_subtitles") filepath, output_path, modifier = _setup_modifier(arguments, "_dynamic_subtitles")
@@ -401,9 +470,13 @@ async def handle_generate_dynamic_subtitles(arguments: dict) -> Sequence[TextCon
# loop to whichever one `self.clips` last indexed, stacking every # loop to whichever one `self.clips` last indexed, stacking every
# clip's captions onto a single wrong spine element instead of each # clip's captions onto a single wrong spine element instead of each
# clip's own. See Engine/docs/05_EXPERIENCIAS.md, entry 2026-08-17. # clip's own. See Engine/docs/05_EXPERIENCIAS.md, entry 2026-08-17.
modifier.remove_generated_subtitles(el, ('dynamic',))
lines = modifier.generate_dynamic_subtitles( lines = modifier.generate_dynamic_subtitles(
el, clip_words, config, segments=clip_segments el, clip_words, configs=configs, segments=clip_segments,
role=single_role_override,
compound_subphrases=True,
) )
modifier.suppress_plain_under_dynamic(el)
added.append((name, len(lines), len(clip_words))) added.append((name, len(lines), len(clip_words)))
if not added: if not added:
@@ -443,6 +516,10 @@ async def handle_generate_plain_subtitles(arguments: dict) -> Sequence[TextConte
clip_filter = arguments.get("clip_name") clip_filter = arguments.get("clip_name")
saved = load_plain_subtitle_config() saved = load_plain_subtitle_config()
saved["role"] = _title_subrole(
arguments.get("role") or saved.get("role"),
_DEFAULT_PLAIN_ROLE,
)
font = arguments.get("font") or saved["font"] font = arguments.get("font") or saved["font"]
font_size = int(arguments.get("font_size", saved["font_size"])) font_size = int(arguments.get("font_size", saved["font_size"]))
font_color = arguments.get("font_color") or saved["font_color"] font_color = arguments.get("font_color") or saved["font_color"]
@@ -479,6 +556,7 @@ async def handle_generate_plain_subtitles(arguments: dict) -> Sequence[TextConte
skipped.append((name, "no words in clip's source range")) skipped.append((name, "no words in clip's source range"))
continue continue
modifier.remove_generated_subtitles(el, ('plain',))
blocks = _plain_subtitle_blocks(clip_words, max_words) blocks = _plain_subtitle_blocks(clip_words, max_words)
created = 0 created = 0
for block in blocks: for block in blocks:
@@ -492,7 +570,7 @@ async def handle_generate_plain_subtitles(arguments: dict) -> Sequence[TextConte
start = max(0.0, min(float(w.get("start", 0.0)) for w in block)) start = max(0.0, min(float(w.get("start", 0.0)) for w in block))
end = max(float(w.get("end", start)) for w in block) end = max(float(w.get("end", start)) for w in block)
duration = max(end - start, modifier.frame_duration_fraction()) duration = max(end - start, modifier.frame_duration_fraction())
modifier.add_text_title( title = modifier.add_text_title(
el, el,
text, text,
offset=f"{start:.6f}s", offset=f"{start:.6f}s",
@@ -506,8 +584,11 @@ async def handle_generate_plain_subtitles(arguments: dict) -> Sequence[TextConte
face=None, face=None,
font_scale=1.0, font_scale=1.0,
size_param=font_size, size_param=font_size,
role=saved["role"],
) )
modifier.mark_generated_subtitle(title, 'plain')
created += 1 created += 1
modifier.suppress_plain_under_dynamic(el)
if created: if created:
added.append((name, created, len(clip_words))) added.append((name, created, len(clip_words)))
@@ -541,13 +622,16 @@ async def handle_generate_plain_subtitles(arguments: dict) -> Sequence[TextConte
async def handle_generate_subtitles_by_emphasis(arguments: dict) -> Sequence[TextContent]: async def handle_generate_subtitles_by_emphasis(arguments: dict) -> Sequence[TextContent]:
"""Generate plain titles for the whole clip and dynamic titles for the """Generate dynamic titles for the emphasis phrases, and plain titles for
emphasis phrases on top, then hide (enabled="0") the plain titles that every OTHER word — a plain block is simply not created where a dynamic
fall under a dynamic phrase — never split the word list between the two. phrase already covers, rather than created and disabled.
Plain always covers every word, so turning emphasis off later (editing A disabled ("enabled=0") title still shows as its own struck-through clip
the phrase review and re-running) never leaves a silent gap: the plain in Final Cut's timeline even though it never renders — a heavily
title was there all along, just disabled. emphasized edit ended up with dozens of dead clips cluttering the track.
Not generating them there trades that clutter for a smaller gap: if the
emphasis is turned off by hand later, the plain line has to be
regenerated rather than just re-enabled.
""" """
model = arguments.get("model", "base") model = arguments.get("model", "base")
language = arguments.get("language") language = arguments.get("language")
@@ -555,37 +639,30 @@ async def handle_generate_subtitles_by_emphasis(arguments: dict) -> Sequence[Tex
clip_filter = arguments.get("clip_name") clip_filter = arguments.get("clip_name")
granularity = arguments.get("granularity", "phrase") granularity = arguments.get("granularity", "phrase")
saved_dynamic = load_dynamic_subtitle_config() # One or more named, active layouts — with 2+ active, each emphasis block
body_color = saved_dynamic["active_color"] # below randomly samples one of them (see generate_dynamic_subtitles).
dynamic_config = DynamicSubtitleConfig( active_dynamic_layouts = get_active_dynamic_subtitle_layouts()
style=WordStyle( dynamic_configs = [
font=saved_dynamic["font"], _build_dynamic_subtitle_config(saved, {"granularity": granularity})
font_size=int(saved_dynamic["font_size"]), for saved in active_dynamic_layouts
active_color=body_color, ]
inactive_color="0.7 0.7 0.7 1", single_dynamic_role = (
emphasis_look=WordLook( active_dynamic_layouts[0]["role"] if len(active_dynamic_layouts) == 1 else None
int(saved_dynamic["emphasis_size"]),
saved_dynamic["emphasis_color"] or body_color,
font=saved_dynamic["emphasis_font"],
face=saved_dynamic["emphasis_face"],
kerning=0.0,
),
body_look=WordLook(
int(saved_dynamic["font_size"]),
body_color,
font=saved_dynamic["font"],
face="Bold",
kerning=1.2,
),
),
band_height=float(saved_dynamic["band_height"]),
block_center_y=float(saved_dynamic["block_center_y"]),
granularity=granularity,
text_scale=float(saved_dynamic["text_scale"]),
line_gap=float(saved_dynamic["line_gap"]),
) )
saved_plain = load_plain_subtitle_config() saved_plain = load_plain_subtitle_config()
dynamic_role, plain_role = _separate_subtitle_roles(
single_dynamic_role or dynamic_configs[0].role,
saved_plain.get("role"),
)
if len(active_dynamic_layouts) == 1:
single_dynamic_role = dynamic_role
else:
for cfg in dynamic_configs:
cfg.role = _title_subrole(cfg.role, _DEFAULT_DYNAMIC_ROLE)
if cfg.role == plain_role:
cfg.role = _DEFAULT_DYNAMIC_ROLE
saved_plain["role"] = plain_role
plain_font = saved_plain["font"] plain_font = saved_plain["font"]
plain_font_size = int(saved_plain["font_size"]) plain_font_size = int(saved_plain["font_size"])
plain_font_color = saved_plain["font_color"] plain_font_color = saved_plain["font_color"]
@@ -615,30 +692,41 @@ async def handle_generate_subtitles_by_emphasis(arguments: dict) -> Sequence[Tex
continue continue
spans = _load_emphasis_spans(media_path) spans = _load_emphasis_spans(media_path)
if not spans: exclude_spans = _load_plain_exclude_spans(media_path)
if not spans and not exclude_spans:
no_review.append(name) no_review.append(name)
clip_source_start = modifier.source_file_start(el).to_seconds() clip_source_start = modifier.source_file_start(el).to_seconds()
clip_duration = modifier._parse_time(el.get("duration", "0s")).to_seconds() clip_duration = modifier._parse_time(el.get("duration", "0s")).to_seconds()
window_end = clip_source_start + clip_duration window_end = clip_source_start + clip_duration
def _clip_relative(span_list: list[dict]) -> list[tuple[float, float]]:
return [
(max(0.0, float(s["start"]) - clip_source_start), min(clip_duration, float(s["end"]) - clip_source_start))
for s in span_list
if float(s["end"]) > clip_source_start and float(s["start"]) < window_end
]
# Clip-relative windows, for deciding which plain titles to hide — # Clip-relative windows, for deciding which plain titles to hide —
# same coordinate space add_text_title's offsets end up in. # same coordinate space add_text_title's offsets end up in. Dynamic
clip_spans = [ # spans hide plain (see the trade-off note below); explicit
(max(0.0, float(s["start"]) - clip_source_start), min(clip_duration, float(s["end"]) - clip_source_start)) # `subtitle_common: false` spans hide it too, even without a dynamic
for s in spans # line covering the gap.
if float(s["end"]) > clip_source_start and float(s["start"]) < window_end clip_spans = _clip_relative(spans)
] clip_hide_plain_spans = clip_spans + _clip_relative(exclude_spans)
all_words = data.get("words", []) all_words = data.get("words", [])
modifier.remove_generated_subtitles(el, ('dynamic', 'plain'))
dynamic_lines = 0 dynamic_lines = 0
dynamic_word_count = 0 dynamic_word_count = 0
emphasis_words = _words_in_spans(all_words, spans) emphasis_words = _words_in_spans(all_words, spans)
clip_emphasis_words = _words_overlapping_clip(emphasis_words, clip_source_start, window_end) clip_emphasis_words = _words_overlapping_clip(emphasis_words, clip_source_start, window_end)
if clip_emphasis_words: if clip_emphasis_words:
all_segments = data.get("segments", []) # Reviewed phrases, not broader Whisper segments, define where
emphasis_segments = _segments_in_spans(all_segments, spans) # a dynamic composition may live. Separate emphasis windows must
# never hold text over the plain speech between them.
emphasis_segments = spans
clip_segments = [ clip_segments = [
{ {
"start": float(s.get("start", 0.0)) - clip_source_start, "start": float(s.get("start", 0.0)) - clip_source_start,
@@ -653,15 +741,22 @@ async def handle_generate_subtitles_by_emphasis(arguments: dict) -> Sequence[Tex
# entry 2026-08-17). # entry 2026-08-17).
dynamic_lines = len( dynamic_lines = len(
modifier.generate_dynamic_subtitles( modifier.generate_dynamic_subtitles(
el, clip_emphasis_words, dynamic_config, segments=clip_segments el, clip_emphasis_words, configs=dynamic_configs, segments=clip_segments,
role=single_dynamic_role,
compound_subphrases=True,
hold_between_sentences=False,
) )
) )
dynamic_word_count = len(clip_emphasis_words) dynamic_word_count = len(clip_emphasis_words)
# Plain covers EVERY word in the clip — never filtered by emphasis. # Plain covers every word OUTSIDE an emphasis span. A block landing
# Titles landing under a dynamic phrase are disabled below instead of # under a dynamic phrase is simply not created there — generating it
# never being created, so turning emphasis off later never leaves a # disabled was tried first, but every disabled title still shows up
# silent gap where neither style is on screen. # as its own clip in Final Cut's timeline (just struck through), so
# a heavily-emphasized edit ended up with dozens of dead clips
# cluttering the track for no visible benefit. The trade-off: if the
# emphasis is later turned off by hand, the plain line under it has
# to be regenerated rather than just re-enabled.
plain_created = 0 plain_created = 0
plain_hidden = 0 plain_hidden = 0
clip_all_words = _words_overlapping_clip(all_words, clip_source_start, window_end) clip_all_words = _words_overlapping_clip(all_words, clip_source_start, window_end)
@@ -676,6 +771,9 @@ async def handle_generate_subtitles_by_emphasis(arguments: dict) -> Sequence[Tex
continue continue
start = max(0.0, min(float(w.get("start", 0.0)) for w in block)) start = max(0.0, min(float(w.get("start", 0.0)) for w in block))
end = max(float(w.get("end", start)) for w in block) end = max(float(w.get("end", start)) for w in block)
if _overlaps_any_span(start, end, clip_hide_plain_spans):
plain_hidden += 1
continue
duration = max(end - start, modifier.frame_duration_fraction()) duration = max(end - start, modifier.frame_duration_fraction())
title = modifier.add_text_title( title = modifier.add_text_title(
el, el,
@@ -691,11 +789,12 @@ async def handle_generate_subtitles_by_emphasis(arguments: dict) -> Sequence[Tex
face=None, face=None,
font_scale=1.0, font_scale=1.0,
size_param=plain_font_size, size_param=plain_font_size,
role=saved_plain["role"],
) )
modifier.mark_generated_subtitle(title, 'plain')
plain_created += 1 plain_created += 1
if _overlaps_any_span(start, end, clip_spans):
title.set("enabled", "0") modifier.suppress_plain_under_dynamic(el)
plain_hidden += 1
if dynamic_lines or plain_created: if dynamic_lines or plain_created:
added.append( added.append(
@@ -721,12 +820,12 @@ async def handle_generate_subtitles_by_emphasis(arguments: dict) -> Sequence[Tex
result += ( result += (
f"- **Clips Captioned**: {len(added)}\n" f"- **Clips Captioned**: {len(added)}\n"
f"- **Dynamic Title Lines (emphasis)**: {total_dynamic}\n" f"- **Dynamic Title Lines (emphasis)**: {total_dynamic}\n"
f"- **Plain Title Blocks (full clip)**: {total_plain}\n" f"- **Plain Title Blocks**: {total_plain}\n"
f"- **Plain Blocks Hidden Under Emphasis (enabled=\"0\")**: {total_hidden}\n" f"- **Plain Blocks Skipped Under Emphasis (not created there)**: {total_hidden}\n"
f"- **Total Words**: {total_words}\n\n" f"- **Total Words**: {total_words}\n\n"
) )
result += _markdown_table( result += _markdown_table(
["Clip", "Dynamic Lines", "Plain Blocks", "Hidden", "Words"], ["Clip", "Dynamic Lines", "Plain Blocks", "Skipped", "Words"],
[[n, str(d), str(p), str(h), str(w)] for n, d, p, h, w in added], [[n, str(d), str(p), str(h), str(w)] for n, d, p, h, w in added],
) )
if no_review: if no_review:
+26
View File
@@ -331,6 +331,32 @@ class TestCutClipRanges:
assert mod._parse_time(seg.get("duration")).to_seconds() == pytest.approx(4.0) assert mod._parse_time(seg.get("duration")).to_seconds() == pytest.approx(4.0)
assert removed.to_seconds() == pytest.approx(0.0) assert removed.to_seconds() == pytest.approx(0.0)
def test_interior_sliver_between_two_cuts_is_dropped_not_kept(self, tmp_path):
"""A keep segment under the threshold BETWEEN two cuts (both
neighbors already removed) is the pause between two things that were
cut, not real content — regression for a real voice-edit where
consecutive short cuts left 0.07-0.23s clips surviving between the
real ones. Unlike the edge case, there is no kept neighbor to widen:
the sliver is just dropped, extending the surrounding cut over it."""
from fcpxml.models import TimeValue
mod = self._make_modifier(tmp_path)
clip = self._spine_clips(mod)[0]
# 4s clip; cut 0.5..1.5 and 1.6..3.5 -> would-be keeps: [0..0.5],
# [1.5..1.6] (0.1s interior sliver), [3.5..4].
removed = mod.cut_clip_ranges(clip, [
(TimeValue(1, 2), TimeValue(3, 2)),
(TimeValue(8, 5), TimeValue(7, 2)),
])
clips = self._spine_clips(mod)
assert len(clips) == 3 # interview x2 real segments + broll; no 0.1s sliver
seg1, seg2, broll = clips
assert broll.get("name") == "broll"
assert mod._parse_time(seg1.get("duration")).to_seconds() == pytest.approx(0.5)
assert mod._parse_time(seg2.get("duration")).to_seconds() == pytest.approx(0.5)
assert removed.to_seconds() == pytest.approx(3.0)
def test_markers_follow_their_segment(self, tmp_path): def test_markers_follow_their_segment(self, tmp_path):
from fcpxml.models import TimeValue from fcpxml.models import TimeValue
+36
View File
@@ -260,6 +260,42 @@ class TestBackToActions:
) )
assert result["actions"] == [] assert result["actions"] == []
def test_consecutive_inactive_phrases_merge_into_one_cut(self):
"""Two deactivated phrases in a row must not leave the pause between
them (2.0-2.3 here) uncut — a phrase-by-phrase cut would strand it as
a tiny surviving sliver clip in the final timeline."""
review = build_phrase_review(
_timeline([_segment(0, 2), _segment(2.3, 4), _segment(4.5, 6)])
)
review["phrases"][0]["active"] = False
review["phrases"][1]["active"] = False
result = phrase_review_to_actions(review)
cuts = [(a["start"], a["end"]) for a in result["actions"] if a["kind"] == "cut"]
assert cuts == [(0.0, 4.0)]
def test_inactive_run_at_the_end_still_flushes(self):
"""A run of deactivated phrases with nothing active after it must
still produce its cut — regression for merging logic that only
flushed on hitting the next active phrase."""
review = build_phrase_review(_timeline([_segment(0, 2), _segment(2.3, 4)]))
review["phrases"][0]["active"] = False
review["phrases"][1]["active"] = False
result = phrase_review_to_actions(review)
cuts = [(a["start"], a["end"]) for a in result["actions"] if a["kind"] == "cut"]
assert cuts == [(0.0, 4.0)]
def test_isolated_inactive_phrases_stay_separate_cuts(self):
"""An active phrase between two inactive ones must not be swallowed —
only truly CONSECUTIVE inactive phrases merge."""
review = build_phrase_review(
_timeline([_segment(0, 2), _segment(2.3, 4), _segment(4.5, 6)])
)
review["phrases"][0]["active"] = False
review["phrases"][2]["active"] = False
result = phrase_review_to_actions(review)
cuts = [(a["start"], a["end"]) for a in result["actions"] if a["kind"] == "cut"]
assert cuts == [(0.0, 2.0), (4.5, 6.0)]
class TestManualZooms: class TestManualZooms:
def test_manual_zoom_becomes_an_action_without_a_scale(self): def test_manual_zoom_becomes_an_action_without_a_scale(self):
@@ -0,0 +1,72 @@
"""Exercise the real subtitle handler across emphasis gaps and regeneration."""
import asyncio
from pathlib import Path
import pytest
from fcpxml.models import DynamicSubtitleConfig
from fcpxml.writer import FCPXMLModifier
from server_tools import subtitles
@pytest.fixture
def caption_job(tmp_path, monkeypatch):
modifier = FCPXMLModifier(Path(__file__).parents[1] / "examples/sample.fcpxml")
parent = modifier._require_clip("Interview_A")
parent.set("start", "0s")
parent.set("duration", "10s")
media = tmp_path / "speech.mov"
media.touch()
modifier.resources[parent.get("ref")]["src"] = media.as_uri()
monkeypatch.setattr(modifier, "_iter_spine_clips", lambda: iter([(0, parent)]))
monkeypatch.setattr(subtitles, "_setup_modifier", lambda *args: (
"input.fcpxml", str(tmp_path / "output.fcpxml"), modifier,
))
monkeypatch.setattr(subtitles, "get_active_dynamic_subtitle_layouts", lambda: [
{"role": "titles.dinamicas"},
])
monkeypatch.setattr(subtitles, "_build_dynamic_subtitle_config", lambda *args: DynamicSubtitleConfig())
monkeypatch.setattr(subtitles, "load_plain_subtitle_config", lambda: {
"role": "titles.convencionais", "font": "Helvetica Neue", "font_size": 48,
"font_color": "1 1 1 1", "max_words": 1, "position_y": -100,
"uppercase": False, "keep_punctuation": True,
})
monkeypatch.setattr(subtitles, "_load_plain_exclude_spans", lambda _: [])
monkeypatch.setattr(subtitles, "_load_emphasis_spans", lambda _: [
{"start": 0, "end": 1}, {"start": 4, "end": 5},
])
monkeypatch.setattr(subtitles, "_load_or_transcribe", lambda *args: ({
"words": [
{"word": "corpo", "start": 0, "end": 1},
{"word": "muda", "start": 2, "end": 3},
{"word": "também", "start": 4, "end": 5},
],
"segments": [{"start": 0, "end": 1}, {"start": 4, "end": 5}],
}, None))
return modifier, parent
def test_dynamic_composition_clears_before_plain_words_in_emphasis_gap(caption_job):
modifier, parent = caption_job
asyncio.run(subtitles.handle_generate_subtitles_by_emphasis({}))
dynamic = parent.findall("ref-clip")
plain = parent.findall("title")
assert dynamic and plain
for composition in dynamic:
start = modifier._parse_time(composition.get("offset"))
end = start + modifier._parse_time(composition.get("duration"))
for title in plain:
plain_start = modifier._parse_time(title.get("offset"))
plain_end = plain_start + modifier._parse_time(title.get("duration"))
assert not (start < plain_end and plain_start < end), "dynamic and plain overlap"
def test_regeneration_replaces_generated_subtitles_and_preserves_manual_titles(caption_job):
modifier, parent = caption_job
manual = modifier.add_text_title(parent, "Nome da médica", role="titles.manual")
asyncio.run(subtitles.handle_generate_subtitles_by_emphasis({}))
initial = len(parent.findall("ref-clip")) + len(parent.findall("title"))
asyncio.run(subtitles.handle_generate_subtitles_by_emphasis({}))
assert len(parent.findall("ref-clip")) + len(parent.findall("title")) == initial
assert manual in list(parent)
+60
View File
@@ -0,0 +1,60 @@
"""ClipDeAjuste deve gerar filtros como filhos diretos do <clip>.
O DTD 1.13 não define nenhum elemento <adjustment> — filter-video/
filter-audio vêm direto no <clip>, depois de audio-channel-source* e antes
de metadata?. Ver Engine/docs/09_MANUTENCAO.md §2.6 e 05_EXPERIENCIAS.md.
"""
import xml.etree.ElementTree as ET
from fcpxml.models.timeline import EfeitoAjuste, ParametroEfeito
from fcpxml.models.timing import TimeValue
from fcpxml.writer.adjustment import ClipDeAjuste
def test_criar_nao_usa_wrapper_adjustment():
resources = ET.Element("resources")
efeito = EfeitoAjuste(
nome="Color Curves",
uid="FFFF0000-0000-0000-0000-000000000000",
tipo="video",
parametros=[ParametroEfeito(nome="Amount", valor="0.5", chave=".../9999")],
)
clip = ClipDeAjuste(
nome="Ajuste de cor", duracao=TimeValue(300, 30), efeitos=[efeito]
).criar(resources)
assert clip.find("adjustment") is None
filtro = clip.find("filter-video")
assert filtro is not None
assert filtro.get("name") == "Color Curves"
assert filtro in list(clip)
def test_filter_video_vem_antes_de_filter_audio():
resources = ET.Element("resources")
efeito_audio = EfeitoAjuste(
nome="Gain", uid="AAAA0000-0000-0000-0000-000000000000", tipo="audio"
)
efeito_video = EfeitoAjuste(
nome="Blur", uid="BBBB0000-0000-0000-0000-000000000000", tipo="video"
)
clip = ClipDeAjuste(
nome="Ajuste misto",
duracao=TimeValue(300, 30),
efeitos=[efeito_audio, efeito_video],
).criar(resources)
tags = [child.tag for child in clip]
assert tags == ["filter-video", "filter-audio"]
def test_criar_registra_recurso_effect_uma_vez_por_uid():
resources = ET.Element("resources")
efeito = EfeitoAjuste(
nome="Color Curves", uid="FFFF0000-0000-0000-0000-000000000000", tipo="video"
)
ClipDeAjuste(nome="A", duracao=TimeValue(300, 30), efeitos=[efeito]).criar(resources)
ClipDeAjuste(nome="B", duracao=TimeValue(300, 30), efeitos=[efeito]).criar(resources)
assert len(resources.findall("effect")) == 1
+89
View File
@@ -0,0 +1,89 @@
# RAG deste projeto (G-ART)
Banco de RAG próprio do G-ART — usado só para a IA indexar código/documentação
e responder consultas gastando menos tokens, sem precisar reler o repositório
inteiro a cada tarefa. **Não é o banco de dados do sistema**: é infraestrutura
de apoio ao desenvolvimento, mantida à parte da aplicação.
Segue o mesmo padrão dos projetos irmãos (Doza, Tigre, Jhonny): **um único
container Postgres + pgvector compartilhado** (`rag-hub-db`) na VPS da
equipe, e **um banco por sistema** dentro dele.
```
rag-hub-db (container único na VPS)
├── rag_doza ← banco do Doza
├── rag_tigre ← banco do Tigre
├── jhonny-rag ← banco do Jhonny
└── rag_gart ← banco deste projeto
```
## Divisão de responsabilidades neste projeto
Diferente dos projetos irmãos, aqui a indexação **já existia antes desta
pasta** e mora em `admin/`, não em `rag/`:
- **Indexação** — [`admin/update_rag.py`](../admin/update_rag.py), chamado
por `admin/update_rag.command` (túnel SSH + execução) e por
`admin/run.command` (roda junto com o app). Varre `INCLUDE_EXTENSIONS`
(`.command .md .py .sh .sql .swift .txt .yml .yaml`) a partir da raiz do
projeto, corta por janela de linhas (`CHUNK_LINES`), grava embeddings via
Ollama e é incremental (hash por arquivo em `gart.indexed_files`).
- **Schema** — [`schema.sql`](schema.sql) nesta pasta: é o que
`admin/update_rag.py` espera encontrar (`gart.code_chunks`,
`gart.file_index`, `gart.indexed_files`). Rodar uma vez para provisionar
um banco novo.
- **Busca** — [`search.py`](search.py) e o wrapper
[`search_gart.sh`](search_gart.sh) nesta pasta: é o que os projetos irmãos
chamam de `search_<projeto>.sh`. Não existia ainda para o G-ART.
- **Credenciais** — reaproveitadas de `admin/gart-rag.env` (mesmo arquivo que
`admin/update_rag.command` já usa), para não duplicar a senha em dois
lugares. Ver `admin/gart-rag.env.example` para o formato.
## Como a busca funciona
Duas listas em paralelo, fundidas com RRF ponderado (parâmetros herdados dos
projetos irmãos, calibrados lá via `rag/bench.py` sobre consultas douradas):
1. **densa** — embedding do trecho de código/texto;
2. **lexical** — `pg_trgm` sobre os símbolos declarados (nomes de
classe/função extraídos por regex em `admin/update_rag.py`), para
consultas que citam o nome exato de algo;
3. **resumo** (`file_index.summary_embedding`) — hoje não é preenchido por
`admin/update_rag.py` (só grava `summary` em texto, sem embedding), então
essa lista fica vazia até alguém adicionar isso ao indexador. A busca
funciona normalmente sem ela.
Cada trecho guarda `start_line`/`end_line`, então o resultado aponta a janela
exata (`code/fcpxml/writer/modifier.py:120-180`) em vez de mandar ler o
arquivo inteiro.
### Modos de saída
| Comando | O que traz |
|---|---|
| `rag/search_gart.sh "consulta"` | caminho, faixa de linhas e uma linha de descrição (padrão) |
| `… --snippet` | + 300 chars do trecho |
| `… --full` | + o trecho inteiro |
| `… --json` | saída estruturada |
| `… --module X` / `--path Y` / `--ext .py` | restringe o escopo |
| `… --map [termo]` | inventário de arquivos, sem nenhum código |
## Arquivos desta pasta
- `README.md` — este arquivo.
- `SETUP.md` — passo a passo para provisionar o banco `rag_gart` na primeira
vez.
- `schema.sql` — schema do G-ART (extensões, tabelas, índices). Estado FINAL
desejado: num banco novo basta rodá-lo.
- `embed.py` — chamada ao Ollama compartilhada entre indexador e busca (só
os prefixos `search_document:`/`search_query:` do nomic-embed-text).
- `search.py` / `search_gart.sh` — busca híbrida e seu wrapper.
- `ensure_tunnel.sh` — abre o túnel SSH até `rag-hub-db` se ainda não estiver
aberto (idempotente).
## Onde ficam as credenciais reais
Nunca nesta pasta. Credenciais de indexação (usuário/senha do Postgres) ficam
em `admin/gart-rag.env` (fora do git). Acesso SSH à VPS e senha do usuário
admin do `rag-hub-db` ficam documentados no `VPS-ACCESS.md` de outro projeto
da equipe que já usa a mesma VPS — peça a quem provisionou o banco.
+71
View File
@@ -0,0 +1,71 @@
# Provisionar o banco `rag_gart`
Passo a passo para criar o banco deste projeto no container compartilhado
`rag-hub-db` (mesma VPS usada por Doza/Tigre/Jhonny). Só precisa ser feito
uma vez (ou para reprovisionar do zero).
Precisa de acesso SSH à VPS (`root@179.197.228.240`, chave já autorizada) e
das credenciais do usuário admin do `rag-hub-db` (`rag_admin` — senha em
`/docker/rag-hub/.env` na própria VPS; não é duplicada em nenhum projeto).
## 1. Criar o banco e a role de indexação
Via túnel SSH ou `docker exec` na VPS, como `rag_admin`:
```sql
CREATE DATABASE rag_gart OWNER rag_admin;
\c rag_gart
CREATE EXTENSION IF NOT EXISTS vector;
CREATE EXTENSION IF NOT EXISTS pg_trgm;
-- Role de indexação, sem DDL (só o que admin/update_rag.py e rag/search.py
-- precisam: SELECT/INSERT/UPDATE/DELETE no schema gart).
CREATE ROLE gart_rag_indexer LOGIN PASSWORD '<senha forte gerada aqui>';
```
## 2. Criar o schema
Rode [`schema.sql`](schema.sql) neste banco (idempotente, `CREATE ... IF NOT
EXISTS` em tudo):
```bash
psql "postgresql://rag_admin@127.0.0.1:55435/rag_gart" -f rag/schema.sql
```
Depois conceda os privilégios ao usuário de indexação:
```sql
GRANT USAGE ON SCHEMA gart TO gart_rag_indexer;
GRANT SELECT, INSERT, UPDATE, DELETE ON ALL TABLES IN SCHEMA gart TO gart_rag_indexer;
GRANT USAGE, SELECT ON ALL SEQUENCES IN SCHEMA gart TO gart_rag_indexer;
ALTER DEFAULT PRIVILEGES IN SCHEMA gart
GRANT SELECT, INSERT, UPDATE, DELETE ON TABLES TO gart_rag_indexer;
```
## 3. Preencher as credenciais locais
```bash
cp admin/gart-rag.env.example admin/gart-rag.env
```
Editar `admin/gart-rag.env` e colocar a senha gerada no passo 1 em
`RAG_DB_PASSWORD`. Esse arquivo é ignorado pelo git — nunca commitar.
## 4. Indexar pela primeira vez
```bash
admin/update_rag.command
```
Abre o túnel SSH (se não estiver aberto), roda `admin/update_rag.py` e grava
os chunks + o mapa de arquivos em `rag_gart`.
## 5. Testar a busca
```bash
rag/search_gart.sh "como funciona o export para DaVinci"
rag/search_gart.sh --map fcpxml
```
Se vier `[RAG vazio, buscando local]`, confira se o passo 4 rodou sem erro e
se `RAG_DB_PASSWORD` está correta.
+37
View File
@@ -0,0 +1,37 @@
"""Embeddings compartilhados entre indexador e busca do RAG do G-ART.
admin/update_rag.py já indexa (chunking por janela de linhas, incremental,
rodado por admin/run.command). Este módulo só isola a chamada ao Ollama para
que rag/search.py use exatamente o mesmo modelo/prefixo na consulta.
Prefixos do embedding
----------------------
O nomic-embed-text espera `search_document: ` no que é indexado e
`search_query: ` no que é consultado — sem isso a qualidade da busca cai.
admin/update_rag.py já indexa com `search_document: ` (ver `_embed` lá).
Trocar esse regime invalida os vetores antigos e exige reindexar tudo.
"""
import os
import requests
OLLAMA_URL = os.environ.get("OLLAMA_URL", "http://127.0.0.1:11434")
EMBED_MODEL = os.environ.get("RAG_EMBED_MODEL", "nomic-embed-text")
EMBED_DIM = int(os.environ.get("RAG_EMBED_DIM", "768"))
DOC_PREFIX = "search_document: "
QUERY_PREFIX = "search_query: "
def _embed(text: str, prefix: str = DOC_PREFIX):
response = requests.post(
f"{OLLAMA_URL.rstrip('/')}/api/embeddings",
json={"model": EMBED_MODEL, "prompt": f"{prefix}{text}"},
timeout=60,
)
response.raise_for_status()
vector = response.json()["embedding"]
if len(vector) != EMBED_DIM:
raise ValueError(f"embedding com {len(vector)} dimensões; esperado {EMBED_DIM}")
return vector
+28
View File
@@ -0,0 +1,28 @@
#!/bin/bash
# Garante que o túnel SSH até o Postgres da VPS (rag-hub-db) está aberto em
# 127.0.0.1:55435. Idempotente: se já estiver escutando, não faz nada.
# Mesmo container compartilhado usado por rag_doza/rag_tigre/jhonny-rag.
HOST="root@179.197.228.240"
LOCAL_PORT=55435
REMOTE_PORT=55435
if nc -z 127.0.0.1 "$LOCAL_PORT" 2>/dev/null; then
echo "[rag] tunel ja aberto em 127.0.0.1:$LOCAL_PORT"
exit 0
fi
echo "[rag] abrindo tunel SSH ate $HOST ($LOCAL_PORT -> $REMOTE_PORT)..."
ssh -f -N -o ExitOnForwardFailure=yes -o ServerAliveInterval=30 -o ServerAliveCountMax=3 \
-L "127.0.0.1:${LOCAL_PORT}:127.0.0.1:${REMOTE_PORT}" "$HOST"
for _ in $(seq 1 10); do
if nc -z 127.0.0.1 "$LOCAL_PORT" 2>/dev/null; then
echo "[rag] tunel ativo."
exit 0
fi
sleep 0.5
done
echo "[rag] AVISO: nao foi possivel confirmar o tunel." >&2
exit 1
+6
View File
@@ -0,0 +1,6 @@
# Dependências só das ferramentas de RAG (dev, não da aplicação).
# admin/update_rag.py usa psycopg2 + requests diretamente (sem dotenv, pois
# admin/update_rag.command já faz `source` do .env antes de chamá-lo).
psycopg2-binary>=2.9
python-dotenv>=1.0
requests
+86
View File
@@ -0,0 +1,86 @@
-- Schema RAG do projeto G-ART (fcp-mcp-server).
-- Idempotente: seguro rodar múltiplas vezes (CREATE ... IF NOT EXISTS).
-- Segue o padrão dos bancos irmãos (rag_doza, rag_tigre, jhonny-rag): um
-- banco por sistema dentro do container compartilhado rag-hub-db, schema
-- próprio. Banco: rag_gart. Schema: gart.
--
-- Colunas e tabelas espelham exatamente o que admin/update_rag.py grava
-- (code_chunks, file_index, indexed_files) e o que rag/search.py lê.
CREATE EXTENSION IF NOT EXISTS vector;
CREATE EXTENSION IF NOT EXISTS pg_trgm;
CREATE SCHEMA IF NOT EXISTS gart;
CREATE TABLE IF NOT EXISTS gart.code_chunks (
id bigserial PRIMARY KEY,
file_path text NOT NULL,
content text NOT NULL,
chunk_index int,
embedding vector(768),
content_hash text,
file_mtime double precision,
-- Faixa de linhas do trecho no arquivo original. É o que permite ao
-- agente ler só a janela relevante em vez do arquivo inteiro.
start_line int,
end_line int,
-- Nomes declarados no trecho (class/def/func...), separados por vírgula —
-- o lado lexical (pg_trgm) da busca híbrida casa contra isto.
symbols text,
-- Módulo derivado do caminho relativo (primeiro segmento, ex: code/fcpxml -> code).
module text,
-- 'window': admin/update_rag.py corta por janela de linhas, não por
-- declaração (o projeto é majoritariamente Python/Swift/Markdown/shell).
kind text,
updated_at timestamp DEFAULT now()
);
-- HNSW, não ivfflat: com poucas centenas/milhares de chunks o ivfflat
-- particiona o espaço em listas quase vazias e a busca com probes baixo
-- varre quase nada (ver rag/README.md para os números de referência).
CREATE INDEX IF NOT EXISTS code_chunks_embedding_hnsw_idx
ON gart.code_chunks USING hnsw (embedding vector_cosine_ops)
WITH (m = 16, ef_construction = 64);
-- Acelera a reindexação incremental (busca por file_path) e a limpeza de
-- chunks de um arquivo antes de reinserir.
CREATE INDEX IF NOT EXISTS idx_code_chunks_file_path
ON gart.code_chunks (file_path);
-- Lado lexical da busca híbrida.
CREATE INDEX IF NOT EXISTS code_chunks_file_path_trgm_idx
ON gart.code_chunks USING gin (file_path gin_trgm_ops);
CREATE INDEX IF NOT EXISTS code_chunks_symbols_trgm_idx
ON gart.code_chunks USING gin (symbols gin_trgm_ops);
CREATE INDEX IF NOT EXISTS code_chunks_module_idx
ON gart.code_chunks (module);
-- Hash de conteúdo por arquivo, usado pelo indexador para pular arquivos
-- que não mudaram desde a última rodada (reindexação incremental).
CREATE TABLE IF NOT EXISTS gart.indexed_files (
file_path text PRIMARY KEY,
content_hash text NOT NULL,
updated_at timestamp DEFAULT now()
);
-- Mapa de arquivos: 1 linha por arquivo. Responde "onde fica X" e "o que
-- tem no módulo Y" sem trazer nenhum corpo de código.
CREATE TABLE IF NOT EXISTS gart.file_index (
file_path text PRIMARY KEY,
module text,
main_type text,
public_symbols text[],
summary text,
n_lines int,
content_hash text,
summary_embedding vector(768),
updated_at timestamp DEFAULT now()
);
CREATE INDEX IF NOT EXISTS file_index_module_idx
ON gart.file_index (module);
CREATE INDEX IF NOT EXISTS file_index_path_trgm_idx
ON gart.file_index USING gin (file_path gin_trgm_ops);
CREATE INDEX IF NOT EXISTS file_index_summary_hnsw_idx
ON gart.file_index USING hnsw (summary_embedding vector_cosine_ops)
WITH (m = 16, ef_construction = 64);
+434
View File
@@ -0,0 +1,434 @@
"""Busca semântica RAG do projeto G-ART.
Consulta `gart.code_chunks` no Postgres (via túnel SSH) combinando dois
sinais e devolvendo faixas de linha, para o agente ler só o trecho relevante
em vez do arquivo inteiro.
Uso:
rag/search_gart.sh "como funciona o export FCPXML"
rag/search_gart.sh "FCPXMLModifier" --snippet
rag/search_gart.sh "voice_actions" --module fcpxml --json
rag/search_gart.sh --map fcpxml # inventário do módulo
Como módulo:
from search import rag_search
rag_search("consulta", top_k=5)
Por que busca híbrida
---------------------
A busca puramente densa erra nomes exatos: procurar `FCPXMLModifier` pode não
trazer `modifier.py` no topo. Por isso rodamos duas listas em paralelo —
densa (embedding) e lexical (pg_trgm sobre os símbolos declarados) — e
fundimos com RRF, que soma 1/(k+posição) de cada lista e portanto não exige
normalizar escalas diferentes. Adaptado do rag/search.py dos projetos irmãos
(Doza, Tigre, Jhonny) — mesmos parâmetros, calibrados lá via rag/bench.py.
"""
import argparse
import json
import os
import sys
import psycopg2
from dotenv import load_dotenv
RAG_DIR = os.path.dirname(os.path.abspath(__file__))
ROOT = os.path.dirname(RAG_DIR)
# Por padrão reaproveita as mesmas credenciais do indexador
# (admin/gart-rag.env), para não duplicar a senha em dois arquivos.
_load_path = os.environ.get("RAG_ENV_FILE", os.path.join(ROOT, "admin", "gart-rag.env"))
load_dotenv(_load_path)
if RAG_DIR not in sys.path:
sys.path.insert(0, RAG_DIR)
from embed import QUERY_PREFIX, _embed # noqa: E402
# Constante do Reciprocal Rank Fusion e pesos por lista — valores herdados
# dos projetos irmãos (calibrados lá via varredura sobre consultas douradas,
# ver rag/bench.py no projeto Doza/Tigre). Ao mexer nesses números, monte um
# conjunto de consultas de referência para o G-ART antes.
RRF_K = 8
W_DENSE = 1.0
W_LEX = 1.0 # multiplicado pela similaridade bruta do casamento
W_SUMMARY = 1.5 # o resumo é curto e preciso: um acerto ali vale mais
# Piso do casamento lexical: abaixo disso o lado lexical se cala em vez de
# afogar a lista densa com ruído.
LEX_MIN = 0.5
# Quantos chunks do mesmo arquivo podem ocupar o top-k, para um arquivo
# grande não tomar todos os lugares.
MAX_PER_FILE = 2
# Quanto código o modo --snippet mostra por resultado.
SNIPPET_CHARS = 300
_EXTRA_COLS = ("start_line", "end_line", "symbols", "module", "kind")
_cols_cache = {}
def _available_cols(cur, schema):
if schema not in _cols_cache:
cur.execute(
"""SELECT column_name FROM information_schema.columns
WHERE table_schema = %s AND table_name = 'code_chunks'""",
(schema,),
)
_cols_cache[schema] = {r[0] for r in cur.fetchall()}
return _cols_cache[schema]
def _select_cols(cur, schema):
have = _available_cols(cur, schema)
extra = ", ".join(f"c.{c}" if c in have else f"NULL AS {c}"
for c in _EXTRA_COLS)
return f"c.id, c.file_path, c.content, {extra}"
def _db_connect():
return psycopg2.connect(
host=os.environ.get("RAG_DB_HOST", "127.0.0.1"),
port=os.environ.get("RAG_DB_PORT", "55435"),
dbname=os.environ.get("RAG_DB_NAME", "rag_gart"),
user=os.environ.get("RAG_DB_USER", "gart_rag_indexer"),
password=os.environ["RAG_DB_PASSWORD"],
connect_timeout=5,
)
def _schema():
return os.environ.get("RAG_DB_SCHEMA", "gart")
def _qid(schema):
"""Schema como identificador SQL seguro (aspas duplas)."""
return '"' + schema.replace('"', '""') + '"'
def _filters(module, path, ext, have=()):
"""Cláusulas de escopo aplicadas antes do ranqueamento."""
clauses, params = [], []
if module and "module" in have:
clauses.append("c.module = %s")
params.append(module)
if path:
clauses.append("c.file_path ILIKE %s")
params.append(f"%{path}%")
if ext:
clauses.append("c.file_path LIKE %s")
params.append(f"%{ext}")
return (" AND " + " AND ".join(clauses) if clauses else ""), params
def _probe_terms(query):
"""Termos que o lado lexical tenta casar contra os símbolos: a consulta
inteira e a versão sem espaços (faz "voice timeline" casar com
`VoiceTimeline`)."""
terms = [query, query.replace(" ", "")]
return list(dict.fromkeys(t for t in terms if t))
def _row_to_dict(row):
return {
"id": row[0], "file_path": row[1], "content": row[2],
"start_line": row[3], "end_line": row[4],
"symbols": row[5], "module": row[6], "kind": row[7],
}
def _dense(cur, schema, q_emb, limit, where, params):
cur.execute("SET LOCAL hnsw.ef_search = 64")
cur.execute(
f"""
SELECT {_select_cols(cur, schema)}, 1 - (c.embedding <=> %s::vector) AS s
FROM {_qid(schema)}.code_chunks c
WHERE c.embedding IS NOT NULL {where}
ORDER BY c.embedding <=> %s::vector
LIMIT %s
""",
[q_emb] + params + [q_emb, limit],
)
return [(_row_to_dict(r), float(r[8])) for r in cur.fetchall()]
def _lexical(cur, schema, terms, limit, where, params):
have = _available_cols(cur, schema)
stem = "regexp_replace(c.file_path, '^.*/|\\.[^.]*$', '', 'g')"
target = f"coalesce(c.symbols, {stem})" if "symbols" in have else stem
tiebreak = "c.start_line" if "start_line" in have else "c.id"
score = f"(SELECT max(word_similarity(t, {target})) FROM unnest(%s::text[]) t)"
cur.execute(
f"""
SELECT {_select_cols(cur, schema)}, {score} AS s
FROM {_qid(schema)}.code_chunks c
WHERE {score} >= %s {where}
ORDER BY s DESC, {tiebreak} ASC
LIMIT %s
""",
[terms] + [terms, LEX_MIN] + params + [limit],
)
return [(_row_to_dict(r), float(r[8])) for r in cur.fetchall()]
def _summary_dense(cur, schema, q_emb, limit, where, params):
"""Terceira lista: busca sobre o resumo do arquivo (file_index), não
sobre o código — resgata arquivos pequenos e precisos que a lista por
trecho, dominada por arquivos grandes, deixa passar."""
cur.execute("SELECT to_regclass(%s)", (f"{_qid(schema)}.file_index",))
if cur.fetchone()[0] is None:
return []
cur.execute("SET LOCAL hnsw.ef_search = 64")
cur.execute(
f"""
SELECT file_path, 1 - (summary_embedding <=> %s::vector) AS s
FROM {_qid(schema)}.file_index
WHERE summary_embedding IS NOT NULL
ORDER BY summary_embedding <=> %s::vector
LIMIT %s
""",
(q_emb, q_emb, limit),
)
hits = cur.fetchall()
if not hits:
return []
order = {fp: i for i, (fp, _) in enumerate(hits)}
raw = {fp: s for fp, s in hits}
cur.execute(
f"""
SELECT DISTINCT ON (c.file_path) {_select_cols(cur, schema)}
FROM {_qid(schema)}.code_chunks c
WHERE c.file_path = ANY(%s) AND c.embedding IS NOT NULL {where}
ORDER BY c.file_path, c.embedding <=> %s::vector
""",
[list(order)] + params + [q_emb],
)
rows = [_row_to_dict(r) for r in cur.fetchall()]
rows.sort(key=lambda r: order[r["file_path"]])
return [(r, raw[r["file_path"]]) for r in rows]
def _fuse(ranked_lists):
"""Reciprocal Rank Fusion ponderada sobre listas já ordenadas."""
scores, best = {}, {}
for results, weight_of in ranked_lists:
for rank, (row, raw) in enumerate(results, 1):
contrib = weight_of(raw) / (RRF_K + rank)
scores[row["id"]] = scores.get(row["id"], 0.0) + contrib
best.setdefault(row["id"], row)
ordered = sorted(scores.items(), key=lambda kv: -kv[1])
return [dict(best[i], score=s) for i, s in ordered]
def _dedupe(rows, top_k, max_per_file=MAX_PER_FILE):
"""Limita chunks por arquivo; o excedente vira uma nota de localização."""
kept, counts, extras = [], {}, {}
for row in rows:
fp = row["file_path"]
if counts.get(fp, 0) < max_per_file:
counts[fp] = counts.get(fp, 0) + 1
row["also_at"] = []
kept.append(row)
else:
extras.setdefault(fp, []).append((row["start_line"], row["end_line"]))
for row in kept:
row["also_at"] = extras.get(row["file_path"], [])[:3]
return kept[:top_k]
def _attach_map(cur, schema, rows):
"""Anexa tipo principal e resumo de `file_index`, quando existir."""
cur.execute("SELECT to_regclass(%s)", (f"{_qid(schema)}.file_index",))
if cur.fetchone()[0] is None or not rows:
return rows
paths = list({r["file_path"] for r in rows})
cur.execute(
f"SELECT file_path, main_type, summary FROM {_qid(schema)}.file_index "
f"WHERE file_path = ANY(%s)",
(paths,),
)
info = {p: (t, s) for p, t, s in cur.fetchall()}
for row in rows:
main_type, summary = info.get(row["file_path"], (None, None))
row["main_type"] = main_type
row["summary"] = summary
return rows
def rag_search(query, top_k=5, module=None, path=None, ext=None, dense_only=False):
"""Busca híbrida. Devolve dicts com file_path, faixa de linhas e score."""
schema = _schema()
pool = max(top_k * 4, 20)
conn = _db_connect()
try:
with conn.cursor() as cur:
where, params = _filters(module, path, ext,
_available_cols(cur, schema))
q_emb = _embed(query, prefix=QUERY_PREFIX)
lists = [(_dense(cur, schema, q_emb, pool, where, params),
lambda raw: W_DENSE)]
if not dense_only:
lists.append((_lexical(cur, schema, _probe_terms(query),
pool, where, params),
lambda raw: W_LEX * (0.5 + raw)))
lists.append((_summary_dense(cur, schema, q_emb, pool,
where, params),
lambda raw: W_SUMMARY))
rows = _dedupe(_fuse(lists), top_k)
rows = _attach_map(cur, schema, rows)
conn.commit()
finally:
conn.close()
return rows
def map_files(term=None, module=None, limit=60):
"""Inventário de arquivos — responde "onde fica X" sem corpo de código."""
schema = _schema()
conn = _db_connect()
try:
with conn.cursor() as cur:
cur.execute("SELECT to_regclass(%s)", (f"{_qid(schema)}.file_index",))
if cur.fetchone()[0] is None:
return []
clauses, params = [], []
if module:
clauses.append("module = %s")
params.append(module)
if term:
clauses.append("(file_path ILIKE %s OR main_type ILIKE %s "
"OR summary ILIKE %s)")
params += [f"%{term}%"] * 3
where = "WHERE " + " AND ".join(clauses) if clauses else ""
cur.execute(
f"""SELECT file_path, module, main_type, summary, n_lines
FROM {_qid(schema)}.file_index {where}
ORDER BY module, file_path LIMIT %s""",
params + [limit],
)
return [
{"file_path": r[0], "module": r[1], "main_type": r[2],
"summary": r[3], "n_lines": r[4]}
for r in cur.fetchall()
]
finally:
conn.close()
# ---------------------------------------------------------------------------
# Formatação
# ---------------------------------------------------------------------------
_NOISE = {'"""', "'''", "# ---", "---", "/*", "*/", "{", "}", "*"}
def _first_doc_line(content):
"""Primeira linha que serve de descrição do trecho, usada quando o
arquivo não tem resumo em `file_index`."""
for line in content.splitlines():
s = line.strip()
if s.startswith(("///", "//", "#", "*")):
s = s.lstrip("/#* ").strip()
if s and s not in _NOISE:
return s
for line in content.splitlines():
s = line.strip().strip("\"'").strip()
if s and s not in _NOISE and not s.startswith("import"):
return s
return ""
def format_results(results, mode="map"):
"""Renderiza o resultado no modo pedido. O padrão é `map`: caminho +
faixa de linhas + uma linha de descrição, sem corpo de código."""
if not results:
return "[RAG vazio, buscando local]\n"
if mode == "json":
return json.dumps(results, ensure_ascii=False, indent=2) + "\n"
out = []
for i, r in enumerate(results, 1):
loc = r["file_path"]
if r.get("start_line") and r.get("end_line"):
loc += f":{r['start_line']}-{r['end_line']}"
out.append(f"{i} {r['score']:.2f} {loc}")
desc = r.get("summary") or _first_doc_line(r["content"])
stem = os.path.splitext(os.path.basename(r["file_path"]))[0]
label = r.get("main_type") or ""
if label == stem:
label = ""
if label and desc:
out.append(f" {label} · {desc[:78]}")
elif desc:
out.append(f" {desc[:88]}")
spans = [f"{a}-{b}" for a, b in r.get("also_at", []) if a and b]
if spans:
out.append(f" (+ tambem em {', '.join(spans)})")
if mode == "snippet":
body, size = [], 0
for ln in r["content"].splitlines():
if body and size + len(ln) > SNIPPET_CHARS:
body.append("…")
break
body.append(ln)
size += len(ln) + 1
out.append("".join(f" | {ln}\n" for ln in body))
elif mode == "full":
out.append("".join(f" | {ln}\n" for ln in r["content"].splitlines()))
return "\n".join(out) + "\n"
def format_map(rows):
if not rows:
return "[RAG vazio, buscando local]\n"
out = []
current = None
for r in rows:
if r["module"] != current:
current = r["module"]
out.append(f"\n{current or '(sem modulo)'}")
name = os.path.basename(r["file_path"])
desc = (r["summary"] or "")[:78]
out.append(f" {name:<38} {r['n_lines']:>5}L {desc}")
return "\n".join(out) + "\n"
def main():
ap = argparse.ArgumentParser(
description="Busca RAG hibrida (densa + lexical, fundidas com RRF) do G-ART")
ap.add_argument("query", nargs="?", help="consulta")
ap.add_argument("top_k", nargs="?", type=int, default=5)
ap.add_argument("--snippet", action="store_true", help="mostra 300 chars do trecho")
ap.add_argument("--full", action="store_true", help="mostra o trecho inteiro")
ap.add_argument("--json", action="store_true", help="saida estruturada")
ap.add_argument("--module", help="restringe a um modulo (ex: fcpxml)")
ap.add_argument("--path", help="restringe a caminhos contendo este texto")
ap.add_argument("--ext", help="restringe a uma extensao (ex: .py)")
ap.add_argument("--map", dest="map_term", nargs="?", const="",
help="inventario de arquivos em vez de busca por trecho")
ap.add_argument("--dense-only", action="store_true",
help="desliga o lado lexical (para comparacao)")
args = ap.parse_args()
if args.map_term is not None:
print(format_map(map_files(term=args.map_term or None, module=args.module)))
return
if not args.query:
ap.error("informe a consulta, ou use --map")
mode = ("json" if args.json else "full" if args.full
else "snippet" if args.snippet else "map")
results = rag_search(args.query, args.top_k, module=args.module,
path=args.path, ext=args.ext, dense_only=args.dense_only)
print(format_results(results, mode=mode))
if __name__ == "__main__":
main()
+19
View File
@@ -0,0 +1,19 @@
#!/bin/zsh
# Busca RAG do G-ART. Abre o túnel SSH se preciso e chama rag/search.py com
# as credenciais de admin/gart-rag.env (mesmas do indexador incremental).
set -euo pipefail
ROOT="$(cd "$(dirname "$0")/.." && pwd -P)"
RAG_DIR="$ROOT/rag"
"$RAG_DIR/ensure_tunnel.sh"
PYTHON="${RAG_PYTHON:-}"
if [[ -z "$PYTHON" ]]; then
for candidate in "$ROOT/admin/.venv/bin/python3" "$ROOT/rag/.venv/bin/python3"; do
if [[ -x "$candidate" ]]; then PYTHON="$candidate"; break; fi
done
fi
PYTHON="${PYTHON:-$(command -v python3)}"
exec "$PYTHON" "$RAG_DIR/search.py" "$@"