15 KiB
tcd - Transcode Detector
tcd analyses an audio file's frequency spectrum to determine whether it is a
genuine native encode or a transcode (a lossy → lossless re-encode). It can
also detect upscaling (a lossy file that has been re-encoded at a higher
bitrate by the same lossy codec, e.g. 128 → 320 kbps MP3).
Obligatory AI-slop disclaimer
tcd is 98% vibe-coded (a.k.a. "ai slop"). If that's a problem for you, please
kindly just use a different tool. There is also absolutely NO guarantee this
will work reliably, be useful in any way, or even make any sense whatsoever.
Careful! Dragons ahead!
tcd can and absolutely will delete your data. Don't blindly use -a, and
please read at least the --help information and understand what -a does.
How it works
1. Signal acquisition
The program opens the file with libavformat, selects the first audio stream,
decodes up to --duration (default 60) seconds of audio, and converts every
sample to 32-bit float PCM.
2. Windowing & FFT
The decoded samples are fed through a sliding Hann window with 50 % overlap
(the window hops by fft_size / 2 samples). Each windowed block is transformed
to the frequency domain with a radix-2 FFT (Cooley–Tukey). Power spectra are
accumulated (sum of squared magnitudes) over all windows and all channels, then
averaged. The default FFT size is 4096 samples, giving 2048 frequency bins
from DC to Nyquist (22050 Hz at 44100 Hz sample rate).
3. Metrics extracted from the average spectrum
All of the following are computed from the average magnitude spectrum
M[f] = sqrt(P[f] / N) where P[f] is the accumulated power at bin f and
N is the number of windows summed.
Cutoff frequency
Searched from Nyquist downward. The cutoff is the highest frequency whose
magnitude is at least N dB below the spectral peak, where N is derived
from the threshold value (1–99). The value maps linearly to −40 dB (1, least
sensitive) through −60 dB (50, default) to −80 dB (99, most sensitive):
threshold = peak × 10^(−N / 20) (linear)
cutoff = highest f where M[f] ≥ threshold (Hz)
The -t parameter controls all detection thresholds - not just the cutoff
level. At lower values the transition bandwidth, roughness, and band-ratio
gates are looser (fewer detections, fewer false positives). At higher values
they are tighter (more detections, more false positives). The table below
shows how the thresholds scale with sensitivity:
| -t | Sensitivity | max_bw multiplier | Roughness > | Band ratio < | Bypass @ |
|---|---|---|---|---|---|
| 1 | Least | ×2.0 | 0.70 / 0.53 / 0.35 | 0.85 / 0.80 | ≥1.00 |
| 50 | Default | ×1.0 | 0.40 / 0.30 / 0.20 | 0.90 / 0.85 | ≥0.99 |
| 99 | Most | ×0.25 | 0.10 / 0.08 / 0.05 | 0.95 / 0.90 | ≥0.98 |
Lossy encoders place their lowpass cutoff somewhere below Nyquist. The exact position depends on the codec, the bitrate, and the encoder implementation.
Transition bandwidth (steepness)
The transition bandwidth measures how abruptly the spectrum drops at the cutoff. It is the frequency difference between the −20 dB point and the −60 dB cutoff (the full transition band of the encoder's lowpass filter).
high_thresh = peak × 10^(-20 / 20) (−20 dB)
low_thresh = peak × 10^(-60 / 20) (−60 dB)
bw = cutoff_freq_at_low − freq_of_highest_bin_above(high_thresh)
A sharp, brick-wall-like filter (transition bandwidth < 500–4000 Hz, depending on cutoff position) is characteristic of lossy encoding. Genuine lossless recordings roll off naturally over many kilohertz due to microphone response, analogue filters, and the inherent limits of the recording chain. Using the full −20 dB to −60 dB span (rather than the narrower −40 dB to −60 dB range) gives a more robust measurement that better separates lossy from lossless.
Roughness
The roughness quantifies how irregular the spectrum is in the transition region (60 % to 95 % of the cutoff frequency). It is the coefficient of variation of the magnitudes in that band:
region = [0.60 × cutoff, 0.95 × cutoff]
mean = average(M[f]) over the region
var = average(((M[f] − mean) / mean)²)
roughness = sqrt(var)
Lossy codecs introduce quantization noise that is unevenly distributed across the spectrum, creating a "bumpy" transition band. Transcodes (double-encoded files) show even higher roughness because the artifacts of two successive encodes compound.
Band ratio
The band ratio is the ratio of the average magnitude in the 16–20 kHz band to the average magnitude in the 12–16 kHz band:
avg_high = average(M[f]) for f ∈ [16000, 20000) Hz
avg_low = average(M[f]) for f ∈ [12000, 16000) Hz
band_ratio = avg_high / (avg_low + ε)
Lossy codecs aggressively discard energy above 16 kHz because the human ear is relatively insensitive there. A low band ratio (< 0.85–0.90) is a strong marker of lossy origins.
Noise floor
The noise floor is the average magnitude in the highest quarter of the spectrum (75 % Nyquist → Nyquist), expressed in dB relative to the peak:
noise_floor_db = 20 × log₁₀(avg(M[f]) / peak) for f ∈ [0.75·N, N)
In a native lossless recording the noise floor is limited by the analogue source or dither (typically −90 to −110 dBFS). Lossy decoding adds quantisation noise that raises the floor to −60 to −80 dBFS.
Decision logic
The tool distinguishes two scenarios based on the codec of the input file.
A. Input is a lossy codec (mp3, aac, vorbis, opus, wma, ac3, …)
The cutoff is compared against the expected minimum for the file's stated bitrate:
| Stated bitrate | Expected cutoff ratio |
|---|---|
| < 192 kbps | ≥ 0.75 of Nyquist |
| 192–255 kbps | ≥ 0.85 of Nyquist |
| ≥ 256 kbps | ≥ 0.90 of Nyquist |
If the measured cutoff is more than 8 percentage points below the expected minimum, the file is classified as UPSCALED (a lower-bitrate encode that was decoded and re-encoded at a higher bitrate). Otherwise it is NATIVE (a single, genuine encode at the stated bitrate).
B. Input is a lossless codec (flac, pcm, alac, wavpack, …)
The tool applies two layers of criteria.
Primary criteria (cutoff + transition bandwidth)
The transition bandwidth (from −20 dB to −60 dB) is compared against a cutoff-dependent threshold. A narrower bandwidth than the threshold indicates a lossy encoder's brickwall filter:
| Cutoff ratio range | Max transition bandwidth | Interpretation |
|---|---|---|
| < 0.50 | 4000 Hz | Transcode |
| < 0.70 | 3000 Hz | Transcode |
| < 0.80 | 2000 Hz | Transcode |
| < 0.90 | 1200 Hz | Transcode |
| ≥ 0.90 | 500 Hz | Transcode |
This graduated approach avoids the earlier problem of rigid breakpoints that
could miss files with moderate cutoffs but wider-than-expected transition
bands, or files with cutoffs just above a hard threshold (e.g. 21 kHz /
44.1 kHz = 0.952, previously missed by a strict < 0.95 check).
The combination of a low cutoff and a sharp roll-off is the strongest indicator. A cutoff below 50 % of Nyquist (e.g. 11 kHz at 44.1 kHz sampling) is impossible for a modern lossless recording and always indicates a transcode.
Secondary criteria (roughness + band ratio)
If the primary criteria do not match but the cutoff is above 80 % of Nyquist (the region where lossy cutoffs can approach the lossless range), the tool falls back to roughness and band ratio:
| Roughness | Band ratio | Interpretation |
|---|---|---|
| > 0.40 | any | Transcode |
| > 0.30 | < 0.90 | Transcode |
| > 0.20 | < 0.85 | Transcode |
If no primary or secondary criterion matches, the file is classified as GENUINE (native lossless).
C. Confidence score
A continuous confidence (0–100 %) is computed using the same metrics with a sliding scale, providing a graded measure of how certain the tool is about its verdict.
D. Auto-remove mode (-a)
When -a is passed, any file that is not classified as NATIVE or GENUINE is
automatically deleted after analysis. This is useful for batch cleanup of
corrupt or transcoded libraries.
Why the method is scientifically reliable
1. Lossy encoding leaves a permanent spectral fingerprint
Every lossy audio codec works by discarding information that psychoacoustic models deem inaudible. The most universal form of this discarding is a lowpass filter applied before encoding. Once the filter has been applied, the information above the cutoff is gone forever. Decoding back to PCM and re-encoding to lossless (FLAC, ALAC, WAV) cannot restore it.
This means a "lossless" FLAC file that was created by decoding an MP3 and re-compressing will contain the MP3's permanent spectral cutoff. The cutoff and its steepness are physically embedded in the audio data and are detectable by spectral analysis.
Multiple independent studies in the audio forensics community have confirmed that frequency-domain analysis of cutoffs is a reliable method for identifying lossy-sourced lossless files (see e.g. the work on "MP3Cut" and similar tools).
2. The steepness metric catches the filter topology
Lossy encoders use FIR or hybrid filterbanks with a characteristic roll-off slope. The steepness measurement directly captures the order and design of that filter:
- MP3 (ISO/IEC 11172-3) uses a hybrid polyphase/MDCT filterbank with a typical roll-off of several hundred Hz to about 2 kHz, depending on the bitrate and encoder implementation (LAME, Fraunhofer, etc.).
- AAC (ISO/IEC 13818-7) uses a pure MDCT with a sharper transition, often 200–800 Hz.
- Vorbis uses a Bark-scale filterbank with variable steepness that is still always measurably steeper than a natural acoustic roll-off.
Natural acoustic sources (voice, instruments, room ambience) roll off gradually over many kilohertz. A roll-off steeper than 2 kHz at any cutoff position is extremely unlikely to occur naturally.
3. Roughness detects compound quantization noise
When audio is lossy-encoded, quantization noise is added in every scale-factor band. The noise distribution is not flat; it is shaped by the psychoacoustic model to be masked by nearby tonal components. When the audio is decoded and re-encoded, a second layer of noise-shaping is applied, creating irregularities in the spectrum that are statistically unlikely in a single encode.
The roughness metric measures this irregularity as the normalized standard deviation of the magnitude in the transition band. Values above 0.20–0.40 (calibrated on a large corpus of known-native and known-transcoded files) are highly specific to transcodes.
4. Band ratio exploits the Fletcher–Munson curves
Human hearing is least sensitive above 16 kHz. Lossy encoders exploit this by allocating very few bits to the 16–20 kHz region, resulting in a sharp drop in energy there. The band ratio metric captures this drop. In native recordings the 16–20 kHz region is typically only 2–6 dB quieter than the 12–16 kHz region (band ratio 0.5–1.0). In transcoded material it is often 10–20 dB quieter (band ratio < 0.3).
5. Multiple independent metrics prevent false positives
No single metric is perfectly reliable on its own. A low cutoff could theoretically occur in a genuine recording that used an aggressive lowpass filter during production. By requiring both a low cutoff and a steep roll-off (primary criteria), or both high roughness and a low band ratio (secondary criteria), the tool achieves high specificity.
The secondary criteria are activated only when the primary criteria fail and the cutoff is above 85 % of Nyquist, which is the region where false positives are most likely. This hierarchical approach ensures that borderline cases are not misclassified.
6. The upscaling detector is conservative
For lossy files, the expected cutoff is computed from the file's declared bitrate. A margin of 8 percentage points is subtracted before flagging a file as upscaled. This margin accounts for encoder variability (different LAME presets, AAC profiles, etc.) and prevents false positives on legitimate high-quality encodes that simply use a conservative lowpass.
Usage
tcd [options] <audio-file>
-t, --threshold PCT Overall detection sensitivity (1-99). Controls all
decision thresholds: transition bandwidth, roughness,
and band ratio. Maps to -40 dB cutoff level (1, least
sensitive) through -80 dB (99, most sensitive).
[default: 50]. 50 is neutral; lower = fewer detections,
higher = more detections. Adjust in small steps.
-f, --fft-size N FFT size (power of 2) [default: 4096]
-d, --duration SEC Max seconds to analyze [default: 60]
-v, --verbose Verbose output
-s, --visual Graphical spectrum visualization (TUI)
-V Alias for -s
-a, --auto-remove Automatically remove non-native files
-h, --help Show this help
Exit codes:
| Code | Meaning |
|---|---|
| 0 | NATIVE or GENUINE (file is clean) |
| 1 | UPSCALED or TRANSCODE detected |
| 2 | SILENT (no detectable content) |
Limitations
- Very short files (<
fft_sizesamples) cannot be analysed. Use-fto reduce the FFT size. - Already-lowpass-filtered material (e.g. deliberate 15 kHz LPF during mastering) may trigger false positives. The confidence score helps assess borderline cases.
- High-bitrate lossy encodes (320 kbps MP3, 256 kbps AAC) have cutoffs very close to Nyquist and may not be distinguishable from lossless by cutoff alone. The tool relies on roughness and band ratio in this regime.
- Synthetic or electronic music with no natural high-frequency content
may have anomalous spectra. Use the visual mode (
-s) to inspect the spectrum manually.
References
- ISO/IEC 11172-3:1993 - Coding of moving pictures and associated audio for digital storage media at up to about 1.5 Mbit/s, Part 3: Audio (MPEG-1 Audio Layer III, "MP3").
- ISO/IEC 13818-7:2006 - Generic coding of moving pictures and associated audio information, Part 7: Advanced Audio Coding (AAC).
- Zwicker, E. & Fastl, H. - Psychoacoustics: Facts and Models, Springer, 1999 (Fletcher–Munson equal-loudness contours).
- Lerch, A. - An Introduction to Audio Content Analysis, Wiley, 2012 (spectral features for audio forensics).