DOCS/Custom Extractors

Custom Extractors

When the built-in plugins don't support your metadata format, you can write a custom extractor. An extractor receives file candidates and returns discovered examples.

Extractor Interface

interface Extractor<TMetadata>

Inputs:

  • candidates — Dirent[] entries (files and directories) matching the scan config
  • options — optional ExtractorOptions with rootPath, exclude[], and signal for cancellation

Output:

interface ExtractorResult<TMetadata>

Step-by-Step Walkthrough

Here's a complete custom extractor that reads TOML metadata files:

/**
 * Custom TOML-based extractor example.
 *
 * This demonstrates how to create your own extractor that:
 * 1. Scans for a specific file pattern (meta.toml)
 * 2. Parses metadata from that file format
 * 3. Claims files and returns Example objects
 */
import {
  ExampleFile,
  type Example,
  type Extractor,
  type ExtractorResult,
} from 'functional-examples';
import { readdirSync, type Dirent } from 'node:fs';
import { readFile } from 'node:fs/promises';
import path, { join } from 'node:path';

/**
 * Metadata structure for TOML examples.
 * You can define any shape that fits your use case.
 */
export interface TomlMetadata {
  id: string;
  title: string;
  description?: string;
  author?: string;
  [key: string]: unknown;
}

/**
 * Create a custom extractor that reads TOML metadata files.
 *
 * Extractors implement a candidate-based pattern: they're called with
 * pre-filtered candidates (files and directories) and decide which to handle.
 */
export function createTomlExtractor(): Extractor<TomlMetadata> {
  return {
    name: 'toml-extractor',

    async extract(
      candidates: Dirent[]
    ): Promise<ExtractorResult<TomlMetadata>> {
      const examples: Example<TomlMetadata>[] = [];
      const claimedFiles = new Set<string>();
      const errors: { path: string; message: string }[] = [];

      // Find meta.toml files from candidates
      const tomlFiles: string[] = [];

      for (const candidate of candidates) {
        const fullPath = path.join(candidate.parentPath, candidate.name);

        if (candidate.isFile()) {
          // Direct file candidate: check if it's a meta.toml
          if (candidate.name === 'meta.toml') {
            tomlFiles.push(fullPath);
          }
        } else if (candidate.isDirectory()) {
          // Directory candidate: look for meta.toml inside
          const metaPath = path.join(fullPath, 'meta.toml');
          try {
            await readFile(metaPath, 'utf-8');
            tomlFiles.push(metaPath);
          } catch {
            // No meta.toml in this directory, skip
          }
        }
      }

      for (const tomlFile of tomlFiles) {
        try {
          const content = await readFile(tomlFile, 'utf-8');
          const metadata = parseSimpleToml(content);

          const exampleDir = path.dirname(tomlFile);

          // Collect all files in the example directory
          const files = collectExampleFiles(exampleDir);

          // Claim all files
          for (const file of files) {
            claimedFiles.add(file);
          }

          examples.push({
            id: metadata.id,
            title: metadata.title,
            description: metadata.description,
            rootPath: exampleDir,
            files: files.map((f) => new ExampleFile({
              absolutePath: f,
              relativePath: path.relative(exampleDir, f),
            })),
            metadata,
            extractorName: 'toml-extractor',
          });
        } catch (err) {
          errors.push({
            path: tomlFile,
            message: `Failed to parse: ${(err as Error).message}`,
          });
        }
      }

      return { examples, errors, claimedFiles };
    },
  };
}

/**
 * Simplified TOML parser for demonstration.
 * In production, use a proper TOML library like @iarna/toml.
 */
function parseSimpleToml(content: string): TomlMetadata {
  const lines = content.split('\n');
  const result: Record<string, string> = {};

  for (const line of lines) {
    // Match: key = "value"
    const match = line.match(/^(\w+)\s*=\s*"(.*)"/);
    if (match) {
      result[match[1]] = match[2];
    }
  }

  if (!result['id'] || !result['title']) {
    throw new Error('TOML must have id and title fields');
  }

  return {
    id: result['id'],
    title: result['title'],
    description: result['description'],
    author: result['author'],
  };
}

function collectExampleFiles(root: string) {
  if (root.endsWith('node_modules')) {
    return [];
  }

  let files: string[] = [];
  const entries = readdirSync(root, { withFileTypes: true });
  for (const entry of entries) {
    if (entry.isDirectory()) {
      files = files.concat(
        collectExampleFiles(join(entry.parentPath, entry.name))
      );
    } else {
      files.push(join(entry.parentPath, entry.name));
    }
  }
  return files;
}

Key Concepts

  1. Scan candidates — The extractor receives pre-filtered file entries. Iterate them to find your metadata files (e.g., meta.toml).

  2. Claim files — Add every file your extractor "owns" to the claimedFiles set. This prevents other extractors from processing the same files.

  3. Return errors gracefully — For recoverable issues (malformed metadata, missing fields), push to the errors array instead of throwing. The scanner collects all errors across extractors.

  4. Set extractorName — Each example includes its extractorName so results show which extractor produced it.

Registering a Custom Extractor

Via Programmatic Scan

/**
 * Scan script using the custom TOML extractor.
 */
import { scan } from 'functional-examples';
import { resolveConfig } from 'functional-examples';
import { createTomlExtractor } from './toml-extractor.js';

async function main() {
  const jsonOutput = process.argv.includes('--json');

  const config = await resolveConfig({
    root: '.',
    plugins: [
      {
        name: 'toml-extractor',
        extractor: createTomlExtractor(),
      },
    ],
    scan: {
      include: ['**/*'],
      exclude: ['**/node_modules/**'],
    },
  });

  const result = await scan({ config });

  if (jsonOutput) {
    console.log(
      JSON.stringify(
        {
          examples: result.examples.map((e) => ({
            id: e.id,
            title: e.title,
            description: e.description,
            files: e.files.map((f) => f.relativePath),
            metadata: e.metadata,
          })),
          errors: result.errors,
          stats: result.stats,
        },
        null,
        2
      )
    );
  } else {
    console.log(`Found ${result.examples.length} TOML example(s):\n`);
    for (const example of result.examples) {
      console.log(`  ${example.id}: ${example.title}`);
      if (example.description) {
        console.log(`    ${example.description}`);
      }
      console.log(
        `    Files: ${example.files.map((f) => f.relativePath).join(', ')}`
      );
      console.log();
    }

    if (result.errors.length > 0) {
      console.log(`Errors (${result.errors.length}):`);
      for (const error of result.errors) {
        console.log(`  - ${error.path}: ${error.message}`);
      }
    }
  }
}

main().catch(console.error);

Via Config Plugin

Wrap your extractor in a plugin object to register it via config:

import type { Plugin } from 'functional-examples';
import { createTomlExtractor } from './toml-extractor.js';

const tomlPlugin: Plugin = {
  name: 'toml',
  extensions: ['.toml'],
  extractor: createTomlExtractor(),
};

export default {
  plugins: [tomlPlugin],
  scan: { include: ['examples/**/*'] },
};

Error Handling

  • Recoverable errors → push to errors[] array (scanner aggregates them)
  • Unrecoverable failures → throw an error (scanner catches and reports it)
  • Always validate required fields (id, title) before creating an Example

Testing Custom Extractors

Write tests that verify your extractor:

  1. Finds examples from valid metadata files
  2. Returns errors for malformed metadata
  3. Claims the correct files
  4. Handles empty candidate lists gracefully

See also: Plugin Authoring for building a full plugin with extractors, parsers, and more.