EXAMPLES/custom-extractor

Custom Extractor

Demonstrates creating a custom Extractor implementation to support alternative metadata formats like TOML. Advanced usage requiring a TypeScript config with inline plugin code.

meta-yml8 FILEScustomextractortomladvanced

Custom Extractor Example

This example demonstrates how to create your own custom extractor for file formats not covered by built-in plugins.

Usage

bash
1# Scan using the custom TOML extractor
2npx tsx scan.ts

The Extractor

The createTomlExtractor function implements the full extractor interface — scanning for meta.toml files, parsing metadata, and claiming files:

export function createTomlExtractor(): Extractor<TomlMetadata> {
  return {
    name: 'toml-extractor',

    async extract(
      candidates: Dirent[]
    ): Promise<ExtractorResult<TomlMetadata>> {
      const examples: Example<TomlMetadata>[] = [];
      const claimedFiles = new Set<string>();
      const errors: { path: string; message: string }[] = [];

      // Find meta.toml files from candidates
      const tomlFiles: string[] = [];

      for (const candidate of candidates) {
        const fullPath = path.join(candidate.parentPath, candidate.name);

        if (candidate.isFile()) {
          // Direct file candidate: check if it's a meta.toml
          if (candidate.name === 'meta.toml') {
            tomlFiles.push(fullPath);
          }
        } else if (candidate.isDirectory()) {
          // Directory candidate: look for meta.toml inside
          const metaPath = path.join(fullPath, 'meta.toml');
          try {
            await readFile(metaPath, 'utf-8');
            tomlFiles.push(metaPath);
          } catch {
            // No meta.toml in this directory, skip
          }
        }
      }

      for (const tomlFile of tomlFiles) {
        try {
          const content = await readFile(tomlFile, 'utf-8');
          const metadata = parseSimpleToml(content);

          const exampleDir = path.dirname(tomlFile);

          // Collect all files in the example directory
          const files = collectExampleFiles(exampleDir);

          // Claim all files
          for (const file of files) {
            claimedFiles.add(file);
          }

          examples.push({
            id: metadata.id,
            title: metadata.title,
            description: metadata.description,
            rootPath: exampleDir,
            files: files.map((f) => new ExampleFile({
              absolutePath: f,
              relativePath: path.relative(exampleDir, f),
            })),
            metadata,
            extractorName: 'toml-extractor',
          });
        } catch (err) {
          errors.push({
            path: tomlFile,
            message: `Failed to parse: ${(err as Error).message}`,
          });
        }
      }

      return { examples, errors, claimedFiles };
    },
  };
}

Key Concepts

Tree-Scan Pattern

Extractors implement the "tree-scan" pattern:

  • Called once with candidate Dirent[] entries
  • Return ALL examples found in those candidates
  • Claim files they've processed via claimedFiles

Claimed Files

The claimedFiles set tells the scanner which files this extractor owns. This is used for:

  • Conflict detection (multiple extractors claiming same file)
  • Path-based conflict resolution

Example Shape

Each example must have:

  • id — Unique identifier
  • title — Display name
  • rootPath — Base directory
  • files — Array of file info
  • metadata — Your custom metadata type
  • extractorName — Your extractor's name

When to Create a Custom Extractor

  • Unique metadata format (TOML, INI, custom JSON schema)
  • Special directory structures
  • Language-specific conventions
  • Integration with other tools

All Example Files

FILE EXPLORER
README.md
1# Custom Extractor Example
2
3This example demonstrates how to create your own custom extractor for file formats not covered by built-in plugins.
4
5## Usage
6
7```bash
8# Scan using the custom TOML extractor
9npx tsx scan.ts
10```
11
12## The Extractor
13
14The `createTomlExtractor` function implements the full extractor interface — scanning for `meta.toml` files, parsing metadata, and claiming files:
15
16<%= region('createExtractor') %>
17
18## Key Concepts
19
20### Tree-Scan Pattern
21
22Extractors implement the "tree-scan" pattern:
23- Called once with candidate `Dirent[]` entries
24- Return ALL examples found in those candidates
25- Claim files they've processed via `claimedFiles`
26
27### Claimed Files
28
29The `claimedFiles` set tells the scanner which files this extractor owns. This is used for:
30- Conflict detection (multiple extractors claiming same file)
31- Path-based conflict resolution
32
33### Example Shape
34
35Each example must have:
36- `id` — Unique identifier
37- `title` — Display name
38- `rootPath` — Base directory
39- `files` — Array of file info
40- `metadata` — Your custom metadata type
41- `extractorName` — Your extractor's name
42
43## When to Create a Custom Extractor
44
45- Unique metadata format (TOML, INI, custom JSON schema)
46- Special directory structures
47- Language-specific conventions
48- Integration with other tools
49