All documentation pages

    9 min read

    Research Guide

    Step through the Research module from sources to results: adding papers, review, configuration, generating and running code, figures, exporting a ZIP, supported models, limits and recovery.

    What Research does

    Research turns a paper into an editable project: it reads the PDFs, shows you what it found and where, lets you settle conflicts and override settings, writes concrete code, runs it and produces measured figures on the Dashboard. It is a reproduction workflow. A generated project, a successful run and a reproduced published result are three different claims, and the module labels each separately.

    Research appears in the sidebar and in New Project → From research papers. To explore it without a paper or an AI key, choose Load original tutorial, a complete public-domain specification you can generate, train and chart.

    The workspace

    One research session is open at a time. The header holds the session switcher, Import and New research. A stage rail runs Sources, Review, Configuration, Generate and run, Results. Each stage shows what is saved ("1 decision needed", "Generated, not run", "Smoke results only"), and only a completed check or run earns a success mark. Experiment and target selectors in the rail apply to every stage. Any Evidence link opens a reader beside the stage showing the cited quote, the line in the extracted text and, when the PDF is stored, the rendered page.

    1. Sources

    Drop PDFs, or fetch from an arXiv identifier, abstract page or PDF address, or supply text for scanned pages. Mark exactly one included document primary, its supplements, and background papers reference. Reference papers cannot supply this experiment's hyperparameters. A source can be excluded or removed; either detaches the current specification until the next analysis.

    Before anything is sent, the stage names the provider and shows the session's request and token budget. Set up the provider under Settings; the key stays in memory unless you choose Remember key.

    2. Analyze and review

    Choose Analyze sources. Each stage is saved as it validates, so a retry reuses finished work. In Review, settle competing values under Decisions, check each experiment's operations and inferred shapes, and read what had to be assumed. Evidence opens the exact preserved span. Quotes must match the supplied text exactly, and one repair attempt is allowed before a failure is reported.

    3. Configuration

    Every setting is listed with where its value came from (the paper, a default, or your override). Filter by name or origin, and open a setting's audit for its rule, reported candidates, override history and the code that consumes it. Edits form a draft: Apply configuration and regenerate creates a new specification revision, and Discard changes drops the draft. An unresolved conflict blocks generation.

    4. Generate and run

    Generate writes the project files. For files you edited, choose Keep or Replace. Open Code to read them; the Model Builder shows the pinned architecture, and Research code stays authoritative. The reproduction ladder beside the stage states separately whether the project is generated, smoke-checked, fully run and published.

    • Run a synthetic smoke check first. It is labelled as smoke wherever it appears.
    • For a full run, bind complete data, the feature order, labels, class order and split protocol. Complete Data Hub CSV, TSV, JSON or derived versions can be copied in; remote, Parquet and preview-only sources must first be materialised into a supported format.
    • A run's measured history is drawn as it arrives.
    • A project holds one generated project. After a newer specification revision, run actions wait until you generate again.

    5. Results and figures

    Results lists runs. Each shows its metrics, training history and confusion counts, which figure intents it could fill (with the reason for each one it could not), and its frozen configuration and hashes. Completed evaluations create Dashboard figures automatically, even when Research and the Dashboard are not on screen. Figures are normal editable widgets; your edits to titles, palettes and position are kept when figures are regenerated. Deleted figures stay deleted unless you choose Restore deleted figures.

    Figure types are learning curves, complete confusion matrices, single-run method comparison bars, regression predictions and residuals, and ROC and precision-recall curves when the positive class is declared. Curves are never smoothed or invented, and numbers reported in the paper stay separate from your measurements.

    What it can generate

    • Tabular MLP classifiers and regressors
    • Small image CNNs
    • Residual or branched classifiers with add and concat
    • An explicit logistic-regression baseline

    in PyTorch, TensorFlow.js (CPU) or scikit-learn, depending on the family. Supported neural operators are dense, conv2d, max and average pooling, flatten, global average pooling, ReLU, tanh, sigmoid, softmax, dropout, batch normalisation, add and concat. Anything else fails explicitly instead of producing a model that looks right and is not. Full tabular runs use numeric features with training-only standardisation; image full runs use train, validation and test class folders at the model's exact dimensions; browser full runs support tabular JSON rows.

    Running outside the browser

    Generated PyTorch and scikit-learn projects run with Python 3.11 or newer:

    code
    pip install -r requirements.txtpython train.py --config configs/smoke.json --run-id smokepython evaluate.py --run-id smoke --split testpython plot_results.py --run-id smokepython self_test.py

    TensorFlow.js projects use Node:

    code
    npm installnode train.mjs --config configs/smoke.json --run-id smokenode self_test.mjs

    For paper data, use configs/effective.json and a fresh run ID. Passing self-tests prove the template works on its fixture, not that it reproduces a published accuracy.

    You can also submit a frozen snapshot to an attached Jupyter kernel. Training runs in two-epoch segments; Resume next segment continues from saved optimizer and random state, and Reconcile remote run checks artifacts after a disconnect without restarting. Snapshots are limited to 1 MiB of encoded text. See Remote compute.

    Export and import

    Export a Research ZIP to move code, configuration, specification revisions, runs and Dashboard formatting and data to another machine; including original PDFs is optional. Import needs no AI connection. ZIPs exclude credentials, and every path and checksum is validated before import.

    Limits

    • Up to 10 sources per session, 25 MiB per PDF, 200 pages per document and 500 pages per session
    • One primary reproduction experiment at a time
    • Textless pages need supplied text or must be excluded
    • Multi-paper comparison, generic attention, OCR extraction and arbitrary model families are not supported yet
    • Local writes require the tab that owns the project; a second tab is blocked

    Recovery

    Interrupted work is durable while the Studio is open. If a folder publication was interrupted, Recover folder publication restores the files safely and can be run repeatedly. Failed figure jobs show their diagnostic and a retry in the Job Centre.

    The shorter overview is Research. Follow Turn a Paper into a Running Experiment for a worked example.