> For the complete documentation index, see [llms.txt](https://docs.owkin.com/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://docs.owkin.com/what-you-can-do-with-k-pro/writing-reproducible-skills.md).

# Writing Reproducible Skills

### 0. What this guide is for

This guide is for anyone who wants to create reliable, reusable **skills** in K Pro.

A **skill** is a packaged set of instructions that teaches an LLM to perform a specific scientific, analytical, or operational task consistently. A good skill lets users avoid re-explaining the same method, assumptions, and quality standards every time they need the task performed.

This guide explains how to write skills that are:

1. **Discoverable** — the agent can select the right skill at the right time.
2. **Scoped** — the skill does one coherent job and knows when not to run.
3. **Reproducible** — the same input produces the same output, within documented tolerances.
4. **Executable** — every step maps to an available runtime capability.

#### Quick overview

**What does a skill mean?**

A **skill** is a reusable instruction package that teaches an LLM to complete a specific task consistently. It is more than a prompt: it defines the task objective, the required inputs, the recommended workflow, the expected output, and the checks the LLM should perform before returning an answer.

In practice, a skill helps users avoid repeating the same instructions every time they need the same type of work done. For example, instead of explaining how to extract adverse events from a clinical document each time, a skill can capture that method once and reuse it.

**When to use a skill**

Use a skill when the task is repeatable and benefits from a standard approach. A skill is useful when:

* The same type of request comes up often.
* The task follows a known method, checklist, or template.
* The output needs to be consistent across users or projects.
* The LLM must respect specific rules, quality checks, or escalation criteria.
* The workflow has several steps that should happen in a defined order.

Good examples include extracting structured information from documents, preparing standard reports, reviewing content against defined criteria, running a recurring analysis, or producing outputs in a required format.

**When not to use a skill**

Do not use a skill when the request is simple, one-off, or too open-ended to follow a reusable method. A skill is usually not needed when:

* The user only needs a quick answer or small rewrite.
* The task changes completely from case to case.
* The user is brainstorming and does not need a fixed output.
* The LLM lacks the required data, tool, or permission to complete the task.
* The task requires expert judgment before any automated workflow should proceed.
* Another existing skill already covers the same need.

A simple rule is: use a skill when you can say, “For this kind of request, the LLM should follow this same method each time.”

**Atomic skills and workflow skills**

An **atomic skill** is a reusable building block. A **workflow skill** combines several actions into an end-to-end process. When in doubt, start with an atomic skill; workflow skills can be created later by combining well-defined atomic skills.

There are two common skill types:

| Skill type         | Meaning                                                           | Use when                                                             | Example                                                                |
| ------------------ | ----------------------------------------------------------------- | -------------------------------------------------------------------- | ---------------------------------------------------------------------- |
| **Atomic skill**   | A focused skill that performs one clear task from start to finish | The output is useful on its own and can be reused in other workflows | Extract action items from meeting notes                                |
| **Workflow skill** | A broader skill that coordinates several steps or smaller skills  | The task has phases, dependencies, or decision points                | Produce a full research report from source collection to final summary |

**Best practices**

When writing a skill, follow these best practices:

* **Use a specific name:** make it clear what the skill does and what domain or input it applies to.
* **Write a clear description:** explain when to use the skill, what it produces, and when not to use it.
* **Define the objective:** state the task the agent is expected to complete.
* **List required inputs:** make clear what the user must provide before the skill can run.
* **Describe the workflow:** provide steps that are precise enough to execute and clear enough for users to understand.
* **Specify expected outputs:** define the format, content, and level of detail expected.
* **Add quality checks:** tell the agent how to verify that the result is complete and reliable.
* **Set boundaries:** include what is in scope, what is out of scope, and when to ask for human review.
* **Keep it reusable:** avoid references to a specific conversation, temporary file, or undocumented context.
* **Use accessible language:** keep the guide understandable for external users while preserving the technical details needed to run the skill correctly.

#### 0.1 What is inside a `SKILL.md` file

A `SKILL.md` file has three main layers: **name**, **description**, and **content**.

```yaml
---
name: differential-expression-analysis-bulk-rna
description: >
  Perform a basic differential expression analysis on bulk RNA-seq count data.
  Use when the user provides a gene count matrix and sample metadata, and asks
  for differentially expressed genes between two conditions. Produces a ranked
  differential expression table, a short interpretation, and optional plots.
  Do not use for single-cell RNA-seq or spatial transcriptomics workflows.
---

# Differential expression analysis for bulk RNA-seq

## Objective

Identify genes that are significantly upregulated or downregulated between two
conditions using a reproducible bulk RNA-seq differential expression workflow.

## Required inputs

- Count matrix with genes as rows and samples as columns
- Sample metadata with one condition column
- Contrast to test, such as `treated` versus `control`

## Workflow summary

1. Validate that sample names match between the count matrix and metadata.
2. Filter low-count genes using the documented threshold.
3. Run differential expression with fixed parameters.
4. Rank genes by adjusted p-value and log2 fold change.
5. Return the results table, summary, and quality notes.

## Expected outputs

- `datatable` — ranked gene-level results with log2 fold change, p-value, and adjusted p-value
- `text` — concise interpretation of the strongest signals
- `plot` — optional volcano plot or MA plot
```

1. **`name` — the skill identifier**

   The `name` is the stable identifier used to route, package, and reference the skill. It should be lowercase, kebab-case, specific enough to distinguish the skill from neighboring skills, and should match the skill folder and packaged `.skill` filename.
2. **`description` — the routing summary**

   The `description` is the short text the agent sees before it reads the full skill. It should explain what the skill does, when to use it, what output it produces, and when not to use it. A good description helps the agent select the right skill without over-triggering on related but different tasks.
3. **Content — the execution guide**

   The body of `SKILL.md` contains the actual instructions for running the skill. It should define the objective, success criteria, scope boundaries, required inputs, runtime dependencies, workflow steps, validation checks, reproducibility controls, expected outputs, and at least one example execution. This is the part that turns the skill from a vague method into a reproducible procedure.

***

### 1. Naming your skill

#### 1.1 Why naming matters

When an agent decides whether to use a skill, it primarily sees the skill’s `name` and `description`. It does **not** read the full `SKILL.md` before making that routing decision.

A vague name creates silent failures: the agent may select a neighboring skill, follow it correctly, and still return a plausible answer to the wrong question.

Treat the name as a routing signal, not as a cosmetic label.

#### 1.2 The recommended naming formula

Use:

```
[action] + [object] + [scope]
```

| Slot       | Answers                         | Examples                                                                                     |
| ---------- | ------------------------------- | -------------------------------------------------------------------------------------------- |
| **Action** | What does the skill do?         | extraction, scoring, profiling, clustering, retrieval, classification, detection, estimation |
| **Object** | What does it operate on?        | adverse-events, binding-pockets, gene-expression, copy-number, phenotypes, trial-outcomes    |
| **Scope**  | What makes this skill specific? | from-drug-labels, via-gwas, in-scrna-seq, across-protein-family                              |

Examples:

| Too ambiguous        | Better                                           |
| -------------------- | ------------------------------------------------ |
| `feature-extraction` | `histomorphological-feature-extraction-from-wsi` |
| `cell-clustering`    | `unsupervised-cell-clustering-scrna`             |
| `network-analysis`   | `gene-regulatory-network-inference-multi-omics`  |
| `trend-detection`    | `pharma-pipeline-trend-detection`                |

#### 1.3 Identifier format rules

The skill `name`, folder name, and packaged `.skill` filename must match exactly.

| Rule                | Convention                       | Example                                            |
| ------------------- | -------------------------------- | -------------------------------------------------- |
| Case                | Lowercase only                   | `gwas-association-lookup`                          |
| Separator           | Hyphens only                     | `adverse-event-extraction-from-drug-labels`        |
| Special characters  | Avoid all special characters     | `km-cox-regression`, not `km/cox-regression`       |
| Version suffixes    | Do not include versions in names | not `gwas-lookup-v2`                               |
| Acronyms            | Lowercase, no dots               | `scrna`, `wsi`, `dea`, `gwas`, `pca`, `umap`, `qc` |
| Length              | Aim for ≤ 60 characters          | abbreviate the scope if needed                     |
| Minimum specificity | At least 3 semantic tokens       | `quality-control-omics`, not `qc`                  |

Recommended folder structure:

```
my-skill-name/
└── SKILL.md

# packaged as:
my-skill-name.skill
```

#### 1.4 Naming rules

1. **Use at least three semantic tokens.**

   Stopwords such as `from`, `of`, and `for` do not count. One- and two-token names are usually ambiguous at scale.
2. **Front-load a discriminating action.**

   If many skills share the same domain noun, vary the leading verb: `parse-ae-from-clinical-trial-text`, `scrape-ae-from-congress-abstracts`, `extract-ae-from-drug-labels-articles-faers`.
3. **Add the data modality when needed.**

   Use `[action]-[object]-[modality]` when the same action and object exist across modalities.

   | Full name                 | Recommended abbreviation |
   | ------------------------- | ------------------------ |
   | Bulk RNA-seq              | `bulk-rna`               |
   | Single-cell RNA-seq       | `scrna`                  |
   | Spatial transcriptomics   | `spt`                    |
   | Whole-slide image         | `wsi`                    |
   | Copy number               | `cna`                    |
   | Mutation / WES            | `mutation`               |
   | Proteomics                | `proteomic`              |
   | Epigenetics               | `epigenetic`             |
   | Clinical trial data       | `clinical-trial`         |
   | Text / literature         | `text`                   |
   | Tabular / structured data | `tabular`                |
4. **Separate method-level and application-level skills.**

   A general statistical method and a domain-specific use of that method should be separate:

   * `survival-analysis-km-cox`
   * `survival-association-testing-target-expression`
5. **Scope by source when skills extract the same entity.**

   Adverse events from clinical trial protocols, congress abstracts, drug labels, and FAERS are different skills.
6. **Standardize formatting.**

   Use lowercase kebab-case for identifiers. Use sentence case for display names, with standard uppercase acronyms where appropriate.
7. **Merge or explicitly split near-duplicates.**

   If two skill names remain confusable after applying the rules above, either merge them or add a clearer differentiator.
8. **Avoid same-surface duplicates.**

   Two skills that reduce to the same platform tool call over the same input table are usually one skill.
9. **Preview the slug.**

   The platform normalizes names before file lookup. Names that differ only by punctuation, spacing, or case can collide.

   ```python
   import re
   slug = re.sub(r"[^a-z0-9]+", "-", name.lower()).strip("-")
   ```

#### 1.5 Naming decision flow

```
1. Draft a name as [action] + [object] + [scope].
2. Count semantic tokens, excluding stopwords. Are there at least three?
   - No: add modality, source, method, or domain context.
3. Does another skill share the same action or object?
   - Yes: add a stronger differentiator.
4. Does another skill use the same data modality?
   - Yes: make the source, method, or output explicit.
5. Convert to lowercase kebab-case.
6. Read the name in isolation.
   - If a colleague could not pick the right skill from the name alone, revise it.
```

***

### 2. Writing the description

#### 2.1 The description is the trigger

The `description` field determines when the skill is selected. It must avoid two failure modes:

* **Under-triggering:** the skill is not selected when it should be.
* **Over-triggering:** the skill is selected when a neighboring skill would be better.

Write the description as a selector, not as a full manual.

#### 2.2 Recommended description structure

Target **60–120 words**.

Use this structure:

```
WHAT: What the skill does.
WHEN: When the agent should use it.
HOW: What outputs it produces.
NOT: When not to use it, with redirects if relevant.
ALSO: Synonyms, informal user phrasing, or edge cases that should still trigger it.
```

**WHAT**

Mirror the name. If the name says `from-wsi`, the description should say “whole-slide images” or “WSI”.

Good:

> Extract and structure adverse events (AEs) from clinical trial text documents, including protocols, clinical study reports (CSRs), and investigator brochures.

Weak:

> Analyze clinical text for safety information.

**WHEN**

Include realistic user requests, input types, and workflow contexts.

Good:

> Use when the user provides trial documents and asks for “AE extraction,” “safety tables,” “CTCAE grading,” or “side effects reported in this study.”

**HOW**

State the expected artifact.

Good:

> Produces a structured table mapping each adverse event to MedDRA preferred term, CTCAE grade, frequency, dose relationship, and causal attribution.

**NOT**

Name the closest confusion partners.

Good:

> Do not use for adverse events from congress abstracts; use `scrape-ae-from-congress-abstracts`. Do not use for FAERS or drug-label analysis; use `extract-ae-from-drug-labels-articles-faers`.

**ALSO**

Add phrases users might use without naming the method.

Good:

> Also trigger for informal requests such as “what toxicities were seen in the phase I?” or “summarize the safety profile from this CSR.”

#### 2.3 Description length by collision risk

| Collision risk | Target length | Required sections                   |
| -------------- | ------------- | ----------------------------------- |
| Low            | 40–60 words   | WHAT + WHEN + HOW                   |
| Medium         | 60–90 words   | Add NOT                             |
| High           | 90–120 words  | Full WHAT / WHEN / HOW / NOT / ALSO |

Avoid descriptions above \~150 words. Put the detailed procedure in the body of `SKILL.md`.

#### 2.4 Confusion-partner audit

Before finalizing the description, compare it against neighboring skills.

| Question                                                 | Action                                 |
| -------------------------------------------------------- | -------------------------------------- |
| Does another skill share ≥ 2 keywords?                   | Add a NOT redirect.                    |
| Does another skill use the same data modality?           | State your modality explicitly.        |
| Does another skill answer a similar biological question? | Make the output format explicit.       |
| Could a user reasonably confuse them?                    | Add the distinguishing phrase to ALSO. |
| Do both skills call the same tool on the same input?     | Consider merging them.                 |

Keep a short audit table during review:

| Confusion partner                            | Shared keywords            | Discriminating signal                             |
| -------------------------------------------- | -------------------------- | ------------------------------------------------- |
| `scrape-ae-from-congress-abstracts`          | adverse events, extraction | Source is congress abstracts, not trial documents |
| `extract-ae-from-drug-labels-articles-faers` | adverse events, safety     | Source is drug labels, publications, or FAERS     |

#### 2.5 Description anti-patterns

| Anti-pattern          | Why it fails                   | Fix                                 |
| --------------------- | ------------------------------ | ----------------------------------- |
| Too generic           | Matches too many skills        | Name source, modality, and output   |
| Too narrow            | Misses valid user phrasings    | Add synonyms and informal triggers  |
| Broader than the name | Creates routing conflicts      | Match the name’s exact scope        |
| No anti-triggers      | Fails near neighbouring skills | Add NOT redirects                   |
| Too long              | Hard to use for selection      | Move detail into `SKILL.md`         |
| Implementation-heavy  | Irrelevant to routing          | Describe inputs and outputs instead |

#### 2.6 Worked metadata example

```yaml
---
name: parse-ae-from-clinical-trial-text
description: >
  Extract and structure adverse events (AEs) from clinical trial text documents,
  including protocols, clinical study reports (CSRs), and investigator brochures.
  Use when the user provides trial documents and asks for "AE extraction",
  "safety tables", "CTCAE grading", or "side effects from this trial". Produces
  a structured table with MedDRA terms, CTCAE grades, frequencies, and causal
  attribution. Do not use for congress abstracts; use
  `scrape-ae-from-congress-abstracts`. Do not use for drug labels, publications,
  or FAERS; use `extract-ae-from-drug-labels-articles-faers`. Also trigger for
  "what toxicities were seen in the phase I?"
---
```

***

### 3. Choosing the right skill scope

Skill granularity is a design decision. If it's too small, the LLM must orchestrate too many fragments. Too large, and the skill becomes hard to test, reuse, or debug.

#### 3.1 Atomic skills

An **atomic skill** performs one coherent unit of work.

Use an atomic skill when:

* The intermediate output is meaningful on its own.
* The step can be reused across workflows.
* Failure can be repaired without restarting the entire analysis.
* Human review is expected between steps.
* The skill has a clear input and output contract.

Examples:

```
quality-control-omics
alignment-and-quantification
differential-expression-analysis-dea
```

#### 3.2 Workflow skills

A **workflow skill** orchestrates multiple atomic skills and defines the logic connecting them.

Use a workflow skill when:

* The sequence is fixed or follows documented branches.
* Later decisions depend on earlier outputs.
* The final result is the main artifact.
* Restarting from the beginning is acceptable if an early step fails.
* The workflow has explicit decision gates.

#### 3.3 Granularity checklist

| Factor            | Atomic skill          | Workflow skill               |
| ----------------- | --------------------- | ---------------------------- |
| Coupling          | Low                   | High                         |
| Reusability       | High                  | Lower                        |
| Decision points   | Few                   | Many                         |
| Failure recovery  | Restart from a step   | Often restart from beginning |
| Human checkpoints | Between steps         | At workflow boundaries       |
| Output meaning    | Each output is useful | Final output matters most    |

#### 3.4 The restart test

Ask:

> If this failed halfway through, where would a human scientist restart?

If a human would restart from an intermediate artifact, split the skill there. If they would rerun the whole sequence, keep it together.

#### 3.5 Workflow skill template

```yaml
---
name: bulk-rnaseq-differential-expression-workflow
description: >
  Run an end-to-end bulk RNA-seq differential expression workflow by
  orchestrating quality control, quantification, and differential expression
  analysis. Use when the user provides bulk RNA-seq count or FASTQ-derived
  inputs and wants a reproducible DEG report. Produces QC summaries, a ranked
  DEG table, plots, and a methods report. Do not use for single-cell RNA-seq;
  use the relevant scRNA-seq workflow skill.
---
```

```markdown
# Bulk RNA-seq differential expression workflow

## Objective

Given bulk RNA-seq data from two experimental conditions, produce a reproducible ranked list of differentially expressed genes.

## Skill dependencies

1. `quality-control-omics`
2. `alignment-and-quantification`
3. `differential-expression-analysis-dea`

## Workflow logic

### Phase 1 — QC and preprocessing

Delegate to `quality-control-omics`.

**Decision gate:** if read retention is below 85%, stop and ask for human review.

### Phase 2 — Quantification

Delegate to `alignment-and-quantification`.

**Decision gate:** if alignment rate is below 70%, stop and recommend contamination or reference-genome checks.

### Phase 3 — Differential expression

Delegate to `differential-expression-analysis-dea`.

Use FDR < 0.05 and |log2FC| > 1 unless the user provides different thresholds.

## Human checkpoints

- After Phase 1: review the QC summary.
- After Phase 3: review the volcano plot and top-ranked genes before finalizing
```

### 4. Defining the objective and success criteria

A reproducible skill needs a precise contract. Every `SKILL.md` should begin with an objective block that states the scientific task, the success criteria, and the scope boundaries.

#### 4.1 SMART criteria

| Criterion    | Requirement                     | Example                                                            |
| ------------ | ------------------------------- | ------------------------------------------------------------------ |
| Specific     | Name the exact task             | “Perform DESeq2 differential expression,” not “analyze data”       |
| Measurable   | Include thresholds or metrics   | “FDR < 0.05”, “AUC > 0.80”                                         |
| Actionable   | Provide executable steps        | “Filter genes with counts per million < 1 in fewer than 3 samples” |
| Reproducible | Pin parameters, versions, seeds | “Use random\_state=42”                                             |
| Testable     | Define validation checks        | “Output must pass schema validation”                               |

#### 4.2 Objective block template

```markdown
## Objective

[One sentence stating the scientific or analytical task.]

## Success criteria

- [ ] Primary output: [description and output type]
- [ ] Quality threshold: [metric and threshold]
- [ ] Validation check: [how correctness is verified]
- [ ] Reproducibility check: [seed, version, deterministic ordering, or tolerance]

## Scope boundaries

**In scope**

- [Inputs, modalities, and analyses this skill handles]

**Out of scope**

- [Inputs, modalities, or questions that require another skill]
- [Cases requiring human review]
```

#### 4.3 Example objective block

```markdown
## Objective

Identify genes significantly upregulated in treated versus control bulk RNA-seq samples using a documented differential expression workflow.

## Success criteria

- [ ] Primary output: ranked differential expression table (`datatable`)
- [ ] Plot output: volcano plot with significant genes highlighted (`plot`)
- [ ] Quality threshold: adjusted p-value < 0.05 and log2 fold change > 1
- [ ] Validation check: output table contains `gene_id`, `gene_name`, `log2FoldChange`, `pvalue`, and `padj`
- [ ] Reproducibility check: all stochastic operations use `random_state=42`

## Scope boundaries

**In scope**

- Bulk RNA-seq count matrices
- Two-condition comparisons
- Basic covariate adjustment if metadata is provided

**Out of scope**

- Single-cell RNA-seq analysis
- Time-series modelling
- Survival association testing
```

***

### 5. Reproducibility requirements

Reproducibility is not optional. A skill is reproducible when another qualified person can execute it independently and obtain the same result, within documented tolerance.

#### 5.1 Required practices

Every skill must:

1. **Pin library versions.**

   Use versions available in the execution environment. The environment lock file or runtime manifest is the source of truth.
2. **Set random seeds.**

   Use `random_state=42` unless another seed is justified. Apply this to train/test splits, sampling, UMAP, bootstrapping, permutations, and any stochastic model.
3. **Document parameters and defaults.**

   Include thresholds, filters, model choices, reference resources, and fallback behavior.
4. **Use deterministic data ordering.**

   Every SQL query must end with `ORDER BY <deterministic_key>`. Without ordering, row order is undefined.
5. **Avoid wall-clock time in result logic.**

   Do not use `datetime.now()` or “today” as an implicit parameter. If a date matters, require the user or calling workflow to pass it explicitly.
6. **Compare floating-point values with tolerance.**

   Never require exact equality for floats. State acceptable tolerances, such as `abs(actual - expected) < 1e-6`.
7. **Document data provenance.**

   State which input files, database tables, filters, and external knowledge sources were used.
8. **Make failure modes explicit.**

   State when to stop, retry, ask for clarification, or escalate to human review.

#### 5.2 Example workflow execution

Every skill should include at least one complete example execution. This is both documentation and a baseline for evaluation.

Minimum coverage:

1. Mock user request.
2. Required inputs.
3. First tool call or runtime capability used.
4. SQL query, if platform data is queried, including deterministic `ORDER BY`.
5. Plotting recipe, if visual output is expected.
6. Expected outputs and their output types.
7. Validation checks.

Example:

```markdown
## Example workflow execution

**Mock user request**

"I have RNA-seq count data from 3 treated and 3 control liver samples. Can you identify significantly upregulated genes?"

**Required inputs**

- Count matrix with genes as rows and samples as columns
- Sample metadata with condition labels
- Optional covariates

**Expected response flow**

1. Validate that the count matrix contains exactly six sample columns.
2. Check for duplicated gene identifiers.
3. Filter low-count genes using the documented threshold.
4. Run differential expression with fixed parameters.
5. Filter results at FDR < 0.05 and log2FC > 1.
6. Return a ranked table and volcano plot.

**Expected outputs**

- `datatable` — ranked differential expression table with `gene_id`, `gene_name`, `baseMean`, `log2FoldChange`, `pvalue`, and `padj`
- `plot` — volcano plot with significant genes highlighted
- `report` — methods and results narrative
- `text` — concise executive summary
```

***

### 6. Tools, scripts, and delegation

A skill can describe reasoning steps, invoke runtime tools, or delegate to another skill. Choose the most deterministic mechanism that fits the task.

#### 6.1 When to implement a tool versus delegate to the agent

| Criterion            | Prefer a tool or script     | Prefer agent reasoning  |
| -------------------- | --------------------------- | ----------------------- |
| Determinism required | Yes                         | No                      |
| Logic complexity     | High                        | Low to moderate         |
| Error sensitivity    | High                        | Low                     |
| Reuse frequency      | Reused often                | One-off                 |
| Domain expertise     | Encoded rules or thresholds | Standard interpretation |
| Output schema        | Strict                      | Flexible                |

Examples:

* Use a tool for schema validation, statistical computation, model scoring, file conversion, and plotting.
* Use agent reasoning to explain results, compare trade-offs, or write a narrative summary after validated outputs are produced.

#### 6.2 Tool documentation template

Every tool or script referenced by a skill must be documented clearly.

```markdown
## Tool: qc_check

**Purpose**

Validate sequencing quality metrics against predefined thresholds.

**Version**

1.2.0

**Inputs**

- `--input`: path to QC summary file, TSV format
- `--thresholds`: path to thresholds JSON file; default `default_thresholds.json`

**Outputs**

- STDOUT: JSON object with pass/fail status per metric
- Exit code 0: all checks passed
- Exit code 1: one or more quality thresholds failed
- Exit code 2: unexpected error such as missing input or malformed config

**Error handling**

- Missing input: print `Input file not found: {path}` and exit 2
- Malformed config: print `Malformed thresholds file: {path}` and exit 2
- Invalid metric: print `Unknown metric: {metric}` and exit 1
```

#### 6.3 Error handling rules

Scripts must never fail silently.

Required behavior:

* Validate inputs before processing.
* Include the path, expected format, and observed problem in error messages.
* Use distinct exit codes for success, expected validation failure, and unexpected runtime failure.
* Log non-critical warnings without halting.
* Return structured errors where possible.
* Do not expose raw stack traces as the primary user-facing error.

#### 6.4 Delegating to other skills

Delegate to another skill when the downstream task is already captured by a reviewed skill.

When delegating:

* Use the target skill’s exact slug.
* Pass a self-contained prompt. Do not assume the delegated skill can see your conversation history, files, or previous tool outputs unless those are explicitly provided.
* State the expected output type.
* State any thresholds, seeds, and constraints that must carry over.
* Validate the delegated output before using it downstream.

Example:

```markdown
Delegate to `quality-control-omics` with:

- Input: path to FASTQ-derived QC metrics table
- Required threshold: read retention ≥ 85%
- Expected output: QC summary (`datatable`) and recommendation (`text`)
- Stop condition: if any required QC metric fails, halt workflow and request human review
```

***

### 7. Runtime contract

This section explains how a skill must map its instructions to the runtime environment where the agent operates. This is important for external users because a scientifically sound skill can still fail if it assumes capabilities that are not available at execution time.

A skill should never say “fetch the data,” “run the analysis,” or “plot the results” unless it also says **which runtime capability does that work, with what inputs, and what output is expected**.

#### 7.1 Execution model

Skills execute inside K through an agent runtime. The skill is an instruction package; it is **not** an independent program with unrestricted system access.

Key implications:

1. **No ambient access.**

   A skill cannot assume direct access to the internet, a local filesystem, internal APIs, notebooks, databases, or prior conversation state.
2. **Every external action must go through an available runtime capability.**

   Data retrieval, computation, literature search, plotting, and file generation must be expressed as explicit tool-backed steps.
3. **A skill must be self-contained.**

   Only the contents of `SKILL.md` and explicitly bundled assets should be assumed available at runtime.

#### 7.2 Runtime capabilities

The runtime exposes capabilities through MCP-backed tools. In workflow steps, refer to tools by their fully qualified runtime identifiers.

| Capability                 | Runtime service                                 | Typical use                                                              |
| -------------------------- | ----------------------------------------------- | ------------------------------------------------------------------------ |
| Plot rendering             | `mcp__mcp-gateway__plotter__...`                | Generate validated plots from tabular results                            |
| Prior biological knowledge | `mcp__mcp-gateway__bioknowledge-navigator__...` | Query curated resources such as DepMap, GTEx, IMPC, or Open Targets      |
| Scientific literature      | `mcp__mcp-gateway__consensus__...`              | Retrieve and summarize literature-backed evidence                        |
| Clinical trials            | `mcp__mcp-gateway__trials-navigator__...`       | Search, filter, and compare trial records                                |
| Digital pathology          | `mcp__mcp-gateway__pathology-explorer__...`     | Analyze pathology datasets and whole-slide-image-derived outputs         |
| Open web search            | `mcp__mcp-gateway__web-search__...`             | Retrieve public web evidence when allowed by the workflow                |
| Sandboxed code execution   | `mcp__mcp-gateway__kodexec__...`                | Run deterministic scripts, validation logic, or statistical calculations |

Use the exact tool identifiers available in the target deployment. If a tool name differs between environments, the skill must either use the deployment-specific name or state the required tool dependency explicitly.

#### 7.3 Writing executable workflow steps

Each workflow step that uses a tool must include:

* **Tool identifier:** the fully qualified runtime tool name.
* **Purpose:** why the tool is being used.
* **Inputs:** files, tables, parameters, filters, thresholds, and seeds.
* **Expected output:** output type and schema.
* **Validation:** how to check the output is acceptable.
* **Failure behavior:** what to do if the tool fails or returns insufficient data.

Weak step:

```markdown
Search the literature for evidence linking the target to the disease.
```

Executable step:

```markdown
Use `mcp__mcp-gateway__consensus__search` to retrieve literature evidence linking `{target_gene}` to `{disease}`.

Inputs:
- `target_gene`: HGNC gene symbol provided by the user
- `disease`: disease name or ontology term provided by the user
- `max_results`: 20

Expected output:
- `datatable` with columns `title`, `year`, `journal`, `claim`, `evidence_strength`, and `url`
- `text` summary of the top supported claims

Validation:
- Exclude records without a source URL or publication year.
- If fewer than three relevant records are returned, state that evidence is insufficient and do not over-interpret.

Failure behavior:
- If the tool is unavailable or returns no results, stop and report that the literature retrieval step could not be completed.
```

#### 7.4 Sandboxed code execution

Use sandboxed code execution for deterministic transformations, validation, statistics, and file generation.

A skill that uses code execution must specify:

* Language and runtime, if relevant
* Required package versions or runtime manifest
* Input file paths or data objects
* Output file paths
* Random seeds
* Expected stdout/stderr behavior
* Exit codes
* Validation checks

Example:

```markdown
Use `mcp__mcp-gateway__kodexec__run` to execute `validate_deg_table.py`.

Inputs:
- `deg_table_path`: path to differential expression CSV
- Required columns: `gene_id`, `gene_symbol`, `log2FoldChange`, `pvalue`, `padj`

Parameters:
- `max_missing_padj_fraction`: 0.01
- `float_tolerance`: 1e-6

Expected output:
- `text` validation summary
- Exit code 0 if the table is valid
- Exit code 1 if validation fails
- Exit code 2 for unexpected runtime errors
```

#### 7.5 Output contract

Every skill output must resolve to exactly one of the supported content types:

```
text · plot · file · data · datatable · report
```

Use the content type intentionally:

| Output type | Use for                                | Examples                                       |
| ----------- | -------------------------------------- | ---------------------------------------------- |
| `text`      | Short narrative answer                 | one-paragraph summary, warning, recommendation |
| `plot`      | Rendered visualization                 | volcano plot, Kaplan-Meier curve, forest plot  |
| `file`      | Downloadable artifact                  | CSV, PDF, zipped outputs, model artifact       |
| `data`      | Structured object for downstream use   | JSON object, metrics dictionary                |
| `datatable` | Tabular result intended for inspection | ranked genes, trial comparison table           |
| `report`    | Longer structured deliverable          | methods + results narrative, evidence report   |

Every expected output section should name the type:

```markdown
## Expected outputs

- `datatable` — ranked DEG table with `gene_id`, `gene_symbol`, `log2FoldChange`, `pvalue`, and `padj`
- `plot` — volcano plot with significance thresholds displayed
- `report` — methods and results narrative, including parameters and limitations
- `text` — concise executive summary
```

#### 7.8 Bundled assets

Only `SKILL.md` is guaranteed to be present at runtime unless additional assets are explicitly bundled with the skill.

If the skill requires supporting files, list them in `SKILL.md`:

```markdown
## Bundled assets

- `deg_table.schema.json` — validates differential expression output
- `methods_report.md` — report template
- `validate_deg_table.py` — deterministic validation script
```

For each bundled asset, state:

* Purpose
* Whether it is required or optional
* How it is invoked

Do not reference local notebooks, private paths, or unpublished files that are not included in the skill package.

#### 7.9 Secrets, credentials, and restricted data

Skills must never include secrets, tokens, passwords, private API keys, or credentials.

If a workflow requires authenticated access:

* State the required runtime service.
* Do not embed credentials in `SKILL.md`.
* Do not ask users to paste secrets into the prompt.
* Fail safely if the service is unavailable.
* Avoid copying restricted data into broader outputs unless the workflow explicitly permits it.

***

### 8. Recommended `SKILL.md` template

Use this template as a starting point.

```markdown
---
name: [lowercase-kebab-case-skill-name]
description: >
  [WHAT: what the skill does.] [WHEN: when to use it, including user phrases.]
  [HOW: expected outputs.] [NOT: when not to use it, with redirects.]
  [ALSO: synonyms or edge cases that should still trigger it.]
---

# [Human-readable skill title]

## Objective

[One sentence describing the task.]

## Success criteria

- [ ] Primary output: [description] (`output_type`)
- [ ] Quality threshold: [metric and threshold]
- [ ] Validation check: [schema, sanity check, or benchmark]
- [ ] Reproducibility check: [seed, ordering, version, or tolerance]

## Scope boundaries

**In scope**

- [Supported inputs]
- [Supported analyses]

**Out of scope**

- [Unsupported inputs]
- [Cases requiring another skill or human review]

## Required inputs

| Input | Required? | Source | Format | Validation |
|---|---|---|---|---|
| `[input_name]` | Yes/No | User / prior step / runtime service / bundled asset | `[format]` | `[check]` |

## Runtime dependencies

| Capability | Tool identifier | Purpose | Required? |
|---|---|---|---|
| `[capability]` | `mcp__mcp-gateway__[service]__[tool]` | `[purpose]` | Yes/No |

## Workflow steps

### Step 1 — [name]

**Tool**

`[fully qualified tool identifier]`

**Purpose**

[Why this step is needed.]

**Inputs**

- `[input]`: [description]

**Parameters**

- `[parameter]`: [value]

**Expected output**

- `[output_type]` — [description]

**Validation**

- [Validation check]

**Failure behaviour**

- [What to do if the step fails]

## Reproducibility controls

- Random seed: `[seed]`
- Package versions or runtime manifest: `[reference]`
- Deterministic ordering: `[ordering rule]`
- Date handling: `[explicit date input, if applicable]`
- Float tolerance: `[tolerance]`

## Expected outputs

- `text` — [short summary]
- `datatable` — [table schema]
- `plot` — [plot description]
- `report` — [long-form report]
- `file` — [downloadable artifact, if any]
- `data` — [structured object, if any]

## Example workflow execution

**Mock user request**

"[Realistic request]"

**Expected response flow**

1. [Step]
2. [Step]
3. [Step]

**Expected outputs**

- `[output_type]` — [description]
```

***

### 9. Further reading

* Anthropic, *The Complete Guide to Building Skills for Claude*: <https://resources.anthropic.com/hubfs/The-Complete-Guide-to-Building-Skill-for-Claude.pdf>
* Claude Code skills documentation: <https://code.claude.com/docs/en/skills>
* *Agent Skills with Anthropic* course: <https://learn.deeplearning.ai/courses/agent-skills-with-anthropic>


---

# Agent Instructions
This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com.

## Querying This Documentation
If you need additional information that is not directly available in this page, you can query the documentation by asking a question.

Perform an HTTP GET request on the following URL with the `ask` and `goal` query parameters:

```
GET https://docs.owkin.com/what-you-can-do-with-k-pro/writing-reproducible-skills.md?ask=<question>&goal=<user_goal>
```

`ask` is the immediate question: it should be specific, self-contained, and written in natural language.
`goal` is what the user is ultimately trying to achieve, the reason they need the answer. Sharing it helps GitBook give you a better, more relevant answer. A goal is most helpful when it describes the outcome the user wants rather than restating the question. For example, with `ask=how do I create an API token`, a goal like `automate deployments from our CI pipeline` lets GitBook tailor the answer to that use case.

The response will contain a direct answer to the question and relevant excerpts and sources from the documentation.

Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections.
