# What is K Pro?

K Pro is an agentic AI platform designed to explore complex biology and accelerate biomedical research. It combines unparalleled access to multimodal data, cutting-edge AI to understand biology, and pioneering agentic AI capabilities to help researchers make faster, more confident, and better-informed decisions.

Currently, K-Pro is available in two versions:

* **K Pro Free:** The standard version for researchers and academics
* **K Pro:** The advanced version with enhanced capabilities for pharmaceutical teams

#### **Core capabilities**

**Consolidated knowledge from multiple sources:**

By pulling from credible data sources (including PubMed and comprehensive biological knowledge bases), K Pro answers scientific questions with synthesized, evidence-based insights. All data and recommendations are fully traceable, allowing source verification.

**Partnering with Consensus, leading AI-powered literature review service**

K Pro users can access literature review content powered by Consensus. Consensus is an AI-powered academic search engine built on a database of over 220 million peer-reviewed research papers. Unlike general AI tools, every response is tied back to a real research paper and is grounded in scientific research.

* [More information on Owkin’s partnership with Consensus](https://www.owkin.com/newsfeed/owkin-announces-partnership-with-consensus-to-strengthen-literature-intelligence-for-owkins-ai-scientist-k-pro)
* [More details on Consensus paper coverage](https://help.consensus.app/en/articles/10055108-consensus-research-database)

<figure><img src="/files/MxeYAZQ9K5TfRdTxbEWY" alt=""><figcaption></figcaption></figure>

**Smart hypothesis testing:**

You can use K Pro to validate scientific hypotheses, leveraging advanced statistical reasoning and up-to-date datasets. The system helps explore whether your ideas are supported by the current body of scientific evidence.

<figure><img src="/files/v5EXDpxKouFPUoMT4MlG" alt=""><figcaption></figcaption></figure>

**Scientific writing assistance:**

K Pro generates high-quality, publication-ready text tailored to your research needs, including review paragraphs, scientific summaries, and manuscript sections.

**Interactive data visualization:**

K Pro creates interactive plots and data visualizations using resources like TCGA and MOSAIC, making it easy to summarize, analyze, and interpret complex patient datasets. You can explore data by slicing and dicing variables to uncover deeper insights.

<figure><img src="/files/O7aNmi7oJx4WU57b2scO" alt=""><figcaption></figcaption></figure>

### Example: How K Pro differentiates from competition in the AI-for-target-prioritization space

1. **Multimodal patient data at depth.** Most platforms sit on top of aggregated public databases (Open Targets, GWAS Catalog, ChEMBL). K Pro is built on Owkin's data network of 164 academic medical institutions, with MOSAIC as the flagship cohort in oncology: one of the largest spatially-resolved multimodal oncology datasets, combining bulk RNA-seq, single-cell RNA-seq, spatial transcriptomics, whole-exome sequencing, H\&E and IHC pathology imaging, and longitudinal clinical follow-up. For target prioritization, this enables spatial co-localization of targets with TLS structures, malignant-cell specificity at single-cell resolution, intra-tumor heterogeneity quantification.
2. **Specialized AI models, wrapped in expert skills.** K Pro orchestrates calls to specialized models like HistoPLUS (Owkin's computational pathology foundation model, Adjadj et al., 2025), which detects 13 cell types including under-studied immune populations like neutrophils and eosinophils, with 14% better F1 classification than prior SOTA at 5× fewer parameters. It can also call Owkin Zero, our proprietary biological reasoning model (32B-parameter post-trained model on \~10M curated Q\&A pairs with reinforcement learning from verifiable biological rewards). These models are called by skills encoding specialized know-how — for example, how to orchestrate relevant data and tools to produce a target characterization package consistent with customer templates. Skills are provided by Owkin and can be customized per team.

For target prioritization specifically, K Pro is designed to improve several measurable outcomes:

* **Time to first target dossier** (literature + expression + competitive landscape): 2–4 weeks of analyst time before K Pro, vs. hours with K Pro, reviewed and refined collaboratively
* **Number of targets evaluated in parallel for a given indication**: a handful before (constrained by analyst capacity), vs. dozens in parallel and consistent campaigns with K Pro
* **Reproducibility of scoring across analysts and projects**: same skill applied consistently with versioned cohort definitions
* **Evidence coverage (modalities per target)**: often bulk RNA-seq plus literature before, vs. bulk + single-cell + spatial RNA + pathology + clinical landscape in addition to literature with K Pro


# Quickstart guide

This guide gets you from zero to a real K Pro result in two prompts. For background on what K Pro does, see What is K Pro?

### Your first prompt

New to K Pro? Try this prompt right now:

> List the top 10 genes over-expressed in the MOSAIC Bladder cancer cohort compared to TCGA-BLCA.

You'll get summary tables and distributions showing the clinical breakdown of the cohort.

<figure><img src="/files/R1FT7x3RhCwMonAonSOY" alt=""><figcaption></figcaption></figure>

### Your follow-up

From there, try a follow-up:

> List the genes unregulated in the MOSAIC cohort and identify from the literature if they are considered as therapeutic targets.

This is the core loop in K Pro: ask a question, get a result, refine with a follow-up.

### Where to go next

* **Want better prompts?** Read the [Prompting guide & prompt library](https://docs.owkin.com/what-you-can-do-with-k-pro/prompting-guide-and-prompt-library) for the A-S-R-T-C framework, common failure modes, and a library of tested prompts organized by use case.
* **Want a tour by goal?** See [What you can do with K Pro ](https://docs.owkin.com/what-you-can-do-with-k-pro/overview)— how to explore a cohort, compare groups, rank candidate targets, review the literature, and more.
* **Curious what data is available?** Browse the [Dataset Catalog](https://k.owkin.com/explore-data/overview).

**Note:** K Pro is designed to be your collaborative research partner, helping you navigate complex biological data and accelerate discoveries in biomedical research. As you become more familiar with the platform, you'll discover increasingly powerful ways to leverage its capabilities for your specific research needs.


# Prompting guide & prompt library

Learn how to get the best results from K-Pro. This guide covers the principles behind effective prompts and provides a ready-to-use library organized by use case.

### Start here

New to K-Pro? Try this prompt right now:

> **Provide clinical characteristics of the TCGA-BRCA cohort. Include the distribution of tumor stage and age at diagnosis.**

<figure><img src="/files/clrFp16Vp9uBgDWXY2h2" alt=""><figcaption></figcaption></figure>

You'll get summary tables and distributions showing the clinical breakdown of the cohort. From there, try a follow-up:

> **Compute PAM50-like subtype scores using key marker genes from bulk RNA-seq and stratify survival accordingly**

This is the core loop in K Pro: ask a question, get a result, refine with a follow-up. Every prompt in this guide works this way.

### Part 1 — How to prompt K Pro effectively

#### The A-S-R-T-C Framework

Every high-performing prompt follows a five-part structure. This framework was built from analysis of 500+ real K-Pro interactions — prompts that follow it succeed consistently, prompts that skip components hit predictable failure modes.

**\[Action] + \[Subject] + \[Resolution] + \[Tool] + \[Context]**

| Component          | What it is                          | Why it matters                                                                    | Keywords & examples                                             |
| ------------------ | ----------------------------------- | --------------------------------------------------------------------------------- | --------------------------------------------------------------- |
| **A — Action**     | The analytical intent               | Prevents the AI from just "describing" — forces it to calculate                   | `Analyze`, `Compare`, `Correlate`, `Perform DEG`, `Generate KM` |
| **S — Subject**    | Gene symbols + aliases (HGNC)       | Avoids "Column Not Found" errors                                                  | `TNFRSF9 (CD137)`, `CD274 (PD-L1)`, `VSIR (VISTA)`              |
| **R — Resolution** | Data granularity                    | Directs the AI to the correct single-cell or bulk layer                           | `Level 3 (refined types)`, `Level 4 (subsets)`, `Bulk RNA`      |
| **T — Tool**       | Visualization type                  | Bypasses tool-specific failure modes (e.g., heatmap crashes on large gene sets)   | `dotplot`, `Violin Plot`, `Kaplan-Meier`, `Volcano plot`        |
| **C — Context**    | Biological or cohort stratification | Filters out irrelevant data (e.g., GBM cells appearing in a bladder cancer query) | `Epithelioid vs. Biphasic`, `TP53 Mut vs. WT`, `TCGA-BRCA`      |

> **Not every prompt needs all five components.** Literature searches typically need only A + S + C. Data exploration needs at least A + S + C. Visualization and spatial queries benefit from the full A-S-R-T-C.

#### From weak to effective — see the difference

Example: Gene expression in mesothelioma

**Weak prompt:**

> "Show me CD137 in mesothelioma."

*What goes wrong:* The AI might pick the wrong data resolution, use a basic bar chart, or fail to find the gene name "CD137" in the database columns.

**Better:**

> "Analyze TNFRSF9 at Level 2 in mesothelioma."

*Improved:* Uses the official gene symbol and specifies resolution. But still lacks the visualization tool and comparison context.

**A-S-R-T-C prompt:**

> "**Analyze** \[A] **TNFRSF9 (CD137)** \[S] at **Single-cell Level 3** \[R] using a **dotplot** \[T] in **Epithelioid vs. Biphasic Mesothelioma** \[C]."

*Why it works:* The agent knows exactly what to calculate, where to find the gene, what resolution to use, which chart to produce, and how to stratify the data.

#### Common failure modes and how to fix them

These are the four most frequent errors observed in real K-Pro sessions. Each one maps to a missing A-S-R-T-C component.

**1. The "Complexity Crash" — too many variables**

|                     |                                                                                                                                                    |
| ------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Failed prompt**   | "Plot expression for AREG, EREG, HBEGF, BTC, TGFA, EGFR, KRAS, BRAF, MET, and HER3 in a heatmap for all bladder cancer types."                     |
| **What went wrong** | Overloaded the heatmap tool with too many genes across too many categories → timeout error                                                         |
| **Fix**             | "**Analyze** the EGFR-ligand and bypass pathway genes at **Bulk RNA level** using a **dotplot** stratified by **Bladder Histological Subgroups**." |
| **Why it works**    | Switching the **Tool** (T) to a dotplot handles high-dimensional data more efficiently than a heatmap                                              |

**2. The "Schema Error" — gene alias not found**

|                     |                                                                                                                                                        |
| ------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------ |
| **Failed prompt**   | "Show me VISTA expression in mesothelioma."                                                                                                            |
| **What went wrong** | The agent couldn't find "VISTA" in column headers → "Column Not Found" error                                                                           |
| **Fix**             | "**Compare** **VSIR (VISTA)** expression at **Single-cell Level 3** using a **Violin Plot** across **Epithelioid and Biphasic mesothelioma subtypes"** |
| **Why it works**    | Using the official HGNC symbol as **Subject** (S) with the alias in parentheses ensures correct mapping                                                |

**3. The "Ambiguity Error" — wrong indication pulled in**

|                     |                                                                                                                    |
| ------------------- | ------------------------------------------------------------------------------------------------------------------ |
| **Failed prompt**   | "What are the top cell types expressing CD137?"                                                                    |
| **What went wrong** | The AI pulled Glioblastoma data (Astrocytes/Microglia) because no cancer type was specified                        |
| **Fix**             | "Analyze TNFRSF9 (CD137) distribution at Single-cell Level 3 specifically within the Bladder Cancer MOSAIC cohort" |
| **Why it works**    | Providing the **Context** (C) filters out non-relevant indications                                                 |

**4. The "Spatial Resolution" failure — too vague for spatial data**

|                     |                                                                                                                           |
| ------------------- | ------------------------------------------------------------------------------------------------------------------------- |
| **Failed prompt**   | "Where is FAP expressed in the slides?"                                                                                   |
| **What went wrong** | Too vague — produced a generic slide view with no quantitative insight                                                    |
| **Fix**             | "Calculate the spatial abundance of FAP within the Invasive Front vs. Tumor Core across ovarian, bladder cancer patients" |
| **Why it works**    | Forces a calculation (A) on specific spatial niches (C) rather than a simple visual search                                |

#### The "Chain of Thought" sequence

Your most productive sessions will follow a logical drill-down path. Don't start with a complex differential expression query — start simple and build up.

| Step            | What to ask                                         | Example                                                                                                            |
| --------------- | --------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------ |
| **1. Survey**   | Broad overview of a gene across indications         | "Show expression of NECTIN4 across all mosaic cancer types"                                                        |
| **2. Focus**    | Zoom into one indication at higher resolution       | "Zoom into MOSAIC-BLCA at Level 3 resolution."                                                                     |
| **3. Spatial**  | Check co-localization in the tumor microenvironment | "Does NECTIN4+ tumor cells co-localize with CD8+ T cells in the invasive front in MOSAIC bladder cancer patients?" |
| **4. Clinical** | Link findings to patient outcomes                   | "Does NECTIN4 overexpression correlate with overall survival in TCGA-BLCA?"                                        |

Each step builds on the context from the previous one — K-Pro maintains conversation history within a session.

#### Quick-reference prompt templates

Copy these templates and fill in the bracketed values.

**Survival Analysis**

```
Generate Kaplan-Meier curves for [OS/PFS/rwPFS] in bladder cancer
patients from TCGA, stratified by FGFR3 expression
(high vs. low using median cutoff).
```

**Differential Expression**

```
Perform a DEG analysis on across immunotherapy responders (complete+partial responders) 
and nonresponders (progressive and stable disease) in MOSAIC NSCLC patients 
excluding unknown/not evaluable patients adjust for [age ,sex]. 
Show top 50 DEGs as a volcano plot with |log2FC| > 1 and padj < 0.05 highlighted.
```

**Spatial Analysis**

```
Visualize FGFR3 and Nectin 4 spatial co-expression in
bladder cancer from MOSAIC, showing correlation within
tumor regions vs. invasive front.
```

**Cell Type Profiling**

```
Compare expression across Level 4 immune cell types in 
PDL1 scRNA-seq data from bladder. 
Display as a dot plot with expression level and percentage of cells expressing.
```

**Target Discovery**

```
Identify differentially expressed genes when 
the primary target TP53 is inhibited in Bladder cancer
```

***

### Part 2 — Prompt Library

Ready-to-use prompts organized by use case. Each prompt has been verified against K-Pro documentation.

#### Literature Review & Biomedical Knowledge

**Summarize publications on a specific gene or pathway**

**Prompt:**

> Summarize the latest PubMed publications on the role of TP53 mutations in non-small cell lung cancer. Focus on findings from the last 3 years and highlight any consensus on prognostic significance.

**Expected result:** A structured summary of key publications with citations, organized by main findings and areas of consensus/controversy.

**Follow-up:**

> Are there any conflicting findings across these studies regarding TP53's role as a predictive biomarker for immunotherapy response in NSCLC?

**Tip:** Be specific about the gene, indication, and timeframe. Vague prompts like "tell me about TP53" return overly broad results.

**Search for publications about a drug target**

**Prompt:**

> Find publications investigating TROP2 as a therapeutic target in triple-negative breast cancer. Include any data on TROP2 expression levels and their correlation with clinical outcomes.

**Expected result:** A curated list of relevant publications with key findings, expression data, and clinical correlations.

**Follow-up:**

> Based on these publications, what is the evidence for using TROP2 expression as a patient selection biomarker for ADC therapies?

**Tip:** Combine the target name with a specific indication and the type of evidence you need (expression, outcomes, mechanisms).

#### Data Exploration & Cohort Analysis

**Explore TCGA cohort demographics**

**Prompt:**

> Provide clinical characteristics of the TCGA-BRCA cohort. Include the distribution of tumor stage and age at diagnosis.

**Expected result:** Summary tables and/or distributions showing the clinical breakdown of the TCGA-BRCA cohort.

**Tip:** Always specify the exact TCGA cohort code (e.g., TCGA-BRCA, TCGA-LUAD) rather than the disease name alone.

**A-S-R-T-C breakdown:** Show \[A] clinical characteristics \[S: cohort variables] at bulk level \[R] as summary tables \[T] in TCGA-BRCA \[C].

**Explore MOSAIC Window multiomics data**

**Prompt:**

> Provide the number of patients that have access to all the available (RNA-seq, Single cell, spatial, H\&E, clinical) modalities for MOSAIC bladder cancer.

**Expected result:** A summary of MOSAIC Window data availability for the requested indication, broken down by modality.

**Follow-up:**

> For ovarian cancer in mosaic window, show the distribution of HRD status." There is no CRC samples in mosaic window.

**Tip:** MOSAIC Window is Owkin's proprietary multimodal dataset. Specify the indication to scope the data landscape before diving into analysis.

**Characterize a patient cohort**

**Prompt:**

> Characterize the TCGA-LUAD cohort: show me the distribution of KRAS mutation status, smoking history, tumor stage, and median overall survival for KRAS wt and mutated patients.

**Expected result:** A multi-variable cohort characterization with summary statistics per subgroup.

**Follow-up:**

> Which of these subgroups has the worst overall survival, and what are their distinguishing molecular features?

**Tip:** List the specific variables you want characterized upfront — K-Pro works best when it knows exactly what you're looking for.

**Assess gene expression across tissues (target prioritization)**

**Prompt:**

> Compare the expression of NECTIN4 across all TCGA cancer types. Show me a pan-cancer overview ranked by median expression level.

**Expected result:** A ranked table or visualization of NECTIN4 expression across TCGA indications.

**Follow-up:**

> For the top 3 indications with highest NECTIN4 overexpression, show me the correlation between NECTIN4 expression and overall survival.

**A-S-R-T-C breakdown:** Compare \[A] NECTIN4 \[S] at Bulk RNA level \[R] as a ranked table \[T] across all TCGA cancer types vs. matched normal tissue \[C].

**Multiomics integration**

**Prompt:**

> In the TCGA-BRCA cohort, integrate RNA-seq gene expression and mutation data. Identify genes that are both differentially expressed and frequently mutated in basal-like versus luminal A patients.

**Expected result:** A multi-omics integration result showing genes that appear significant across both data modalities.

**Follow-up:**

> Perform an individual gene expression comparison of MSIGDB DNA repair, oxidative stress, and EMT pathway signatures at Bulk RNA level using a violin plot with p-values comparing Basal-like and Luminal TCGA-BRCA patients.

**Tip:** Multiomics queries are computationally intensive. Start with a specific comparison (two subgroups) rather than "analyze everything."

#### Visualization

**Agent:** Data Explorer (with visualization capabilities)

**Generate a Kaplan-Meier survival curve**

**Prompt:**

> Create a Kaplan-Meier survival curve for TCGA-BRCA patients stratified by TP53 mutation status (mutated vs. wild-type). Include the p-value from a log-rank test and the number of patients at risk.

**Expected result:** A Kaplan-Meier plot with two curves, log-rank p-value, and at-risk table.

**Tip:** K-Pro supports real-time plot iteration. Ask for specific formatting changes (colors, labels, font size) in follow-up prompts rather than trying to specify everything at once.

**A-S-R-T-C breakdown:** Create \[A] TP53 mutation survival analysis \[S] at bulk level \[R] as a KM curve with log-rank test \[T] in TCGA-BRCA, mutated vs. wild-type \[C].

**Create comparative visualizations**

**Prompt:**

> Create a box plot comparing the expression of CD274 (PD-L1) across the five molecular subtypes in TCGA-BRCA. Add individual data points and mark statistically significant pairwise comparisons.

**Expected result:** A box plot with overlaid data points and significance brackets between groups.

**Follow-up:**

> Now create a heatmap of the top 50 differentially expressed genes between these molecular subtypes.

**Tip:** You can request multiple plot types in sequence. Each builds on the data context from the previous query.

#### Statistical Analysis

**Compute survival statistics**

**Prompt:**

> In the TCGA-STAD cohort, test whether there is a statistically significant difference in overall survival between microsatellite-instable (MSI-H) and microsatellite-stable (MSS) patients. Report the hazard ratio, 95% confidence interval, and log-rank p-value.

**Expected result:** Statistical test results with HR, CI, and p-value, plus a supporting KM curve.

**Follow-up:**

> Run a multivariate Cox regression adjusting for age, stage, and MSI status to confirm whether MSI status is an independent prognostic factor.

**Tip:** Always specify the statistical test you want (log-rank, Cox, t-test) for precise results. K-Pro backs every analysis with p-values and population-level data.

#### Cross-Dataset Discovery

**Link datasets to find novel biological mechanisms**

**Prompt:**

> In the TCGA-KIRC cohort, identify genes whose expression is significantly associated with both VHL mutation status and response to immune checkpoint inhibitors. Cross-reference findings with published literature on VHL-related immune evasion mechanisms.

**Expected result:** A list of candidate genes with statistical associations, linked to supporting literature evidence.

**Follow-up:**

> For the top 3 candidate genes, show me their spatial expression patterns in the tumor microenvironment using MOSAIC data if available.

**Tip:** This is a multi-step, multi-agent use case. Be explicit about both the data analysis you want AND the literature cross-reference.

### Do's and don'ts

#### Do

* **Use official gene symbols** with common aliases in parentheses: `VSIR (VISTA)`, not just `VISTA`
* **Specify the dataset** using the cohort name and the indication.
* **Name the statistical test** you want: log-rank, Cox regression, t-test
* **Iterate in follow-ups** — refine plots, add filters, drill deeper
* **Start simple, then drill down** — follow the Survey → Focus → Spatial → Clinical sequence
* **Specify the visualization type** — dotplot, violin, KM curve — to avoid tool-selection errors

#### Don't

* **Don't write "search engine" prompts** — "Tell me about TP53" is too vague
* **Don't overload a single prompt** — stacking multiple plots or analyses into one request is more likely to fail than running a complex one. Break the work into sequential steps, the way you would during a scientific deep-dive.&#x20;
* **Don't skip the indication/cohort** — without context, the AI may pull irrelevant data from other cancer types
* **Don't try to specify everything in one prompt** — ask for the analysis first, then adjust formatting in follow-ups


# K Pro for teams

K Pro is an AI scientist for biopharma R\&D. It works across multimodal data so your teams can move from question to evidence faster, without writing code.

How you'll use K Pro depends on your role. Find your team below to see the workflows it supports and the kinds of decisions it helps with.

***

### Scientific & Translational teams

**This is for you if** you work in translational medicine, biomarker sciences, precision medicine, computational biology, or discovery biology, and you sit where early discovery meets clinical development.

**What you're working toward**

* Discovering and validating biomarkers
* Prioritizing targets
* Connecting molecular biology to clinical outcomes

**Where K Pro helps**

* Test biomarker hypotheses across multiple modalities in hours instead of weeks
* Work with curated multimodal data, including MOSAIC real patient data, without spending weeks getting it analysis-ready
* Run cross-modality analysis in one place instead of stitching tools together, for example a spatial transcriptomics plot in about a minute rather than three days
* Review the literature at scale
* Produce publication-ready evidence packages, with outputs traceable back to their source

**Challenges this addresses**

* Exploring a single biomarker hypothesis across several modalities can take one to two months
* Data is siloed and slow to bring to the right quality
* Workflows are split across fragmented tools with no shared view
* Literature review at scale is a bottleneck
* Turning analysis into a clear evidence package adds days of effort

***

### Clinical Development & Medical teams

**This is for you if** you lead clinical development or clinical science, work in medical affairs, or own clinical biomarker strategy, sitting between research and the clinic.

**What you're working toward**

* Increasing probability of success
* Accelerating biomarker strategy
* De-risking late-stage decisions

**Where K Pro helps**

* Build the multimodal patient evidence to define the right population before a trial starts, anchored on MOSAIC real patient data
* Run responder analysis and multimodal feature discovery to understand who responds and why
* Connect molecular signals to clinical outcomes without long bioinformatics turnaround
* Assemble evidence packages for internal committees faster
* Keep outputs interpretable and traceable to source data

K Pro informs and accelerates clinical strategy. It does not replace regulatory validation.

**Challenges this addresses**

* Defining the right patient population often relies on incomplete biological evidence
* Linking genomic features to response data can take months
* Responder analysis is a bottleneck and expensive to outsource
* Early go/no-go calls can feel like judgment calls without a real-world patient baseline
* Evidence packages for committees take too long to assemble

***

### Strategy, BD & Portfolio teams

**This is for you if** you work in portfolio strategy, external innovation, business development, or corporate strategy, bridging scientific teams and executive leadership.

**What you're working toward**

* Maximizing portfolio value by prioritizing the right assets
* Sizing the opportunity and validating clinical viability
* Building a differentiated, defensible position

**Where K Pro helps**

* Run competitive landscaping, indication expansion, and asset prioritization in one workflow
* Turn weeks of manual diligence into structured, biology-grounded analysis
* Surface white space by combining scientific, clinical, and competitive signals
* Build investment cases faster, with outputs that are traceable, auditable, and ready to present to leadership

**Challenges this addresses**

* Asset diligence is slow and manual, stitching together fragmented sources
* True white space is hard to find when signals live in different places
* Portfolio prioritization is inconsistent across teams using different assumptions and datasets
* Competitive intelligence is reactive rather than continuous
* Building investment cases takes weeks

AstraZeneca has licensed K Pro under a three-year agreement to build biopharma agents for competitive intelligence across pharmaceutical targets, assets, and trials.

***

### Data, IT & Digital teams

**This is for you if** you lead data science, bioinformatics, AI/ML, or sit in data and digital leadership, owning how new platforms get integrated and governed.

**What you're working toward**

* Reducing data friction and time to insight
* Deploying AI securely at enterprise scale
* Enabling teams to run analyses autonomously

**Where K Pro helps**

* Deploy securely within your existing infrastructure
* Integrate with your data stack as an intelligence layer rather than a replacement
* Cut data onboarding from weeks to hours
* Extend the platform with your own data (BYOD)
* Keep work reproducible and auditable, with transparency and control over system behavior

K Pro integrates within enterprise IT infrastructure and decision workflows, as in the three-year AstraZeneca licensing agreement.

**Challenges this addresses**

* Data preparation is the main bottleneck before any analysis can start
* Every new tool adds integration, authentication, and governance work
* Data science teams are overloaded with repetitive data requests
* A lack of standardization across tools and data versions limits reproducibility and scale
* Valuable internal datasets stay siloed and underused

***

### Executive Leadership

**This is for you if** you sit at the top of the organization (CEO, CSO, CMO, EVP/SVP R\&D, CDO) with final authority over R\&D and AI strategy.

**What you're working toward**

* Maximizing portfolio value and R\&D productivity
* Accelerating time to market
* De-risking pipeline decisions to reduce late-stage failure

**Where K Pro helps**

* Turn fragmented scientific data into faster, higher-confidence decisions, for example up to 30x faster target prioritization
* Surface signal earlier to support go/no-go calls
* Show enterprise-wide impact rather than isolated use cases
* Give leadership a clearer view of where data and AI are improving speed, probability of success, and cost efficiency

**Challenges this addresses**

* R\&D productivity is under pressure from rising costs and high clinical attrition
* High-stakes portfolio decisions are made under persistent uncertainty
* Pipeline value is lost through late-stage failures that better signal could have flagged earlier
* AI and digital investments are fragmented and don't always translate into measurable impact
* There's no consolidated view of where AI is materially moving the needle


# What's included in your plan

<br>


# How K Pro works

When you ask K Pro a question, it routes your request to the right agent, runs the analysis, and returns an interactive result. This page covers the three ways you can ask, the agents available to handle your request, and the role of skills.

### Three ways to ask K Pro a question

K Pro is designed as a no-code platform for wet-lab scientists and disease-area experts. It supports three modes of interaction along a spectrum of control, and a scientist moves between them as the task requires.

#### Co-pilot: Conversational exploration

**Ask. Explore. Iterate.**

Ask a question in natural language. Get back an interactive artifact you can click into, drill down, and build on.

Example: The scientist opens K Pro and asks a question in plain English: "Show me NECTIN4 expression across malignant cell types in MOSAIC bladder cancer cohorts, and compare to GTEx healthy tissues." K Pro returns an interactive cell-type expression plot, a healthy-tissue comparison heatmap, and a narrative summary linking to the underlying data slices.

The scientist can click any data point to see the patient-level data and the statistical test that produced the value. A plain-English follow-up ("Now show the spatial distribution of NECTIN4-expressing tumor cells relative to CD8+ T-cells") triggers the next analysis automatically.

<figure><img src="/files/u7XZZspWTMoOMUkGLe5s" alt=""><figcaption></figcaption></figure>

#### Semi-autonomous: Skilled workflows

**Run expert workflows on demand.**

Trigger a skill (Owkin-certified or custom) for complex multi-step analyses where you want a constrained, reproducible workflow rather than free-form conversation.

Example: The scientist knows what analytical workflow they want and invokes a pre-built skill directly with a slash command: e.g. `/target-characterization NECTIN4 in BLCA` or `/drug-positioning <asset> in <indication>`. K Pro runs the full multi-step workflow end-to-end and returns a structured report. Skills come from three sources:

* Owkin-certified workflows built from a decade of pharma engagements
* KOL-derived skills encoding leading domain experts' methodologies
* Customer-authored skills built by client teams

#### Broad autonomous campaigns

**Scale across hypotheses.**

Launch multiple hypotheses in parallel through dedicated sandboxed agents — useful when you want to evaluate several strategies side by side.

Example: The scientist launches multiple hypotheses in parallel through sandboxed agents, each running the same skill or different skills against different parameters. Each agent produces its own report; results are aggregated for comparison. This is where K Pro scales beyond a single analysis to portfolio-level exploration.

In none of these modes does the scientist write code or specify which datasets to query. K Pro makes those decisions transparently and the scientist can audit any of them.

<figure><img src="/files/zBCBRfR7SHI95TarMHne" alt=""><figcaption></figcaption></figure>

### What's behind K Pro

K Pro draws on a 104-hospital patient data network covering latest standard of care, with **1.2M patients** characterized across multimodal data (bulk, single-cell, spatial RNA, pathology, longitudinal follow-ups). Indications covered include oncology, immunology & inflammation, cardiometabolic, and neuro.

### Agents & tools available in K Pro Free

Each question is routed to the agent or tool best suited to handle it. The agents below are available in K Pro Free.

| **Agent / tool**         | **What it does**                                                                                                                                             |
| ------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| Multi-Omics              | Analyzes multi-omics data from MOSAIC and TCGA datasets. Handles gene expression, signatures, survival analysis, and visualization.                          |
| Consensus                | AI-powered scientific search engine for finding relevant papers, literature-backed hypothesis validation, and scientific background research.                |
| Literature Review        | Retrieves scientific documents using semantic search over 30M+ PubMed abstracts. Answers scientific questions with cited sources.                            |
| Prior Knowledge          | Queries gene-level features from public databases — protein families, localization, essentiality, oncogenicity, safety/toxicity, target tractability scores. |
| Discoverability          | Reference information about Owkin, MOSAIC, and TCGA datasets — useful for onboarding and orienting before analysis.                                          |
| Pathology Explorer (MCP) | Pathology AI capabilities: TCGA slide-level histomics, cell-type filtering, slide visualization, survival analysis on cell-type features.                    |
| Plotter                  | Generates plots/charts from your data, with schema inspection and SQL validation.                                                                            |
| Code Execution           | Runs custom analyses on customer data in a secure, isolated workspace. Upload files to process your own data and answer questions beyond the built-in tools. |

### What is a skill?

A **skill** is a pre-packaged set of instructions and tools designed for a high-stakes task. When you trigger a skill, K Pro switches into a specific mode where it applies the right filters, checks the right thresholds, and returns a defined output format. Skills make analyses reproducible: same input, same kind of output, every time.

Three types of skills are available:

* **Owkin-certified skills:** validated scientific workflows built from Owkin projects.
* **KOL-derived skills**: domain experts' methods and reasoning encoded as executable workflows.
* **Your own skills**: your internal best practices, layered on top of the existing library.

###


# Target identification & indication discovery

Use K Pro to identify and prioritize therapeutic targets, and to characterize the patient subgroups in which they're most relevant.

* **Prioritize targets**: rank therapeutic targets and gene candidates using statistical evidence from TCGA and MOSAIC datasets.
* **Characterize patient subgroups**: identify distinct populations based on multi-omics profiles and clinical outcomes.

The pages below walk through each task with a tested example prompt.

{% embed url="<https://vimeo.com/1183008004?fe=ci&fl=sv&share=copy>" %}

Here is a concrete worked example of how K Pro navigates target prioritization for Antibody-Drug Conjugates (ADCs) or, analogously, targeted alpha therapies (TATs). A similar workflow is used to identify radio-ligand therapies (RLTs), except that protein localization on the cell membrane is not required.

**1. User request.** A researcher prompts K Pro: *"Identify and prioritize top ADC targets for bladder cancer (BLCA)."*

**2. K Pro reasons about the constraints.** K Pro translates the query into a list of biological constraints and respective analyses. For a viable ADC target, the system knows the protein must be expressed on the cell surface, be highly prevalent in BLCA tumors, and have minimal (optimally, zero) expression in essential healthy tissues to ensure a safe therapeutic window. It also reasons that expression heterogeneity might drive resistance, and hence good targets are expressed in a high proportion of cells/regions of each tumor, and that non-expressing regions should optimally be impacted through the bystander effect.

**3. K Pro analyses bulk + single-cell + spatial transcriptomics.** The system analyzes bulk transcriptomics (TCGA, MOSAIC) as well as single-cell and spatial transcriptomics (MOSAIC) data, scoring each target according to the above criteria. Targets score highly if they are expressed highly in tumor and not expressed in normal tissue (using GTEx and single-cell normal tissue atlases), and expressed either ubiquitously within the tumor (all single cells / spatial regions) or expressed in a pattern that allows the bystander effect to kill nearby tumor cells that do not express the target (based on a bystander score feature Owkin has developed).

**4. K Pro outputs targets ranked by quality.** The system provides a series of both textual and visual (plots) outputs that provide target ranking/recommendations and per-target evidence for target quality.

**5. Optional competitive intelligence.** If opted & prompted by the user, the system can provide recommendations rooted not only in biological data but also on competitive landscape analysis. The system conducts in-depth competitive intelligence analysis (including trial data, patent data, news feeds, conference abstracts, etc.) through its competitive intelligence agentic capabilities.

Since this example is about expression-based targets (ADCs, TATs), some features typically associated with small molecules like genetic association, dependency/addiction, biochemical pathways, for example, are not relevant and hence the system correctly does not apply them.

At the end, the user can ask to download the analyses as a comprehensive report. Alternatively, the sequence of investigation can be bundled in a skill that helps reproduce the exact same flow without repeating all the sequential prompts.


# Explore a cohort

**What to do:** Start by asking K Pro about a cohort's composition to understand what's available before diving into specifics. Always specify the exact TCGA cohort code (e.g., TCGA-BRCA, TCGA-LUAD) rather than the disease name alone.

**Which agent/skill:** MultiOmics Agent

**Example prompt:**

> Provide clinical characteristics of the TCGA-BRCA cohort. Include the distribution of tumor stage and age at diagnosis.

<figure><img src="/files/AEpFQre7wO7f5fVWHqia" alt=""><figcaption></figcaption></figure>

**Expected result:** Summary tables and/or distributions showing the clinical breakdown of the TCGA-BRCA cohort.

[**A-S-R-T-C breakdown**](https://docs.owkin.com/getting-started/prompting-guide-and-prompt-library)**:** Show \[A] clinical characteristics \[S: cohort variables] at bulk level \[R] as summary tables \[T] in TCGA-BRCA \[C].

<br>


# Compare groups

**What to do:** Define two patient subgroups and ask K Pro to compare them across a specific data type. Always specify the statistical test you want (log-rank, Cox, t-test) for precise results.

**Example prompt:**

> In the TCGA-STAD cohort, test whether there is a statistically significant difference in overall survival between microsatellite-instable (MSI-H) and microsatellite-stable (MSS) patients. Report the hazard ratio, 95% confidence interval, and log-rank p-value.

**Expected result:** Statistical test results with HR, CI, and p-value, plus a supporting KM curve.

**Follow-up:**

> Run a multivariate Cox regression adjusting for age, stage, and MSI status to confirm whether MSI status is an independent prognostic factor.


# Rank candidate targets

**What to do:** Compare a candidate gene's expression across cancer types to identify indications where it's overexpressed. Specify "compared to matched normal tissue" for more meaningful results.

**Example prompt:**

> Compare the expression of NECTIN4 across all TCGA cancer types. Show me a pan-cancer overview ranked by median expression level, and highlight which indications show overexpression compared to matched normal tissue.

<figure><img src="/files/eDqCuFhu3xHLfOUQv55F" alt=""><figcaption></figcaption></figure>

**Expected result:** A ranked table or visualization of NECTIN4 expression across TCGA indications, with tumor vs. normal comparisons.

**Follow-up:**

> For the top 3 indications with highest NECTIN4 overexpression, show me the correlation between NECTIN4 expression and overall survival.

<figure><img src="/files/CRrQZfFYPlsMaTsx7Wmj" alt=""><figcaption></figcaption></figure>

[**A-S-R-T-C breakdown**](https://docs.owkin.com/getting-started/prompting-guide-and-prompt-library)**:** Compare \[A] NECTIN4 \[S] at Bulk RNA level \[R] as a ranked table \[T] across all TCGA cancer types vs. matched normal tissue \[C].


# Indication prioritization & target validation

Use K Pro to assess whether a target is worth advancing, and to validate that biomarkers behave consistently across data modalities.

* **Assess target tractability & druggability**: evaluate target tractability, safety profiles, and therapeutic potential using curated databases.
* **Validate biomarkers across modalities**: test biomarkers and hypotheses across bulk RNA-seq, single-cell, spatial, and clinical data.

**What to do:** Combine multiple data modalities (RNA-seq, mutation data, etc.) within a cohort to identify candidates that show convergent evidence across layers. Multiomics queries are computationally intensive: start with a specific comparison (two subgroups) rather than "analyze everything."

**Example prompt:**

> In the TCGA-BRCA cohort, integrate RNA-seq gene expression and mutation data. Identify genes that are both differentially expressed and frequently mutated in basal-like versus luminal A patients.

**Expected result:** A multi-omics integration result showing genes that appear significant across both data modalities.

**Follow-up:**

> Perform an individual gene expression comparison of MSIGDB DNA repair, oxidative stress, and EMT pathway signatures at Bulk RNA level using a violin plot with p-values comparing Basal-like and Luminal TCGA-BRCA patients.

{% embed url="<https://vimeo.com/1183008489?amp;fe=ci&amp;fl=sv&share=copy>" %}


# Drug positioning & treatment landscape

Use K Pro to position your asset in its competitive landscape and connect biology to clinical outcomes.

* **Connect biology to clinical outcomes**: integrate genomic, transcriptomic, and immune data to identify therapeutic opportunities and predict treatment response.

The pages below cover competitive positioning and clinical-trial navigation.


# Competitive positioning

**What to do:** Ask K Pro to summarize the competitive landscape for a target or therapeutic class, drawing from recent conference abstracts and trial reports.

**Which agent/skill:** Knowledge Agent

**Example prompt:**

> Summarize the competitive positioning of TIGIT inhibitors based on recent ESMO/ASCO abstracts.


# Navigate clinical trials

## What is Trial Navigator?

Trial Navigator is a specialized agent in Owkin K that helps pharmaceutical and clinical teams explore the clinical trial landscape using plain-language questions. You ask the way you'd ask a colleague; it interprets the intent, filters a large clinical-trials dataset, and returns the right output: a visual timeline, a data table, a statistical projection, or a direct answer.

Use it to:

* **Map competitive landscapes** across an indication, mechanism of action, or drug class
* **Track competitor pipelines** and development timelines
* **Answer quantitative questions** about the trial landscape (counts, averages, distributions)
* **Compare trials** against a reference trial (e.g. "PFS better than trial X")
* **Deep-dive a single trial's** design assumptions and projected readout times

## How it works

You don't need to know the internals to use Trial Navigator, but a quick mental model helps you write better queries. Behind the scenes it:

{% stepper %}
{% step %}

### Plans

Breaks a complex question into sub-questions and, when needed, runs them in sequence, so "find trials better than X" first finds X, then compares the rest of the landscape against it.
{% endstep %}

{% step %}

### Translates

Turns your words into structured filters: indication, phase, sponsor, dates, mechanism of action, enrollment thresholds, and more.
{% endstep %}

{% step %}

### Reconciles terminology

Maps your wording to the dataset's standard names via a built-in thesaurus, so "immunotherapy", "IO", and "PD-1 inhibitor" resolve to the right underlying terms.
{% endstep %}

{% step %}

### Answers

Returns a Gantt chart and table, a statistical projection, and/or a written summary, depending on what you asked for.
{% endstep %}
{% endstepper %}

{% hint style="info" %}
You can be conversational and use common abbreviations, the tool is designed to resolve them. The more specific you are, the sharper the result.
{% endhint %}

## See it in action

{% embed url="<https://vimeo.com/1183008489>" %}

## Writing effective queries

* Name the **therapeutic area / indication** (e.g. NSCLC, TNBC, glioblastoma).
* Name the **mechanism of action or drug class** when relevant (PD-1/PD-L1, EGFR TKI, ADC, VEGFR, CLDN18.2).
* Name **sponsors, drugs, or NCT IDs** when you have them.
* Add a **phase** to narrow the development stage (Phase I, II, III, I/II…).
* Add a **timeframe** ("started after 2020", "between 2018 and 2023").
* Add **thresholds** ("enrollment greater than 500", "accrual target of exactly 500").
* Use **exclusions** freely ("excluding Phase I", "excluding breast and prostate cancer").
* Ask for **top / latest / longest** when you want a ranked shortlist ("the 5 most recent…", "longest-running…").

## What you can ask

### Competitive landscape / list trials

Returns an interactive Gantt chart, a detailed results table, and summary statistics.

```
What are the clinical trials for immunotherapy in lung cancer?
Show me immunotherapy trials for breast cancer started in 2020 or later.
Find PD-1 or PD-L1 antagonist trials for colorectal cancer
Find EGFR tyrosine kinase inhibitor trials for lung cancer in Europe
Show me all Phase 1/2 studies for CLDN18.2 ADCs
Show me gene therapy trials for glioblastoma started after 2018
```

### Complex multi-filter queries

The tool handles several constraints at once, combining indication, phase, geography, dates, mechanism, and thresholds.

```
Find Phase II and III immuno-oncology trials for breast or lung cancer in
Europe or North America, started between 2018-2023, with enrollment at least
500 patients
Show me VEGFR tyrosine kinase inhibitor trials for prostate or colorectal
cancer in Asia, started between 2010-2020, excluding Phase I trials
Show me all Phase 3 Merck & Co. studies in NSCLC evaluating ADCs
Chart all Roche P3 studies for atezolizumab in TNBC that have published data
```

### Ranked / "top N" / "latest N" / "longest-running"

Ask for a ranked shortlist instead of the full set.

```
Show me the 5 most recent AstraZeneca-sponsored clinical trials
What are the longest-running clinical trials for melanoma?
Show me recent small trials with high objective response rate
```

### Aggregate & quantitative questions

For counts, averages, percentages, and "what's covered" questions, the tool answers directly in text, with supporting figures where useful.

```
How many clinical trials are there for lung cancer started between 2015 and 2024?
How many trials are sponsored by Roche that started in 2020?
What is the average accrual target for Phase III trials?
What percentage of ADC trials are Phase III?
What is the most common indication for HER2-targeting ADC clinical trials?
```

### Comparative queries (against a reference trial)

The tool can find a reference trial first, then compare the rest of the landscape against it.

```
Find trials with progression-free survival greater than trial NCT00490139
Find trials with overall survival greater than trial NCT00490139 for breast cancer
```

### Single-trial statistical analysis

Deep-dive the statistical assumptions behind a specific trial's design and projected readout times. Returns a projection plot, predicted median OS/PFS, predicted enrollment duration, and the reference trials used for the prediction.

```
Show me the details for trial NCT00490139
Perform a statistical analysis for trial NCT00338247
What are the underlying assumptions for read out times of trial NCT00490139?
```

## Understanding the outputs

### Gantt chart (timeline)

* **X-axis:** time (years)
* **Y-axis:** individual trials, grouped by phase and sorted by start date
* **Colour:** trial phase; bars show duration with mechanism-of-action abbreviations
* Includes both industry-sponsored and publicly/academically sponsored trials.

<figure><img src="https://1034377358-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FTmCjUj6W0sltjW7q2vBp%2Fuploads%2F80eNT1X3EnqRlsCuSm5G%2Fimage.png?alt=media&#x26;token=4e78405e-82da-4277-91da-8bc459f28f58" alt=""><figcaption></figcaption></figure>

### Results table

A sortable, filterable list of every matching trial: NCT ID (linked to [ClinicalTrials.gov](http://clinicaltrials.gov)), sponsor, phase, indication, drug/mechanism, and key results.

<figure><img src="https://1034377358-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FTmCjUj6W0sltjW7q2vBp%2Fuploads%2F9Thpl1bYCG2EIIxPib8x%2Fimage.png?alt=media&#x26;token=3ceebee3-18a4-4617-8076-a9c01333606b" alt=""><figcaption></figcaption></figure>

### Summary statistics

An automatically computed answer summary, for example total trial count, phase distribution, and top indications represented.

<figure><img src="https://1034377358-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FTmCjUj6W0sltjW7q2vBp%2Fuploads%2Fr3kvNEV29PzeSw9LkSSx%2Fimage.png?alt=media&#x26;token=147b83b9-594a-4ca1-9ceb-a775863cd739" alt=""><figcaption></figcaption></figure>

### Statistical projection plot

For single-trial analysis: an enrollment curve and survival projections, with the reference trials (by NCT ID) used to generate the predictions.

<figure><img src="https://1034377358-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FTmCjUj6W0sltjW7q2vBp%2Fuploads%2FEYk00s1kgXQtfgnhme0o%2Fimage.png?alt=media&#x26;token=d22534ba-e94d-4856-9706-83b4e7e9a72c" alt=""><figcaption></figcaption></figure>

### PowerPoint export

Trial Navigator can export a competitive-landscape view into a PowerPoint slide.

{% hint style="info" %}
PowerPoint export is a newer, deployment-gated capability, availability depends on your K environment.
{% endhint %}

## Searchable data fields

Trial Navigator is built on a comprehensive **oncology** clinical-trials dataset sourced from ClinicalTrials.gov. You can filter and search across the following dimensions.

### Identification

| Field           | Description                                                |
| --------------- | ---------------------------------------------------------- |
| NCT ID          | [ClinicalTrials.gov](http://clinicaltrials.gov) identifier |
| Brief Title     | Short trial title                                          |
| Official Title  | Full formal trial title                                    |
| Commercial Name | Sponsor's commercial trial name                            |

### Status & timeline

| Field                   | Description                                                    |
| ----------------------- | -------------------------------------------------------------- |
| Overall Status          | Recruiting, Active Not Recruiting, Completed, Terminated, etc. |
| Trial Start Date        | Date the trial began                                           |
| Primary Completion Date | Date of primary outcome measurement                            |
| Completion Date         | Full study completion date                                     |
| Trial Duration          | Duration in days                                               |

### Design

| Field          | Description                                               |
| -------------- | --------------------------------------------------------- |
| Study Type     | Interventional, Observational                             |
| Trial Phase    | I, II, III, IV, I/II, II/III, III/IV                      |
| Randomization  | Whether the trial is randomized                           |
| Number of Arms | Count of treatment arms                                   |
| Group Types    | Experimental, Active Comparator, Placebo Comparator, etc. |

### Sponsors

| Field                       | Description                                                             |
| --------------------------- | ----------------------------------------------------------------------- |
| Private / Industry Sponsors | Pharmaceutical and biotech companies (Pfizer, Roche, AstraZeneca, etc.) |
| Public / Academic Sponsors  | Universities, hospitals, research centers, government agencies          |
| Industry Sponsored          | Boolean flag for industry involvement                                   |

### Enrollment & accrual

| Field               | Description                        |
| ------------------- | ---------------------------------- |
| Enrollment          | Number of participants             |
| Accrual Target      | Target number of subjects          |
| Accrual Actual      | Actual enrolled subjects           |
| Enrollment Duration | Time to complete enrollment (days) |

### Geography

| Field           | Description                                         |
| --------------- | --------------------------------------------------- |
| Trial Countries | Countries where the trial is conducted              |
| Trial Regions   | Regions (North America, Western Europe, Asia, etc.) |
| Number of Sites | Count of trial sites                                |

### Eligibility criteria

| Field            | Description                                                                 |
| ---------------- | --------------------------------------------------------------------------- |
| Gender           | Both, Female Only, Male Only                                                |
| Age Range        | Minimum and maximum age for eligibility                                     |
| ECOG Status      | Performance status requirements (0-4 scale)                                 |
| Disease Stages   | Localized, locally advanced, advanced, metastatic, resectable, unresectable |
| Disease Paths    | Persistent, recurrent, refractory, relapsed, resistant                      |
| Lines of Therapy | Prior treatment requirements                                                |

### Indications (cancer types)

Trial Navigator covers a comprehensive oncology landscape, including:

**Solid tumors:**

* Breast, Lung (NSCLC, SCLC), Colorectal, Pancreas
* Melanoma, Renal, Bladder, Prostate
* Head/Neck, Esophageal, Gastric, Liver
* Ovarian, Endometrial, Cervical
* CNS tumors (Glioblastoma, Astrocytoma, Meningioma, etc.)
* And many more

**Hematological malignancies:**

* Leukemia (ALL, AML, CML, CLL)
* Lymphoma (Hodgkin's, Non-Hodgkin's)
* Multiple Myeloma
* Myelodysplastic Syndrome
* Myeloproliferative Neoplasms

### Treatment & intervention

| Field                | Description                                                                    |
| -------------------- | ------------------------------------------------------------------------------ |
| Drug Names           | Investigational drug names                                                     |
| Drug Type            | ADC, Antibody Bispecific, Antibody Multispecific, Gene Therapy, Small Molecule |
| Mechanisms of Action | JAK inhibitor, CDK inhibitor, PD-1/PD-L1 inhibitor, etc.                       |
| Targets (HGNC)       | Gene symbols (KRAS, TP53, BRCA1, EGFR, ERBB2, etc.)                            |
| Biomarkers           | Biomarkers used in trial selection                                             |
| Combination Therapy  | Whether the trial tests a drug combination                                     |

### Endpoints

| Field             | Description                                        |
| ----------------- | -------------------------------------------------- |
| Primary Endpoints | OS, PFS, DUAL, or other                            |
| Has OS            | Overall Survival as endpoint                       |
| Has PFS           | Progression-Free Survival as endpoint              |
| Has EFS/DFS/RFS   | Event-free, Disease-free, Recurrence-free Survival |

### Outcomes (actual & predicted)

| Field                         | Description                              |
| ----------------------------- | ---------------------------------------- |
| Objective Response Rate       | ORR from trial results                   |
| Disease Control Rate          | DCR from trial results                   |
| Median OS (months)            | Actual median overall survival           |
| Median PFS (months)           | Actual median progression-free survival  |
| Predicted Median OS           | AI-predicted median OS                   |
| Predicted Median PFS          | AI-predicted median PFS                  |
| Predicted Enrollment Duration | AI-predicted time to complete enrollment |

## Tips & good-to-knows

{% hint style="info" %}

* **Be conversational, but specific.** Abbreviations and synonyms are resolved by the thesaurus; specificity narrows results.
* **Smart name resolution.** The tool understands common synonyms and variations, for example "Keytruda" maps to pembrolizumab, "Herceptin" to trastuzumab, gene-name variants (ERBB2 / HER2 / neu), and sponsor name variations (Merck & Co., MSD, Merck Sharp & Dohme).
* **Large landscapes are shown, not silently truncated.** The chart focuses on the most relevant trials for readability while the underlying query covers the full matching set. If a very large landscape is capped for display, the tool tells you.
* **Statistical projections are model-based estimates**, derived from historically comparable reference trials. Treat them as directional, not definitive, and note they are available only for Phase II and III trials with sufficient historical data.
* **Coverage is oncology-focused.** Non-oncology questions may return limited or no results.
* **Data reflects ClinicalTrials.gov**, so completeness depends on what sponsors have registered and published. For real-time status, verify directly on [ClinicalTrials.gov](http://clinicaltrials.gov).
  {% endhint %}


# Literature review


# Review & synthetize literature

**What to do:** Be specific about the gene, indication, and timeframe. Vague prompts like "tell me about TP53" return overly broad results.

**Which agent/skill:** Consensus (AI-powered scientific search engine)

**Example prompt:**

> Summarize the latest PubMed publications on the role of TP53 mutations in non-small cell lung cancer. Focus on findings from the last 3 years and highlight any consensus on prognostic significance.

<figure><img src="/files/E62ce2btkz4S77xrDK9j" alt=""><figcaption></figcaption></figure>

**Expected result:** A structured summary of key publications with citations, organized by main findings and areas of consensus/controversy.

**Follow-up:**

> Are there any conflicting findings across these studies regarding TP53's role as a predictive biomarker for immunotherapy response in NSCLC?

**Second example — searching for publications about a drug target:**

> Find publications investigating TROP2 as a therapeutic target in triple-negative breast cancer. Include any data on TROP2 expression levels and their correlation with clinical outcomes.

*Expected result:* A curated list of relevant publications with key findings, expression data, and clinical correlations.

*Follow-up:*

> Based on these publications, what is the evidence for using TROP2 expression as a patient selection biomarker for ADC therapies?


# Explore the patent landscape

## Overview

&#x20;The Patent LLM system helps translational research teams assess druggability and competitive landscape around specific gene targets by:

* **Target Validation Acceleration:** Quickly determining if a target is druggable by finding patents that have already attempted to target the gene of interest
* **Competitive Intelligence:** Identifying where a target falls in the landscape ("sweet spot" with few patents, "no patent" higher risk, or "many patents" requiring strategy pivot)
* **Research Efficiency:** Automating patent searches that would otherwise require significant manual effort
* **Informed Decision Making:** Supporting decisions about which targets to pursue, whether to build in-house or seek in-licensing, and competitive positioning

## Output formats

The Patent LLM agent returns **natural language summaries with structured elements** embedded in conversational text. The output is designed to provide both high-level overview and detailed patent-by-patent analysis, with references and links embedded within the narrative.

<figure><img src="/files/FqVFrIAPubL6nYY7QXPs" alt=""><figcaption></figcaption></figure>

## Query examples by use case

### IP landscape analysis

**Purpose:** Map the patent landscape around a target or mechanism to understand who owns what IP and identify crowded vs. open spaces.

#### **Example Prompt:**

> "I'm interested in **DGAT2** as a therapeutic target. Are there patents related to this target? In which disease areas? Please cover as many modalities as possible (small molecules, antibodies, etc.), discard abandoned patents, and include a summary of the claims as well as the entity submitting the patent. Include correct links (no hallucinations).

### Target druggability assessment

**Purpose:** Determine if a target has been successfully pursued before (evidence of druggability) and understand what classes of molecules have been tested.

#### **Example Prompt:**

> "Are there patents covering assets against target CD73? I want to understand if this target is druggable and already have small molecules or antibodies been developed against it?"

### Freedom-to-operate research

**Purpose:** Assess whether developing a therapeutic against a specific target will infringe existing patents; identify potential blocking patents.

#### **Example Prompt:**

> "I'm developing a **small molecule inhibitor** targeting **KRAS G12C**. What US patents exist that cover KRAS G12C inhibitors? Please identify the key claims and the companies holding these patents. Highlight any patents that are still active (not expired or abandoned)."

### Competitive intelligence

**Purpose:** Understand the competitive landscape—who is working on the same target, what development stage, and what's the strategic positioning.

#### **Example Prompt:**

> "What is the competitive landscape for **PD-L1** as a therapeutic target? Which companies hold patents? What are the key mechanistic approaches (checkpoint blockade, antibody-drug conjugates, bispecifics)?"

### Early research (target ID and prioritization)

**Purpose:** During target discovery/validation, use patent data as one signal to prioritize targets—understanding which have precedent and which are novel.

#### **Example Prompt:**

> "I have a list of candidate targets for **ovarian cancer**: **FOLR1**, **NECTIN4**, **TROP2**, and **CLDN6**. For each target, identify how many US patents exist that claim therapeutic assets. I want to understand which targets have the most precedent (druggability signal) vs. which are under-explored."

### In-licensing (asset assessment)

**Purpose:** During due diligence for in-licensing or acquisition, assess the IP landscape around the asset's target and mechanism to understand competitive risk and differentiation.

#### **Example Prompt:**

> "I'm assessing a **NECTIN4-targeted ADC** for in-licensing. Are there existing patents covering NECTIN4 ADCs? What do the claims cover—target specificity, linker chemistry, payload, or combination strategies?"

## Limitations

* Currently, the tool only leverages US patents (USPTO via PatentsView).


# Interactive data visualisations


# Prompting for a plot

**What to do:** Specify the chart type, variables, and dataset. K Pro supports real-time plot iteration: ask for specific formatting changes (colors, labels, font size) in follow-up prompts rather than trying to specify everything at once.

**Which agent/skill:** Visualization Agent

**Example prompt:**

> Create a Kaplan-Meier survival curve for TCGA-BRCA patients stratified by TP53 mutation status (mutated vs. wild-type). Include the p-value from a log-rank test and the number of patients at risk.

<figure><img src="/files/UT9dKDyV17XG10g8cFAn" alt=""><figcaption></figcaption></figure>

**Expected result:** A Kaplan-Meier plot with two curves, log-rank p-value, and at-risk table.

[**A-S-R-T-C breakdown**](https://docs.owkin.com/getting-started/prompting-guide-and-prompt-library)**:** Create \[A] TP53 mutation survival analysis \[S] at bulk level \[R] as a KM curve with log-rank test \[T] in TCGA-BRCA, mutated vs. wild-type \[C].

#### Create comparative visualizations

**Prompt:**

> Create a box plot comparing the expression of CD274 (PD-L1) across the five molecular subtypes in TCGA-BRCA. Add individual data points and mark statistically significant pairwise comparisons.

**Expected result:** A box plot with overlaid data points and significance brackets between groups.

**Follow-up:**

> Now create a heatmap of the top 50 differentially expressed genes between these molecular subtypes.

**Tip:** You can request multiple plot types in sequence. Each builds on the data context from the previous query.


# Limitations & how to evaluate results


# Example: Visualisations by modalities

## Overview

* Clinical: 3 plots
* Bulk RNA-Seq: 6 plots
* Single-Cell RNA-Seq: 7 plots
* Spatial Transcriptomics: 12 plots
* Histomics: 1 plot
* Multi-Modal: 3 plot

## Clinical Plots

### 1. Clinical Endpoint Plot

**Modality:** Clinical

**Description:** The clinical endpoint plot can be used to answer questions about the distribution of a clinical variable across patient groups or perform survival analysis when the variable is a survival endpoint. It can be used in two different settings:

1. **Survival Analysis**: If the variable is a survival endpoint (death or progression), it uses a Kaplan-Meier plot displaying survival probability over time
2. **Variable Description**: If the variable is not a survival endpoint, it displays the distribution of the variable across patient groups

**Parameters:**

* `title`: Plot title
* `group`: Grouping criterion (default: indication)
* `variable_name`: The name of the variable to plot (must be “death” or “progression” for Kaplan-Meier)
* `use_kaplan_meier`: Whether to use the Kaplan-Meier plot (must be True for survival endpoints

**Use Cases (Survival Analysis):**

* How does survival differ between MOSAIC patients with and without KRAS mutations in Non-Small Cell Lung Cancer?

<figure><img src="/files/4GYEoOURiz6uEHUsvZBu" alt=""><figcaption></figcaption></figure>

* How does survival differ between patients with and without KRAS mutations in colorectal cancer?
* What is the progression-free survival for patients with different tumor stages in ovarian cancer?
* How does age at diagnosis affect survival outcomes in lung cancer patients?
* What is the survival probability for patients with different molecular subtypes of breast cancer?
* How does treatment response impact overall survival in patients with advanced melanoma?

**Use Cases (Variable Description):**

* What is the patient age distribution across MOSAIC indications?

<figure><img src="/files/FKzBXQqjilwpyPWMJzkQ" alt=""><figcaption></figcaption></figure>

* What is the mutation status of KRAS across MOSAIC indications?
* Are there more smokers among Bladder or in Ovarian MOSAIC patients?
* How are the tumor stages distributed across different indications?
* How well do lung cancer patients respond to chemotherapy based on KRAS mutation status?

***

### 2. Gantt Plot

**Modality:** Clinical

**Description:** Gantt chart for patient treatment timeline visualization at the patient level. This displays horizontal bars representing treatment periods for each patient, with markers indicating when samples were collected during treatment. It utilizes treatment data, clinical data, and sample metadata to provide a comprehensive view of each patient’s treatment journey.

**Parameters:**

* `title`: Plot title
* `nb_patients`: Number of patients to display (default: 8)

**Use Cases:**

* Show the treatment timelines of KRAS mutated MOSAIC ovarian cancer patients with sample collection points

<figure><img src="/files/A8ud5bLY69ibhrWdCv09" alt=""><figcaption></figcaption></figure>

* What are the treatment durations across all patients in the cohort?
* Display the treatment timeline for patients in the responder group
* When were samples collected relative to treatment start and end dates?
* Show the temporal relationship between treatment periods and sampling events for each patient

***

### 3. Sankey Plot

**Modality:** Clinical

**Description:** The Sankey diagram summarizes treatment sequences at the cohort level. Each node represents a treatment type at a given treatment line. Each link shows how many patients moved from one treatment to another between successive lines of therapy. Links are colored according to the patients’ group, allowing you to see which treatment paths are most common and how they differ across categories.

**Parameters:**

* `title`: Plot title
* `group`: Optional group to color the Sankey links by (e.g., gender, mutation status, response status)

**Use Cases:**

* Visualize the treatment flow of MOSAIC Ovarian cancer patients grouped by M Stage status

<figure><img src="/files/EoFPqPgLXobMYXzmpJtQ" alt=""><figcaption></figcaption></figure>

* Show treatment flow for MOSAIC patients stratified by TP53 mutation status
* Color the Sankey plot by EGFR status for MOSAIC lung cancer patients
* Create a Sankey diagram grouped by best response for MOSAIC patients
* Show MOSAIC treatment transitions split by lymphocyte density

***

## Bulk RNA-Seq Plots

### 1. Bulk Violin Plot

**Modality:** Bulk RNA-Seq (BKRNASEQ)

**Description:** The violin plot displays variations in gene expression through violin plots. It can be used in two configurations:

1. Compare one gene or gene signature expression across groups
2. Compare several genes or gene signatures expression without another grouping criterion

The plot supports both regular grouping and stratified grouping. Each violin plot represents either the distribution of a gene or the distribution of signature scores.

**Parameters:**

* `title`: Plot title
* `gene_input_list`: List of gene inputs (single genes or gene signatures)
* `group`: Grouping criterion (required if gene\_input\_list length is 1, must be None if length > 1)
* `stratify_by`: Optional stratification criterion (e.g., gender, mutation status)

**Important Rules:**

1. When several genes/gene signatures are requested, group MUST be None
2. When a single gene or gene signature is requested, group MUST be filled
3. Stratification can be applied to compare expression across groups while stratifying by another variable

**Use Cases:**

* What's the expression of Nectin4 across MOSAIC indications?

<figure><img src="/files/zHJ5QTqYGOjPb7BHBVWe" alt=""><figcaption></figcaption></figure>

* How is the expression of TP53 across smoking status?
* Is IL21 more highly expressed in patients with lung cancer than in patients with breast cancer?
* Can you compare the expression of IL21 across tumor stages for patients with gastric cancer?
* How is the expression of the Cytotoxic\_T\_Cell\_Signature \[CD8A, GZMB, PRF1, CXCL9] varying across smoking status?
* How is the expression of TP53 across smoking status for patients with or without KRAS mutation?
* Is IL21 more highly expressed for male or female patients across indications?

***

### 2. Bulk Heatmap Plot

**Modality:** Bulk RNA-Seq (BKRNASEQ)

**Description:** This tool displays a heatmap of bulk RNA-Seq data to compare the expression of multiple genes across different groups. This is equivalent to the bulk violin plot but for simultaneously displaying multiple genes and/or gene signatures across multiple groups.

**Parameters:**

* `title`: Plot title
* `gene_input_list`: List of single genes and gene signatures (required, max 20 items)
* `group`: Grouping criterion (default: indication)

**Use Cases:**

* How does the expression of genes ROR1, ERBB2 and CDK4 vary across MOSAIC indications?

<figure><img src="/files/jMcSx7fHQtrNjNKxK2Sn" alt=""><figcaption></figcaption></figure>

* What tumor stages express WNT1 but not TP53?
* What is the expression of gene CDK4 and gene signature \[ERBB2, ROR1] across indications?
* How does the expression of gene signature \[ERBB2, ROR1] vary across different indications?

***

### 3. Bulk UMAP Gene Expression Plot

**Modality:** Bulk RNA-Seq (BKRNASEQ)

**Description:** The Bulk UMAP plots bulk RNA-Seq data in reduced dimension space (2D UMAP). Each dot matches a sample, and the color is determined by gene expression of selected genes. This plot is relevant for viewing gene expression across one group of patients or bulk RNA-Seq samples, not for comparing across groups.

**Parameters:**

* `title`: Plot title
* `gene_input`: The gene or gene signature to plot

**Requirements:**

* Requires a single indication

**Use Cases:**

* Display IL6 expression on a UMAP of MOSAIC bladder indication BulkRNAseq data

<figure><img src="/files/xMTRQVpcYBiEWnXQu4Wz" alt=""><figcaption></figcaption></figure>

* How is TP53 expressed in the bulk RNA-Seq data?
* How is YAP1 expressed for this one group of patients?
* Is SOX2 expressed in breast cancer?
* How is TNFa expressed in MOSAIC?
* What is the expression of IL6 in TCGA?
* What is the expression of gene signature CD8A, GZMB, PRF1, CXCL9 in TCGA?

***

### 4. Bulk UMAP Metadata Plot

**Modality:** Bulk RNA-Seq (BKRNASEQ)

**Description:** The Bulk UMAP plots bulk RNA-Seq data in reduced dimension space (2D UMAP). Each dot matches a sample, and the color is determined by the group parameters (clinical data or bulk RNA-Seq metadata, by default indication).

**Parameters:**

* `title`: Plot title
* `group`: Grouping criterion (default: patient\_id)

**Requirements:**

* Requires a single indication

**Use Cases:**

* Show me a UMAP of the bulk RNA-Seq data for MOSAIC patients with bladder cancer grouped by smoking status

<figure><img src="/files/Ph4TmjmzdgpjjwHgKcfE" alt=""><figcaption></figcaption></figure>

* Show me a UMAP of the bulk RNA-Seq data for patients over 60 years old with breast cancer
* Show me a UMAP of the bulk RNA-Seq data for dataset TCGA
* What is the transcriptomic profile of patients in MOSAIC with ovarian cancer?

***

### 5. Bulk Pairwise Correlation Plot

**Modality:** Bulk RNA-Seq (BKRNASEQ)

**Description:** The pairwise correlation plot displays correlations between gene expression levels in a heatmap format. It shows pairwise correlations between genes or gene signatures, with correlation coefficients displayed in each cell and p-values available on hover.

**Parameters:**

* `title`: Plot title
* `gene_input_list`: List of genes or gene signatures to correlate (required, max 15 items)
* `correlation_method`: Method for correlation calculation (“pearson” or “spearman”, default: “pearson”)

**Use Cases:**

* What is the correlation between ERBB2 and ESR1 expression in MOSAIC NSCLC?

<figure><img src="/files/tzZVoV6X1KXhWhVEJp99" alt=""><figcaption></figcaption></figure>

* Show me the pairwise correlations between TP53, BRCA1, and BRCA2 in TCGA breast cancer
* What is the correlation between an immune checkpoint signature \[PDCD1, CTLA4, LAG3] and a DNA repair signature \[BRCA1, BRCA2, ATM]?
* How correlated are the expression levels of HER2, ESR1, and PGR in MOSAIC breast cancer?
* Show correlations between oncogenes MYC, KRAS, and PIK3CA in MOSAIC ovarian cancer

***

### 6. Bulk DEA Plot (Differential Expression Analysis)

**Modality:** Bulk RNA-Seq (BKRNASEQ)

**Description:** This tool performs bulk differential expression analysis (DEA) on bulk RNA-Seq data. It compares gene expression between different groups to identify differentially expressed genes. This analysis identifies genes that are significantly up- or down-regulated between conditions, and can be used to identify top differential genes between two groups of patients. It can be adjusted for available covariates and can include user-specified genes in the results.

**Parameters:**

* `title`: Plot title
* `group`: Grouping criterion for differential expression analysis (e.g., gender, condition)
* `gene_input_list`: Optional list of single genes or gene signatures to include in results
* `covar_adj_variable`: Optional variables to adjust for in the analysis

**Use Cases:**

* In MOSAIC glioblastoma patients who received chemotherapy, which genes are differentially expressed between complete responders and patient with progressive disease based on BulkRNAseq data?

<figure><img src="/files/YLpJPN4HmrZYr5hXFOC2" alt=""><figcaption></figcaption></figure>

* What are the top 10 most up-regulated genes in condition A compared to condition B?
* Are gene X and Y differentially expressed between condition A and B?
* Perform differential expression analysis comparing tumor stages
* Which genes show significant expression differences across treatment groups?
* Find differentially expressed genes between responders and non-responders
* What genes are upregulated in high-risk patients compared to low-risk patients?

***

## Single-Cell RNA-Seq Plots

### 1. Single-Cell Cell Type Proportion Plot

**Modality:** Single-Cell RNA-Seq (SCRNASEQ)

**Description:** This plot displays variations in cell-type proportions for each cell type at the sample level. Each violin plot represents a cell type, and if a grouping parameter is requested, each violin is colored according to the group parameter.

**Parameters:**

* `title`: Plot title
* `group`: Optional grouping criterion
* `cell_group`: Cell type level (“cell\_type\_level\_1\_major” or “cell\_type\_level\_2\_mid”, default: “cell\_type\_level\_2\_mid”)

**Use Cases:**

* What is the cell type composition of patients with lung versus bladder cancer?

<figure><img src="/files/smbwd6G19skYim8hmE0m" alt=""><figcaption></figcaption></figure>

* Do smokers have more T cells than non-smokers in gastric cancer?
* Do we see more immune cells in later stages of MESO?
* Do patients with KRAS mutation have more stromal cells than patients without?

***

### 2. Single-Cell Patient Cell Expressing Plot

**Modality:** Single-Cell RNA-Seq (SCRNASEQ)

**Description:** This plot shows the percentage of cells expressing a gene across different groups (such as indications, tumor stages, or patients) for each cell type. Results per patient are shown by hovering over the scatter points.

**Parameters:**

* `title`: Plot title
* `gene_input`: Single gene to filter on
* `group`: Grouping criterion (default: indication)
* `cell_type`: Cell type level (“cell\_type\_level\_1\_major” or “cell\_type\_level\_2\_mid”, default: “cell\_type\_level\_2\_mid”)

**Use Cases:**

* How often is TP53 expressed across MOSAIC indications for each cell type?

<figure><img src="/files/sTDGCnpVXoCu4YQy2BRJ" alt=""><figcaption></figcaption></figure>

* How often is TP53 expressed across tumor stages for each cell type?
* How often is TP53 expressed across indications for each cell type, restricting to patients older than 65?

***

### 3. Single-Cell Dot Plot

**Modality:** Single-Cell RNA-Seq (SCRNASEQ)

**Description:** This plot is equivalent to a signature heatmap plot but for single-cell data, incorporating both average expression and the proportion of cells expressing the gene. It can be used in two main configurations:

1. Multiple genes/signatures on x-axis, groups on y-axis: Compare expression of MULTIPLE genes across different groupings
2. Single gene/signature with patient groups on x-axis, cell groups on y-axis: Display expression of EXACTLY ONE gene across different patient and cell groupings

**Parameters:**

* `title`: Plot title
* `xgroup`: Either a list of genes/signatures OR a GroupRequest for patient groupings
* `ygroup`: GroupRequest for grouping data (default: cell\_type\_level\_2\_mid)
* `gene_input`: Optional single gene or gene signature for metadata dot plot case

**Use Cases:**

* How are YAP1 and TGFB1 expressed in different cell types of MOSAIC patients?

<figure><img src="/files/9jB1fGPPS83TgZvGbvLW" alt=""><figcaption></figcaption></figure>

* Do T cells express more ERBB2 or TNFa than B cells?
* Compare the expression of CD4, CD8 and CD19 in different indications in single-cell data
* How is WNT1 expressed in different cell types across different indications?
* Show the expression of ERBB2 gene across different cell types and indications
* Do T cells express more TP53 than B cells across all tumor stages?

***

### 4. Single-Cell Coexpression Dot Plot

**Modality:** Single-Cell RNA-Seq (SCRNASEQ)

**Description:** This plot visualizes the co-expression patterns of two gene or gene signature inputs across different cell groups. It shows the proportion of cells co-expressing both inputs and compares it to the proportion expressing each one individually. The comparison is done through the Jaccard index between the two inputs.

**Parameters:**

* `title`: Plot title
* `gene_input_list`: List of exactly 2 single genes or gene signatures
* `group`: Grouping criterion for the y-axis (default: cell\_type\_level\_2\_mid)

**Use Cases:**

* How are EGFR and ERBB2 co-expressed in different cell types in MOSAIC bladder patients?

<figure><img src="/files/vw95NZJfR1KvjrYGWS5W" alt=""><figcaption></figcaption></figure>

* In which cell type is the gene pair (ERBB2,EGFR) most co-expressed?
* Investigate the gene association patterns of the target pair ERBB2 and EGFR
* How are gene signatures \[CD4, CD8, CD19] and \[TGFB1, TGFB2, TGFB3] co-expressed in different cell types?
* What are the synergies between gene signatures in different cell types?

***

### 5. Single-Cell UMAP Gene Expression Plot

**Modality:** Single-Cell RNA-Seq

**Description:** This plot displays single-cell RNA-Seq data in a 2D UMAP where each dot represents a cell. The color of the dots is determined by gene expression of the selected genes. If more than one gene is selected, the color will be a gradient of the mean expression. This plot automatically includes a corresponding UMAP showing cell type annotations for context.

**Parameters:**

* `title`: Plot title
* `gene_input`: Gene or gene signature to plot

**Requirements:**

* Requires a single indication

**Use Cases:**

* Display IL6 expression on a UMAP of MOSAIC bladder indication ScRNAseq data

<figure><img src="/files/kvIk10p4EWepGA6jR7Ap" alt=""><figcaption></figcaption></figure>

* How is TP53 expressed in the single-cell RNA-Seq data?
* How is TP53 expressed across all cells?
* Is SOX2 expressed in all cells in breast cancer?
* How is gene signature CD8A, GZMB, PRF1, CXCL9 expressed across all cells?

***

### 6. Single-Cell UMAP Metadata Plot

**Modality:** Single-Cell RNA-Seq

**Description:** This plot displays single-cell RNA-Seq data in a 2D UMAP where each dot represents a cell. The color of the dots is determined by the group parameters (clinical data or single-cell metadata, by default indication).

**Parameters:**

* `title`: Plot title
* `group`: Grouping criterion (default: cell\_type\_level\_3\_granular)

**Requirements:**

* Requires a single indication

**Use Cases:**

* Show me a UMAP of MOSAIC patients with bladder cancer grouped by cell types

<figure><img src="/files/BDVQ18Zd9Q6Z4r6IWnxE" alt=""><figcaption></figcaption></figure>

* Show me a UMAP of the single-cell RNA-Seq data for patients with lung cancer
* Show me a UMAP of the single-cell RNA-Seq data for patients over 60 years old
* Show me a UMAP of the single-cell RNA-Seq data for dataset MOSAIC
* Do patients with lung cancer have a different overall gene expression profile than patients with breast cancer?
* What is the transcriptomic profile of patients in MOSAIC?

***

### 7. Single-Cell Correlation Heatmap Plot

**Modality:** Single-Cell RNA-Seq

**Description:** Creates a correlation dot matrix showing pairwise correlations between genes or gene signatures in single-cell RNA-seq data, filtered by a specific cell type. The dot matrix displays correlation coefficients (color), percentage of cells co-expressing both genes (size), and statistical significance.

**Parameters:**

* `title`: Plot title
* `gene_input_list`: List of genes or gene signatures to analyze (max 20 items)
* `correlation_method`: Correlation method (“pearson” or “spearman”, default: “pearson”)
* `cell_level`: Cell type level column (default: “cell\_type\_level\_2\_mid”)

**Use Cases:**

* Create a correlation dot matrix for genes ERBB2, ROR1, and TGFB1 in T cells of MOSAIC patients

<figure><img src="/files/u4CXz7EXmxlat9rKFfJo" alt=""><figcaption></figcaption></figure>

* Show me the correlation and co-expression between CD4, CD8, and CD19 genes in B cells
* What are the gene expression correlations and co-expression patterns in T cells for genes CD4, CD8 and CD19?
* Compare correlations and co-expression between genes and gene signatures in different cell types
* Generate a correlation dot matrix for the ERBB2\_ROR1 signature in macrophages

***

## Spatial Transcriptomics Plots

### 1. Spatial Transcriptomics Dot Plot

**Modality:** Spatial Transcriptomics

**Description:** This plot is equivalent to a signature heatmap plot or single-cell dot plot but for spatial transcriptomics data. It can be used to compare the expression of MULTIPLE genes or signatures when grouping spots by cell types or patient-derived groups.

**Parameters:**

* `title`: Plot title
* `gene_input_list`: List of genes or gene signatures (required, max 20 items)
* `group`: Grouping criterion for the y-axis (default: DominantCellType\_cohort\_level\_2)

**Important:** This plot handles MULTIPLE genes or signatures. For ONE gene or signature, use SptMetadataDotPlot instead.

**Use Cases:**

* How are SYPL1 and TGFB1 expressed in different cell types of MOSAIC ovarian cancer patients measured by Visium?

<figure><img src="/files/gRGVEZl74AcAuFUOGxeO" alt=""><figcaption></figcaption></figure>

* Compare the expression of CD4, CD8 and CD19 across different cell types in Visium data
* How are KLRC1 and CD8A expressed in different cell types measured by Visium?
* How spatially heterogeneous is the expression of CD8A in female patients?
* How are gene YAP1 and gene signature \[TGFB1, TGFB2] expressed in different cell types in Visium data?

***

### 2. Spatial Transcriptomics Metadata Dot Plot

**Modality:** Spatial Transcriptomics

**Description:** This function displays the gene expression of exactly ONE gene or gene signature across different groupings of spots (by default cell-type). This plot is equivalent to a bulk violin (stratified) plot but for spatial transcriptomics data, incorporating both average information and the proportion of spots expressing the gene.

**Parameters:**

* `title`: Plot title
* `gene_input`: Single gene or gene signature to filter on
* `group`: Grouping criterion (default: indication)
* `spot_group`: Metadata column to group on at spot level (“DominantCellType\_cohort\_level\_2” or “tumor\_region”, default: “DominantCellType\_cohort\_level\_2”)

**Important:** This plot handles ONLY ONE gene or gene signature. For multiple genes or signatures, use SptDotPlot instead. It shows expression across TWO groups: one for patients and one for cell types.

**Use Cases:**

* How is SYPL1 expressed in different cell types across different MOSAIC indications for Visium data?

<figure><img src="/files/r1APCRbMDCS1u0EVRLhK" alt=""><figcaption></figcaption></figure>

* Do T cells express more TP53 than B cells across all tumor stages in spatial transcriptomics data?
* Compare the expression of CD4 for KRAS positive and negative patients in meso across cell types for Visium data
* How is gene signature PDCD1, PDCD1LG2 expressed in MESO patients across cell types for Visium data?
* How is gene signature PDCD1, PDCD1LG2 expressed in MESO patients across tumor regions for Visium data?

***

### 3. Spatial Transcriptomics Slide Display Gene Expression Plot

**Modality:** Spatial Transcriptomics

**Description:** This plot displays gene expression of spots on a spatial representation of the tissue. Spots are colored according to selected gene expression. The plot can include the H\&E image as a background to visualize tissue morphology alongside gene expression. This plot automatically includes a corresponding plot showing dominant cell type distribution.

**Parameters:**

* `title`: Plot title
* `gene_input`: Gene or gene signature to filter on
* `patient_id`: Optional patient ID to filter on (if not provided, one will be selected automatically)

**Requirements:**

* Requires filtering to select a single patient\_id

**Use Cases:**

* How is CD4 spatially distributed in the tissue of the oldest patient in MOSAIC mesothelioma indication?

<figure><img src="/files/gxYFEgkzJrhtMkngyw9k" alt=""><figcaption></figcaption></figure>

* How is CD8A spatially distributed in the tissue of patient xxx?
* Is WNT1 expressed in specific regions of the tissue of patient xxx?
* How is ROR1, ERBB2 signature spatially distributed in the tissue of the oldest patient in MESO?

***

#### 4. Spatial Transcriptomics Slide Display Cell Types Level 2 Plot

**Modality:** Spatial Transcriptomics

**Description:** This plot displays the cell type deconvolution of spots on a spatial representation of the tissue. Each spot is represented by a pie chart drawn according to the fraction of each cell type. The plot can include the H\&E image as a background to visualize tissue morphology alongside cell type information.

**Parameters:**

* `title`: Plot title
* `patient_id`: Optional patient ID to filter on (if not provided, one will be selected automatically)

**Requirements:**

* Requires filtering to select a single patient\_id

**Use Cases:**

* How are the cell types distributed in the tissue for the oldest patient in MOSAIC mesothelioma indication?

<figure><img src="/files/XPdGAEM0svpOyHY6kKQ9" alt=""><figcaption></figcaption></figure>

* Are there signs of immune infiltration in the tissue of patient xxx?
* Are T cells and tumor cells colocalized in the tissue of patient xxx?
* Can you show me the cell type deconvolution in the tissue for a patient in BRCA?

***

### 5. Spatial Transcriptomics Slide Display Colocalization Plot

**Modality:** Spatial Transcriptomics

**Description:** This plot displays the co-expression of two genes or the colocalization of a gene and a cell type at every spot on a spatial representation of the tissue. It can display:

1. Spatial similarity between expressions of two genes
2. Spatial similarity between expression of one gene and proportion of a particular cell type

Every spot is colored according to the co-expression/colocalization. Permutation-based p-values are calculated and can be viewed by hovering over spots.

**Parameters:**

* `title`: Plot title
* `gene_input_list`: List of 1 or 2 single genes (depending on use case)
* `patient_id`: Optional patient ID to filter on
* `cell_type`: Optional cell type level for gene-cell type colocalization (must be None when gene\_input\_list has 2 items)

**Requirements:**

* Requires filtering to select a single patient\_id

**Use Cases:**

* How is the co-expression of ERBB2 and CD4 spatially distributed in the tissue of the oldest patient in MOSAIC mesothelioma indication?

<figure><img src="/files/TmkFTcgbkHNswJethbKP" alt=""><figcaption></figcaption></figure>

* Show the spatial co-expression of ERBB2 and CD8A in patient xxx
* How are ERBB2 and WNT1 spatially co-expressed on slide xxx?
* Show me the spatial association of ERBB2 and CD4 in the spatial representation of tissue of patient xxx
* How is the expression of ERBB2 spatially correlated to the proportion of T\_NK cells in tissue of the oldest patient?
* Show the spatial association of ERBB2 and Malignant cells in tissue of patient xxx
* How are ERBB2 and B\_cells colocalized on slide xxx?
* Is gene ROR1 spatially associated with the presence of MoMac cells in tissue of patient xxx?

***

### 6. Spatial Transcriptomics Slide Display Metadata Plot

**Modality:** Spatial Transcriptomics

**Description:** This plot displays metadata information (categorical variables) of spots on a spatial representation of the tissue. Spots are colored according to the selected grouping variable (by default tumor region). The plot can include the H\&E image as a background to visualize tissue morphology.

**Parameters:**

* `title`: Plot title
* `patient_id`: Optional patient ID to filter on (if not provided, one will be selected automatically)
* `spot_group`: Grouping criterion for spots (default: tumor\_region)

**Requirements:**

* Requires filtering to select a single patient\_id

**Use Cases:**

* Can you show me the tumor region annotation for the oldest patient in MOSAIC mesothelioma indication?

<figure><img src="/files/wPZH5Upce7jiV63cm5V3" alt=""><figcaption></figcaption></figure>

* How are the dominant cell types distributed in the tissue for the oldest patient in MESO?
* What are the different tumor regions in the tissue of patient xxx?
* Show me the spatial distribution of tumor regions for a patient in BRCA?

***

### 7. Spatial Transcriptomics Coexpression Dot Plot

**Modality:** Spatial Transcriptomics

**Description:** This plot visualizes the spatial co-expression patterns of two genes or gene signatures across different groups. It shows the proportion of spatial spots (and their nearest neighbors) co-expressing both genes. The comparison is done through the Jaccard index, and is particularly useful for identifying groups where genes are spatially co-expressed.

**Parameters:**

* `title`: Plot title
* `gene_input_list`: List of exactly 2 single genes or gene signatures
* `group`: Grouping criterion for the y-axis (default: DominantCellType\_cohort\_level\_2)

**Use Cases:**

* How are EGFR and ERBB2 spatially co-expressed in different dominant cell types of MOSAIC bladder cancer patients?

<figure><img src="/files/iEEC2Z4gl2DWs5mxWEii" alt=""><figcaption></figcaption></figure>

* In which cell type is the gene pair (ERBB2,EGFR) most co-expressed spatially?
* Investigate the gene association patterns of the target pair ERBB2 and EGFR
* Assess the spatial coexpression of CD19 and MS4A1 in ovarian cancer patients
* Show the spatial co-expression for TP53 and MDM2 expression in GBM female patients older than 65 years

***

### 8. Spatial Transcriptomics Cell Type Proportion Plot

**Modality:** Spatial Transcriptomics

**Description**: Displays variations in cell-type deconvolution fractions for each cell type at the sample level using box plots. One box per cell type. The optional `color_group` parameter colors boxes by any clinical or spatial metadata variable. Includes statistical testing between groups when a grouping is provided. Analogous to the documented Single-Cell Cell Type Proportion Plot, but derived from spatial deconvolution rather than scRNA-seq.

**Parameters**:

* `color_group_display_name`: Optional display name for a color grouping variable (e.g., indication, gender)
* `title`: Optional custom title

**Use Cases**:

* What is the cell type composition of patients with lung versus bladder cancer in spatial transcriptomics?

<figure><img src="/files/m8bkBmxnXdEKuRcBDZZk" alt=""><figcaption></figcaption></figure>

* Do smokers have more T cells than non-smokers in gastric cancer based on spatial transcriptomics?
* Do we see more immune cells in later stages of mesothelioma in spatial transcriptomics data?
* Do patients with KRAS mutation have more stromal cells than patients without in spatial transcriptomics?

***

### 9. Spatial Transcriptomics Patient-Level Cell Type Co-occurrence Plot (Jaccard Index)

**Modality**: Spatial Transcriptomics

**Description**: Summarises spatial co-localisation between two selected cell types across patients using violin plots. For each sample, a Jaccard index is computed from spot-level deconvolution fractions (using the 30th percentile as a presence threshold). Per-sample Jaccard scores are displayed as violins per patient group (e.g., indication), with individual sample points overlaid. A gene-stratified variant further splits the violins by gene expression strata (gene+/gene−).

**Parameters**:

* `cell_type_1`: First cell type (default: `Malignant`)
* `cell_type_2`: Second cell type (default: `B_cell`; must differ from `cell_type_1`)
* `x_group_display_name`: Display name for the grouping variable
* `gene_input` *(stratified variant only)*: Gene or gene signature to derive strata from (default: ERBB2)
* `title`: Optional custom title

**Use Cases**:

* Compare colocalization of Malignant and B cells across MOSAIC indications

<figure><img src="/files/8CuiTqrnWl6a4Y1iregT" alt=""><figcaption></figcaption></figure>

* How does Malignant vs T\_NK colocalization vary between lung and breast?
* Show colocalization of DC and MoMac by tumor stage
* Do ERBB2-high samples show more Malignant/B\_cell co-occurrence than ERBB2-low samples?

***

### 10. Spatial Transcriptomics Patient-Level Moran’s I Plot

**Modality**: Spatial Transcriptomics

**Description**: Summarises the spatial autocorrelation (Moran’s I) of a gene or gene signature across patients using violin plots. Within each sample, Moran’s I is computed on spot-level log-normalised counts using spatial coordinates. The per-sample Moran’s I scores are displayed as violins per patient group (e.g., indication), with individual sample points overlaid. A high Moran’s I indicates the gene tends to be expressed in spatially clustered regions within the tissue.

**Parameters**:

* `gene_input`: Single gene or gene signature (required; default: ERBB2)
* `x_group_display_name`: Display name for the grouping variable
* `title`: Optional custom title

**Use Cases**:

* Compare Moran’s I of ERBB2 across MOSAIC indications

<figure><img src="/files/AFe4vF9UbzWpVYw65tz2" alt=""><figcaption></figcaption></figure>

* Is TP53 spatially autocorrelated in lung vs breast cancer?
* Show Moran’s I for the Cytotoxic\_T signature across tumor stages

***

### 11. Spatial Transcriptomics Cell Type Proportion Stacked Bar Chart

**Modality**: Spatial Transcriptomics

**Description**: Displays proportions of all detected cell types as a stacked bar chart, with each bar representing a clinical group (e.g., indication, cancer stage). Proportions are first averaged across spots per patient, then averaged across patients per group. A minimum proportion threshold collapses low-abundance cell types into an “Other” category. Cell types are ordered by biological hierarchy (Immune, Malignant, Stromal, Epithelial, Other). Includes statistical testing.

**Parameters**:

* `x_group_display_name`: Display name for the grouping variable
* `include_cell_types`: Optional list of specific cell types to include (default: all detected)
* `min_proportion_threshold`: Minimum mean proportion threshold to show a cell type separately (default: 0.01); cell types below this are grouped as “Other”
* `show_all_cell_types`: Set to `true` to override the threshold and show all cell types
* `proportion_method`: `deconvolution_mean` (default, average deconvolution fraction across spots) or `dominant_fraction` (fraction of spots where the cell type is dominant)
* `title`: Optional custom title

**Use Cases**:

* What cell types are detected in the spatial transcriptomics data for this group of patients?
* How is the spatial transcriptomics-derived cell type composition distributed across indications?
* Show in a stacked bar chart ho the spatial transcriptomics-derived cell type composition are distributed across MOSIAC breast and lung indications

<figure><img src="/files/7x8QI3Aqrxdxg9K4Q1Z1" alt=""><figcaption></figcaption></figure>

***

### 12. Spatial Transcriptomics Cell Type Proportion Groups Plot

**Modality**: Spatial Transcriptomics

**Description**: Shows the deconvolution fraction for a single selected cell type across patient groups (x-axis), with samples split into two gene-expression strata (gene+/gene−). The stratification is derived from a gene or gene signature using either a median split or a spot-presence method. Two grouped box plots are shown side-by-side within each x-axis group. This plot is for **cell type composition questions**, not gene expression questions.

**Parameters**:

* `cell_type`: Cell type to plot (e.g., `malignant`, `T_NK`, `B_cell`; must match a spatial deconvolution level 2 column)
* `x_group_display_name`: Display name for the grouping variable
* `gene_input`: Gene or gene signature used to split samples into two strata (default: PTGES)
* `stratification_method`: `median_split` (default) or `presence_any_spot`
* `title`: Optional custom title

**Use Cases**:

* Compare malignant deconvolution scores across cohorts, stratified by PTGES expression
* Do Malignant proportions differ by cohort code for ERBB2-high vs ERBB2-low samples?
* Within each cohort, are fibroblast proportions different for a cytotoxic T-cell signature high vs low?
* Compare T\_NK deconvolution fractions across indications, split by PDCD1 expression (+/-)

***

## Histomics Plots

### 1. Histomics Cell Type Proportion Plot

**Modality:** Histomics

**Description:** Given the selected filters, the histomics cell type proportion plot displays the proportions of different cell types in tumor samples as a stacked bar chart. Each bar represents a slide or a group of slides. This plot is relevant to answer questions about the cell type composition of samples based on H\&E slides.

**Parameters:**

* `title`: Plot title
* `group`: Grouping criterion (default: indication)
* `min_proportion_threshold`: Minimum proportion threshold for cell types to be displayed individually (default: 0.01)
* `show_all_cell_types`: Whether to show all cell types or group low-proportion ones into ‘Other’ (default: False)
* `max_slides`: Maximum number of slides to display (default: 20)
* `sort_by`: Optional column to sort cell types by
* `include_cell_types`: Optional list of cell types to include in the plot
* `legend_title`: Title of the legend (default: “Cell Types”)

**Use Cases:**

* How is the H\&E-derived cell type composition distributed across MOSAIC indications?

<figure><img src="/files/8UP4sTlh9NVuJQZNrShC" alt=""><figcaption></figcaption></figure>

* What cell types are detected in the H\&E slides for this group of patients?
* What is the cell type composition based on histology for this group of patients?
* Based on the H\&E slide, do patients with lung cancer have a different cell type composition than patients with breast cancer?
* Based on the H\&E slide, do patients with lung cancer have more tumor cells than patients with breast cancer?

***

## Multi-Modal Plots

### 1. Bulk RNA-Seq vs Single-Cell Expression Concordance Plot

**Modality**: Multi-Modal (Bulk RNA-Seq + Single-Cell RNA-Seq)

**Description**: Scatter plot combining bulk RNA-seq and single-cell RNA-seq data. Each point represents a sample. The x-axis shows bulk expression log2(TPM+1) and the y-axis shows the percentage of cells of a selected cell type expressing the gene. A dropdown allows switching between cell type granularity levels. Points are optionally colored by a patient grouping variable. Accepts a single gene only.

**Parameters**:

* `gene_input`: Single gene to display (required)
* `color_group_display_name`: Display name for the optional color grouping variable (default: indication)
* `cell_type`: Cell type granularity level — `cell_type_level_1_major`, `cell_type_level_2_mid` (default), or `cell_type_level_3_granular`
* `title`: Optional custom title

**Use Cases**:

* Compare bulk versus single-cell expression of CD274 in B cells across MOSAIC indications

<figure><img src="/files/0aZhoGwArmz7UTTQyxDa" alt=""><figcaption></figcaption></figure>

* How does bulk expression of ERBB2 relate to the percentage of T cells expressing it?
* Show the relationship between bulk TP53 expression and malignant cell expression percentage
* How does bulk TP53 expression correlate with malignant cell expression percentage in lung vs breast cancer?
* Bulk tissue expression compared to single cell malignant expression for CD274

***

### 2. Bulk RNA-Seq vs Spatial Transcriptomics Expression Concordance Plot

**Modality**: Multi-Modal (Bulk RNA-Seq + Spatial Transcriptomics)

**Description**: Scatter plot combining bulk RNA-seq and spatial transcriptomics data. Each point represents a sample. The x-axis shows bulk expression log2(TPM+1) and the y-axis shows the percentage of spots expressing the gene for a selected tumor region. A dropdown allows switching between tumor regions (e.g., Tumor, Stroma, Interface). Points are optionally colored by a patient grouping variable. Accepts a single gene only.

**Parameters**:

* `gene_input`: Single gene to display (required)
* `color_group_display_name`: Display name for the optional color grouping variable (default: indication)
* `title`: Optional custom title

**Use Cases**:

* Compare bulk versus spatial transcriptomics expression of CD274 across indications
* Show the relationship between bulk TP53 expression and spatial spot expression percentage
* How does bulk TP53 expression correlate with spatial spot expression percentage in lung vs breast cancer?
* Bulk tissue expression compared to spatial transcriptomics expression for CD274
* How does bulk expression of ERBB2 relate to the percentage of spots expressing it in MOSAIC breast cancer patients?

<figure><img src="/files/lZQHVMdcDjOOOrj78MT9" alt=""><figcaption></figcaption></figure>

***

### 3. Bulk Expression vs Spatial Moran’s I Plot

**Modality**: Multi-Modal (Bulk RNA-Seq + Spatial Transcriptomics)

**Description**: Scatter plot showing the relationship between bulk gene expression (log2 TPM+1, x-axis) and spatial autocorrelation (Moran’s I, y-axis) for a selected gene or gene signature. Moran’s I is computed per sample from spot-level log-normalised counts using spatial coordinates. Each point represents a sample, colored by a patient grouping variable (e.g., indication). Useful for identifying genes that are highly expressed at the bulk level and also spatially organised within the tissue.

**Parameters**:

* `gene_input`: Single gene or gene signature (required)
* `color_group_display_name`: Display name for the color grouping variable (default: indication)
* `title`: Optional custom title

**Use Cases**:

* Show the relationship between bulk TP53 expression and spatial autocorrelation
* Concordance between bulk expression and spatial Moran’s I for a gene signature
* Show the relationship between bulk expression and spatial Moran’s I for ERBB2 in MOSAIC breast cancer patients


# Use cases library

Each use case below shows a complete K Pro workflow from question to result. Prompts are copy-pasteable. Follow them in sequence: each step builds on the context of the previous one.

### Use case 1: Pan-cancer target prioritization

**Goal.** Identify cancer indications where a candidate gene is overexpressed, then check whether that overexpression correlates with patient outcomes. Useful for ADC target prioritization.

**Datasets used.** TCGA (pan-cancer).

**Prerequisites.** A candidate gene (HGNC symbol). The example uses NECTIN4.

#### Step 1: Pan-cancer ranking

*Prompt:*

> &#x20;Compare the expression of NECTIN4 across all TCGA cancer types. Show me a pan-cancer overview ranked by median expression level.

*Expected result:* A ranked table or visualization of NECTIN4 expression across TCGA indications.

#### Step 2: Survival correlation in top indications

*Prompt:*

> For the top 3 indications with highest NECTIN4 overexpression, show me the correlation between NECTIN4 expression and overall survival

*Expected result:* Survival analyses for each of the top 3 indications.

[*A-S-R-T-C breakdown for Step 1*](https://docs.owkin.com/getting-started/prompting-guide-and-prompt-library)*:* Compare \[A] NECTIN4 \[S] at Bulk RNA level \[R] as a ranked table \[T] across all TCGA cancer types \[C].

*Next steps to explore:* Drill into single-cell expression in the top indication (use the Chain of Thought sequence in the Prompting guide).

***

### Use case 2: Biomarker exploration

**Goal.** Characterize a TCGA cohort across clinical and molecular variables, then identify the worst-prognosis subgroup and its molecular features. Useful for biomarker hypothesis generation.

**Datasets used.** TCGA-LUAD.

**Prerequisites.** None.

#### Step 1: Literature summary

*Prompt:*

> I have a drug targeting both EP2 and EP4 in cancer. Identify most relevant genes for the prostaglandin pathway activity in oncology that I could use as a biomarker.&#x20;

*Expected result:* A comprehensive stratification panel covering gene biosynthesis, receptor expression, degradation and transport.

#### Step 2: Biomarker exploration using BulkRNA

*Prompt:*

> Which TCGA cancer indication expresses the highest PTGS2 expression, and what is the relative expression level compared to other indications?

*Expected result:* Gene expression across cancer types in TCGA datasets with population summary.

*Tip:* List the specific variables you want characterized upfront in Step 1 — K Pro works best when it knows exactly what you're looking for.

#### Step 3: Biomarker Exploration using ScRNA

*Prompt:*

> Show me the percentage of cells expressing PTGS2 per patient per cell type across mosaic indications?

*Expected result:* Gene expression across cancer types in TCGA datasets with population summary.

*Tip:* List the specific variables you want characterized upfront in Step 1 — K Pro works best when it knows exactly what you're looking for.

***

### Use case 3: Literature review on a drug target

**Goal.** Build an evidence base for a candidate target by surveying recent publications and probing for predictive-biomarker evidence. Useful for in-licensing or target validation due-diligence.

**Datasets used.** PubMed (via Consensus).

**Prerequisites.** A target gene + indication. The example uses TROP2 in triple-negative breast cancer.

#### Step 1: Target landscape in indication

*Prompt:*

> Find publications investigating TROP2 as a therapeutic target in triple-negative breast cancer. Include any data on TROP2 expression levels and their correlation with clinical outcomes.

*Expected result:* A curated list of relevant publications with key findings, expression data, and clinical correlations.

#### Step 2: Predictive-biomarker evidence

*Prompt:*

> Based on these publications, what is the evidence for using TROP2 expression as a patient selection biomarker for ADC therapies?

*Expected result:* A synthesis of biomarker-relevant evidence from the publications surfaced in Step 1.

*Tip:* Combine the target name with a specific indication AND the type of evidence you need (expression, outcomes, mechanisms). Vague prompts like "tell me about TROP2" return overly broad results.

***

### Use case 4: Cross-dataset discovery linking genomics, immunotherapy response, and literature

**Goal.** Identify genes that are both statistically associated with a mutation and an immunotherapy outcome, cross-reference with literature on the relevant biology, and then check spatial expression patterns. Useful for novel-mechanism discovery in a known molecular context.

**Datasets used.** TCGA-HNSC

**Prerequisites.** A mutation status of interest + an outcome variable. The example uses VHL mutation status + immune checkpoint inhibitor response.

#### Step 1: Identify candidate genes

*Prompt:*

> In the TCGA-HNSC cohort, identify genes whose expression is significantly associated with both TP53 mutation. Cross-reference findings with published literature on TP53-related immune evasion mechanisms.

*Expected result:* A list of candidate genes with statistical associations, linked to supporting literature evidence.

#### Step 2: Spatial expression of top candidates

*Prompt:*

> Can you plot the expression of the top5 candidate genes across Tumor islets, edge, stroma stratified by TP53 mutation status in mosaic HNSC patients?

*Expected result:* Spatial expression maps for the top 3 candidate genes (where MOSAIC data is available for the relevant indication).

*Tip:* This is a multi-step, multi-agent use case. Be explicit about both the data analysis you want AND the literature cross-reference — K Pro maintains conversation history within a session, so context from Step 1 carries into Step 2.


# Case studies library

### Case study 1: Multimodal target & population characterization for asset positioning

The brief was to characterize two candidate targets for an IO asset in a solid-tumor indication: expression profile, intra-tumoral heterogeneity, spatial behavior in the tumor microenvironment, and outcome signal in the IO-treated subgroup.

The data substrate was Owkin's MOSAIC cohort for the indication, covering 490 real-world patients across five treatment cohorts, with 208 patients having all four modalities (bulk RNA-seq, single-cell RNA-seq, spatial transcriptomics, clinical follow-up).&#x20;

The deliverable integrated all four modalities in a single report. Cross-modality concordance gave the client more confidence than any single modality could provide. Intra-tumoral heterogeneity and spatial autocorrelation surfaced patterns that bulk-only platforms cannot produce.

The skills underlying this analysis (cohort building, multimodal target characterization, spatial TME analysis, and outcome association) are the same skills K Pro orchestrates as composable workflows today.

***

### Case study 2: Drug to clinically actionable patient subgroups

The brief was to identify the patient subgroup most likely to respond to an immune checkpoint asset (an inhibitory NK-receptor antagonist combined with an anti-PD1) across two solid-tumor indications, and to distill the answer into clinically implementable selection rules a clinician could apply at trial screening.

*Approach.* A drug-relevant feature set spanning the target axis, checkpoints, NK and T cell signatures, and the broader microenvironment was scored per patient. Consensus clustering yielded three subgroups along a gradient of drug-enabling context (cold, warm, hot). Cluster labels transferred to two independent private cohorts at Jaccard 0.87 and Cohen Kappa 0.87 (n=278 on the larger cohort). The subgroups were distilled into a three-gene transcriptomic signature and a histomics decision tree (PDL1 status, lymphocyte density, cancer-cell density, TLS) reaching sensitivity 0.80 and specificity 0.74 on H\&E. Trial simulation estimated approximately 30% reduction in recruitment needed for an equivalent-powered Phase III in the larger indication. The client is now validating these biomarkers prospectively in Phase II.

K Pro's drug-positioning skill suite covers this exact five-step workflow: cohort building, drug-context feature engineering, patient clustering, cluster characterization, and biomarker distillation across transcriptomics, histomics, and clinical modalities.


# Patent LLM

## Overview

&#x20;The Patent LLM system helps translational research teams assess druggability and competitive landscape around specific gene targets by:

* **Target Validation Acceleration:** Quickly determining if a target is druggable by finding patents that have already attempted to target the gene of interest
* **Competitive Intelligence:** Identifying where a target falls in the landscape ("sweet spot" with few patents, "no patent" higher risk, or "many patents" requiring strategy pivot)
* **Research Efficiency:** Automating patent searches that would otherwise require significant manual effort
* **Informed Decision Making:** Supporting decisions about which targets to pursue, whether to build in-house or seek in-licensing, and competitive positioning

## Output formats

The Patent LLM agent returns **natural language summaries with structured elements** embedded in conversational text. The output is designed to provide both high-level overview and detailed patent-by-patent analysis, with references and links embedded within the narrative.

<figure><img src="/files/FqVFrIAPubL6nYY7QXPs" alt=""><figcaption></figcaption></figure>

## Query examples by use case

### IP landscape analysis

**Purpose:** Map the patent landscape around a target or mechanism to understand who owns what IP and identify crowded vs. open spaces.

#### **Example Prompt:**

> "I'm interested in **DGAT2** as a therapeutic target. Are there patents related to this target? In which disease areas? Please cover as many modalities as possible (small molecules, antibodies, etc.), discard abandoned patents, and include a summary of the claims as well as the entity submitting the patent. Include correct links (no hallucinations).

### Target druggability assessment

**Purpose:** Determine if a target has been successfully pursued before (evidence of druggability) and understand what classes of molecules have been tested.

#### **Example Prompt:**

> "Are there patents covering assets against target CD73? I want to understand if this target is druggable and already have small molecules or antibodies been developed against it?"

### Freedom-to-operate research

**Purpose:** Assess whether developing a therapeutic against a specific target will infringe existing patents; identify potential blocking patents.

#### **Example Prompt:**

> "I'm developing a **small molecule inhibitor** targeting **KRAS G12C**. What US patents exist that cover KRAS G12C inhibitors? Please identify the key claims and the companies holding these patents. Highlight any patents that are still active (not expired or abandoned)."

### Competitive intelligence

**Purpose:** Understand the competitive landscape—who is working on the same target, what development stage, and what's the strategic positioning.

#### **Example Prompt:**

> "What is the competitive landscape for **PD-L1** as a therapeutic target? Which companies hold patents? What are the key mechanistic approaches (checkpoint blockade, antibody-drug conjugates, bispecifics)?"

### Early research (target ID and prioritization)

**Purpose:** During target discovery/validation, use patent data as one signal to prioritize targets—understanding which have precedent and which are novel.

#### **Example Prompt:**

> "I have a list of candidate targets for **ovarian cancer**: **FOLR1**, **NECTIN4**, **TROP2**, and **CLDN6**. For each target, identify how many US patents exist that claim therapeutic assets. I want to understand which targets have the most precedent (druggability signal) vs. which are under-explored."

### In-licensing (asset assessment)

**Purpose:** During due diligence for in-licensing or acquisition, assess the IP landscape around the asset's target and mechanism to understand competitive risk and differentiation.

#### **Example Prompt:**

> "I'm assessing a **NECTIN4-targeted ADC** for in-licensing. Are there existing patents covering NECTIN4 ADCs? What do the claims cover—target specificity, linker chemistry, payload, or combination strategies?"

## Limitations

* Currently, the tool only leverages US patents (USPTO via PatentsView).


# Customize K Pro

K Pro is customizable at three levels.

**User specific, Per-team and per-org via custom skills.** Skills are both how users invoke complete multi-step workflows and how customer bioinformatics and biology teams encode internal best practice. Customer-specific skills can capture internal scoring methodologies, naming conventions, regulatory submission templates, or the methodologies of internal KOLs. Custom skills are versioned, promoted through a publish-and-review lifecycle, and can take precedence over Owkin-built skills under per-customer governance.

**Custom data and tool integration.** Proprietary datasets connect via the BYOD pattern (data lake, MCP server wrapping an internal database). Internal tools (compound libraries, biomarker pipelines, RWD systems) register as MCP servers and become first-class agents in K Pro's orchestrator. Custom scoring dimensions can be added through internal MCP servers (e.g. "K Pro, also penalize targets that failed our internal toxicity screen").

**Custom views and reports.** Filter and re-rank dynamically. Export to report formats matching your team's structure or branding. Persistent project workspaces with custom dashboards.

Skill authoring today benefits from Owkin-side support during onboarding. We are investing in self-service skill-creation tooling so customer teams can author skills directly. Status: in active development.


# Collaboration features

K Pro collaborative features are being built with the following principles in mind:

**Shared workspaces.** Project-level workspaces centralize a team's analytical work for a dedicated team, specific target, indication, or program. Cohort definitions are shared, versioned, and reusable across team members and analyses. Analytical artifacts (plots, reports, intermediate results) accumulate within a project as a shared body of evidence.

**Role-based access.** Our K-platform support fine-grained permissions, we are able to define roles such as Admin for an Organisation, Admin for a project and users within projects. Tiered access for collaborators (internal team members, external academic partners, CROs) is supported with strict level of isolation.

**Decision and audit trail.** All agent, users actions and tool invocations within a project are logged for audit purpose. Cohort and artifact version history are preserved. *"Why was this target prioritized?"* can be reconstructed from the audit trail at any point.

**Reporting.** Teams collaboratively assemble target assessment reports that combine K Pro's findings with team annotations and decisions. Reports export to standard formats for sharing with non-K Pro users.


# Browse the dataset catalog

K Pro comes with a curated set of datasets ready to use from day one, and gives you the ability to discover and request access to additional datasets from the Owkin catalog. This section describes what is available out of the box in K Pro Free, and how to explore further datasets that can be integrated into your project.

Owkin's data coverage is designed around **depth and multimodality** while maximizing breadth across all domains.

By default, the K Pro comes with a **foundation of public datasets that can support user’s research**: TCGA, GTEx, CPTAC. Other public datasets can be either uploaded into the product, or integrated at your request depending of the volume and complexity of the dataset.

To go beyond our readily available and sublicensable data products are currently concentrated in **oncology** (11 indications, including NSCLC, breast, ovarian, DLBCL, bladder, GBM, pancreatic, head & neck, and mesothelioma), where we offer one of the most comprehensive multimodal patient-level dataset catalogs available for licensing. This focus reflects deliberate curation rather than a gap, and **ensures high data quality, rich annotation, and clinical-grade metadata across these indications**.

Beyond our core catalog, **our data sourcing offering based on a network of 2.5M+ patient data points** **extends coverage to additional oncology indications as well as key therapy areas including Inflammation & Immunology** (e.g. IBD, SLE, RA), Neurology (Alzheimer's disease), and CVRM.

Regarding **data recency**, the majority of our datasets include patients enrolled from 2012 onwards, reflecting the period of most significant advances in molecular profiling and digital pathology. Our network infrastructure can support active refresh cycles or any *de novo* access through our data sourcing offering, with access to data collected through the current year for select partners and indications.

Geographically, our **proprietary data products draw from both US and EU cohorts**, providing transatlantic representation that supports regulatory-relevant diversity. Our sourcing network extends to the APAC region for targeted data acquisition when geographic diversity is a project requirement.

{% content-ref url="/pages/2uwOqESrXC0iFDj8MSGN" %}
[Datasets available in K Pro Free](/explore-and-analyse-data/data-catalog/datasets-available-in-k-pro-free)
{% endcontent-ref %}

{% content-ref url="/pages/y9KJm7yd1lPhdY2mf5lY" %}
[Browse for additional datasets of interests](/explore-and-analyse-data/data-catalog/browse-for-additional-datasets-of-interests)
{% endcontent-ref %}

The following pages further specify data assets available on K Pro:

{% content-ref url="/pages/8hJ2bLOWEh1QI1mtpGvs" %}
[Broken mention](broken://pages/8hJ2bLOWEh1QI1mtpGvs)
{% endcontent-ref %}

{% content-ref url="/pages/Tt9DSGf8Ajmp3mwBbfYT" %}
[Mosaic Window](/explore-and-analyse-data/data-catalog/browse-for-additional-datasets-of-interests/mosaic-window)
{% endcontent-ref %}

{% content-ref url="/pages/aYf7gOKLERsVCZTj5mNP" %}
[MOSAIC](/explore-and-analyse-data/data-catalog/browse-for-additional-datasets-of-interests/mosaic-dataset)
{% endcontent-ref %}


# Datasets available in K Pro Free

Below is the list of public and private datasets currently available in the Owkin K-Pro Free platform, along with their respective licenses and versions.

**1- Datasets that can be explored the MultiOmics Agent in K Pro Free**

The MultiOmics Agent analyzes complex biological datasets across different modalities, enabling you to explore relationships between molecular features and clinical outcomes.

| Dataset Name | Type   | Website                                                      | License Link                                                                                                                        | License Name                                                 | Version |
| ------------ | ------ | ------------------------------------------------------------ | ----------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------ | ------- |
| **TCGA**     | Public | <https://www.cancer.gov/ccg/research/genome-sequencing/tcga> | <https://creativecommons.org/licenses/by-nc-nd/4.0/>                                                                                | CC Attribution-NonCommercial-NoDerivatives 4.0 International | 4.0     |
| **GTEx**     | Public | <https://www.gtexportal.org/home/>                           | <https://www.gtexportal.org/home/license>                                                                                           | GTEx Portal data                                             | N/A     |
| **CPTAC**    | Public | <https://portal.gdc.cancer.gov/>                             | [https://github.com/PayneLab/cptac?tab=License-1-ov-file#readme](< https://github.com/PayneLab/cptac?tab=License-1-ov-file#readme>) | Apache                                                       | N/A     |

**2- Datasets used to generate pre aggregated insights for the Knowledge agent**

Knowledge data represent a set of (gene-level) features that are based on publicly available databases and ressources, and provide information on the general biology of the target, independently of the specific disease or discovery context. Currently covered databases and sources can be found below:

| Dataset Name                  | Type   | Website                                                         | License Link                                                                                                                 | License Name                                                                 | Version |
| ----------------------------- | ------ | --------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------- | ------- |
| **ChEMBL**                    | Public | <https://www.ebi.ac.uk/chembl/>                                 | [Deed - Attribution-ShareAlike 3.0 Unported - Creative Commons](https://creativecommons.org/licenses/by-sa/3.0/)             | Deed Attribution-Share                                                       | 3.0     |
| **CollecTRI**                 | Public | <https://github.com/saezlab/CollecTRI?tab=readme-ov-file>       | [The GNU General Public License v3.0 - GNU Project - Free Software Foundation](https://www.gnu.org/licenses/gpl-3.0.en.html) | The GNU General Public License v3.0 - GNU Project - Free Software Foundation | 3.0     |
| **Complex Portal**            | Public | <https://www.ebi.ac.uk/complexportal/home>                      | [Deed - CC0 1.0 Universal - Creative Commons](https://creativecommons.org/publicdomain/zero/1.0/)                            | Deed - CC0 1.0 Universal - Creative Commons                                  | 1.0     |
| **DepMap**                    | Public | <https://depmap.org/portal/>                                    | [Deed - Attribution 4.0 International - Creative Commons](https://creativecommons.org/licenses/by/4.0/)                      | Deed - Attribution 4.0 International - Creative Commons                      | 4.0     |
| **Ensembl (BioMart)**         | Public | <https://www.ensembl.org/info/data/biomart/index.html>          | [Deed - Attribution 4.0 International - Creative Commons](https://creativecommons.org/licenses/by/4.0/)                      | Deed - Attribution 4.0 International - Creative Commons                      | 4.0     |
| **gnomAD**                    | Public | <https://gnomad.broadinstitute.org/>                            | [Deed - CC0 1.0 Universal - Creative Commons](https://creativecommons.org/publicdomain/zero/1.0/)                            | Deed - CC0 1.0 Universal - Creative Commons                                  | 1.0     |
| **GTEx**                      | Public | <https://gtexportal.org/home/>                                  | [GTExPOrtal Data License](https://gtexportal.org/home/license)                                                               | GTExPOrtal Data License                                                      | 4.0     |
| **Hallmarks of Cancer**       | Public | <https://pubmed.ncbi.nlm.nih.gov/21376230/>                     | No licence                                                                                                                   | N/A                                                                          | N/A     |
| **Human Protein Atlas (HPA)** | Public | <https://www.proteinatlas.org/>                                 | [Deed - Attribution-ShareAlike 3.0 Unported - Creative Commons](https://creativecommons.org/licenses/by-sa/3.0/)             | Deed Attribution-Share                                                       | 3.0     |
| **IntOgen**                   | Public | <https://www.intogen.org/search>                                | [Deed - CC0 1.0 Universal - Creative Commons](https://creativecommons.org/publicdomain/zero/1.0/)                            | Deed - CC0 1.0 Universal - Creative Commons                                  | 1.0     |
| **MsigDB**                    | Public | <https://www.gsea-msigdb.org/gsea/msigdb/human/collections.jsp> | [Deed - Attribution 4.0 International - Creative Commons](https://creativecommons.org/licenses/by/4.0/)                      | Deed - Attribution 4.0 International - Creative Commons                      | 4.0     |
| **Open Targets**              | Public | <https://platform.opentargets.org/>                             | [Deed - CC0 1.0 Universal - Creative Commons](https://creativecommons.org/publicdomain/zero/1.0/)                            | Deed - CC0 1.0 Universal - Creative Commons                                  | 1.0     |
| **Reactome**                  | Public | <https://reactome.org/>                                         | [Deed - Attribution 4.0 International - Creative Commons](https://creativecommons.org/licenses/by/4.0/)                      | Deed - Attribution 4.0 International - Creative Commons                      | 4.0     |
| **Uniprot**                   | Public | <https://www.uniprot.org/>                                      | [Deed - Attribution 4.0 International - Creative Commons](https://creativecommons.org/licenses/by/4.0/)                      | Deed - Attribution 4.0 International - Creative Commons                      | 4.0     |


# Browse for additional datasets of interests

K Pro allows you to discover multimodal datasets that are part of the Owkin data catalog and that could be integrated to your project based on your research needs. These dataset span various modalities (bulk, single cell, spatial transcriptomics, WES) but also various therapeutic area such as Oncology, I\&I, Cardiovascular and Neurodegenerative diseases.

Datasets will be displayed in the “[Explore data](https://k.owkin.com/explore-data/overview)” page, where a user can register interest, and be directed to contact our sales team.

<figure><img src="/files/TagAxhK7WCEdl1qDEiD8" alt=""><figcaption></figcaption></figure>

<figure><img src="/files/I5gt72gDJlqhND7TdkAh" alt=""><figcaption></figcaption></figure>


# Mosaic Window

{% hint style="info" %}
**Availability**

The use of this dataset is free but requires to users to be habilitated through a research objective that is reviewed by the producing consortium. This dataset can be made available for a trial on K Pro tool at your request.
{% endhint %}

MOSAIC Window is a curated subset of the groundbreaking MOSAIC dataset, available through K Pro. It includes spatial omics and multimodal data from **100 patients** across 8 cancer types:

* **BLCA (Bladder Cancer):** 15 patients
* **OV (Ovarian Cancer):** 15 patients
* **GBM (Glioblastoma):** 10 patients
* **DLBCL (Diffuse Large B-cell Lymphoma):** 10 patients
* **MESO (Mesothelioma):** 10 patients
* **NSCLC (Non-Small Cell Lung Cancer):** 15 patients
* **HNSCC (Head & Neck Squamous Cell Carcinomas):** 10 patients
* **BRCA (Breast Cancer):** 15 patients

This unique resource enables researchers to explore tumor biology at near single-cell resolution, providing detailed insight into tumor and immune cell interactions.

<figure><img src="/files/jFskSBHYLJvFcnpbijhj" alt=""><figcaption></figcaption></figure>

**MOSAIC Window Patient Details**

**BLCA (15 patients):**

* Stage II and III urothelial carcinoma (one squamous cell carcinoma)
* Derived from upfront cystectomies
* Most patients receiving complete lymph node dissection
* Some treated with adjuvant chemotherapy, chemotherapy, or immune checkpoint inhibitors (ICI) at relapse, with or without radiotherapy

**OV (15 patients):**

* FIGO stage II to IV high grade serous carcinoma
* 7 baseline, 2 post-NACT, 6 relapse lesions
* Treated by upfront or post-NACT interval debulking surgeries and chemotherapy
* Additional treatments: Bevacizumab, PARP inhibitors, and ICI in first or later lines

**GBM (10 patients):**

* Glioblastoma as per WHO 2021 definition, all IDH wildtype
* 6 unmethylated and 4 methylated MGMT promoter samples
* Obtained from baseline surgery before standard adjuvant therapy
* 1 to 3 brain tumor sites, tumor sizes 27 to 80mm diameter
* 5 patients received Bevacizumab at relapse

**DLBCL (10 patients):**

* Ann Arbor stage III or IV
* 3 Activated B-cell-like type (ABC), 6 Germinal center B-cell group (GCB), 1 unknown subtype
* Obtained at baseline before R-CHOP therapy
* Targeted interventions at relapse

**MESO (10 patients):**

* Stage I to III pleural mesotheliomas
* 5 epithelioid, 5 biphasic
* 9 baseline biopsies/surgeries, 1 post-NACT sample
* 4 patients treated by NACT (1 partial response, 3 progressive diseases)
* 4 patients received anti-PD-1 ICI at relapse (monotherapy or combination with anti-CTLA4)

**NSCLC (15 patients):**

* 10 Adenocarcinoma and 5 squamous cell carcinoma
* Male-predominant (10/15)
* 8 patients ≥ 70 years old
* Mostly early-to-intermediate stage at diagnosis (Stage I: 5, Stage II: 4, Stage III: 2, Stage IV: 1)
* 50% former smokers, 29% active smokers and 21% never smokers 21.4%
* 14 patients received surgical resection
* 4 patients with KEAP1 mutations, 2 patients with TP53 mutations and 1 patient with a co-mutation of KEAP1+STK11
* 9 patients had upfront surgery, 3 patients had neoadjuvant chemotherapy (NACT) before surgery and 1 patient had initial 1L immunotherapy combination with chemotherapy for a metastatic disease.
* 11 samples are from baseline surgery, 3 from post-NACT surgical samples and 1 baseline sample for the patient treated with upfront chemotherapy + immunotherapy

**HNSCC (10 patients):**

* 6 male patients , mostly between 50 and 69 years old
* Majority of patients were advanced-stage at diagnosis — Stage IV (5), Stage III (2)
* All the patients received immunotherapy in the Recurrent Unresectable and/or Metastatic (R/M) setting.
* 9/10 patients had surgical resection.
* 4 active smokers, 5 former smokers, 1 never-smoker
* 3 patients had CPS > 20
* 1 patient has positive HPV status by p16 IHC
* 6 patients had upfront surgery with various adjuvant strategies and 2 patients had initial 1L anti-PD1 treatment
* 6 patients have paired samples from primary tumour and recurrence, including 1 patient with an additional post-Immunotherapy progression sample

**BRCA (15 patients):**

* 10 patients had infiltrating ductal adenocarcinoma
* 10 patients had triple-negative breast cancer (TNBC)
* Mostly early/intermediate stage disease at diagnosis— Stage I (3), II (7), III (2)
* 8 patients were postmenopausal and 7 were premenopausal
* All patients are BRCA1 / BRCA2 wild-type, and 2 patients had PIK3CA mutations
* 9 patients were treated with neoadjuvant chemotherapy (NACT) before surgery
* 2 patients have paired samples at baseline and post-NACT surgery, including 1 patient with an additional sample at disease recurrence.
* 4 other patients have paired samples at post-NACT surgery and disease recurrence.
* 1 samples is from a baseline metastatic site


# MOSAIC

## Overview

MOSAIC is a flagship Owkin data asset: a large spatially resolved dataset with **6 data modalities per sample across 11 cancer indications** in a centralized platform.

**11 cancer types covered:** NSCLC, Ovarian, Bladder, Mesothelioma, Glioblastoma, Breast, DLBCL, HNSCC, Pancreas, CRC, Gastric.

**6 data modalities per sample:**

* Clinical data: medical files and consent, clinically validated
* Spatial transcriptomics: subsequent slides from a FFPE block, pathology validated
* Single cell transcriptomics
* Bulk RNA-Seq
* Whole Exome Sequencing (WES)
* Digitized H\&E

**Sample breakdown**

**2,716 patients in the study.**

* \~15% of patients have multiple samples
* \~80% of samples are pre-treatment
* \~10% of samples are post-treatment
* \~10% of samples are relapse / recurrence

## **Clinical data collected**

**Common forms** (patient-related information):

* Demographics (date of diagnosis, date of last follow-up or death, general demographic information, cancer indication)
* Consent & Eligibility
* Subject history
* Treatment form (oncological treatment types, dosage, routes, dates and response)
* Oncologic events before inclusion in MOSAIC (progression/recurrence, other cancer; includes OS and PFS calculations)
* 'End of study' form (cause of end of study or death, if applicable)
* Follow-up forms (yearly occurrence of novel oncologic events and/or death)

**Cancer-type-specific forms:**

* Clinical (height, weight, date of diagnosis, tumor and metastasis location, cTNM, stage; plus tumor-specific information when applicable)
* Pathology (histological type & subtype, (y)pTNM, prognostic histological features, IHC and FISH results when applicable)
* Mutations (all known genetic alterations)

## **Data quality**

**Consistent and rigorous data generation:** uniform sample processing, centralised NGS, rigorous QC steps at every stage.

MOSAIC uses a dynamic, multi-role QC approach across the full workflow:

* **Clinician review** — patient selection (validation of cohort inclusion criteria), clinical record review (eCRF completeness, coherence, accuracy), workflow adaptation based on QC of existing database.
* **Pathologist review** — block selection (tissue of origin, histological subtype, sample timepoint, tumor content), histology slides QC (cuts, tumor content, scanning and staining artefacts), spatial transcriptomics QC.
* **Biologist review** — single-cell annotation (cohort-level cell type annotation, validation of automatic label transfer), clinical record completeness review.

## **Partner institutions**

MOSAIC is generated through partnerships with leading academic medical centers at the forefront of spatial omics research:

* **Gustave Roussy** — PI: Fabrice André (ESMO President-elect, h-index 116). World's top 15 hospitals (Newsweek, 2023). Spatial transcriptomics pioneers via Center for Experimental Therapies platform.
* **CHUV Lausanne** — PI: Raphaël Gottardo (h-index 60). Strong expertise in spatial biology via PETRA platform.
* **University of Pittsburgh** — PI: Robert Ferris (h-index 107). Ranked 3rd in 2022 NIH funding (behind only Johns Hopkins and UCSF). Among the largest academic medical centers in the US.
* **Uniklinikum Erlangen** — PI: Arndt Hartmann (h-index 108). Top-11 German hospitals (2023). Oncology cluster of excellence "NCT" with strong expertise in DLBCL, MM, Bladder, Breast, GBM.
* **Charité** — PI: Ulrich Keilholz (h-index 79). World's top-10 best hospitals and smart hospitals (Newsweek, 2023). #1 in Germany. Spatial transcriptomics pioneers via MDC center.

## **Access**

The broader MOSAIC dataset is part of K Pro paid tiers. The K Pro Free subset is available as **MOSAIC Window** above.


# Understand your data (QC, methods)

For your data to be fully exploitable by K Pro's AI agents, it needs to meet a defined set of structural and quality requirements. This section explains how Owkin thinks about AI data readiness, which biological modalities and file formats K Pro currently supports, and the naming conventions and ontologies data must conform to. Whether you are preparing your own dataset for integration or evaluating a third-party dataset, this section is the technical reference you need.

***

K Pro uses a multi-layer quality and provenance framework for public data.

**Data quality.** Public datasets are transformed into a unified K Data Model with schema harmonization, ontology mapping, normalization, and multi-level structuring. Quality controls include relational mapping to avoid orphaned records, pre-computed metadata, and other filtering steps for integrity and fast exploration.

**Provenance tracking.** Every analysis is anchored in authoritative sources and validated biological knowledge bases. The platform keeps complete provenance for outputs and logs data sources, model decisions, and reasoning steps. For literature-backed content, K Pro uses RAG to verify cited PubMed articles.

**Version control / traceability.** The AI-Readiness Maturity Model defines version history at Level 2 and full data lineage at Level 4. Level 5 adds reproducibility with code + environment and detailed audits. Public dataset pages also expose dataset versions where available — for example, the TCGA entry lists a version in the dataset catalog.

**Proprietary data (MOSAIC and similar).** Owkin computational biology teams have developed data processing pipelines following gold standards and have, where necessary, optimized the pipelines for the specific dataset. Biomedical experts have been included in the development loop for conducting confirmatory analyses with the data, annotations of single-cell clusters, etc., additionally ensuring a high quality data for discovery and other typical uses.

{% content-ref url="/pages/RLzBO1KSapCNhQnq1RdP" %}
[The AI-maturity model](/explore-and-analyse-data/k-pro-data-model-and-technical-references/the-ai-maturity-model)
{% endcontent-ref %}

{% content-ref url="/pages/plui2xjJVbZAi6NSMbak" %}
[Supported modalities](/explore-and-analyse-data/k-pro-data-model-and-technical-references/supported-modalities)
{% endcontent-ref %}

{% content-ref url="/pages/px5r2MV8w8ih27rA0Xy0" %}
[Preferred ontologies and nomenclatures](/explore-and-analyse-data/k-pro-data-model-and-technical-references/preferred-ontologies-and-nomenclatures)
{% endcontent-ref %}


# About Owkin-managed data

An opinionated structuring framework designed to bridge the gap between bioinformatics outputs and high-performance analytical queries.

### Overview

The K Data Model optimizes information retrieval across multiple biological scales by transforming raw files (e.g. h5ad, csv, zarr) into a unified, SQL-queryable parquet structure. It enables seamless cross-modality analysis—from clinical longitudinal trends down to single-cell gene expression and spatial spots—ensuring every data point is mapped to a common coordinate system and ontology.

### Pipeline overview

```mermaid
flowchart LR
    A("Raw Bioinformatics Files<br/>(h5ad, CSV, Zarr)") --> B(K Data Model Transformation)
    B --> C(Standardized Parquet Store)
    C --> D(SQL-Ready K Pro Platform)
```

### Processing & standardization

These steps transform heterogeneous data into a unified, high-performance schema.

| Step                        | What happens                                                         | Why                                                                   |
| --------------------------- | -------------------------------------------------------------------- | --------------------------------------------------------------------- |
| **Multi-Level structuring** | Data is organized into Patient, Sample, Slide, and Cell/Spot levels. | Optimizes query performance for different scales of analysis.         |
| **Schema harmonization**    | Conversion of raw formats to **Parquet** files.                      | Enables universal SQL querying and high-speed data loading.           |
| **Ontology mapping**        | Gene names and disease labels are matched to a single reference.     | Ensures consistency and enables cross-dataset comparisons.            |
| **Normalization**           | Counts for bulk, single-cell, and spatial data are normalized.       | Removes technical variation to allow biological interpretation.       |
| **Feature enrichment**      | Computation of UMAPs (bulk & single-cell) and pseudobulk.            | Provides instant, high-quality visualizations without on-the-fly lag. |

#### Data filtering & Quality control

The K Data Model applies specific rules to ensure data integrity and visualization speed.

| Filter / QC Step           | Criteria                                                          | Impact                                                                  |
| -------------------------- | ----------------------------------------------------------------- | ----------------------------------------------------------------------- |
| **Relational mapping**     | Uses common keys to link patients, samples, and cells.            | Prevents "orphaned" data; ensures clinical context is always available. |
| **Pre-computed metadata**  | Ranges, completeness, and cohort counts are calculated upfront.   | Accelerates UI responsiveness for filtering and exploration.            |
| **Visual optimization**    | Data is partitioned/saved specifically for heatmaps and dotplots. | Delivers near-instant rendering of complex large-scale matrices.        |
| **Longitudinal alignment** | Chronological ordering of treatments and events.                  | Enables accurate survival and treatment-response analysis.              |

#### Technologies used

| Technology                    | Purpose                                                         |
| ----------------------------- | --------------------------------------------------------------- |
| **AnnData / Scanpy**          | Handling single-cell and spatial transcriptomics objects.       |
| **NumPy / Pandas**            | Core matrix manipulation and numerical processing.              |
| **Parquet**                   | Columnar storage format for efficient, large-scale data access. |
| **AWS SDK (or other vendor)** | Cloud-native data management and storage interface.             |
| **SQL**                       | Universal query language for all standardized data levels.      |

What this means for your analyses:

* **Granular access:** Query clinical data (e.g., all patients on immunotherapy) and drill down to single-cell gene expression or immune cell density.
* **Ready-to-view:** Pre-computed visualizations like UMAPs and heatmaps reduce wait times.
* **Ontological truth:** Synonymous gene names or mismatched disease labels are aligned automatically.


# The AI-maturity model

An AI-ready dataset is a collection of biomedical data specifically prepared so that both humans and K-Pro can seamlessly use it for analytics, model training and research. To establish a systematic way of assessing the value of a dataset, Owkin has introduced an AI-Readiness Maturity Model on a 6-level scale for its own datasets:

* **Level 0 - Uncontrolled data:** Data lacks governance, compliance, or minimal metadata for cataloging.
* **Level 1 - Storage & Legal compliance:** Raw data with minimal metadata; stored securely, ISO 27001 compliant, license & IRB in place.
* **Level 2 – Discoverability:** Data dictionary, schema, programmatic metadata access, and version history available.
* **Level 3 – Exploration:** Quality checks documented, summary tables provided, with manifest/ReadMe for dataset exploration.
* **Level 4 – Interoperability:** Standard formats, cross-modality links, automated QC, and full data lineage.
* **Level 5 – Full traceability + Optimized for AI/ML:** Optimized views, precomputed features, reproducibility (code + environment), and detailed audits.

In order for a third-party dataset to be computed by K-Pro, it has to meet some strict requirements (layout, schema, dictionary) picked amongst the ones above, and described in this document. Owkin is keen to support your journey to achieve this.


# Supported modalities

As of March 2026, only the modalities below are supported by K-Pro Data Model:

* Clinical data
* Molecular data:
  * Bulk RNA seq
  * Single / nuclei cell RNA seq
  * Spatial transcriptomics (VisiumSD)
  * Whole Exome Sequencing
  * Proteomics
* Histology - slide annotation linked to a patient

{% hint style="info" %}
Extra modalities / technologies are available upon request. By default K-Pro expects to receive processed data (in the format of count matrixes for example for bulk and single cell data), however Owkin has also developed in-house processing pipeline that can process raw files for both bulk RNAseq, Chromium single-cell and Visium spatial sequencing.
{% endhint %}

#### Clinical data

Standard clinical data formats are supported (.csv, .tsv, .xlsx, …).

#### Molecular data

Here is a table of supported modalities and source formats.

| **Modalities**          | **Formats**                                                                                           |
| ----------------------- | ----------------------------------------------------------------------------------------------------- |
| Bulk RNAseq             | Count matrix (.txt / .tsv / .csv); AnnData (.h5ad)                                                    |
| Single cell RNAseq      | Matrix Market (.mtx + .tsv); HDF5 (.h5); AnnData (.h5ad); Seurat object (.rds)                        |
| Spatial transcriptomics | Matrix Market (.mtx + barcodes.tsv + features.tsv); HDF5 (.h5); AnnData (.h5ad); Seurat object (.rds) |
| WES / WGS               | VCF                                                                                                   |
| Proteomics              | Normalized intensity matrix (.txt / .tsv / .csv); AnnData (.h5ad)                                     |

#### Imaging data

Here is a table of supported modalities and formats.

| **Modalities**             | **Formats**                                 |
| -------------------------- | ------------------------------------------- |
| Histology (H\&E WSI)       | .tif, .tiff, .svs, .dcm, .svs, .ndpi, .mrxs |
| Immunohistochemistry (IHC) | .tif, .tiff, .svs, .dcm, .svs, .ndpi, .mrxs |

Note that Owkin developed several imaging processing pipelines that can be applied to imaging data, including:

* Cell segmentation and cell annotation via the [HIPE](https://arxiv.org/abs/2508.09926) model . Resulting cell-type quantifications are then saved (format: .csv) and used by K pro.
* IHC score extraction with proprietary models (*e.g.*, HER2 score, NMR, etc.). Resulting patient-level scores are then saved and used by K-pro.

Self Supervised Features extraction via [Foundation Models](https://arxiv.org/html/2501.16239v1#S4): slides are divided into tiles from which latent representations are computed, saved (format: .npy)


# Preferred ontologies and nomenclatures

The full schemas and data dictionaries of tables used in K-Pro can be provided upon request. Data loaded in K-Pro follows a OMOP like schema with the following notable categorical standards :

| Category Type                        | Possible Values                                                                                                                                                                            |
| ------------------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| QC Status                            | pass \| fail \| flag                                                                                                                                                                       |
| LDH Level Bins                       | Normal \| Elevated \| Low                                                                                                                                                                  |
| Protein Measure Type                 | global \| phospho                                                                                                                                                                          |
| Tumor Region                         | Stroma \| Tumor Edge \| Tumor Islet                                                                                                                                                        |
| Vital Status                         | Alive \| Deceased                                                                                                                                                                          |
| Gender                               | Male \| Female                                                                                                                                                                             |
| ECOG Performance Status              | 0 \| 1 \| 2 \| 3 \| 4 \| 5                                                                                                                                                                 |
| Smoking Status                       | Current \| Former \| Never                                                                                                                                                                 |
| Treatment Response (RECIST)          | Complete Response \| Partial Response \| Stable Disease \| Progressive Disease                                                                                                             |
| Treatment Line                       | Neoadjuvant \| Adjuvant \| First-line \| Second-line \| Third-line                                                                                                                         |
| Treatment Setting                    | Neoadjuvant \| Adjuvant \| Palliative \| Curative                                                                                                                                          |
| Treatment Type                       | chemotherapy \| immunotherapy \| surgery \| radiotherapy \| targeted therapy                                                                                                               |
| Sample Tissue                        | Tumor \| Normal \| Blood \| Metastasis                                                                                                                                                     |
| Histological Sample Type             | Primary Tumor \| Adjacent Normal \| Metastatic \| Recurrent                                                                                                                                |
| Sample Collection Chronology         | Baseline \| On-treatment \| Post-treatment \| Progression                                                                                                                                  |
| TNM T Stage                          | T0 \| T1 \| T1a \| T1b \| T1c \| T2 \| T2a \| T2b \| T3 \| T3a \| T3b \| T4 \| T4a \| T4b \| Tx \| Tis                                                                                     |
| TNM N Stage                          | N0 \| N1 \| N1a \| N1b \| N1c \| N2 \| N2a \| N2b \| N2c \| N3 \| N3a \| N3b \| Nx                                                                                                         |
| TNM M Stage                          | M0 \| M1 \| M1a \| M1b \| M1c \| Mx                                                                                                                                                        |
| TNM Stage Groups                     | Stage 0 \| Stage I \| Stage IA \| Stage IB \| Stage II \| Stage IIA \| Stage IIB \| Stage IIC \| Stage III \| Stage IIIA \| Stage IIIB \| Stage IIIC \| Stage IV \| Stage IVA \| Stage IVB |
| FIGO Stage                           | Stage I \| Stage IA \| Stage IB \| Stage IC \| Stage II \| Stage IIA \| Stage IIB \| Stage III \| Stage IIIA \| Stage IIIB \| Stage IIIC \| Stage IV \| Stage IVA \| Stage IVB             |
| Ann Arbor Stage                      | Stage I \| Stage II \| Stage III \| Stage IV (may include A/B suffix)                                                                                                                      |
| International Prognostic Index (IPI) | Low \| Low-Intermediate \| High-Intermediate \| High                                                                                                                                       |
| HER2 IHC Status                      | 0 \| 1+ \| 2+ \| 3+                                                                                                                                                                        |
| ER/PR IHC Status                     | Positive \| Negative \| (percentage values)                                                                                                                                                |
| Microsatellite Instability           | MSI-High \| MSI-Low \| MSS (Microsatellite Stable)                                                                                                                                         |
| MGMT Promoter Status                 | Methylated \| Unmethylated                                                                                                                                                                 |
| Platinum Sensitivity                 | Sensitive \| Resistant \| Refractory                                                                                                                                                       |
| Homologous Recombination Deficiency  | Positive \| Negative \| Deficient \| Proficient                                                                                                                                            |
| Surgery Type                         | Primary resection \| Debulking \| Biopsy only \| Mastectomy \| Lumpectomy                                                                                                                  |
| Residual Disease                     | R0 (complete resection) \| R1 (microscopic residual) \| R2 (macroscopic residual)                                                                                                          |
| Menopausal Status                    | Premenopausal \| Postmenopausal \| Perimenopausal                                                                                                                                          |
| Race                                 | White \| Black or African American \| Asian \| Native Hawaiian or Pacific Islander \| American Indian or Alaska Native \| Other \| Unknown                                                 |
| Alcohol Intake                       | Current \| Former \| Never                                                                                                                                                                 |
| COPD Stage                           | Stage 0 \| Stage I \| Stage II \| Stage III \| Stage IV                                                                                                                                    |
| End of Study Reason                  | Completed \| Death \| Disease Progression \| Adverse Event \| Withdrawal \| Lost to Follow-up                                                                                              |
| Variant Classification               | Missense\_Mutation \| Nonsense\_Mutation \| Frame\_Shift\_Del \| Frame\_Shift\_Ins \| In\_Frame\_Del \| In\_Frame\_Ins \| Splice\_Site \| Silent                                           |


# Connecting your data sources

Once data has been prepared, a secure connection must be established between the client's data repositories and K. The preferred access model should be agreed upon between the client and Owkin during the pre-sales phase.

Three connectivity options are available:

* **Transfer to Owkin-managed storage** Data is transferred into an Owkin-managed environment where it is stored and served directly to K. This is the simplest option and is well suited for clients who prefer a fully managed approach.
* **Expose customer-managed storage** Data remains in the client's own infrastructure. The client provides the necessary credentials and configuration so that K can access the data remotely. This option preserves the client's existing data residency and governance controls.
* **Connect an existing data platform** K integrates directly with platforms such as Databricks or Snowflake, leveraging the client's existing governance policies and data pipelines. This is the preferred option for organizations that have already invested in a centralized data platform.

For proprietary assets like screening results, assay data, target assessments, or compound libraries, the same pattern applies: bring the data in your preferred tabular or platform-backed structure, then align it to the supported model.

Owkin can also provide tooling and services for transformation, and optional AI enrichment can extract additional features before analysis.

***

K Pro's connection can be tailored to the use case. For instance, when connecting to a client's data platform, we can establish a live connection to a specific view of the data, or alternatively connect to a safeguarded version that requires manual updates — the choice depends entirely on the use cases being served. The same flexibility applies to other integration methods, such as connecting to storage (whether Owkin-managed or client-managed).


# Preparing your data

K is designed to work flexibly with the data clients provide. Within its supported modalities, K does not require data to conform to a single standard, ontology, or processing pipeline. However, the more standardized and homogeneous the data, the greater the analytical value K can deliver.

To help clients choose the right level of preparation, we define three tiers of data standardization, each with a different value-to-effort trade-off:

**Data Standardization Tiers**

* **Basic:** A small set of lightweight rules that the data must satisfy, for example, every clinical table must include a column with a unique patient identifier. These rules are minimal and easy to meet. The client is responsible for enforcing them. *(A full specification of these rules will be available soon and can be shared upon request.)*
* **Recommended:** An entry-level standardization effort focused on aligning key columns: demographics, procedural fields, and modality-specific measures (e.g., cell counts), to a shared data model. This makes the most common analytical operations directly comparable across datasets. Owkin can provide tooling (data contracts, methodology) and services to support these transformations.
* **Full:** Complete alignment to a common data model, ontology, and preparation pipeline. This level is typically required for use cases that involve cross-dataset cohorts or analyses run on merged datasets. Delivering this level of standardization requires a high-touch professional-services engagement from Owkin.

> **Optional enrichment:** Owkin has developed proprietary models that can extract additional features from raw data (for example, detecting cell types or adding bulk-deconvolution signatures) thereby augmenting the dataset before analysis.

**New modalities**

Adding aggregated tabular data for a new modality (such as a flow cytometry table) will not cause a technical failure in K. However, without a dedicated integration, the platform cannot guarantee the depth of analysis or reproducibility that a fully supported modality provides.

For this reason, we recommend clients to engage with the Owkin team before onboarding a new modality. This allows us to assess the data structure, confirm analytical coverage, and, where needed, implement the modality-specific logic required to deliver high-value, reproducible insights for the client's use case.


# Data enrichment

K Pro's AI toolkit enriches raw data across three stages: **data generation** (lab protocols and tissue sourcing), **data processing** (SOTA cloud-based QC & ETL pipelines), and **data augmentation** where AI transforms unstructured biological data into quantified, analysis-ready biology.

Four augmentation axis are currently available, each described below.

***

#### AI cell detection: Histomics

Histomics is Owkin's AI-based digital pathology tool for cell detection and segmentation, including tumour-infiltrating lymphocytes (TILs) and tertiary lymphoid structures (TLS).

**Key capabilities:**

* Detects **13 cell types**, including understudied immune populations such as neutrophils and eosinophils
* Trained across **5 cancer types**, leveraging transfer learning to maximise efficiency
* Achieves **24% better F1 classification** of cells and 5% better detection using 5× fewer parameters
* Built on **200,000 consensus annotations** from 10 pathologists

> Reference: Adjadj et al. arXiv 2025

***

#### AI spatial prediction

K Pro can predict gene expression at each spatial spot of a spatial transcriptomics cohort using the associated H\&E tile, enabling near single-cell resolution through model distillation.

The model uses a spatial neighbourhood attention architecture (multi-head attention over tile embeddings from neighbouring spots), and was benchmarked on the HEST dataset:

| Feature extractor | Training data | HEST Average (Pearson) |
| ----------------- | ------------- | ---------------------- |
| Baseline iBOT     | FFCD          | 0.246                  |
| H0                | FFCD          | 0.286                  |
| H0-mini           | FFCD          | 0.344                  |
| **H0-mini**       | **MOSAIC**    | **0.381**              |

> Reference: Schmauch et al. arXiv 2024

***

#### AI enhanced resolution: Deconvolution

K Pro applies deconvolution algorithms to increase the resolution of Visium spatial transcriptomics data down to single-cell level, leveraging paired modalities (H\&E + scRNA-seq + spatial).

Two outputs are supported:

* **Spot-level cell type deconvolution:** answers specific tumour microenvironment (TME) questions by identifying dominant cell types per spot
* **Spatialization of tumour transcriptomic clusters:** maps distinct tumour areas by learning cell signatures from single-cell RNA-seq on paired samples within the same cohort

For reference-free deconvolution, K Pro uses **MixUpVI**, a joint probabilistic model of pseudobulk and single-cell transcriptomics that estimates cell-type proportions without requiring a reference. Published at ICML 2025 (Grouard, Ouardini, Rodriguez, Vert, Espin-Perez).

***

#### AI cell-cell communication

K Pro models local ligand-receptor (LR) interactions using spatial data, without relying on a reference dataset. The pipeline computes LRI values across three diffusion modes — cell contact (no diffusion), secreted signalling (one-neighbour diffusion), and hormone signalling (two-neighbour diffusion) — using prior knowledge tables of ligand-receptor pairs.

Outputs include:

* **Ligand expression by cell type** (dot plot per programme)
* **Cellular communication network** (chord diagram of sender/receiver cell types)
* **Spatial map of LRI values** overlaid on the tissue slide


# File Upload

Upload files in chat for analysis with Code Execution.

Upload files from the chat surface to analyze your own data with Code Execution.

### Upload a file

Open the upload modal in chat and select a file. Each file can be up to **1 GB**.

Supported formats:

* PDF (`.pdf`)
* Text and Markdown (`.txt`, `.md`)
* Tabular data (`.csv`, `.tsv`, `.parquet`)

### Use uploaded files in Code Execution

Each upload is stored in your chat sandbox. Code Execution can access files in this location during your analysis.

Uploads are available only to the user who added them. Upload only low-sensitivity data.


# Pathology Explorer MCP : AI-Powered Tissue Analysis

Pathology Explorer is an advanced AI tool that transforms standard hematoxylin and eosin (H\&E) histology slides into detailed, queryable insights. By analyzing the complex spatial organization of the tumor microenvironment (TME), it enables researchers to uncover patterns that traditional histology analysis often misses—patterns that are critical for predicting treatment response and understanding disease progression.

### Key Capabilities

* **Comprehensive Cell Segmentation and Classification:** Trained on over 200,000 expert annotations, Pathology Explorer automatically segments and classifies all cells within a tissue sample in minutes, providing unprecedented granularity in tissue analysis.
* **Actionable Biomarker Discovery:** Transform raw histology data into interpretable, clinically relevant biomarkers that can inform research decisions and therapeutic strategies.
* **State-of-the-Art AI Architecture:** Powered by Owkin's advanced deep learning models, rigorously benchmarked against 6+ leading encoders and best-in-class architectures to ensure accuracy and reliability.
* **Large-Scale Database Integration:** Seamlessly analyze H\&E slides from major databases including The Cancer Genome Atlas (TCGA), enabling large-scale retrospective and prospective studies.

### Applications

Pathology Explorer empowers researchers and clinicians to:

* Quantify spatial relationships within the tumor microenvironment
* Identify predictive biomarkers for treatment response
* Accelerate histopathology research with automated, reproducible analysis
* Generate hypotheses about disease mechanisms based on tissue architecture


# Getting started

Use Pathology Explorer in Claude, via MCP (Model Context Protocol), to query TCGA H\&E slides. Browse cohorts, view thumbnails and tiles, run survival analyses, and export features.

### Connect Pathology Explorer to Claude

#### Prerequisites

* An Owkin account (create a free account at [k.owkin.com](https://k.owkin.com/auth/signup?next=%2Fchat)).
* A paid Claude plan (Pro, Max, Team, or Enterprise).
* Claude Custom Connectors enabled for your workspace.

The integration uses Claude’s **remote MCP custom connector** flow. Claude Free does not support this. See Anthropic’s guide: [Getting started with custom connectors using remote MCP](https://support.claude.com/en/articles/11175166-getting-started-with-custom-connectors-using-remote-mcp).

#### Claude.ai (web)

{% stepper %}
{% step %}

#### Add the connector (workspace admin)

Go to **Admin settings** → **Connectors** → **Add custom connector**.

Enter:

* **Name:** `Owkin`
* **URL:** `https://mcp.k.owkin.com/mcp`

Leave **OAuth Client ID** and **OAuth Client Secret** empty.
{% endstep %}

{% step %}

#### Connect and authenticate

Go to **Settings** → **Connectors**.

Click **Connect** next to **Owkin**.

Approve the access request. Then sign in to Owkin.

![](https://docs.owkin.com/~gitbook/image?url=https%3A%2F%2F1398098133-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FsQnMGEZUFMazkLv5a4BR%252Fuploads%252Fgit-blob-b6d1aea382214f2352ccdc3196675f8fab3f4187%252FOwkinMCPLogin.png%3Falt%3Dmedia\&width=768\&dpr=3\&quality=100\&sign=492f89d3\&sv=2)
{% endstep %}

{% step %}

#### Verify the connection

Start a new chat and ask:

```
Can you list the available TCGA cohorts?
```

If it works, Claude returns a cohort list from Owkin.
{% endstep %}
{% endstepper %}

#### Claude Desktop

If you’re on a Team/Enterprise workspace, the connector added by an admin should appear on Desktop too.

1. Open **Settings** → **Connectors**.
2. Click **Connect** next to **Owkin**.
3. Sign in to Owkin when redirected.
4. Run the same verification prompt.

{% hint style="info" %}
If you don’t see **Connectors** (or **Owkin** is missing), your workspace may not allow custom connectors. Ask your Claude admin.
{% endhint %}

### Troubleshooting

{% hint style="warning" %}
**Sign-in page looks stuck**

The auth page can appear to hang. The connection may still succeed.

Go back to Claude and run the verification prompt.
{% endhint %}

{% hint style="warning" %}
**Error: “Invalid or expired transaction”**

This usually happens after clicking **Allow access** twice.

Retry in a different browser. This is a known Chrome issue with `claude://` redirects.
{% endhint %}

{% hint style="warning" %}
**Connected, but tools don’t show up**

* Confirm you connected the **Owkin** connector in **Settings → Connectors**.
* Re-check the URL: `https://mcp.k.owkin.com/mcp`.
* Restart Claude after changes.
  {% endhint %}

{% hint style="warning" %}
**Session expired**

Sessions currently last **15 minutes**.

Disconnect and reconnect the Owkin connector in **Settings → Connectors**.
{% endhint %}

If you still can’t connect, submit a support request using this [form](https://owkinkhelp.zendesk.com/hc/en-us/requests/new?ticket_form_id=41391636133649).

### What you can do with Pathology Explorer

Pathology Explorer is built for cohort exploration and hypothesis testing on TCGA histology.

It is powered by Owkin’s model (see paper: <https://arxiv.org/abs/2508.09926>). It supports slide-level features and tile visualizations.

#### Example prompts

**Find cohorts and cases**

```
List the TCGA cohorts you support. Then show me the cohorts with the most slides.
```

**Stratify patients by immune infiltration**

```
In TCGA_LUAD, find slides with low lymphocyte density. Summarize patient-level trends.
```

**Find “most enriched” slides and plot**

```
In TCGA_BRCA, find the slide most enriched in eosinophils and show the thumbnail.
```

**Run survival analysis**

```
In TCGA_BLCA, is plasmocyte density associated with overall survival? Use OS and stratify patients.
```

**Export features for reproducibility**

```
Export histomics data for cohort TCGA_LUAD as parquet.
```

### Data and outputs

#### Available data

You can query TCGA cohorts available in this integration. Use:

```
List the available TCGA cohorts.
```

For the current cohort coverage list, also see [Extended features description for Pathology Explorer](/connect-and-integrate/pathology-explorer-mcp-ai-powered-tissue-analysis/understanding-pathology-explorers-analysis-capabilities).

#### Cell types

The model detects **six cell types** on H\&E. Ask for the exact list:

```
List the available cell types.
```

#### Outputs you can retrieve

* Slide thumbnails.
* Tile mosaics with predictions.
* Slide-level histomics features.
* Parquet exports for downstream analysis.

### Tool reference (advanced)

<details>

<summary>Show the MCP tools exposed by the connector</summary>

* **Pathology Explorer description**: summary of model and capabilities.
* **List available cell types**: returns the supported cell types.
* **List TCGA cohorts**: returns supported TCGA cohorts.
* **Describe slide-level histomics features**: describes feature names and meanings.
* **Filter slides by histomics features**: returns slide IDs matching feature criteria.
* **Get histomics features for slides**: returns features for specific slide IDs.
* **Display slide thumbnail**: renders a slide thumbnail.
* **Perform survival analysis**: runs OS/PFS survival analysis on a cohort.
* **Download histomics data**: presigned URL for Parquet download.
* **Filter tiles by histomics**: selects tiles by feature criteria.
* **Display tiles with predictions**: renders a tile mosaic with predictions.

</details>


# Understanding Pathology Explorer's Analysis Capabilities

Pathology Explorer exposes slide-level histomics features derived from TCGA H\&E slides.

This page documents the feature families and naming conventions.

### Cell quantification and distribution <a href="#cell-quantification-and-distribution" id="cell-quantification-and-distribution"></a>

For each supported `cell_type` (for example: lymphocytes, neutrophils, plasmocytes, fibroblasts, eosinophils, cancer cells), the model provides:

* `count_{cell_type}`: Total number of detected cells in the slide.
* `global_density_{cell_type}`: Cell density per unit tissue area.

{% hint style="info" %}
The exact `cell_type` tokens are returned by the **List available cell types** tool.
{% endhint %}

### Nuclear morphology metrics <a href="#nuclear-morphology-metrics" id="nuclear-morphology-metrics"></a>

For each `cell_type`, the model provides:

* `mean_area_{cell_type}`: Mean nuclear area.
* `mean_circularity_{cell_type}`: Mean nuclear circularity.
* `mean_perimeter_{cell_type}`: Mean nuclear perimeter.

### Spatial organization and tissue architecture <a href="#spatial-organization-and-tissue-architecture" id="spatial-organization-and-tissue-architecture"></a>

Pathology Explorer also measures how cells are distributed across tissue regions.

#### Regional density analysis <a href="#regional-density-analysis" id="regional-density-analysis"></a>

For three region types (tumor, tumor core, and stroma in tumor core), the model provides:

* `density_{cell_type}_in_{region}`: Cell density for a given cell type in a given region.

#### Regional area measurements <a href="#regional-area-measurements" id="regional-area-measurements"></a>

For each region, the model provides:

* `area_{region}`: Region area within the slide.

#### Cell–cell interaction analysis <a href="#cell-cell-interaction-analysis" id="cell-cell-interaction-analysis"></a>

For selected pairs of cell types, the model provides:

* `average_co_occurrence_{cell_type}_{cell_type2}_rad_20.0um`

This feature answers:

*How many `cell_type2` nuclei are found, on average, within 20 µm of each `cell_type` nucleus?*

Interpretation:

* **0** means no local co-occurrence at 20 µm.
* Larger values mean denser local neighborhoods of `cell_type2` around `cell_type`.

#### Tumor-infiltrating lymphocyte assessment <a href="#tumor-infiltrating-lymphocyte-assessment" id="tumor-infiltrating-lymphocyte-assessment"></a>

The model also computes:

* `tils_diffusivity`: a metric that quantifies how diffusely TILs are distributed in the slide.

### Supported TCGA cohorts <a href="#supported-tcga-cohorts" id="supported-tcga-cohorts"></a>

Histomics features are available for the following TCGA cohorts:

* TCGA\_ACC
* TCGA\_BLCA
* TCGA\_BRCA
* TCGA\_CESC
* TCGA\_CHOL
* TCGA\_COAD
* TCGA\_DLBC
* TCGA\_ESCA
* TCGA\_HNSC
* TCGA\_KICH
* TCGA\_KIRC
* TCGA\_KIRP
* TCGA\_LIHC
* TCGA\_LUAD
* TCGA\_LUSC
* TCGA\_MESO
* TCGA\_OV
* TCGA\_PAAD
* TCGA\_PRAD
* TCGA\_READ
* TCGA\_SARC
* TCGA\_STAD
* TCGA\_THCA
* TCGA\_THYM
* TCGA\_UCEC
* TCGA\_UCS


# API documentation

**Architecture pattern.** K Pro exposes two programmable surfaces: a REST API for synchronous job submission, status polling, and result retrieval (structured JSON with provenance), and an MCP server endpoint that lets external agents discover and invoke K Pro platform directly.

Long-running work (e.g. literature mining, multi-omic queries, federated analyses) runs asynchronously through the orchestrator. Submissions return a job ID; completion may be delivered via webhook callback or polled status.

**Representative use cases.**

* **Batch target scoring.** A differential-expression pipeline submits a gene list; the orchestrator runs the target-evaluation agent over each entry and returns a ranked JSON payload with confidence and evidence links.
* **Pipeline step.** Nextflow or Snakemake stages call K Pro as a typed REST node, blocking on a webhook callback before downstream steps proceed.
* **Embedded reasoning.** A customer's own agent calls K Pro tools over MCP to compose target-evaluation reasoning into its own workflow without round-tripping data through a UI.


# MCP & Agent tools

**Architecture pattern.** To maximize reasoning and agentic interoperability, K Pro integrates with customer tooling through MCP configurations rather than bespoke per-vendor connectors. Customer systems (e.g. ELN, LIMS, data lake, BI) are connected to K Pro as MCP servers, either Owkin-provided wrappers for common tools or the customer's own servers. Authentication and per-tool authorization flow through OIDC/SAML or service accounts/API keys (especially for bulk data querying where it may be more efficient).

Data stays in customer systems; K Pro queries in place. New integrations are MCP-server additions.

**Representative integrations and complementary functionality:**

* **Benchling, Dotmatics (ELN/LIMS).** Pull assay and entity data on demand into a reasoning session; push predictions and annotated records back through the same MCP surface.
* **Snowflake, S3, Databricks, GCS.** Read-only querying with customer credentials; no duplication into K Pro. The customer may run data processing components to augment their data with Owkin and industry standard models, or render it more compatible with agentic reasoning.
* **KNIME, Pipeline Pilot.** K Pro callable from REST or Python nodes as a pipeline step.
* **Spotfire, Tableau.** K Pro writes structured outputs back to the customer warehouse; existing dashboards refresh against those tables.


# SSO and authentication

To streamline access across all integrated tools, K Pro supports SAML 2.0 and OIDC for Single Sign-On (e.g. Okta, Azure Entra ID, AWS Cognito). This ensures that user permissions and data access controls are consistently maintained whether the user is accessing the platform directly or interacting with data via Benchling, Spotfire, or similar.

Permissions are enforced consistently across direct K Pro access and MCP-mediated tool calls.

Where necessary or preferred, API keys or service accounts may be used.


# Model governance


# Trust and verification

K Pro is engineered with multiple layers of safeguards to ensure scientific rigor, transparency, and accountability. These measures work together to minimize errors while empowering researchers to critically evaluate AI-generated insights.

### Evidence Provenance and Traceability

Every conclusion generated by K Pro is anchored in evidence from authoritative sources, including PubMed literature and validated biological knowledge bases. The platform maintains complete provenance for all outputs, allowing users to trace recommendations back to their original sources. Comprehensive logging captures every data source, model decision, and reasoning step, creating a fully auditable trail of the analysis process.

### Technical Safeguards Against Hallucinations

K Pro implements several technical mechanisms to combat the hallucination risks inherent in large language models:

* **Retrieval-Augmented Generation (RAG)**: A RAG system verifies the existence and relevance of cited PubMed articles, ensuring that literature references are genuine and pertinent to the scientific question
* **Tool-based grounding**: K Pro's architecture relies heavily on specialized tools—including modality-specific AI models and data query systems—whose proper execution is continuously monitored
* **Data-anchored analysis**: One of K Pro's core strengths is its foundation in real patient data. Analyses and visualizations are generated from actual queried datasets, making it impossible to fabricate results

### The Role of Human Expertise

While K Pro incorporates robust verification mechanisms, we acknowledge that no AI system is infallible. Scientific oversight and expert review remain essential components of responsible research. K Pro is designed to augment—not replace—human expertise and decision-making. Like any scientific tool, its outputs require thoughtful interpretation and validation by qualified researchers.This multi-layered approach to trust and verification enables K Pro to maintain the highest standards of scientific integrity while providing researchers with transparent, accountable AI assistance.


# Model versioning

\[REVIEW NEEDED] — Specific model version numbers, release dates, and changelog for K Pro agents and underlying models not documented in sources. Model versioning policy needed.

K Pro's performance is fundamentally influenced by the underlying large language model's capabilities, particularly its ability to accurately interpret user questions and execute appropriate tool calls. Industry benchmarks consistently demonstrate that newer LLM versions deliver superior performance on tool-calling tasks, as evidenced by [agent performance leaderboards](https://galileo.ai/blog/agent-leaderboard-v2).

#### Model Transitions and Optimization

Transitioning between different LLMs or upgrading to newer versions requires careful recalibration of the system. Each model has distinct characteristics that necessitate adjustments in prompting strategies and context engineering to achieve optimal results. Our evaluation automation framework assesses these configurations to ensure that each LLM integration meets K Pro's performance standards.

As models evolve and improve, K Pro benefits from enhanced reasoning capabilities, more accurate tool selection, and better interpretation of complex scientific queries—ultimately leading to more reliable and relevant outputs for researchers.


# Explainability and traceability

### Ensuring citation and reference integrity

K Pro leverages Retrieval-Augmented Generation (RAG) on PubMed abstracts to provide accurate and verifiable citations. A key component of this system is the validation that all returned PubMed IDs correspond to actual published articles. While a minimal risk exists that the language model may generate responses that don't fully align with the retrieved article content, this probability is kept low through our RAG architecture.

Our commitment to citation accuracy extends beyond basic validation. We conduct internal evaluations against established public benchmarks for literature review tasks, and we continuously refine our daily evaluation protocols. Future enhancements will include more sophisticated analysis to verify that generated answers appropriately incorporate and reflect the content of retrieved PubMed IDs.

### Measuring and preventing hallucinations

K Pro implements comprehensive monitoring systems designed to detect and mitigate hallucinations at multiple stages of the response generation process.

**Tool call accuracy monitoring:** Daily automated tracking uses metrics such as Tool Call Accuracy (TCA) to measure how frequently the system correctly identifies and invokes the appropriate tools. This monitoring enables early detection of systemic issues, including cases where the agent incorrectly requests a tool, fails to recognize when a tool is necessary, or selects a suboptimal tool for the task at hand.

**Parameter validation:** Correct tool selection alone is insufficient—the parameters passed to those tools must also be accurate and complete. When incorrect parameters are supplied, the resulting actions can produce erroneous outputs that appear as hallucinations in the final response. To address this, we continuously monitor parameter accuracy and completeness through automated testing against a carefully curated set of evaluation questions, ensuring that tool invocations are not only appropriate but also properly configured.

***

**Logging & audit trail.** All events are recorded in the audit trail. K Pro uses end-to-end request tracing. Logs are structured and include correlation fields to support investigations and auditability. LLM usage is logged and tool-call sequencing is traceable at the orchestrator level.

**Access controls (for audit/log systems).** Observability systems are restricted to authorized personnel, behind internal network controls and authenticated via corporate identity.

**Traceability of AI recommendations.** Every AI-generated answer can be traced to:

1. The user prompt / session context
2. The tools invoked and data services queried
3. The step-by-step orchestration trace

Tracing is segregated by customer and is not used for model training without explicit consent from the customer.


# On K Pro skills

K Pro does not offer a single static way of discovering and prioritizing new target hypotheses. Instead, it provides a broad spectrum to enable both rule-based discovery/ranking of target lists and AI-driven ("self-drive") modes. Much effort of the Owkin team has been invested in deriving features from multimodal data, enabling the representation of targets in the respective feature space, and the relative ranking of targets in respective representations. On top of that, Owkin has developed different workflows (implemented as agentic skills) that utilize these features to generate ranked lists of novel target hypotheses. Most of these workflows aim to at the least cover typical rules used at pharma for target discovery, such as the ADC target discovery exemplified above. In addition to biological evidence, K Pro's workflows also consider competitive landscapes regarding each target hypothesis, population sizing (relevant for market sizing), druggability/tractability, assayability, and other, context dependent factors (e.g. tumor essentiality / addiction in the case of cell-intrinsic, i.e. oncogenic signaling-related targets). Notably, K Pro's workflows are editable/extendable/editable by the user, requiring no coding expertise due to the fact they are written in natural language.

In addition to the rules-based approach, K Pro is currently being developed in the direction of a self-driving AI scientist: after light prompting by the user, it uses data and prior knowledge to generate hypotheses and iteratively engages more and orthogonal evidence sources to deepen, reject or prioritize those, generating structured reports for pharma scientists. In these autonomous discovery campaigns, K Pro taps on the entirety of knowledge, data, tools, and agentic skills to build the workflows in run-time. Importantly, the autonomous AI scientist capabilities are not meant to replace but to complement K Pro's rule-based discovery capabilities to offer the user a broader spectrum of capabilities and user control.

**K Pro's scoring is hybrid and skill-based, not a single black-box ML model.**

* **Skills are the unit of methodology.** A target prioritization skill (e.g. "ADC target prioritization for indication X") is a versioned package. It defines which features to compute, which data sources to query, which tools and models to call, and how to aggregate the results into a ranking. Skills are composable and inspectable. The methodology is auditable.
* **Per-skill features** handle things like differential expression, healthy-tissue expression rank, malignant-cell specificity score, gene essentiality from DepMap, HistoPLUS-derived cell-type composition from H\&E, spatial co-localization scores from spatial transcriptomics etc.
* **Aggregation uses weighted ranking and/or data-driven feature selection.** Recent internal work on ADC target prioritization used expert-based weighting given the scarcity of ground truth labels, but past work on small molecules implemented end-to-end data-driven weighting of the features and ranking of the targets. Weights are configurable per skill and per use case, and can be customized for customer strategic preferences (different weights for ADC vs. small-molecule vs. radioligand targets, for example).
* **User customization sits on top of this.** Users can adjust weights at run time ("downweight literature, upweight spatial heterogeneity"). They can add custom constraints ("exclude any target with moderate or higher CNS expression"). They can register customer-internal scoring dimensions through custom MCP servers. They can customize existing skills or author new ones.

As a consequence, we do not have one and only proprietary scoring algorithm that produces a calibrated success probability across all use cases. The scoring is the composition of explicit, inspectable skills. This is a strength for auditability and for working inside customer existing scientific frameworks. It is intentionally not a black box.


# AI ethics and oversight


# AI ethical framework

\[REVIEW NEEDED — The excluded marketing/SEO content contains ethical positioning, but detailed ethical framework documentation—principles, review process, decision criteria—was not found in technical


# External oversight

### External Expert Collaboration and Community Engagement

K Pro's development is shaped through active collaboration with external stakeholders across regulatory, academic, and clinical domains. We maintain ongoing engagement with regulatory bodies and bioethics groups to ensure our platform meets the highest standards of scientific and ethical rigor. The academic community plays a particularly vital role in K Pro's evolution. Biologists and researchers have been invited to participate as beta testers and early users, embodying our philosophy of "biologists building for biologists." This collaborative approach ensures that K Pro's capabilities and interface align with the real-world needs and workflows of scientific researchers.

### Bias Prevention and Equitable Performance

K Pro incorporates comprehensive bias mitigation strategies throughout its lifecycle, from initial development through ongoing deployment. During model development and testing, we conduct systematic bias audits that evaluate performance across diverse demographic and clinical datasets. This rigorous evaluation process ensures that recommendations remain robust and equitable across all patient populations.Our commitment to bias prevention extends beyond internal testing. We partner with leading academic and clinical collaborators who independently validate K Pro's performance in real-world settings. These partnerships provide critical external perspectives that help us identify and address potential disparities that might not be apparent in controlled testing environments.Post-deployment monitoring forms a crucial final layer of our bias prevention strategy. Continuous surveillance allows us to quickly detect and correct any new or unforeseen sources of bias that may emerge as the platform is used in diverse contexts. This proactive approach, combined with K Pro's transparency features that enable users to trace data sources and recommendation rationale, ensures ongoing accountability and equitable performance.


# Security architecture


# Security measures

We adhere to a 15-point Security Principle Framework that prioritizes proactive design, strict access controls, and resilience. The strategy is built on the philosophy that security is everyone's responsibility, compliance is merely a baseline, and systems must be designed assuming that breaches can occur (Zero Trust).

#### Pillar 1: Secure Architecture & Infrastructure

Focus: Building a hardened foundation that minimizes the blast radius of any potential attack.

* Security by Design & Defense in Depth: We do not rely on a single control. Security is integrated during the architecture phase to prevent costly rework, using layered defenses (e.g., App auth + Network ACLs).
* Zero Trust & Immutable Infrastructure: We trust nothing by default. Every request is verified regardless of origin. Infrastructure is deployed via code (IaC) rather than manual patches to prevent configuration drift.
* Resilience: We design systems to degrade gracefully, assuming failure is inevitable, and prioritize fast recovery (MTTR).

#### Pillar 2: Identity & Access Management

Focus: Ensuring only the right people and services have access to the right resources.

* Least Privilege: Access is restricted to the absolute minimum required for a role.
* Strong IAM: We enforce centralized identity management and Multi-Factor Authentication (MFA) to protect against credential theft.

#### Pillar 3: The Secure Development Lifecycle (SDLC)

Focus: Automating security to catch vulnerabilities before they reach production.

* Shift Left: Security testing (static analysis) happens early in the CI/CD pipeline, not just before deployment.
* Secure Defaults: Systems launch with the most secure settings enabled automatically (e.g., encryption on by default).
* Supply Chain Security: We actively scan and validate third-party dependencies and libraries to prevent upstream attacks.

#### Pillar 4: Visibility & Data Protection

Focus: Knowing what we have, protecting it, and watching it closely.

* Data Classification: Sensitive data (PII/PHI) is identified, tagged, and encrypted according to its risk level.
* Auditability & Monitoring: We implement comprehensive logging and real-time behavioral analytics to detect anomalies immediately.
* Incident Readiness: We don't just watch; we practice. Tabletop exercises ensure we are ready to respond to incidents effectively.

#### Pillar 5: Culture & Compliance

Focus: Making security a human norm rather than just a technical requirement.

* Shared Responsibility: Security is an organizational norm; engineers are trained to own the security of their code.
* Compliance as Baseline: We view regulatory requirements as the "floor," not the "ceiling," effectively going beyond what is legally required to ensure true safety.

For more information please visit our[ Trust Centre on Vanta](https://app.vanta.com/owkin/trust/qq8guymgbci1jnk49kjbc)

Email: <security@owkin.com>


# API documentation

tbd

\[REVIEW NEEDED — No API documentation currently available. API development is planned post-October release. Specific endpoints, authentication methods, rate limits, and SDK documentation not yet published


# Enterprise security

At Owkin, keeping your data secure is our highest priority. While much of our technology is developed and managed in-house, we also partner with select, highly reputable vendors who must meet our stringent privacy, security, and ethics standards. Each partner is carefully vetted through rigorous due diligence, including detailed security assessments and contractual requirements aligned with our own commitments.

To ensure the highest standards of information protection, we employ robust organizational and technical measures, conduct regular internal and external audits, and perform comprehensive Security Risk Assessments with every major change to our systems. When integrating large language models or other third-party components, we choose hosting options that guarantee privacy and confidentiality for all data and outputs. This privacy-first approach ensures full compliance with **GDPR** and **HIPAA** requirements.

Owkin is certified to **ISO 27001:2022** for information security and **ISO 13485:2016** for medical device quality, reflecting our ongoing dedication to safeguarding your data. With these measures in place, you can be confident that your information is protected at every stage.&#x20;

All data in K Pro is segregated by customer to ensure confidentiality. Access to data by Owkin employees is limited to those who have an operational role requiring maintenance access. System integrity and information security is maintained through multiple layers including 24/7 monitoring.

***

Owkin's platform architecture aligns with enterprise security assessment standards by being ISO 27001 certified since November 2021, regularly undergoing internal and external audits, and performing security risk assessments across the organization. Data is encrypted at rest (AES-256) and in transit, and third-party audits and penetration tests are conducted to validate security controls. Additionally, Owkin's cloud provider (AWS) holds certifications such as ISO 27001, supporting compliance with industry standards.


# Vendor management

### Technology Stack and Vendor Management

Owkin's technology infrastructure combines proprietary in-house development with carefully selected partnerships. While we maintain direct control over much of our core technology, we collaborate with highly reputable vendors who must demonstrate alignment with our stringent privacy, security, and ethics standards.

#### Vendor Vetting and Due Diligence

Every potential partner undergoes rigorous evaluation before integration into our technology ecosystem. Our vetting process includes comprehensive security assessments and detailed due diligence reviews. Contractual agreements with all vendors incorporate strict requirements that mirror our own commitments to data protection and security, ensuring consistent safeguards across our entire stack.

#### Security and Privacy Framework

Data security is paramount in K Pro's design and operation. We implement robust organizational and technical measures to protect information at every stage, supported by regular internal and external audits. Major system changes trigger comprehensive Security Risk Assessments to proactively identify and address potential vulnerabilities.When integrating large language models or other third-party components, we exclusively select hosting configurations that guarantee privacy and confidentiality for all data inputs and outputs. This privacy-first architecture ensures full compliance with both GDPR and HIPAA requirements, providing robust protection for sensitive health and research data.

#### Certifications and Standards

Owkin's commitment to security and quality is validated through internationally recognized certifications:

* **ISO 27001:2022** for information security management
* **ISO 13485:2016** for medical device quality management

These certifications reflect our ongoing dedication to maintaining the highest standards of data protection and operational excellence throughout K Pro's lifecycle.


# Privacy architecture

K Pro is designed with privacy and security at its core to protect your data:

* **Personal chat history:** Your chat history is securely stored and always associated with your user account and organization
* **Access control:** Only authenticated users can access their own chat history - no one else can view your conversations
* **Secure infrastructure:** Our database is hosted on a managed, secure cloud infrastructure with strict access controls and network policies
* **Data isolation:** Data uploaded to K Pro is only visible to you and members of your organization

Owkin's platform enforces data boundaries through a multi-account architecture, ensuring that each customer's data is isolated from others. Effectively, each instance of K Pro is a single-tenant deployment in an isolated account.

Owkin's platform enforces centralized authorization checks before any dataset/tool access using Role-Based Access Control (RBAC).

Customer data is accessed only through Owkin K services/tools (query-based + APIs/SDK); end users do not get direct DB access.

***

K Pro draws on several categories of knowledge sources to contextualize responses to users:

* **Underlying data assets analyzed by K Pro.** Held either in secured object storage or in connected data platforms (Snowflake, Databricks, etc.). The data model varies by asset and is preserved as-is from the source system; K Pro accesses these through purpose-built connectors rather than re-modeling the data centrally.
* **User information.** Stored in a relational database (RDBMS), capturing identity and account attributes such as name, email, and organization.
* **Interaction logs.** Logged in a relational database (RDBMS). Each entry records the user question, tool calls invoked, latency, timestamp, and supporting metadata required for observability and auditability.
* **Evaluation and observability telemetry.** Streamed to dedicated observability and evaluation platforms, where logs and traces support quality monitoring, regression testing, and incident investigation.

K Pro does not currently maintain a central knowledge graph / vector store as part of the core data model.

User traces can be accessed programmatically via a backend API to support observability. This API feeds the observability platform.


# Access controls and permissions

### Access controls and permissions

K Pro ensures the highest level of data security and confidentiality through comprehensive permissions management:

#### Data Segregation

All data in K Pro is segregated by customer to ensure complete confidentiality. Your data remains isolated and accessible only to authorized users within your organization.

#### Controlled Access

Access to customer data by Owkin employees is strictly limited to personnel with operational roles requiring maintenance access. Access is granted only for customer-focused activities and improvements, such as troubleshooting, performance optimization, and platform reliability. All access is logged and monitored to maintain transparency and accountability.

#### Security Monitoring

System integrity and information security are maintained through multiple layers of protection.


# Infrastructure and hosting

### K Pro deployment models

Owkin supports **multiple deployment models for K Pro**, including Owkin-hosted managed service (single-tenant SaaS), and dedicated cloud environment (BYOC). The platform can be deployed in public clouds (such as AWS, Azure and GCP). Custom deployment may involve coordinated IT workshops, appropriate contract sign-off, and alignment with the IT specifications required for each deployment model.

### Data privacy and security

For customers running K Pro on their own cloud subscription, we maintain strict separation of your data from our infrastructure. Your storage is set up in a cloud account that is entirely separate from the K Pro account. This ensures that your data remains under your control and protected within your own cloud environment.

To enable K Pro to access your data while maintaining this security boundary, we utilize secure cross-account access patterns, such as shared network configurations. This allows your storage to be securely mounted and accessed by K Pro without compromising the isolation between accounts, preserving both your privacy and security throughout the entire process.


# Legal


# Terms and conditions

**Effective Date: date of approval of the Terms and Conditions by the User or the Platform access by the User.**

*Version of the terms and conditions: v.4 date October 30, 2025.* Previous versions of the Terms and Conditions can be consulted [here](https://docs.owkin.com/terms-and-support/legal/terms-and-conditions/what-are-the-previous-versions-of-owkin-k-pro-terms-and-conditions).

**PREAMBLE**

The following Terms and Conditions apply solely to the use of the **K-Pro Free** version made available by Owkin. All other versions, editions, or products provided by Owkin — including paid, enterprise, or customized solutions — are governed by separate terms and conditions or contractual agreements, as applicable. Use of any other Owkin product or service constitutes acceptance of the specific terms governing that product or service.

“**K-Pro Free**” is the lite version of Owkin’s commercial product “K Pro”, an artificial intelligence platform dedicated to biomedical research (the “**Platform**”, as defined below) developed by Owkin Inc., a Delaware company having a business address at 185 Alewife Brook Parkway, #210, Cambridge, MA 02138, United States

The Platform enables its user (the “**User**”, “**You**”, as defined below) to query a conversational agent in natural language with the aim to help generate and validate scientific hypotheses through the exploration of multi-modal data with visualisation, scientific literature and specific public biological knowledge.

By accessing and using the Platform, You:

* acknowledge having read, understood and accepted these Terms and Conditions, without reservation, limitation or condition;
* represent and warrant that You (i) are duly authorized to accept these Terms and Conditions on Your behalf or on behalf of your employer if applicable (ii) are not under any restriction to access and use the Platform, and in particular that your status does not imply additional formalities to access and use the Platform; (iii) have informed ifYou work in an organisation (both for-profit or non-profit entities), obtained approval from - any appropriate board or authorities (including hierarchical authority if applicable) of Your access to the Platform and the Terms and Conditions (collectively the “**Criteria**”). You shall be able to provide all necessary documentation to demonstrate compliance with these Criteria to Owkin at the time of acceptance of these Terms;
* acknowledge that Access is granted based on Your declarations. As a reminder, the Platform is accessible to any researcher interested in biology, subject to fulfilment of the Criteria and the conditions herein, and Owkin will grant an Access to the User based on the declarations made by the User. However, Owkin reserves the right to verify whether the Criteria are met, and may suspend access to the Platform until fulfilment of the Criteria or terminate the access without additional formalities.

Please note that access to K Pro, meaning the premium version of the Platform is subject to specific additional business terms and conditions.

**PLEASE READ THESE TERMS AND CONDITIONS CAREFULLY BEFORE ACCESSING OR USING THIS PLATFORM. Any access to the Platform involves the irrevocable and unreserved knowledge and acceptance of these Terms and Conditions by the User. If You do not agree with these Terms and Conditions, You may not access and use the Platform. These Terms and Conditions may be amended from time to time and we recommend that You review these Terms regularly.**

**1 - Definitions**

For the purposes of these Terms:

“Access Period”:refers to the period during which the User is authorized to use and access the Platform in accordance with these Terms.

**“Agent”:**

refers to the conversational agent that analyses and interprets the Inputs and selects the right tool and dataset or combination of tools and datasets available on the Platform or accessible to the Agent with the relevant parameters to provide Outputs to the User.

**“Applicable Laws and Regulations”:**

refers to any laws, regulations guidelines and other requirements (regulatory or other) in any jurisdiction applicable to Owkin and/or the Users during its/their use of the Platform.

**“Commercial Use”:**

refers to any use intended for commercial advantage and/or monetary compensation, including without limiting the foregoing any research project conducted for and/or on behalf of an industrial and/or commercial third party and/or the provision of commercial services with (or using) Outputs, sales of Outputs (or products incorporating such Outputs, for instance commercialisation of a diagnostic kit containing a biomarker to the extent the biomarker was an Output), or marketing for commercial services on, or sales of, Outputs, (ii) developing the Outputs (for instance clinical trial on an asset which is the Outputs) with a view to directly commercialize such Outputs (or products incorporating such Outputs), and (iii) manufacturing an Output (or products incorporating such Outputs), with a view to directly commercialize such Outputs (or products incorporating such Outputs).

**“Data”:**

refers to the Public Dataset and Non-Public Dataset.

**“Documentation”:**

refers to the instructions for use of the Platform, which are available at the following link: [https://docs.owkin.com/user-guide/getting-started](https://owkinkhelp.zendesk.com/hc/en-us/articles/33938116119697-Owkin-K-User-Guide-Getting-started).

**“Inputs”:**

refers to the information, data, knowledge, content, intellectual reflection, idea formalization, displayed, uploaded or provided in any manner whatsoever to the Platform by the Users in particular while making queries.

**“Intellectual Property Rights” or “IPRs”:**

refers to trademarks, patents, designs and models, copyrights, trade names, trade secret, signs, logos, graphics, domain names all rights of whatsoever nature in computer programs, codes, databases, algorithms, Platform, data (including Outputs and Inputs), and any other signs and intellectual property rights owned by Owkin and/or a User, whether or not protectable under Applicable Laws and Regulations, whether registered or unregistered, including all granted registrations and applications of the same

**“Login Information”:**

refers to a unique identifier assigned to a User which, combined with a password, enables the User to authenticate himself or herself to access the Platform.

**“Outputs”:**

refers to any analysis, predictions, recommendations, decisions, content (including, texts, images, photos, videos, graphics), results, under any format, generated by the Platform to answer a query of the User.

**“Owkin Proprietary Knowledge”:**

refers collectively to the Platform, any proprietary content of Owkin uploaded or made available by Owkin on the Platform, its Documentation and the corresponding IPRs.

**“Platform”:**

refers to the systems operated and made available by Owkin to the User under these Terms, collectively branded as “K-Pro”, which include the Agent, Software and any Data that could be used for the generation of the Outputs. The Platform made available under these Terms is named K-Pro Free and is the lite version of K Pro.

**“Non-Public Dataset”:**

refers to anonymized, de-identified and/or pseudonymized data used by or via the Platform, which is under Owkin’s control, that could be accessed by the User under certain restricted conditions, depending on the version of the Platform, and used to generate Outputs.

**“Public Dataset”:**

refers to the anonymized, de-identified data and/or pseudonymized data that have been made available to the public by a third-party natural person or entity, potentially under specific terms of use, that could be accessed and used by the User via the Platform to generate Outputs.

**“Software”:**

refers to any public or private software, algorithms and/or libraries used by Owkin to enable Users to access and use the Platform.

**“Terms and Conditions” or “Terms”:**

refers to these Terms and Conditions and their Annexes.

**“Territory”:**

refers to the countries, listed in the Annex B, from where the User can access and use the Platform, subject to the terms of the Section 8.

**“User”:**

refers to any natural person over 18 years of age with access to the Platform authorised by Owkin and, where applicable, its employer and/or any appropriate board or authorities.

The definitions referred to in this Section apply to both singular and plural terms.

**2 - Purpose of the Terms**

The purpose of these Terms is to set out the rights and obligations of Owkin and the User regarding the use and access to the Platform and its content during the Access Period. These Terms apply each time the User is accessing the Platform.

**3 - Right of access - Use of the Platform.**

**3.1 -** Subject to the limitation set forth in these Terms, Owkin grants to any User, a personal, limited, for free, non-exclusive, non-transferable, non-sublicensable, revocable, in the Territories (and without prejudice to Section 8), right to access and use the Platform, and the Documentation during the Access Period. The rights set forth in this Section 3.1 are subject to the compliance with the Criteria and these Terms during the Access Period. Owkin may suspend User’s access to the Platform, in whole or in part, at its sole discretion, if these Criteria are not (or no longer) satisfied, and access will be reinstated upon User’s fulfilment of the Criteria. In case of breach of the Terms, Owkin could also terminate the access at any time without any prior notice to the User.

**3.2 -** The above mentioned right of use excludes any right to use the Platform (i) on human subjects, including for research involving human beings (including without limitation in the context of clinical trials and/or any other kind of biomedical research which, all or in part of the purpose, is to develop and/or compare and/or assess the performance of medical devices and/or providing prognostic, diagnosis and/or therapeutics purposes), or in routine, for care or diagnostic purposes, or (ii) with or on behalf of third parties or (iii), if the User is an employee under public or private contract with its employer, in connection with the missions assigned to the User by its organisation or with its organisation’s data except as specifically accepted by its employer informed of the Terms and in particular the absence of confidentiality of Inputs and Outputs and related ownership and/or (iv) for Commercial Use.

**3.3** - The User may access the Platform via the Internet. The User shall use the Platform at his or her own risk and under his or her sole responsibility. The User is aware that the Platform is hosted by a “cloud” provider and made available to the User via the Platform dedicated website or a link from Owkin’s website to the Platform website.

**3.4 -** When accessing the Platform for the first time, the User will register to the Platform by entering the information requested by Owkin, accepting these Terms and if any, providing all necessary document demonstrating compliance with the Criteria as described in the Preamble, and then at each connection, the User shall use the Login Information that Owkin provided so that he or she can connect to the Platform. The Login Information is strictly personal and confidential.

**3.5 -** The User expressly undertakes to keep his or her Login Information confidential and not to communicate it to any third parties. The User shall bear all consequences that may result from the voluntary or involuntary disclosure of its Login Information, and Owkin shall not be held liable for any use of the account by a third party who has gained access to the User's Login Information, in any manner whatsoever. The User shall be solely liable to Owkin for all use of the Platform in violation of these Terms. The User undertakes to immediately inform Owkin of any suspicious or fraudulent use of his or her Login Information.

**4 - Public and Non-Public Datasets**

**4.1** – The User can only use the Public Datasets through the Platform, subject to the provisions set forth below. Except as specifically identified in the list of Public and Non-Public Datasets in Annex A of these Terms, no Non-Public Datasets will be made available to the User except if the User accesses another version of the Platform under specific additional terms.

**4.2 -** The Public and Non-Public Datasets used by the Platform are listed in Annex A of these Terms. If the User wishes to directly access and use any Public or Non-Public Dataset, the User will access it under his or her own responsibility and under the terms of any license potentially associated with such Dataset. Furthermore, the User understands and agrees that:

* Use and access to the Public and Non-Public Dataset is made at his or her own risk and shall be done in compliance (i) with the Applicable Laws and Regulations; (ii) the specific terms and conditions of use of each Public and Non-Public Dataset.
* Owkin cannot guarantee that the (i) anonymization, de-identification and/or pseudonymization of the Public and Non-Public Datasets will have the same meaning or requirements under any Users’ Applicable Laws and Regulations (ii) the Outputs is only composed of data that never permits individualisation of patient or research participant by the User.
* Owkin, at the Effective Date, has performed a good faith check of the Public and Non-Public Datasets’ terms for their material alignment with the non-Commercial Use of the Public and Non-Public Dataset on the Platform only.
* User is solely responsible to (i) verify if his or her use of the Public and Non-Public Datasets is compliant with (a) the Datasets’ licences (b) the Applicable Laws and Regulations applying to such User, and (ii) where applicable, to comply with any additional formalities.

**4.3** - The User acknowledges and agrees that:

* Data hosted on the Platform must remain within the Platform and should not be exfiltrated or transferred outside without proper written authorization of Owkin or in application to the licenses applying to this Data;
* it shall not engage in any activities that facilitate or attempt to facilitate the unauthorized extraction, copying, or removal of content, including Data, from the Platform. For clarity, taking a screenshot, a picture and/or a video of the Platform is not allowed.

**4.4 -** The Platform may use or be used in connection with third-party software, product, services or integration (“**Third-Party Services**”). Owkin will determine in its sole discretion which Third Party Services may be used by the Platform or accessible to the Users via the Platform. Notwithstanding any contrary provisions, use of Third-Party Services is at User own risk and subject to User’s compliance with any terms, conditions or policies applicable to such Third-Party Services. Owkin does not control or accept any liability for claims and resulting liabilities, damages, losses and expenses, including reasonable attorneys' fees, arising out of or resulting from use of Third-Party Services.

**5 - Obligations of the User**

**5.1 -** The User acknowledges being familiar with the purpose, functionalities and operating procedure of the Platform and any content accessible to this User, having ensured that it conforms with its needs and skills, in particular based on the information provided in the Documentation.

**5.2 -** The User may contact Owkin, by email to [support@owkin.com](mailto:Product-support@owkincom) or by completing and sending the dedicated form on the Platform, for any types of requests related to the Platform such as service requests, errors and other operating anomalies, as well as all suggestions for improvement that he or she deems desirable.

**5.3 -** The User undertakes to:

* use the Platform and its content in accordance with the provisions of these Terms, as of the Documentation and of the Applicable Laws and Regulations;
* use the Platform and its content in accordance with their intended purposes and in a manner compatible with reasonable use.

**5.4 -** The User shall refrain from using the Platform and its content for illicit or illegal purposes. The User undertakes, in particular to not: (i) lease, sell, sub-license, assign, grant access to or otherwise transfer the Platform nor its content (in particular the Software and the Agents) and/or the Documentation to any third party for any purpose whatsoever, including for evaluation or testing purposes; (ii) decipher, decompile, disassemble, reverse assemble, modify, translate, reverse engineer or otherwise attempt to derive source code, algorithms, tags, specifications, architectures, structures or other elements of the Platforms, in whole or in part; (iii) attempt to re-identify, by any way whatsoever, the data subjects from which the Data may come or to investigate the origins of populations and/or make research ancestors ; (iv) represent that the Outputs has been generated by a human being while this not the case; (v) use the Outputs to generate any technology that competes with all/or part of the Platform.

**5.5 -** The User is solely responsible for the Inputs and any information provided to Owkin and/or the Platform. The User may not provide or make available on the Platform in any manner whatsoever any Inputs, feedback or any other information or content that is (i) false, inaccurate, not up-to-date or incomplete; (ii) offensive, defamatory, violent, illegal, obscure or contrary to public order, to morality or to Owkin’s or its affiliates’ interests or in breach of Applicable Laws and Regulations; (iii) directly or indirectly identifying a data subject in particular a patient and/or a participant to a research involving human beings.

**5.6** - If the User were to find that it has access to Data or Outputs, while using the Platform, that can be directly attributed to a specific data subject, it will immediately (i) stop accessing such Data or Outputs; and (ii) within a delay which shall not exceed twenty-four (24) hours notify Owkin by email to [support@owkin.com](mailto:Product-support@owkincom) adding <dpo@owkin.com> in copy.

**5.7** - The User shall not damage, interfere with or disrupt the Platform, including circumvent any rate limits or restrictions or bypass any protective measures or safety mitigations Owkin puts in the Platform, including for example overloading the Platform, introducing viruses or malware, spamming or DDoSing Services.

**5.8 -** The User undertakes to provide regular feedback on his or her use of the Platform, which may include ideas, suggestions for improvement and ratings requested by Owkin on the Platform through Thumbs Up or Thumbs Down or through discussions with the support team (“**Feedback**”), to help Owkin improve the Platform as stated in Section 6.4. Access to the Platform is conditional upon the provision of the User’s Feedback, which may be used or not by Owkin, and - in any case - shall never be conditioned to any further obligation or payment in return.

**6 - Intellectual Property**

**6.1 - Owkin Proprietary Knowledge**

**6.1.1 –** As the owner and/or guardian of the Owkin Proprietary Knowledge, Owkin guarantees that it can grant access to the Platform in the Territory and for the intended purpose of the Platform.

**6.1.2 -** Owkin Proprietary Knowledge is and remains wholly owned or controlled by Owkin. Any improvement of Owkin Proprietary Knowledge, its associated content and all Intellectual Property Rights attached thereto (including their previous, current and future developments, updates and versions) is and remains owned or controlled by Owkin. The User’s right to access and use the Platform granted under these Terms does not entail any transfer and/or assignment of any Intellectual Property Right attached to the Owkin Proprietary Knowledge.

**6.1.3** - The User acknowledges that Owkin may freely use any information related to the use of Owkin Proprietary Knowledge that it receives as Feedback from the User. The User shall refrain from claiming any IPRs, to said Feedback, or to any development or improvement that Owkin and/or its affiliates could make in connection with User’s Feedback either during or after the term of the Access Period.

**6.1.4 -** The User shall (i) respect Owkin’ IPRs; (ii) refrain from directly or indirectly infringing Owkin’s ownership rights, title and interest to Owkin Proprietary Knowledge and any IPRs attached thereto, in any manner and by any means whatsoever; and (iii) not delete or alter the copyright and/or proprietary notice appearing on the Platform, its content and any Documentation of any type on which this notice features.

**6.1.5 -** Except as expressly provided in the Article 6.3.4, the User is not allowed to use the name, trademark and logo of Owkin (including K-Pro Free and/or K Pro name(s), trademark(s) and logo(s)).

**6.2 - Inputs**

**6.2.1** - The User is and remains the owner and/or guardian of any content that he or she provided on the Platform, including Inputs.

**6.2.2** - The User grants Owkin a worldwide, non-exclusive, transferable, sub-licensable, free of charge licence to use, reproduce, represent and exploit the Inputs for the improvement of Owkin Proprietary Knowledge, the Platform, the Documentation and any other services or products of Owkin or its affiliated companies, perpetually or where applicable for the duration of the intellectual property rights.

**6.2.3** - The User represents and warrants that it holds all the necessary rights, consents and/or authorizations to provide and use the Inputs on the Platform and grant the rights set forth in Section 6.2.2.

**6.2.4** - The Platform is not designed to store User’s or third-parties’ confidential information and Owkin does not guarantee the confidentiality of the Inputs. The User shall not provide his or her confidential information nor the one of his or her employer, as applicable, within Inputs.

**6.2.5 -** The User is solely responsible for the storage and back up copy of the Inputs, specifically before closing his or her browser. Owkin shall not be held liable for any User’ claims related to the copy and/or transfer of the Inputs.

**6.3 - Outputs**

**6.3.1 -** Owkin does not claim any rights, interests, or title on the Outputs which are provided “as is” to the User, with the limitations described in Section 7.2. For clarity, such Outputs may be used by the User, at its sole risk, for internal research purposes only (excluding any Commercial Use). The User grants Owkin with a non-exclusive, worldwide, royalty-free, irrevocable, transferable and sublicensable license to use, reproduce, represent and exploit the Outputs for the improvement of Owkin Proprietary Knowledge, the Platform, the Documentation and any other services or products of Owkin or its affiliated companies, perpetually or where applicable for the duration of the intellectual property rights.

**6.3.2 -** The User understands and agrees that the Outputs:

* may not be accurate and/or not an exact representation of the reality and therefore needs to be independently validated by the User by following the relevant and applicable scientific, ethical and legal practice;
* do not reflect Owkin’s opinions, analysis and/or interpretations of the Inputs;
* may not be unique and other users may obtain similar Outputs, by using the Platform and/or any similar other platforms and/or tools (if any). Consequently, the User acknowledges that the rights granted over the Outputs, in accordance with the Article 6.3.1, does not extend to the Outputs obtained by other Users or third parties using the Platform and/or any similar other platforms and/or tools (if any).

**6.3.3 -** The User is solely responsible for the storage and back up copy of the Outputs, specifically before closing his or her browser and/or before the term of the Access Period. Owkin shall not be held liable for any User’ claims related to the copy and/or transfer of the Outputs.

**6.3.4 -** In consideration for the access to the Platform granted by Owkin to the User, the User undertakes to credit Owkin in any scientific publication referring to the Outputs obtained in connection with the use of the Platform, as follows: “*The results \<published or shown> here are in whole or part based upon results generated by using K-Pro, a Platform by Owkin*”.

**6.4 - Feedbacks**

**6.4.1 -** By accepting these Terms, the User grants Owkin with a non-exclusive, worldwide, sublicensable, transferable, royalty-free, and irrevocable license to use, reproduce, represent and exploit his or her Feedback for the improvement of Owkin Proprietary Knowledge, the Platform, the Documentation and any other services or products of Owkin perpetually or for the duration of the applicable intellectual property rights.

**7 - Warranties and exclusions**

**7.1 -** The User acknowledges that its use of the Platform is voluntary and will not be subject to any remuneration in any form whatsoever.

**7.2 -** The Platform, the Documentation, the Data and the Outputs are made available “as is” and Owkin does not provide or make any representation or warranty, either express or implied, including but not limited to (i) quality, fitness or suitability for any particular purpose, including to review and validate the scientific hypothesis of the User; and/or (ii) non-infringement of third parties’ rights; (iii) absence of any errors, defects, design flaws or other problems (iv) accuracy or reliability of the Platform and/or Outputs; (v) uninterrupted, timely or secure provision of the Platform.

**7.3 -** Owkin shall not be held liable for material damages caused by the Platform and/or the Data to the User and/or any third parties. The User is solely responsible for its use of the Inputs, Platform, Documentation, Data and the Outputs and shall bear all risks associated with this use, including but not limited to: instability, malfunctioning, loss of content and in particular Inputs and Outputs, and other damage or loss.

**7.4 -** To the fullest extent permitted by Applicable Laws and Regulations, Owkin shall not be held liable to the User for any direct or indirect damage or loss whatsoever (such as commercial or financial damage or operating losses that may affect the User) resulting from (i) any use of, or inability to use, the Platform, the Documentation, the Outputs and the Data or (ii) the use or reliance on any information or results, including Outputs, displayed or available in the Platform. The User shall indemnify and hold harmless Owkin and its affiliates, and their personnel, agents and representatives, and successors and permitted assigns, against any and all third party claims and resulting liabilities, damages, losses and expenses, including reasonable attorneys' fees, arising out of or resulting from (i) negligence or wilful misconduct in connection with these Terms, or (ii) a breach of these Terms, each by the User and, when its legal person, its affiliates, personnel, agents or representatives.

**7.5 -** Owkin shall not be held liable in the event of legal proceedings initiated by third parties against the User due to the illicit or illegal use of the Platform, the Documentation, the Data, the Inputs and the Outputs.

**7.6** - Nothing in these Terms limits or excludes any other warranty or liability that cannot be limited or excluded under Applicable Laws or Regulation, or death or personal injury caused by negligence.

**8 - Export control**

The User may not export or provide access to the Platform to persons or entities or into countries or for uses where it is prohibited under U.S., European Union Law, Applicable Laws or Regulations or other applicable international law. Without limiting the foregoing sentence, this restriction applies (a) to countries where export from the US, and/or the European Union or into such a country would be prohibited or illegal without first obtaining the appropriate license, and (b) to persons, entities, or countries covered by U.S. sanctions and/or European Union sanctions.

**9 - Transparency**

The User is not obligated, directly or indirectly, to prescribe, purchase or recommend the prescription or the purchase of any products manufactured or distributed by any of the entities of Owkin, or to make any referrals of patients to such entities, to otherwise generate purchases for or direct any purchasers to such entities. The access rights to the Platform granted to the User is not intended to compensate the User for past, present, or future prescriptions, referrals or recommendations of Owkin products.

**10 - Term and termination**

**10.1 –** You have the right to access and use the Platform from the Effective Date to the termination of the Terms pursuant to Sections 10.2 and 10.3. You may also choose to terminate the application of the Terms at any time by ceasing to access and use the Platform and requesting the deletion of your account by sending a request to [support@owkin.com](mailto:Product-support@owkincom). The Terms will apply to you until the complete deletion of your account by Owkin.

If your employer is interested in the premium version of Owkin’s Platform: K Pro, please contact us by email.

**10.2 -** Owkin reserves the right to suspend or terminate the Platform and/or any of its features, at any time, without having to bring the case before the competent court, at Owkin’s discretion, without any justification, nor prior notice.

**10.3** In case of suspected or actual breach to these Terms by the User, Owkin reserves the right to immediately, without prior notice, without having to bring the case before the competent court: (i) delete any breaching content, including Inputs or Outputs ; (ii) suspend or restrict access to the Platform, in total or in part; and (iii) suspend or cancel the User account at its own discretion; without prejudice of any other damages Owkin may claim.

**10.4** In case of termination, the User account will be deactivated and archived for a duration necessary for the applicable statute of limitations, if needed, for evidence purposes, and the applicable legal retention periods. The Inputs and Outputs are not stored into the User account and will not be provided to the User at the end of the Terms; it is the User’s sole responsibility to ensure that such Inputs and Outputs are regularly saved by the User following their generation.

**11 - Personal data protection**

**11.1 -** The Privacy Policy available [here](https://www.owkin.com/policies/privacy-policy) explains how Owkin is processing personal data of the User in connection with their access and use of the Platform. The Cookies Policy available [here](https://k.owkin.com/legal/cookie-policy) explains how and for which purposes the Platform may include technical devices (cookies or other tracking technologies) that allows sending to Owkin and its processors information about the use and navigation of the User. The Privacy Policy and Cookies Policy can be updated from time to time and do not form part of these Terms.

**11.2** Owkin and the User recognize that they have full and entire knowledge of the obligations under the regulation applicable to personal data provided under any provision of a legislative or regulatory, European or national nature, resulting in particular from Regulation 2016/679/EU of 27 April 2016 on the protection of individuals with regard to the processing of personal data and on the free movement of such data as well as any other regulations applicable in this field that may be added to or subsequently replaced by it ("**Personal Data Regulation**"), that may be applicable to them, in their respective capacity as independent data controllers. To the extent that the Personal Data Regulation would be applicable, Owkin is the data controller of the personal data processing of the User’s personal data of the Platform, including the contact details and any information resulting, automatically or not, from the use and navigation on the Platform of the User. The Data does not include personal data relating to EU data subjects. In this regard, each Party shall take, for its own processing activities, any appropriate measures to ensure the compliance with this Regulation. To the extent that personal data transfers should occur, these Terms incorporate the European Commission Standard Contractual Clauses as made available by the European Commission on its website in a non-modifiable form available here: [Standard Contractual Clauses](https://eur-lex.europa.eu/eli/dec_impl/2021/914/oj?uri=CELEX:32021D0914\&locale=en), and completed as set forth in Annex C (the “**Clauses**”).

**12 - General provisions**

**12.1 - Changes to the Terms.** Owkin reserves the right to make changes to these Terms at any time. Any change will come into force at the time of new Terms availability on Owkin or the Platform dedicated websites. Owkin advises the User to consult frequently such websites to be informed on the new Terms. If the User disagrees with the new Terms, User shall stop using the Platform.

**12.2 - Assignment**. The User shall not assign all or part of the Terms to a third party without the prior written consent of Owkin. However, Owkin may assign these Terms to its affiliates and/or any third party, notably in the event of a change of control of Owkin and/or its affiliates.

**12.3 - Force majeure**. Owkin may not be held liable for any failure or delay in the performance of their obligations under these Terms and Conditions resulting from circumstances that are unforeseeable, beyond its reasonable control, and which cannot be overcome despite its reasonable endeavors (“**Force Majeure Event**”). However, under such circumstances, Owkin undertakes to notify the Force Majeure Event to the User as soon as possible.

**12.4 - Independence of the parties**. The parties acknowledge and agree that they may not under any circumstances make a commitment in the name of and/or on behalf of one another. Under no circumstances shall these Terms be interpreted as creating a binding relationship or de facto partnership between the parties, each of which shall be considered an independent co-contractor.

**12.5 - Relations between the parties.** The parties acknowledge and accept that their collaboration can in no case be considered as a de facto company, a joint venture, or any other situation entailing a representation of reciprocity or solidarity towards their respective creditors.

**12.6 - Entire agreement.** These Terms express the entirety of the obligations of the parties and supersedes any previous agreement between the parties relating to the subject matter hereof.

**12.7 - Non-waiver**. If, in the event of a breach by either party of its obligations under the Terms, the non-defaulting Party does not waive its rights resulting from said breach. In no event, the failure to exercise its rights shall be construed as a waiver of the exercise of said rights, either in the future, or in the event of a similar breach by the defaulting party of its obligations under the Terms.

**12.8 - Severability.** If any provision of the Terms is held to be invalid or unenforceable pursuant to a prevailing rule of law or a final court decision, the provision shall be either amended to render it valid or deemed removed, without this leading to the invalidity of these Terms or affecting the validity of its other provisions, which shall retain their full force and scope, provided that the balance of these Terms remains unaltered.

**12.9 - Applicable law, dispute resolution.** These Terms are subject to the laws of the State of New York. In the event of a dispute relating to the interpretation, performance or validity of the Terms, express jurisdiction is attributed to the competent court within the jurisdiction of the State of New York.

\*\*\*

**Annex A: List of Public and Non-Public Datasets**

The list of Public and Non-Public Datasets currently available in the K-Pro Platform, along with their respective licenses and versions, is available [here](https://owkinkhelp.zendesk.com/hc/en-us/articles/34965323359377-What-is-the-list-of-public-and-private-datasets-available-in-Owkin-K-Navigator).

**Annex B: Territories where the access and use of the Platform is allowed**

Any country except for the countries subject to sanctions and/or restrictions notably in application of the Export Control laws and regulations.

Without limiting the foregoing, the access and use of the Platform is formally prohibited in the following countries: Afghanistan, Yemen, Mali, Syria, Somalia, Central African Republic, South Sudan, Libya, Bangladesh, North Korea, Russia, Iran, Egypt, China, Belarus, Indonesia, Taiwan.

**Annex C – Standard Contractual Clauses**

In the event the User is based in the EU, Owkin and the User agree that these Clauses apply to them as such, depending on the scope that relates to them and for the purposes of the Clauses, the following specific terms have been agreed.

1. The Parties to these Clauses are Owkin as the importer and data controller (DPO contact details: <dpo@owkin.com>) and the User as the exporter and the data controller (with the relevant details of both parties as set out in the Terms) (Appendix I of the Clauses).
2. These Clauses are signed by the Parties on the Effective Date and as part of their acceptance of the Terms (Appendix I of the Clauses).
3. The Clauses apply with respect to the data transfer outside the EU for the activities of providing the Platform to the User, and covering the following transfer (Appendix I of the Clauses):

* Data subjects: User.
* Data Categories: User identification and contact details as well as professional information, usage and navigation data.
* Duration of the transfer: for the entire duration of the Terms and for the duration identified in Owkin’s Privacy Policy following the Terms.
* Nature of the transfer: use, management, transmission and storage.
* Purpose of the transfer: performance of the Terms.
* Data retention periods: until the end of contractual relationship under the Terms and, if necessary, archived for the applicable prescription period.
* The technical and organizational measures implemented by the data importer are described in the Terms (Appendix II of the Clauses).

1. The nature of the data transfer outside the EU is a transfer Controller to Controller (Module 1).
2. Owkin and the User agree to opt-out of the “Docking clause” (Clause 7).
3. For the purposes of Redress (Clause 11), the Parties do not agree to the option proposed in the second paragraph of Clause 11(a).
4. For Supervision (Clause 13), the competent supervisory authority is the Data Protection Authority of the country where the data exporter is established.
5. For Governing Law (Clause 17), the Clauses shall be governed by the law of a country allowing for third-party beneficiary rights. The Parties agree that this shall be the law of the country where the data exporter is established.
6. For the Choice of forum and jurisdiction (Clause 18), any dispute arising from these Clauses shall be resolved by the courts of the country where the data exporter is established.


# What are the previous versions of Owkin K-Pro Terms & Conditions?

Bellow you will find previous versions of K-Pro Free Terms and Conditions. Please refer to the latest version of the [Terms and Conditions](broken://pages/8ijtAWS7lRtuSSFcYL1H).

***

**Effective Date: approval of the Terms and Conditions by the User.**

*Version of the terms and conditions: v.1 date April 18, 2025.*

**PREAMBLE**

“K Pro Free” is an artificial intelligence platform dedicated to biomedical research (the “**Platform**”, as defined below) developed by Owkin Inc., a Delaware company having a business address at NeueHouse Madison Square 110 E 25th St, New York, NY 10010, United States of America.

The Platform enables its user (the “**User**”, “**You**”, as defined below) to query a conversational agent in natural language with the aim to help generate and validate scientific hypotheses through the exploration of multi-modal data with visualisation, scientific literature and specific public biological knowledge.

By accessing and using the Platform, You:

* acknowledge having read, understood and accepted these Terms and Conditions, without reservation, limitation or condition;
* represent and warrant that You (i) are duly authorized to accept these Terms and Conditions on Your behalf or on behalf of your employer if applicable; (ii) are not under any restriction to access and use the Platform, and in particular that your status of healthcare professional or public agent does not imply additional formalities to access and use the Platform; (iii) have informed - and if You are healthcare professional and/or public agent, obtained approval from - any appropriate board or authorities (including hierarchical authority if applicable) of Your access to the Platform, if required by law (collectively the “**Criteria**”). Your shall provide all necessary documentation to demonstrate compliance with these Criteria to Owkin at the time of acceptance of this Terms;
* acknowledge that Access is granted based on Your declarations the Platform is accessible to any researcher interested in biology for free, subject to fulfilment of the Criteria, and Owkin will grant an Access to the User based on the declarations made by the User. However Owkin reserves the right to verify whether the Criteria are met, and depending on User’s professional status, may suspend access to the Platform until fulfilment of the Criteria. In addition, please note that access by employees of any type of for-profit entities shall be subject to prior written approval of Owkin under specific business conditions.

**PLEASE READ THESE TERMS AND CONDITIONS CAREFULLY BEFORE ACCESSING OR USING THIS PLATFORM. Any access to the Platform involves the irrevocable and unreserved knowledge and acceptance of these Terms and Conditions by the User. If You do not agree with these Terms and Conditions, You may not access and use the Platform. These Terms and Conditions may be amended from time to time and we recommend that You review these Terms regularly.**

**1 - Definitions**

For the purposes of these Terms:

| “Access Period”:                              | refers to the period during which the User uses and accesses the Platform in accordance with these Terms.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                          |
| --------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| **“Agent”:**                                  | refers to the conversational agent that analyses and interprets the Inputs and selects the right tool and dataset or combination of tools and datasets available on the Platform or accessible to the Agent with the relevant parameters to provide Outputs to the User.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                           |
| **“Applicable Laws and Regulations”:**        | refers to any laws, regulations guidelines and other requirements (regulatory or other) in any jurisdiction applicable to Owkin and/or the Users during its/their use of the Platform.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                             |
| **“Commercial Use”:**                         | refers to any use intended for commercial advantage and/or monetary compensation, including without limiting the foregoing any research project conducted for and/or on behalf of an industrial and/or commercial third party and/or the provision of commercial services with (or using) Outputs, sales of Outputs (or products incorporating such Outputs, for instance commercialisation of a diagnostic kit containing a biomarker to the extent the biomarker was a Outputs), or marketing for commercial services on, or sales of, Outputs, (ii) developing the Outputs (for instance clinical trial on an asset which is the Outputs) with a view to directly commercialize such Outputs (or products incorporating such Outputs), and (iii) manufacturing a Outputs (or products incorporating such Outputs), with a view to directly commercialize such Outputs (or products incorporating such Outputs). |
| **“Data”:**                                   | refers to the Public Dataset and Private Dataset.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                  |
| **“Documentation”:**                          | refers to the instructions for use of the Platform, which are available at the following link: <https://owkinkhelp.zendesk.com/hc/en-us/articles/33938116119697-Owkin-K-User-Guide-Getting-started>.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                               |
| **“Inputs”:**                                 | refers to the information, data, knowledge, content, intellectual reflection, idea formalization, displayed, uploaded or provided in any manner whatsoever to the Platform by the Users in particular while making queries.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                        |
| **“Intellectual Property Rights” or “IPRs”:** | refers to trademarks, patents, designs and models, copyrights, trade names, trade secret, signs, logos, graphics, domain names all rights of whatsoever nature in computer programs, codes, databases, algorithms, Platform, data (including Outputs and Inputs), and any other signs and intellectual property rights owned by Owkin and/or a User, whether or not protectable under Applicable Laws and Regulations, whether registered or unregistered, including all granted registrations and applications of the same                                                                                                                                                                                                                                                                                                                                                                                        |
| **“Login Information”:**                      | refers to a unique identifier assigned to a User which, combined with a password, enables the User to authenticate himself or herself to access the Platform.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                      |
| **“Outputs”:**                                | refers to any analysis, predictions, recommendations, decisions, content (including, texts, images, photos, videos, graphics), results, under any format, generated by the Platform to answer a query of the User.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                 |
| **“Owkin Proprietary Knowledge"**             | refers collectively to the Platform, any proprietary content of Owkin uploaded or made available by Owkin on the Platform, its Documentation and the corresponding IPRs.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                           |
| **“Platform”:**                               | refers to the systems operated and made available by Owkin to the User under these Terms, collectively branded as “K-Pro Free”, which include the Agent, Software and any Data that could be used for the generation of the Outputs. The Platform made available under these Terms is the free of charge version of K-Pro Free.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                    |
| **“Private Dataset”:**                        | refers to anonymized, de-identified and/or pseudonymized data used by or via the Platform, which is under Owkin’s control, that could be accessed by the User under certain restricted conditions, depending on the version of the Platform, and used to generate Outputs.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                         |
| **“Public Dataset”**                          | refers to the anonymized, de-identified data and/or pseudonymized data that have been made available to the public by a third-party natural person or entity, potentially under specific terms of use, that could be accessed and used by the User via the Platform to generate Outputs.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                           |
| **“Software”:**                               | refers to any public or private software, algorithms and/or libraries used by Owkin to enable Users to access and use the Platform.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                |
| **“Terms and Conditions” or “Terms”:**        | refers to these Terms and Conditions and its Annexes.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                              |
| **“Territory”**                               | refers to the countries, listed in the Annex B, from where the User can access and use the Platform, subject to the terms of the Section 8.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                        |
| **“User”:**                                   | refers to any natural person over 18 years of age with access to the Platform authorised by Owkin and, where applicable, its employer and/or any appropriate board or authorities.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                 |

The definitions referred to in this Section apply to both singular and plural terms.

**2 - Purpose of the Terms**

The purpose of these Terms is to set out the rights and obligations of Owkin and the User regarding the use and access to the Platform and its content during the Access Period. These Terms apply each time the User is accessing the Platform.

**3 - Right of access - Use of the Platform.**

**3.1 -** Subject to the limitation set forth in these Terms, Owkin grants to the User a personal, limited, non-exclusive, non-transferable, non-sublicensable, revocable, in the Territories (and without prejudice to Section 8 ), free of charge right to access and use the Platform, and the Documentation during the Access Period. The rights set forth in this Section 3.1 are subject to the compliance with the Criteria during the Access Period. Owkin may suspend User’s access to the Platform, in whole or in part, at its sole discretion, if these Criteria are not (or no longer) satisfied, and access will be reinstated upon User’s fulfilment of the Criteria.

**3.2 -** The above mentioned right of use excludes any right to use the Platform on human subjects, including for research involving human beings (including without limitation in the context of clinical trials and/or any other kind of biomedical research which, all or in part of the purpose, is to develop and/or compare and/or assess the performance of medical devices and/or providing prognostic, diagnosis and/or therapeutics purposes), or in routine, for care or diagnostic purposes, as well as any use of the Platform with or on behalf of third parties and/or for Commercial Use.

**3.3** - The User may access the Platform via the Internet. The User shall use the Platform at his or her own risk and under his or her sole responsibility. The User is aware that the Platform is hosted by a “cloud” provider and made available to the User via K-Pro Free dedicated website or a link from Owkin’s website to K-Pro Free website.

**3.4 -** When accessing the Platform for the first time, the User will register to the Platform by entering the information requested by Owkin, accepting these Terms and if any, providing all necessary document demonstrating compliance with the Criteria as described in the Preamble, and then at each connection, the User shall use the Login Information that Owkin provided so that he or she can connect to the Platform. The Login Information is strictly personal and confidential.

**3.5 -** The User expressly undertakes to keep his or her Login Information confidential and not to communicate it to any third parties. The User shall bear all consequences that may result from the voluntary or involuntary disclosure of its Login Information, and Owkin shall not be held liable for any use of the account by a third party who has gained access to the User's Login Information, in any manner whatsoever. The User shall be solely liable to Owkin for all use of the Platform in violation of these Terms. The User undertakes to immediately inform Owkin of any suspicious or fraudulent use of his or her Login Information.

**4 - Public and Private Datasets**

**4.1** – The User can only use the Public Datasets through the Platform, subject to the provisions set forth below. Except as specifically identified in the list of Public and Private Datasets, no Private Datasets will be made available to the User except if the User accesses a for-fee version of the Platform under specific additional terms.

**4.2 -** The Public and Private Datasets used by the Platform are listed in the Annex A of these Terms. If the User wishes to directly access and use any Public and Private Dataset, the User will access it under his or her own responsibility and under the terms of any license potentially associated with such Dataset. Furthermore, the User understands and agrees that:

* Use and access to the Public and Private Dataset is made at his or her own risk and shall be done in compliance (i) with the Applicable Laws and Regulations; (ii) the terms and conditions of use of each Public and Private Dataset.
* Owkin cannot guarantee that the (i) anonymization, de-identification and/or pseudonymization of the Public and Private Datasets will have the same meaning or requirements under any Users’ Applicable Laws and Regulations (ii) the Outputs is only composed of data that never permits individualisation of patient or research participant by the User.
* Owkin, at the Effective Date, has performed a good faith check of the Public and Private Datasets’ terms for their material alignment with the non-Commercial Use of the Public and Private Dataset on the Platform only.
* User is solely responsible to (i) verify if his or her use of the Public and Private Datasets is compliant with (a) the Datasets’ licences (b) the Applicable Laws and Regulations applying to such User, and (ii) where applicable, to comply with any additional formalities.

**4.3** - The User acknowledges and agrees that:

* Data hosted on the Platform must remain within the Platform and should not be exfiltrated or transferred outside without proper written authorization of Owkin or in application to the licenses applying to this Data;
* it shall not engage in any activities that facilitate or attempt to facilitate the unauthorized extraction, copying, or removal of content, including Data, from the Platform.

**4.4 -** The Platform may use or be used in connection with third-party software, product, services or integration (“**Third-Party Services**”). Owkin will determine in its sole discretion which Third Party Services may be used by the Platform or accessible to the Users via the Platform. Notwithstanding any contrary provisions, use of Third-Party Services is at User own risk and subject to User’s compliance with any terms, conditions or policies applicable to such Third-Party Services. Owkin does not control or accept any liability for claims and resulting liabilities, damages, losses and expenses, including reasonable attorneys' fees, arising out of or resulting from use of Third-Party Services.

**5 - Obligations of the User**

**5.1 -** The User acknowledges being familiar with the purpose, functionalities and operating procedure of the Platform and any content accessible to this User, having ensured that it conforms with its needs and skills, in particular based on the information provided in the Documentation.

**5.2 -** The User may contact Owkin, by email to [support@owkin.com](mailto:Product-support@owkincom) or by completing and sending the dedicated form on the Platform, for any types of requests related to the Platform such as service requests, errors and other operating anomalies, as well as all suggestions for improvement that he or she deems desirable.

**5.3 -** The User undertakes to:

* use the Platform and its content in accordance with the provisions of these Terms, as of the Documentation and of the Applicable Laws and Regulations;
* use the Platform and its content in accordance with their intended purposes and in a manner compatible with reasonable use;
* provide regular feedback on his/her use of the Platform to help Owkin improve the Platform as stated in section 6.4.

**5.4 -** The User shall refrain from using the Platform and its content for illicit or illegal purposes. The User undertakes, in particular to not: (i) lease, sell, sub-license, assign, grant access to or otherwise transfer the Platform nor its content (in particular the Software and the Agents) and/or the Documentation to any third party for any purpose whatsoever, including for evaluation or testing purposes; (ii) decipher, decompile, disassemble, reverse assemble, modify, translate, reverse engineer or otherwise attempt to derive source code, algorithms, tags, specifications, architectures, structures or other elements of the Platforms, in whole or in part; (iii) attempt to re-identify, by any way whatsoever, the data subjects from which the Data may come or to investigate the origins of populations and/or make research ancestors ; (iv) represent that the Outputs has been generated by a human being while this not the case; (v) use the Outputs to generate any technology that competes with all/or part of the Platform.

**5.5 -** The User is solely responsible for the Inputs and any information provided to Owkin and/or the Platform. The User may not provide or make available on the Platform in any manner whatsoever any Inputs, feedback or any other information or content that is (i) false, inaccurate, not up-to-date or incomplete; (ii) offensive, defamatory, violent, illegal, obscure or contrary to public order, to morality or to Owkin’s or its affiliates’ interests or in breach of Applicable Laws and Regulations; (iii) directly or indirectly identifying a patient and/or a participant to a research involving human beings.

**5.6** - If the User were to find that it has access to Data or Outputs, while using the Platform, that can be directly attributed to a specific data subject, it will immediately (i) stop accessing such Data or Outputs; and (ii) within a delay which shall not exceed twenty-four (24) hours notify Owkin by email to <dpo@owkin.com>.

**5.7** - The User shall not damage, interfere with or disrupt the Platform, including circumvent any rate limits or restrictions or bypass any protective measures or safety mitigations Owkin puts in the Platform, including for example overloading the Platform, introducing viruses or malware, spamming or DDoSing Services.

**5.8 -** The User undertakes to provide regular feedback on his or her use of the Platform, which may include ideas, suggestions for improvement and ratings requested by Owkin on the Platform through Thumbs Up or Thumbs Down or through discussions with the support team (“**Feedback**”). Free access to the Platform is conditional upon the provision of the User’s Feedback, which may be used or not by Owkin, and - in any case - shall never be conditioned to any further obligation or payment in return.

**6 - Intellectual Property**

**6.1 - Owkin Proprietary Knowledge**

**6.1.1 –** As the owner and/or guardian of the Owkin Proprietary Knowledge, Owkin guarantees that it can grant access to the Platform in the Territory and for the intended purpose of the Platform.

**6.1.2 -** Owkin Proprietary Knowledge is and remains wholly owned or controlled by Owkin. Any improvement of Owkin Proprietary Knowledge, its associated content and all Intellectual Property Rights attached thereto (including their previous, current and future developments, updates and versions) is and remains owned or controlled by Owkin. The User’s right to access and use the Platform granted under these Terms does not entail any transfer and/or assignment of any Intellectual Property Right attached to the Owkin Proprietary Knowledge.

**6.1.3** - The User acknowledges that Owkin may freely use any information related to the use of Owkin Proprietary Knowledge that it receives as feedback from the User. The User shall refrain from claiming any IPRs, to said feedback, or to any development or improvement that Owkin and/or its affiliates could make in connection with User’s feedback either during or after the term of the Access Period.

**6.1.4 -** The User shall (i) respect Owkin’ IPRs; (ii) refrain from directly or indirectly infringing Owkin’s ownership rights, title and interest to Owkin Proprietary Knowledge and any IPRs attached thereto, in any manner and by any means whatsoever; and (iii) not delete or alter the copyright and/or proprietary notice appearing on the Platform, its content and any Documentation of any type on which this notice features.

**6.1.5 -** Except as expressly provided in the Article 6.3.4, the User is not allowed to use the name, trademark and logo of Owkin.

**6.2 - Inputs**

**6.2.1** - The User is and remains the owner and/or guardian of any content provided on the Platform, including Inputs.

**6.2.2** - The User grants Owkin a worldwide, non-exclusive, transferable, sub-licensable, free of charge licence to use, reproduce, represent and exploit the Inputs for the improvement of Owkin Proprietary Knowledge, the Platform, the Documentation and any other services or products of Owkin or its affiliated companies, perpetually or where applicable for the duration of the intellectual property rights.

**6.2.3** - The User represents and warrants that it holds all the necessary rights, consent and/or authorizations to provide and use the Inputs on the Platform and grant the rights set forth in Section 6.2.2.

**6.2.4** - The Platform is not designed to store User’s or third-parties’ confidential information and Owkin does not guarantee the confidentiality of the Inputs. The User shall not provide his or her confidential information nor the one of his or her employer, as applicable, within Inputs.

**6.2.5 -** The User is solely responsible for the storage and back up copy of the Inputs, specifically before closing his or her browser. Owkin shall not be held liable for any User’ claims related to the copy and/or transfer of the Inputs.

**6.3 - Outputs**

**6.3.1 -** Owkin does not claim any rights, interests, or title on the Outputs which are provided “as is” to the User, with the limitations described in Section 7.2. For clarity, such Outputs may be used by the User, at its sole risk, for internal academic research purposes only (excluding any Commercial Use). However, as the User is using a free version of the Platform, the User grants Owkin with a non-exclusive, worldwide, royalty-free, irrevocable, transferable and sublicensable license to use, reproduce, represent and exploit the Outputs for the improvement of Owkin Proprietary Knowledge, the Platform, the Documentation and any other services or products of Owkin or its affiliated companies, perpetually or where applicable for the duration of the intellectual property rights.

**6.3.2 -** The User understands and agrees that the Outputs:

* may not be accurate and/or not an exact representation of the reality and therefore needs to be independently validated by the User by following the relevant and applicable scientific, ethical and legal practices;
* do not reflect Owkin’s opinions, analysis and/or interpretations of the Inputs;
* may not be unique and other users may obtain similar Outputs, by using the Platform and/or any similar other platforms and/or tools (if any). Consequently, the User acknowledges that the rights granted over the Outputs, in accordance with the Article 6.3.1, does not extend to the Outputs obtained by other Users or third parties using the Platform and/or any similar other platforms and/or tools (if any).

**6.3.3 -** The User is solely responsible for the storage and back up copy of the Outputs, specifically before closing his or her browser and/or before the term of the Access Period. Owkin shall not be held liable for any User’ claims related to the copy and/or transfer of the Outputs.

**6.3.4 -** In consideration for the free of charge access to the Platform granted by Owkin to the User, the User undertakes to credit Owkin in any scientific publication referring to the Outputs obtained in connection with the use of the Platform, as follows: “*The results \<published or shown> here are in whole or part based upon results generated by using K-Pro Free, a Platform by Owkin*”.

**6.4 - Feedbacks**

**6.4.1 -** By accepting these Terms, the User grants Owkin with a non-exclusive, worldwide, sublicensable, transferable, royalty-free, and irrevocable license to use, reproduce, represent and exploit his or her Feedback for the improvement of Owkin Proprietary Knowledge, the Platform, the Documentation and any other services or products of Owkin perpetually or for the duration of the applicable intellectual property rights.

**7 - Warranty and exclusions**

**7.1 -** The User acknowledges that its use of the Platform is voluntary and will not be subject to any remuneration in any form whatsoever.

**7.2 -** The Platform, the Documentation, the Data and the Outputs are made available “as is” and Owkin does not provide or make any representation or warranty, either express or implied, including but not limited to (i) quality, fitness or suitability for any particular purpose, including to review and validate the scientific hypothesis of the User; and/or (ii) non-infringement of third parties’ rights; (iii) absence of any errors, defects, design flaws or other problems (iv) accuracy or reliability of the Platform and/or Outputs; (v) uninterrupted, timely or secure provision of the Platform.

**7.3 -** Owkin shall not be held liable for material damages caused by the Platform and/or the Data to the User and/or any third parties. The User is solely responsible for its use of the Inputs, Platform, Documentation, Data and the Outputs and shall bear all risks associated with this use, including but not limited to: instability, malfunctioning, loss of content and in particular Inputs and Outputs, and other damage or loss.

**7.4 -** To the fullest extent permitted by Applicable Laws and Regulations, Owkin shall not be held liable to the User for any direct or indirect damage or loss whatsoever (such as commercial or financial damage or operating losses that may affect the User) resulting from (i) any use of, or inability to use, the Platform, the Documentation, the Outputs and the Data or (ii) the use or reliance on any information or results, including Outputs, displayed or available in the Platform. The User shall indemnify and hold harmless Owkin and its affiliates, and their personnel, agents and representatives, and successors and permitted assigns, against any and all third party claims and resulting liabilities, damages, losses and expenses, including reasonable attorneys' fees, arising out of or resulting from (i) negligence or wilful misconduct in connection with these Terms, or (ii) a breach of these Terms, each by the User and, when its legal person, its affiliates, personnel, agents or representatives.

**7.5 -** Owkin shall not be held liable in the event of legal proceedings initiated by third parties against the User due to the illicit or illegal use of the Platform, the Documentation, the Data and the Outputs.

**7.6** - Nothing in these Terms limits or excludes any other warranty or liability that cannot be limited or excluded under Applicable Laws or Regulation, or death or personal injury caused by negligence.

**8 - Export control**

The User may not export or provide access to the Platform to persons or entities or into countries or for uses where it is prohibited under U.S., European Union Law, Applicable Laws or Regulations or other applicable international law. Without limiting the foregoing sentence, this restriction applies (a) to countries where export from the US, and/or the European Union or into such a country would be prohibited or illegal without first obtaining the appropriate license, and (b) to persons, entities, or countries covered by U.S. sanctions and/or European Union sanctions.

**9 - Transparency**

**9.1 -** The User acknowledges that in accordance with applicable laws, regulations and/or ethical codes certain services contracts, consultancy contracts, sponsorship contracts and/or collaboration contracts, terms and conditions, between on the one hand health care companies (as Owkin), and on the other hand (associations of) health care professionals (HCPs) and/or health care or academic institutions or hospitals (HCOs), may be subject to mandatory notification/authorization to the competent authorities (“**Anti-Gift Obligations**”) and publication, including but not limited to the publication of amounts paid and personal data (such as names, location, etc.), (“**Transparency Obligations**”). The User hereby acknowledges that in accordance with the Transparency Obligations, Owkin and its affiliates may need to make public information of the aforementioned nature in relation to these Terms, including but not limited to the identity of the employer of the User and the market value for the employer of the User of the access to and the use of the Platform. The User hereby expressly (i) consents on behalf of his/her institution to such publication; and (ii) represents and warrants having the authorization of its employer to enable Owkin to use its name for the compliance with such Transparency Obligations.

**9.2 -** The User is not obligated, directly or indirectly, to prescribe, purchase or recommend the prescription or the purchase of any products manufactured or distributed by any of the entities of Owkin, or to make any referrals of patients to such entities, to otherwise generate purchases for or direct any purchasers to such entities. The access rights to the Platform granted to the User is not intended to compensate the User for past, present, or future prescriptions, referrals or recommendations of Owkin products.

**10 - Term and termination**

**10.1 -** The Terms shall enter into force on the Effective Date and shall remain effective until the 6th of May 2025, except as terminated by Owkin in application of this Section.

**10.2 -** Owkin reserves the right to suspend or terminate the Platform and/or any of its features, at any time, without having to bring the case before the competent court, at Owkin’s discretion, without any justification, nor prior notice.

**10.3** In case of suspected or actual breach to these Terms by the User, Owkin reserves the right to immediately, without prior notice, without having to bring the case before the competent court: (i) delete any breaching content, including Inputs or Outputs ; (ii) suspend or restrict access to the Platform, in total or in part; and (iii) suspend or cancel the User account at its own discretion; without prejudice of any other damages Owkin may claim.

**10.4** In case of termination, the User account will be deactivated and archived for a duration necessary for the applicable statute of limitations, if needed, for evidence purposes, and the applicable legal retention periods. The Inputs and Outputs are not stored into the User account and will not be provided to the User at the end of the Terms; it is the User’s sole responsibility to ensure that such Inputs and Outputs are regularly saved by the User following their generation.

**11 - Personal data protection**

**11.1 -** The Privacy Policy available [here](https://www.owkin.com/policies/privacy-policy) explains how Owkin is processing personal data of the User in connection with their access and use of the Platform. The Cookies Policy available [here](https://k.owkin.com/legal/cookie-policy) explains how and for which purposes the Platform may include technical devices (cookies or other tracking technologies) that allows sending to Owkin and its processors information about the use and navigation of the User. The Privacy Policy and Cookies Policy can be updated from time to time and do not form part of these Terms.

**11.2 -** Owkin and the User recognize that they have full and entire knowledge of the obligations under the regulation applicable to personal data provided under any provision of a legislative or regulatory, European or national nature, resulting in particular from Regulation 2016/679/EU of 27 April 2016 on the protection of individuals with regard to the processing of personal data and on the free movement of such data as well as any other regulations applicable in this field that may be added to or subsequently replaced by it ("**Personal Data Regulation**"), that apply to them, in their respective capacity as independent data controllers. Owkin is the data controller of the personal data processing for the development, provision, maintenance and enhancement of the Platform, including the Inputs, Outputs and any information resulting, automatically or not, from the use and navigation on the Platform of the User. As such, each Party shall take, for its own processing activities, any appropriate measures to ensure the compliance with this Regulation. To the extent that personal data transfers should occur, these Terms incorporate the European Commission Standard Contractual Clauses as made available by the European Commission on its website in a non-modifiable form available here: [Standard Contractual Clauses](https://eur-lex.europa.eu/eli/dec_impl/2021/914/oj?uri=CELEX:32021D0914\&locale=en), and completed as set forth in Annex C (the “**Clauses**”)..

**12 - General provisions**

**12.1 - Changes to the Terms.** Owkin reserves the right to make changes to these Terms at any time. Any change will come into force at the time of new Terms availability on Owkin or K-Pro Free dedicated websites. Owkin advises the User to consult frequently such websites to be informed on the new Terms. If the User disagrees with the new Terms, User shall stop using the Platform.

**12.2 - Assignment**. The User shall not assign all or part of the Terms to a third party without the prior written consent of Owkin. However, Owkin may assign these Terms to its affiliates and/or any third party, notably in the event of a change of control of Owkin and/or its affiliates.

**12.3 - Force majeure**. Owkin may not be held liable for any failure or delay in the performance of their obligations under these Terms and Conditions resulting from circumstances that are unforeseeable, beyond its reasonable control, and which cannot be overcome despite its reasonable endeavors (“**Force Majeure Event**”). However, under such circumstances, Owkin undertakes to notify the Force Majeure Event to the User as soon as possible.

**12.4 - Independence of the parties**. The parties acknowledge and agree that they may not under any circumstances make a commitment in the name of and/or on behalf of one another. Under no circumstances shall these Terms be interpreted as creating a binding relationship or de facto partnership between the parties, each of which shall be considered an independent co-contractor.

**12.5 - Relations between the parties.** The parties acknowledge and accept that their collaboration can in no case be considered as a de facto company, a joint venture, or any other situation entailing a representation of reciprocity or solidarity towards their respective creditors.

**12.6 - Entire agreement.** These Terms express the entirety of the obligations of the parties and supersedes any previous agreement between the parties relating to the subject matter hereof.

**12.7 - Non-waiver**. If, in the event of a breach by either party of its obligations under the Terms, the non-defaulting Party does not waive its rights resulting from said breach. In no event, the failure to exercise its rights shall be construed as a waiver of the exercise of said rights, either in the future, or in the event of a similar breach by the defaulting party of its obligations under the Terms.

**12.8 - Severability.** If any provision of the Terms is held to be invalid or unenforceable pursuant to a prevailing rule of law or a final court decision, the provision shall be either amended to render it valid or deemed removed, without this leading to the invalidity of these Terms or affecting the validity of its other provisions, which shall retain their full force and scope, provided that the balance of these Terms remains unaltered.

**12.9 - Applicable law, dispute resolution.** These Terms are subject to the laws of the State of New York. In the event of a dispute relating to the interpretation, performance or validity of the Terms, express jurisdiction is attributed to the competent court within the jurisdiction of the State of New York.

* \*\*

**Annex A: List of Public and Private Datasets**

| Dataset Name                                        | Type                                  | Website                                                         | License Link                                                                                                                 | License Name                                                                 | Version |
| --------------------------------------------------- | ------------------------------------- | --------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------- | ------- |
| **ChEMBL**                                          | Public                                | <https://www.ebi.ac.uk/chembl/>                                 | [Deed - Attribution-ShareAlike 3.0 Unported - Creative Commons](https://creativecommons.org/licenses/by-sa/3.0/)             | Deed Attribution-Share                                                       | 3.0     |
| **CollecTRI**                                       | Public                                | <https://github.com/saezlab/CollecTRI?tab=readme-ov-file>       | [The GNU General Public License v3.0 - GNU Project - Free Software Foundation](https://www.gnu.org/licenses/gpl-3.0.en.html) | The GNU General Public License v3.0 - GNU Project - Free Software Foundation | 3.0     |
| **COMPARTMENTS**                                    | Public                                | <https://compartments.jensenlab.org/Search>                     | [Deed - Attribution 4.0 International - Creative Commons](https://creativecommons.org/licenses/by/4.0/)                      | Deed - Attribution 4.0 International - Creative Commons                      | 4.0     |
| **Complex Portal**                                  | Public                                | <https://www.ebi.ac.uk/complexportal/home>                      | [Deed - CC0 1.0 Universal - Creative Commons](https://creativecommons.org/publicdomain/zero/1.0/)                            | Deed - CC0 1.0 Universal - Creative Commons                                  | 1.0     |
| **DepMap**                                          | Public                                | <https://depmap.org/portal/>                                    | [Deed - Attribution 4.0 International - Creative Commons](https://creativecommons.org/licenses/by/4.0/)                      | Deed - Attribution 4.0 International - Creative Commons                      | 4.0     |
| **Ensembl (BioMart)**                               | Public                                | <https://www.ensembl.org/info/data/biomart/index.html>          | [Deed - Attribution 4.0 International - Creative Commons](https://creativecommons.org/licenses/by/4.0/)                      | Deed - Attribution 4.0 International - Creative Commons                      | 4.0     |
| **gnomAD**                                          | Public                                | <https://gnomad.broadinstitute.org/>                            | [Deed - CC0 1.0 Universal - Creative Commons](https://creativecommons.org/publicdomain/zero/1.0/)                            | Deed - CC0 1.0 Universal - Creative Commons                                  | 1.0     |
| **GTEx**                                            | Public                                | <https://gtexportal.org/home/>                                  | [GTExPOrtal Data License](https://gtexportal.org/home/license)                                                               | GTExPOrtal Data License                                                      | 4.0     |
| **Hallmarks of Cancer**                             | Public                                | <https://pubmed.ncbi.nlm.nih.gov/21376230/>                     | No licence                                                                                                                   | N/A                                                                          | N/A     |
| **Human Protein Atlas (HPA)**                       | Public                                | <https://www.proteinatlas.org/>                                 | [Deed - Attribution-ShareAlike 3.0 Unported - Creative Commons](https://creativecommons.org/licenses/by-sa/3.0/)             | Deed Attribution-Share                                                       | 3.0     |
| **IntOgen**                                         | Public                                | <https://www.intogen.org/search>                                | [Deed - CC0 1.0 Universal - Creative Commons](https://creativecommons.org/publicdomain/zero/1.0/)                            | Deed - CC0 1.0 Universal - Creative Commons                                  | 1.0     |
| **MsigDB**                                          | Public                                | <https://www.gsea-msigdb.org/gsea/msigdb/human/collections.jsp> | [Deed - Attribution 4.0 International - Creative Commons](https://creativecommons.org/licenses/by/4.0/)                      | Deed - Attribution 4.0 International - Creative Commons                      | 4.0     |
| **Open Targets**                                    | Public                                | <https://platform.opentargets.org/>                             | [Deed - CC0 1.0 Universal - Creative Commons](https://creativecommons.org/publicdomain/zero/1.0/)                            | Deed - CC0 1.0 Universal - Creative Commons                                  | 1.0     |
| **RCSB PDB**                                        | Public                                | <https://www.rcsb.org/>                                         | [Deed - CC0 1.0 Universal - Creative Commons](https://creativecommons.org/publicdomain/zero/1.0/)                            | Deed - CC0 1.0 Universal - Creative Commons                                  | 1.0     |
| **Reactome**                                        | Public                                | <https://reactome.org/>                                         | [Deed - Attribution 4.0 International - Creative Commons](https://creativecommons.org/licenses/by/4.0/)                      | Deed - Attribution 4.0 International - Creative Commons                      | 4.0     |
| **TCGA**                                            | Public                                | <https://www.cancer.gov/ccg/research/genome-sequencing/tcga>    | <https://creativecommons.org/licenses/by-nc-nd/4.0/>                                                                         | CC Attribution-NonCommercial-NoDerivatives 4.0 International                 | 4.0     |
| **TMHMM**                                           | Public                                | <https://services.healthtech.dtu.dk/services/TMHMM-2.0/>        | <https://opensource.org/license/mit>                                                                                         | MIT                                                                          | N/A     |
| **Uniprot**                                         | Public                                | <https://www.uniprot.org/>                                      | [Deed - Attribution 4.0 International - Creative Commons](https://creativecommons.org/licenses/by/4.0/)                      | Deed - Attribution 4.0 International - Creative Commons                      | 4.0     |
| **Tabula Sapiens**                                  | Public                                | <https://tabula-sapiens.sf.czbiohub.org/>                       | <https://github.com/czbiohub-sf/tabula-sapiens/blob/master/LICENSE>                                                          | BSD 3-Clause License                                                         |         |
| **MOSAIC Window**                                   | Private with restricted public access | <https://www.mosaic-research.com/mosaic-window>                 | Terms and conditions                                                                                                         | N/A                                                                          | N/A     |
| **University Hospital Erlangen anonymized dataset** | Private with restricted public access | <https://www.uk-erlangen.de/en/>                                | Terms and conditions                                                                                                         | N/A                                                                          | N/A     |

**Annex B: Territories where the access and use of the Platform is allowed**

Any country except for the countries subject to sanctions and/or restrictions notably in application of the Export Control laws and regulations.

Without limiting the foregoing, the access and use of the Platform is formally prohibited in the following countries: Afghanistan, Yemen, Mali, Syria, Somalia, Central African Republic, South Sudan, Libya, Bangladesh, North Korea, Russia, Iran, Egypt, China, Belarus, Indonesia, Taiwan.

**Annex C – Standard Contractual Clauses**

Owkin and the User agree that these Clauses apply to them as such, depending on the scope that relates to them and for the purposes of the Clauses, the following specific terms have been agreed.


# Cookie policy

This cookie policy describes what kinds of cookies and similar technologies Owkin (hereinafter as “**Owkin**” “**us**” or “**we**”) uses in connection with our Platform (as defined in our [Terms and Conditions](broken://pages/8ijtAWS7lRtuSSFcYL1H)), and how you can manage them (hereinafter as the “**Cookie Policy**”).

This Cookie Policy aims to inform you of

1. the ways in which we use cookies on your terminal,
2. their purposes; and
3. your rights regarding the use of these cookies.

This Cookie Policy is compliant with the requirements of the General Data Protection Regulation (EU) 2016/679 of April 27th, 2016 (“**GDPR**”) regarding the use of cookies.

This Cookie Policy is applicable by the sole fact of its publication on our website and does not replace the [Terms and Conditions](broken://pages/8ijtAWS7lRtuSSFcYL1H) of our Platform, nor [Owkin’s Privacy Policy](https://www.owkin.com/policies/privacy-policy).

**What are cookies?**

Cookies are small text files stored and/or read by your browser on your terminal when you are visiting the Website. These cookies allow us to store information and identify the terminal you are using, information about your visit, such as your language settings or when you logged in, which can improve your experience when you revisit the Platforms.

Cookies are used on various platforms (website, e-mail, mobile applications, other). The cookies used by our Platforms have the following purposes:

* improving your experience by configuring the use of the environment according to your preferences.
* improve your experience through the study of user behavior on the Platform.

Cookies allow us to process some of your personal data, such as, but not limited to your IP address or the date and time of your connection. The Platform uses different types of cookies which are placed by us (first party cookies) or by third party services used by our Platform, which are cookies from a domain different from the domain of our Platform (third party cookies).

**What cookies do we use?**

* **Necessary cookies**

Necessary cookies are cookies which enable us to provide you with access to the Platform and to use it correctly. Without such cookies, our Platform cannot function properly.

| Source | Cookie name             | Duration             | Purposes                                                                                        |
| ------ | ----------------------- | -------------------- | ----------------------------------------------------------------------------------------------- |
| Owkin  | userPreferences.cookies | Expires after 1 year | This cookie is used by Owkin to record the choice of the user regarding cookies placing or not. |

* **Analytics cookies**

Analytics cookies help Owkin to understand how visitors interact with the Platform by collecting and reporting information.

| Source        | Cookie name | Duration              | Purposes                                                                                                                                                                                                                                                                                                                                                                                                                                                            |
| ------------- | ----------- | --------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Fullstory     | fs\_uid     | Expires after 1 year  | The 'fs\_uid' cookie can be thought of as the capture cookie. When a user visits the Platform , that cookie is used to track the user across sessions and pages. The same user may visit the Platform multiple times and may navigate to many pages within a single session. This cookie ensures that all captured session traffic is associated with one user. A session cannot be captured without this cookie and the users anonymized visit will not be logged. |
| Fullstory     | fs\_lua     | Expires after 30 mins | Captures the timestamp of the last user action. It is used to assist with the Fullstory session lifecycle, ensuring user activity extends the session.                                                                                                                                                                                                                                                                                                              |
| Surveysparrow | ssUserId    | Expires after 1 year  | User id cookie for SurveySparrow. Allows to track user answers to surveys created with this tool.                                                                                                                                                                                                                                                                                                                                                                   |
| Surveysparrow | cookieyesID | Expires after 1 year  | User id for cookieYes, the cookie consent solution used by SurveySparrow.                                                                                                                                                                                                                                                                                                                                                                                           |

**Recipients**

Access to personal data that we collect through cookies is limited to both internal and external recipients:

* internally, personal data is only accessible to authorized departments of Owkin, if necessary;
* externally, personal data is only accessible to our technical service providers who can place cookies on your terminal and/or administer them and access the related personal data.

\\

* **Consent**

You have the right to consent to the use of any cookies studying your behavior (cookies linked to audience measurement operations). However, your consent is not required for cookies necessary for technical use, since such cookies are necessary to use the Platform.

The collection of your consent takes the form of an information banner with an acceptance button and the option to click on the link in order to read the Cookie Policy.

We draw your attention to the fact that disabling cookies may prevent you from accessing certain features on our Platform.

* **Cookie settings**

If you would like to change your consent to accept or refuse cookies on this website, you can do it any time by consulting the link available below in this Cookie Policy.

[Change my consent](https://github.com/owkin/k/blob/main-docs/settings/cookie-preferences/README.md)

You have the right to set the use of cookies on your terminal in order to accept or refuse all or part of the cookies that may be read or stored on your terminal, regardless of their nature and origin. However, this does not apply for the cookies necessary for proper functioning of the Platform.

To ensure that you have full control over the cookies stored in your terminal, you can check your browser settings.

Most browsers allow you to choose or refuse all or part of the cookies, or even to select only those you wish to keep. To that end, each browser has their own specificities, and the settings remain relatively accessible via the « Help Menus » of each of them.

For example:

* **Google Chrome:** <https://support.google.com/accounts/answer/61416?co=GENIE.Platform%3DDesktop&hl=fr>;
* **Internet Explorer:** <https://support.microsoft.com/fr-fr/help/17442/windows-internet-explorer-delete-manage-cookies>;
* **Mozilla Firefox:** <https://support.mozilla.org/fr/kb/activer-desactiver-cookies-preferences>; <https://support.mozilla.org/fr/kb/effacer-les-cookies-pour-supprimer-les-information>;
* **Opera:** <http://help.opera.com/Windows/10.20/fr/cookies.html>.

For information about how to manage cookies on the browser of your mobile device, you will need to consult the device manual.

Owkin cannot guarantee the durability of these URLs, nor the quality of the information contained therein.

* **Updates**

Owkin may at any time amend or supplement this Cookie Policy, including in the event of regulatory changes or recommendations from the supervisory authorities, including the CNIL and/or to reflect the changes of Owkin’s practices. Any new version of this Cookie Policy will be available on this page. Therefore, Owkin invites you to check this Cookie Policy on a regular basis.

**Contact us**

If you have any questions about this Cookie Policy, please do not hesitate to contact us via the contact form available at the following address: <https://owkin.com/contact>.


# Subprocessors

At Owkin, we engage a limited number of carefully vetted third-party service providers (“subprocessors”) to support the delivery, maintenance, and enhancement of our services. Each subprocessor may have access to or process personal data on our behalf in the course of providing its services. We maintain a rigorous due diligence process to assess their technical and organizational measures, ensuring they meet the standards required under the **EU General Data Protection Regulation (GDPR)** and the **UK GDPR**. This page identifies our current subprocessors, describes the nature of their processing activities and safeguards in place.

| Subprocessor                                                       | Purpose of Processing                                                                                                | Data Categories Processed                                                                                                                             | Location                                    | Safeguards                                                                                                                                                                           |
| ------------------------------------------------------------------ | -------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| **Amazon Web Services EMEA SARL (and its approved subprocessors)** | Cloud infrastructure, hosting, storage, networking, and related operational services supporting our SaaS platform.   | Customer data, including personal data contained in customer accounts and application usage data.                                                     | EU (Ireland)                                | Data Processing Addendum (DPA) in place; certified under ISO 27001, SOC 2, and GDPR-compliant Standard Contractual Clauses (SCCs) for international transfers; CCPA-compliant terms. |
| **Anthropic PBC (via AWS Marketplace)**                            | AI and large language model processing for product features involving natural language understanding and generation. | Text input data submitted by users for AI-based processing.                                                                                           | EU (Ireland) - hosted on AWS infrastructure | DPA and SCCs in place via AWS Marketplace; data access limited to processing context; CCPA-compliant terms.                                                                          |
| **Qdrant GmbH (via AWS Marketplace)**                              | Managed vector database services for semantic search, embeddings storage, and similarity queries.                    | Metadata, vector embeddings derived from public data (no raw personal data stored directly).                                                          | EU (Ireland) - hosted on AWS infrastructure | DPA in place via AWS Marketplace; GDPR-compliant data transfer mechanisms and SCCs; CCPA-compliant data processing terms.                                                            |
| **PostHog Ltd.**                                                   | Product analytics and usage tracking to improve user experience, feature adoption, and platform performance.         | Usage data (e.g., feature interactions, session metadata, device/browser info, timestamps). No customer content or sensitive personal data collected. | EU (Germany, hosted on AWS infrastructure)  | GDPR-compliant DPA; SCCs for data transfers outside the EEA; CCPA-compliant processing terms.                                                                                        |
| **Weave Wandb, Inc**                                               | Observability and tracing for AI workflows, including monitoring of model inputs, outputs, and performance metrics.  | Application logs, metadata, and pseudonymized user interaction data. No sensitive personal data intentionally processed.                              | US (hosted on GCP infrastructure)           | GDPR-compliant DPA; SCCs for international transfers as applicable; CCPA-compliant terms.                                                                                            |
| **Zendesk, Inc**                                                   | Customer support, helpdesk, and ticketing platform used to manage and respond to customer service inquiries.         | Customer contact information (e.g., name, email), communication content, and related metadata.                                                        | US (hosted on AWS infrastructure)           |                                                                                                                                                                                      |

Questions or feedback? Contact <support@owkin.com>


# Acceptable Use Policy

This Acceptable Use Policy applies to access and use of the Services provided to Client by Owkin in accordance with the Agreement. Owkin reserves the right to update this Acceptable Use Policy from time to time to reflect:

* developments in Owkin’s Services;
* technological progress and evolving market standard practices, particularly in relation to artificial intelligence architecture, safety, and security; and
* changes in Applicable Laws.

Owkin will provide quarterly written notice to the Client by email, indicating the changes (if any) to the Acceptable Use Policy provided in Schedule 6 vs. the then current version of the Acceptable Use Policy.

This Acceptable Use Policy is designed to ensure the Client's lawful and responsible access to the Platform and use of the Services. By executing the Agreement, Client agrees to comply with the following undertakings.

#### 1. Prohibited uses

Client must not use the Services to engage in, facilitate, or promote unlawful, deceptive, manipulative, or abusive activities, including but not limited to the following activities:

1. Violation of any Applicable Laws (e.g., intellectual property law, data privacy and artificial intelligence regulations);
2. Engaging in illegal activities (e.g., human trafficking, real money gambling, sale or purchase of illegal drugs or any illegal or controlled substances or goods or services, design or sale of weapons or dangerous and chemical materials);
3. Promotion of violence, hate speech, harm, extremism, or terrorism (e.g., any use that glorifies or facilitates self-harm or suicide, is sexually exploitative or abusive, constitutes hate speech or promotes racism, discrimination, hatred, or abuse, against any group or individual based on protected characteristics like race, ethnicity, national origin, religion, disability, sexual orientation, gender, or gender identity, promotes any act of violence or intimidation against people or animals);
4. Child sexual exploitation and sexually or violent explicit content involving minors (e.g., any use that exploits, abuses, sexualizes or endangers children, or otherwise compromises the health and safety of children; generation, creation, sharing, or facilitation of sexually explicit content involving minors, including pornographic content, exposing minors to violent content);
5. Individuals’ surveillance and predictive activities (e.g., illegal profiling or surveillance, including spyware or communications surveillance, untargeted scraping of facial images to create or expand a facial recognition database, or predictive policing, assessing the risks of a person committing a criminal offence);
6. Compromising individuals’ privacy or identity (e.g., violation of a person’s privacy rights, unauthorized access to personal information, unlawful access to or tracking of a person’s physical location, unlawful social scoring, real-time identification of a person or inference of emotions or protected characteristics of a person such as race or political opinions based on biometric data (including facial recognition);
7. Misinformation (e.g., generating or promoting misinformation, disinformation, defamatory content, or political propaganda; manipulating public opinion on health, safety, policy, or political matters; or undermining democratic processes through misrepresentation of voting procedures or voter suppression);
8. Fraudulent, abusive, misleading, or deceptive activities (e.g., enabling counterfeiting or distribution of illegal goods; promoting spam generation or distribution; creating content for fraud, scams, phishing, or malware causing financial or psychological harm; or producing deceptive digital content such as fake reviews or media);
9. Compromising or circumventing Owkin or a third-party security, cybersecurity safeguards or computer or network systems (e.g., exploiting vulnerabilities in IT systems, networks, or applications without authorization of the IT system owner; unauthorized access to IT systems, networks, applications, or devices through technical attacks or social engineering; creating or distributing malware, ransomware, or other types of malicious code; unauthorized access to the Services through any malicious activity including cybersecurity attacks or data poisoning).

#### 2. High-Risk Uses

Unless expressly authorized by Owkin and agreed between the Parties through a separate addendum to the Agreement, Client shall not use the Services for High-Risk Uses.

"High-Risk Uses" means any use of the Services that poses a significant risk of harm to the health, safety, fundamental rights or legal interests of natural persons, including use cases classified as high-risk under:

* AI Systems that are safety components of products covered by EU harmonisation legislation as set out in Annex I of the EU AI Act (Regulation (EU) 2024/1689);
* AI Systems listed in Annex III of the EU AI Act (Regulation (EU) 2024/1689); or
* Any other applicable artificial intelligence, data protection or sector-specific legislation in force in the relevant jurisdiction.

High-Risk Uses include, but are not limited to, deployment of the Services in the following areas:

* biometric identification and categorisation;
* critical infrastructure management;
* education and vocational training (including admissions and assessments);
* employment, recruitment and workers management;
* access to essential private and public services, credit and benefits;
* law enforcement and predictive policing;
* migration, asylum and border control;
* administration of justice and democratic processes; and
* any other use case designated as high-risk by applicable legislation or competent authorities.

Client acknowledges that High-Risk Uses require compliance with applicable regulatory requirements and that unauthorized deployment may result in immediate termination of the Agreement for material breach of this Acceptable Use Policy.

#### 3. Specific Use Cases

For any deployment of the Platform in automated, integrated, or programmatic contexts (including API usage, agentic workflows, skills, or code execution features), the Client shall implement appropriate human oversight and governance controls where the Platform is used to:

* support or automate decisions affecting individual rights;
* process personal data at scale;
* generate content intended for public distribution;
* execute financial transactions;
* interact autonomously with third-party systems; or
* operate in safety-critical or regulated environments.

Such controls shall be proportionate to the risk of the use case and may include, where applicable:

* monitoring and override capabilities;
* documented human approval processes for high-impact actions;
* user training; and
* incident record-keeping mechanisms.

These obligations do not apply to standard interactive use of the Platform without autonomous or system-level integration.

#### 4. Notification

If Owkin becomes aware that Client has violated this Acceptable Use Policy or is otherwise misusing the Services, Owkin may immediately restrict, suspend or terminate Client's right to use the Services without liability.

Client shall promptly notify Owkin if it becomes aware of any violation of this Acceptable Use Policy by Client or any Use.


# Support


# Standard support & availability

This page describes the **standard support** available to all K Pro users. Specific commitments.&#x20;

Availability targets, formal priority levels, and response times are defined separately under the applicable **commercial agreement** and take precedence where one is in place.

### 1. Overview & Availability

K Pro is supported during standard business hours.

* **Support Hours:** 9:00 AM to 6:00 PM (Paris Time)

Availability targets and uptime commitments are defined as part of the applicable commercial agreement rather than as a general public commitment.

### 2. Contacting Support

For any questions or issues, please submit a request through our [support form](https://owkinkhelp.zendesk.com/hc/en-us/requests/new). This is the fastest way to reach our team. Your request is routed to the right contact, and updates are shared by email.

### 3. Submitting a Request

To help us investigate quickly, include:

* **Contact Information:** Your email address (and a phone number if you'd like us to be able to reach you directly).
* **Issue Description:** A detailed explanation of the problem.
* **Reproduction Steps:** Step-by-step instructions to reproduce the error (if applicable).
* **Timestamp:** The date and time the incident occurred.
* **Attachments:** You can attach files directly to the request.

If direct upload fails, email files to <support@owkin.com> with your Service Request Number in the subject line.

### 4. Issue prioritization

We review each request based on impact, scope, and whether a workaround exists. This helps us prioritize critical incidents first, especially issues affecting platform access or core workflows.

Formal priority levels and service level commitments are defined separately as part of enterprise agreements.

### 5. Communication & Updates

We keep you informed throughout the lifecycle of your request.

For critical incidents affecting platform availability:

* **Initial Alert:** We notify affected customers as soon as an outage is confirmed.
* **Updates:** We share progress updates until the issue is resolved.
* **Resolution:** We send a confirmation once service is restored.
* **Post-incident Follow-up:** We may share a summary with key contacts for major incidents.

For general support requests:

* **Confirmation:** You receive a Service Request reference number by email.
* **Updates:** We share progress as the request moves forward.
* **Closure:** We close the request once the issue is resolved or the question is answered.

### **6. Maintenance Policy**

When K Pro is hosted by Owkin, we handle all infrastructure and software updates.

* **Planned Maintenance:** Scheduled to minimize impact. Enterprise customers are notified in advance in accordance with their commercial agreement.
* **Unplanned Maintenance:** Reserved for urgent issues. We notify affected enterprise customers as soon as possible when it is required, and again once it is completed.


# Changelog

## August 2026

### 2028-08-12

**Pseudobulk differential expression:** one cell type vs. rest. K-Pro can now compare a chosen cell type against all other cell types in a single pseudobulk differential expression run, and supports patient\_id as a covariate so patients missing from a contrast group are handled correctly instead of skewing the aggregation.

**Chat file uploads and downloads.** Resolved an issue where file uploads and downloads in chat could stall at 0 bytes. Transfers now stream reliably and support files up to 1 GB, and selecting a large batch of files uploads them in controlled groups so a single failed file no longer holds up the rest.

**File previews and downloads.** Any previewed file can now be saved with a download button in the panel, not just PDFs.

**Skill name shown in chat titles.** Chats started by invoking a skill now keep the skill’s name in their auto-generated title, making it clear from the chat list which skill was used.

### 2026-08-03

**Clinico-Pathological Group Characterization Analysis.** K-Pro can now produce a publication-style clinico-pathological group characterization table summarising a cohort: continuous variables as mean (SD) or median \[Q1, Q3], categorical variables as counts and percentages, with a column flagging missing data. Add a grouping variable to compare groups side by side, with a between-group p-value, standardized mean differences, and per-group distribution histograms.

**Consistent significance markers on plots.** Significance is now shown consistently in both the results table and the corresponding plot.&#x20;

## July 2026

### 2026-07-23

**Copy-number variation analysis (InferCNV)** You can now infer large-scale copy-number gains and losses from single-cell or spatial transcriptomics data against a normal reference. Results display as a chromosome heatmap, with optional clustering of the inferred profiles into putative tumour clones.

**Oncoprint visualisation** K-Pro can now generate an oncoprint showing the mutation and copy-number-alteration landscape across a cohort for a chosen gene panel, with samples ordered so co-altered cases group together.

**Classification result visualisations** Classification analyses now include built-in visualisations of their results.

**Gantt chart export improvements** Resolved an issue where exported Gantt charts could lose bar labels, and added an overall-survival / progression-free-survival maturity legend to PowerPoint exports.

**Smoother response rendering** Improved markdown formatting and spacing while responses stream, and fixed a case where the fullscreen panel stayed open after being closed.

### 2026-07-08

**Dataset access is now project-scoped** Reading a dataset is now limited to members of that dataset's project, giving clearer, project-based control over who can see what.

**Clearer plots with legends and wide labels** Charts with legends or wide axis labels now keep their full margins instead of clipping, fixing cases where swimlane legends were missing from Gantt charts.

### 2026-07-02

**Pathway and gene-signature enrichment analysis** K-Pro can now score pathways and gene signatures across your cohort. Single-sample scoring (ssGSEA) returns an enrichment score per sample for each signature, comparable across patient groups, and pre-ranked GSEA scores a ranked gene list against gene sets, returning a normalized enrichment score and significance per pathway. Signatures can be supplied manually or taken from curated pathway databases.

### June 2026

### 2026-06-26

**Interactive whole-slide viewer** Pathology Explorer now lets you open a whole-slide image and pan, zoom, and view it fullscreen. Tiles stream on the fly, so large slides open in seconds instead of waiting on a full download.

**Streaming responses** Answers and result cards now appear progressively as they are generated, so you can start reading and reviewing results while the rest is still being produced.

**Richer data catalog** The Data Generation catalog is now organized around curated profiles, product details no longer repeat across panels, and the map auto-zooms to your current filters as you refine them.

### 2026-06-19

**scRNAseq Pseudobulk Differential Expression Analysis** K-Pro can now compare gene expression between two patient groups across all cell types aggregated per patient, or within a specific cell type such as malignant cells..

**Export for reports and plots** Reports and plots now include an export button directly in the interface.

### 2026-06-12

**Pathology Explorer** Explore whole-slide pathology images directly in K-Pro, view slide thumbnails, inspect image tiles alongside model predictions, and overlay spatial maps.\
\
**Biomarker Extraction** A new analysis identifies a small set of features that best separates a chosen group of patients and presents them as a simple decision tree, giving you candidate biomarkers you can act on.\
\
**Analysis Rationale** Every plot now includes a short explanation of why that analysis answers your question.\
\
Resolved an issue where logout could be blocked during onboarding.

## May 2026

### 2026-05-28

**Unified Agent Orchestration** All K-Pro users now run on the next-generation orchestration pipeline, producing more direct and consistent answers across multi-step questions.

**Gene–Pathway Lookup** K-Pro can now map a gene to the biological pathways it belongs to, and list the genes within a given pathway, using curated MSigDB gene sets.

**Consensus Clustering for Patient Grouping** Multi-omics analyses now include consensus agglomerative clustering to group patients, returning a consensus matrix and cluster labels. Silhouette and purity scores are reported alongside scatter and consensus clustering analyses to help assess cluster quality.

**Gene Availability Checks** Analyses now verify that each requested gene is present in the dataset and notify you when one is missing, instead of failing silently.

**Pathology Explorer: Larger Slide Uploads** Whole-slide images up to 4 GB are now supported (previously 2 GB), preventing large slides from being silently rejected.

**Clearer Context Compaction** When long conversations are summarised to stay within context limits, a live progress bar is now shown instead of the raw summariser output.

### 2026-05-20

**Pathology Explorer** Flexible cohort name formats Pathology Explorer tools now accept both TCGA-LUAD and TCGA\_LUAD cohort name formats, preventing silent tool call failures when the standard hyphenated format is used.

**Multi-omics** Corrected oncogenic variant handling Mutation frequency analysis now correctly treats the oncogenic column as a boolean flag, fixing plots that previously misgrouped oncogenic vs. non-oncogenic variants.

**Login reliability** Resolved a rare login loop that could occur when multiple authentication requests fired in parallel.

### 2026-05-06

**Prior Knowledge Database: Tool Restored** Fixed an outage affecting the Bioknowledge Navigator / Prior Knowledge Database tool that prevented queries against curated knowledge sources between late April and early May. Authentication to underlying data has been hardened to prevent recurrence.

**Claude 4.x Model Upgrade** All K-Pro services (Bioknowledge Navigator, Multi-Omics, orchestration) now run exclusively on Claude 4.x. Earlier Sonnet 3 / 3.5 / 3.7 and Haiku 3 references have been deprecated, providing more capable reasoning across the platform.

## April 2026

### 2026-04-28

**Consensus Replaces Literature Review** A new evidence search tool, Consensus, is now live for K-Pro Free users. Compared to the previous Literature Review service: significantly wider research article coverage, with most articles accessible as full text rather than abstracts only; new filters for recency (e.g. publications from the last year) and top-tier journals. The Explore Data tab now points to Consensus as the literature source.

### 2026-04-27

**Improved Agent Reasoning for K-Pro Free** K-Pro Free users now run on a next-generation agentic pipeline with flat orchestration. The new pipeline streamlines how queries are routed across tools, producing more direct, consistent answers — especially for multi-step clinical questions. Plotter generation has been updated in parallel: cleaner clinical summaries, better population plots, and improved alignment between the data you ask about and the visualization returned.

**Onboarding Refresh** The post-login onboarding experience has been redesigned with a single-video modal and a new onboarding quiz that captures user persona for tailored prompt suggestions. Both new and returning users see the redesign.

**Personalized Suggested Prompts** Suggested prompts on the chat landing page are now personalized based on user persona, surfacing more relevant starter queries.

**Longer Conversations Stay Coherent** New context compaction automatically summarises older conversation history when approaching the token limit, preserving context across long analytical sessions.

**Pathology Explorer: GigaTIME Immune Biomarker Features** New tools expose GigaTIME virtual multiplex immunofluorescence (mIF) data for TCGA-LUAD: 21 mean channel expression values per slide via `get_gigatime_features`, marker grid display via `display_gigatime_map`, and HE-vs-marker side-by-side overlays via `display_gigatime_overlay`. Survival analysis and cohort description tools have been extended to support categorical Kaplan-Meier stratification on Gigatime data.

### 2026-04-09

**Streaming Responses** Agent responses now stream in real time as they are generated, reducing perceived latency for long answers.

**Explore Data: More Datasets and Refined Filters** Additional datasets are now surfaced in the Explore Data tab, with refined display and filtering for inclusion/exclusion (I/E) criteria.

**MOSAIC Data Access Clarity** K-Pro now correctly explains MOSAIC data access: only MOSAIC Window (a 60-patient public subset) is hosted on EGA. Access to the full MOSAIC dataset requires a data licensing conversation with Owkin.

### 2026-04-03

**TA-Aware Population Plots** Clinical summaries now adapt to the dataset's Therapeutic Area. Oncology queries surface a clinical summary table, an indication-specific summary, and Kaplan-Meier curves for Overall Survival (OS) and Progression-Free Survival (PFS). Immunology & Inflammation studies (Crohn's disease, UC) display a clinical summary table alongside a sample-level disease activity score summary and biopsy site information.\
\
**Authentication Upgrade:** Login flows have been migrated from Amplify to a standard OIDC protocol, improving security and reliability. Microsoft Entra ID is now supported as an identity verifier for enterprise environments.\
\
**Chat History: Infinite Scrolling** The chat history page now supports infinite scrolling for easier navigation through long conversation histories.

## March 2026

### 2026-03-20

**Therapeutic-Area-Aware Clinical Summaries** Clinical summaries now automatically select the most relevant plots based on the therapeutic area — Kaplan-Meier curves for oncology, variable descriptions for immunology and inflammation datasets.

**Variant-Level Table** A new variant-level table is now available in multi-omics analyses, enabling exploration of genomic variant data at the individual variant level.

**Procedure Event Table** Patient procedure events are now available as a dedicated table, expanding the clinical data accessible for analysis.

### 2026-03-02

**K Data Catalog: Improved Dataset Discovery** Enhanced dataset metadata and improved navigation make it easier to explore and understand available datasets directly from the platform.

**Multi-Omics: Cross-Center Cohort Grouping** Multi-center cohorts can now be grouped at the indication level, removing center-level fragmentation for cleaner population-level analyses.

## February 2026

### 2026-02-18

**Clinical Summaries: Multi-Dataset Queries Fixed** Resolved an issue preventing clinical summaries from querying across multiple datasets in a single session.

### 2026-02-12

**Improved PDF Report Export** Reports now render special characters correctly and plots are placed in order for more accurate exports.

## January 2026

### 2026-01-15

**Pathology MCP Explorer: Tile-Level Visualization**

* Display tiles with overlay of cell predictions.
* 13 MCP tools total (up from 12)

### 2026-01-14

**Improved chart formatting** Chart titles and axis labels now use consistent capitalization across all visualizations.

## December 2025

### 2025-12-31

**Pathology MCP Explorer: Tile-Level Analysis**

* Filter and sort individual tiles within a slide by any histomics feature.
* Get detailed statistics (mean, std, min, max, percentiles) for any feature across a cohort.

**Guided Prompts** — 7 pre-built prompts to help you get started:

* Refine patient stratification
* Visualize whole slide images
* Assess biomarker potential
* Export quantitative data
* Explore Pathology Explorer capabilities
* Visualize tiles with predictions
* Get cohort-level statistics

Improvements:

* Feature descriptions now link to detailed documentation.
* Better timeout handling for large slide operations.
* 12 MCP tools total (up from 9)

### 2025-12-30

**PDF report export** Reports can now be downloaded as PDF files for easy sharing and archiving.

**Bulk vs. spatial transcriptomics concordance** New visualization compares bulk RNA-seq with spatial transcriptomics expression for the same genes.

**Gene expression stratification in survival analysis** Kaplan-Meier plots now support patient stratification by bulk gene expression levels (high vs. low expressors).

**Resolved visualization timeouts** Gantt visualizations no longer time out on large datasets.

### 2025-12-17

**MOSAIC V5 single-cell data** Added V5 single-cell RNA-seq for colorectal cancer with improved UMAP dimensionality reduction computed at the indication level.

**Enhanced visualization clarity** Sankey and Gantt plots now include legends with consistent color coding. Sankey plots auto-prune less relevant pathways. Single-cell plots have improved colorbar layouts.

**Improved spatial and single-cell plot accuracy** Updated plot descriptions improve how K Pro interprets analysis requests.

**Fixed Sankey plot grouping** Corrected patient grouping errors in treatment pathway visualizations.

**Fixed single-cell annotations** Resolved cell type mapping errors.

### 2025-12-05

**Treatment pathway visualization** New Sankey diagrams show patient flows across lines of therapy with transitions and clinical responses.

**Corrected volcano plot captions** Captions now correctly display gene/protein counts instead of sample counts.

### 2025-12-01

**Pathology MCP ExplorerInitial release.**

* 9 MCP tools for histomics analysis
* Support for 6 cell types across TCGA cohorts
* Survival analysis with Kaplan-Meier curves
* Slide thumbnail visualization
* Histomics data download via presigned URLs

## September 2025

### 2025-09-16

**Cell expression percentages in spatial analysis** Spatial transcriptomics now shows the percentage of cells expressing a target per patient, broken down by cell type.

**Improved platform capability guidance** K Pro now more accurately describes its own capabilities when asked.

### 2025-09-02

**In-platform feedback button** Added "Contact us" button for direct feedback and support requests.

**Resolved multi-omics display errors** Fixed visualization rendering issues.

## August 2025

### 2025-08-07

**MOSAIC datasets in data explorer** MOSAIC datasets now appear on the Explore Data page for easier discovery.

**Clearer product identification** UI now clearly indicates which product you are using.

**Improved multi-omics grouping** Fixed cohort comparison reliability.

## July 2025

### 2025-07-30

**Explore Data tab** New interface for browsing available datasets (TCGA, MOSAIC Window) before starting analyses.

**Advanced filtering and grouping** Filter on any indication-specific clinical variable. Group patients by histomics variables such as cell density from H\&E images.

**MOSAIC V5.3 release** Added clinical variables and manual annotations for the full MOSAIC cohort.

### 2025-07-12

**Sample-level multi-omics analysis** Plots can now be generated at the sample level to reveal intra-patient heterogeneity.

### 2025-07-03

**Gene signature support across all plot types** All visualizations now support custom multi-gene signatures. Define pathway-level scores (e.g., `[PTGS2, PTGES, PTGES2, PTGES3]`) and analyze combined expression.

## June 2025

### 2025-06-20

**Single-cell co-expression analysis** New dotplot visualizes gene pair co-expression at single-cell resolution with Jaccard index scoring.

**Chat session deletion** Delete individual sessions for better organization.

**Gene signature support (initial release)** Bulk violin, stratified violin, and UMAP plots now support gene signatures.

**Suggested prompts** Landing page now shows example queries for faster onboarding.

**Fixed bulk RNA-seq calculations** Corrected log2(TPM+1) transformation errors.

**Improved population summaries** Resolved demographic table display issues.

Questions or feedback? Contact <support@owkin.com>


# FAQ

Below are answers to the most commonly asked questions about K Pro. If you don't find your answer here, please contact us via the [support form](https://owkinkhelp.zendesk.com/hc/en-us/requests/new?tf_anonymous_requester_email=ewa.kondratowicz-ext@owkin.com\&tf_34057044510353=https://k.owkin.com/chat).

### Getting Started

**What is K Pro and how is it different from a general-purpose LLM?** K Pro is Owkin's agentic AI platform purpose-built for biomedical research. Unlike generic LLMs, K Pro (1) grounds every answer in queried patient data and validated scientific sources rather than parametric memory alone, (2) provides access to exclusive multimodal datasets such as MOSAIC, (3) generates interactive, publication-ready visualizations in real time, (4) coordinates specialized AI agents — each designed for a distinct research task — through an orchestration layer, and (5) is engineered by scientists for biomedical workflows.

**What are the available service tiers?** K Pro is available in four tiers — Free, Light, Standard, and Premium — each unlocking additional agents, datasets, and collaboration features. K Pro Free gives individual access to the Analyze Agentic Space and public datasets (MOSAIC Window, TCGA). Light adds custom data upload (BYOD) and CCPA compliance. Standard introduces team collaboration and the Activate space. Premium unlocks the Amplify space, dedicated onboarding, SSO, RBAC, and priority support. See the [Service tiers](https://docs.owkin.com/getting-started/service-tiers) page for a full comparison.

**What are Agentic Spaces and K Agents?** Agentic Spaces are thematic groupings of AI agents aligned to phases of the research and drug development pipeline. The three spaces are **Analyze** (literature review, multimodal exploration, gene knowledge), **Activate** (biomarker identification, patient stratification), and **Amplify** (digital twin enrichment, clinical trial strategy). Each space contains specialized K Agents — for example, the Literature Navigator, Multimodal Explorer, and Knowledge Explorer in Analyze. Agents are orchestrated automatically based on the user's query.

**How do I get the best results from K Pro?** Formulate specific, precise queries rather than broad questions. Specify the data types you want to explore (genomic, clinical, spatial, etc.), request visualizations for complex relationships, and use follow-up suggestions to deepen your analysis. Progressing from exploratory to detailed questions yields the best outcomes.

### Data & Datasets

**What datasets are available in K Pro?** All tiers include access to TCGA (20,000+ samples across 33 cancer types) and the MOSAIC Window (60 patients across 5 cancer types). Paid tiers can access the full MOSAIC dataset (2,200+ patients, 9 cancer types, 5 modalities) and additional datasets from the data catalog. From the Light tier onward, users can also upload their own proprietary data (BYOD).

**What data modalities and file formats does K Pro support?** K Pro supports clinical data (`.csv`, `.tsv`, `.xlsx`), bulk RNA-seq (`.txt`, `.tsv`, `.csv`, `.h5ad`), single-cell/nuclei RNA-seq (`.mtx`, `.h5`, `.h5ad`, `.rds`), spatial transcriptomics (`.mtx`, `.h5`, `.h5ad`, `.rds`), whole exome/genome sequencing (`.vcf`), proteomics (`.txt`, `.tsv`, `.csv`, `.h5ad`), H\&E histology (`.tif`, `.tiff`, `.svs`, `.dcm`, `.ndpi`, `.mrxs`), and immunohistochemistry (same imaging formats). See the [Supported modalities](https://docs.owkin.com/data-in-k-pro/k-pro-data-model-and-technical-references/supported-modalities) page for full details.

**How do I upload and prepare my own data?** From the Light tier onward, you can upload proprietary datasets through the Bring Your Own Data (BYOD) capability. Data preparation can be handled in two ways: (1) the **Data Transformation Agent (DTA)**, which provides automated validation and standardization for supported formats, or (2) **expert curation services** for complex datasets, custom experimental formats, or modalities not yet supported by the DTA (e.g., certain imaging data or spatial transcriptomics). See [Preparing your data](https://docs.owkin.com/data-in-k-pro/integrating-your-data-to-k-pro/preparing-your-data) for guidance.

**Can I connect K Pro to my own cloud storage or data platform?** Yes. K Pro supports integration with customer-managed storage solutions, including dedicated cloud accounts and buckets (e.g., AWS), as well as enterprise data platforms such as Databricks and Snowflake. These connections use secure cross-account access patterns to maintain data isolation and governance. See [Connecting your data sources](https://docs.owkin.com/data-in-k-pro/integrating-your-data-to-k-pro/connecting-your-data-sources) for details.

**What is the AI-Readiness Maturity Model?** Owkin's 6-level framework (Level 0–5) for assessing how well a dataset is prepared for use with K Pro. It ranges from uncontrolled data with no governance (Level 0) to fully traceable, AI/ML-optimized datasets with reproducibility standards (Level 5). The model helps organizations understand what steps are needed to make their data AI-ready. See [The AI maturity model](https://docs.owkin.com/data-in-k-pro/k-pro-data-model-and-technical-references/the-ai-maturity-model) for the full framework.

**How does K Pro enrich my data?** Through its data enrichment pipeline, K Pro can augment datasets with AI-derived features including spatial gene expression predictions from H\&E images, cell-type deconvolution at near single-cell resolution, nuclear morphology analysis (cell counts, densities, shape metrics), and ligand-receptor interaction modeling for cell communication analysis. See [Data enrichment](https://docs.owkin.com/data-in-k-pro/integrating-your-data-to-k-pro/data-enrichment) for details.

### Pathology Explorer

**What is Pathology Explorer?** Pathology Explorer is an AI-powered tool that transforms H\&E whole-slide images into granular, queryable insights. Trained on over 200,000 annotations, it detects and classifies 6 cell types (lymphocytes, neutrophils, eosinophils, plasmocytes, fibroblasts, and cancer cells) and produces quantitative features including cell counts, densities, nuclear morphology, spatial co-occurrence, and TIL diffusivity. It currently supports 27 TCGA tumor cohorts.

**How do I access Pathology Explorer through Claude?** Pathology Explorer is available as an MCP (Model Context Protocol) integration. To connect it, add Owkin's MCP server (`https://mcp.k.owkin.com/mcp`) as a Custom Connector in Claude.ai or Claude Desktop (requires a paid Claude plan). You will be prompted to authenticate with your Owkin account via OAuth. Once connected, you can invoke Pathology Explorer tools directly from your Claude interface. See the [Pathology Explorer getting started](https://docs.owkin.com/core-features-and-usage/pathology-explorer-mcp-ai-powered-tissue-analysis/getting-started) guide for step-by-step instructions.

**Can I export data from Pathology Explorer?** Yes. Pathology Explorer supports data export in Parquet format, and slide images can be downloaded through presigned URLs.

### Visualizations

**What types of visualizations does K Pro support?** K Pro supports over 17 chart types spanning clinical data (Kaplan-Meier survival plots, Gantt treatment timelines, Sankey treatment flows), bulk RNA-seq (violin plots, heatmaps, UMAP, pairwise correlation, differential expression), single-cell RNA-seq (cell-type proportion, dot plots, co-expression, UMAP), spatial transcriptomics (slide displays, Moran's I, cell-type co-occurrence, deconvolution bar charts), histomics (cell-type proportion), and cross-modal concordance plots. See [Visualisation capabilities](https://docs.owkin.com/core-features-and-usage/overview/visualisation-capabilities) for a full catalog.

**Can I export plots and analyses from K Pro?** Yes. K Pro generates interactive Plotly visualizations that can be reviewed and exported. The underlying code used to generate each plot is also available for inspection and reproducibility.

### Data Confidentiality & Protection

**How does K Pro protect confidential data?** Only chat history is saved, and it is linked to authenticated users within their organization. The database is a managed RDS instance on Owkin's AWS account, accessible solely through secure service credentials and network policies. Data is stored in an efficient, column-oriented file format optimized for secure storage and retrieval.

**How does K Pro guard against IP leakage?** Data is segregated by customer. Employee access is restricted to those with operational maintenance roles. Security includes 24/7 monitoring across multiple protective layers.

**Can anyone see the data uploaded to K Pro?** No. Uploaded data is visible only to you and your organization members — never shared externally.

**Where is my data hosted?** Data is stored on Owkin's managed, secure cloud infrastructure. For customers with their own cloud subscriptions, data can reside in a separate cloud account maintained distinct from K Pro's infrastructure, with access enabled through secure cross-account access patterns. See [Infrastructure and hosting](https://docs.owkin.com/data-in-k-pro/privacy-and-compliance/infrastructure-and-hosting) for more details.

### Transparency & Explainability

**Can I access the codebase and decision-making process of K Pro?** While K Pro's codebase is proprietary, users can inspect the reasoning process undertaken by agents. Generated plots and the code behind them are available for review, providing full transparency into how results are produced.

**How can I make sure that the scientific conclusions generated by K Pro are evidence-based and accountable?** Results are grounded in authoritative sources like PubMed and validated biological knowledge bases. K Pro provides complete provenance for all outputs — users can trace recommendations back to original data sources and citations through explainability features that log all data sources, model decisions, and reasoning steps.

**How does K Pro handle data integration and standardization across diverse datasets?** K Pro uses specialized tools and agents designed for dataset diversity. The orchestration layer captures user intent and calls appropriate tools to query scientific literature, gene databases, and multi-omics cohorts. Under the hood, data is harmonized following an OMOP-like schema and preferred ontologies to ensure consistent querying across heterogeneous sources.

### Citations & References

**How does Owkin ensure proper and deterministic retrieval of citations and references?** A RAG (Retrieval-Augmented Generation) system verifies the existence and relevance of cited PubMed articles. K Pro's Literature Navigator queries over 22 million PubMed abstracts through semantic search, and internal evaluation against public benchmarks monitors alignment between generated answers and cited sources.

### Trust & Scientific Rigor

**How does Owkin measure and prevent hallucinations?** K Pro employs multiple layers of protection: (1) daily monitoring of Tool Call Accuracy (TCA) — the percentage of interactions where the correct tool is identified and called with accurate parameters, (2) RAG-based citation validation ensuring all referenced PubMed IDs correspond to real articles, (3) tool-based grounding that anchors analysis in modality-specific AI models and actual patient data rather than LLM parametric memory, and (4) automated evaluation tests using curated questions run on a regular basis.

**How does Owkin prevent bias in K Pro recommendations?** Bias audits occur during development and testing phases. Models are evaluated across diverse demographic and clinical datasets with independent validation from academic and clinical partners. Post-deployment continuous monitoring identifies emerging sources of bias. Transparency features enable users to trace data sources and recommendation rationale.

**Who oversees K Pro's AI ethics?** Owkin engages in active collaboration with external stakeholders across regulatory, academic, and clinical domains, including regulatory bodies and bioethics groups. The platform is built with a "biologists building for biologists" philosophy, involving researchers as beta testers and early users to ensure practical alignment with scientific workflows and ethical standards.

### LLM & Technology

**How would different LLMs or versions of the same LLM affect final outcomes?** Outcomes depend on the LLM's capability to understand questions and call appropriate tools. Newer LLM versions typically perform better. Switching between LLMs requires adjustments to prompting and context engineering. K Pro's agentic architecture means that the specialized tools and data pipelines remain constant regardless of the underlying LLM — the model orchestrates, but the scientific analysis is performed by dedicated tools and models.

**Does Owkin own its full end-to-end technology stack? If not, how does Owkin vet its vendors?** Much of the technology is developed in-house, including Owkin's proprietary foundation models (iBOT, H0, H0-mini) and the HIPE model for cell segmentation. Select vendor partners undergo rigorous security assessments and contractual alignment. Owkin maintains ISO 27001:2022 (information security) and ISO 13485:2016 (medical device quality) certifications, ensuring GDPR and HIPAA compliance.

### Compliance & Certifications

**What security certifications does Owkin hold?** Owkin is certified to ISO 27001:2022 (information security management) and ISO 13485:2016 (medical device quality management). The platform is designed for full compliance with GDPR (EU/UK), HIPAA (US), and CCPA (California, from Light tier onward).

**Can K Pro be deployed in my own infrastructure?** K Pro supports flexible deployment models. For customers with their own cloud environments, data can reside in a dedicated cloud account separate from K Pro's infrastructure, with secure cross-account access. The implementation lifecycle includes a specification phase to define deployment architecture and security protocols tailored to your requirements. See the [Implementation lifecycle](https://docs.owkin.com/terms-and-support/support/implementation-lifecycle-and-communication-protocol) page for the full process.

### Ownership & Legal

**Who owns the output from user prompts in K Pro?** Owkin does not claim rights over user-submitted content or outputs. Owkin retains a license to use the content and output to improve its products and services, while protecting academic freedom. See the [Terms and conditions](https://docs.owkin.com/terms-and-support/legal/terms-and-conditions) for the complete legal framework.

### Support & Account

**How do I get support?** Submit requests via the [support form](https://docs.owkin.com/terms-and-support/support/support-and-sla). Support operates during standard business hours (9:00 AM – 6:00 PM Paris Time). Response times depend on priority: P1 (critical, full outage) — 1-hour response, 4-hour resolution; P2 (major issue, no workaround) — 2-hour response, 8-hour resolution; P3 (workaround available) — 2-hour response, 16-hour resolution; P4 (general inquiries) — 2-hour response, 75-hour resolution.

**How do I delete my account?** To permanently delete your K Pro account, contact the Customer Success Team by submitting a request through the [support portal](https://docs.owkin.com/terms-and-support/support/support-and-sla). The team will process your request in a timely manner. Note that certain information may be retained when necessary to meet legal requirements or legitimate operational needs, in compliance with GDPR. See [Account management](https://docs.owkin.com/terms-and-support/support/account-management) for details.


# Implementation lifecycle & communication protocol

This document defines the end-to-end lifecycle for deploying Owkin solutions. It outlines the responsibilities, required documentation, and communication triggers from initial specification through to decommissioning.

### 1. Roles & Responsibilities

* Owkin Deployment Team: Responsible for infrastructure setup, software installation, and initial verification.
* Customer Point of Contact (POC): Responsible for providing access/permissions, coordinating internal validation, and signing off on deliverables.

### 2. The Lifecycle Workflow

#### Phase 1: Definition & Specification (Step 1)

Before technical work begins, the scope and methodology are formally agreed upon.

* Action: Owkin and Customer collaborate to define the deployment architecture (On-Premise vs. Hosted) and support tier.
* Documentation Generated:
* Deployment Specifications Document: Details hardware requirements, network flows, and security protocols.
* Support Service Definition: Defines SLAs, support hours, and incident reporting channels.
* Specific UAT Definition: Defines the acceptance criteria of the implementation
* Communication: Kick-off meeting and formal exchange of the Specification Document.

#### Phase 2: Contractualization (Step 2)

Both parties agree on the terms of the contract, including the scope and methodology

* Action: Owkin initiates a contract to sign with the customer, including the documents generated in the previous step in the appendix
* Documentation Generated:
* Contract: Details the mutual obligations of the two parties (i.e. scope, duration, pricing)
* Communication: Contract signed and sent to both parties

#### Phase 3: Readiness & Access (Step 3)

The Customer prepares the environment based on the methodology agreed in Phase 1.

* Scenario A (Customer IS): Customer provisions cloud capabilities and grants agreed IAM permissions to Owkin.
* Scenario B (Owkin IS): Customer provides necessary data feeds or integration credentials.
* Documentation Required:
* Access Credentials: (Transmitted via secure channel).
* Communication Trigger:
* Sender: Customer
* Message: "Readiness Confirmation: Permissions granted and environment ready."

#### Phase 4: Deployment Execution (Step 4)

Owkin takes control to execute the technical setup.

* Action: Owkin’s team provisions infrastructure (if applicable) and installs the software suite according to the Deployment Specifications Document.
* Documentation Generated:
* Installation Log: Internal record of configuration steps.
* Communication Trigger:
* Sender: Owkin
* Message: "Deployment Complete. The system is ready for validation."

#### Phase 5: Validation & Acceptance (Step 5)

This is the critical gate for moving to production.

* Action: The Customer performs User Acceptance Testing (UAT) to verify conformity.
* The Loop:
* If Conformity = Fail: Owkin iterates on the configuration. Deployment returns to Phase 3.
* If Conformity = Pass: The process moves to sign-off.
* Documentation Generated:
* Deployment Acceptance Certificate (DAC): A formal document signed by the Customer acknowledging the specification criteria have been met.
* Communication: Validation feedback calls; Submission of the signed DAC.

#### Phase 6: Support & Operations (Step 6)

Upon signing the DAC, the project transitions from "Build" to "Run."

* Action: Service runs according to the Support Service Definition agreed upon in Step 1 for the duration of the contract.
* Documentation Maintained:
* Incident Reports: Generated per ticket.
* Maintenance Logs: Records of updates/patches.
* Communication:
* Standard Support Channels (Ticketing System / Support Email).
* Quarterly Reviews (if applicable to cover business and service topics).

#### Phase 7: Decommissioning (Step 7)

At the contract end, data privacy and resource cleanup are prioritized.

* Action: Owkin initiates the takedown of deployed software and infrastructure.
* Documentation Generated:
* Decommissioning Report: Confirmation that all data has been wiped and access removed.
* Communication Trigger:
* Sender: Owkin
* Message: "Decommissioning complete. Contract formally concluded."

<br>


# Account management

If you would like to permanently delete your K Pro account and associated personal data, contact our Customer Success Team by [**submitting a request**](https://owkinkhelp.zendesk.com/hc/en-us/requests/new).

Our team will review your request and process it promptly.

In some cases, we may need to retain specific information to comply with legal obligations or for legitimate business purposes. This is done in accordance with applicable data protection regulations, including GDPR.


# Glossary

This glossary provides definitions for key terms used throughout the K Pro documentation. It is non-exhaustive; terms are listed alphabetically within thematic groups for ease of navigation.

### **Platform & Product Terms**

**Activate** An Agentic Space in K Pro designed for translational research tasks such as biomarker identification, patient selection, and cohort stratification. Available from the Standard tier onward, it contains the Deep Patient Explorer and Population Optimizer agents.

**Agentic AI** An AI system capable of autonomously planning and executing multi-step tasks by selecting and coordinating specialized tools and agents in response to a user's intent. K Pro uses an agentic AI architecture to orchestrate complex biomedical analyses.

**Agentic Space** A grouping of K Pro agents designed to address a specific phase of the research or drug development pipeline. The three agentic spaces are Analyze, Activate, and Amplify, each available from a specific service tier.

**Amplify** An Agentic Space in K Pro targeting advanced drug development workflows, including digital twin enrichment and clinical trial strategy. Available in the Premium tier, it contains the Multimodal Twin Enricher and Trial Strategy Explorer agents.

**Analyze** An Agentic Space in K Pro focused on foundational research capabilities, including literature review, multimodal patient data exploration, and gene-level biological knowledge retrieval. Available from the Free tier onward, it contains the Literature Navigator, Multimodal Explorer, and Knowledge Explorer agents.

**Data Enrichment** The process by which K Pro augments uploaded or existing datasets with AI-derived features, including spatial gene expression prediction from H\&E images, cell-type deconvolution, nuclear morphology analysis, and ligand-receptor interaction modeling.

**Deep Patient Explorer** A K Agent within the Activate space that supports biomarker identification and patient stratification by performing in-depth analyses across multimodal datasets.

**Histomics** A digital pathology capability within K Pro that performs AI-driven cell detection, segmentation, and classification on H\&E whole-slide images. It identifies cell types such as lymphocytes, neutrophils, eosinophils, plasmocytes, fibroblasts, and cancer cells, and produces quantitative features including cell counts, densities, and nuclear morphology metrics.

**K Agents** Specialized AI modules within K Pro, each designed to perform a specific class of research task (e.g., literature review, multiomics analysis, clinical data exploration). Agents are orchestrated by K Pro's underlying system to collaborate in response to user queries.

**K Pro** Owkin's agentic AI platform for biomedical research. K Pro provides access to curated multimodal datasets, specialized AI agents, and visualization tools through a natural language interface. Available in four tiers: Free, Light, Standard, and Premium.

**K Pro Free** The no-cost entry tier of K Pro, providing individual access to the Analyze Agentic Space — including the Literature Navigator, Multimodal Explorer, and Knowledge Explorer agents — and public datasets including the MOSAIC Window and TCGA. Does not include data upload, collaboration, or enterprise features.

**K Pro Light** A paid K Pro tier providing individual access to core agentic capabilities, custom data upload (BYOD), and CCPA compliance, in addition to all features available in K Pro Free.

**K Pro Standard** A paid K Pro tier offering team collaboration, multi-user support, and access to the Activate Agentic Space (Deep Patient Explorer and Population Optimizer), in addition to all Light-tier features.

**K Pro Premium** The highest K Pro tier, providing access to the Amplify Agentic Space (Multimodal Twin Enricher and Trial Strategy Explorer), dedicated onboarding, priority support, and enterprise-grade capabilities, in addition to all Standard-tier features.

**Knowledge Explorer** A K Agent within the Analyze space that retrieves gene- and target-level biological knowledge from curated databases, covering protein families, oncogenicity, immune pathways, tractability, and expression profiles.

**Literature Navigator** A K Agent within the Analyze space that performs literature review and gap analysis by querying PubMed abstracts through semantic search and retrieval-augmented generation (RAG), returning citation-backed summaries.

**MCP (Model Context Protocol)** An open protocol that enables AI assistants such as Claude to connect to external tools and data sources through a standardized interface. Owkin's MCP integration exposes the Pathology Explorer toolset to Claude.ai and Claude Desktop users.

**MCP Connector** The configuration entry point used to link Claude.ai or Claude Desktop to Owkin's MCP server at `https://mcp.k.owkin.com/mcp`. Once connected, users can invoke Pathology Explorer tools directly from their Claude interface.

**Multimodal Explorer** A K Agent within the Analyze space that enables interactive exploration of multimodal patient datasets — including clinical, transcriptomic, spatial, and histological data — through natural-language queries and automated visualization.

**Multimodal Twin Enricher** A K Agent within the Amplify space that enriches datasets by generating AI-derived features — such as predicted spatial transcriptomics, cell-type deconvolution, and cell communication profiles — from existing modalities.

**Orchestration** The process by which K Pro's underlying system interprets a user's intent, selects the appropriate agents and tools, sequences their execution, and consolidates the results into a coherent response. Orchestration incorporates trust and verification mechanisms, including RAG-based citation validation, tool-based grounding against modality-specific AI models, and data-anchored analysis to ensure outputs are derived from actual patient data.

**Pathology Explorer** An AI-powered tool available through K Pro and as an MCP integration, designed to transform H\&E whole-slide images into granular, queryable insights. Trained on 200,000+ annotations, it detects and classifies 6 distinct cell types (lymphocytes, neutrophils, eosinophils, plasmocytes, fibroblasts, and cancer cells), supports spatial analysis of the tumor microenvironment across 27 TCGA tumor cohorts, and enables data export in Parquet format.

**Population Optimizer** A K Agent within the Activate space that assists with patient selection and cohort optimization for clinical development.

**Trial Strategy Explorer** A K Agent within the Amplify space that supports clinical trial de-risking by analyzing competitive clinical trial landscapes and real-world multimodal patient data.

### **Data & Datasets**

**Bring Your Own Data (BYOD)** A K Pro capability (available from Light tier) allowing users to upload their own proprietary datasets for analysis within the platform. Data must conform to K Pro's supported modalities and file format specifications.

**Bulk RNA-seq (bRNAseq)** Bulk RNA sequencing. A method to assess the overall transcriptome of a tissue sample by sequencing RNA from a mixture of cells. Supported formats in K Pro: `.txt`, `.tsv`, `.csv`, `.h5ad`.

**Deconvolution** A computational method for inferring the proportions of distinct cell types within a mixed sample — such as a bulk RNA-seq measurement or a spatial transcriptomics spot — by leveraging reference single-cell expression profiles. K Pro uses deconvolution extensively in spatial transcriptomics analyses and data enrichment.

**Gene Signature** A curated set of genes whose combined expression pattern represents a specific biological process, cell state, or phenotype. Gene signatures can be used in K Pro visualizations in place of individual genes for more robust biological characterization.

**H\&E (Haematoxylin and Eosin)** A standard histological staining technique applied to whole-slide tissue images. H\&E staining highlights cell nuclei and cytoplasm and is the primary imaging modality used by Pathology Explorer.

**IHC (Immunohistochemistry)** A laboratory technique that uses antibodies to detect specific protein markers in tissue sections. K Pro supports IHC slide images and can extract marker-specific scores (e.g., HER2, ER/PR) through proprietary models. Supported formats: `.tif`, `.tiff`, `.svs`, `.dcm`, `.ndpi`, `.mrxs`.

**Jaccard Index** A similarity metric used in K Pro to quantify co-expression overlap between two genes or gene signatures. Calculated as the proportion of cells (or spots) co-expressing both inputs relative to cells expressing either, it ranges from 0 (no overlap) to 1 (complete overlap).

**Moran's I** A spatial autocorrelation statistic that measures the degree to which a gene's expression is spatially clustered within a tissue section. A high Moran's I indicates that the gene tends to be expressed in localized regions rather than randomly across the tissue.

**MOSAIC (Multi Omics Spatial Atlas In Cancer)** A landmark multi-institutional dataset developed by Owkin in collaboration with academic centers (University of Pittsburgh, Gustave Roussy, Lausanne University Hospital, Erlangen University Hospital, and Charité Berlin). It contains clinical data, bulk RNA-seq, single-nuclei RNA-seq, spatial transcriptomics, WES, and H\&E histology across 9 cancer types and 2,200+ patients.

**MOSAIC Window** A curated, publicly accessible subset of the MOSAIC dataset, included in all K Pro tiers. Contains multimodal data from 60 patients across five cancer types: BLCA, OV, GBM, DLBCL, and MESO.

**Multimodal Data** Data combining multiple biological measurement types (modalities) from the same patient or sample, such as clinical records, genomics, transcriptomics, histology, and spatial omics.

**OMOP (Observational Medical Outcomes Partnership)** A standardized data model for organizing and harmonizing observational health data. K Pro follows an OMOP-like schema to ensure interoperability across heterogeneous clinical datasets.

**Proteomics** The large-scale study of proteins, including their expression levels, modifications, and interactions. K Pro supports proteomics data as a modality, accepting normalized intensity matrices (`.txt`, `.tsv`, `.csv`) and AnnData (`.h5ad`) formats.

**RECIST (Response Evaluation Criteria in Solid Tumors)** A standardized set of rules for assessing tumor response to treatment, classifying outcomes as Complete Response (CR), Partial Response (PR), Stable Disease (SD), or Progressive Disease (PD).

**Single-cell RNA-seq (scRNA-seq)** Single-cell RNA sequencing. A technique that measures gene expression at the level of individual cells, enabling identification of distinct cell populations and states within a tissue. Supported formats in K Pro: `.mtx`, `.h5`, `.h5ad`, `.rds`.

**Spatial Transcriptomics (ST)** A technique that measures gene expression while preserving the spatial location of cells or spots within a tissue section. K Pro supports the Visium Cytassist protocol from 10X Genomics. Supported formats: `.mtx`, `.h5`, `.h5ad`, `.rds`.

**TCGA (The Cancer Genome Atlas)** A publicly funded, comprehensive cancer genomics resource comprising data from 20,000+ primary cancer samples across 33 cancer types. Available in all K Pro tiers.

**TILs (Tumor-Infiltrating Lymphocytes)** Immune cells — primarily T cells and B cells — that have migrated from the bloodstream into a tumor. Their density and distribution, quantifiable through K Pro's Histomics and Pathology Explorer tools, are important prognostic and predictive biomarkers in oncology.

**TLS (Tertiary Lymphoid Structures)** Organized aggregates of immune cells that form within or near tumors, resembling lymph nodes. Their presence is generally associated with favorable anti-tumor immune responses and can be detected through K Pro's digital pathology capabilities.

**TME (Tumor Microenvironment)** The cellular and molecular environment surrounding a tumor, including immune cells, stromal cells, blood vessels, and signaling molecules. Characterizing the TME is a key use case for Pathology Explorer and the MOSAIC dataset.

**TNM Staging** A cancer staging system that classifies tumors based on three factors: the size and extent of the primary Tumor (T), involvement of regional lymph Nodes (N), and presence of distant Metastasis (M). Used alongside stage groupings (I–IV) in K Pro's clinical data model.

**WES (Whole Exome Sequencing)** A sequencing technique that targets the protein-coding regions (exons) of the genome, approximately 1–2% of the total genome, which harbors the majority of disease-causing mutations. Supported format in K Pro: `.vcf`.

**WGS (Whole Genome Sequencing)** A sequencing technique that determines the complete DNA sequence of an organism's genome, including both coding and non-coding regions. Supported alongside WES in K Pro via VCF file format.

### **AI & Technical Terms**

**AI-Readiness Maturity Model** Owkin's 6-level framework (0–5) for assessing how well a dataset is prepared for use with K Pro. Levels range from uncontrolled data (Level 0) to fully traceable and AI/ML-optimized datasets with reproducibility standards (Level 5).

**Coding Agent** An AI capability that can create and execute custom analysis steps to answer a user's question. Unlike predefined workflows, a coding agent adapts its approach to the task at hand, which makes it more flexible for open-ended analysis.

**DESeq2** A widely used statistical method for differential gene expression analysis from count-based RNA-seq data. K Pro employs DESeq2 as part of its bulk RNA-seq differential expression analysis (DEA) workflow.

**Embedding / Latent Representation** A compressed, numerical representation of high-dimensional biological data (e.g., gene expression profiles) produced by machine learning models. Used in K Pro for dimensionality reduction plots (PCA, UMAP, t-SNE).

**Foundation Model** A large-scale AI model trained on broad data that can be adapted to many downstream tasks. Owkin deploys foundation models (including iBOT, H0, and H0-mini) for self-supervised feature extraction from whole-slide images, which are used in data enrichment and spatial prediction workflows.

**GTEx (Genotype-Tissue Expression)** A public resource cataloging gene expression levels across human tissues from healthy donors. K Pro references GTEx data to provide tissue-specificity context for genes of interest.

**Hallucination** In the context of LLMs, the generation of plausible-sounding but factually incorrect or fabricated information. K Pro implements monitoring (Tool Call Accuracy) and technical guardrails (PubMed ID verification via RAG) to detect and reduce hallucinations.

**Harmonization** The process of standardizing and normalizing heterogeneous datasets — across modalities, formats, and sources — so they can be queried consistently by K Pro's agents and tools.

**HIPE Model** An Owkin-developed AI model for cell segmentation and annotation on H\&E whole-slide images. It detects and classifies individual cell nuclei into types (e.g., lymphocytes, cancer cells, fibroblasts), producing quantitative histomics features stored as CSV files.

**Ligand-Receptor (LR) Interaction** A cell communication analysis that models signaling between cells by identifying paired ligand and receptor molecules. K Pro's data enrichment pipeline maps these interactions across diffusion modes (cell contact, secreted, and hormone signaling) and visualizes them as dot plots, chord diagrams, and spatial overlays.

**LLM (Large Language Model)** A deep learning model trained on large text corpora, capable of understanding and generating natural language. K Pro is built on top of LLMs, which interpret user queries and coordinate agent responses.

**OpenTargets** An open-access platform integrating public domain data to enable systematic identification and prioritization of drug targets. K Pro uses OpenTargets data for baseline expression profiling and tractability assessments in gene knowledge exploration.

**PubMed** A freely accessible database maintained by the National Library of Medicine containing over 36 million citations and abstracts of biomedical literature. K Pro's Literature Navigator queries PubMed through semantic search and RAG to provide citation-backed research summaries.

**RAG (Retrieval-Augmented Generation)** A technique that combines a generative LLM with a retrieval system to ground responses in factual, source-verified content. K Pro uses RAG over 22M+ PubMed abstracts to ensure literature citations are valid and relevant.

**Reactome** A curated, open-source database of biological pathways and reactions. K Pro references Reactome immune pathways when providing gene-level biological context through the Knowledge Explorer agent.

**Skill** A predefined product capability designed for a specific task, such as running an analysis, producing a visualization or a complete scientific workflow. Skills use known inputs and standardized logic, which makes them consistent, repeatable, and easier to trust for routine work.

**TCA (Tool Call Accuracy)** An internal monitoring metric used by Owkin to measure the proportion of agent interactions in which the correct tool is identified and called with appropriate parameters. TCA is used to detect systematic agent errors.

**Tool** A specialized product component that performs a defined operation inside K Platform, such as retrieving data, running an analysis, or generating a chart. Tools execute work in a structured and standardized way.

### **Visualization Terms**

**Kaplan-Meier Plot** A statistical visualization of time-to-event data (e.g., overall survival) that shows the estimated survival probability over time for one or more patient groups. Typically includes log-rank test p-values and hazard ratios.

**Oncoprint** A matrix visualization displaying the pattern of genetic alterations (mutations, copy number variants, etc.) across a set of patients and genes. Useful for identifying co-occurring or mutually exclusive alterations.

**UMAP (Uniform Manifold Approximation and Projection)** A dimensionality reduction algorithm used to project high-dimensional biological data (e.g., single-cell gene expression) into 2D or 3D for visual exploration. Also available: PCA and t-SNE.

**Volcano Plot** A scatter plot used to visualize differential gene expression, with statistical significance (−log p-value) on the y-axis and effect size (log fold-change) on the x-axis. Highlights genes that are both statistically significant and biologically meaningful.

**WSI (Whole-Slide Image)** A high-resolution digital scan of a complete tissue section on a glass slide. Supported formats in K Pro: `.tif`, `.tiff`, `.svs`, `.dcm`, `.ndpi`, `.mrxs`.

**Security & Compliance Terms**

**CCPA (California Consumer Privacy Act)** A California state privacy law granting consumers rights over their personal data, including the right to know, delete, and opt out of its sale. K Pro compliance with CCPA is available from the Light tier onward.

**GDPR (General Data Protection Regulation)** The European Union's primary data protection regulation, governing the collection, storage, processing, and transfer of personal data. K Pro is GDPR compliant for EU and UK users.

**HIPAA (Health Insurance Portability and Accountability Act)** A US federal law establishing standards for the protection of health information. K Pro's security architecture is designed to support HIPAA compliance requirements.

**IAM (Identity and Access Management)** The framework of policies and technologies used to control user access to systems and data. K Pro uses IAM to enforce customer-level data segregation and role-based permissions.

**ISO 27001:2022** The international standard for information security management systems (ISMS). Owkin is certified to ISO 27001:2022, demonstrating systematic controls for protecting data confidentiality, integrity, and availability.

**ISO 13485:2016** The international standard for quality management systems in medical device development. Owkin holds this certification, applicable to its AI model development practices.

**RBAC (Role-Based Access Control)** A security model in which system access is granted based on a user's role within an organization. Available in the Premium tier of K Pro.

**SSO (Single Sign-On)** An authentication scheme that allows users to log in once and access multiple applications without re-entering credentials. K Pro supports SSO integration in the Premium tier.

**Note:** This glossary is maintained as a living reference. If you encounter a term not defined here, or believe a definition requires updating, please contact <support@owkin.com>.


# Help Center

<h2 align="center">What can we help you find?</h2>

<p align="center">Browse the topics below or use the GitBook Assistant to ask anything you need help with.</p>

<p align="center"><button type="button" class="button primary" data-action="ask" data-icon="gitbook-assistant">How can we help?</button><a href="https://owkinkhelp.zendesk.com/hc/en-us/requests/new" class="button secondary" data-icon="paper-plane">Contact support</a></p>

&#x20;

<table data-view="cards"><thead><tr><th></th><th></th><th data-hidden data-card-target data-type="content-ref"></th></tr></thead><tbody><tr><td><h4><strong>Getting started</strong></h4></td><td>Everything you need to begin</td><td><a href="/pages/M79vEy5pPyA2JgPF0NTA">/pages/M79vEy5pPyA2JgPF0NTA</a></td></tr><tr><td><h4><strong>Running analyses and working with results</strong></h4></td><td>Get the most out of your prompts</td><td><a href="/pages/XvdaniKQcuMdlkDOigqg">/pages/XvdaniKQcuMdlkDOigqg</a></td></tr><tr><td><h4><strong>Troubleshooting</strong></h4></td><td>Fix common errors and issues</td><td><a href="/pages/sY809Cp8jUdtKtqwMxGt">/pages/sY809Cp8jUdtKtqwMxGt</a></td></tr></tbody></table>

&#x20;

&#x20;

&#x20;


# How do I sign in to K-Pro for the first time?

{% hint style="info" %}
**What this article covers** This article explains how to activate your account and log in to K-Pro, including what to do if your provisional password or activation link has expired.
{% endhint %}

### Step-by-step: First-time login

1. Open the welcome email from Owkin (check your spam folder if you don't see it in your inbox).
2. Click the activation link, or navigate to your K-Pro instance URL (shared by your Owkin contact).
3. On the login page, click "Forgot password" and enter your institutional email address.
4. You will receive a password reset email. Click the link inside to set your own password.
5. Log in with your email and new password.

{% hint style="warning" %}
**Important** Use the institutional email address that your Owkin contact used to create your account. Using a different email will result in a "user not found" error.
{% endhint %}

### My activation link or provisional password doesn't work

Activation links and provisional passwords expire after a limited time. If yours has expired:

* Reply to your welcome email and ask your Owkin contact to resend the activation link.
* Alternatively, go to the K-Pro login page, click "Forgot password", and request a new link directly.
* If neither option works, contact K-Pro support via the "Need Help" button in the platform interface.

***

**✉️ Need more help?**\
If this article did not answer your questions, please [contact our support team](https://owkinkhelp.zendesk.com/hc/en-us/requests/new?ticket_form_id=33609591380241\&brand_id=33635288496017) and we will be happy to help.


# What data is available in K-Pro?

{% hint style="info" %}
**What this article covers** This article describes the datasets available in K-Pro, clarifies how biological data is labelled, and explains what to do when you can't find expected data.
{% endhint %}

### Standard datasets included in K-Pro

* **TCGA** — The Cancer Genome Atlas: pan-cancer genomic and clinical data
* **MOSAIC Window** — Owkin's proprietary multi-omics dataset. Covers **11 cancer indications, 6 data modalities, and 2,716 patients** (clinical, genomics, transcriptomics, spatial transcriptomics, proteomics, histology). *Numbers reflect the May 2026 dataset snapshot — see the live Dataset Catalog for current totals.* Full dataset catalog available at docs.owkin.com
* **Scientific literature search via Consensus** — AI-powered scientific search engine (consensus.app) integrated into K-Pro. Enables literature-backed hypothesis validation, paper discovery on a gene/disease, and scientific background research.

### K-Pro free vs. add-on data scope

K-Pro free includes the public MOSAIC and TCGA scope. Some datasets, indications, and modalities (e.g., additional proteomics for Epkin, client-specific cohorts for Servier) are only enabled as add-ons for specific organizations. Your Owkin contact will have shared the exact list of datasets included in your instance during onboarding.

### Why can't I find certain data?

{% hint style="info" %}
**"I can't find blood or adjacent non-tumor tissue data"**

The MOSAIC dataset focuses on tumoral biopsies. It does not include blood transcriptomics data. Tissue types that might seem equivalent to "adjacent non-tumor" are labelled by cell type (e.g., epithelial, stromal, immune cells) rather than by tissue origin label.

**Example prompt:** "Plot the expression of EGFR in malignant and epithelial cells in lung patients from MOSAIC"
{% endhint %}

{% hint style="info" %}
**"I can't find data for my cancer indication (e.g., lung cancer, PAAD)"**

MOSAIC Window covers 11 cancer indications in total. However, in some trial configurations, only a subset of indications is enabled for your instance. To see the full, up-to-date list of indications available in K-Pro, visit **docs.owkin.com → Dataset Catalog**. If you believe your trial should include an indication that you cannot query, contact your Owkin project manager.
{% endhint %}

***

**✉️ Need more help?**\
If this article did not answer your questions, please [contact our support team](https://owkinkhelp.zendesk.com/hc/en-us/requests/new?ticket_form_id=33609591380241\&brand_id=33635288496017) and we will be happy to help.


# My queries keep failing or returning only literature summaries

{% hint style="info" %}
**What this article covers** This article explains why K-Pro queries may fail or fall back to literature summaries, and provides guidance on how to write prompts that yield multi-omics analysis results.
{% endhint %}

### Why does this happen?

* **The query intent is ambiguous** — if K-Pro cannot clearly map your question to a dataset and analysis type, it may return a literature-based answer instead of running a data analysis.
* **The requested data is not available in your instance** — for example, querying a cancer indication not included in your subscription scope.
* **A temporary service issue** — the underlying analysis services can experience intermittent errors.

### How to write better prompts

Be as specific as possible about the dataset, the analysis type, and the biological entities involved.

{% hint style="success" %}
**Good prompt examples:**

* "Plot EGFR expression in lung adenocarcinoma patients from TCGA"
* "Compare survival outcomes for high vs. low CD8+ T-cell infiltration in TCGA breast cancer"
* "Show the spatial distribution of CD36 expression in GBM samples from MOSAIC"
  {% endhint %}

{% hint style="danger" %}
**Prompts that may trigger fallback to literature:**

* "Tell me about EGFR" → too vague, no dataset or analysis specified
* "What happens in lung cancer?" → general question, likely routed to literature
  {% endhint %}

### What to do if errors persist

1. Try rephrasing your query using the examples above.
2. Check the data availability for the indication or dataset you are querying (see Article 2).
3. Wait a few minutes and retry — intermittent service errors often resolve automatically.
4. If the issue persists, submit a support request via the "Need Help" button with the exact prompt and error message.

***

**✉️ Need more help?**\
If this article did not answer your questions, please [contact our support team](https://owkinkhelp.zendesk.com/hc/en-us/requests/new?ticket_form_id=33609591380241\&brand_id=33635288496017) and we will be happy to help.


# How do I refine, iterate, and export a chart?

{% hint style="info" %}
**What this article covers** This article explains how to modify a chart after it has been generated, iterate on it using natural language follow-up prompts, and export it for use in presentations or publications.
{% endhint %}

When K-Pro returns a chart or visualisation, it is fully interactive — you are not locked into the initial result. You can refine, filter, or re-format it simply by typing a follow-up instruction in the chat.

### How to iterate on a chart

After a chart is generated, type a follow-up instruction in the chat to modify it. K-Pro understands context from the previous message, so you do not need to repeat the full prompt.

{% hint style="success" %}
**Iteration prompt examples:**

* "Filter this plot to only show stage III patients"
* "Split by EGFR mutation status"
* "Change the colour palette to be colour-blind friendly"
* "Remove the outliers and re-plot"
* "Add a statistical significance annotation"
* "Show the same data as a violin plot instead of a box plot"
  {% endhint %}

### How to export a chart

Once you are satisfied with a chart, you can download it for use in presentations, reports, or publications:

1. Hover over the chart in the K-Pro interface.
2. Click the **download icon** (or the ⋮ menu) that appears in the top-right corner of the chart.
3. Select your preferred format: **PNG** (for presentations) or **SVG** (for publication-quality vector graphics).

{% hint style="warning" %}
**Troubleshooting:** If the download icon is not visible, try hovering directly over the chart area. On some browsers, the icon only appears on hover. If the issue persists, try a different browser or submit a support request.
{% endhint %}

### Tips for better charts

* **Start specific**: name the gene, dataset, cancer type, and chart type in your first prompt for the most relevant first result.
* **Iterate progressively**: one change at a time works better than combining multiple instructions in a single follow-up.
* **Reference "this chart" or "this plot"** in follow-up prompts so K-Pro knows you are modifying the existing result rather than starting a new analysis.

***

**✉️ Need more help?**\
If this article did not answer your questions, please [contact our support team](https://owkinkhelp.zendesk.com/hc/en-us/requests/new?ticket_form_id=33609591380241\&brand_id=33635288496017) and we will be happy to help.


# "User quota exceeded" error

{% hint style="info" %}
**What this article covers** This article explains K-Pro's usage limits and what to do when you reach your query quota.
{% endhint %}

If you see a "User quota exceeded" or similar error message when submitting a query in K-Pro, it means you have reached the usage limit for your account during the current period.

### Why is there a usage limit?

K-Pro enforces per-account usage limits (measured in queries or tokens) to ensure fair access across all users. These limits are defined in your subscription agreement and are not individually adjustable by default.

### What can I do?

* **Wait for your quota to reset.** Limits may reset on a rolling basis. Contact support to find out when your quota resets.
* **Contact support.** If you believe you have hit the limit unexpectedly or need continued access for urgent research, submit a support ticket via the "Need Help" button.

{% hint style="info" %}
We are working on making usage limits clearer within the platform interface so you can track your remaining quota. Thank you for your patience.
{% endhint %}

***

**✉️ Need more help?**\
If this article did not answer your questions, please [contact our support team](https://owkinkhelp.zendesk.com/hc/en-us/requests/new?ticket_form_id=33609591380241\&brand_id=33635288496017) and we will be happy to help.


# "Operation not allowed" error

{% hint style="info" %}
**What this article covers** This article explains the "operation not allowed" error in K-Pro, the most common reasons it appears, and what to do when it does.
{% endhint %}

{% hint style="warning" %}
**Important note for support teams** Permissions were migrated from the legacy `permissions.yaml` to OpenFGA (`backend/permissions/local_seed.yaml`) in the April–May 2026 restructure. If the user-facing error wording has changed from "operation not allowed" to a different string, update this article's title and first sentence accordingly.
{% endhint %}

If K-Pro returns a popup or message saying **"Operation not allowed"**, it means the action or query you attempted is outside the boundaries of what your account is authorised to do. This is a permissions or scope error, not a technical fault.

### Common causes

* **Querying data you do not have access to** — K-Pro enforces strict data segregation by organisation. If you attempt to access a dataset, indication, or modality not included in your subscription, the request will be blocked.
* **Requesting an action outside the platform's scope** — Certain operations (e.g. writing to external systems, accessing non-cancer literature in some configurations, or uploading files) may not be enabled for your account.
* **A prompt that previously worked now fails** — This can happen if the underlying data or your subscription scope has changed, or if the query is now routing to a different agent.

### What to do

1. Check whether the data or feature you are requesting is included in your subscription.
2. Try rephrasing your request to be more specific about the dataset and type of analysis.
3. If the same prompt worked before and now fails, contact your Owkin project manager — your subscription scope or data configuration may have changed.
4. If none of the above applies, submit a support request via the **Need Help** button with the exact prompt and the error message. The support team can check your permissions configuration.

{% hint style="info" %}
This error is not a platform crash. Your data and previous analyses are unaffected. You can continue using K-Pro for other queries while waiting for a resolution.
{% endhint %}

***

**✉️ Need more help?**\
If this article did not answer your questions, please [contact our support team](https://owkinkhelp.zendesk.com/hc/en-us/requests/new?ticket_form_id=33609591380241\&brand_id=33635288496017) and we will be happy to help.


