Create and register proteins in Benchling Biologics

Denise
Denise
  • Updated

Benchling Biologics registration takes DNA or amino acid sequences as input and automatically creates and links all component entities at once, including domains, chains, variable pairs, and the protein entity itself, so you don't have to build and link each one manually. Use this article to prepare your spreadsheet, register a batch of antibodies, and review the entities Benchling creates.

During registration, Benchling annotates sequences with complementary-defining regions (CDRs), framework regions (FRs), and germline genes. It then validates each sequence for structural soundness, checking factors like CDR length, domain type, and light chain pairing. If any sequence fails validation, nothing in that batch is registered, and you'll receive an error report to fix and resubmit. 

Note: Benchling Biologics requires a separate license and must be configured before registration is possible. See Configure Benchling Biologics schemas and formats to get started.

Understand the registration flow

Protein registration is an asynchronous process. After you submit, Benchling runs a pipeline that creates and links all entities, annotates sequences, and validates them against the selected germline. You'll receive an email notification when the job is complete — either with a link to your summary dataset or a CSV of errors to fix. Registration is designed to be all-or-nothing — if Benchling detects errors during validation, nothing is created and you'll receive an error CSV. 

Benchling creates the following entity types during a single registration run:

  • DNA domains and/or AA domains
  • DNA chains and/or AA chains
  • DNA variable pairs and/or AA variable pairs
  • Protein entities

Starting points: You can start from either domains or chains, and provide sequences as DNA bases, amino acid residues, or Registry IDs of existing entities. Benchling detects the input type automatically. If starting from domains, Benchling concatenates chains from those domains. If starting from chains, Benchling decomposes chains into domains. In both cases, DNA inputs are automatically translated to AA.

Note: AA inputs are not back-translated to DNA. If you need DNA sequences for your components, prepare and register them before or after protein registration.

Registering proteins with tags or cleavage sites: If your sequences include a tag (for example, a His-tag) or a cleavage site (for example, a protease recognition sequence), Benchling can identify and label these as sub-elements during registration.

  • Starting from domains: provide the sub-element sequence in its dedicated column in the template (for example, an N-terminal sub-element column).
  • Starting from chains: you don't need a separate column. Benchling identifies the sub-element sequence by exclusion — any part of the chain that doesn't match a known domain — and matches it against sub-elements already registered on your tenant.

In both cases, the sub-element must already be registered on your tenant for Benchling to recognize and label it. If you're using a new tag or cleavage site that hasn't been registered yet, register it first as an AA or DNA sequence using your tenant's sub-element schema. See Create and manage entities with the Registry for details. Once a sub-element is registered, you don't need to re-register it for future antibody registrations.

Note: Labeling a sub-element as a cleavage site marks its position on the chain. It does not create a separate cleaved protein entity — Benchling does not currently support automatically generating a cleaved child entity from a cleavage site.

Pre-requisites & limitations 

Each format requires a specific spreadsheet template with format-specific column headers. Download the template from within the registration flow.

  • Column headers must match the template exactly. Column order doesn't matter
  • Each column must use a single input type across all rows — raw DNA bases, amino acid residues, or an existing registered entity name or Registry ID. Input types can vary between columns, but not within a single column
  • You can register a maximum of 1,000 antibodies per run

Custom fields on protein entities can't be set during registration. You can update additional metadata on registered entities after registration.

Notebook entry sections must be enabled on your tenant for the receipt feature during registration to work.

Register Antibodies

Configure your registration

  1. Click Global create, hover over Protein, and click Create proteins
  2. Select a format from the format picker
  3. Select the associated schema — Benchling auto-selects the schema if only one exists for that format
  4. Select a germline and numbering scheme from the Germline dropdown. Available germlines and schemes are shown in the dropdown.
  5. Select your starting protein component: Domains or Chains
  6. If starting from domains, select whether you'll provide variable pair entities or individual VH and VL domains
  7. Click Next
  8. Click Download template to get the prescribed columns for your format and starting component type
  9. Fill in the spreadsheet with your sequence data

Note: If your format includes tags or cleavage sites and you're starting from domains, the template includes a column for each sub-element position (for example, an N-terminal sub-element column).

Upload your spreadsheet

Using the spreadsheet you prepared previously, upload it to complete registration.

  1. Upload the completed spreadsheet. Benchling performs upfront validation and flags missing columns or incorrect headers before proceeding
  2. Click Next

Map columns

  1. Review Benchling's column type detection — Benchling identifies whether each column contains existing entities or raw sequences, and whether raw sequences are DNA bases or amino acid residues
  2. Correct any column types if needed, then click Confirm column types
  3. Click Next

Finalize & create

  1. Review the entity creation summary — this shows each component type, the number of entities to create, and the schema they'll be created in
  2. Select the appropriate schemas for any component types not auto-assigned
  3. If your spreadsheet includes DNA inputs, select the genetic code for DNA-to-AA translation
  4. Select the project folder where proteins and components will be saved
  5. Optionally, search for and select an existing Notebook entry to receive a summary receipt when registration completes 
  6. Click Create and register

Note: After clicking Create and register, a notification appears with a link to track registration progress. Don't edit anything on that page — it's for monitoring status only.

Review registration results

Benchling sends an email notification when registration completes, whether it succeeded or encountered errors.

If registration encounters errors:

  1. Open the error notification email
  2. Click Download errors to download a .csv (protein_creation_errors.csv) listing the issues
  3. Review the errors — each row identifies the issue and the affected sequence
  4. Fix the issues in your spreadsheet
  5. Restart the registration flow from Global create

If registration succeeds:

  1. Open the success notification email
  2. Click View summary to open the summary Dataset — this lists all entities created, reused, or merged during registration, with a column per component type indicating whether each entity was created, referenced, or merged
  3. If you sent a receipt to a Notebook entry, open the entry to find a new section with a link to the Dataset

Note: If Benchling finds an existing entity that matches a sequence you provided, it marks that entity with a yellow icon in the summary.

View registered proteins and sequences

Once registration is complete, open a protein entity to explore its structure and metadata.

Each protein entity includes:

  • A Format field showing the VERITAS string for the format
  • Chain and domain fields linking to each registered sequence
  • Fieldset fields, including any computed biochemical property fields configured on your tenant
  • A protein map (glyph) showing the 2D structural visualization of the molecule

Navigate the protein map: Hover over a domain in the glyph and click to open a sequence preview of that component. In the glyph, distinct chains are color-coded, variable domains appear darker than constant domains, and hinges and linkers are represented as black lines. Click a chain name in the legend to open the full chain sequence.

Review annotation data on sequences: During registration, Benchling annotates all new sequences and writes structured metadata to the relevant schema fields. The following data is available on domain sequence entities:

  • CDR1, CDR2, CDR3 sequences
  • Framework region annotations
  • Closest-related V gene, J gene, and C gene, with percent identity
  • Liabilities (asparagine deamidation, methionine oxidation, N-linked glycosylation, aspartate isomerization, lysine glycation, aspartic acid–proline cleavage, hydrolysis)
  • Mutations and warnings (truncated framework regions, constant domain mutations)

Note: Benchling only annotates new sequences during registration. If you reference an existing entity by name or Registry ID, or if Benchling matches a sequence to an existing entity via uniqueness checks, that existing sequence is not re-annotated or updated.

 

Characterizing and validating protein entities 

While registering protein entities, Benchling annotates and validates the sequences that make up proteins. Effectively, Benchling is confirming that it matches the format you specified and that it looks like a real antibody sequence. Benchling also adds annotation metadata to new sequences that you can use when designing and evaluating antibodies.

During registration, findings come back in three tiers:

 Tier

 Effect on registration

 Where you see it

Errors

Blocks registration

Error report, delivered by email

Warnings

Does not block registration

Warnings field on the domain

Liabilities

Does not block registration

Liabilities field on the domain, plus annotations on the sequence

Note: Existing sequences used in protein registration are validated but never modified. If your import spreadsheet references an existing Benchling sequence, or we find an identical chain/domain sequence during registration while respecting uniqueness checks, we reuse it as is. We do not re-annotate it or update its metadata, so that existing data is not overwritten without your knowledge.

 

Characterization and validation overview 

For a given a format, Benchling looks for all expected domains in each chain, in the proper order and each domain and chain is matched independently. Additionally, with the exception of element domains, all domains in a format are required. 

During this process, Benchling checks for any annotation or validation errors across all entities. While there is no strict validation cutoff for domain identity, our testing shows that sequences at roughly 85% or greater identity with the reference germline annotate correctly.

During the registration process, if errors are detected it is reported and nothing is imported. In this way, registration is all or nothing. If everything succeeds: all new entities are created and registered while existing entities are reused. 

There are rare exceptions where a one-off failure or a Registry validation issue can leave an import partially created. When that happens, Benchling registers some antibodies but not all, and an email notification contains a breakdown of which succeeded and which errored.

Upon registration, most metadata is written to new sequences in two places (see example in the image below):

  • As annotations on the domain and chain sequences (see the sequence map), so you can see it positioned on the sequence

  • As structured schema fields on the domain entities (see the Metadata tab), so it is searchable and reportable across your registry

image-20260821-230515.png

 

How domains are annotated 

 Domains

 How Benchling annotates it

 What you get

Variable - VH, VL, VHH, VA, VB

Search against V and J genes of the germline you select, using IMGT or Kabat numbering (based on your germline choice)

Region annotated with variable domain type, plus FR1–FR4 and CDR1–CDR3 annotations, closest V and J gene with percent identity, mutations relative to germline, and liabilities

Constant - CL, CH1, Hinge, CH2, CH3, CH4, CA, CB

Search against C genes of the germline you select, using EU numbering

Region annotated with constant domain type, including closest C gene with percent identity, and mutations

Fusion

If provided within chains, search against any preregistered fusion domains; if provided as domains, created automatically

Region annotated as “fusion,” with percent identity and mutations relative to the matched fusion domain

Linker and element

Found by exclusion, leaving regions not annotated by variable, constant, or fusion domains; not aligned to a germline

Region annotated as “element,” with part links to any preregistered sub-elements* (tags, cleavage sites) found inside them

*Note: Tags and cleavage sites must be preregistered in an AA and/or DNA sub-element schema in order to be found and annotated on element domains, and must appear as an exact match within the element domain. Tags and cleavage sites cannot be identified otherwise. 

 

How CDRs and FRs are identified 

Benchling’s annotation algorithm follows the following process to find CDRs and FRs:

  1. First, Benchling locates FRs against germline

    • After the chain is decomposed into its domains, FR1–FR3 are located against the framework segments of the germline V genes and FR4 against the germline J genes. This is evaluated on amino acid sequences; any DNA input is translated using the genetic code you selected

  2. Then, Benchling recovers anything the first step missed and call CDRs by exclusion

    • For DNA input, if an expected FR is missing from the initial germline alignment (often due to divergence from the germline through engineering or humanization), Benchling searches the interval between the neighboring FRs for the consensus amino-acid sequence of that framework in the germline you selected 

    • Each CDR is then the span between two consecutive framework regions, so CDR boundaries follow from the framework boundaries, and therefore from the germline database you choose

  3. Finally, sequences are numbered and reported

    • Positional numbering (IMGT or Kabat for variable domains, depending on the germline you chose) is applied and mutations are labeled in that numbering scheme

    • Benchling also records the closest V and J gene with percent identity

List of possible errors

Any of errors below block registration of a sequence and, as a result, any protein imported with it.

 

 Error

 Details

 Applies to

Expected domain could not be annotated

Triggered when the format expects a required domain we could not place in the chain

All domains and chains

Domain out of order

Triggered when a domain matches at a position that conflicts with the order declared by the format

All domains and chains

Two domains overlap

Triggered when adjacent domains overlap by any number of residues

All domains and chains

Chain doesn’t match its domains

Triggered when a chain’s sequence can’t be reconciled with the domains it’s declared to contain

All chains

Domain type doesn’t match

Reported when a previously registered domain has a “domain type” value that doesn’t match the position in which its used in a chain

All domains

Light chain includes kappa and lambda

Reported when the VL and CL domains in a light chain are not all kappa or all lambda

VL and CL domains within a single chain

Same CH3s in an asymmetric format

Requires that two interacting CH3s in an asymmetric format must link to different entities

Interacting CH3 domains in an asymmetric center

CDR length outside the allowed range

Evaluated within all variable domains, where CDR1 must be 3–30 AA, CDR2 3–30 AA, and CDR3 3–35 AA

Variable domains

Variable domain too long

Triggered when a variable domain spans more residues than a variable domain can occupy

Variable domains

No FR or CDR regions can be found at all

Evaluated within all variable domains

Variable domains

CDR or FR region cannot be placed

Evaluated within all variable domains

Variable domains

Ambiguous bases/residues

Includes degenerate DNA bases or non-standard amino acids in the input

All domains

DNA doesn’t translate

Triggered when a DNA sequence isn’t a multiple of 3, or doesn’t translate to its corresponding AA sequence

All DNA domains and chains

Frameshift mutation

Checked per domain and across the whole coding sequence

All domains

Internal stop codon

Checked within each domain; a trailing terminal stop codon at the end of a chain is not flagged

All domains and chains

 

List of warnings

Warnings never block registration and are reported in the Warnings field of registered domains. The message column shows an example of the text you’ll see.

 Warning

 Details

 Message

 Applies to

Truncated FR region

Reported when a FR region is shorter than its germline sequence

'FR-H1' is truncated (start) or (end)

Variable domains

Truncated constant domains

Reported when a constant domain is shorter than its reference sequence

'CH1' is truncated (end)

Constant domains

Imperfect reference sequence match

Reported when percent identity to the matched reference is below 100%

CH2: No exact reference sequence match

Constant and fusion domains

No reference sequence match

Reported when a constant domain has no reference sequences available at all; this can occur for germlines that don’t have complete constant region coverage (e.g. the llama and alpaca germlines include no CL, and the llama germline includes no CH4) 

'CL': no reference sequences found - check germline and reference document configuration

Constant domains

Domain extension

Reported when reference alignments leave a short span of residues uncovered between two resolved domains; when that gap sits next to a constant domain, Benchling assigns the span to that adjacent constant domain so the annotations stay contiguous and marks the added span with positioned annotations

Domain extension (CH2)

Constant domains

Extra residues found 

Reported when residues are uncovered by the annotated domains at the start, before the next domain, or at the end. Extra residues do not need to fall inside an element domain to register, and there is no size threshold; registration fails only if the chain can’t be reconciled with its domains or a required domain is missing.

2 extra residue(s) were found at the start (or before the next domain, at the end; DNA reports base(s))

Any chain

 

Types of liabilities

A liability is a sequence motif associated with a known chemical or physical degradation risk. Benchling reports liabilities as sequence annotations and metadata, not as a blocker to registration. Liabilities are always reported with the region it was found in – for example, “Deamidation (CDR-H3).” Each liability is listed once per region in the Liabilities field, even if the motif occurs more than once there; every occurrence is annotated on the sequence.

Liabilities are detected on variable domains only and separately within each CDR and FR. By default, they are not detected on constant domains, hinges, linkers, element domains, or fusion domains. Any liabilities that straddle a boundary (e.g. an N-glycosylation sequon split between the end of FR3 and the start of CDR-H3) are not reported.

Each variable domain also has a Liability score. It starts at 0 and decreases by 10 for each liability or annotation warning found on that domain, so it is a simple tally, not a severity ranking or developability measure.

Motif liabilities

 Liability

 Reported as

 Motif

 Applies to

Asparagine deamidation

Deamidation (CDR-H3)

N[GSTH]

All CDRs and FRs

Aspartate isomerization

Asp isomerization (CDR-H3)

D[GS]

All CDRs and FRs

N-linked glycosylation

Glycosylation (CDR-H3)

N–X–[ST], X ≠ P

All CDRs and FRs

Lysine glycation

Lys glycation (CDR-H3)

KD, KE, EK

All CDRs and FRs

Aspartic acid – proline cleavage

Cleavage (CDR-H3)

DP

All CDRs and FRs

Hydrolysis

Hydrolysis (CDR-H3)

NP

All CDRs and FRs

Methionine oxidation

Oxidation (CDR-H3)

M

CDRs only

 

Cysteine liabilities

Benchling also flags cysteines that fall outside the expected disulfide bond pattern. Each message gives the number of cysteines found, the region, and what was expected.

 Reported as

 Flagged when

 Applies to

3 cysteines in VH (odd count, possible unpaired cysteine)

The total cysteine count is odd

Whole variable domain

2 cysteines in FR-H1 (at most 1 expected)

More than one cysteine (one canonical framework cysteine is expected)

FR1, FR3

1 cysteine in CDR-H3 (none expected)

Any cysteine at all

FR2, FR4, every CDR

 

Other reported metadata

 Field

 What it contains

 Applies to

V gene / J gene

Closest germline V and J gene

Variable domains

V gene percent identity / J gene percent identity

Percent identity to the closest V and J gene

Variable domains

CDR1 / CDR2 / CDR3

CDR sequences

Variable domains

C gene

Closest germline C gene; if several alleles match equally, the first is recorded

Constant domains

C gene percent identity

Percent identity to the matched reference

Constant domains

Mutations

Mutations relative to the germline or matched reference, also shown as annotations on sequences

Variable, constant and fusion domains

Liabilities

Motif and cysteine liabilities (see above), also shown as annotations on sequences

Variable domains

Liability score

Tally of liabilities and annotation warnings (see above)

Variable domains

Warnings

Non-blocking warnings (see above)

All domains

 

FAQ

Q: Can I use registration tables to register protein entities? 

A: No. Protein registration is only supported via Global create or the API. Registration tables are not supported for protein schemas. Component entities (domains, chains, pairs) can be pre-registered via registration tables if needed.

How long does registration take? 

Registration time varies by format and batch size. As a general estimate, registration takes 15–20 seconds per antibody after startup. Registration may take longer in larger registries.

What happens if registration encounters an error? 

Registration is all-or-nothing — if any error is detected, nothing is created and you'll receive an email with a CSV summarizing all errors. Fix the issues and restart the flow.

Can I include sequences that are already registered in Benchling? 

Yes. In your spreadsheet, provide the entity name or Registry ID instead of raw sequence data. Benchling reuses the existing entity rather than creating a new one. Existing sequences are not re-annotated during registration.

How many antibodies can I register in a single run? 

Up to 1,000 antibodies per run, with the following additional limits:

  • 4,000 chains per run
  • 14,000 domains per run
  • 45 domains per protein
  • 18 domains per chain
  • 1,200 amino acid residues per chain
  • 3,600 DNA residues per chain
  • 41,000 total entities per run

Will Benchling re-annotate existing sequences if I reference them in my spreadsheet? 

No. Benchling only annotates new sequences during registration. Existing sequences referenced by name, Registry ID, or matched via uniqueness checks are not updated or re-annotated.

Was this article helpful?

Have more questions? Submit a request