Smart Sequence/Methodology
TECHNICAL METHODOLOGY

How Smart Sequence decides what the evidence supports.

This is the implementation-facing methods record for sequence interpretation, adaptive similarity search, biological ranking and evidence arbitration. It describes what the resolver actually does, including deterministic thresholds and known boundaries.

01

Scope

Smart Sequence accepts nucleotide or amino-acid sequence input, performs local interpretation, selects an appropriate BLAST-family search plan, ranks returned biological evidence, resolves a constrained molecular/taxonomic interpretation when supported, and then attaches compatible gene/protein/structure context.

Not inferred automatically:

Similarity alone does not establish orthology, gene expression, pathogenicity, causality, a unique source organism, or clinical significance.

02

Input normalization

FASTA headers, whitespace, digits and common formatting separators are removed before sequence interpretation. Molecule class is inferred from the cleaned alphabet. Inputs containing only nucleotide/IUPAC symbols are treated as nucleotide; sequences containing amino-acid-only symbols are treated as protein.

Nucleotide alphabet: A C G T U R Y S W K M B D H V N
Protein alphabet: standard amino-acid symbols plus ambiguity symbols

The UI exposes the inferred molecule class as a read-only status. Manual DNA/protein selection is not required in the normal workflow.

03

Local interpretation

Nucleotide input is inspected locally before any external submission. DEVNU computes basic composition and candidate open reading frames using the installed standard translation table. ORF presence is interpreted as coding potential only.

Internal ORF conditionCurrent thresholdEffect
Plausible ORF≥60 aa and ≥45% of nucleotide query represented by ORF length × 3Allows supporting protein evidence to contribute to coding-sequence promotion.
Compelling ORF≥100 aa and ≥60% of nucleotide query represented by ORF length × 3Supports coding potential. Eligible ORF probes are scheduled independently of BLASTN and BLASTX validation; a probe never replaces whole-query evidence.
Boundary:

An ORF is not proof that a locus is expressed, translated in vivo, or functionally equivalent to a database protein.

04

Adaptive search planning

The recommended plan depends on the observed molecule class and local coding evidence.

Protein input
  ├─ blastp  → direct protein similarity
  └─ tblastn → nucleotide records encoding similar proteins

Nucleotide input
  ├─ blastn  → direct nucleotide similarity
  ├─ compelling ORF (≥100 aa, ≥60% query) → blastp on translated ORF
  └─ otherwise → blastx translated-frame evidence

Recommended mode currently submits blastn/blastp/blastx through EMBL-EBI Job Dispatcher when compatible and uses NCBI BLAST URL API for tblastn or provider fallback. Advanced mode deliberately pins submission to NCBI so requested database/scoring parameters are preserved exactly.

05

Search parameters

Simple mode uses provider-compatible defaults. Advanced search exposes E-value threshold, target database profile, maximum retained targets, protein scoring matrix, gap costs and low-complexity filtering.

ControlMeaningImplementation
E-valueMaximum expected number of chance matches retained by BLAST.EXPECT
MatrixProtein substitution scoring model.MATRIX; BLOSUM45/50/62/80/90 and PAM30/70/250 are accepted by the NCBI URL API.
Gap costsGap-open and gap-extension penalties.GAPCOSTS; the route validates the allowed NCBI URL API pairs for nucleotide or protein-family programs.
Database profileFocused, RefSeq or broad search space.Protein: Swiss-Prot / RefSeq protein / nr. Nucleotide: core_nt / RefSeq RNA / nt.
Max targetsMaximum database sequences retained.HITLIST_SIZE
Low complexityQuery masking/filtering to reduce spurious low-complexity hits.FILTER

Unsupported parameter combinations are rejected with HTTP 400 rather than silently rewritten. Recommended mode is not converted to Advanced mode unless the user explicitly enables those overrides.

06

Similarity metrics

Identity is the fraction of aligned positions containing identical residues/bases. Query coverage is the fraction of the submitted query represented in the alignment. E-value and bit score retain the provider's statistical interpretation.

Important:

100% query coverage means the entire submitted query is aligned. It does not mean the query spans the complete target protein, gene or genome record.

07

Biological ranking

Provider order is retained as provenance but is not treated as biological truth. Each hit receives an internal biological score derived from sequence metrics plus transparent record-quality signals. The score is used to rank retained evidence, not to manufacture statistical significance.

Ranking contributionCurrent implementation
Identity0.42 × identity%
Query coverage0.34 × coverage%
E-value strengthUp to +18 from -log10(E)/7; E-values above 1e-5, 1e-3 and 1e-2 receive progressively larger penalties.
Bit scoreUp to +4 from bitScore / 100.
Curated RefSeq+14
Curated/recognized protein record+10
Natural organism label+6
Complete coding/full-length wording+3
Predicted RefSeq+3
Gene consensus agreement+8
Synthetic construct−24
Clone-derived non-curated record−2

Gene consensus is computed from up to the first 30 provider hits. A gene-bearing hit contributes only when identity is at least 45%, query coverage at least 40%, and E-value at most 1e-2; curated/natural records receive additional weight. Protein and nucleotide hit lists remain separate evidence streams.

08

Promotion gate

Similarity results are converted into evidence-strength bands. These are explicit deterministic thresholds in the current resolver, not a learned classifier.

Protein-space evidence: blastp / blastx

ClassE-valueQuery coverageIdentity
Decisive≤1e-20≥60%≥30%
Strong≤1e-10≥50%≥25%
Supporting≤1e-5≥40%≥25%
WeakPasses E ≤1e-2 but misses thresholds above——
Insignificant>1e-2——

Nucleotide-space evidence: blastn / tblastn

ClassE-valueQuery coverageIdentity
Decisive≤1e-50≥80%≥95%
Strong≤1e-20≥70%≥90%
Supporting≤1e-5≥40%No additional identity cutoff in this band
WeakPasses E ≤1e-2 but misses thresholds above——
Insignificant>1e-2——
Promotion behavior:

For protein input, supporting-or-better direct protein evidence can unlock representative protein annotation. For nucleotide input, strong protein evidence is promotable; supporting protein evidence is promotable only when a plausible local ORF (≥60 aa and ≥45% query) is present and the query is not classified as non-coding RNA. Strong non-coding nucleotide evidence blocks translated protein evidence from driving identity.

09

Taxonomic arbitration

Taxonomy is resolved from natural-organism nucleotide evidence rather than from the representative protein record. Synthetic/vector/unclassified labels are excluded from taxonomic resolution.

After ranking, DEVNU forms an equivalence set from strong-or-decisive natural hits that are within 0.25 percentage points of the top identity and within 2 percentage points of the top query coverage, considering up to 12 equivalent hits.

Equivalent-set patternReturned taxonomic state
More than one genusAmbiguous across genera; no unique genus/species call.
One genus, multiple species or unresolved species labelsGenus-level resolution only.
Exactly one species representedSpecies-level taxonomic context.
Insufficient natural labelsTaxonomy unresolved.

When protein records are tied at the same evidence-strength band and within the same 0.25% identity / 2% coverage window, nucleotide species evidence may choose the representative protein annotation among those ties. This is a representative-record tie-break only; it is not converted into a species claim.

10

Functional enrichment

External annotation is explicitly distinguished from sequence-derived evidence. Current enrichment layers include UniProt protein annotation, Ensembl/NCBI gene context, Gene Ontology, InterPro, Reactome, PDB and AlphaFold where available.

Smart Sequence first requests a lightweight UniProt core record so function and subcellular localization do not depend on slower InterPro/Reactome/AlphaFold calls. Full enrichment remains available in dedicated workbenches.

11

Partial failures

FailureWhat remains valid
BLAST submission/result failsLocal interpretation only.
Protein enrichment failsCompleted similarity identity remains valid; annotation is marked unavailable.
Gene context failsSequence/protein identity remains valid.
Taxon image/media failsTaxonomic evidence is unaffected.
Structure lookup failsNo structure claim is made; protein identity can remain valid.
12

Reproducibility

Structured exports capture the input, SHA-256 input fingerprint, generation timestamp, provider/job identifiers, program, provider-reported version, database metadata when returned, retrieval timestamp and advanced settings. Reports link back to this methodology page.

Database versions are provider metadata: DEVNU records them when the provider exposes them and does not fabricate a version string when it does not.

Because the ranking and promotion rules above are implementation rules, publication-grade use should record a DEVNU software/resolver version together with the exported run manifest.

13

Known limitations

Interpretation is especially vulnerable for very short queries, low-complexity sequence, recent paralogs, large conserved families, partial loci, pseudogenes, horizontal gene transfer, recombination/reassortment, synthetic constructs, sparse environmental references and incorrectly annotated public records.

The numeric promotion thresholds are conservative heuristics for evidence gating; they are not universal biological constants and should not be interpreted as substitutes for domain-specific validation.

Phylogenetic inference additionally depends on homolog selection, comparable sequence regions, alignment quality, model choice and support estimation. A tree produced from non-comparable fragments can be computationally valid and biologically meaningless.

14

Use-case tutorials

KNOWN PROTEIN

HBB

Use the built-in HBB example to inspect direct protein identity, reverse nucleotide support and the distinction between molecular identity and organism evidence.

PARTIAL HUMAN PROTEIN

OCA2 fragment

Demonstrates why 100% query coverage does not mean a full-length target and how gene function/locus context is attached independently.

AMR PROTEIN

TEM-1 β-lactamase

Use the built-in TEM-1 example to follow protein identity into antimicrobial-resistance function and comparative homolog evidence without turning a similarity result alone into a phenotype claim.

ENVIRONMENTAL / 16S

Unknown bacterial marker

Use a curated 16S reference as a tutorial control, then compare against an environmental fragment and watch taxonomic resolution degrade as the query becomes less discriminative.

DISEASE-GENE VARIANT WORKFLOW

Sequence → gene → variant context

Start from a disease-gene sequence fragment, resolve the gene without treating the sequence match as a pathogenicity claim, then continue into the Variant and Disease workbenches. Variant significance must come from variant-specific evidence, not from Smart Sequence similarity alone.

15

References & services

DEVNU Science · Smart Sequence methodology · implementation behavior should be interpreted together with the resolver/software version recorded in an exported analysis.