AraBase User Guide
AraBase is a genomic resource for Arabidopsis thaliana that provides access to TAIR12 and TAIR10 / Araport11 annotations, locus search, type-aware locus cards, sequence retrieval, genome browsing, Gene Ontology tools, expression maps, protein information, and downloadable genomic datasets.
The availability of individual panels and actions depends on the selected annotation release, feature type, source-data availability, and external service availability.
Overview
Use AraBase when you need to move between curated Arabidopsis genome annotations, downloadable files, sequence retrieval, GO analysis, and locus-level summary pages. The main site navigation contains Home, Tools, Data, Docs, Contact, and the search box.
Use the same release throughout a coordinate-based analysis unless you have explicitly converted coordinates between releases.
Quick Start
Useful Search Locus examples include AT1G01010, RDR2, AT1TE99104, Copia, and Satellite. Select the correct assembly before searching for a locus or browser interval.
Homepage Navigation
The homepage is also the Arabidopsis thaliana species page. Its hero buttons jump to major site tasks.
| Button | Function |
|---|---|
| Genome | Jumps to the homepage Genome Information section. |
| Tools | Opens the available AraBase tools: Search Locus, Genome Browser, Sequence Extract, GO Enrichment Analysis, and GO Term Search. |
| News | Jumps to News & Updates. |
| Data Portal | Opens downloadable datasets at /data/. |
| User Guide | Opens this documentation page at /docs/. |
| Contact | Opens /contact/. |
Search Locus
Search Locus provides a unified interface for finding supported genomic features in TAIR12 and TAIR10 / Araport11. Choose a release, enter a search term, and open a type-aware locus card from the results.
Supported Query Types
Locus ID searches include parent loci and transcript-suffixed identifiers such as AT1G01010, AT1G01010.1, and AT1TE99104. Transcript suffixes are normalized to the parent AGI locus where the static index supports that relationship.
Alias and gene-symbol matching is case-insensitive for indexed aliases, for example RDR2, FLC, or AGO4.
Annotation-type searches can find indexed feature classes and biotypes such as protein-coding gene, pseudogene, miRNA, tRNA, and rRNA.
TE classification searches can find indexed classes and families such as Class I, LTR, Copia, Gypsy, and Helitron.
Non-TE repeat classification searches can find indexed repeat classes such as Satellite, Simple repeat, Tandem repeat, and Low-complexity region when those records are present in the selected release.
Ortholog identifiers are searched only when orthology data are present in the static search index. No separate orthology source should be assumed for a record that does not show an ortholog match.
Search Filters
Search results use these filters in interface order: All, Genes, Transposable Elements, Non-TE Repeats, Non-coding RNAs, and Pseudogenes.
Child features such as mRNA, exon, CDS, and UTR are generally displayed inside the parent locus card rather than returned as independent top-level loci.
Search Result Fields
Result cards summarize the primary locus ID, feature class, subtype or biotype, alias or name, chromosome and coordinates, annotation release, and match reason. Current match reasons include locus ID, alias, annotation type, TE family or classification, repeat class, and ortholog ID when the relevant index data exist.
Direct locus-card links can be shared, for example /tools/search-locus/?release=TAIR12&id=AT1G01010.
Locus Cards
AraBase uses type-aware locus cards. Common genomic information is shown for all supported features, while gene-, TE-, repeat-, ncRNA-, and pseudogene-specific panels are displayed only when relevant data are available.
Common Information
Most locus cards show the locus ID, feature type or biotype, selected release, chromosome, position, strand, feature length, aliases or names when available, description when available, and Genome Browser view.
Where a sequence store exists, the Information section includes a Copy sequence button. Gene-like loci may also include View the locus, ePlant, TAIR, SNP Viewer, and T-DNA Express buttons, depending on the identifier and feature type.
Gene Cards
Protein-coding gene cards can include Information, Transcripts, Protein, TAIR12 - Araport11 Comparison, Functional Annotation, Gene Ontology, Gene Expression Atlas (eFP), Structure Prediction (AlphaFold), and Genome Browser.
The Transcripts table shows Transcript ID, Type, Length, Exons, CDS length, UTR annotation, and available copy actions for transcript or CDS sequence.
The Protein table shows Protein ID, Transcript ID, Length (aa), Molecular weight (kDa), and Copy protein.
Protein molecular weight is reported for each protein isoform in kilodaltons (kDa). NA indicates that an exact value could not be calculated because the sequence contains unresolved amino acids such as X. An asterisk indicates that gap or placeholder characters were removed before calculation. Values represent the unmodified protein sequence and do not account for post-translational modifications.
The Gene Ontology panel groups annotations by Biological Process, Molecular Function, and Cellular Component. It shows GO IDs, GO terms, relationships, evidence, with/from values, references, annotators, dates, and GO-slim summaries where present. GO IDs link to AmiGO, and AGI with/from identifiers link back to AraBase Search Locus.
The Gene Expression Atlas (eFP) panel is collapsed by default. Expand it, choose a dataset, choose Absolute or Relative mode, then click Load expression map. The image is loaded only after that user action. You can open the full-size BAR image or the selected gene in ePlant. Availability differs among genes and datasets, and images are requested from BAR.
The Structure Prediction (AlphaFold) panel is shown when an AraBase protein has an appropriate exact mapping to a UniProt protein with an available AlphaFold model. The displayed model is based on the mapped UniProt protein and is not calculated directly by AraBase. The panel is collapsed by default and loads Mol*, static AlphaFold metadata, and external structure files only after expansion. Available actions include Open AlphaFold DB, Download Image as JPG, SVG, or PDF, Download mmCIF, and Download PDB.
The TAIR12 - Araport11 Comparison panel summarizes corresponding IDs, coordinates, feature length, transcript count, transcript correspondence, and comparison status where comparison data exist. A retained status does not mean exon, CDS, UTR, or coordinate structures are identical.
Transposable-Element Cards
Transposable-element cards show TE ID, feature type, release, chromosome, position, strand, length, TE classification, TE family, method, identity, name or alias, sequence action where available, and Genome Browser view.
Gene-only panels such as transcripts, protein, GO annotations, eFP, and AlphaFold are omitted for TE cards.
Non-TE Repeat Cards
Non-TE repeats are repetitive genomic features that are not classified as mobile genetic elements. Cards can show repeat ID, repeat type, repeat family or class, coordinates, strand, length, sequence action where available, and Genome Browser view.
Available categories depend on the release and index, and may include satellite, simple repeat, tandem repeat, low-complexity, centromeric, or telomeric repeat records.
Non-coding RNA Cards
Non-coding RNA cards are shown for supported RNA classes such as miRNA, tRNA, rRNA, snRNA, snoRNA, lncRNA, and other ncRNA records when present. Available fields vary by RNA class and release, but may include RNA class, locus ID, coordinates, strand, sequence action, and Genome Browser view.
Pseudogene Cards
Pseudogene cards can show pseudogene ID, pseudogene type, aliases, genomic coordinates, strand, length, sequence action where available, and Genome Browser view. Protein-specific panels are normally omitted.
GO Term Search
GO Term Search allows users to search the current TAIR Gene Ontology annotation dataset by GO ID, term name, keyword, or GO-slim category and retrieve associated Arabidopsis genes.
Accepted examples include GO:0006355, 0006355, DNA methylation, flowering, RNA binding, and DNA binding. GO ID input is normalized, so go:0006355 and 0006355 are treated as GO:0006355.
Ontology Filters
GO Term Search filters are All, Biological Process, Molecular Function, and Cellular Component. Filter counts reflect the current query results.
Search Results
Each result shows the GO ID, term name, ontology, unique annotated gene count, annotation-row count, and GO-slim categories where present. Expanding a result loads the associated gene list.
The gene list shows Gene ID, Symbol, Description, Relation, and Evidence. Available actions are Copy gene IDs, Download TSV, Show more, and Show all.
Data Limitations
This version searches GO term names and GO-slim categories. GO definitions and synonyms are not included unless an ontology file has been incorporated.
Gene counts represent unique genes after duplicate annotation rows are merged. Associations explicitly qualified with NOT are excluded from positive gene lists. GO annotations are shown from the direct source associations and are not automatically propagated through the GO hierarchy.
GO Enrichment Analysis
GO Enrichment Analysis identifies Gene Ontology terms that are statistically overrepresented in a submitted Arabidopsis gene list compared with the selected background.
| Tool | Input | Output |
|---|---|---|
| GO Term Search | GO ID, term name, or keyword | Matching GO terms and associated genes |
| GO Enrichment Analysis | Arabidopsis gene list | Statistically overrepresented GO terms |
Input Gene List
Paste gene or transcript IDs into the text area. The parser accepts one ID per line and also handles comma-, space-, semicolon-, and tab-separated lists. AGI gene IDs are case-normalized, transcript suffixes are normalized to parent genes, duplicates are removed, and invalid identifiers are reported.
Example:
AT1G01010
AT2G40030
AT3G18780Background Gene Set
The default background is all genes with GO annotations. A custom background gene list can be supplied from the Analysis settings panel. The disabled protein-coding background option is not currently available.
Ontology Categories
Available ontology choices are All ontologies, Biological Process, Molecular Function, and Cellular Component.
Statistical Method
AraBase runs the analysis locally in a web worker. It uses a one-sided hypergeometric over-representation test and Benjamini-Hochberg false-discovery-rate correction.
Results
The Input summary reports submitted identifiers, unique submitted identifiers, unique normalized genes, genes in the selected background, genes with GO annotations, genes without GO annotations, invalid identifiers, genes excluded from background, and tested GO terms.
The Enriched GO terms table shows GO ID, GO term, Aspect, Query genes, Query total, Background genes, Background total, Gene ratio, Background ratio, Fold enrichment, P-value, and FDR. Each GO term can expose matching input genes and per-term copy or download actions.
Plot
The Enrichment plots panel is collapsed by default. When opened, the plot can use Significance: -log10(FDR), Fold enrichment, or Gene ratio. Separate plot colors are configurable for Biological Process, Molecular Function, and Cellular Component.
Downloads
The Downloads section provides Input report, Complete results TSV, Significant results TSV, and Enrichment plots as JPG, SVG, or PDF. Plot image exports are generated at 600 dpi.
Interpretation
A significant GO term indicates that genes associated with that term occur more frequently in the submitted list than expected from the selected background. It does not by itself demonstrate that the process is activated, repressed, or causally responsible for a phenotype.
Related GO terms often overlap, parent and child terms are not statistically independent, results depend on annotation coverage and background choice, small lists may have limited power, and adjusted P values should normally be preferred over raw P values.
Genome Browser
Genome Browser uses JBrowse2 to inspect TAIR12 and TAIR10 / Araport11 annotations. Select an assembly using the buttons above the browser, search a locus ID or interval in the JBrowse location field, zoom and pan across the assembly, toggle tracks from Available tracks, and click features to inspect annotations.
TAIR12 opens on Chr1 with protein-coding genes with UTRs and TAIR12 transposable elements loaded. TAIR10 / Araport11 opens on Chr1 with Araport11 protein-coding genes and classified transposable elements loaded.
Coordinates are release-specific. Confirm the selected assembly before interpreting or exporting a region.
JBrowse2 can also load user tracks from File or Add track controls. Supported custom-track file types include FASTA with indexes, GFF3, GTF, BED, BEDGraph, BigWig, BigBed, BAM, CRAM, VCF, and tabix-indexed text tracks. Compressed and indexed files should include the matching index required by JBrowse2, such as .fai, .gzi, .tbi, .csi, .bai, or .crai depending on the data type.
Sequence Extraction
Sequence Extract retrieves sequences from static per-release stores. Select Assembly, choose Feature Type, paste up to 1000 feature IDs, and click Search.
Supported assemblies are TAIR12 and TAIR10 / Araport11. Supported feature types are Gene, Transcript / mRNA, CDS, Protein, and TE / repeat.
Feature IDs may be entered one per line or separated by spaces, commas, semicolons, or tabs. Matching is case-insensitive through the sequence index. The result summary reports indexed records, input IDs, found records, and not found records. The results table shows Input ID, Status, Matched ID, Feature Type, Length, and Description.
Use Export CSV to download the result table. Use Download to download matched sequences in FASTA format. FASTA headers are generated from the matched records and selected feature type.
Sequence extraction follows the selected release. TAIR12 and TAIR10 / Araport11 may use different feature boundaries, transcript models, or annotation structures.
Data Portal
Data Portal provides downloadable genome assemblies, annotations, sequences, and related resources for TAIR12 and TAIR10 / Araport11.
Available Releases
TAIR12 and TAIR10 / Araport11 files are maintained separately. Use files from the same release for coordinate-based analysis unless coordinates have been converted explicitly.
Genome Assemblies
Available genome FASTA files include Athaliana_TAIR12_genomic_Chr_softmasked.fa.gz for TAIR12 and TAIR10_chr_all.fa.gz for TAIR10 / Araport11.
Gene Annotations
TAIR12 gene annotation files include TAIR12_protein_coding.gff3.gz, TAIR12_protein_coding_with_UTRs.gff3.gz, TAIR12_pseudogene.gff3.gz, TAIR12_non_coding_RNA.gff3.gz, and TAIR12_genome_annotation_TAIR.gff3.gz.
TAIR10 / Araport11 annotation files include Araport11_GFF3_genes_transposons.20250813_protein_coding.gff3.gz, Araport11_GFF3_genes_transposons.20250813.gff3.all_ncRNAs.gff3.gz, and Araport11_GFF3_genes_transposons.20250813.gff3.gz.
TE and Repeat Annotations
TAIR12 TE and repeat annotation files include TAIR12_All_TE.gff3.gz, TAIR12_ClassI_TE.gff3.gz, TAIR12_ClassII_TE.gff3.gz, TAIR12_LTR.gff3.gz, TAIR12_LINE.gff3.gz, TAIR12_SINE.gff3.gz, TAIR12_TIR.gff3.gz, TAIR12_RC.gff3.gz, and TAIR12_Non_TE_repeats.gff3.gz.
TAIR10 / Araport11 TE annotation files include Araport11_GFF3_genes_transposons.20250813_classified_transposable_element.gff3.gz and Araport11_GFF3_genes_transposons.20250813_transposable_element_gene.gff3.gz.
Sequence Files
TAIR12 sequence files include TAIR12_protein_coding_gene_sequences.fasta.gz, TAIR12_transcripts.fasta.gz, TAIR12_CDS.fasta.gz, TAIR12_proteome.fasta.gz, TAIR12_All_TE_sequences.fasta.gz, and Araport11_ChrM_ChrC.fa.gz.
TAIR10 / Araport11 sequence files include Araport11_gene_sequences.fa.gz, Araport11_transcriptome.fa.gz, Araport11_CDS.fa.gz, Araport11_proteome.fa.gz, and Araport11_TE_sequences.fa.gz.
Index Files
Companion index files such as .fai, .gzi, and .tbi are used internally by JBrowse2, indexed FASTA access, and tabix-indexed annotation access. The Data Portal intentionally lists downloadable data files, not standalone index records.
File Formats
FASTA stores nucleotide or protein sequences. GFF3 and GTF store genome annotations. TSV stores tabular results or mappings. JSON stores static application indexes used by AraBase tools. .gz indicates compressed files.
Choosing the Correct Release
Use TAIR12 files for TAIR12 coordinates, locus cards, and browser views. Use TAIR10 / Araport11 files when working with the legacy TAIR10 reference or Araport11 annotations.
FAQs
Why can't I find a locus?
The selected release may be wrong, the feature may be absent from that release, the alias may not be indexed, the feature type may not be supported, or the record may not be included in the current static index.
Why are coordinates different between TAIR12 and TAIR10 / Araport11?
The selected release controls the annotation record, coordinates, sequence, and browser assembly. Do not compare coordinates across releases without confirming compatibility.
Why are some locus-card panels missing?
Panels are displayed only when relevant data exist for the selected feature. Expression and AlphaFold panels are gene-specific, TE and repeat cards omit protein-specific panels, and external mappings may be unavailable.
Why is protein molecular weight shown as NA?
The protein sequence contains unresolved amino acids such as X, or no valid molecular-weight record is available.
What does the asterisk after a molecular weight mean?
Gap or placeholder characters were removed before molecular weight was calculated.
Why does an eFP image fail to load?
The selected atlas may lack the gene, the selected mode may be unavailable, BAR may be temporarily unavailable, third-party images may be blocked, or the network request may fail.
Why is an AlphaFold model unavailable?
A model is displayed only when the AraBase protein has an appropriate exact UniProt mapping with an available AlphaFold structure.
Why does GO Term Search return no result?
The GO ID may be absent from the current TAIR GO dataset, the query may be a definition or synonym not included in the current index, an ontology filter may exclude the term, or the spelling may not match indexed text.
What is the difference between GO Term Search and GO Enrichment Analysis?
GO Term Search starts with a GO ID or keyword and returns matching GO terms and genes. GO Enrichment Analysis starts with a gene list and identifies GO terms that occur more frequently than expected.
Why are there no significant GO enrichment results?
The gene list may be small, few input genes may have GO annotations, the biological signal may be weak, the background choice may reduce contrast, multiple-testing correction may remove nominal significance, or invalid IDs may have been removed.
Why does JBrowse open the wrong region or fail to locate a feature?
The wrong assembly may be selected, chromosome naming may not match the release, the locus may be absent from the selected release, the coordinate format may be invalid, or old browser state may need to be cleared.
Why do downloaded sequences differ between releases?
TAIR12 and TAIR10 / Araport11 may use different feature boundaries, transcript models, or annotation structures. Sequence retrieval follows the selected release.
Can I share a locus card or tool result?
Locus-card URLs include the selected release and locus ID and can be bookmarked or shared directly, for example /tools/search-locus/?release=TAIR12&id=AT1G01010. Expanded panels may not be preserved, unsaved analysis input may not be encoded in the URL, and GO or enrichment state is shareable only when the implementation encodes it.
How can I contact AraBase or report a problem?
Use the AraBase contact page. Please include the page URL, selected release, locus or GO ID, browser, screenshot, exact error message, and steps to reproduce.
Data Sources and Citations
AraBase integrates redistributed static files, derived static indexes, and external links or images. TAIR and Araport11 provide genome and annotation resources. Gene Ontology annotations are used for locus cards, GO Term Search, and GO Enrichment Analysis. BAR and ePlant provide external expression views and eFP images. UniProt and AlphaFold DB provide protein mapping and predicted-structure resources where exact mappings exist. JBrowse2 powers the embedded Genome Browser. NCBI and Ensembl are linked as external reference resources from the homepage.
Please cite AraBase together with the relevant genome assembly, annotation release, tool, and external resource used in your analysis.
AraBase
Chen, H., Emmerson, R., and Mosher, R. 2026. Near-gapless and haplotype-resolved Capsella genomes enable investigation into genomic consequences of mating system shifts. bioRxiv, 2026.07.10.737683. https://doi.org/10.64898/2026.07.10.737683
TAIR12
Reiser, L., Proia, A., Bakker, E., Subramaniam, S., Khosa, K., Sawant, S., Chen, X., Prithvi, T., and Berardini, T. Z. 2026. Recent major changes to TAIR: updates to the database, website, and Arabidopsis genome. Genetics 232(4):iyaf248. https://doi.org/10.1093/genetics/iyaf248
Wlodzimierz, P., Rabanal, F. A., Burns, R., et al. 2023. Cycles of satellite and transposon evolution in Arabidopsis centromeres. Nature 618:557–565. https://doi.org/10.1038/s41586-023-06062-z
TAIR10
Lamesch, P., Berardini, T. Z., Li, D., Swarbreck, D., Wilks, C., Sasidharan, R., Muller, R., Dreher, K., Alexander, D. L., Garcia-Hernandez, M., Karthikeyan, A. S., Lee, C. H., Nelson, W. D., Ploetz, L., Singh, S., Wensel, A., and Huala, E. 2012. The Arabidopsis Information Resource (TAIR): improved gene annotation and new tools. Nucleic Acids Research 40(D1):D1202–D1210. https://doi.org/10.1093/nar/gkr1090
Araport11
Cheng, C.-Y., Krishnakumar, V., Chan, A. P., Thibaud-Nissen, F., Schobel, S., and Town, C. D. 2017. Araport11: a complete reannotation of the Arabidopsis thaliana reference genome. The Plant Journal 89(4):789–804. https://doi.org/10.1111/tpj.13415
JBrowse 2
Diesh, C., Stevens, G. J., Xie, P., De Jesus Martinez, T., Hershberg, E. A., Leung, A., Guo, E., Dider, S., Zhang, J., Bridge, C., Hogue, G., Wang, X., Liu, G., Dunn, M., Holmes, I. H., and Buels, R. M. 2023. JBrowse 2: a modular genome browser with views of synteny and structural variation. Genome Biology 24:74. https://doi.org/10.1186/s13059-023-02914-z
AlphaFold
Jumper, J., Evans, R., Pritzel, A., et al. 2021. Highly accurate protein structure prediction with AlphaFold. Nature 596:583-589. https://doi.org/10.1038/s41586-021-03819-2
UniProt
The UniProt Consortium. 2025. UniProt: the Universal Protein Knowledgebase in 2025. Nucleic Acids Research 53(D1):D609-D617. https://doi.org/10.1093/nar/gkae1010
Gene Ontology
Ashburner, M., Ball, C., Blake, J., et al. 2000. Gene Ontology: tool for the unification of biology. Nature Genetics 25:25-29. https://doi.org/10.1038/75556
BAR / ePlant
Sullivan, A., Lombardo, M. N., Pasha, A., Lau, V., Zhuang, J. Y., Christendat, A., Pereira, B., Zhao, T., Li, Y., Wong, R., Qureshi, F. Z., and Provart, N. J. 2025. 20 years of the Bio-Analytic Resource for Plant Biology. Nucleic Acids Research 53(D1):D1576-D1586. https://doi.org/10.1093/nar/gkae920
Licensing
AraBase website content, AraBase-generated pages, documentation, and site-specific assets are licensed under the Creative Commons Attribution-NonCommercial 4.0 International License (CC BY-NC 4.0).
Commercial use is not permitted under this license. For commercial licensing terms, please contact the maintainer.
Third-party datasets, external services, software dependencies, themes, external logos, and linked resources remain subject to their own licences and terms. AraBase documentation does not relicense third-party resources.