-
Notifications
You must be signed in to change notification settings - Fork 29
Add prefix consistency check #646
New issue
Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.
By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.
Already on GitHub? Sign in to your account
base: release/mvp
Are you sure you want to change the base?
Conversation
There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
Looks good to me.
There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
stable_id_prefix_consistency_check should be selected based on 'genebuild.method' metakey
|
|
||
| $self->translation_stable_id_check($species_id); | ||
| # NEW: check base prefix consistency across feature types | ||
| $self->stable_id_prefix_consistency_check($species_id); |
There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
Could you condition it with a genebuild.method ?
Like we have it here
| SKIP: { |
| $self->stable_id_check('gene', $species_id); | ||
| $self->stable_id_check('transcript', $species_id); | ||
| $self->stable_id_check('exon', $species_id); | ||
| $self->translation_stable_id_check($species_id); |
There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
Need to check this for metazoa/plants data. Most likely is not true for microbes if translations have the same sequences, as they are using sequence derived hashes as IDs.
Upd. Checked. Seems to be hold.
@Ensembl/plantazoa can you please check that this new consistency check works for your stable IDs? If not, I can update so that it only runs on genebuild cores.
Updated Stable ID check to check that all prefixes per features are consistent.
Tested:
changed prefix for one gene from ENSIMR -> ENSXXX, then test fails
Info about the new check:
Given a stable_id, we first drop any .version, then we strip trailing digits. From the remaining letters/underscores, if the last character is one of G T E P (the feature-type letter), we remove that one letter to get the base prefix. Examples: