Data Preparation and Embedding

implementationChallenge

Prompt Content

Write a Python script to load genomic sequences from a FASTA file and tokenize them for StarCoder 2 using the Hugging Face Transformers library. Ensure that you handle the 4-letter DNA alphabet by mapping tokens correctly and implementing a sliding window approach for long sequences.

Try this prompt

Open the workspace to execute this prompt with free credits, or use your own API keys for unlimited usage.

Related Prompts

Explore similar prompts from our community

Usage Tips

Copy the prompt and paste it into your preferred AI tool (Claude, ChatGPT, Gemini)

Customize placeholder values with your specific requirements and context

For best results, provide clear examples and test different variations

Data Preparation and Embedding

Prompt Content

Related Prompts

Initial Assessment Prompt

Content Classification Prompt

Problem Definition and Requirements Gathering

Usage Tips