REPOGEO REPORT · LITE
princeton-nlp/LESS
Default branch main · commit 8abf9628 · scanned 6/13/2026, 11:27:56 AM
GitHub: 528 stars · 47 forks
Action plan is what to do next — copy-pasteable changes prioritized by impact. Category visibility is the real GEO test: when a user asks an AI a brand-free question that should surface princeton-nlp/LESS, does the AI actually recommend you — or your competitors? Objective checks verify the metadata signals AI engines weight first. Self-mention check detects whether AI even knows you exist by name.
Action plan — copy-paste fixes
3 prioritized changes generated by gemini-2.5-flash. Mark items done after you ship the fix.
- highreadme#1Clarify the repo's purpose as a data selection tool for LLM instruction tuning in the README's opening
Why:
CURRENTThis repo contains the code for our ICML 2024 paper LESS: Selecting Influential Data for Targeted Instruction Tuning. In this work, we propose a data selection method to select influential data to induce a target capability.
COPY-PASTE FIXLESS provides a practical, code-based method for **selecting influential training data to enhance targeted instruction tuning of large language models (LLMs)**. This repository contains the official implementation for our ICML 2024 paper, 'LESS: Selecting Influential Data for Targeted Instruction Tuning,' which proposes this novel data selection approach to induce specific target capabilities.
- mediumtopics#2Add more specific topics related to LLM fine-tuning and data optimization
Why:
CURRENTdata, data-selection, influence, instruction-tuning, llama, llm, mistral
COPY-PASTE FIXdata, data-selection, influence, instruction-tuning, llama, llm, mistral, llm-fine-tuning, data-optimization, model-tuning
- lowhomepage#3Add the paper's arXiv link as the repository homepage
Why:
COPY-PASTE FIXhttps://arxiv.org/abs/2402.06020
Category GEO backends resolved for this scan: google/gemini-2.5-flash, deepseek/deepseek-v4-flash
Category visibility — the real GEO test
Brand-free queries asked to google/gemini-2.5-flash. Did AI recommend you, or someone else?
Same questions for every model — switch tabs to compare answers and rankings.
- argilla-io/argilla · recommended 1×
- heartexlabs/label-studio · recommended 1×
- snorkel-team/snorkel · recommended 1×
- cleanlab/cleanlab · recommended 1×
- OpenAI's API · recommended 1×
- CATEGORY QUERYHow can I select the most effective training data for targeted LLM instruction tuning?you: not recommendedAI recommended (in order):
- Argilla (argilla-io/argilla)
- Label Studio (heartexlabs/label-studio)
- Snorkel (snorkel-team/snorkel)
- Cleanlab (cleanlab/cleanlab)
- OpenAI's API
- Claude
- Gemini
- Cohere's Embeddings
- OpenAI's `text-embedding-ada-002`
- GPT-4
- Claude 3
- Llama 3
- NLPAug (makcedward/nlpaug)
- TextAttack (textattack/textattack)
- Alpaca (tatsu-lab/stanford_alpaca)
- Dolly 2.0 (databrickslabs/dolly)
- ShareGPT
- FLAN
AI recommended 18 alternatives but never named princeton-nlp/LESS. This is the gap to close.
Show full AI answer
- CATEGORY QUERYWhat methods exist for identifying influential data points to enhance large language model fine-tuning?you: not recommendedAI recommended (in order):
- PyTorch-Influence-Functions
- Shapley
- PyTorch
- TensorFlow
- Hugging Face Transformers
- modAL
- UMAP
- t-SNE
- PCA
AI recommended 9 alternatives but never named princeton-nlp/LESS. This is the gap to close.
Show full AI answer
Objective checks
Rule-based audits of metadata signals AI engines weight most.
- Metadata completenesswarn
Suggestion:
- README presencepass
Self-mention check
Does AI even know your repo exists when asked about it directly?
- Compared to common alternatives in this category, what is the core differentiator of princeton-nlp/LESS?passAI named princeton-nlp/LESS explicitly
AI answers can be confidently wrong. Read for accuracy: does it match your actual tech stack, audience, and differentiator?
- If a team adopts princeton-nlp/LESS in production, what risks or prerequisites should they evaluate first?passAI named princeton-nlp/LESS explicitly
AI answers can be confidently wrong. Read for accuracy: does it match your actual tech stack, audience, and differentiator?
- In one sentence, what problem does the repo princeton-nlp/LESS solve, and who is the primary audience?passAI named princeton-nlp/LESS explicitly
AI answers can be confidently wrong. Read for accuracy: does it match your actual tech stack, audience, and differentiator?
Embed your GEO score
Drop this badge into the README of princeton-nlp/LESS. It auto-updates whenever the report is rescanned and links back to the latest report — easy public proof that you care about AI discoverability.
[](https://repogeo.com/en/r/princeton-nlp/LESS)<a href="https://repogeo.com/en/r/princeton-nlp/LESS"><img src="https://repogeo.com/badge/princeton-nlp/LESS.svg" alt="RepoGEO" /></a>Subscribe to Pro for deep diagnoses
princeton-nlp/LESS — Lite scans stay free; this card itemizes Pro deep limits vs Lite.
- Deep reports10 / month
- Brand-free category queries5 vs 2 in Lite
- Prioritized action items8 vs 3 in Lite