INCI Names & CAS Numbers

Five INCI Mistakes AI Tools Make on Ingredient Lists

The specific ways generic AI chatbots get INCI names, CAS numbers, and blend math wrong, and why those errors are easy to miss on a filing.

Diane R.4 min read

Someone pasted a formula into a general-purpose AI chatbot last month, asked it to generate the INCI list for a notification, and got back something that looked completely plausible. Correct-sounding names, a tidy table, confident formatting. Two of the entries were wrong in ways that wouldn't show up unless you already knew the ingredient. That's the trap with AI-generated INCI lists: the output is confident regardless of whether it's right, and confidence is exactly the thing that makes people stop double-checking.

Here are the specific failure patterns worth watching for.

1. Inventing plausible-sounding INCI names

Large language models are pattern generators. When asked for the INCI name of an obscure botanical extract or a newer synthetic ingredient, a model will sometimes produce something that follows INCI naming conventions perfectly, Latin binomial plus plant part plus extract, without that exact name existing in any actual nomenclature dictionary. It reads exactly like a real entry because the model learned the pattern, not the specific ingredient.

This is the hardest failure mode to catch because it doesn't look wrong. The fix is boring but necessary: verify every INCI name against an actual reference, not against how plausible it sounds.

2. Assigning a CAS number that belongs to a different substance

CAS numbers are precise identifiers, and models frequently confuse similarly named substances or grab a CAS number associated with a related but distinct compound. Some ingredients legitimately have multiple CAS numbers depending on manufacturing process, and some botanical extracts have no CAS number at all. A model asked to "give me the CAS number" for a botanical will sometimes produce one anyway rather than saying there isn't one, because producing an answer is what it's optimized to do.

If a filing carries the wrong CAS number, that's a real error in your documentation trail, even if the INCI name itself was correct.

3. Failing to expand supplier blends into components

This is probably the single most consequential mistake. A supplier trade name like "Emulsense HC" or "Plantapon SPS" isn't an INCI name and never belongs on a notification directly. It has to be expanded into its actual INCI components, and each component's real concentration in the finished product is the blend's use level multiplied by its share within the blend.

Generic AI tools frequently either skip this step entirely, listing the trade name as though it were fileable, or they expand it using generic assumed percentages instead of the specific supplier's actual blend composition. Both produce a filing that doesn't reflect the real formula.

4. Misreading concentration as the finished-product percentage

Related to the blend problem: models sometimes take a percentage mentioned anywhere in a formula description and apply it directly to the wrong ingredient, particularly when a formula is described in prose rather than a clean table. If your notes say "5% of a preservative blend that's 30% phenoxyethanol," a careless read can produce "phenoxyethanol 5%" instead of the correct 1.5%. That's a five-fold overstatement of an actual restricted ingredient's concentration, which is exactly the kind of number a Hotlist screening cares about.

5. Treating trade names as safe to submit as-is

Trade names carry marketing weight and brand identity for the supplier, but a filing needs the standardized INCI name underneath. AI tools trained broadly on internet text will sometimes recognize a trade name and produce a description of it, but stop short of mapping it fully to INCI, especially for smaller or regional suppliers whose documentation isn't widely represented in training data. The gap gets filled with a best guess that reads fluently and isn't verified against the supplier's own SDS.

A quick comparison of failure types

Mistake Why it happens What it looks like on a filing
Invented INCI name Pattern completion without a real match A name that looks right but doesn't exist
Wrong CAS number Confusing similar substances Wrong identifier tied to a real ingredient
Unexpanded blend Trade name treated as final Non-INCI name submitted directly
Miscalculated concentration Percentage applied to wrong component Overstated or understated restricted ingredient level
Trade name left as-is Supplier-specific mapping missing from training data Marketing name where INCI should be

Where the actual safeguard is

None of this means AI has no place in cosmetic regulatory work. It means unverified AI output shouldn't be the last step before a submission. The useful version of this workflow pairs automated matching against real INCI and CAS references with a human reviewer checking the result, especially the blend expansion math, before anything gets filed.

That's the structure Cosmetic Comply is actually built around: it matches your ingredients to INCI names and CAS numbers, expands supplier blends into their real components with carried-through percentages, screens each one against the relevant prohibited and restricted lists with a confidence score, and then a real compliance reviewer checks the result before the notification goes out. The automation speeds up the tedious matching work. The reviewer is there specifically because automation alone, of any kind, still gets things confidently wrong sometimes.

READY TO FILE?

Send your ingredients and we take it from here

A short intake form is all it takes to start. Every ingredient gets checked against your market's prohibited and restricted lists, then we file your notification and hand you a number you can track.

Start a filing

Keep reading