Complex odors prove easier to map than expected with machine learning
If you want to describe a particular color, you could look to the Pantone color wheel to find its exact hue, saturation and brightness, and how it compares with other colors.
Key points
- Focus: If you want to describe a particular color, you could look to the Pantone color wheel to find its exact hue, saturation and brightness, and how it
- Detail: Science reporting: verify primary technical documentation
- Editorial reading: science reporting; whenever possible, verify the cited primary source.
If you want to describe a particular color, you could look to the Pantone color wheel to find its exact hue, saturation and brightness, and how it compares with other colors. But nothing like that has existed for complex odors. The science-journalism coverage adds useful context, while the strongest evidential footing still comes from the underlying data, papers or institutional documentation.
That matters because chemistry gains force when a claimed structure or process can be described with enough precision to be reproduced by others. Synthetic routes, spectroscopic signatures, yield under defined conditions and stability under realistic operating parameters are the currency of credibility in chemistry, and a result that lacks these details cannot be evaluated independently. The distance between a discovery on a laboratory bench and a process that works reliably at scale is measured in years of optimization, and each step reveals constraints that were invisible at smaller scale. This article has been reviewed according to Science X's editorial process and policies. In 2015, IBM ran an open investigator DREAM challenge to take a single odor molecule and predict what it smells like based on its chemical structure, he said.
First, they standardized and compiled six data sets of odor-similarity measurements from three different studies into a new data set, comprising 168 unique single molecules, 731. The mixture-pair distances were mapped onto a continuous perceptual scale from 0 (indistinguishable) to 1 (most distinct).
Next, over a three-month period, 26 teams competed to predict how similar the paired scents would be on a hidden test set of 46 mixture pairs. Following the challenge, they validated the model on an independent set of 50 olfactory mixture pairs.
It achieved a median RMSE (root mean square error) of 0.08, where 0 is a perfect score, meaning the model had good accuracy in predicting scent similarities. It also achieved a Pearson correlation of 0.57 on the test set, meaning the model had a moderate to strong positive correlation in predicting the characteristics of the data set.
The broader interest lies in whether the claimed property or reaction pathway can be characterized with enough precision to support replication by other groups. Chemistry has a replication problem that is less discussed than the one in psychology or medicine, but it is real: synthetic procedures that work reliably in one laboratory sometimes fail to transfer, for reasons ranging from impure starting materials to undocumented temperature sensitivities. A result that comes with full experimental detail and a clear characterization of the product is far more valuable than one that reports a discovery without the procedural backbone.
While many in the sensory science field believed that using science to predict similarity among mixtures would be much more difficult than predicting similarity among single. Investigators are now writing up results from a third DREAM Olfaction challenge to submit for journal publication.
Because this item comes through Phys. org Chemistry as science journalism, it should be treated as contextual reporting rather than primary evidence. Good science reporting can identify why a result matters, connect it to the wider literature and make technical work readable, but the decisive evidence remains in the original paper, dataset, mission release or technical record. That distinction is especially important when a story is later repeated by aggregators, because repetition increases visibility, not evidential strength.
The next step is to see whether independent groups working with orthogonal techniques reach compatible conclusions, and whether the result scales beyond the conditions used in the original study. Chemical discoveries that matter tend to be ones whose key properties can be measured by multiple spectroscopic, crystallographic or computational methods that are unlikely to share the same blind spots. Scalability, cost and long-term stability under realistic operating conditions are additional filters that come into play before any practical application becomes viable.

Original source: Phys. org Chemistry