Featured Post

best posts

Click the links below  >complexity and AI: why LLMs succeeded where generative linguistics failed >the sociology of false beliefs >...

Thursday, October 1, 2026

a hierarchy of data

David Deutsch in his first book identifies four essential theories of understanding the world: quantum physics (a reductionist theory), natural selection (a theory of development), computer theory (information structure) and falsificationism (a theory of theories). Simple. He chose these four presumably because of their explanatory power. 

That's looking at what these theories accomplish, how they explain. A different approach might be to look at the data that these theories range over, a kind of bottom-up perspective to see what's going on with these theories. Differences in complexity are more likely to emerge from the data bottom-up than from the theory top-down (complexity and AI, the shallowness of AI is its strength), as we'll see, so observing the distinctions between data types may improve our understanding of their complexity.

I want to be as simple as possible here, just a sketch:

Five categories of scientific data:
Natural science data set: the phenomena
Human sciences, besides the phenomena, add another data set, reflexive data: what the objects of the science say about the phenomena
The symbols (words) with which the objects (humans) say and think about their reflexive data
The noumena: emotions and qualia that are attached to the data set and the symbols
The theoretical predictive conjectures, explanations and understandings about the preceding.

1. the natural sciences -- physics, chemistry, astronomy, biology etc. -- each identify the entities of concern (an ontology) by ascribing to them their distinctive properties and in virtue of those properties, categorizing them; observing the behaviors of the entities, their interactions and the outcomes of these interactions.

2. the social sciences -- sociology, anthropology, economics, political science -- also identify the entities of concern (humans), observing their behaviors, their interactions and the outcomes of their interactions, just like the natural sciences. 

2.a Among the behaviors, however, of human entities are what the entities say about themselves. Planets and chemicals don't talk about themselves or each other, so this data set is unique to the social sciences. What makes this interesting is that what the entities say about themselves often does not match what they are actually doing and who they actually are, and what they think their goals are -- the outcomes -- may have little relation to their actions. As well, what they say about their environment and how they react to it may be disjoint from what's actually around them or how they react to them. In anthropology this is called the difference between the etic (the bare facts about them and their environment) and the emic (how they interpret themselves and their environment). We generaly accept that four-letter words are bad, and use euphemisms to avoid them. But do we really feel they are bad? We use them with good friends, and good friends are by definition good. Euphemisms are hypocritical -- we use them when we are trying to appear to be better than we actually are -- and hypocrisy is undeniably bad. Consider, why hide your honest views if they aren't bad? "Shit" has little offensiveness: "I gotta go to the supermarket to buy some shit," "Not going out tonight. I got too much shit to do,' and the ever handy "What's this shit?" or "Look at this shit!" Meanwhile the euphemism "feces" is strictly disgusting as it refers always and only to a pile of shit (http://euideism and euphemism: distinct semantic strategies for French and Latin borrowed words). Or to take a broader example, we go to college to learn, yet real learning occurs on the job or elsewhere in real life. It's actually difficult to explain why we seem to demand or expect young people go to college. Is it a jobs program for otherwise useless PhDs? In any case, this reflexive talk is an additional data set of the social sciences that the natural sciences don't have. 

The social sciences take different approaches to this data set. Sociology relies more on natural science methodology than anthropologists. So a sociologist might quantify the use of skirts among males and females in western society and concluding that they are more frequent among females than males. That's a behavioral, statistical fact. An anthropologist, dealing with symbols that are intrinsically tied to meanings and values, will intuit that the skirt is a cultural symbol in the West meaning "female, not male" a symbolism of gender in the culture that explains the statistical distribution. 

Economists also use the quantitative methods of the natural sciences, but with the inherent implication that economic behaviors are value-laden. A price is not just a number, it's a reflection on human desire and an equilibrium between opportunity costs, another set of valued desires. 

So every sociologist studies statistics; economists statistics and math; neither studies linguistics which is required of anthropologists, since language is as symbol system.

3. This brings us to the third data set, the symbols that constitute the language that the humans use to say things and express their thoughts. The science of the symbol is distinct from all other sciences. For one thing, the other sciences define the entities their theories range over by the properties of those objects. For symbols, those properties are kind of irrelevant. The word "dog" doesn't sound like a dog  or smell like a dog or look like one either. Its relation to dogs is arbitrary, not essential as the entities of other sciences are. Also, the word is not just related to an object <dog>. It's tied to an idea or meaning or thought, and those things are not phenomena at all. They are supranatural. Symbols are the core example of an emergent property that transcends the laws of physics (information faster than light). Symbols parse the phenomena of the world, pattern them and lead to their understanding of each phenomenon and its relation to the others. Each language is, as Deutsch himself points out, a theory of the world. 

4. Beyond the etic identities, behaviors, interactions and outcomes (goals), and the emic interpretations and the symbols and symbol systems through which the interpretations are rendered and communicated, there's the emotions and sensibilities of the humans, their qualia, their noumenal experience. There are a few theories about these, but this stuff is somewhat elusive. The emotions have recently been given a reductionist account: they are all degrees of arousal on a scale of value of good-desirable to bad-undesirable. On this account, one might describe anger as bad arousal directed at a person or thing that has presented an obstacle to one's goals. Something like that. All I want to say here is that aside from emotions, our personal qualia are the only data or information that we access immediately -- that is, without any mediation of symbol or interpretation. Notice that while we can name these qualia, we can't describe them as we might the emotions. What does the color green look like? Green. What does spicy taste like? Peppery. Okay, so what does peppery taste like? Spicy. There's just no way around these. These feelings have no formal nor material cause. They have an effective cause in the brain, and if we believe the explanations of evolutionary psychology, we can assess their purpose towards survival or reproduction. But what is their form? No doubt this lack of form and its logical privacy is what prevents us from describing these qualia. 

All the other data humans gather are mediated as information, just as LLMs learn from words through the mediation of digital weights, directions and distances between other words and texts. The difference between us and AIs is not that we have access to reality -- our access to reality is just information from the senses. LLMs are also just information, but from read-only texts. We play around with the phenomena -- it's not read-only. That's the big diff. Except for qualia, which we get immediately. And can't describe : )

5. all the above data are theory-driven, so theory is a meta-data category, a rich source of understanding, there being so many different kinds of theories (theory theory).   

No comments:

Post a Comment