Data Types Across Programming Languages: What You Actually Need to Know

Every introduction to programming eventually reaches data types.

An integer is a whole number. A string contains text. A Boolean is true or false. A floating-point number represents numbers with fractional components.

Then you learn another programming language and discover that apparently none of this was quite as simple as advertised.

An integer in Python is not the same thing as an integer in C. JavaScript’s number covers values other languages separate. Rust would rather make absence explicit than let null wander through a program. R has several different ways for something to be missing.

The definitions are not wrong. They are incomplete. A data type does more than tell a computer how to store a value. It expresses what a language assumes you should be able to represent, which distinctions matter, and which mistakes should be caught before the program runs.

So instead of memorizing another list of types, let’s give several languages the same small collection of facts and see what happens.

Meet Alex.

Alex is about to have a complicated afternoon, mostly because of one missing middle name.

The deceptively simple data model

Suppose our program needs to represent a customer:

name = "Alex"
age = 42
discount = 0.15
middle_name = no value
favorite_numbers = 3, 7, 11
status = ACTIVE

Nothing exotic. Text, an integer, a decimal value, an absent value, a collection, and a status chosen from a known set.

If data types were merely different spellings for the same concepts, moving this model between languages would be syntax conversion.

It isn’t. And watch which field causes the most trouble.

Python: let the values tell us what they are

Python makes our first version pleasantly straightforward.

name = "Alex"
age = 42
discount = 0.15
middle_name = None
favorite_numbers = [3, 7, 11]
status = "ACTIVE"

At runtime those values have types such as str, int, float, NoneType, and list. We did not declare them before assigning because Python is dynamically typed.

That makes Alex easy to get into the system. It also means constraints exist only if we deliberately add them.

Nothing about status = "ACTIVE" prevents a later status = "ACTVE". Python does not know our domain recognizes ACTIVE, SUSPENDED, and DELETED and has never heard of ACTVE. We can model that with an Enum, type hints, validation, or classes, but the bare string does not carry the rule.

Python’s integers hold another surprise for anyone arriving from C or Java. A normal Python int is not limited to a fixed 32-bit range.

age = 999999999999999999999999999999999999

An absurd customer age, and a perfectly valid Python integer.

What nothing looks like in Python

Spelling What it means Who makes you deal with it
None The only nothing available Nobody

Already, “integer” has stopped meaning one universal thing. And None is a single, undifferentiated way of saying nothing is here, which will turn out to be a choice rather than an inevitability.

JavaScript: one number, and two different nothings

const name = "Alex";
const age = 42;
const discount = 0.15;
const middleName = null;
const favoriteNumbers = [3, 7, 11];
const status = "ACTIVE";

Similar enough, until we ask JavaScript about the values.

typeof age;         // "number"
typeof discount;    // "number"
typeof middleName;  // "object"

Yes, typeof null returns "object", a famous historical artifact that cannot be fixed without breaking existing code.

More important for our comparison, 42 and 0.15 share the ordinary type number, based on IEEE 754 double-precision floating point. JavaScript also has BigInt for integers beyond the safe range of number.

Absence is where it gets interesting. An uninitialized variable is undefined. A variable deliberately set to nothing is null. The language provides two spellings for emptiness and then declines to tell you which one means what. In practice one tends to mean “nobody has assigned this yet” and the other “somebody decided this is empty,” except when a codebase uses them the other way around, or interchangeably.

Loose equality shows what the language really thinks. Zero, the empty string, false, and an empty array all agree they are the same thing. The two values that actually mean nothing is here do not join in:

0 == "";             // true
"" == false;         // true

null == undefined;   // true
null === undefined;  // false
null == 0;           // false
null == "";          // false

null and undefined are loosely equal to each other and to nothing else in the language. They form their own small island.

So JavaScript does distinguish absence from emptiness, and the rule is arguably the right one: a missing thing is not a zero thing. The problem is where that rule lives. It sits inside an equality operator famous for being unpredictable, next to a typeof that reports null as "object". Alex Dorey’s JavaScript Equality Table renders every comparison as a grid, if you want the full picture and a certain amount of despair.

What nothing looks like in JavaScript

Spelling What it means Who makes you deal with it
undefined Never assigned Nobody
null Deliberately empty, by convention Nobody

A type system is already more than a vocabulary list.

C: representation becomes difficult to ignore

#include <stdint.h>

typedef enum {
    ACTIVE,
    SUSPENDED,
    DELETED
} Status;

typedef struct {
    char *name;
    int32_t age;
    double discount;
    char *middle_name;
    int favorite_numbers[3];
    Status status;
} Customer;

Representation is visible almost everywhere. We chose an explicitly 32-bit signed integer for age, double for the discount, an array of exactly three integers, pointers to characters for names, and an enumeration for status.

A null pointer might represent an absent middle name:

customer.middle_name = NULL;

But C does not attach our intended meaning to that pointer. NULL here could mean Alex has no middle name, that we have not loaded it yet, that the allocation failed, or that somebody forgot to initialize the struct. The program establishes the convention and the program is responsible for respecting it. The compiler will not help.

What nothing looks like in C

Spelling What it means Who makes you deal with it
NULL pointer Whatever this program decided it means Convention, and only convention

C also forces us to care directly about ranges and representation. A fixed-width signed integer cannot grow forever. This is not C being gratuitously difficult. C was designed around a different relationship between programmer and machine, and its type system reflects that.

Java: primitives, references, and the null question

enum Status {
    ACTIVE,
    SUSPENDED,
    DELETED
}

class Customer {
    String name;
    int age;
    double discount;
    String middleName;
    int[] favoriteNumbers;
    Status status;
}

Java distinguishes primitive types such as int, double, and boolean from reference types such as String, Customer, and arrays.

An int cannot be null. A String can.

That single asymmetry does a lot of work. It means absence is available for some of Alex’s fields and structurally impossible for others, and the choice was made by whoever picked the field types rather than by anyone thinking about the domain. Unless we add conventions or tooling, String middleName does not tell us whether null is an expected state or a bug waiting for the right execution path.

What nothing looks like in Java

Spelling What it means Who makes you deal with it
null An empty reference Nobody
(not available) Primitives cannot be absent at all The compiler, by refusing

Java also has wrapper classes such as Integer, which means two things that sound like “an integer” behave differently:

int age;
Integer optionalAge;

One of those can be absent. The other cannot. Learning the keyword is only the beginning; the useful question is what guarantees come with it.

Rust: make some invalid states harder to represent

enum Status {
    Active,
    Suspended,
    Deleted,
}

struct Customer {
    name: String,
    age: u32,
    discount: f64,
    middle_name: Option<String>,
    favorite_numbers: Vec<i32>,
    status: Status,
}

Look at middle_name: Option<String>.

Instead of allowing a String to secretly contain null, Rust represents the possibility of absence in the type itself. The value is either Some(...) or None, and code consuming it has to deal with that possibility before it can get at the string.

This is the same information every other language has been carrying around implicitly, finally written down somewhere the compiler can read it. The other languages let you forget. Rust makes forgetting a compile error.

The Status enum similarly makes the valid alternatives explicit. We cannot accidentally assign ACTVE, because no such variant exists.

What nothing looks like in Rust

Spelling What it means Who makes you deal with it
None, inside Option<T> This value may legitimately be absent The compiler

A type system can encode valid states, not merely storage categories. That does not make incorrect programs impossible. It moves some classes of error from runtime behaviour into representations the compiler can reject.

The type system is participating in the design.

Go: useful zero values and deliberate simplicity

type Status string

const (
    Active    Status = "ACTIVE"
    Suspended Status = "SUSPENDED"
    Deleted   Status = "DELETED"
)

type Customer struct {
    Name            string
    Age             int
    Discount        float64
    MiddleName      *string
    FavoriteNumbers []int
    Status          Status
}

One important Go idea appears before we populate Alex at all: types have zero values. The zero value of an int is 0, of a bool is false, of a string is "". Pointers, slices, maps, functions, channels, and interfaces can be nil.

That creates a question. If Age contains 0, is Alex zero years old, or did nobody supply an age?

The type cannot answer. If the application uses a zero value to stand for “not supplied,” that state becomes indistinguishable from a legitimately supplied zero. If the distinction matters, the model has to represent it deliberately, which is why MiddleName above is a *string rather than a string.
What nothing looks like in Go

Spelling What it means Who makes you deal with it
0, "", false Zero/default value; cannot by itself tell us whether a domain value was supplied The model must distinguish this if it matters
nil Empty pointer, slice, map, channel, or interface Nobody

Convenient defaults acquire meaning only when they enter a domain.

R: “missing” is not one simple state

name <- "Alex"
age <- 42L
discount <- 0.15
middle_name <- NA_character_
favorite_numbers <- c(3L, 7L, 11L)
status <- "ACTIVE"

R was designed around statistical computing, and its types and missing-value semantics reflect that.

NA represents a missing value, with typed forms such as NA_integer_ and NA_character_. R also has NULL, which is not another spelling of NA, and numeric computation can produce NaN.

These answer different questions. A missing observation in a dataset is not the same thing as the absence of an object. Neither is the same thing as a mathematically undefined result. R treats those as three distinct facts, because in statistics they are three distinct facts, and conflating them produces wrong answers rather than crashes.

Consider:

mean(c(10, 20, NA))
mean(c(10, 20, NA), na.rm = TRUE)

What nothing looks like in R

Spelling What it means Who makes you deal with it
NA A missing observation Functions ask, via na.rm
NULL No object at all Nobody
NaN An undefined numeric result Nobody

The second call makes an explicit methodological choice to exclude missing observations. The language made you say so.

The field that caused all the trouble

Alex’s missing middle name has now been through seven general-purpose languages. Stack those seven tables on top of each other and the third column tells the whole story.

Python gave it None, one undifferentiated nothing. JavaScript offered two and left the meaning to you. C gave a null pointer whose meaning lives entirely in a convention the compiler cannot see. Java made absence available for reference types and impossible for primitives, as a side effect of a decision about memory. Rust put it in the type signature and refused to let you ignore it. Go made it indistinguishable from a valid value unless you reach for a pointer. R split it into three concepts because its users needed three.

Not one of those languages is wrong. They disagree about a prior question: how much does the programmer have to admit they do not know?

Representations of absence carry semantics whether or not the programmer acknowledges them. The difference between these languages is who is required to do the acknowledging, and when.

Which is why the missing middle name turned out to be more interesting than the integer.

The table gets weird quickly

Language Typing Ordinary integers 42 and 0.15 same type? How absence is spelled
Python Dynamic Arbitrary size No None
JavaScript Dynamic Fixed (BigInt for more) Yes, both number null and undefined
C Static Fixed width, chosen Usually no Null pointer, by convention
Java Static Fixed width No null, references only
Rust Static Fixed width No Option<T>, in the type
Go Static Fixed width No Zero values, nil for some types
R Dynamic Fixed width Usually distinct NA, NULL, NaN

Closed sets of alternatives vary just as much. Rust, Java, and C have real enums, Python offers Enum if you reach for it, Go builds them from a named type plus constants, JavaScript models them by hand, and R uses factors. Collections differ too: C, Java, Rust, and Go enforce an element type in the declaration, Python and JavaScript do not, and R’s atomic vectors quietly coerce everything toward a common type.

There is no universal translation saying:

Python int = JavaScript number = C int = Java int = Rust i32 = Go int = R integer

Those types overlap. They do not carry identical guarantees, ranges, operations, failure modes, or assumptions.

The same name does not imply the same contract.

If you were designing a language, what would you choose?

This gets much more interesting when you reverse the problem.

Suppose you are designing a language.

You need an integer.

Easy.

Except now you have questions. Fixed widths, or one integer that grows until memory runs out? What happens on overflow? Should 3 become 3.0 on its own? Should "42" + 1 be legal, and if so, what should it mean?

Then absence, which by now you can see is the hard one. Do you have null? Can every reference be null, or do you require an explicit optional type? Do you distinguish a missing observation from a non-existent object? Whichever you choose, every program ever written in your language inherits that decision.

Then collections, where you decide whether elements must share a type and whether values can be moved or consumed. Then domain state, where you decide whether a programmer can say a value is only ever ACTIVE, SUSPENDED, or DELETED, and whether the compiler can prove every case was handled.

Finally, timing. When should a bad type relationship surface: while parsing, during compilation, at runtime, or only when the offending path finally executes?

There is no universally correct set of answers. There are consequences.

Every choice makes some programs easier to express, some mistakes harder to make, some implementations simpler, some runtimes more complicated, and some programmers wonder why on Earth the language designer did that.

Sometimes the answer is historical compatibility. Sometimes performance. Sometimes safety. Sometimes the problem domain.

And sometimes the language designer has a very good explanation and would appreciate it if everyone stopped bringing up typeof null.

One language where absence is not the question

Everything above is a general-purpose language, and every one of them had to decide what nothing means. That is not a universal preoccupation. It is what happens when a language expects you to model a world.

The e language, commonly associated with Cadence Specman, was designed for hardware verification. Verification code describes legal ranges of values, generates constrained test data, and explores systems whose state spaces are far too large to enumerate by hand. So its declarations carry constraints rather than storage categories:

struct customer {
    age : uint (bits: 8);
    keep age in [18..120];

    status : [ACTIVE, SUSPENDED, DELETED];
    keep soft status == ACTIVE;
};

There is no absence card for e in this article, and that is the point. In ordinary application programming we ask what type of value this is. A verification environment asks what values are legal here, and how valid combinations should be generated so we can explore what the system does.

If you look only at Python, JavaScript, Java, and C, it is easy to assume programming languages are all solving approximately the same problem with different punctuation.

They aren’t.

Types are compressed language philosophy

At the beginner level, “a type tells us what kind of data a value contains” is perfectly serviceable. Eventually that runs out of road. A type can tell us what values exist, how they are stored, which operations are permitted, which conversions happen silently, what counts as absence, whether absence must be acknowledged, whether alternatives form a closed set, and in some languages how long a value lives and who owns it. “Integer, string, Boolean, float” is not the definition of a type system. It is the introductory vocabulary.

Which is why learning the types in a new language beats memorizing a conversion table.

Python’s arbitrary-precision integers are a tradeoff. So is C asking you to care about width. Java’s primitive and reference split quietly decides which of your fields are allowed to be empty. Go buys simplicity with zero values and pays for it with ambiguity at the domain boundary. R distinguishes forms of missingness because missing observations are a normal feature of statistical work, so its problem domain is visible in its types. A verification language making constrained generation first-class tells you what its designers expected programmers to spend their days doing.

A language’s types are partly a record of what its designers thought mattered, and partly a record of what they were willing to let you avoid thinking about.

They are compressed language philosophy.

Alex survived

After all of this, Alex is still 42, still has no recorded middle name, still likes 3, 7, and 11, and remains active. Nothing about the underlying person changed. But every language forced different decisions about how those facts became computation, and the fact that changed most was the one that was not there.

Knowing that Python has int, Java has int, and Rust has i32 is useful when you are trying to get a program running. Understanding why they are not interchangeable is useful when you are trying to understand languages.

Once you start asking what a language lets you represent, what it makes you acknowledge, and what it refuses to let you say, data types stop being the boring chapter before loops. They become one of the quickest ways to see what a language believes programming should be.


What is the strangest, most useful, or most frustrating type decision you’ve encountered in a programming language? More importantly, what problem do you think the language designer was trying to solve with it? Or, if you were designing your own language, which of these decisions would you make differently?

Facebooktwitterredditlinkedinmail

Leave a Reply

Your email address will not be published. Required fields are marked *

This site uses Akismet to reduce spam. Learn how your comment data is processed.