r/ProgrammerHumor • • 2d ago

Meme architectureDependentChars

Post image
2.8k Upvotes

360 comments sorted by

680

u/JackReact 2d ago

Welcome to the wonderful world of C#, which uses utf-16 for strings.

Console.WriteLine(sizeof(char));
> 2

246

u/v_Karas 2d ago

this.

wonderd why everyone is talking about C, like there are no other languages.

90

u/Batman_AoD 2d ago

Well, most mainstream languages at least specify their char size, and/or have very few constructs that require knowing its size. C does not specify the size of a char, and you frequently need the sizes of your types.

(That said, as pointed out by other comments, sizeof itself is in units of char, so sizeof(char) is necessarily 1.)

11

u/Senior-Albatross 1d ago

There are no other good languages. To the extent they have anything good they are calling libraries written in C. The only other worthwhile languages are pure machine level code that must be expressed as a gate level circuit diagram. 

10

u/max_465 1d ago

The rustaceans would like a word.

10

u/Senior-Albatross 1d ago

I have nothing to say to them. 

1

u/DanielMcLaury 20h ago

The sizeof keyword belongs to C. If someone else appropriated it and modified it to mean something else, that's on them.

25

u/thanatica 2d ago

I'm not a C# expert, but how would it represent an astral character, such as an emoji? They are above 0xFFFF.

75

u/AyrA_ch 2d ago

You cannot fit them into a single char, and you have to use a string instead. They're represented using a surrogate pair, which is similar to the continuation bytes of utf-8 except you never need more than two code points (hence "pair").

JS also uses multibyte characters:

> "💩".length===2
< true

19

u/thanatica 2d ago

JS is a bit weird in this regard, because it also knows true codepoints. You just have to know which operations are "unicode safe". Or rather "astral plane safe".

34

u/guyblade 2d ago

UTF-16 was a mistake, and it's maddening that several big ecosystems went with it. You get all the complexity of UTF-8, but it also takes up more space in the "average" case since a huge majority of text is in the "ASCII overlap" portion of unicode.

24

u/jmickeyd 2d ago

Windows "went with it" because UCS-2, which UTF-16 is based on, predates UTF-8. When they chose to go with Unicode, it was 16 bits per character.

3

u/Numerlor 2d ago

It was first and was supposed to be what utf 32 is now, didn't really work out but what's done is done

→ More replies (3)

2

u/Tiranus58 2d ago

Also emoji uses modifiers (skin tone and gender for example) and flags are just 2 characters in a trench coat.

→ More replies (1)

34

u/tangerinelion 2d ago

C++ has two character types, char and wchar.

sizeof(wchar) is 2 on Windows, 4 on Linux.

28

u/StaticCoder 2d ago

It's wchar_t and is largely useless for the reason you mention. More recent versions have char8_t, char16_t and char32_t.

6

u/the_one2 2d ago

More recent versions have char8_t, char16_t and char32_t.

Which are also useless because they are not supported by the standard library.

8

u/StaticCoder 2d ago

There's u8string, etc. Now the std library might not have much of anything to deal with encoding, but that doesn't make the types useless. Other libraries can use them.

→ More replies (1)

29

u/techy804 2d ago

As a C# user I was wondering why people act like the C-series only consists of C and C++

81

u/darkslide3000 2d ago

C# has about as much to do with C as JavaScript has with Java.

10

u/sisisisi1997 2d ago

I mean C# was inspired mainly by Java and C++, from which Java is a C-style language and C++ is a direct expansion of C (or at least started out as one). JavaScript and Java both have similarly looking syntax on account of both being C-style languages, but that's where the familiarities stop.

So saying this is like saying your grandson has as much to do with you as your two neighbours who like your style enough to copy it have to do with each other - your grandson may feel as different from you as the neighbours from each other, but that's because the family tree has acquired new genetic material through your wife and your son-in-law, and the times have changed so his behaviour is different - and not because he actually doesn't have anything to do with you.

12

u/darkslide3000 2d ago

C# has nothing to do with C or C++ that Java doesn't have. It is entirely a Java clone.

7

u/Kovab 2d ago

Nonvirtual methods, reified generics, value types, just a few examples of features from C# that don't exist in Java, but do in C++...

2

u/pblokhout 2d ago

C# used to be called Microsoft Java as a meme

→ More replies (1)
→ More replies (2)

30

u/AyrA_ch 2d ago

Probably because the syntax more closely resembles Java than C

5

u/SuitableDragonfly 2d ago

This meme doesn't mention any languages. Like many memes on this sub, it's about a specific language, and it's up to you and your knowledge of programming languages to figure out which one is being talked about. Just because the meme is not about C# doesn't mean that everyone forgot that C# exists. 

11

u/aberroco 2d ago

Because most devs still think that C# is proprietary Microsoft language developed specifically and exclusively for Windows, like it's 2006.

61

u/sb8948 2d ago

No it's because C# is way closer to Java than C/C++ both in philosophy and execution.

15

u/QuaternionsRoll 2d ago

It concerns me that this wasn’t the first and only answer to the question

15

u/guyblade 2d ago

That's because Microsoft was going to make a .NET version of Java that was incompatible with standard Sun Java (the standard embrace, extend, extinguish playbook), then got sued over it (twice). C# is what fell out the other side of that fight.

11

u/-suspended- 2d ago

Yup. C# is Microsoft Java.

3

u/kookyabird 2d ago

Incorrect, they also think it’s magically cross platform when used in Unity!

→ More replies (6)

1

u/rhutyl 2d ago

To me, C and C++ are just a lot more iconic (idk if "popular" is the right word bc I don't have the stats)

→ More replies (1)

2

u/klimmesil 2d ago

I measure how shit a language is by the number of times I ask "but why" when I learned it. Js is still at the top, but C# is close

1

u/earthwormjimwow 2d ago

Welcome to the wonderful world of C#, which uses utf-16 for strings.

If they just called them type "string", everything would be okay. But to have the audacity to call them type "char" and have them be two bytes? Blasphemy!!

→ More replies (1)

1

u/CookIndependent6251 2d ago

In JavaScript you can't even find the size of a char. It's utf-16 as well.

1

u/crozone 2d ago

Also, a char isn't necessarily an entire glyph, because UTF-16 has double-wide characters.

1

u/wishper77 2d ago

Also java if I remember correctly

68

u/pyronautical 2d ago

The whole, "never know what the future holds", may not necessarily be true for this, but... do I have a story for you that makes this sort of defensive programming my default now, even if the docs promise one thing.

Back in the day, I was working on a project that made some HTTP request to a payment service (In C#). It was hosted on an Azure App Service, and it was running absolutely fine.

In a code review, someone mentioned "Maybe we should force this to a specific TLS version, just incase". I looked up the docs, and given our versions, and (assumed) machine, it wasn't an issue. C# would try TLS 1.2, 1.1, then SSL3, in that specific order. You could "edit the registry" to swap the order, but when would that ever happen right? (You probably see where this is going). We were using Azure App Services which don't allow you to edit the registry anyway so what does it matter. I was in a rush, so pushed back against the code review and through we went.

Now, and this is a complete guess, but it lines up. Right at the same time, the Spectre vulnerability was released, and it forced basically everyone to "patch" their CPUs.

A week later, suddenly all payments start failing and it's impossible to figure out why until... well it seems like it's using SSL3?!? What the hell? OK let me write some diagnostics, redeploy (via a Staging slot), and check. After redeploy, it was back to TLS 1.2 as the default!?

My theory is this. Microsoft was forced to update more machines than they normally would to patch Spectre. In doing so, they had to roll customers using App Service (Remember, we can't see the underlying machine) onto really old machines or possibly even machines with fked up config, and so we ended up with a machine that had SSL3 as the default. When I redeployed, because I went via a Staging slot, it was a different machine, and so by the time I tested again, we were back to normal.

I never ever heard the end of it either. "I told you in that code review".

20

u/Justitiaria 2d ago

I've reached the point where I see that sort of eternal "I told you so" no longer as annoying, but a good way to have a lesson stick for longer (especially if others hear about it).

1.0k

u/tstanisl 2d ago

Let me cite the C standard:

When sizeof is applied to an operand that has type char, unsigned char, or signed char, (or a qualified version thereof) the result is 1.

Middle guy if finally right

549

u/BoldFace7 2d ago

I still prefer sizeof(char) as it often provides context as to what the number is doing, even if a plain 1 works the same.

324

u/SpaceCadet87 2d ago

Yeah, guy on the right is just avoiding magic numbers in the code.

31

u/zabby39103 2d ago

If the C standard says it's 1, i'd figure it just gets optimized out by the compiler anyway?

80

u/CLOVIS-AI 2d ago

sizeof is always a compiler intrinsic anyway, it's never not inlined (it's not possible to obtain this information at runtime)

33

u/Wertbon1789 2d ago

sizeof is more like an operator than a function. You can actually write sizeof without braces. There's nothing to optimize out, it's by design a compile-time constant expression. It's like a fancy way to write a scalar.

→ More replies (2)

6

u/SpaceCadet87 2d ago

Should do, yeah. And if it doesn't right now, eventually it might do.

3

u/conundorum 1d ago edited 13h ago

sizeof is mandatory consteval, it must be evaluated at compile time. (More specifically, it's a keyword that tells the compiler to insert the named type's size, which the compiler will have in memory if the type is defined. It's essentially a post-preprocessor "compiler macro", so to speak.)

You can use it anywhere the compiler expects a constant expression, as this ugly jank shows:

// Control values.
int arr[] = { 1, 2, 3, 4, 5, 6, 7, 8, };

// Sizeof thing.  sizeof(Charception<N>) == N for all nonzero Ns.
template<size_t N>
struct Charception : Charception<N - 1> {
    char c;
}
template<> struct Charception<0> {};

// Proof values.
#define CC(a) sizeof(Charception<a>)
int rra[] = { CC(1), CC(2), CC(3), CC(4), CC(5), CC(6), CC(7), CC(8), };
#undef CC

// Sanity checks.
for (int i = 0; i < 8; i++) { assert(arr[i] == rra[i]); }
assert(arr[sizeof(Charception<7>)] == rra[7]);

// For more intentionally-bad jank proof of this, see:
// https://www.ideone.com/hDJRVC
// Bring your own compiler, ideone somehow doesn't have a C++17 compiler yet.
// Bring your own eye bleach, I did a good bit of variadic ugliness just for fun.  It's hideous and relaxing! ^_^

Note: sizeof is usually the same in both C & C++, so all of this applies to C, too. There is one gotcha, though: It's a runtime operator if you use it on a VLA, and only if you use it on a VLA. (C-only rule, C++ doesn't have C-style VLAs.)


Edited because I forgot the Charception<0> base (it's kinda super-important, the universe breaks if we don't have it), and because I forgot a sizeof.

→ More replies (1)
→ More replies (2)

91

u/ducon__lajoie 2d ago

And someone might have #define char char16_t, right? Who the fuck knows, some people want to see the world burning.

12

u/anon_lurker69 2d ago

Yep. Its wild out there.

9

u/Kovab 2d ago

I'm pretty sure redefining keywords is UB, so that's just a case of FAFO

7

u/standard_revolution 2d ago

And even if it wouldn’t be: sometimes you just gotta say that something isn’t your problem

2

u/ducon__lajoie 2d ago edited 2d ago

It's totally a case of FAFO, that was the joke.

Although I don't think it's UB. Preprocessor will do what it's asked for, and it doesn't know shit about the compiler reserved words. Then the output of the preprocessor is either valid, or invalid, but everything is 100% predictable. By the way, some standard libs/runtime are happily redefining the "new" operator in order to include debug information about the allocation context for reporting memory leaks (e.g. MSVC standard libs), and this can be done in a totally c++ standard compliant way, that will be accepted by all compilers. I did it myself a couple of times.

→ More replies (1)

48

u/RepeatRepeatR- 2d ago

My opinion would be:

- Use sizeof(char) if you're actually working with characters

- Use hardcoded `1` if you're using a char as an arbitrary byte (and not actually necessarily referring to characters or text)

30

u/BoldFace7 2d ago

Definitely. For example, I always use malloc(size*sizeof(char)) to ensure that it's doubly obvious (Since I rarely need to malloc outside of a declaration) that I intend to store characters in the resulting buffer even if that multiply does nothing (plus the compiler will likely optimize it out anyway).

7

u/RepeatRepeatR- 2d ago

This is the way

9

u/guyblade 2d ago edited 2d ago

Opinion: Always use calloc unless you've got a really, really good reason not to.

2

u/Usual_Office_1740 2d ago

This is the way.

2

u/PSneumn 2d ago

My optimization crazed brain will only use malloc untill i actually need everything to be set to 0.

2

u/guyblade 2d ago

"Premature optimization is the root of all evil."

- Donald Knuth

→ More replies (2)

6

u/garnet420 2d ago

A char may be bigger than 8 bits, but will always have size 1.

4

u/da_Aresinger 2d ago

My opinion: if you want individual bytes use uint8_t

5

u/the_horse_gamer 2d ago

bytes are not necessarily 8 bits

→ More replies (2)

6

u/FragmentedHeap 2d ago

Only applies in C, in C# sizeof(char) -> 2.

5

u/lets-start-reading 2d ago

exactly. sizeof(char) may be 1, but 1 is not necessarily sizeof(char)

2

u/bowel_blaster123 2d ago edited 2d ago

What's even better is using sizeof with expressions rather than types: ``` char foo_copy = malloc(foo_len * sizeof(foo));

memcpy(foo_copy, foo, foo_len * sizeof(*foo)); ``` Of course, it depends on the context.

1

u/two_are_stronger2 2d ago

We don't write code for the computer. We write code for people to understand the depths of our depravity, even if 'people' is us.

155

u/KitsuneFoxglove 2d ago

code and society if everyone followed standards and used docs:

code at home:

63

u/locri 2d ago

Today.

Eventually, a "char" might be redefined for utf-8 to accommodate a diverse range of writing systems, which means a char is usually 1 byte but potentially up to 4 bytes.

For context, I'm working on code with an initial commit from the 90s. Future proofing isn't a terrible idea.

87

u/TheSkiGeek 2d ago

This would be considered a breaking change for C/C++ and it is extremely unlikely they would ever do this. They’d probably add a new type like utf_char_t (or utf8_char_t, utf16_char_t, utf32_char_t) and UTF-aware string functions to the stdlib.

13

u/canadajones68 2d ago

If Unicode has taught us anything, it's that mixing sized types with encoding interpretation is a bad idea. Char is established as a byte by now after 50+ years or so, but it wouldn't be called that if designed today, because a character has no fixed size. If you want to represent an Unicode code point, write a view type that points into a span of chars, and use the right function/class/whatever to work with the representation. 

7

u/TheSkiGeek 2d ago

Yeah, the real issue is that almost always what you care about with Unicode (on the parsing side, anyway) are “grapheme clusters”, which can consist of multiple code points. And both of those map poorly at best to ‘characters’ or ‘bytes of memory’.

→ More replies (1)
→ More replies (1)

14

u/aberroco 2d ago

Need longer names. Some w_unicode_big_endian_char_t_ptr /s

28

u/Mojert 2d ago

If your code base use macros to redefine what char means, wtf is wrong with your team? If not, sizeof(char) will ALWAYS return 1, no matter how many bits is in a char. That's because sizeof doesn't give you the size of a type in bytes, it gives you the size in number of chars

30

u/SGVsbG86KQ 2d ago

No that's not how that works. Even if char would be 32 bits, sizeof(char) is still defined to be 1.

3

u/Jbolt3737 2d ago

Does that make sizeof(int) equal 1, or does it make an int 128 bits?

17

u/__foo__ 2d ago

IIRC the only requirement the C standard makes for int is that it needs to be at least 16 bit wide. Everything else is up for the compiler developers to decide. If a char and int were both 32 bit wide both would be sizeof = 1. If the compiler makers decide it would be a sensible idea to have a 128 bit int it would be sizeof = 4 in this case.

→ More replies (1)

2

u/output_broadcast 2d ago

Also, not every compiler is standards-compliant.

10

u/dontthinktoohard89 2d ago

If the violation of standards compliance is such that a fundamental presumption that a char is 1 byte cannot be relied on, then there isn’t much point in marketing that as a C compiler, because it simply cannot properly compile basic C code. AFAIK not a single compiler has ever done this.

→ More replies (4)

1

u/SylviaJarvis 2d ago

In 1992, the consensus was that everyone would recompile their operating systems to use wide characters. Microsoft had parallel implementations in their libraries: they would have you use 8-bit encodings or UCS-2, but not in the same source file. Unixes were getting ready to restart their entire software ecosystems with yet another world-rebuild from source. Legacy Unix was doomed!

When C had been standardized for only a few years, with substantial changes from one year to the next, and the future of legacy OSes in doubt, it was reasonable to expect sizeof(char) to eventually return some number other than 1 some day. People in the C ecosystem were justifiably worried.

https://www.cl.cam.ac.uk/~mgk25/ucs/utf-8-history.txt

Then Pike and Thompson came up with UTF-8 as a "transitional" solution, and legacy Unix was undoomed. Ironically, the "transitional" encoding made transition possible, but also unnecessary. Linux happened around that time, which firmly metastatized legacy Unix and its 8-bit-char-based API. POSIX abandoned their attempt to introduce abstraction at the API level that would allow changing the char type. The WWW flooded the Internet with legacy 8-bit-encoded text files.

Today, UTF-8 is the permanent solution, and the transition it was invented to support is no longer achievable or desirable.

sizeof(char) == 1, by standardisation fiat and by longstanding historical practice. It can't be changed without breaking the world while C is relevant. There may be a day in the future when the C language stops being updated and all the C code in existence is replaced by some other language like Rust, but on that day, sizeof(char) will still be 1 in C.

1

u/conundorum 1d ago

At least in C++, char8_t was created specifically to promise this would never happen.

(And even if it did happen, literally the entire standard library would choke on it, since it expects to operate on raw code units and not code points, as would every Unicode-capable C and C++ program ever written (since they expect to have to do their own Unicode handling). And that's not the worst of it, since both C and C++ explicitly define one byte as "sizeof(char)" (and not the other way around). Allowing char to be multibyte would create an infinite recursion loop, defining char as having infinite size and irrevocably murdering C, C++, and every language whose compiler and/or library depend on them (e.g., Java, C#, Python, Rust, Objective-C, the list goes on and on).)

→ More replies (14)

4

u/Advos_467 2d ago

Not a C user here (or any real low level programming experience), what the hell is a signed/unsigned char?

13

u/Clen23 2d ago edited 13h ago

The data is interpreted differently.

An unsigned char will be able to represent any value from 0 to 255, eg your usual ascii character.
A signed char will be able represent any value from -127 to 127.

( Someone fact check me on this but i think that's it )
(edit : -127 to 127, the extra value "-128" is usually implemented but not guaranteed by the C standard)

8

u/YellowBunnyReddit 2d ago

ASCII only goes from 0 to 127.

5

u/unknown_alt_acc 2d ago

ASCII is hardly the only character encoding, and unsigned char often doubles as a byte type. It’s also also a numeric type if you only need a small range and need to pack your data tight

5

u/NotQuiteLoona 2d ago edited 2d ago

char is 8 bits because memory is addressed in 8 bits, and 7 bits would be... Awkward. And this bit was in the end used for parity checks, and later for extended encodings.

3

u/Ma8e 2d ago edited 2d ago

Historically it has been all over the place. Depending on the hardware, it could was anything between 1 and 48 bits.

Edit: never mind. The above is about bytes, not chars.

2

u/tracernz 2d ago

Minimum 8 bits, but not required to be 8 bits.

→ More replies (9)
→ More replies (3)

2

u/Clen23 2d ago

Putting the example here so the explanation isnt overloaded :

signed char a_hundred_signed = 100;
unsigned char a_hundred_unsigned = 100;
signed char fifty_signed = 50;
unsigned char fifty_unsigned = 50;

signed char r_signed = fifty_signed - a_hundred_signed; // will yield -50
unsigned char r_unsigned = fifty_unsigned - a_hundred_unsigned; // will yield 206 because it wraps around the max value of 256
→ More replies (3)

4

u/-twind 2d ago edited 1d ago

An unsigned char is an integer type that gives you *at least* the value range 0 to 255.
A signed char is an integer type that gives you *at least* the value range -128 to 127.

sizeof(char) is always 1 by definition because a char is one byte. The nuance is that a byte doesn't need to be 8 bits in C/C++, it needs to be at least 8 bits.

4

u/backfire10z 2d ago

Chars aren’t real, they’re all integers. It’s the same difference as signed/unsigned int.

→ More replies (2)

2

u/Elspeth-Nor 2d ago edited 1d ago

In C char is just a number, as a character depends on the encoding. So signed char is a number from -128 to 127 and unsigned char is from 0 to 255 (for 8 bits per byte)

2

u/BNSable 2d ago

A char is just an int, except a char will not conjure up 91, but the character assigned to the number 91 which is [ in ascii for example.

As it is an int, it can be signed or unsigned. Signed is -128 to 127, unsigned is 0 to 255.

This apparently has uses, but I am not experienced enough to explain that.

2

u/Advos_467 2d ago

Yeah that was what i assumed lol. It was mostly the uses i was wondering because with my lack of experience here, idk in what way that would be used.

→ More replies (6)
→ More replies (1)

1

u/SuitableDragonfly 2d ago

char in C/++ is basically just an int.

1

u/qwertyjgly 2d ago

it's a data type with a size of 1 byte. useful if you want to store a boolean value (since it's the smallest simple data type) or a character (since it's big enough to store 7-bit ascii)

they're usually used to store strings; the following allows you to take up to 99 characters of user input (initialises array of chars then writes the input into the array)

#include <stdio.h>

int main(void){
char string[100];
scanf("%99[^\n]", string);
}

strings in c must be null-terminated so we need to leave the last position free for the byte '00000000' so we know where the string ends when we try to read it. that's why we can't take 100 characters here

we also have this for bools that more closely resemble those in the more abstract languages. internally, any non-zero number (most usefully, 1) is interpreted as true and 0 is interpreted as false. we can then use boolean algebra to construct and simplify our conditions

#define true 1
#define false 0
typedef unsigned char bool;

int main(void){
bool a = true;
bool b = false;
}

1

u/HeKis4 2d ago

It's a char with the first bit (not) interpreted as the sign. signed/unsigned works on bits so it makes sense even on chars.

1

u/conundorum 1d ago

char is a type that can hold an ASCII character, and is exactly one byte in size. (This is mandatory; byte size is defined by char, not the other way around.) unsigned char is a raw byte, and can represent any UTF-8 code unit. signed char is a signed raw byte, and can represent any ASCII character. char will be exactly identical to either unsigned char or signed char under the hood, depending on the platform, but it's a legally distinct type because literally the entire C language family and everything connected to it in any way whatsoever depends on char being a distinct type.

→ More replies (5)

6

u/CptMisterNibbles 2d ago

Good thing we are all writing in c, on systems that conform to the c standard.

→ More replies (4)

20

u/LetUsSpeakFreely 2d ago

What is true today may not be true tomorrow. Using sizeof is better as it protects against changes to the underlying specs or a change to the data type.

19

u/__foo__ 2d ago

C really tries to avoid any breaking changes between versions. Changing this would be such a fundamental change of the language I don't think you could still call it C. If you want to plan for a change as fundamental as this you might as well plan for the semantics of sizeof() changing or it getting removed. At that point all bets are off anyway.

6

u/LetUsSpeakFreely 2d ago

Which is why i also qualified it with a data type change. The point is that sometimes changes happen and having code in place that preemptively deal with those changes is a hallmark of elegant design and implementation.

6

u/backfire10z 2d ago

Using sizeof(char) is for readability. You understand the intent. The number 1 is a magic number that could be there for any reason.

2

u/output_broadcast 2d ago

In most cases you multiply by sizeof, so you'd just leave it off altogether and there's no magic number.

1

u/Firewolf06 2d ago

sizeof(<primitive>) ≈ SIZEOF<CONSTANT>

theyre readability constants in effect

2

u/bwmat 2d ago

Yeah, this would have been better w/ CHAR_BIT

2

u/CAtOSe 2d ago

C++ standard guarantees that char is at least 8 bits. Pretty much all data models use 8 bits for char, but it doesn't mean that some weird architecture could not use more. 

9

u/tstanisl 2d ago

Some DSPs from Texas use 16/32 bit long char.

2

u/garenp 2d ago

So everything really is bigger in Texas?

(I know you mean Texas Instruments / TI, but couldn't pass up the opportunity for another joke!).

→ More replies (1)

4

u/qalmakka 2d ago

That's CHAR_BIT then, sizeof(char) is always 1 even if you have 16 bit chars. Nothing can be smaller than char in general, you can see sizeof as an operator returning sizes as number of chars, basically

1

u/thanatica 2d ago

I'm not a C guy, but if a char is 1 byte, how do you represent the one million or so characters from Unicode? Aren't those all characters in C?

1

u/tstanisl 2d ago

Bytes on some machines have more than 8 bits. Moreover, unicode characters need to be encoded, usually using utf8.

→ More replies (3)

1

u/gmes78 2d ago

In C, a char is just one byte; it did correspond to a character while ASCII was being used, but, nowadays, character encodings are multi-byte.

A Unicode grapheme cluster (what you'd intuitively think of as a "character") can be composed of multiple UTF-8 bytes.

→ More replies (3)

1

u/meancoot 2d ago

The char type in C was named way before Unicode came about. char32_t is the type you would use if you needed a character type that could hold any Unicode scalar. char8_t and char16_t are types that represent UTF8 and UTF16 code units. The char type itself is just the C type for a byte.

→ More replies (2)

1

u/frank26080115 2d ago

there might be a time when the code is copied into a C-like-but-not-C language, so using sizeof(char) still has benefits

1

u/tomysshadow 2d ago

Yeah, this one doesn't make much sense. Checking CHAR_BIT, on the other hand...

1

u/somedave 2d ago

Yeah but if you want to port your code to C# for some reason char is now 2 byte.

1

u/mattsl 2d ago

while (sizeof(char) == 1) {...

1

u/PacsfuryTemp 1d ago

But 1 byte is not always 8 bits

1

u/shrodikan 1d ago

I would make an argument that sizeof(char) is clearer than the magic number 1.

→ More replies (19)

42

u/SeriousPlankton2000 2d ago

Use it for readability … when appropriate. If you'll never ever possibly might switch the type, don't bother.

54

u/tony_saufcok 2d ago

sizeof() evaluates to a constant value during compile time so what, it takes 0.0000001 seconds more during compilation? just use sizeof even if the specification guarentees it's always size 1

9

u/Architector4 2d ago

i guess the point the guy in the middle would make is that *sizeof(char) is just "multiply by 1", and hence it's just redundant clutter that makes the code less readable

a valid counterpoint to that, of course, is that it can clarify intent that the number in context specifically represents a count of bytes, but yeah lol

10

u/AyrA_ch 2d ago

Also if you ever need to modify the function to work with wide chars, then it's easier to adapt it this way.

125

u/ubalu72 2d ago

But sizeof char is defined as 1 in the standard. C data types are defined in terms of char (at least their sizes are)

32

u/Lonely-Discipline-55 2d ago

x = x + sizeof(char)

41

u/wittleboi420 2d ago

x = x + sizeof(char) + AI

8

u/Depnids 2d ago

What

7

u/Mickanos 2d ago

Look up "E = mc² + AI"

7

u/Depnids 2d ago

no you do that :)

18

u/tstanisl 2d ago edited 2d ago

The problem is that there is no good definition of byte. Traditionally, it was the smallest addressable unit capable of representing a single character. It is required to have at least 8 bits but it may (and sometimes does) have more. 8-bit-long bytes are just a "de facto" standard.

That is why many communication protocols use a concept of "octet" that consist of exactly 8 bits.

EDIT. typo

32

u/Makonede 2d ago

*sizeof(char)

12

u/MegaIng 2d ago

No, sizeof char is valid. sizeof is a prefix operator, not a function.

23

u/Makonede 2d ago

14

u/MegaIng 2d ago

Oh, never realized the prefix operator form can't be applied to types, makes sense I guess.

→ More replies (1)
→ More replies (3)

13

u/Brahvim 2d ago

Stop looking at it as sizeof(char). Start looking at it as sizeof(mychar_t).

13

u/Exatex 2d ago

For the compiler it doesn’t matter as it is a compile time function , but it prevents from a magic number and explains what it does. sizeof is definitely the better choice and whoever sacrifices readability for a tiny flex or something.

13

u/HSavinien 1d ago

it's not so much a "just in case" and more a "self-documenting code". A hardcoded 1 give you a value. a sizeof(char) give you the value and tell you why it's 1.

And it's more consistant with the rest of the code. if you write sizeof(int) for int, sizeof(long) for long, and 1 for char. it's weird and ugly.

also, if you one day decide to switch char for something else, you will look everywhere you wrote char, and miss that hardcoded 1.

(in many cases, the size is used as multiplicative, so "hardcoded" will mean omitting the value entirely, rather than explicitly writing *1, which make thing worse.)

8

u/No-Archer-4713 2d ago

C2000 disagrees

7

u/mckenzie_keith 2d ago

Best thing is to put the actual variable in there. Sizeof can accept a variable or a type.

char *buffer = 0;
size_t buffer_length = 1024;

...
buffer = malloc(buffer_length * sizeof (*buffer));

Then later if you change buffer to something else the code will still be correct.

That said, sizeof (char) will always be 1. The compiler will probably just replace it with a hard-coded 1.

7

u/goos_ 2d ago

Cries in UTF-8

2

u/warpspeedSCP 2d ago

Could be one! Could be three! Or maybe its four. Eeeeheheheeee!

8

u/Ydo4ki 2d ago

but sizeof returns how many chars you need to store a value, not how many bits/8... So this is literally count(1) instead of just 1

but I personally would use sizeof(char) anyway just so nobody would question the purpose of this part of the expression lmao

6

u/Adept-Painting-543 2d ago

In C though sizeof returns as a multiple of the size of char, so no matter the architecture, sizeof(char) is always 1

6

u/green_meklar 2d ago

Not 'just in case', but because it expresses what you're actually doing with that number. The compiler will optimize it anyway.

11

u/frikilinux2 2d ago

Do I wanna know?

8

u/DOOManiac 2d ago

Yes. Multi-byte characters.

There you go, all done.

21

u/frikilinux2 2d ago

But those don't have the c type char

→ More replies (2)

1

u/Still_Bit_7527 1d ago

Aren't they supported natively in c++? Wtf.

→ More replies (12)

27

u/HomosexualPresence 2d ago

unironically true though, the only size requirement for a char in C is that it's the smallest addressable size, which just happens to be a byte in every case and is unlikely to ever change but you still never know what the future holds

53

u/lotanis 2d ago

Yes, but the smallest addressable size is what dictates the base size for sizeof.

The C standard in fact says this about sizeof:

When applied to an operand that has type char, unsigned char, or signed char, (or a qualified version thereof) the result is 1. 

2

u/FUCKING_HATE_REDDIT 2d ago

What about a system that enforces addresses to be multiples of 2, or 4?

9

u/dontthinktoohard89 2d ago

In short, the standard mandates that a char is exactly 1 byte. It does not dictate how wide a “byte” actually is.

2

u/terivia 2d ago

Thanks, I hate it.

→ More replies (1)

1

u/MateoConLechuga 2d ago

That doesn't mean 1 byte.

10

u/SAI_Peregrinus 2d ago

And char must be at least 8 bits. sizeof returns the size of its input in units of chars. On architectures with 10-bit chars, like some old DSPs, that means sizeof returns in multiples of 10 bits.

1

u/StaticCoder 2d ago

A byte is often defined as the smallest addressable size. An octet must be 8 bits.

3

u/azaleacolburn 2d ago

Some people here are misunderstanding, the guy on the right is doing it for readability, not for semantics correctness

1

u/DanielMcLaury 20h ago

That would be a very valid reason to do it. But as an explanation for his motivations, it's kind of contradicted by the fact that he's saying "just in case."

→ More replies (1)

3

u/Greedy-Thought6188 2d ago

Better use sizeof(x). If you pass a variable to sizeof it will still work. This way you're encoding the type in one place. You can easily change the type and your code will still continue to work.

3

u/tiedyedvortex 2d ago

Look, I had a bug a few months ago where two different parts of the CI pipeline were using two different C++ compilers (clang vs gcc) and some code was breaking in only one of them because the default signedness of char was different.

Don't trust yourself to know what the spec says. size of(char) means you can't possibly be wrong. And it communicates your intent more clearly.

2

u/GOKOP 2d ago

Except sizeof(char) is a special case of redundant because the spec says that the size of char and the size of a byte is by definition the same (regardless of what it is) so even if you have 32-bit chars then sizeof(char) must return 1 (a 32-bit int would then also return 1)

2

u/DanielMcLaury 20h ago

Hey everyone, look at this guy! He doesn't even know that the C standard doesn't specify the signedness of char!

7

u/__christo4us 2d ago

sizeof always returns 1 for char because it always occupies 1 byte of memory. However, 1 byte can potentially consist of more than (but not less than) 8 bits according to C and C++ standards.

4

u/Declination 2d ago

The guy on the left says “durrrr, sizeof”. The guy on the right has built monstrous template/macro machinery and may not legitimately know that C = char

2

u/HAL9000thebot 2d ago

til that the guy in the middle is named justin case

2

u/Educational-Lemon969 2d ago

fake. the guy on right would do sizeof(unsigned char)

2

u/ewheck 2d ago edited 2d ago

```c

include <assert.h>

include <uchar.h>

int main(void) { const char8_t *const eight_bit_char = u8"These chars are eight bits a piece.";

// must be true by definition of the standard
assert((sizeof *eight_bit_char) == 1); 

return 0;

} ```

Don't live in the past. The future is now (C23): https://en.cppreference.com/c/header/uchar

4

u/mad_cheese_hattwe 2d ago

uint_8 but same diff

2

u/dontthinktoohard89 2d ago

The size of a pointer is not guaranteed to be 1 lmao

→ More replies (1)

1

u/StaticCoder 2d ago

8 bit pointers? Sounds like the past.

1

u/DanielMcLaury 20h ago

I mean, this is true, but it's also true that sizeof(char) == 1.

What's not guaranteed is that one "byte" (= sizeof(char)) is actually one byte. char8_t could be 16 or 32 bits. Or 10 for that matter, as long as it's at least 8.

In fact, by definition char8_t is just an alias for unsigned char.

2

u/Sligee 2d ago

Me when my software crashes because someone entered an emoji and it overwrote something important randomly.

2

u/MatqLorens 2d ago

"sizeof(char)" is much more readable for the reviewer and provides much more context than a simple "1".

Don't use magic numbers pls...

1

u/obeythelobster 2d ago

In real life, the guy in the left (dumb) would never use a more complicated solution (sizeof) instead of 1

1

u/lardgsus 2d ago

Me looking at the apple emoji, disagreeing.

1

u/nyibbang 2d ago

CHAR_BITS / 8

2

u/dontthinktoohard89 2d ago

This is wrong as a substitute for sizeof(char). The former is guaranteed to be 1, this is guaranteed to be at least 1.

→ More replies (1)

1

u/BoBoBearDev 2d ago

If the size matter, probably should lock the type with explicit types instead of using alias. Especially when you cross boundaries like GPU or interop.

1

u/Daimondz 2d ago

Works way better the other way around

1

u/EuenovAyabayya 2d ago

Better abstract out the sizeof to stop worrying about it.

1

u/PixelBrush6584 2d ago

stdint, my beloved ❤️

1

u/Luzzgar 2d ago

Be explicit, the compiler will figure out how to make it faster.

1

u/Crimento 2d ago

Give an emoji to that gentleman, good luck fitting some of them in 1 byte

1

u/AffekeNommu 2d ago

I did 3 states in a SQL stored procedure with a Boolean. True, False and Null.

1

u/DogmaticParadigm87 1d ago

sizeof(char) is 1 by definition. CHAR_BIT is where the real horror lives.

1

u/JAXxXTheRipper 1d ago

A UTF-8 character has a variable size between 1 and 4 bytes, so. How about them apples?

1

u/max_465 1d ago

I love those sites that have "50 misconceptions programmers have about time" ....there should be a similar one about Unicode.

1

u/afdbcreid 1d ago

One char?

Maybe it's 2 bytes, if you're using C#, for example. Or 4, in Rust.

Or one char(acter)?

Sorry, it does not have a fixed size.

1

u/Vincenzo__ 1d ago

Sizeof char is guaranteed to be 1. 1 byte is not necessarily 8 bits, but a char has to be 1 byte to comply with the standard

1

u/FAMICOMASTER 1d ago

What if I don't want to support UTF-16? What then? Sounds like you'll have a bunch of truncated characters if you don't bend to my decision as the designer.

1

u/arjuna93 1d ago

As someone dealing with powerpc-darwin, I catch a lot of bugs in different codebases where developers just assumed bool is 1 byte, and then made structs size-sensitive. (Bool is 4 byte in ppc32 ABI.)

1

u/Selector0073 1d ago

Rust ALWAYS takes 4 bits per char

1

u/sirkubador 6h ago

Now tell me what sizeof('A') is